Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -1189,3 +1189,8 @@ git diff --check
end-to-end Android application latency or a cross-hardware speed ranking.
- Distribution-shift and extrapolation robustness were not measured in the
reported experiments.

## iOS implementation update

See [iOS implementation and validation boundary](docs/ios-implementation-update.md)
for background execution, raw evidence export, integration tests and device steps.
32 changes: 32 additions & 0 deletions docs/ios-implementation-update.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# iOS implementation and validation boundary

The SwiftUI app executes the generated aircraft-design Core ML surrogate. It does
not execute the separate DASHlink real-flight model. Benchmarking now runs off the
UI thread and saves raw warm latencies and inputs alongside model provenance,
thermal state and simulator identity. The contract rejects invalid normalization
and nonfinite inputs. The Python validator rejects nonfinite summaries and checks
raw sample consistency when samples are supplied.

From the repository root:

```sh
pip install -e '.[neural,coreml,dev]'
python scripts/prepare_ios_resources.py
cd ios
xcodegen generate
xcodebuild -project EdgeGenBenchDemo.xcodeproj -scheme EdgeGenBenchDemo -destination 'platform=iOS Simulator,name=iPhone 16' test
```

Choose an installed simulator from `xcrun simctl list devices available`. The new
CoreMLIntegrationTests loads the bundled model, executes inference and attaches
100-run evidence to the Xcode result. A simulator test checks integration only.
For hardware evidence, choose a physical iPhone and signing team in Xcode, run the
app and export its JSON through Share. Use `scripts/validate_ios_evidence.py --help`
to validate against the source artifact hashes. No physical iPhone timings have
been collected in this update. The local Xcode build was blocked by the execution
sandbox; the integration test has not yet been confirmed passing.

Core ML `.all` requests available compute units; it does not prove ANE placement.
Use retained Instruments measurements for placement or power claims. Cold latency
includes model loading plus first inference. This is a model benchmark, not a
validated aircraft-design recommendation system.
8 changes: 7 additions & 1 deletion ios/EdgeGenBenchDemo/BenchmarkEvidence.swift
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,9 @@ struct IOSBenchmarkEvidence: Codable {
let latency: LatencySummary
let outputMaxAbsDrift: Double
let outputs: [PredictionValue]
var warmLatencySamplesMs: [Double]? = nil
var inputValues: [Double]? = nil
var inputCategory: String? = nil
}

struct PredictionValue: Codable {
Expand Down Expand Up @@ -100,7 +103,10 @@ enum IOSBenchmarkRunner {
warmRuns: warmRuns
),
outputMaxAbsDrift: maxDrift,
outputs: coldOutput.map { PredictionValue(name: $0.name, value: $0.value) }
outputs: coldOutput.map { PredictionValue(name: $0.name, value: $0.value) },
warmLatencySamplesMs: latencies,
inputValues: numericValues,
inputCategory: category
)
}

Expand Down
77 changes: 61 additions & 16 deletions ios/EdgeGenBenchDemo/ContentView.swift
Original file line number Diff line number Diff line change
Expand Up @@ -7,22 +7,33 @@ struct ContentView: View {
"hybridization_ratio"
]
@State private var values = [4.0, 250.0, 180.0, 300.0, 0.65, 0.5]
@State private var category = "battery_electric"
@State private var category = "conventional_turboprop"
@State private var predictions: [Prediction] = []
@State private var message = "Run the bundled Core ML model and capture cold + warm evidence."
@State private var inputWarning: String?
@State private var evidence: IOSBenchmarkEvidence?
@State private var evidenceURL: URL?
@State private var isRunning = false

var body: some View {
NavigationStack {
Form {
Section("Aircraft design") {
Section("Aircraft design — generated-data model") {
ForEach(featureNames.indices, id: \.self) { index in
TextField(featureNames[index], value: $values[index], format: .number)
TextField(featureNames[index].replacingOccurrences(of: "_", with: " "), value: $values[index], format: .number)
.keyboardType(.decimalPad)
}
TextField("propulsion_architecture", text: $category)
Picker("Propulsion architecture", selection: $category) {
ForEach(["conventional_turboprop", "fuel_cell_electric", "parallel_hybrid", "series_hybrid"], id: \.self) { value in
Text(value.replacingOccurrences(of: "_", with: " ")).tag(value)
}
}
Button("Reset to reference inputs", action: resetInputs)
if let inputWarning {
Label(inputWarning, systemImage: "exclamationmark.triangle")
.font(.footnote)
.foregroundStyle(.orange)
}
}
Section {
Button(isRunning ? "Benchmarking…" : "Run cold + warm benchmark", action: runBenchmark)
Expand Down Expand Up @@ -50,23 +61,57 @@ struct ContentView: View {
}
}
.navigationTitle("EdgeGenBench")
.disabled(isRunning)
.onChange(of: values) { _, _ in updateInputWarning() }
}
.task { updateInputWarning() }
}

private func runBenchmark() {
guard values.count == featureNames.count, values.allSatisfy(\.isFinite) else {
message = "Enter finite numeric values for every feature."
return
}
updateInputWarning()
isRunning = true
do {
let result = try IOSBenchmarkRunner.run(numericValues: values, category: category)
evidence = result
predictions = result.outputs.map { Prediction(name: $0.name, value: $0.value) }
evidenceURL = try result.writeTemporaryJSON()
message = "Core ML benchmark completed (1 cold + \(result.latency.warmRuns) warm runs)."
} catch {
predictions = []
evidence = nil
evidenceURL = nil
message = error.localizedDescription
evidence = nil
evidenceURL = nil
let inputValues = values
let inputCategory = category
Task {
do {
let result = try await Task.detached(priority: .userInitiated) {
try IOSBenchmarkRunner.run(numericValues: inputValues, category: inputCategory)
}.value
evidence = result
predictions = result.outputs.map { Prediction(name: $0.name, value: $0.value) }
evidenceURL = try result.writeTemporaryJSON()
message = "Completed 1 model-load + first-prediction and \(result.latency.warmRuns) warm runs."
} catch {
predictions = []
message = error.localizedDescription
}
isRunning = false
}
}

private func resetInputs() {
values = [65.0, 950.0, 535.0, 527.0, 0.57, 0.24]
category = "conventional_turboprop"
message = "Reference inputs restored."
updateInputWarning()
}

private func updateInputWarning() {
let means = [64.93524, 950.4023, 534.7646, 526.7014, 0.57418, 0.24301]
let scales = [14.40655, 318.0289, 66.45373, 130.1412, 0.072085, 0.215283]
guard values.count == means.count, values.allSatisfy(\.isFinite) else {
inputWarning = "Some inputs are not finite."
return
}
isRunning = false
let maximumZ = zip(values, zip(means, scales)).map { pair in
abs((pair.0 - pair.1.0) / pair.1.1)
}.max() ?? 0
inputWarning = maximumZ > 3 ? "Inputs are outside the training distribution (max |z| = \(maximumZ.formatted(.number.precision(.fractionLength(1))))." : nil
}
}
30 changes: 28 additions & 2 deletions ios/EdgeGenBenchDemo/SurrogatePredictor.swift
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,7 @@ final class SurrogatePredictor {
let contractData = try Data(contentsOf: contractURL)
contract = try JSONDecoder().decode(ModelContract.self, from: contractData)
contractSHA256 = SHA256.hash(data: contractData).map { String(format: "%02x", $0) }.joined()
try contract.validate()
guard contract.featureMean.count + contract.categories.count == contract.inputDimension,
contract.featureScale.count == contract.featureMean.count,
contract.targets.count == contract.outputDimension,
Expand All @@ -72,7 +73,8 @@ final class SurrogatePredictor {
}

func predict(numericValues: [Double], category: String) throws -> [Prediction] {
guard numericValues.count == contract.featureMean.count,
guard numericValues.allSatisfy({ $0.isFinite }),
numericValues.count == contract.featureMean.count,
let categoryIndex = contract.categories.firstIndex(of: category) else {
throw SurrogateError.invalidContract("input values do not agree")
}
Expand All @@ -89,8 +91,32 @@ final class SurrogatePredictor {
normalized.count == contract.outputDimension else {
throw SurrogateError.invalidOutput
}
return contract.targets.indices.map { index in
guard (0..<normalized.count).allSatisfy({ normalized[$0].doubleValue.isFinite }) else {
throw SurrogateError.invalidOutput
}
let predictions = contract.targets.indices.map { index in
Prediction(name: contract.targets[index], value: normalized[index].doubleValue * contract.targetScale[index] + contract.targetMean[index])
}
guard predictions.allSatisfy({ $0.value.isFinite }) else {
throw SurrogateError.invalidOutput
}
return predictions
}
}


extension ModelContract {
func validate() throws {
guard schemaVersion == "1.0" || schemaVersion == "1.1", !numericFeatures.isEmpty,
numericFeatures.count == featureMean.count,
featureMean.count == featureScale.count,
Set(numericFeatures).count == numericFeatures.count,
!categories.isEmpty, Set(categories).count == categories.count,
featureMean.allSatisfy({ $0.isFinite }),
featureScale.allSatisfy({ $0.isFinite && $0 > 0 }),
targetMean.allSatisfy({ $0.isFinite }),
targetScale.allSatisfy({ $0.isFinite && $0 > 0 }) else {
throw SurrogateError.invalidContract("nonfinite, duplicate, or invalid preprocessing values")
}
}
}
22 changes: 22 additions & 0 deletions ios/EdgeGenBenchDemoTests/CoreMLIntegrationTests.swift
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
import XCTest
@testable import EdgeGenBenchDemo

final class CoreMLIntegrationTests: XCTestCase {
func testBundledModelRunsAndRejectsInvalidInputs() throws {
let predictor = try SurrogatePredictor()
let values = predictor.contract.featureMean
let category = try XCTUnwrap(predictor.contract.categories.first)
let result = try predictor.predict(numericValues: values, category: category)
XCTAssertEqual(result.count, predictor.contract.outputDimension)
XCTAssertTrue(result.allSatisfy { $0.value.isFinite })
XCTAssertThrowsError(try predictor.predict(numericValues: [.nan], category: category))
XCTAssertThrowsError(try predictor.predict(numericValues: values, category: "unknown"))
let evidence = try IOSBenchmarkRunner.run(numericValues: values, category: category)
XCTAssertEqual(evidence.warmLatencySamplesMs?.count, 100)
XCTAssertLessThanOrEqual(evidence.outputMaxAbsDrift, 1e-6)
let attachment = XCTAttachment(data: try JSONEncoder().encode(evidence), uniformTypeIdentifier: "public.json")
attachment.name = "CoreML execution evidence"
attachment.lifetime = .keepAlways
add(attachment)
}
}
16 changes: 16 additions & 0 deletions reports/iphone/run-01-report.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# EdgeGenBench Core ML execution report

- Status: `validated_physical_iphone_coreml`
- Device: `iPhone17,1`
- OS: `iOS 26.6.2`
- Backend: `CoreML` (requested compute units: `all`)
- Cold latency: `154.524125 ms`
- Warm mean latency: `0.036253 ms`
- Warm p95 latency: `0.042750 ms`
- Warm runs: `100`
- Output max absolute drift: `0`
- Thermal state: `nominal` → `nominal`
- Power: `not measured`
- Apple Neural Engine placement: `not measured`

> Physical iPhone latency; not proof of Apple Neural Engine placement and not a power measurement.
25 changes: 25 additions & 0 deletions reports/iphone/run-01-summary.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
{
"status": "validated_physical_iphone_coreml",
"captured_at_utc": "2026-09-13T00:42:55Z",
"app_version": "0.1.0",
"device": {
"model": "iPhone17,1",
"simulator": false,
"systemName": "iOS",
"systemVersion": "26.6.2"
},
"backend": "CoreML",
"requested_compute_units": "all",
"latency": {
"coldMs": 154.524125,
"warmMeanMs": 0.03625252,
"warmP95Ms": 0.04275,
"warmRuns": 100
},
"output_max_abs_drift": 0,
"thermal_state_before": "nominal",
"thermal_state_after": "nominal",
"power_measurement": "not_measured",
"neural_engine_placement": "not_measured",
"claim_boundary": "Physical iPhone latency; not proof of Apple Neural Engine placement and not a power measurement."
}
Loading
Loading