Skip to content

Revalidate the current ONNX model on Qualcomm QNN - #49

Merged
triasha72 merged 4 commits into
mainfrom
feature/final-hybrid-acceptance
Aug 26, 2026
Merged

Revalidate the current ONNX model on Qualcomm QNN#49
triasha72 merged 4 commits into
mainfrom
feature/final-hybrid-acceptance

Conversation

@triasha72

Copy link
Copy Markdown
Owner

Summary

  • rerun the repository's current ONNX model through Qualcomm AI Hub on Snapdragon 8 Elite QRD
  • compile batch 1/32/256 graphs, link one QNN context binary, and validate held-out inference
  • update the tracked QNN report and portfolio matrix with matching model provenance
  • add a resumable authenticated rerun script and reproducibility documentation

Physical-device results

Batch Latency Throughput Peak memory Placement
1 0.038 ms 26,315.8 samples/s 122,855,424 B NPU × 9
32 0.040 ms 800,000.0 samples/s 122,880,000 B NPU × 9
256 0.047 ms 5,446,808.5 samples/s 122,896,384 B NPU × 9

Source ONNX SHA-256: 191927e05b5f82a539f0ad35c78dafe7f969cb4e57e9c76556cb9f35053e658e

Compile jobs: jpyxnmd05, jp0jk610g, jp8x813qg
Link job: jgk4d8lvp
Profile jobs: j567vw3np, jpyxnmvr5, jglx7lllg
Inference jobs: jg9m8d3q5, jpez2yl7p, jgk4d82op

Verification

  • ruff format --check .
  • ruff check .
  • 26 Qualcomm/QNN and acceptance tests passed
  • regenerated acceptance outputs match tracked files byte-for-byte

The full local pytest collection was blocked by a stale macOS scikit-learn binary in the existing virtual environment; GitHub's clean Linux CI remains the authoritative full-suite run.

Power remains unmeasured, and AI Hub model-profile latency is not presented as Android application end-to-end latency.

@triasha72 triasha72 closed this Aug 26, 2026
@triasha72 triasha72 reopened this Aug 26, 2026
@triasha72
triasha72 merged commit b63bc34 into main Aug 26, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant