Skip to content

fix: stop a cancelled Qwen3-TTS audio decode at its next chunk on native v0.4.1-1 - #630

Merged
leehack merged 8 commits into
mainfrom
fix/86-tts-cancel-wiring-v0411-1
Sep 23, 2026
Merged

leehack merged 8 commits into
mainfrom
fix/86-tts-cancel-wiring-v0411-1

Conversation

@leehack

@leehack leehack commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

Adopts leehack/llamadart-native@v0.4.1-1 and wires its TTS cancel fix (llamadart-native#87, for llamadart-native#86), part of #322.

Commit What
f37c5ae71 The sync workflow's commit from #622, included unchanged: same SHA, bot author and committer. #622 will be closed as superseded.
3de7c7e34 The cancel wiring. Every mtmd context gets llama_dart_tts_eval_callback as cb_eval. Each synthesis gets a one-byte cancel flag owned by the main isolate. The worker reads the flag before native setup and attaches it with llama_dart_tts_set_cancel_flag after start. The flag is freed after the terminal response, after a dispose ack, or when its request was never sent. Both lookups are optional.
643f72a6d Post-bump changes. The integration test requires both exports whenever the runtime has the TTS API; the no-TTS-API branch stays. The e2e between-steps test picks its branch from llama_dart_tts_set_cancel_flag, so the pinned runtime takes the attach branch. The CHANGELOG and guide name the floor as llamadart-native v0.4.1-1 or later.
b61136de3 Turns green the two gates that #622 left red (see Maintainer decision).
dc5e95385 The integration test requires the TTS API lookup to record no startup diagnostics. Windows x64 CI runs it on the pinned DLL.
2340e6f5b Merges main at 8f6651f35 (#606, #616). See Merge with main.
e2e9fe957 The CHANGELOG entry links the TTS tracking issue #322 instead of #325.
4c14e3c38 Docs only: website/docs/changelog/recent-releases.md mirrors the two Unreleased entries above (the v0.4.1-1 pin and the TTS cancel), text unchanged.

Maintainer decision in this PR

On #622, check_webgpu_bridge_tag fails: native_release_pins.dart pins v0.4.1-1, expected the bridge-qualified native anchor v0.4.1. The v0.1.44 Web bridge manifest records native v0.4.1, and the gate required the native pin to equal it. b61136de3 adds bridgeApprovedNativePin = 'v0.4.1-1', separate from the anchor. --verify-manifest still checks the anchor against the published manifest, and passes. The two bridge docs now name both tags.

This approves v0.4.1-1 beside the v0.1.44 bridge assets without re-qualifying them. Both native release manifests record llama.cpp v0.4.1 at b29c606e28. The other option is a new bridge assets release, which is out of scope here.

b61136de3 also fixes website/docs/guides/performance-tuning.md:151, which verify_release_docs_versions flagged on #622.

Merge with main (2340e6f5b)

  • Conflicts: CHANGELOG.md keeps both sides' entries. llama_cpp_backend_test.dart keeps all 48 tests from both sides and both fake-worker dispose holds (holdNextDispose from here, holdDispose from #606), which share one _heldDispose field.
  • Against main, this PR's added and removed lines under lib/, tool/, packages/, doc/, website/, README.md, test/e2e/, test/integration/, worker_test.dart and test/unit/tooling/ are the same as before the merge.
  • #606 against the cancel wiring, line numbers at e2e9fe957:
    • Flag lifecycle: #606 does not touch synthesizeTextToSpeech, cancelTextToSpeech, dispose or _disposeWorker. llama_cpp_backend.dart still frees the flag once: on startup failure (:906), in the synthesis finally (:960), or after the dispose ack (:805, captured at :770). The one new dispose-state check, the early return in decisionHeadFree (:1036), does not touch the flag.
    • mtmd contexts: createMultimodalContext is still the only mtmd_init_from_file caller (llama_cpp_service.dart:6889), through _mtmdContextParamsFor (:6908). #606 adds no mtmd path.
    • DecisionEngine: the encoder is a llama context from llama_context_default_params() and applyDecisionContextParams (llama_cpp_service.dart:7754–7770), with no cb_eval. The head's ggml scheduler (decision_head.dart:421) gets no eval callback.
    • Worker requests: the four decision requests (worker_messages.dart:376–424, worker.dart:315–338) carry no flag address and do not cancel TTS. TextToSpeechSynthesizeRequest is still the only carrier (worker_messages.dart:360).
  • Nothing in the merge needed a code change.

Production-readiness scope

  • User-facing scope: on native llama.cpp, cancelling a Qwen3-TTS synthesis stops the audio decode at its next chunk boundary instead of after the whole native step. The default native pin moves to v0.4.1-1.
  • Supported platforms/paths: every v0.4.1-1 bundle exports both symbols (table below). Model runs were on macOS arm64 only, CPU and Metal. CI ran the integration test against the pinned bundles on Linux x64, Windows x64 and macOS arm64.
  • Unsupported or intentionally unavailable paths: Web/WebGPU and LiteRT-LM are unchanged. A native override older than v0.4.1-1 keeps the old behaviour; the lookups are optional, so it keeps working. A runtime without the TTS API gets no callback. No public API changes.
  • Out of scope / follow-ups: see Follow-ups.

Completeness checklist

  • Declared scope is fully implemented, or explicitly reduced above.
  • Unsupported platform/option combinations fail loudly with actionable diagnostics or are clearly documented as unavailable. Older runtimes keep the previous behaviour, as the CHANGELOG and guide state.
  • Public API docs, README/website docs, examples, support matrices, and changelog entries are updated where relevant.
  • Regression coverage covers the original issue plus key negative/version-skew paths where relevant (Mutation proofs).
  • Security/privacy review completed: no secrets, bearer tokens, signed URLs, or raw secret-bearing paths leak through logs, cache keys, metadata, errors, or snapshots.
  • Follow-up work that is useful but not required for this PR is tracked in GitHub Issues. See Follow-ups.

PR type guidance

  • Feature PR: N/A.
  • Bugfix PR: root cause is in llamadart-native#86: a cancel was only checked between native steps, and one step can run the whole audio decode. Regression tests: native_symbol_integration_test.dart, text_to_speech_e2e_test.dart, and the backend and worker unit tests. The platform matrix is below.
  • Docs-only PR: N/A.

Release checks (v0.4.1-1)

  • The release is a prerelease published 2026-09-23T12:29:11Z, targeting main. assets.json records:
    • tag and native_release_tag: v0.4.1-1
    • llama_cpp_tag: v0.4.1 at b29c606e28a01b1bc8c1351026a0fa6e616bf6c4
    • native_commit: 14d4f0b7e1
    • hook_contract_version: 1
    • smoke_conclusion: passed (workflow run 35850331498)
  • python3 tool/native/sync_native_release_pins.py --llama-cpp-tag v0.4.1-1 --litert-lm-tag keep --dry-run on main returned rc 0. A non-dry run on main modifies 8 files, and git diff f37c5ae71 over everything except bindings.dart is empty.
  • bindings.dart was not regenerated locally. The workflow's Linux ffigen also changed the FILE typedef to _IO_FILE.
  • tool/native/sync_native_headers_and_bindings.sh --tag v0.4.1-1 --skip-ffigen returned rc 0. The v0.4.1-1 header declares llama_dart_tts_set_cancel_flag and llama_dart_tts_eval_callback.
  • Checksums:
    • All 11 bundle archives, fetched through the build hook (testCodeBuildHook, one run per target), match SHA256SUMS.
    • SHA256SUMS matches the manifest digests for all 13 artifacts.
    • The xcframework was resolved through the companion's SwiftPM binaryTarget, whose checksum is bab353ce…5418.
  • Exports: every bundle's wrapper has 14 llama_dart_tts_* exports, including both new symbols.
Bundle Tool File Both exports
macos-arm64, macos-x86_64, ios-arm64, ios-arm64-sim, ios-x86_64-sim nm libllamadart.dylib yes
apple-xcframework: ios-arm64, simulator arm64 and x86_64, macos arm64 and x86_64 nm per arch llama.framework/llama yes
linux-arm64, linux-x64, android-arm64, android-x64 llvm-readelf --dyn-syms libllamadart.so yes
windows-arm64, windows-x64 llvm-readobj --coff-exports llamadart.dll yes

Latency and output

  • Host: macOS arm64 (M4 Max), released bundles, default hook path.
  • Before: main ec6ad6f9b on v0.4.1.
  • After: this branch on v0.4.1-1.
  • Model: Qwen3-TTS-12Hz-1.7B-Base Q4_K_M (8d18c94a) with mmproj Q8_0 (6fd65188), from the pack lock.
  • Measure: the tts pack's in-flight cancel latency. Each cancel is issued at 0.5× the reference generation time.
Cell n Median Range Runs (loadavg at start)
CPU before 12 1156.2 ms 990.1–1621.2 ms 3 runs, each fails cancel_latency_bound (9.55, 8.51, 14.32)
CPU after 12 25.7 ms 1.8–72.5 ms 3 runs, all 15 checks pass (9.47, 9.77, 9.12)
Metal before 8 12.7 ms 0.3–20.6 ms 2 runs, all pass (17.79, 14.92)
Metal after 8 16.0 ms 1.3–24.3 ms 2 runs, all pass (11.59, 15.75)
  • Immediate cancel: CPU median 0.15 ms before and 0.17 ms after.
  • Merged head e2e9fe957, CPU, 3 runs (loadavg 9.32, 9.33, 7.16 at start): in-flight cancel median 25.9 ms, 8.7–40.4 ms (n=12); cancel_latency_bound passes in all 3. Runs 2 and 3 pass all 15 checks. Run 1 fails only peak_memory_bound: its RSS after generate, the baseline, was 3.71 GB instead of the 5.05–5.70 GB of every other run, so its 5.52 GB peak, inside the other runs' 5.42–5.73 GB, measured growth 1.487 against the 1.1 budget. 11.1 of 12 GB of swap was in use. All 3 runs give the same WAV hashes as the runs below.
  • Uncancelled PCM is identical to main:
    • The fresh-load WAVs (speech-0 and speech-3..6) hash ac7ff03787ed in all 6 CPU runs and 37fd838eca9b in all 4 Metal runs.
    • The text-to-speech-smoke WAV is a97f472fb7dc2ee1 on both sides.
  • STT on CPU (Qwen3-ASR 0.6B Q8_0, bca25981) passes all 20 checks on both sides, with loadavg 15.59 on the branch and 12.85 on main.
  • Metal decode-path limitation: on Metal the pack cancels 712–800 ms into a 1525–1594 ms generation, and main already passes there. These samples do not show a cancel landing in the Metal audio decode, so that path is unmeasured.
  • Post-cancel output: with the pack's fixed seed, a synthesis's output depends on its position on the loaded engine, the same on main (#632). On CPU, uncancelled syntheses 1, 2 and 3 give ac7ff03787ed, e7a2275b83d9 and 1844169b2f3d, identically on main and this branch. speech-1, the after_cancel output, is 1844169b2f3d in all 3 main runs, in 2 of 3 branch runs and in all 3 runs at e2e9fe957, so it differs from speech-0 because of its position, not because of the cancel before it. The third branch run gives e45fb58db4c6, which none of those uncancelled syntheses gives. On Metal, uncancelled syntheses 1 to 5 give the same hashes on both sides, while speech-1 varies on both sides and matches none of them. after_cancel passes in every run.
Every pack check, verbatim
== tts-cpu-main-1 functional_pass=False backend=cpu
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound FAIL [1065.415, 1128.678, 1086.811, 1240.211] | immediate_cancel_latency_bound PASS [1.268, 0.144, 0.107, 0.151] | peak_memory_bound PASS | dispose PASS
== tts-cpu-main-2 functional_pass=False backend=cpu
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound FAIL [990.114, 1621.24, 1320.329, 1246.972] | immediate_cancel_latency_bound PASS [1.546, 0.159, 0.214, 0.134] | peak_memory_bound PASS | dispose PASS
== tts-cpu-main-3 functional_pass=False backend=cpu
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound FAIL [1019.278, 1183.643, 1071.71, 1285.159] | immediate_cancel_latency_bound PASS [1.403, 0.137, 0.146, 0.114] | peak_memory_bound PASS | dispose PASS
== tts-cpu-branch-1 functional_pass=True backend=cpu
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [72.534, 37.96, 29.208, 27.22] | immediate_cancel_latency_bound PASS [1.463, 0.125, 0.16, 0.174] | peak_memory_bound PASS | dispose PASS
== tts-cpu-branch-2 functional_pass=True backend=cpu
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [68.351, 4.142, 23.927, 5.157] | immediate_cancel_latency_bound PASS [1.452, 0.131, 0.142, 0.131] | peak_memory_bound PASS | dispose PASS
== tts-cpu-branch-3 functional_pass=True backend=cpu
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [1.828, 35.103, 11.192, 24.206] | immediate_cancel_latency_bound PASS [1.571, 0.257, 0.173, 0.148] | peak_memory_bound PASS | dispose PASS
== tts-metal-main-1 functional_pass=True backend=metal
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [19.164, 14.14, 7.916, 11.341] | immediate_cancel_latency_bound PASS [1.422, 0.126, 0.118, 0.149] | peak_memory_bound PASS | dispose PASS
== tts-metal-main-2 functional_pass=True backend=metal
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [4.883, 0.254, 18.747, 20.569] | immediate_cancel_latency_bound PASS [1.353, 0.177, 0.127, 0.115] | peak_memory_bound PASS | dispose PASS
== tts-metal-branch-1 functional_pass=True backend=metal
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [16.379, 24.282, 5.717, 1.295] | immediate_cancel_latency_bound PASS [1.575, 0.176, 0.116, 0.12] | peak_memory_bound PASS | dispose PASS
== tts-metal-branch-2 functional_pass=True backend=metal
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [15.623, 17.756, 16.601, 9.339] | immediate_cancel_latency_bound PASS [1.472, 0.116, 0.126, 0.132] | peak_memory_bound PASS | dispose PASS
== tts-cpu-merged-1 functional_pass=False backend=cpu
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [40.43, 27.209, 8.7, 30.249] | immediate_cancel_latency_bound PASS [1.33, 0.117, 0.116, 0.131] | peak_memory_bound FAIL [growth 1.487] | dispose PASS
== tts-cpu-merged-2 functional_pass=True backend=cpu
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [19.511, 21.393, 19.725, 24.622] | immediate_cancel_latency_bound PASS [1.637, 0.123, 0.11, 0.105] | peak_memory_bound PASS | dispose PASS
== tts-cpu-merged-3 functional_pass=True backend=cpu
   load PASS | generate PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [19.847, 35.993, 29.336, 34.628] | immediate_cancel_latency_bound PASS [1.626, 0.138, 0.15, 0.132] | peak_memory_bound PASS | dispose PASS
== stt-cpu-main-1 functional_pass=True backend=cpu
   load PASS | generate PASS | bytes_input PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | edge_silence PASS | edge_truncated_riff PASS | edge_stereo_44100 PASS | edge_long_boundary PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [63.556, 56.447, 52.778, 54.144] | immediate_cancel_latency_bound PASS [0.647, 0.326, 0.426, 0.368] | peak_memory_bound PASS | dispose PASS
== stt-cpu-branch-1 functional_pass=True backend=cpu
   load PASS | generate PASS | bytes_input PASS | cancel_immediate PASS | cancel PASS | after_cancel PASS | invalid_input PASS | after_invalid PASS | edge_silence PASS | edge_truncated_riff PASS | edge_stereo_44100 PASS | edge_long_boundary PASS | reload PASS | cleanup_cycle_1 PASS | cleanup_cycle_2 PASS | cleanup_cycle_3 PASS | cancel_latency_bound PASS [64.264, 76.732, 62.698, 49.891] | immediate_cancel_latency_bound PASS [0.729, 0.284, 0.295, 0.347] | peak_memory_bound PASS | dispose PASS

Verify items

  • e2e keying: the between-steps test now branches on debugTtsCancelExportsForTesting()?.setCancelFlag, the service's own resolved API. N9, which reports setCancelFlag: false, turns it red on v0.4.1-1, because the no-attach branch then sees the synthesis cancelled. N1 (no cb_eval) is caught by the integration test.
  • Windows diagnostics: _resolveTtsApi now runs at the first mtmd context creation, including vision and STT. It runs once per service and latches both success and failure.
    • On a runtime with the TTS API, which includes every v0.4.1-1 bundle, the lookup stops at the first wrapper candidate that provides the API. dc5e95385 requires it to record no startup diagnostics, and Windows x64 CI passed that assertion.
    • On a runtime without the TTS API, the lookup tries every wrapper candidate. Reaching this needs a native override older than the TTS ABI. If an existing absolute candidate fails LoadLibraryExW, the lookup records one deduplicated "Failed to preload Windows backend module …" entry.
    • That entry reaches the user only inside a later model or draft-model load failure message from the same worker.
    • Not fixed: resolving only for TTS projectors is not possible, because a TTS projector cannot be identified before mtmd_init_from_file. This case was not run on a pre-TTS Windows runtime.

Mutation proofs

Each mutation below was applied alone. Every control run was green.

Unit, fake worker. Tests from llama_cpp_backend_test.dart and worker_test.dart:

Mutation Result
U1: cancel does not raise the flag 4 red
U2: no free in the synthesis finally 2 red
U3: no free after the dispose ack 1 red
U4: free before the dispose ack 1 red
U5: cancel also frees 4 red
U6: dispose reads the flag after its startup await 1 red
U7: dispose frees the current flag, not the captured one 1 red
U8: no free on startup failure 1 red
U9: flag not zeroed at allocation 1 red
U10: free not idempotent 7 red
U11: free keeps the pointer SIGABRT (rc −6) in routes text-to-speech capability, progress, result, and cancel
U12: raise after free throws 1 red
U13: flag not stored on the backend 4 red
U14: allocator parameter ignored 4 red
U15: worker forwards 0 1 red
U16: request carries 0 SIGSEGV, si_addr=0x0
U17: free never frees 4 red

Service, real Qwen3-TTS on CPU.

Mutation Runtime Result
N1: cb_eval not set v0.4.1-1 integration red (Actual: <0>)
N2: createMultimodalContext bypasses the params helper v0.4.1-1 integration green, as expected; the tts pack fails cancel_latency_bound [1084.169, 1064.475, 1051.808, 1063.309] (loadavg 28.55)
N3: flag not attached v0.4.1-1 e2e between-steps red
N4: no flag check before native init v0.4.1 e2e before-start red
N5: eval-callback lookup reports absent v0.4.1-1 integration red, evalCallback: false
N6: set-cancel-flag lookup reports absent v0.4.1-1 integration red; e2e between-steps green, as expected
N7: set-cancel-flag lookup required v0.4.1 public-API smoke red: TTS unsupported
N7b: eval-callback lookup required v0.4.1 public-API smoke red: TTS unsupported
N8: _ttsEvalCallback without its catch pre-TTS b10173 integration red: Native text-to-speech is unavailable in this runtime bundle
N9: helper reports no setCancelFlag v0.4.1-1 e2e between-steps red
N10: helper without its catch pre-TTS b10173 integration red, same error as N8
N11: service allocates its own flag v0.4.1-1 e2e before-start and between-steps red
N12: helper returns address 0 v0.4.1-1 integration red
N13: helper reports no evalCallback v0.4.1-1 integration red

Gate.

Mutation Result
G1: pin checked against the anchor 4 red
G2: provenance expects the anchor 4 red
G3: doc regex drops the -N suffix, doc/webgpu_bridge.md 4 red
G4: the same, website doc 4 red

Re-run at e2e9fe957, each alone and reverted: U1 (4 red), U3, U4, U6, U7 and U8 (1 red each), U15 (1 red in worker_test.dart), N1 (integration red, Actual: <0>), N5 (integration red, evalCallback: false) and N3 (e2e between-steps red). Every red test is the one it hit before the merge. Controls: backend +48, worker +16, integration +18, e2e +3.

Not covered, N14: this mutation removes the moved ctxParams.use_gpu line.

  • The CPU pack still passes (loadavg 18.06), but the WAV hash changes to 1c4067e6149b and generation drops from 2.4–2.8 s in the other CPU runs to 1.68 s. That means the projector ran on the GPU.
  • The line is carried over from main, where it had no test either.

Test Plan

All at the merged head e2e9fe957 on v0.4.1-1, macOS arm64, unless noted.

4c14e3c38 changes only website/docs/changelog/recent-releases.md (11 added lines); runtime behaviour is unchanged. At 4c14e3c38: prepare_workspace rc 0 with a clean tree, build_site.sh printed Site build ready, validate_links.sh printed Link validation passed, and verify_release_docs_versions returned rc 0 with the same pending companion bump.

  • dart run tool/prepare_workspace.dart: rc 0, tree clean.
  • dart format --output=none --set-exit-if-changed .: Formatted 720 files (0 changed).
  • dart analyze --fatal-infos: No issues found!
  • dart test -p vm -j 1 --exclude-tags local-only, with LLAMADART_TEST_MODELS_DIR holding stories15M: +2963 ~78: All tests passed! (15:17:58–15:19:35, loadavg 9.61→9.94). Before the merge: +2598 ~78.
  • dart test -p chrome --exclude-tags local-only: CI Test Web (Chrome), see Review Notes.
  • Other targeted/local-only validation:
    • ./tool/docs/build_site.sh printed Site build ready, and ./tool/docs/validate_links.sh printed Link validation passed.
    • dart run tool/testing/verify_release_docs_versions.dart returned rc 0 and reported a pending bump: llamadart_llama_cpp_flutter 0.0.19 publishes v0.4.1, but Package.swift pins v0.4.1-1. With --release-prep it fails on that bump, which belongs to the release-prep PR.
    • dart run tool/testing/check_webgpu_bridge_tag.dart --verify-manifest returned rc 0.
    • classify_high_risk_changes on the diff against main printed Classification: high-risk / Surfaces: artifactConsumer, backendRuntime.
    • text_to_speech_e2e_test.dart with the Qwen3-TTS model gave +3: All tests passed!: the public API test plus both service tests on the attach path. WAV a97f472fb7dc2ee1. Main gave +1 before the merge.
    • native_symbol_integration_test.dart: +18. decision_unsupported_model_test.dart: +2.
    • decision_engine_e2e_test.dart, CPU and Metal: +4 each, 24 rows, 0 failures. Worst logit, probability and score differences are 0.1422, 0.0356 and 0.0609 on CPU and 0.1642, 0.0436 and 0.0253 on Metal (head on MTL0), the values doc/decision_engine.md records for laya-Q8_0.gguf. The encoder was a local Q8_0 conversion of the cached convaiinnovations/laya@1c5edc17 checkpoint, run with that checkpoint's model.safetensors and rl_agent_config.json. The pinned fr0stbit3/laya-gguf files are not cached, so they were not run.

Matrix Evidence

Matrix row Scope covered Platform / model / backend Result Evidence / notes
validation-speech tts pack on CPU and Metal, stt pack on CPU macOS arm64 M4 Max; Qwen3-TTS 1.7B Q4_K_M + mmproj Q8_0; Qwen3-ASR 0.6B Q8_0 PASS, one memory-bound failure Before the merge: TTS 15/15 in 3/3 CPU runs and 2/2 Metal runs; STT 20/20. At e2e9fe957, CPU: 2 of 3 runs 15/15; run 1 failed only peak_memory_bound, from a low RSS baseline (Latency and output).
text-to-speech-smoke public API WAV plus 2 service cancel tests macOS arm64, CPU PASS +3 at e2e9fe957; WAV a97f472fb7dc2ee1
static-format-analyze format, analyze macOS arm64; CI PASS Test Plan; CI Analyze & Lint
root-vm VM suite macOS arm64 locally; CI Linux x64, macOS arm64, Windows x64 PASS local +2963 ~78 at e2e9fe957
root-chrome browser suite CI ubuntu Chrome PASS CI Test Web (Chrome)
coverage-lib line coverage CI Linux PASS LCOV line coverage: 83.76% (15963/19059) at e2e9fe957
docs-site site build, links macOS arm64; CI PASS Test Plan; CI Docs Build Check
release-doc-version-consistency docs against the v0.4.1-1 pin local; CI Analyze & Lint PASS companion bump pending for release prep
webgpu-bridge-tag-consistency bridge tag, anchor, approved pin local with --verify-manifest; CI Analyze & Lint PASS see Maintainer decision
native-hook-bundles hook fetch of all 11 bundles, checksums, exports local; CI PASS see Release checks
macos-arm64-runtime-smoke llama.cpp CPU and Metal on this host macOS arm64; GLM-OCR Q4_K_M + mmproj Q8_0 PASS for speech and vision The speech packs above. Multimodal chat's mtmd contexts now get cb_eval: GLM-OCR vision output was 864c7b54f620 on main and this branch, CPU and Metal, with comparable timings. At e2e9fe957 against main 8f6651f35, n=4 each: CPU 3658–3985 vs 3679–3780 ms, Metal 262–353 vs 257–386 ms. Text-only chat has no mtmd context; the chat smoke was not run.
windows-x64-ci-runtime TTS API resolution and startup diagnostics on the pinned DLL CI windows-latest PASS Test Native (windows-latest); Native smoke (windows-latest)
linux-x64-ci-runtime the same on Linux CI ubuntu PASS Test Linux VM (with Coverage); Native Prompt Reuse Parity; Native smoke (ubuntu-latest)
windows-arm64-hook-coverage, linux-arm64-runtime-smoke, android-release-device-pool, ios-arm64-device-smoke, physical-ios-speech-e2e runtime on those targets none N/A No hardware was used. The exports were checked statically (Release checks).
web-text-to-speech-smoke Web TTS none N/A Web is unchanged; the bridge assets stay at v0.1.44.
decision-model-smoke decision_engine_e2e_test.dart macOS arm64, CPU and Metal; local Q8_0 conversion of cached convaiinnovations/laya@1c5edc17 + its head and config PASS +4 each, 0 failures; see Test Plan
high-risk-exact-head-independent-qa independent review macOS arm64; Qwen3-TTS 1.7B on v0.4.1-1 and v0.4.1, CPU; GLM-OCR and DecisionEngine on CPU and Metal PASS accepted at 4c14e3c38 / 8f6651f35 by a fresh independent reviewer, 0 blocking findings; the reviewer's model runs are at e2e9fe957, which differs only in website/docs/changelog/recent-releases.md; readiness evaluator exit 2 (unverifiedPrerequisites); see the high-risk block

Follow-ups

  • #628: LlamaEngine.dispose() and unloadModel() wait out a running TTS synthesis instead of cancelling it.
  • #627: the tts pack's half-way cancel rarely lands in the Metal audio decode, and its results do not record where it landed.
  • #629: the llama.cpp backend stops watching its worker isolate after startup; if the worker exits, generate, TTS and dispose() never complete.
  • #609 covers a synthesis or contextCreate started on NativeLlamaBackend in the same tick as dispose(), which never completes. Pre-existing on main.
  • #632: with a fixed seed, Qwen3-TTS gives different audio per synthesis on one engine, and the code predictor ignores the seed. Pre-existing on main.
  • #633: the tts pack's peak_memory_bound fails with no memory growth when the RSS sample after generate is low. Pre-existing on main.
  • Already tracked:
    • llamadart-native#88: a llama-context abort callback, so a slow backbone decode can be interrupted.
    • llamadart-native#89: STT and vision clip-encode cancel, which will share the one cb_eval slot.

Review Notes

  • Independent review status: accepted at 4c14e3c38 against 8f6651f35 by a fresh independent reviewer, 0 blocking findings; see the high-risk block.
  • CI status / head SHA: 4c14e3c38891e415cbb316c4b6babd2444b7e363. 25 checks pass and 5 are skipped, 0 fail. Documentation version and pin checks is skipped by path selection; Analyze & Lint runs both of its gates and passes. Test Native (windows-latest) passed the TTS integration test on the pinned DLL. LCOV 83.79% (15969/19059).

High-risk regression review

  • Status: accepted at 4c14e3c38891e415cbb316c4b6babd2444b7e363 / 8f6651f35f12abb1967c3b1f74f106c1e53361b7.
  • Classification: high-risk (artifactConsumer, backendRuntime)
  • Implementation task: the author of this PR
  • Independent blocking QA task: accepted, 0 blocking findings. The review is by a fresh independent reviewer from the maintainer's operator session with no part in writing this PR or the merge. One attempt with a second review tool, at high reasoning effort, also found no blocking bug; the review does not rely on it.
    • Delta from e2e9fe957: 4c14e3c38 adds 11 lines to website/docs/changelog/recent-releases.md. Every other file is byte-identical, so the runtime review of e2e9fe957 against the same base carries over (below).
    • Mirror: the two new website entries equal the first two CHANGELOG.md Unreleased entries after whitespace and bullet normalisation, and the next ten still match. git diff --check is clean. Both linked issues exist.
    • Docs gates at the head: build_site.sh printed Site build ready (loadavg 21.18→24.66) and validate_links.sh passed (loadavg 24.23). The built Unreleased list is still a tight list. verify_release_docs_versions returned rc 0 with the same pending companion bump; check_webgpu_bridge_tag returned rc 0.
    • CI at the head: 30 check runs: 25 success, 5 skipped, 0 failed. LCOV 83.79% (15969/19059).
    • Runtime review, accepted at e2e9fe957 / 8f6651f35:
      • Merge: outside CHANGELOG.md and the backend test, the PR's added and removed lines against main equal those of ec6ad6f9b..dc5e95385 (524 lines, 21 files). The backend test has 48 tests: all 41 from this branch and all 42 from main, none duplicated.
      • #606 against the cancel wiring: it changes none of synthesizeTextToSpeech, cancelTextToSpeech, dispose or _disposeWorker. _mtmdContextParamsFor is the only cb_eval site; the decision encoder, the head scheduler and the decision worker requests get no callback or flag.
      • Flag lifecycle: traced through a same-tick dispose (#609), a synthesis claimed during dispose, startup failure and the dispose timeout. No use-after-free, double free or lost cancel.
      • Runs on v0.4.1-1: prepare_workspace, format and analyze clean; VM suite +2963 ~78; TTS e2e +3, WAV a97f472fb7dc2ee1; decision e2e +4 on CPU and Metal with the documented worst diffs. CPU tts pack 15/15 (loadavg 13.44): in-flight cancel 36.8, 76.2, 38.9 and 31.6 ms, memory growth 1.049. GLM-OCR output 864c7b54f620 on e2e9fe957 and main: CPU 4361–4470 vs 4537–5125 ms (loadavg 19–23), Metal 335–346 vs 342–345 ms.
      • Version skew, this code on v0.4.1: TTS e2e +3 with the same WAV, through the non-attach branch. The tts pack fails only cancel_latency_bound [1200, 2061, 25, 1488] ms, the old behaviour, with unchanged WAV hashes. The cancel-exports integration test is red there by design.
      • Mutations: U2, U5 and U10 match the table. N3 turns the e2e between-steps test red; N4 on v0.4.1 turns the before-start test red. M2 (the merge keeps only main's dispose hold) turns dispose frees a running synthesis cancel flag after the worker acks red. M1 (keeps only this branch's hold) stays green, as does removing that hold on main. D1 (TTS callback on the decision encoder) leaves the decision e2e numbers identical.
    • Non-blocking notes:
      • CHANGELOG.md has one Unreleased entry with no website mirror, the move of native release pins. It is pre-existing on main 8f6651f35.
      • LCOV is 83.79% at 4c14e3c38 and 83.76% at e2e9fe957, with lib/ byte-identical: CI coverage varies slightly between runs.
      • Vulkan, CUDA and OpenCL vision and STT timings with cb_eval are unmeasured.
      • The review clones fetched public release assets through the build hook: the v0.4.1-1 bundle at e2e9fe957, and the macOS arm64 v0.4.1-1 bundle and the LiteRT-LM v0.17.0-6 runtime at 4c14e3c38. No result depended on them.
  • Exact head / current base: 4c14e3c38891e415cbb316c4b6babd2444b7e363 / 8f6651f35f12abb1967c3b1f74f106c1e53361b7 (GitHub baseRefOid, which is main and an ancestor of the head). Both queried again for the readiness evidence; unchanged, 0 review threads.
  • Production-branch deletion, bypass, or miswire proof: U1–U17, N1–N13 and G1–G4 above. N14 is uncovered and disclosed.
    • Re-run at 4c14e3c38 for the readiness evidence, each alone and reverted: U1 (4 red), U2 (2), U8 (1), M2 (1), U15 (1), G1 (4), P1 (2, restores the base pin v0.4.1 in native_release_pins.dart), N1 and N5 (integration red), N3 (e2e between-steps red). Controls: backend +48, worker +16, gate +58, integration +18, e2e +3 with WAV a97f472fb7dc2ee1. Same results as at e2e9fe957.
  • Readiness evaluator: dart run tool/testing/high_risk_readiness.dart, with repository, PR, author, head and base from a fresh GitHub API query: exit 2, unverifiedPrerequisites (externalPrerequisitesUnavailable), over 24 changed files. That is the best local result: the evidence is internally consistent, and the external prerequisites do not exist.
  • Affected-family real-model/artifact evidence: Qwen3-TTS 1.7B on the released v0.4.1-1 and v0.4.1 macOS arm64 bundles, CPU; GLM-OCR vision and the DecisionEngine e2e on CPU and Metal. All 11 released bundles and the xcframework were checked statically.
  • Explicit unavailable-family or other N/A evidence: no model runs on Linux, Windows, Android or iOS. On Linux x64 and Windows x64, CI checked export resolution and the eval-callback address.
  • Known PR-caused P1 regressions: 0
  • Unresolved review threads: 0

github-actions Bot and others added 7 commits September 23, 2026 16:43
For [llamadart-native#86](leehack/llamadart-native#86) and [#325](#325).

A text-to-speech cancel reached the worker only as a message, so it waited in the worker isolate's queue until the native step in progress returned. On CPU that step can be the whole Qwen3-TTS audio decode.

When the runtime exports `llama_dart_tts_eval_callback`, every mtmd context now gets it as `cb_eval`, so a cancelled task stops the audio decode at its next chunk boundary. The backend allocates a one-byte cancel flag per synthesis, raises it on cancel and still sends the cancel message. The worker reads the flag before native task setup and, when the runtime exports `llama_dart_tts_set_cancel_flag`, attaches it after startup so native code reads it during a step. The flag is freed after the worker's terminal response, after a dispose acknowledgement, or when its request was never sent; a dispose frees only the flag of a synthesis claimed before it began.

Both exports are optional lookups: v0.4.1 and older runtimes keep today's behaviour, and a runtime without the text-to-speech API gets no callback.

Qwen3-TTS 1.7B, tts pack on an M4 Max CPU: in-flight cancels took 17.8-4376.9 ms on main with v0.4.1 and 2.3-118.6 ms with this change on llamadart-native 14d4f0b.
Every llamadart-native v0.4.1-1 bundle exports `llama_dart_tts_eval_callback` and `llama_dart_tts_set_cancel_flag`. The integration test now requires both whenever the runtime has the text-to-speech API, and still accepts a runtime without that API.

`debugTtsCancelExportsForTesting` reports the text-to-speech API the service itself resolves. The e2e between-steps test picks its branch from `llama_dart_tts_set_cancel_flag`, not from the eval callback, so the pinned runtime takes the attach branch.

The CHANGELOG and the text-to-speech guide name llamadart-native v0.4.1-1 or later as the floor.
The v0.4.1-1 sync left two gates red. performance-tuning.md still called v0.4.1 the package-pinned runtime, and the Web/native gate required the native pin to equal the native release the v0.1.44 bridge assets were qualified against (v0.4.1).

v0.4.1-1 rebuilds v0.4.1 from the same upstream llama.cpp commit; both release manifests record b29c606e28. The gate now records v0.4.1-1 as the approved native pin, separate from the qualified anchor that `--verify-manifest` still checks against the published bridge manifest, and the two bridge docs name both. Any other native pin still fails the gate.
Every mtmd context creation now resolves the text-to-speech API, including for vision and speech-to-text. On Windows that lookup preloads each existing absolute wrapper candidate it tries, and a failed preload records a startup diagnostic. The integration test now requires the lookup to record none on the runtime it runs against.
Brings in #606 and
#616. CHANGELOG.md keeps both
sides' Unreleased entries. The backend test's fake worker keeps both
dispose holds: holdNextDispose with acknowledgeHeldDispose from this
branch and holdDispose with releaseDispose from main, sharing one
_heldDispose field.
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Chat app preview removed for leehack/llamadart-chat-pr-630.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant