fix: stop a cancelled Qwen3-TTS audio decode at its next chunk on native v0.4.1-1 - #630
Merged
Merged
Conversation
For [llamadart-native#86](leehack/llamadart-native#86) and [#325](#325). A text-to-speech cancel reached the worker only as a message, so it waited in the worker isolate's queue until the native step in progress returned. On CPU that step can be the whole Qwen3-TTS audio decode. When the runtime exports `llama_dart_tts_eval_callback`, every mtmd context now gets it as `cb_eval`, so a cancelled task stops the audio decode at its next chunk boundary. The backend allocates a one-byte cancel flag per synthesis, raises it on cancel and still sends the cancel message. The worker reads the flag before native task setup and, when the runtime exports `llama_dart_tts_set_cancel_flag`, attaches it after startup so native code reads it during a step. The flag is freed after the worker's terminal response, after a dispose acknowledgement, or when its request was never sent; a dispose frees only the flag of a synthesis claimed before it began. Both exports are optional lookups: v0.4.1 and older runtimes keep today's behaviour, and a runtime without the text-to-speech API gets no callback. Qwen3-TTS 1.7B, tts pack on an M4 Max CPU: in-flight cancels took 17.8-4376.9 ms on main with v0.4.1 and 2.3-118.6 ms with this change on llamadart-native 14d4f0b.
Every llamadart-native v0.4.1-1 bundle exports `llama_dart_tts_eval_callback` and `llama_dart_tts_set_cancel_flag`. The integration test now requires both whenever the runtime has the text-to-speech API, and still accepts a runtime without that API. `debugTtsCancelExportsForTesting` reports the text-to-speech API the service itself resolves. The e2e between-steps test picks its branch from `llama_dart_tts_set_cancel_flag`, not from the eval callback, so the pinned runtime takes the attach branch. The CHANGELOG and the text-to-speech guide name llamadart-native v0.4.1-1 or later as the floor.
The v0.4.1-1 sync left two gates red. performance-tuning.md still called v0.4.1 the package-pinned runtime, and the Web/native gate required the native pin to equal the native release the v0.1.44 bridge assets were qualified against (v0.4.1). v0.4.1-1 rebuilds v0.4.1 from the same upstream llama.cpp commit; both release manifests record b29c606e28. The gate now records v0.4.1-1 as the approved native pin, separate from the qualified anchor that `--verify-manifest` still checks against the published bridge manifest, and the two bridge docs name both. Any other native pin still fails the gate.
Every mtmd context creation now resolves the text-to-speech API, including for vision and speech-to-text. On Windows that lookup preloads each existing absolute wrapper candidate it tries, and a failed preload records a startup diagnostic. The integration test now requires the lookup to record none on the runtime it runs against.
leehack
marked this pull request as ready for review
September 23, 2026 20:48
Contributor
|
Chat app preview removed for |
This was referenced Sep 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adopts
leehack/llamadart-native@v0.4.1-1and wires its TTS cancel fix (llamadart-native#87, for llamadart-native#86), part of #322.f37c5ae713de7c7e34llama_dart_tts_eval_callbackascb_eval. Each synthesis gets a one-byte cancel flag owned by the main isolate. The worker reads the flag before native setup and attaches it withllama_dart_tts_set_cancel_flagafter start. The flag is freed after the terminal response, after a dispose ack, or when its request was never sent. Both lookups are optional.643f72a6dllama_dart_tts_set_cancel_flag, so the pinned runtime takes the attach branch. The CHANGELOG and guide name the floor as llamadart-native v0.4.1-1 or later.b61136de3dc5e953852340e6f5bmainat8f6651f35(#606, #616). See Merge with main.e2e9fe9574c14e3c38website/docs/changelog/recent-releases.mdmirrors the twoUnreleasedentries above (the v0.4.1-1 pin and the TTS cancel), text unchanged.Maintainer decision in this PR
On #622,
check_webgpu_bridge_tagfails:native_release_pins.dart pins v0.4.1-1, expected the bridge-qualified native anchor v0.4.1. The v0.1.44 Web bridge manifest records nativev0.4.1, and the gate required the native pin to equal it.b61136de3addsbridgeApprovedNativePin = 'v0.4.1-1', separate from the anchor.--verify-manifeststill checks the anchor against the published manifest, and passes. The two bridge docs now name both tags.This approves v0.4.1-1 beside the v0.1.44 bridge assets without re-qualifying them. Both native release manifests record llama.cpp
v0.4.1atb29c606e28. The other option is a new bridge assets release, which is out of scope here.b61136de3also fixeswebsite/docs/guides/performance-tuning.md:151, whichverify_release_docs_versionsflagged on #622.Merge with main (
2340e6f5b)CHANGELOG.mdkeeps both sides' entries.llama_cpp_backend_test.dartkeeps all 48 tests from both sides and both fake-worker dispose holds (holdNextDisposefrom here,holdDisposefrom #606), which share one_heldDisposefield.main, this PR's added and removed lines underlib/,tool/,packages/,doc/,website/,README.md,test/e2e/,test/integration/,worker_test.dartandtest/unit/tooling/are the same as before the merge.e2e9fe957:synthesizeTextToSpeech,cancelTextToSpeech,disposeor_disposeWorker.llama_cpp_backend.dartstill frees the flag once: on startup failure (:906), in the synthesisfinally(:960), or after the dispose ack (:805, captured at :770). The one new dispose-state check, the early return indecisionHeadFree(:1036), does not touch the flag.createMultimodalContextis still the onlymtmd_init_from_filecaller (llama_cpp_service.dart:6889), through_mtmdContextParamsFor(:6908). #606 adds no mtmd path.llama_context_default_params()andapplyDecisionContextParams(llama_cpp_service.dart:7754–7770), with nocb_eval. The head's ggml scheduler (decision_head.dart:421) gets no eval callback.worker_messages.dart:376–424,worker.dart:315–338) carry no flag address and do not cancel TTS.TextToSpeechSynthesizeRequestis still the only carrier (worker_messages.dart:360).Production-readiness scope
Completeness checklist
PR type guidance
native_symbol_integration_test.dart,text_to_speech_e2e_test.dart, and the backend and worker unit tests. The platform matrix is below.Release checks (v0.4.1-1)
main.assets.jsonrecords:tagandnative_release_tag: v0.4.1-1llama_cpp_tag: v0.4.1 atb29c606e28a01b1bc8c1351026a0fa6e616bf6c4native_commit:14d4f0b7e1hook_contract_version: 1smoke_conclusion:passed(workflow run 35850331498)python3 tool/native/sync_native_release_pins.py --llama-cpp-tag v0.4.1-1 --litert-lm-tag keep --dry-runon main returned rc 0. A non-dry run on main modifies 8 files, andgit diff f37c5ae71over everything exceptbindings.dartis empty.bindings.dartwas not regenerated locally. The workflow's Linux ffigen also changed theFILEtypedef to_IO_FILE.tool/native/sync_native_headers_and_bindings.sh --tag v0.4.1-1 --skip-ffigenreturned rc 0. The v0.4.1-1 header declaresllama_dart_tts_set_cancel_flagandllama_dart_tts_eval_callback.testCodeBuildHook, one run per target), matchSHA256SUMS.SHA256SUMSmatches the manifest digests for all 13 artifacts.binaryTarget, whose checksum isbab353ce…5418.llama_dart_tts_*exports, including both new symbols.nmlibllamadart.dylibnmper archllama.framework/llamallvm-readelf --dyn-symslibllamadart.sollvm-readobj --coff-exportsllamadart.dllLatency and output
ec6ad6f9bon v0.4.1.8d18c94a) with mmproj Q8_0 (6fd65188), from the pack lock.cancel_latency_bound(9.55, 8.51, 14.32)e2e9fe957, CPU, 3 runs (loadavg 9.32, 9.33, 7.16 at start): in-flight cancel median 25.9 ms, 8.7–40.4 ms (n=12);cancel_latency_boundpasses in all 3. Runs 2 and 3 pass all 15 checks. Run 1 fails onlypeak_memory_bound: its RSS aftergenerate, the baseline, was 3.71 GB instead of the 5.05–5.70 GB of every other run, so its 5.52 GB peak, inside the other runs' 5.42–5.73 GB, measured growth 1.487 against the 1.1 budget. 11.1 of 12 GB of swap was in use. All 3 runs give the same WAV hashes as the runs below.speech-0andspeech-3..6) hashac7ff03787edin all 6 CPU runs and37fd838eca9bin all 4 Metal runs.text-to-speech-smokeWAV isa97f472fb7dc2ee1on both sides.bca25981) passes all 20 checks on both sides, with loadavg 15.59 on the branch and 12.85 on main.ac7ff03787ed,e7a2275b83d9and1844169b2f3d, identically on main and this branch.speech-1, theafter_canceloutput, is1844169b2f3din all 3 main runs, in 2 of 3 branch runs and in all 3 runs ate2e9fe957, so it differs fromspeech-0because of its position, not because of the cancel before it. The third branch run givese45fb58db4c6, which none of those uncancelled syntheses gives. On Metal, uncancelled syntheses 1 to 5 give the same hashes on both sides, whilespeech-1varies on both sides and matches none of them.after_cancelpasses in every run.Every pack check, verbatim
Verify items
debugTtsCancelExportsForTesting()?.setCancelFlag, the service's own resolved API. N9, which reportssetCancelFlag: false, turns it red on v0.4.1-1, because the no-attach branch then sees the synthesis cancelled. N1 (nocb_eval) is caught by the integration test._resolveTtsApinow runs at the first mtmd context creation, including vision and STT. It runs once per service and latches both success and failure.dc5e95385requires it to record no startup diagnostics, and Windows x64 CI passed that assertion.LoadLibraryExW, the lookup records one deduplicated "Failed to preload Windows backend module …" entry.mtmd_init_from_file. This case was not run on a pre-TTS Windows runtime.Mutation proofs
Each mutation below was applied alone. Every control run was green.
Unit, fake worker. Tests from
llama_cpp_backend_test.dartandworker_test.dart:finallyfreenot idempotentfreekeeps the pointerroutes text-to-speech capability, progress, result, and cancelraiseafter free throwssi_addr=0x0freenever freesService, real Qwen3-TTS on CPU.
cb_evalnot setActual: <0>)createMultimodalContextbypasses the params helpercancel_latency_bound[1084.169, 1064.475, 1051.808, 1063.309] (loadavg 28.55)evalCallback: false_ttsEvalCallbackwithout its catchNative text-to-speech is unavailable in this runtime bundlesetCancelFlagevalCallbackGate.
-Nsuffix,doc/webgpu_bridge.mdRe-run at
e2e9fe957, each alone and reverted: U1 (4 red), U3, U4, U6, U7 and U8 (1 red each), U15 (1 red inworker_test.dart), N1 (integration red,Actual: <0>), N5 (integration red,evalCallback: false) and N3 (e2e between-steps red). Every red test is the one it hit before the merge. Controls: backend +48, worker +16, integration +18, e2e +3.Not covered, N14: this mutation removes the moved
ctxParams.use_gpuline.1c4067e6149band generation drops from 2.4–2.8 s in the other CPU runs to 1.68 s. That means the projector ran on the GPU.Test Plan
All at the merged head
e2e9fe957on v0.4.1-1, macOS arm64, unless noted.4c14e3c38changes onlywebsite/docs/changelog/recent-releases.md(11 added lines); runtime behaviour is unchanged. At4c14e3c38:prepare_workspacerc 0 with a clean tree,build_site.shprintedSite build ready,validate_links.shprintedLink validation passed, andverify_release_docs_versionsreturned rc 0 with the same pending companion bump.dart run tool/prepare_workspace.dart: rc 0, tree clean.dart format --output=none --set-exit-if-changed .:Formatted 720 files (0 changed).dart analyze --fatal-infos:No issues found!dart test -p vm -j 1 --exclude-tags local-only, withLLAMADART_TEST_MODELS_DIRholding stories15M:+2963 ~78: All tests passed!(15:17:58–15:19:35, loadavg 9.61→9.94). Before the merge:+2598 ~78.dart test -p chrome --exclude-tags local-only: CITest Web (Chrome), see Review Notes../tool/docs/build_site.shprintedSite build ready, and./tool/docs/validate_links.shprintedLink validation passed.dart run tool/testing/verify_release_docs_versions.dartreturned rc 0 and reported a pending bump:llamadart_llama_cpp_flutter 0.0.19 publishes v0.4.1, but Package.swift pins v0.4.1-1. With--release-prepit fails on that bump, which belongs to the release-prep PR.dart run tool/testing/check_webgpu_bridge_tag.dart --verify-manifestreturned rc 0.classify_high_risk_changeson the diff againstmainprintedClassification: high-risk / Surfaces: artifactConsumer, backendRuntime.text_to_speech_e2e_test.dartwith the Qwen3-TTS model gave+3: All tests passed!: the public API test plus both service tests on the attach path. WAVa97f472fb7dc2ee1. Main gave+1before the merge.native_symbol_integration_test.dart:+18.decision_unsupported_model_test.dart:+2.decision_engine_e2e_test.dart, CPU and Metal:+4each, 24 rows, 0 failures. Worst logit, probability and score differences are 0.1422, 0.0356 and 0.0609 on CPU and 0.1642, 0.0436 and 0.0253 on Metal (head onMTL0), the valuesdoc/decision_engine.mdrecords forlaya-Q8_0.gguf. The encoder was a local Q8_0 conversion of the cachedconvaiinnovations/laya@1c5edc17checkpoint, run with that checkpoint'smodel.safetensorsandrl_agent_config.json. The pinnedfr0stbit3/laya-gguffiles are not cached, so they were not run.Matrix Evidence
validation-speeche2e9fe957, CPU: 2 of 3 runs 15/15; run 1 failed onlypeak_memory_bound, from a low RSS baseline (Latency and output).text-to-speech-smoke+3ate2e9fe957; WAVa97f472fb7dc2ee1static-format-analyzeAnalyze & Lintroot-vm+2963 ~78ate2e9fe957root-chromeTest Web (Chrome)coverage-libLCOV line coverage: 83.76% (15963/19059)ate2e9fe957docs-siteDocs Build Checkrelease-doc-version-consistencyAnalyze & Lintwebgpu-bridge-tag-consistency--verify-manifest; CIAnalyze & Lintnative-hook-bundlesmacos-arm64-runtime-smokecb_eval: GLM-OCR vision output was864c7b54f620on main and this branch, CPU and Metal, with comparable timings. Ate2e9fe957against main8f6651f35, n=4 each: CPU 3658–3985 vs 3679–3780 ms, Metal 262–353 vs 257–386 ms. Text-only chat has no mtmd context; the chat smoke was not run.windows-x64-ci-runtimeTest Native (windows-latest);Native smoke (windows-latest)linux-x64-ci-runtimeTest Linux VM (with Coverage);Native Prompt Reuse Parity;Native smoke (ubuntu-latest)windows-arm64-hook-coverage,linux-arm64-runtime-smoke,android-release-device-pool,ios-arm64-device-smoke,physical-ios-speech-e2eweb-text-to-speech-smokedecision-model-smokedecision_engine_e2e_test.dartconvaiinnovations/laya@1c5edc17+ its head and config+4each, 0 failures; see Test Planhigh-risk-exact-head-independent-qa4c14e3c38/8f6651f35by a fresh independent reviewer, 0 blocking findings; the reviewer's model runs are ate2e9fe957, which differs only inwebsite/docs/changelog/recent-releases.md; readiness evaluator exit 2 (unverifiedPrerequisites); see the high-risk blockFollow-ups
LlamaEngine.dispose()andunloadModel()wait out a running TTS synthesis instead of cancelling it.dispose()never complete.contextCreatestarted onNativeLlamaBackendin the same tick asdispose(), which never completes. Pre-existing on main.peak_memory_boundfails with no memory growth when the RSS sample aftergenerateis low. Pre-existing on main.cb_evalslot.Review Notes
4c14e3c38against8f6651f35by a fresh independent reviewer, 0 blocking findings; see the high-risk block.4c14e3c38891e415cbb316c4b6babd2444b7e363. 25 checks pass and 5 are skipped, 0 fail.Documentation version and pin checksis skipped by path selection;Analyze & Lintruns both of its gates and passes.Test Native (windows-latest)passed the TTS integration test on the pinned DLL. LCOV 83.79% (15969/19059).High-risk regression review
4c14e3c38891e415cbb316c4b6babd2444b7e363/8f6651f35f12abb1967c3b1f74f106c1e53361b7.artifactConsumer,backendRuntime)e2e9fe957:4c14e3c38adds 11 lines towebsite/docs/changelog/recent-releases.md. Every other file is byte-identical, so the runtime review ofe2e9fe957against the same base carries over (below).CHANGELOG.mdUnreleasedentries after whitespace and bullet normalisation, and the next ten still match.git diff --checkis clean. Both linked issues exist.build_site.shprintedSite build ready(loadavg 21.18→24.66) andvalidate_links.shpassed (loadavg 24.23). The builtUnreleasedlist is still a tight list.verify_release_docs_versionsreturned rc 0 with the same pending companion bump;check_webgpu_bridge_tagreturned rc 0.e2e9fe957/8f6651f35:CHANGELOG.mdand the backend test, the PR's added and removed lines againstmainequal those ofec6ad6f9b..dc5e95385(524 lines, 21 files). The backend test has 48 tests: all 41 from this branch and all 42 frommain, none duplicated.synthesizeTextToSpeech,cancelTextToSpeech,disposeor_disposeWorker._mtmdContextParamsForis the onlycb_evalsite; the decision encoder, the head scheduler and the decision worker requests get no callback or flag.+2963 ~78; TTS e2e+3, WAVa97f472fb7dc2ee1; decision e2e+4on CPU and Metal with the documented worst diffs. CPU tts pack 15/15 (loadavg 13.44): in-flight cancel 36.8, 76.2, 38.9 and 31.6 ms, memory growth 1.049. GLM-OCR output864c7b54f620one2e9fe957andmain: CPU 4361–4470 vs 4537–5125 ms (loadavg 19–23), Metal 335–346 vs 342–345 ms.+3with the same WAV, through the non-attach branch. The tts pack fails onlycancel_latency_bound[1200, 2061, 25, 1488] ms, the old behaviour, with unchanged WAV hashes. The cancel-exports integration test is red there by design.main's dispose hold) turnsdispose frees a running synthesis cancel flag after the worker acksred. M1 (keeps only this branch's hold) stays green, as does removing that hold onmain. D1 (TTS callback on the decision encoder) leaves the decision e2e numbers identical.CHANGELOG.mdhas oneUnreleasedentry with no website mirror, the move of native release pins. It is pre-existing onmain8f6651f35.4c14e3c38and 83.76% ate2e9fe957, withlib/byte-identical: CI coverage varies slightly between runs.cb_evalare unmeasured.e2e9fe957, and the macOS arm64 v0.4.1-1 bundle and the LiteRT-LM v0.17.0-6 runtime at4c14e3c38. No result depended on them.4c14e3c38891e415cbb316c4b6babd2444b7e363/8f6651f35f12abb1967c3b1f74f106c1e53361b7(GitHubbaseRefOid, which ismainand an ancestor of the head). Both queried again for the readiness evidence; unchanged, 0 review threads.4c14e3c38for the readiness evidence, each alone and reverted: U1 (4 red), U2 (2), U8 (1), M2 (1), U15 (1), G1 (4), P1 (2, restores the base pinv0.4.1innative_release_pins.dart), N1 and N5 (integration red), N3 (e2e between-steps red). Controls: backend+48, worker+16, gate+58, integration+18, e2e+3with WAVa97f472fb7dc2ee1. Same results as ate2e9fe957.dart run tool/testing/high_risk_readiness.dart, with repository, PR, author, head and base from a fresh GitHub API query: exit 2,unverifiedPrerequisites(externalPrerequisitesUnavailable), over 24 changed files. That is the best local result: the evidence is internally consistent, and the external prerequisites do not exist.