TextToSpeechTask.cancel() is honoured, but the caller still waits for the whole synthesis to finish. Cancelling costs as much as not cancelling.
Observed
cancel() produces the correct terminal state: no TextToSpeechFinalEvent is emitted and completion.state == TextToSpeechCompletionState.cancelled. Only the latency is wrong.
Measured 2026-09-22 on macOS arm64 (Apple M4 Max), locked Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M + mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0, via packages/llamadart_validation/bin/speech.dart --pack tts.
Interval from cancel() to terminal state, 10 runs each:
| Backend |
Cancel latency |
| Metal |
1245-1278 ms |
| CPU |
2231-2565 ms |
Across 30 repeated cancel/dispose/load/generate cycles, cancellation cost 93-104% of the full generation that followed it in the same run. The same ratio for the Qwen3-ASR stt pack is 0.04-0.31% — STT aborts essentially immediately.
CPU latency is roughly 2x Metal, i.e. it tracks backend synthesis cost. That is the giveaway that the in-flight call runs to completion rather than being interrupted.
Expected
Cancellation returns in bounded time, independent of how much synthesis remains — as the STT path already does.
What is verified in the code
_runTask in lib/src/core/speech/text_to_speech.dart checks task.isCancellationRequested only immediately before the backend call (text_to_speech.dart:583), immediately after it (text_to_speech.dart:625), and in the catch (text_to_speech.dart:641). Nothing observes cancellation during await _engine.synthesizeTextToSpeechBackend(...) (text_to_speech.dart:588).
- The cancel hook is wired end to end:
text_to_speech.dart:512 → LlamaEngine.cancelTextToSpeechBackend (engine.dart:1320) → BackendTextToSpeech.cancelTextToSpeech() → llama_cpp_backend.dart:935 posts TextToSpeechCancelRequest to the worker isolate, handled at worker.dart:149.
- The native synthesis loop is not a single opaque call:
llama_cpp_service.dart:7741-7786 steps, checks for LLAMA_DART_TTS_STATUS_CANCELLED / LLAMA_DART_TTS_STATE_CANCELLED, and yields with await Future<void>.delayed(Duration.zero) on every iteration.
So the machinery to abort mid-flight appears to exist on the native path, yet the measurement shows it does not shorten the call. The gap is somewhere between the posted TextToSpeechCancelRequest and that loop.
Where the defect is not
The native side is exonerated by source. In the pinned llamadart-native v0.4.1
(lib/src/hook/native_release_pins.dart:8 → const llamaCppTag = 'v0.4.1'),
src/llama_dart_wrapper.cpp:877-960 implements llama_dart_tts_step as one frame per call
— a single llama_sampler_sample, a single mtmd_helper_gen_audio_step_gen, a single
++frames_generated, return — and it tests the cancel flag before doing any of that work
(:886):
const bool active = tts->state == LLAMA_DART_TTS_STATE_PROCESSING_PROMPT ||
tts->state == LLAMA_DART_TTS_STATE_GENERATING;
if (active && tts->cancel_requested.load(std::memory_order_acquire)) {
tts->state = LLAMA_DART_TTS_STATE_CANCELLED;
tts->error = "TTS task cancelled";
llama_dart_tts_release_task_resources(tts);
llama_dart_tts_write_progress(tts, out_progress);
return LLAMA_DART_TTS_STATUS_CANCELLED;
}
llama_dart_tts_cancel (:970) only stores an atomic bool. So once the flag is set, the
next step aborts — the native path can cancel within a single frame.
Step granularity is corroborated in-repo: physical_ios_speech_e2e_test.dart:1203 cancels
on a progress event with framesGenerated > 0 and fails explicitly if synthesis settled
before emitting one, so multiple steps demonstrably complete during a normal synthesis.
The contradiction to resolve
Every Dart link also reads as correct, and the service step loop yields to the event loop on
every frame (llama_cpp_service.dart:7786), so the worker isolate should service the posted
TextToSpeechCancelRequest within one frame. Read as source, cancellation should land in
one frame — tens of ms. It measures 1245 ms.
Source reading is exhausted; the remaining hop has to be observed. Three timestamps isolate
it to exactly one link: at cancel(), at worker.dart:149 when the message is received,
and at each api.step return.
Reproduction
dart packages/llamadart_validation/bin/speech.dart --pack tts
with the locked Qwen3-TTS model and projector above. Start a synthesis, call TextToSpeechTask.cancel() mid-generation, and time await task.done.
Scope
Native llama.cpp path on macOS arm64, Metal and CPU. Not checked on other platforms or on the WebGPU backend, which has its own cancelTextToSpeech at webgpu_backend.dart:2186.
Related
Qwen3-TTS runtime hardening is tracked by leehack/llamadart#322, which lists cancellation among the behaviours to record on Android Vulkan but carries no latency measurement and is scoped to that device qualification. This is a separate, reproducible defect on the native macOS path.
Context
Found while implementing item 2 of the remaining work in leehack/llamadart#325, in leehack/llamadart#594, which adds a 50 ms cancellation-latency bound to the speech validation pack. That PR makes bin/speech.dart --pack tts exit 1 on cancel_latency_bound alone; every other check in the pack still passes.
TextToSpeechTask.cancel()is honoured, but the caller still waits for the whole synthesis to finish. Cancelling costs as much as not cancelling.Observed
cancel()produces the correct terminal state: noTextToSpeechFinalEventis emitted andcompletion.state == TextToSpeechCompletionState.cancelled. Only the latency is wrong.Measured 2026-09-22 on macOS arm64 (Apple M4 Max), locked
Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M+mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0, viapackages/llamadart_validation/bin/speech.dart --pack tts.Interval from
cancel()to terminal state, 10 runs each:Across 30 repeated cancel/dispose/load/generate cycles, cancellation cost 93-104% of the full generation that followed it in the same run. The same ratio for the Qwen3-ASR
sttpack is 0.04-0.31% — STT aborts essentially immediately.CPU latency is roughly 2x Metal, i.e. it tracks backend synthesis cost. That is the giveaway that the in-flight call runs to completion rather than being interrupted.
Expected
Cancellation returns in bounded time, independent of how much synthesis remains — as the STT path already does.
What is verified in the code
_runTaskinlib/src/core/speech/text_to_speech.dartcheckstask.isCancellationRequestedonly immediately before the backend call (text_to_speech.dart:583), immediately after it (text_to_speech.dart:625), and in thecatch(text_to_speech.dart:641). Nothing observes cancellation duringawait _engine.synthesizeTextToSpeechBackend(...)(text_to_speech.dart:588).text_to_speech.dart:512→LlamaEngine.cancelTextToSpeechBackend(engine.dart:1320) →BackendTextToSpeech.cancelTextToSpeech()→llama_cpp_backend.dart:935postsTextToSpeechCancelRequestto the worker isolate, handled atworker.dart:149.llama_cpp_service.dart:7741-7786steps, checks forLLAMA_DART_TTS_STATUS_CANCELLED/LLAMA_DART_TTS_STATE_CANCELLED, and yields withawait Future<void>.delayed(Duration.zero)on every iteration.So the machinery to abort mid-flight appears to exist on the native path, yet the measurement shows it does not shorten the call. The gap is somewhere between the posted
TextToSpeechCancelRequestand that loop.Where the defect is not
The native side is exonerated by source. In the pinned
llamadart-nativev0.4.1(
lib/src/hook/native_release_pins.dart:8→const llamaCppTag = 'v0.4.1'),src/llama_dart_wrapper.cpp:877-960implementsllama_dart_tts_stepas one frame per call— a single
llama_sampler_sample, a singlemtmd_helper_gen_audio_step_gen, a single++frames_generated, return — and it tests the cancel flag before doing any of that work(
:886):llama_dart_tts_cancel(:970) only stores an atomic bool. So once the flag is set, thenext step aborts — the native path can cancel within a single frame.
Step granularity is corroborated in-repo:
physical_ios_speech_e2e_test.dart:1203cancelson a progress event with
framesGenerated > 0and fails explicitly if synthesis settledbefore emitting one, so multiple steps demonstrably complete during a normal synthesis.
The contradiction to resolve
Every Dart link also reads as correct, and the service step loop yields to the event loop on
every frame (
llama_cpp_service.dart:7786), so the worker isolate should service the postedTextToSpeechCancelRequestwithin one frame. Read as source, cancellation should land inone frame — tens of ms. It measures 1245 ms.
Source reading is exhausted; the remaining hop has to be observed. Three timestamps isolate
it to exactly one link: at
cancel(), atworker.dart:149when the message is received,and at each
api.stepreturn.Reproduction
with the locked Qwen3-TTS model and projector above. Start a synthesis, call
TextToSpeechTask.cancel()mid-generation, and timeawait task.done.Scope
Native llama.cpp path on macOS arm64, Metal and CPU. Not checked on other platforms or on the WebGPU backend, which has its own
cancelTextToSpeechatwebgpu_backend.dart:2186.Related
Qwen3-TTS runtime hardening is tracked by leehack/llamadart#322, which lists cancellation among the behaviours to record on Android Vulkan but carries no latency measurement and is scoped to that device qualification. This is a separate, reproducible defect on the native macOS path.
Context
Found while implementing item 2 of the remaining work in leehack/llamadart#325, in leehack/llamadart#594, which adds a 50 ms cancellation-latency bound to the speech validation pack. That PR makes
bin/speech.dart --pack ttsexit 1 oncancel_latency_boundalone; every other check in the pack still passes.