llama_dart_tts_step checks cancel_requested only on entry. The step that ends speech also runs the Qwen3-TTS code2wav decode of every codec frame still buffered. A cancel issued during that step is not seen by native: the step finishes the decode and returns COMPLETED. That final step takes 1.3-1.6 s on CPU and 0.37-0.38 s on Metal. A cancel during any other step is observed when the next step starts.
Measured on macOS arm64 (Apple M4 Max):
v0.4.1 (a4ee6b9fa7), as pinned by llamadart aed0bd2b52.
- Qwen3-TTS-12Hz-1.7B-Base Q4_K_M + mmproj Q8_0.
- Text "Hello from llamadart. The answer is forty two." (25-38 frames).
- 1-min loadavg 8.30-39.72.
Four timestamps
- T0:
TextToSpeechTask.cancel() on the main isolate.
- T1: the worker isolate receives the cancel.
- T2: entry of the
llama_dart_tts_step call that returns CANCELLED.
- T3: the task is done.
Offsets from T0, in ms:
| backend |
cancel lands in |
n |
T1 |
T2 |
T3 |
| CPU |
frame step |
2 |
3.5-5.4 |
3.7-5.4 |
4.9-5.6 |
| CPU |
final step |
17 |
415.3-1158.5 |
never |
415.4-1158.5 |
| Metal |
frame step |
13 |
0.34-19.0 |
0.36-19.1 |
0.48-19.2 |
| Metal |
final step |
2 |
149.8-292.4 |
never |
149.8-292.4 |
In all 19 final-step samples, the step returns OK with state COMPLETED, and T1 comes 0.14-0.49 ms after it returns.
In a scratch build, calling llama_dart_tts_cancel from the main isolate during the step did not shorten it (loadavg 65.57-117.02). All 6 final steps ran 1296.5-1343.9 ms and returned COMPLETED. The same call during a frame step makes the next step return CANCELLED with no worker message.
Where the delay sits
The delay is inside the native step. A stack profile of the worker thread puts the final step in llama_dart_tts_step → llama_dart_tts_finish_output → qwen3tts_gen_audio_pipeline::get_output → flush_gen_wav → mtmd_gen_audio_process → clip_encode → … → ggml_graph_compute. Metal shows the same chain down to clip_encode.
At a4ee6b9fa7:
src/llama_dart_wrapper.cpp:886: the only cancel check, on entry.
:936 and :957: the end-of-speech branches call llama_dart_tts_finish_output, which calls mtmd_helper_gen_audio_get_output (:202).
:944: each frame step calls mtmd_helper_gen_audio_step_gen.
In llama.cpp b29c606e28, tools/mtmd/mtmd-helper-gen.cpp:
:438: Qwen3-TTS buffers codec frames in windows of 72.
:276: step_gen decodes a full window.
:305: get_output decodes the remainder.
- Neither path has a cancel hook.
Past 72 frames, the same thing happens mid-synthesis. In 152-181-frame runs on CPU, the steps that produced frames 72 and 144 took 1.3-1.7 s. So did the final step, including one with only 8 frames left.
In llamadart at aed0bd2b52, the cancel reaches native only through the worker isolate (lib/src/backends/llama_cpp/worker.dart:149). That isolate stays inside this FFI call (lib/src/backends/llama_cpp/llama_cpp_service.dart:7761) until the call returns. The direct-call run shows that delivering the cancel sooner does not help while the call has no check.
Impact
- On CPU, a cancel that lands in the final step waits for the rest of it, up to 1.16 s in the table above. The task is still reported as cancelled.
- leehack/llamadart#594 (open) adds a 500 ms
cancel_latency_bound to the speech validation pack. With it:
- The
tts pack on CPU fails that bound on every run: 24 of 24 samples at 862.5-1168.4 ms over 6 runs.
- The pack cancels at half the previous generation time. On CPU that point falls inside the 1.3-1.6 s final step.
dart run tool/testing/run_local_e2e.dart --scenario validation-speech-tts runs the pack with --backend cpu and exits 1. cancel_latency_bound is its only failing row.
- Metal passes, 16 of 16 at 0.50-21.1 ms, because its half-way cancel lands in a frame step before the final step.
Related
llama_dart_tts_stepcheckscancel_requestedonly on entry. The step that ends speech also runs the Qwen3-TTS code2wav decode of every codec frame still buffered. A cancel issued during that step is not seen by native: the step finishes the decode and returnsCOMPLETED. That final step takes 1.3-1.6 s on CPU and 0.37-0.38 s on Metal. A cancel during any other step is observed when the next step starts.Measured on macOS arm64 (Apple M4 Max):
v0.4.1(a4ee6b9fa7), as pinned by llamadartaed0bd2b52.Four timestamps
TextToSpeechTask.cancel()on the main isolate.llama_dart_tts_stepcall that returnsCANCELLED.Offsets from T0, in ms:
In all 19 final-step samples, the step returns
OKwith stateCOMPLETED, and T1 comes 0.14-0.49 ms after it returns.In a scratch build, calling
llama_dart_tts_cancelfrom the main isolate during the step did not shorten it (loadavg 65.57-117.02). All 6 final steps ran 1296.5-1343.9 ms and returnedCOMPLETED. The same call during a frame step makes the next step returnCANCELLEDwith no worker message.Where the delay sits
The delay is inside the native step. A stack profile of the worker thread puts the final step in
llama_dart_tts_step→llama_dart_tts_finish_output→qwen3tts_gen_audio_pipeline::get_output→flush_gen_wav→mtmd_gen_audio_process→clip_encode→ … →ggml_graph_compute. Metal shows the same chain down toclip_encode.At
a4ee6b9fa7:src/llama_dart_wrapper.cpp:886: the only cancel check, on entry.:936and:957: the end-of-speech branches callllama_dart_tts_finish_output, which callsmtmd_helper_gen_audio_get_output(:202).:944: each frame step callsmtmd_helper_gen_audio_step_gen.In llama.cpp
b29c606e28,tools/mtmd/mtmd-helper-gen.cpp::438: Qwen3-TTS buffers codec frames in windows of 72.:276:step_gendecodes a full window.:305:get_outputdecodes the remainder.Past 72 frames, the same thing happens mid-synthesis. In 152-181-frame runs on CPU, the steps that produced frames 72 and 144 took 1.3-1.7 s. So did the final step, including one with only 8 frames left.
In llamadart at
aed0bd2b52, the cancel reaches native only through the worker isolate (lib/src/backends/llama_cpp/worker.dart:149). That isolate stays inside this FFI call (lib/src/backends/llama_cpp/llama_cpp_service.dart:7761) until the call returns. The direct-call run shows that delivering the cancel sooner does not help while the call has no check.Impact
cancel_latency_boundto the speech validation pack. With it:ttspack on CPU fails that bound on every run: 24 of 24 samples at 862.5-1168.4 ms over 6 runs.dart run tool/testing/run_local_e2e.dart --scenario validation-speech-ttsruns the pack with--backend cpuand exits 1.cancel_latency_boundis its only failing row.Related
llama_dart_tts_stepas one frame per call does not hold for the final step.