Skip to content

TextToSpeech cancel() waits for the full synthesis (1.2s Metal / 2.5s CPU) instead of aborting #595

Description

@leehack

TextToSpeechTask.cancel() is honoured, but the caller still waits for the whole synthesis to finish. Cancelling costs as much as not cancelling.

Observed

cancel() produces the correct terminal state: no TextToSpeechFinalEvent is emitted and completion.state == TextToSpeechCompletionState.cancelled. Only the latency is wrong.

Measured 2026-09-22 on macOS arm64 (Apple M4 Max), locked Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M + mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0, via packages/llamadart_validation/bin/speech.dart --pack tts.

Interval from cancel() to terminal state, 10 runs each:

Backend Cancel latency
Metal 1245-1278 ms
CPU 2231-2565 ms

Across 30 repeated cancel/dispose/load/generate cycles, cancellation cost 93-104% of the full generation that followed it in the same run. The same ratio for the Qwen3-ASR stt pack is 0.04-0.31% — STT aborts essentially immediately.

CPU latency is roughly 2x Metal, i.e. it tracks backend synthesis cost. That is the giveaway that the in-flight call runs to completion rather than being interrupted.

Expected

Cancellation returns in bounded time, independent of how much synthesis remains — as the STT path already does.

What is verified in the code

  • _runTask in lib/src/core/speech/text_to_speech.dart checks task.isCancellationRequested only immediately before the backend call (text_to_speech.dart:583), immediately after it (text_to_speech.dart:625), and in the catch (text_to_speech.dart:641). Nothing observes cancellation during await _engine.synthesizeTextToSpeechBackend(...) (text_to_speech.dart:588).
  • The cancel hook is wired end to end: text_to_speech.dart:512 → LlamaEngine.cancelTextToSpeechBackend (engine.dart:1320) → BackendTextToSpeech.cancelTextToSpeech() → llama_cpp_backend.dart:935 posts TextToSpeechCancelRequest to the worker isolate, handled at worker.dart:149.
  • The native synthesis loop is not a single opaque call: llama_cpp_service.dart:7741-7786 steps, checks for LLAMA_DART_TTS_STATUS_CANCELLED / LLAMA_DART_TTS_STATE_CANCELLED, and yields with await Future<void>.delayed(Duration.zero) on every iteration.

So the machinery to abort mid-flight appears to exist on the native path, yet the measurement shows it does not shorten the call. The gap is somewhere between the posted TextToSpeechCancelRequest and that loop.

Where the defect is not

The native side is exonerated by source. In the pinned llamadart-native v0.4.1
(lib/src/hook/native_release_pins.dart:8 → const llamaCppTag = 'v0.4.1'),
src/llama_dart_wrapper.cpp:877-960 implements llama_dart_tts_step as one frame per call
— a single llama_sampler_sample, a single mtmd_helper_gen_audio_step_gen, a single
++frames_generated, return — and it tests the cancel flag before doing any of that work
(:886):

const bool active = tts->state == LLAMA_DART_TTS_STATE_PROCESSING_PROMPT ||
                    tts->state == LLAMA_DART_TTS_STATE_GENERATING;
if (active && tts->cancel_requested.load(std::memory_order_acquire)) {
  tts->state = LLAMA_DART_TTS_STATE_CANCELLED;
  tts->error = "TTS task cancelled";
  llama_dart_tts_release_task_resources(tts);
  llama_dart_tts_write_progress(tts, out_progress);
  return LLAMA_DART_TTS_STATUS_CANCELLED;
}

llama_dart_tts_cancel (:970) only stores an atomic bool. So once the flag is set, the
next step aborts — the native path can cancel within a single frame.

Step granularity is corroborated in-repo: physical_ios_speech_e2e_test.dart:1203 cancels
on a progress event with framesGenerated > 0 and fails explicitly if synthesis settled
before emitting one, so multiple steps demonstrably complete during a normal synthesis.

The contradiction to resolve

Every Dart link also reads as correct, and the service step loop yields to the event loop on
every frame (llama_cpp_service.dart:7786), so the worker isolate should service the posted
TextToSpeechCancelRequest within one frame. Read as source, cancellation should land in
one frame — tens of ms. It measures 1245 ms.

Source reading is exhausted; the remaining hop has to be observed. Three timestamps isolate
it to exactly one link: at cancel(), at worker.dart:149 when the message is received,
and at each api.step return.

Reproduction

dart packages/llamadart_validation/bin/speech.dart --pack tts

with the locked Qwen3-TTS model and projector above. Start a synthesis, call TextToSpeechTask.cancel() mid-generation, and time await task.done.

Scope

Native llama.cpp path on macOS arm64, Metal and CPU. Not checked on other platforms or on the WebGPU backend, which has its own cancelTextToSpeech at webgpu_backend.dart:2186.

Related

Qwen3-TTS runtime hardening is tracked by leehack/llamadart#322, which lists cancellation among the behaviours to record on Android Vulkan but carries no latency measurement and is scoped to that device qualification. This is a separate, reproducible defect on the native macOS path.

Context

Found while implementing item 2 of the remaining work in leehack/llamadart#325, in leehack/llamadart#594, which adds a 50 ms cancellation-latency bound to the speech validation pack. That PR makes bin/speech.dart --pack tts exit 1 on cancel_latency_bound alone; every other check in the pack still passes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingpriority:P3Watch or strategic work blocked by upstream/runtime/design dependencies

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions