Skip to content

Export cancel-byte callbacks so llamadart can stop an in-flight llama_decode and clip encode #88

Description

@leehack

Follow-up from leehack/llamadart#600 (for leehack/llamadart#599).

What happens

llamadart cancels a llama.cpp generation by setting a caller-owned byte (malloc<Int8>(1), llama_cpp_backend.dart:388). Its worker isolate reads the byte only between native calls. So a cancel waits for the llama_decode in flight: an STT audio-embedding decode, a text prefill, or a token decode.

llama_set_abort_callback could stop that decode, but it needs a native bool (*)(void *). libllamadart has none that reads a byte:

  • src/llama_dart_wrapper.h at main 448857da1 declares no such function.
  • nm -gU libllamadart.dylib from the v0.4.1 macOS arm64 bundle: the only abort or cancel symbols are upstream setters (llama_set_abort_callback, ggml_backend_cpu_set_abort_callback, ggml_backend_metal_set_abort_callback, ggml_metal_set_abort_callback, ggml_set_abort_callback, llama_context::set_abort_callback), ggml_abort, and llama_dart_tts_cancel.

The Dart alternatives were rejected in leehack/llamadart#600. NativeCallable.isolateLocal aborts the process when called from another thread, so it works only because upstream calls the callback on the decode caller thread. NativeCallable.isolateGroupBound is documented as experimental.

Upstream at the pin

third_party/llama.cpp is b29c606e28 at main and at v0.4.1.

Measured

Setup:

  • Apple M4 Max, macOS arm64.
  • llamadart-native-macos-arm64-v0.4.1.tar.gz (sha256 41d0a529…b1487, matches the release asset).
  • A stand-in reader in its own dylib:
    bool i600_flag_abort(void *data) {
      return __atomic_load_n((const uint8_t *)data, __ATOMIC_RELAXED) != 0;
    }
  • A Dart FFI driver: CPU device only, n_gpu_layers 0, op_offload and offload_kqv false, 8 threads, n_batch = n_ubatch = batch size.
  • Another pthread sets the byte 30-66% into the decode. Latency runs from that store to the llama_decode return.
  • 10 runs per row. 1-min loadavg 7.56-7.79.
model, batch full decode, n=5 (ms) rc store → return (ms)
Qwen3-ASR-0.6B Q8_0, 104-row embedding batch (the size of a Qwen3-ASR audio-chunk decode) 48.4-51.2 2 (10/10) 0.15-0.59, median 0.31
Qwen3-ASR-0.6B Q8_0, 512 tokens 262.4-314.2 2 (10/10) 0.20-2.42, median 0.94
stories15M F32, 1024 tokens 24.5-26.6 2 (10/10) 0.15-1.56, median 0.50
  • With the reader registered and the byte at 0, full decodes took 50.0-60.6, 257.7-270.8 and 25.3-27.3 ms.
  • With n_ubatch at half the batch and the byte set at 80%, all 15 runs returned 2. llama_memory_seq_pos_max was the first ubatch's last position (51, 255, 511).
  • The leehack/llamadart#600 thread probe used the same bundle, plus CPU builds of the pinned source linked to libomp and to libgomp. It covered CPU with 1, 4 and 8 threads, and Metal. Across 26 runs of 255-101,216 calls each, every call ran on the thread that called llama_decode.

Not known

  • Not measured on Linux, Windows, Android or iOS. With OpenMP only the calling thread was checked, on macOS builds; latency was not measured.
  • How llamadart would handle rc 2 from mtmd_helper_decode_image_chunk and its token loop.

Fix direction: export a bool (*)(void *) that returns a relaxed atomic load of the byte at data, nonzero meaning abort.

Activity

  1. leehack commented on Sep 25, 2026

    @leehack
    OwnerAuthor

    Merged #89 into this issue. Both ask libllamadart to export a callback that reads llamadart's cancel byte, and it is probably one PR.

    • llama_decode (this issue). Export a bool (*)(void *) abort callback for llama_set_abort_callback. It reaches only the CPU backend.
    • Clip encode for STT and vision (from No exported cb_eval stops an in-flight STT or vision clip encode on a cancel byte #89). llama_set_abort_callback does not reach mtmd_encode_chunk, because clip creates its own backends. The hook is mtmd_context_params.cb_eval, whose slot #87 already uses for llama_dart_tts_eval_callback. Outside a TTS step, that callback should stop a clip encode when the caller-owned byte is nonzero. Measured in No exported cb_eval stops an in-flight STT or vision clip encode on a cancel byte #89: CPU and Metal stops within one node (at most 29.7 ms on CPU), and contexts stay usable. Asking for every node costs 13–27% on CPU and 4.8–6.3× on Metal, so the ask rate is a design decision. A stopped encode returns 0, so llamadart must re-read the byte before the decode.

    Consumer side: leehack/llamadart#660 (cancel during prompt evaluation).

  2. changed the title [-]No exported abort callback reads a cancel byte, so llamadart cannot stop an in-flight CPU llama_decode[/-] [+]Export cancel-byte callbacks so llamadart can stop an in-flight llama_decode and clip encode[/+] on Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority:P3Useful cleanup or longer-term work

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions