Follow-up from leehack/llamadart#600 (for leehack/llamadart#599).
What happens
llamadart cancels a llama.cpp generation by setting a caller-owned byte (malloc<Int8>(1), llama_cpp_backend.dart:388). Its worker isolate reads the byte only between native calls. So a cancel waits for the llama_decode in flight: an STT audio-embedding decode, a text prefill, or a token decode.
llama_set_abort_callback could stop that decode, but it needs a native bool (*)(void *). libllamadart has none that reads a byte:
src/llama_dart_wrapper.h at main 448857da1 declares no such function.
nm -gU libllamadart.dylib from the v0.4.1 macOS arm64 bundle: the only abort or cancel symbols are upstream setters (llama_set_abort_callback, ggml_backend_cpu_set_abort_callback, ggml_backend_metal_set_abort_callback, ggml_metal_set_abort_callback, ggml_set_abort_callback, llama_context::set_abort_callback), ggml_abort, and llama_dart_tts_cancel.
The Dart alternatives were rejected in leehack/llamadart#600. NativeCallable.isolateLocal aborts the process when called from another thread, so it works only because upstream calls the callback on the decode caller thread. NativeCallable.isolateGroupBound is documented as experimental.
Upstream at the pin
third_party/llama.cpp is b29c606e28 at main and at v0.4.1.
Measured
Setup:
- Apple M4 Max, macOS arm64.
llamadart-native-macos-arm64-v0.4.1.tar.gz (sha256 41d0a529…b1487, matches the release asset).
- A stand-in reader in its own dylib:
bool i600_flag_abort(void *data) {
return __atomic_load_n((const uint8_t *)data, __ATOMIC_RELAXED) != 0;
}
- A Dart FFI driver: CPU device only,
n_gpu_layers 0, op_offload and offload_kqv false, 8 threads, n_batch = n_ubatch = batch size.
- Another pthread sets the byte 30-66% into the decode. Latency runs from that store to the
llama_decode return.
- 10 runs per row. 1-min loadavg 7.56-7.79.
| model, batch |
full decode, n=5 (ms) |
rc |
store → return (ms) |
| Qwen3-ASR-0.6B Q8_0, 104-row embedding batch (the size of a Qwen3-ASR audio-chunk decode) |
48.4-51.2 |
2 (10/10) |
0.15-0.59, median 0.31 |
| Qwen3-ASR-0.6B Q8_0, 512 tokens |
262.4-314.2 |
2 (10/10) |
0.20-2.42, median 0.94 |
| stories15M F32, 1024 tokens |
24.5-26.6 |
2 (10/10) |
0.15-1.56, median 0.50 |
- With the reader registered and the byte at 0, full decodes took 50.0-60.6, 257.7-270.8 and 25.3-27.3 ms.
- With
n_ubatch at half the batch and the byte set at 80%, all 15 runs returned 2. llama_memory_seq_pos_max was the first ubatch's last position (51, 255, 511).
- The leehack/llamadart#600 thread probe used the same bundle, plus CPU builds of the pinned source linked to libomp and to libgomp. It covered CPU with 1, 4 and 8 threads, and Metal. Across 26 runs of 255-101,216 calls each, every call ran on the thread that called
llama_decode.
Not known
- Not measured on Linux, Windows, Android or iOS. With OpenMP only the calling thread was checked, on macOS builds; latency was not measured.
- How llamadart would handle rc 2 from
mtmd_helper_decode_image_chunk and its token loop.
Fix direction: export a bool (*)(void *) that returns a relaxed atomic load of the byte at data, nonzero meaning abort.
Follow-up from leehack/llamadart#600 (for leehack/llamadart#599).
What happens
llamadart cancels a llama.cpp generation by setting a caller-owned byte (
malloc<Int8>(1),llama_cpp_backend.dart:388). Its worker isolate reads the byte only between native calls. So a cancel waits for thellama_decodein flight: an STT audio-embedding decode, a text prefill, or a token decode.llama_set_abort_callbackcould stop that decode, but it needs a nativebool (*)(void *). libllamadart has none that reads a byte:src/llama_dart_wrapper.hat main448857da1declares no such function.nm -gU libllamadart.dylibfrom the v0.4.1 macOS arm64 bundle: the only abort or cancel symbols are upstream setters (llama_set_abort_callback,ggml_backend_cpu_set_abort_callback,ggml_backend_metal_set_abort_callback,ggml_metal_set_abort_callback,ggml_set_abort_callback,llama_context::set_abort_callback),ggml_abort, andllama_dart_tts_cancel.The Dart alternatives were rejected in leehack/llamadart#600.
NativeCallable.isolateLocalaborts the process when called from another thread, so it works only because upstream calls the callback on the decode caller thread.NativeCallable.isolateGroupBoundis documented as experimental.Upstream at the pin
third_party/llama.cppisb29c606e28at main and at v0.4.1.llama_set_abort_callbackpasses the callback to each backend whose registry exposesggml_backend_set_abort_callback(llama-context.cpp#L1147-L1162). Only the CPU backend exposes it (ggml-cpu.cpp#L663-L665).llama.hdocuments it as "currently works only with CPU execution" (llama.h#L393-L396).truereturn stops the graph before the next node (ggml-cpu.c#L3150-L3153). Thread 0 is the thread that calledggml_graph_compute, with or without OpenMP (ggml-cpu.c#L3419-L3453).llama_decodereturns 2 and removes the aborted ubatch's positions. Earlier ubatches stay in memory (llama-context.cpp#L1831-L1858, llama.h#L982-L989).llama_set_abort_callbackdoes not reach a clip encode, which runs on clip's own backends (clip.cpp#L184): neither the encode half of an STT audio chunk nor the code2wav decode in leehack/llamadart-native#86.mtmd_context_params.cb_evaldoes (clip.cpp#L224-L225): leehack/llamadart-native#87 uses it for the code2wav decode, and leehack/llamadart-native#89 tracks the STT and vision encode.Measured
Setup:
llamadart-native-macos-arm64-v0.4.1.tar.gz(sha25641d0a529…b1487, matches the release asset).n_gpu_layers0,op_offloadandoffload_kqvfalse, 8 threads,n_batch=n_ubatch= batch size.llama_decodereturn.n_ubatchat half the batch and the byte set at 80%, all 15 runs returned 2.llama_memory_seq_pos_maxwas the first ubatch's last position (51, 255, 511).llama_decode.Not known
mtmd_helper_decode_image_chunkand its token loop.Fix direction: export a
bool (*)(void *)that returns a relaxed atomic load of the byte atdata, nonzero meaning abort.