Follow-up from leehack/llamadart#600, for leehack/llamadart#599. leehack/llamadart-native#88 covers the in-flight llama_decode; this issue covers the clip encode.
What happens
With leehack/llamadart#600, llamadart's multimodal ingest reads the cancel byte before each chunk and between a media chunk's mtmd_encode_chunk and its embedding decode. The encode in flight still runs to the end. At that PR's CPU f=0.1 lead, a Qwen3-ASR cancel lands in the 104-token audio encode, waits it out, and takes 106.6-109.6 ms (median 108.0).
Upstream can stop it. mtmd_context_params.cb_eval is installed on the scheduler of every clip context, vision and audio (clip.cpp#L224-L225, mtmd.cpp#L575-L576). When it asks for a node and then returns false for it, the scheduler stops that split (ggml-backend.cpp#L1802-L1831).
llamadart creates the mtmd context with cb_eval null (createMultimodalContext). libllamadart has no callback for it that reads a cancel byte: src/ at main 448857da1 never mentions cb_eval.
Measured
Setup:
- Apple M4 Max, macOS arm64.
llamadart-native-macos-arm64-v0.4.1.tar.gz (sha256 41d0a529…b1487), llama.cpp b29c606e28.
- Qwen3-ASR-0.6B Q8_0 + mmproj Q8_0,
jfk.wav: a 104-token and a 52-token audio chunk.
- A Dart FFI driver sets only
cb_eval and cb_eval_user_data in mtmd_context_params_default() (4 threads). Each callback asks for every node and returns false for it when the byte is nonzero: a C function in its own dylib (relaxed atomic load), or a Dart NativeCallable.isolateLocal.
- CPU runs set
GGML_METAL_DEVICES=0. 1-min loadavg 8.45-11.40.
Full encode, ms, min-max (median), n=5:
| backend, chunk |
no callback |
C callback, byte 0 |
Dart callback, byte 0 |
| CPU, 104 tokens |
149.1-151.8 (150.5) |
167.9-172.4 (169.8) |
167.9-171.4 (169.4) |
| CPU, 52 tokens |
74.0-75.6 (74.5) |
88.8-101.1 (94.4) |
90.2-97.0 (93.7) |
| Metal, 104 tokens |
15.5-15.6 (15.6) |
72.2-76.5 (75.4) |
73.0-75.9 (74.2) |
| Metal, 52 tokens |
11.5-11.7 (11.5) |
70.7-76.3 (72.6) |
70.9-73.7 (71.7) |
Stopped encode, ms. "Before": the byte is set before mtmd_encode_chunk (n=5 per callback). "During": another thread sets it after 25%, 50% or 75% of the median full encode; time from that store to the return (C n=17, Dart n=5 at 50%).
| backend, chunk |
before, C / Dart |
during, C |
during, Dart |
| CPU, 104 tokens |
0.13-0.14 / 0.14-0.17 |
0.05-12.44 |
0.04-0.40 |
| CPU, 52 tokens |
0.12-0.14 / 0.13-0.15 |
0.07-6.64 |
0.06-0.14 |
| Metal, 104 tokens |
0.45-0.49 / 0.48-0.51 |
0.05-0.29 |
0.08-0.14 |
| Metal, 52 tokens |
0.44-0.53 / 0.47-0.54 |
0.04-0.25 |
0.04-0.15 |
- A stop waits for the node in flight. The longest time between two checks, each time a
MUL_MAT, was 29.1-29.7 ms on CPU (104 tokens), 14.4-16.2 ms (52 tokens) and at most 4.3 ms on Metal (n=5 each).
- A stopped encode returns 0 after 2 callback calls; a full one makes 1,214 (607 nodes). No later split called back, so CPU and Metal each ran the encode as one split.
- The output is not the chunk's embeddings. After a full encode of the 52-token chunk, a stop before the 104-token encode left the first 53,248 values of
mtmd_get_output_embd equal to the 52-token output. No value of an output stopped during the encode matched the full encode's.
- The contexts stay usable. After 21 stops on each callback's context, a full transcription on the same mtmd and llama contexts gave the same 29 tokens as before them (
language English<asr_text>And so, my fellow Americans, ask not …), and full re-encodes were bit-identical to the earlier ones. CPU and Metal, both callbacks.
llama_set_abort_callback does not reach the encode, because clip creates its own backends (clip.cpp#L184). With the byte set and a C abort callback on the llama context, the CPU 104-token encode ran in full (149.9-166.5 ms, n=3) with 0 abort-callback calls; the next llama_decode returned 2.
Why a native callback
- leehack/llamadart#600 rejected a Dart callback for thread-safety.
NativeCallable.isolateLocal aborts the process when called from another thread; the probe's Dart callback works only because the scheduler calls back on the thread that runs the encode. NativeCallable.isolateGroupBound is documented as experimental.
cb_eval is one slot per mtmd context. leehack/llamadart-native#87 adds llama_dart_tts_eval_callback for it, and the llamadart follow-up it names sets that from createMultimodalContext. Outside a TTS step it asks for nothing. So the encode stop belongs in that exported callback, reading a caller-owned cancel byte.
Costs and caveats
- Asking for every node slowed full encodes (medians) by 13-27% on CPU and 4.8-6.3× on Metal, about the same for C and Dart. The scheduler computes each asked node as its own graph view and synchronizes after it. Metal transcriptions took 267.1-275.9 ms with a callback (n=4) against 133.5-138.4 ms without (n=2).
- Fewer asks cost less, and a stop then waits for the running chunk. Inside a TTS step, leehack/llamadart-native#87 ends a chunk at the first
MUL_MAT after 2.5e9 of work; its encoder overhead is unmeasured.
- On Metal, per-node views changed the output bits: the embeddings differed from those without a callback, though the 29 tokens matched. On CPU they were bit-identical.
- From leehack/llamadart-native#87: on a multi-split backend a break stops only the current split, and each later split computes one node. CUDA fuses groups that continue past a
MUL_MAT, so a chunk end can split one and change output.
- A stopped encode returns 0, so the caller must re-read the byte before the decode, as leehack/llamadart#600's chunk loop does.
mtmd_helper_eval_chunks, which that PR keeps as the fallback when the chunk-level functions are missing, decodes whatever the encode left (mtmd-helper.cpp#L251-L261).
- Vision encodes go through the same hook and were not measured.
Fix direction: outside a TTS step, let the exported cb_eval stop a clip encode when a caller-owned byte is nonzero. llamadart then attaches its per-generation cancel byte and keeps its post-encode check.
Follow-up from leehack/llamadart#600, for leehack/llamadart#599. leehack/llamadart-native#88 covers the in-flight
llama_decode; this issue covers the clip encode.What happens
With leehack/llamadart#600, llamadart's multimodal ingest reads the cancel byte before each chunk and between a media chunk's
mtmd_encode_chunkand its embedding decode. The encode in flight still runs to the end. At that PR's CPU f=0.1 lead, a Qwen3-ASR cancel lands in the 104-token audio encode, waits it out, and takes 106.6-109.6 ms (median 108.0).Upstream can stop it.
mtmd_context_params.cb_evalis installed on the scheduler of every clip context, vision and audio (clip.cpp#L224-L225, mtmd.cpp#L575-L576). When it asks for a node and then returns false for it, the scheduler stops that split (ggml-backend.cpp#L1802-L1831).llamadart creates the mtmd context with
cb_evalnull (createMultimodalContext). libllamadart has no callback for it that reads a cancel byte:src/at main448857da1never mentionscb_eval.Measured
Setup:
llamadart-native-macos-arm64-v0.4.1.tar.gz(sha25641d0a529…b1487), llama.cppb29c606e28.jfk.wav: a 104-token and a 52-token audio chunk.cb_evalandcb_eval_user_datainmtmd_context_params_default()(4 threads). Each callback asks for every node and returns false for it when the byte is nonzero: a C function in its own dylib (relaxed atomic load), or a DartNativeCallable.isolateLocal.GGML_METAL_DEVICES=0. 1-min loadavg 8.45-11.40.Full encode, ms, min-max (median), n=5:
Stopped encode, ms. "Before": the byte is set before
mtmd_encode_chunk(n=5 per callback). "During": another thread sets it after 25%, 50% or 75% of the median full encode; time from that store to the return (C n=17, Dart n=5 at 50%).MUL_MAT, was 29.1-29.7 ms on CPU (104 tokens), 14.4-16.2 ms (52 tokens) and at most 4.3 ms on Metal (n=5 each).mtmd_get_output_embdequal to the 52-token output. No value of an output stopped during the encode matched the full encode's.language English<asr_text>And so, my fellow Americans, ask not …), and full re-encodes were bit-identical to the earlier ones. CPU and Metal, both callbacks.llama_set_abort_callbackdoes not reach the encode, because clip creates its own backends (clip.cpp#L184). With the byte set and a C abort callback on the llama context, the CPU 104-token encode ran in full (149.9-166.5 ms, n=3) with 0 abort-callback calls; the nextllama_decodereturned 2.Why a native callback
NativeCallable.isolateLocalaborts the process when called from another thread; the probe's Dart callback works only because the scheduler calls back on the thread that runs the encode.NativeCallable.isolateGroupBoundis documented as experimental.cb_evalis one slot per mtmd context. leehack/llamadart-native#87 addsllama_dart_tts_eval_callbackfor it, and the llamadart follow-up it names sets that fromcreateMultimodalContext. Outside a TTS step it asks for nothing. So the encode stop belongs in that exported callback, reading a caller-owned cancel byte.Costs and caveats
MUL_MATafter 2.5e9 of work; its encoder overhead is unmeasured.MUL_MAT, so a chunk end can split one and change output.mtmd_helper_eval_chunks, which that PR keeps as the fallback when the chunk-level functions are missing, decodes whatever the encode left (mtmd-helper.cpp#L251-L261).Fix direction: outside a TTS step, let the exported
cb_evalstop a clip encode when a caller-owned byte is nonzero. llamadart then attaches its per-generation cancel byte and keeps its post-encode check.