Skip to content

No exported cb_eval stops an in-flight STT or vision clip encode on a cancel byte #89

Description

@leehack

Follow-up from leehack/llamadart#600, for leehack/llamadart#599. leehack/llamadart-native#88 covers the in-flight llama_decode; this issue covers the clip encode.

What happens

With leehack/llamadart#600, llamadart's multimodal ingest reads the cancel byte before each chunk and between a media chunk's mtmd_encode_chunk and its embedding decode. The encode in flight still runs to the end. At that PR's CPU f=0.1 lead, a Qwen3-ASR cancel lands in the 104-token audio encode, waits it out, and takes 106.6-109.6 ms (median 108.0).

Upstream can stop it. mtmd_context_params.cb_eval is installed on the scheduler of every clip context, vision and audio (clip.cpp#L224-L225, mtmd.cpp#L575-L576). When it asks for a node and then returns false for it, the scheduler stops that split (ggml-backend.cpp#L1802-L1831).

llamadart creates the mtmd context with cb_eval null (createMultimodalContext). libllamadart has no callback for it that reads a cancel byte: src/ at main 448857da1 never mentions cb_eval.

Measured

Setup:

  • Apple M4 Max, macOS arm64. llamadart-native-macos-arm64-v0.4.1.tar.gz (sha256 41d0a529…b1487), llama.cpp b29c606e28.
  • Qwen3-ASR-0.6B Q8_0 + mmproj Q8_0, jfk.wav: a 104-token and a 52-token audio chunk.
  • A Dart FFI driver sets only cb_eval and cb_eval_user_data in mtmd_context_params_default() (4 threads). Each callback asks for every node and returns false for it when the byte is nonzero: a C function in its own dylib (relaxed atomic load), or a Dart NativeCallable.isolateLocal.
  • CPU runs set GGML_METAL_DEVICES=0. 1-min loadavg 8.45-11.40.

Full encode, ms, min-max (median), n=5:

backend, chunk no callback C callback, byte 0 Dart callback, byte 0
CPU, 104 tokens 149.1-151.8 (150.5) 167.9-172.4 (169.8) 167.9-171.4 (169.4)
CPU, 52 tokens 74.0-75.6 (74.5) 88.8-101.1 (94.4) 90.2-97.0 (93.7)
Metal, 104 tokens 15.5-15.6 (15.6) 72.2-76.5 (75.4) 73.0-75.9 (74.2)
Metal, 52 tokens 11.5-11.7 (11.5) 70.7-76.3 (72.6) 70.9-73.7 (71.7)

Stopped encode, ms. "Before": the byte is set before mtmd_encode_chunk (n=5 per callback). "During": another thread sets it after 25%, 50% or 75% of the median full encode; time from that store to the return (C n=17, Dart n=5 at 50%).

backend, chunk before, C / Dart during, C during, Dart
CPU, 104 tokens 0.13-0.14 / 0.14-0.17 0.05-12.44 0.04-0.40
CPU, 52 tokens 0.12-0.14 / 0.13-0.15 0.07-6.64 0.06-0.14
Metal, 104 tokens 0.45-0.49 / 0.48-0.51 0.05-0.29 0.08-0.14
Metal, 52 tokens 0.44-0.53 / 0.47-0.54 0.04-0.25 0.04-0.15
  • A stop waits for the node in flight. The longest time between two checks, each time a MUL_MAT, was 29.1-29.7 ms on CPU (104 tokens), 14.4-16.2 ms (52 tokens) and at most 4.3 ms on Metal (n=5 each).
  • A stopped encode returns 0 after 2 callback calls; a full one makes 1,214 (607 nodes). No later split called back, so CPU and Metal each ran the encode as one split.
  • The output is not the chunk's embeddings. After a full encode of the 52-token chunk, a stop before the 104-token encode left the first 53,248 values of mtmd_get_output_embd equal to the 52-token output. No value of an output stopped during the encode matched the full encode's.
  • The contexts stay usable. After 21 stops on each callback's context, a full transcription on the same mtmd and llama contexts gave the same 29 tokens as before them (language English<asr_text>And so, my fellow Americans, ask not …), and full re-encodes were bit-identical to the earlier ones. CPU and Metal, both callbacks.
  • llama_set_abort_callback does not reach the encode, because clip creates its own backends (clip.cpp#L184). With the byte set and a C abort callback on the llama context, the CPU 104-token encode ran in full (149.9-166.5 ms, n=3) with 0 abort-callback calls; the next llama_decode returned 2.

Why a native callback

  • leehack/llamadart#600 rejected a Dart callback for thread-safety. NativeCallable.isolateLocal aborts the process when called from another thread; the probe's Dart callback works only because the scheduler calls back on the thread that runs the encode. NativeCallable.isolateGroupBound is documented as experimental.
  • cb_eval is one slot per mtmd context. leehack/llamadart-native#87 adds llama_dart_tts_eval_callback for it, and the llamadart follow-up it names sets that from createMultimodalContext. Outside a TTS step it asks for nothing. So the encode stop belongs in that exported callback, reading a caller-owned cancel byte.

Costs and caveats

  • Asking for every node slowed full encodes (medians) by 13-27% on CPU and 4.8-6.3× on Metal, about the same for C and Dart. The scheduler computes each asked node as its own graph view and synchronizes after it. Metal transcriptions took 267.1-275.9 ms with a callback (n=4) against 133.5-138.4 ms without (n=2).
  • Fewer asks cost less, and a stop then waits for the running chunk. Inside a TTS step, leehack/llamadart-native#87 ends a chunk at the first MUL_MAT after 2.5e9 of work; its encoder overhead is unmeasured.
  • On Metal, per-node views changed the output bits: the embeddings differed from those without a callback, though the 29 tokens matched. On CPU they were bit-identical.
  • From leehack/llamadart-native#87: on a multi-split backend a break stops only the current split, and each later split computes one node. CUDA fuses groups that continue past a MUL_MAT, so a chunk end can split one and change output.
  • A stopped encode returns 0, so the caller must re-read the byte before the decode, as leehack/llamadart#600's chunk loop does. mtmd_helper_eval_chunks, which that PR keeps as the fallback when the chunk-level functions are missing, decodes whatever the encode left (mtmd-helper.cpp#L251-L261).
  • Vision encodes go through the same hook and were not measured.

Fix direction: outside a TTS step, let the exported cb_eval stop a clip encode when a caller-owned byte is nonzero. llamadart then attaches its per-generation cancel byte and keeps its post-encode check.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority:P3Useful cleanup or longer-term work

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions