A grammar-constrained createCompletion can abort the Wasm core instead of returning text or an error:
Aborted(undefined). Build with -sASSERTIONS for more info.
In worker mode the abort also moves the bridge to the main thread for the rest of the session (__llamadartBridgeWorkerFallbackReason holds the abort text).
Repro (same on v0.1.44 and v0.1.47), Qwen3.5-0.8B Q4_K_M on the CPU:
await bridge.loadModelFromUrl(modelUrl, { nCtx: 1024, nThreads: 2, nGpuLayers: 0 });
await bridge.createCompletion(
'<|im_start|>user\nIs the sky blue? Answer yes or no.<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n',
{ nPredict: 8, temp: 0, topK: 1, seed: 1, grammar: 'root ::= "yes" | "no"' },
);
- The same grammar also aborts with
topK: 40, temp: 0.8.
- A small JSON grammar aborts with
topK: 1 and succeeds with topK: 40.
- stories15M with the yes/no grammar and
topK: 1 aborts too.
Likely cause: create_sampler (src/llama_webgpu_core.cpp:807-822 at 64ba825) adds the grammar sampler after top-k and top-p. When truncation keeps only tokens the grammar rejects, no candidate is left. llama.cpp's common_sampler avoids this: it checks the sampled token against the grammar and, if rejected, resamples with the grammar applied first.
Expected: a grammar-constrained completion returns text the grammar accepts, or rejects with an error the caller can handle.
Found while validating the bridge pin bump in leehack/llamadart#620. llamadart passes GenerationParams.grammar and topK through unchanged.
A grammar-constrained
createCompletioncan abort the Wasm core instead of returning text or an error:In worker mode the abort also moves the bridge to the main thread for the rest of the session (
__llamadartBridgeWorkerFallbackReasonholds the abort text).Repro (same on
v0.1.44andv0.1.47), Qwen3.5-0.8B Q4_K_M on the CPU:topK: 40, temp: 0.8.topK: 1and succeeds withtopK: 40.topK: 1aborts too.Likely cause:
create_sampler(src/llama_webgpu_core.cpp:807-822at 64ba825) adds the grammar sampler after top-k and top-p. When truncation keeps only tokens the grammar rejects, no candidate is left. llama.cpp'scommon_sampleravoids this: it checks the sampled token against the grammar and, if rejected, resamples with the grammar applied first.Expected: a grammar-constrained completion returns text the grammar accepts, or rejects with an error the caller can handle.
Found while validating the bridge pin bump in leehack/llamadart#620. llamadart passes
GenerationParams.grammarandtopKthrough unchanged.