Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
038dea6
feat: add the pure-Dart core for Laya decision models
leehack Sep 23, 2026
332bf37
feat: run Laya decision models on native llama.cpp with DecisionEngine
leehack Sep 23, 2026
0a42cdc
fix: test decision CPU context flags and correct unload-during-call docs
leehack Sep 23, 2026
5c51654
test: start the decision context fixture from a pooled struct
leehack Sep 23, 2026
fd46e6d
feat: run DecisionEngine on Web through the WebGPU bridge decision API
leehack Sep 23, 2026
fcdcd42
refactor: stream the decision head load and trim hand-written native …
leehack Sep 23, 2026
d9abfc3
feat: add typed decision keys and mark DecisionEngine experimental
leehack Sep 23, 2026
0848acd
fix: snapshot decision batches and reject a model swapped in during load
leehack Sep 23, 2026
5ea1e68
merge: bring #606's typed decision keys into the Web decision backend
leehack Sep 23, 2026
b379dcd
Merge branch 'feat/decision-native' into feat/decision-web
leehack Sep 23, 2026
68210de
docs: note that Web writes integral doubles in decision JSON as integers
leehack Sep 23, 2026
d4f439b
Merge remote-tracking branch 'origin/main' into feat/decision-native
leehack Sep 23, 2026
fc8fa46
Merge branch 'feat/decision-native' into feat/decision-web
leehack Sep 23, 2026
ffd5ac3
Merge main (#606 squash) into feat/decision-web
leehack Sep 23, 2026
2e35a19
Merge remote-tracking branch 'origin/main' into feat/decision-web
leehack Sep 23, 2026
4b78405
chore: adopt Web bridge v0.1.47 with the decision API
leehack Sep 23, 2026
2548420
docs: record the v0.1.47 Web decision check
leehack Sep 23, 2026
d0d422a
Merge remote-tracking branch 'origin/main' into feat/decision-web
leehack Sep 23, 2026
541a577
Merge remote-tracking branch 'origin/main' into feat/decision-web
leehack Sep 23, 2026
587e319
Merge remote-tracking branch 'origin/main' into feat/decision-web
leehack Sep 23, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,14 @@
- Add a notebook in `example/laya_tetris/training/` that fine-tunes a Laya
decision head for the Tetris example and exports it for `DecisionEngine`
([#604](https://github.com/leehack/llamadart/issues/604)).
- Run `DecisionEngine` on WebGPU through the bridge decision API
(apiVersion 1), which bridge assets `v0.1.47+` include
([#604](https://github.com/leehack/llamadart/issues/604)).
* Aligned the default WebGPU bridge assets to `v0.1.47` for the decision API,
retaining Web/native llama.cpp
`v0.4.1@b29c606e28a01b1bc8c1351026a0fa6e616bf6c4` parity and Web
`@litert-lm/core@0.15.0`. Immutable Web asset manifest:
`9c5e9008d187690e283f37b2b892396da03e3c71bf7c3434d7c87ceca8e6a4bc`.
- Extend the GGUF speech-to-text validation pack with four synthetic edge
fixtures built in-process, so no extra audio is stored: generated digital
silence, plus a truncated RIFF, a stereo 44.1 kHz re-encode and a 33-second
Expand Down
6 changes: 4 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,9 @@ models through LiteRT-LM.
- Experimental Laya-style decision models on native llama.cpp through
`DecisionEngine`: typed choice, score, and yes/no answers from a ModernBERT
encoder GGUF and a safetensors head, one encoder pass per question; validated
on macOS (Metal, CPU), other native platforms untested.
on macOS (Metal, CPU), other native platforms untested. Web runs it through
WebGPU bridge assets `v0.1.47+`, which the default Web pin includes; checked
only in headless Chromium on macOS.

Unsupported runtime/option combinations are rejected explicitly instead of
silently degrading. Check the support matrix before relying on a capability for
Expand Down Expand Up @@ -158,7 +160,7 @@ Current default runtime pins:
| --- | --- |
| Native llama.cpp / GGUF | `leehack/llamadart-native@v0.4.1-1` |
| Native LiteRT-LM / `.litertlm` | `leehack/litert-lm-native@v0.17.0-6` |
| Web llama.cpp / GGUF | `leehack/llama-web-bridge-assets@v0.1.44` |
| Web llama.cpp / GGUF | `leehack/llama-web-bridge-assets@v0.1.47` |
| Web LiteRT-LM / `.litertlm` | `@litert-lm/core@0.15.0` |

Native overrides accept stable `vMAJOR.MINOR.PATCH` releases and preserve
Expand Down
118 changes: 111 additions & 7 deletions doc/decision_engine.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,8 @@ DecisionEngine (core, pure Dart)
BackendDecision (backend.dart, web-safe value types)
NativeAutoBackend -> NativeLlamaBackend -> worker isolate -> LlamaCppService
private encoder llama_context + safetensors head + ggml head graph
WebAutoBackend -> WebGpuLlamaBackend -> WebGpuDecisionHeads -> llama-web-bridge
bridge decision API 1: private encoder context + head in the WASM core
raw marker logits + raw act logits -> decoder (core) -> DecisionResult
```

Expand Down Expand Up @@ -89,16 +91,14 @@ abstract class BackendDecision {
- `BackendDecisionOutput`: per sequence, raw marker `logits` and raw
`actLogits` (`Float32List`).

`NativeAutoBackend` implements and forwards it; the LiteRT-LM delegate reports
unsupported. `WebAutoBackend` does not implement it until the bridge ships the
module, so the engine hook reports unsupported and `DecisionEngine.load` throws
`LlamaUnsupportedException`.
`NativeAutoBackend` and `WebAutoBackend` implement and forward it; their
LiteRT-LM delegates report unsupported.

### Engine hooks (`lib/src/core/engine/engine.dart`)

Plain public methods documented as low-level integration hooks, like the TTS
trio. The capabilities hook checks `is! BackendDecision` before readiness, so
Web reports a stable reason without a model.
a backend without the contract reports a stable reason without a model.

Backend handles are not unique over an engine's life: the worker numbers
handles from 1, and a new worker starts after `LlamaEngine.dispose` followed by
Expand Down Expand Up @@ -184,6 +184,12 @@ static helper that unit tests cover:

The run path rejects a sequence longer than `llama_n_ubatch` before
`llama_encode`, whose `GGML_ASSERT` would abort the process.
`validateDecisionSequences` checks every sequence before the first encoder
pass: 1 to `n_ubatch` tokens inside the vocabulary, and 1 to token-count
markers inside the sequence. The bridge core runs the same checks with the same
messages, plus a question type check that `DecisionQuestionType` always passes,
so both runtimes reject the same input; the marker-count bound comes from the
bridge, whose head graph sizes its buffers by marker count.

Windows: `llama.dll` exports no `ggml_*` graph symbols; they live in
`ggml-base.dll` (ops, graph, sched, buffers) and `ggml.dll` (registry). The head
Expand All @@ -195,6 +201,63 @@ and the generated bindings on other platforms. The bindings leave out
on their default asset, `package:llamadart/llamadart`, there too. Generated
bindings are not edited.

### Web (`lib/src/backends/webgpu/`)

`WebGpuLlamaBackend` implements `BackendDecision` through `WebGpuDecisionHeads`
(`webgpu_decision.dart`), which calls the llama-web-bridge decision API:
`getDecisionCapabilities`, `loadDecisionHead`, `runDecision` and
`freeDecisionHead` (bridge `docs/api.md`, "Decision heads"). The bridge runs the
head on WebGPU when the model loaded with GPU layers and on the CPU otherwise,
and reports which as `deviceName`.

- Capability probe: a bridge object without all four methods reports
unsupported with "Web decision models need llama-web-bridge assets v0.1.47+
with the decision API (apiVersion 1)", from the
`webGpuDecisionBridgeRequirement` constant. A capability or head response
with an `apiVersion` other than 1 is unsupported too, and such a head is
freed first. Bridge assets `v0.1.47+`, the default pin among them, have the
API.
- Paths are URLs, resolved in Dart against `document.baseURI` before any
fetch, so a page's `<base href>` applies to both in both bridge modes. The
bridge fetches `headPath`. It takes the config only as text, so `configPath`
is fetched in the page with `fetch`, before the head, and passed as
`configJson`; with both a missing config and a bad head, Web reports the
config where native reports the head. A failed fetch or an HTTP error is
`LlamaModelException` "Cannot read the decision head config at <url>." with
the status or error in `details`. URLs in error messages and details drop
user info, query and fragment, including URLs that browser and bridge
errors quote.
- Handles: the backend numbers heads itself, never reusing a number, and maps
each to the bridge instance and bridge handle that loaded it. `modelFree` and
`dispose` dispose the bridge, and a model load on the same bridge frees every
bridge head, so the backend forgets all heads at each. A run with a forgotten
head, or with a head whose bridge is no longer active, throws
`LlamaStateException` without calling the bridge; a free does nothing.
- Errors: the bridge rejects with plain `Error`s that carry the core's message
and no status code, so the mapping reads the message after stripping the
bridge's `Failed to load decision head: ` or `Decision run failed: ` prefix.
"Load the decision head again" (a freed head, or one lost to a worker
failure, which also forgets the head), "No model loaded", "Bridge has been
disposed", "was cancelled" and "during active generation" map to
`LlamaStateException`, from the capability probe too; "decision encoder
context" to `LlamaContextException`; anything else to unsupported for the
probe, `LlamaModelException` for a load (head URL in `details`),
`LlamaInferenceException` for a run and `LlamaStateException` for a free.
Load errors that ask bridge callers to pass `configJson` name `configPath`
or the config URL instead. Without an active bridge, the probe reports
unsupported and a load throws `LlamaStateException`, as native does for an
unloaded model. A malformed head description or output is
`LlamaDecisionException`, like native's unexpected worker responses.
- Numbers: Web numbers cannot tell `30.0` from `30`, so `pythonJsonDumps`
writes an integral double in a non-`String` state, instructions, criteria or
levels as an int (`30` where Python writes `30.0`). The tokens then differ
from native and Laya; the guide's Web section tells users to pass such
values as `String`s when parity matters.
- The bridge serializes decision calls with its other operations and cannot
cancel a run. When its worker fails during a run, it reloads the model on the
main thread and rejects the run; the engine keeps its model, and the
`DecisionEngine` must be loaded again.

## Parity rules

Sequence (`build_sequence`, `max_len` 512, `head_max_len` 192):
Expand Down Expand Up @@ -261,7 +324,8 @@ JSON-like (null, bool, num, String, List, Map with String keys).
| Linux | CPU, Vulkan, CUDA | expected, untested |
| Windows | CPU, Vulkan, CUDA | expected through the `ggml-base` twins, untested |
| Native LiteRT-LM | - | `LlamaUnsupportedException` |
| Web (WebGPU bridge) | - | `LlamaUnsupportedException` until the bridge module ships |
| Web (WebGPU bridge) | WebGPU or CPU (WASM) | bridge assets `v0.1.47+` (apiVersion 1), the default pin among them; older assets report `LlamaUnsupportedException`. CI uses a fake bridge; checked locally with a real model ([Web check](#web-check)) |
| LiteRT-LM Web | - | `LlamaUnsupportedException` |

Real-model evidence is macOS only. The CPU head unit tests carry no
`local-only` tag, so CI's Linux VM job and its macOS and Windows native test
Expand Down Expand Up @@ -317,6 +381,32 @@ On Metal, disposing the engine with a head still loaded exits cleanly; skipping
the head frees in `freeModel` and `dispose` makes the same exit abort in
`ggml_metal_rsets_free`.

### Web check

Local only, not in CI: `DecisionEngine` through `LlamaEngine(LlamaBackend())`
in Playwright's headless Chromium on the same machine, with the pinned bridge
assets (bridge source `64ba8250`), the 24 fixture rows,
`laya-head.safetensors` and the tolerances of `decision-model-smoke`. Token ids
and markers matched on every row.

| Backbone | Bridge runtime | Head device | Logit diff | Probability diff | Score diff |
| --- | --- | --- | --- | --- | --- |
| `laya-Q8_0.gguf` | WebGPU; worker and main thread on wasm64 and wasm32 | WebGPU | 0.1636 | 0.0436 | 0.0247 |
| F16 (local conversion) | WebGPU; worker and main thread on wasm64 | WebGPU | 0.0169 | 0.0046 | 0.0013 |
| F16 (local conversion) | WASM CPU; worker and main thread on wasm64 | CPU | 0.0149 | 0.0039 | 0.0028 |
| `laya-Q8_0.gguf` | WASM CPU; worker and main thread on wasm64 | CPU | 0.2326 | 0.0628 | 0.1224 |

Q8_0 on the WASM CPU misses the probability and score tolerances on one row,
`plain_text/urgency5`, with the same top option. The bridge's own smoke, which
calls the bridge directly, gets the same worst logit difference on wasm32 and
wasm64 in both bridge modes, so the drift comes from the bridge's WASM CPU
Q8_0 path rather than llamadart.
Typed key reads with the question identity check, sequence validation
messages, error mapping, URL redaction, `<base href>` resolution, and heads
freed or bridges disposed behind the engine's back were checked against the
same assets. The previous pin, `v0.1.44`, which lacks the API, reported
unsupported with the actionable reason in both bridge modes.

## Known limits

User-facing limits are listed under
Expand Down Expand Up @@ -346,7 +436,21 @@ guide.
order; the service's load-time check helpers, head device choice, sequence
validation, and run order through a substituted encoder; worker,
backend-client and router routing with fakes; engine hooks and facade with a
fake backend; Web unsupported path under `@TestOn('browser')`.
fake backend.
- Unit (Chrome): `WebGpuDecisionHeads` against a fake bridge
(`test/support/fake_webgpu_decision_bridge.dart`): the capability probe for
old assets, API version skew, bridge reasons and state rejections; head
loading with page-fetched configs, unreadable ones, URLs resolved against a
`<base href>`, and credentials and queries kept out of errors; error mapping,
including the `configJson` wording; handle scoping to the loading bridge;
malformed responses. `WebGpuLlamaBackend` without an active bridge, and
forgetting heads on `modelFree`, a same-bridge model load and `dispose`;
`WebAutoBackend` forwarding and LiteRT-LM Web reporting unsupported; the
engine hook without a model.
- Integration (Chrome, fake bridge): `DecisionEngine` through `LlamaEngine`,
`WebAutoBackend` and `WebGpuLlamaBackend`: answers, typed key reads with the
question identity check, sequence layout, page-fetched config, old assets,
API version skew, a cancelled capability probe, and a model unload.
- Integration (VM, CI's `stories15M.gguf`): a llama-architecture model is
reported unsupported and `DecisionEngine.load` fails before reading the head.
- Local-only E2E `test/e2e/backends/decision_engine_e2e_test.dart`: real GGUF
Expand Down
33 changes: 20 additions & 13 deletions doc/webgpu_bridge.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,21 +19,22 @@ pipelines.
`https://cdn.jsdelivr.net/gh/leehack/llama-web-bridge-assets@<tag>/llama_webgpu_bridge.js`
2. Local fallback: `./webgpu_bridge/llama_webgpu_bridge.js`

Default pinned tag in the example is `v0.1.44`.
Default pinned tag in the example is `v0.1.47`.

That release embeds llama.cpp `v0.4.1`, matching the `hook/build.dart` native pin
(`v0.4.1-1`, both built from upstream llama.cpp `v0.4.1@b29c606e28a01b1bc8c1351026a0fa6e616bf6c4`)
even though the bridge asset tag `v0.1.44` differs from the native runtime tag
`v0.4.1-1`. Provenance for this immutable consumer artifact: release `389783936`,
tag commit `fdafd9f8cbdb9bf99c359536595eff9a23095379`, bridge source
`89178be67c3c84300bc1b129182bd5bc5a8e21fc`, and manifest SHA-256
`8d61f453753ac7a7d839ac12318b70986a814748d86029993118c19454293aa9`. The bridge
even though the bridge asset tag `v0.1.47` differs from the native runtime tag
`v0.4.1-1`. Provenance for this immutable consumer artifact: release `394986324`,
tag commit `ee45e864641648a99411128f9bb82b7897fad221`, bridge source
`64ba8250871bf2472cc2064c6a00fff050783f02`, and manifest SHA-256
`9c5e9008d187690e283f37b2b892396da03e3c71bf7c3434d7c87ceca8e6a4bc`. The bridge
assets were qualified against native `v0.4.1`; `v0.4.1-1` rebuilds it from the
same upstream commit. It retains the Qwen3-ASR typed speech-to-text contract
introduced in `v0.1.30` and provisions the explicit 1 MiB Wasm stack needed for
memory64 context construction in direct and worker modes. The chat bootstrap opts
`SpeechToTextEngine` into that contract from the immutable tag; older or custom
assets remain disabled unless the host explicitly sets
same upstream commit. The assets add the decision API (apiVersion 1), retain the
Qwen3-ASR typed speech-to-text contract introduced in `v0.1.30`, and provision
the explicit 1 MiB Wasm stack needed for memory64 context construction in
direct and worker modes. The chat bootstrap opts `SpeechToTextEngine` into that
contract from the immutable tag; older or custom assets remain disabled unless
the host explicitly sets
`window.__llamadartBridgeSpeechToTextSupported = true` after equivalent
validation.

Expand All @@ -48,7 +49,7 @@ model bytes.
To vendor pinned assets into local app web files:

```bash
WEBGPU_BRIDGE_ASSETS_TAG=v0.1.44 ./scripts/fetch_webgpu_bridge_assets.sh
WEBGPU_BRIDGE_ASSETS_TAG=v0.1.47 ./scripts/fetch_webgpu_bridge_assets.sh
```

Optional compatibility env vars:
Expand Down Expand Up @@ -129,7 +130,7 @@ You can override CDN source/version before the bridge loader runs:
```html
<script>
window.__llamadartBridgeAssetsRepo = 'leehack/llama-web-bridge-assets';
window.__llamadartBridgeAssetsTag = 'v0.1.44';
window.__llamadartBridgeAssetsTag = 'v0.1.47';
</script>
```

Expand Down Expand Up @@ -179,6 +180,10 @@ window.LlamaWebGpuBridge = class LlamaWebGpuBridge {
- `cancel()`
- `dispose()`
- `applyChatTemplate(messages, addAssistant, customTemplate)`
- `getDecisionCapabilities()`
- `loadDecisionHead(url, { configJson })`
- `runDecision(handle, sequences)`
- `freeDecisionHead(handle)`
- `isGpuActive()`
- `getBackendName()`

Expand All @@ -187,6 +192,8 @@ window.LlamaWebGpuBridge = class LlamaWebGpuBridge {
- Web backend remains GGUF URL-based (`modelLoadFromUrl`).
- If bridge activation fails, model loading fails (no alternate web backend).
- Embeddings on web require bridge assets with embedding APIs (`v0.1.7+`).
- `DecisionEngine` on web requires bridge assets with the decision API
(apiVersion 1, `v0.1.47+`).
- State persistence on web requires bridge assets with state APIs (`v0.1.15+`);
paths are bridge WASMFS virtual paths and are not durable across page reloads.
Durable browser storage currently requires app-level export/import outside the
Expand Down
2 changes: 1 addition & 1 deletion example/chat_app/web/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -211,7 +211,7 @@
typeof configuredRepo === 'string' && configuredRepo.length > 0
? configuredRepo
: 'leehack/llama-web-bridge-assets';
const defaultBridgeAssetsTag = 'v0.1.44';
const defaultBridgeAssetsTag = 'v0.1.47';
const defaultBridgeLlamaCppTag = 'v0.4.1';
const bridgeAssetsTag =
typeof configuredTag === 'string' && configuredTag.length > 0
Expand Down
13 changes: 11 additions & 2 deletions lib/src/backends/llama_cpp/llama_cpp_service.dart
Original file line number Diff line number Diff line change
Expand Up @@ -7859,8 +7859,10 @@ class LlamaCppService {
/// Checks [sequences] against a decision head's limits.
///
/// Each sequence needs 1 to [tokenLimit] tokens, each in `[0, vocabSize)`,
/// at least one marker, and every marker a position in its tokens. Throws
/// [LlamaInferenceException] naming the first sequence that fails.
/// 1 to token-count markers, and every marker a position in its tokens.
/// Throws [LlamaInferenceException] naming the first sequence that fails.
/// The llama-web-bridge decision core runs these checks with the same
/// messages, so both runtimes reject the same sequences.
static void validateDecisionSequences(
List<BackendDecisionSequence> sequences, {
required int tokenLimit,
Expand Down Expand Up @@ -7888,6 +7890,13 @@ class LlamaCppService {
'Decision sequence $i has no option markers.',
);
}
if (sequence.markers.length > tokens.length) {
throw LlamaInferenceException(
'Decision sequence $i has ${sequence.markers.length} markers for '
'its ${tokens.length} tokens; a sequence holds at most one marker '
'per token.',
);
}
for (final marker in sequence.markers) {
if (marker < 0 || marker >= tokens.length) {
throw LlamaInferenceException(
Expand Down
Loading
Loading