Skip to content

feat: run DecisionEngine on Web through the WebGPU bridge decision API - #620

Merged
leehack merged 20 commits into
mainfrom
feat/decision-web
Sep 23, 2026
Merged

leehack merged 20 commits into
mainfrom
feat/decision-web

Conversation

@leehack

@leehack leehack commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • WebGpuLlamaBackend implements BackendDecision through WebGpuDecisionHeads (lib/src/backends/webgpu/webgpu_decision.dart), which calls the llama-web-bridge decision API version 1: getDecisionCapabilities, loadDecisionHead, runDecision and freeDecisionHead. WebAutoBackend forwards these calls; LiteRT-LM Web reports unsupported.
  • Pin: the default bridge assets move from v0.1.44 to v0.1.47 (release 394986324, tag commit ee45e86, manifest 9c5e9008…), built from llama-web-bridge 64ba825 (llama-web-bridge#112). Web and native llama.cpp stay on v0.4.1; native pin v0.4.1-1 remains approved, as #630 set up.
  • Capability check: a bridge without all four methods, or reporting any apiVersion other than 1, is unsupported and names the assets needed. A head reporting another version is freed first.
  • URLs: headPath and configPath resolve against document.baseURI. The page fetches the config and passes its text to the bridge. URLs in errors drop user info, query and fragment.
  • Errors: bridge rejections map to LlamaStateException, LlamaContextException, LlamaModelException or LlamaInferenceException, as on native; malformed bridge responses throw LlamaDecisionException.
  • Handles: heads are numbered by the backend and never reused. Each head belongs to the bridge instance that loaded it; heads are forgotten on bridge dispose, on model load, and when the bridge reports a head lost.
  • Validation: native validateDecisionSequences now rejects more markers than tokens, with the bridge core's message.
  • Numbers: Web can't tell 30.0 from 30, so an integral double in a JSON-like state or question text reaches the model as 30 (native writes 30.0). The guide says so.
  • Docs: Dartdoc, doc/decision_engine.md (Web section and Web check), decision guide, support matrix, WebGPU bridge pages, the TTS guide's pin note, README, CHANGELOG Unreleased.

Part of #604.

Production-readiness scope

  • User-facing scope: with the default pin, DecisionEngine loads and runs on Web, on WebGPU or the bridge WASM CPU, in worker or main-thread mode.
  • Supported platforms/paths: native llama.cpp (unchanged); Web WebGPU bridge assets v0.1.47+ (decision API 1).
  • Unsupported or intentionally unavailable paths:
    • Older assets (for example WEBGPU_BRIDGE_ASSETS_TAG=v0.1.44): capabilitiesFor reports "Web decision models need llama-web-bridge assets v0.1.47+ with the decision API (apiVersion 1); the loaded bridge does not expose it." and DecisionEngine.load throws LlamaUnsupportedException.
    • Another decision API version throws LlamaUnsupportedException.
    • LiteRT-LM Web and native LiteRT-LM throw LlamaUnsupportedException.
  • Out of scope / follow-ups:
    • Repo decision-model-web-smoke scenario: #621.
    • Stable error codes on bridge rejections: llama-web-bridge#114.
    • Relative model URLs 404 in worker mode (predates this PR): llama-web-bridge#113.
    • Pre-existing bridge bugs found in the pin-bump check: llama-web-bridge#115, llama-web-bridge#116, llama-web-bridge#117.
    • laya-Q8_0.gguf on the bridge WASM CPU misses the probability and score tolerances on one fixture row (same top option); a bridge-side drift, documented in the guide and doc/decision_engine.md.
    • P3 test gaps D05 and D06 (below): not filed yet. Two fake-bridge tests would close them: one for a bridge missing only runDecision, and one for {apiVersion: 1} without supported.

Completeness checklist

  • Declared scope is fully implemented.
  • Unsupported combinations fail loudly with actionable, typed errors.
  • Dartdoc, README, website docs, support matrix and CHANGELOG are updated.
  • Regression coverage includes the happy path, old assets, API version skew, LiteRT-LM, error mapping, handle scoping, dispose and reload cascades, lost heads, and URL redaction.
  • Security/privacy: credentials, signed-URL queries and fragments stay out of messages, details and toString().
  • Follow-ups are tracked above; D05/D06 still need an issue or tests.

PR type guidance

  • Feature PR: docs, platform matrix, unsupported-path behavior, and happy and negative tests are covered above.

Test Plan

At head 587e3191c; CI ran on merge 18c898e, whose tree equals the head tree.

  • dart run tool/prepare_workspace.dart: exit 0, tree clean
  • dart format --output=none --set-exit-if-changed .: 0 changed
  • dart analyze: no issues
  • dart test -p vm -j 1 --exclude-tags local-only: CI Linux 3039 passed, 3 skipped; macOS 2964 / 78; Windows 2949 / 89. Local decision/WebGPU/web/tooling/service suites: 1092 passed, 4 skipped.
  • dart test -p chrome --exclude-tags local-only: CI 1370 passed. Local decision/WebGPU/web suites: 469 passed.
  • Other:
    • lib coverage 83.76%; docs build, which fails on broken links (CI Docs Build Check)
    • check_platform_boundaries OK; check_webgpu_bridge_tag --verify-manifest: 17 pins on v0.1.47
    • verify_release_docs_versions passes; --release-prep fails only on the pending companion bump from #630, identically on main 03bca335f
    • dart2wasm on the decision test files: 54 passed
    • exact-head mutation proof and real-model Web check (below)

Matrix Evidence

Matrix row Scope covered Platform / model / backend Result Evidence / notes
static-format-analyze format, analyzer, import boundaries macOS, Flutter 3.47.1; CI Analyze & Lint PASS
root-vm VM suite CI Linux, macOS, Windows PASS counts above
root-chrome browser suite, fake-bridge decision tests CI Chrome PASS 1370 passed
coverage-lib lib coverage CI Linux VM PASS 83.76%
docs-site site build, links CI PASS
release-doc-version-consistency macOS PASS default mode; --release-prep note above
webgpu-bridge-tag-consistency pins, manifest bytes macOS PASS release 394986324, manifest 9c5e9008…
web-production-artifact-smoke production build, mock and real worker/GGUF smokes CI Web Chat Contract, stories15M, bridge v0.1.47 PASS at 587e3191c
Web real-model decision check (local, 587e3191c) DecisionEngine in Chromium, 24 fixture rows, 24 typed reads per run Apple M4 Max, Chromium 151; laya-Q8_0.gguf and F16 + laya-head.safetensors; published v0.1.47; wasm64 PASS, except the known Q8_0 WASM CPU row worst logit/prob/score: Q8_0 WebGPU worker and main thread 0.1636/0.0436/0.0247; F16 WebGPU 0.0169/0.0046/0.0013; F16 WASM CPU 0.0149/0.0039/0.0028; Q8_0 WASM CPU 0.2326/0.0628/0.1224 (fails plain_text/urgency5, same top option). All equal the documented numbers.
Web decision lifecycle (local, 587e3191c) lost head, reload, bridge dispose; old assets same; v0.1.47 and v0.1.44; worker and main thread PASS 0 failures; v0.1.44 reports unsupported and load throws LlamaUnsupportedException
Web decision parity, other modes (local, 4b784051a) wasm32 core (Q8_0 WebGPU); main thread on F16 WebGPU and on WASM CPU same PASS (Q8_0 WASM CPU row as above) not rerun; no Web-path file changed since
web-mock-chat-smoke, web-bridge-smoke, web-real-model-smoke, web-speech-to-text-smoke, web-text-to-speech-smoke, gemma4-webgpu-mem64, webgpu-multimodal-regression (local, 4b784051a) chat app and bridge on v0.1.47 headless Chromium / Chrome, macOS; stories15M, Qwen3.5-0.8B (+ mmproj), Qwen3-ASR-0.6B, Qwen3-TTS-1.7B, Gemma 4 E2B PASS Multimodal gate script run directly on local models. Bridge A/B v0.1.44 vs v0.1.47: identical tokenize, greedy output, state round trip, embed, template, cancel and error texts. Since 4b784051a no Web-path or chat-app file changed; lib/ changed only in native llama.cpp files and the native pin.
decision-model-smoke native parity Apple M4 Max, laya-Q8_0.gguf, CPU and Metal PASS at fc8fa4637 native pin then v0.4.1 (now v0.4.1-1 from #630); native change here is one validation branch, pinned by a unit test
gemma4-litert-web N/A LiteRT-LM Web only gains the unsupported report, covered with fakes
structured-output-adversarial N/A classifier: no structured-output surface
high-risk-exact-head-independent-qa exact-head audit and evidence 587e3191c / 03bca335f PASS see below

Review Notes

  • Independent review status:
    • Implementation: a contract-and-skew review found 1 major and 9 minor issues, all fixed except the dispose-clear survivor below. The review of the #606 merge found a P1 (a conflict with that PR's head) and a P2 (integral doubles on Web); both fixed.
    • Pin bump: provenance, the v0.1.44→v0.1.47 bridge JS diff, chat-app Web regressions and a bridge A/B; it found the three pre-existing bridge bugs filed above.
    • Exact head: a fresh Claude Opus 5.5 workflow subagent (audit:620) with no part in writing the PR, blocking-only: ACCEPT, P1 0, P2 0, P3 1 (the D06 test gap below).
  • CI status / head SHA: 587e3191c165ab725535b35d5c1794b40cce62cc: 23 checks pass, 5 skipped, none failing.

High-risk regression review

  • Classification: high-risk (classify_high_risk_changes.dart: artifactConsumer, backendRuntime)
  • Implementation task: the implementation and fixer sessions
  • Independent blocking QA task: audit:620, operator-owned fresh Claude Opus 5.5 subagent (not Codex, not a human). high_risk_readiness.dart on the evidence: exit 2, unverifiedPrerequisites (the repository-local ceiling).
  • Exact head / current base: 587e3191c165ab725535b35d5c1794b40cce62cc / 03bca335fe499104a6067ec3515ac06208d048a1 (main, the merge base)
  • Production-branch deletion, bypass, or miswire proof: 34 of 38 single-guard mutants at the exact head are killed, each applied alone and restored to a clean tree. Killed: decision API version checks on capabilities and heads; the method-presence probe (all methods, and freeDecisionHead alone); bridge-identity checks in run and free; handle reuse; head forgetting on clear, model load and lost head; URL redaction and base-URL resolution; state, context and busy error mapping; config forwarding; malformed-head release; questionType/configJson wire names; WebAutoBackend routing and LiteRT-LM rejection; the native marker-count check; the four v0.1.47 provenance pins. Survivors:
    • B02 (no head clear on bridge dispose) behaves the same: a null active bridge and the bridge-identity check already reject stale heads. B03 (_activeBridge ignoring _usingBridge) differs only inside bridge activation.
    • D05 (probe skips runDecision) and D06 (a capability response without a boolean supported) are test gaps for inputs the published bridge does not produce. P3, non-blocking.
    • The audit's own run killed 15 of 18 (the same D06, B02, B03 survivors); implementation-time mutation killed 33 of 34 at fc8fa4637.
  • Affected-family real-model/artifact evidence: Laya laya-Q8_0.gguf and a local F16 conversion with laya-head.safetensors on the published v0.1.47 assets at the exact head (table above).
  • Explicit unavailable-family or other N/A evidence:
    • The official checkpoint was not run on Web. Loading through configPath ran with the local rl_agent_config.json (under <base href>), and a head without laya.config fails with the expected LlamaModelException.
    • LiteRT-LM Web is covered only through the unsupported path, with fakes.
    • Decision runs were measured only in headless Chromium on macOS (Apple M4 Max).
  • Known PR-caused P1 regressions: 0
  • Unresolved review threads: 0

Question and answer types, Python json.dumps parity, Laya 0.3.5 sequence
assembly and answer decoding, with a parity fixture from the official
checkpoint. Not exported yet; the native backend and DecisionEngine facade
follow in the next PR. Design: doc/decision_engine.md.
Adds the BackendDecision contract, the native path (safetensors reader,
ggml head graph with Windows ggml-base twins, private pooling-NONE encoder
context, worker messages, router forwarding), LlamaEngine hooks with
engine-owned head handles, and the public DecisionEngine facade. Web and
LiteRT-LM report unsupported. Includes a local-only E2E against the laya
0.3.5 fixture, the decision-model-smoke runner scenario, and docs.
Move the decision encoder context setup into applyDecisionContextParams with
unit tests, state that a call already sent to the backend finishes on the
unloaded model, and add local F16 backbone measurements.
…code

- Stream head tensors into the upload buffer; one head load peaks about
  100 MiB lower.
- Allocate head weights with ggml_backend_alloc_ctx_tensors instead of a
  hand-written layout, dropping six GgmlGraphApi entries and their Windows
  twins.
- Replace the fdlibm erf port with Abramowitz and Stegun 7.1.26 (within
  1.4e-7).
- Compute the head's last layer only for the CLS and marker rows.
- Type BackendDecisionSequence.questionType as DecisionQuestionType.
- Parse the head config once, with headLayers on DecisionHeadConfig.
- Return the encoder output as a view, reuse the service's device matching,
  and test the head's device choice.
- Tighten tests that could not fail and share the synthetic head and
  fixture helpers.
- Add ChoiceKey, ScoreKey and NoulKey: typed question handles read back with
  answerOf, including enum and arbitrary option values, checked against the
  request that produced the result. The string and JSON forms are unchanged.
- Accept structured instructions in the question constructors.
- Check empty question ids in DecisionRequest, tokenize each text once per
  call, clamp temperatures once, and reject non-string JSON keys.
- Docs: typed questions guide section, Experimental label, corrected accuracy
  and Laya-compatibility wording, and one home for the measurements.
- systemOneBatch answers the requests as they were when it started, so
  changing the list mid-call can no longer mislabel answers.
- DecisionEngine.load rejects a model unloaded or replaced during the
  capability probe instead of returning an engine for the old model.
- The real-model smoke fails when the requested GPU backend fell back to
  CPU, and checks that a config longer than the encoder was trained for is
  rejected.
- The guide says the head falls back to CPU when no device of the model's
  backend is available.
- Merge feat/decision-native (d9abfc3): typed ChoiceKey, ScoreKey and
  NoulKey, the Experimental label, and BackendDecisionSequence.questionType
  as DecisionQuestionType.
- Web passes questionType.index to the bridge; drop the int32 range check
  and its test, which only guarded the int field.
- Keep the marker-count check next to #606's validation, without the
  removed question type check.
- Test typed key reads and the question identity check through the Web
  backend; docs combine the Experimental label with Web support.
# Conflicts:
#	website/docs/guides/decision-models.md
# Conflicts:
#	lib/src/backends/webgpu/webgpu_backend.dart
Base automatically changed from feat/decision-native to main September 23, 2026 18:18
# Conflicts:
#	CHANGELOG.md
#	website/docs/changelog/recent-releases.md
Fill the Web check with the numbers measured on the published v0.1.47 assets, name the score tolerance Q8_0 on the WASM CPU also misses, scope the guide's F16/GPU claim to the fixture, and state that Web was checked only in headless Chromium on macOS.
# Conflicts:
#	doc/webgpu_bridge.md
#	website/docs/platforms/webgpu-bridge.md
# Conflicts:
#	CHANGELOG.md
#	website/docs/changelog/recent-releases.md
# Conflicts:
#	CHANGELOG.md
#	website/docs/changelog/recent-releases.md
@leehack
leehack marked this pull request as ready for review September 23, 2026 23:59
@leehack
leehack merged commit 838143e into main Sep 23, 2026
29 of 30 checks passed
@leehack
leehack deleted the feat/decision-web branch September 23, 2026 23:59
@github-actions

Copy link
Copy Markdown
Contributor

Chat app preview removed for leehack/llamadart-chat-pr-620.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant