Repository navigation
feat: run DecisionEngine on Web through the WebGPU bridge decision API - #620
Merged
Merged
Conversation
Question and answer types, Python json.dumps parity, Laya 0.3.5 sequence assembly and answer decoding, with a parity fixture from the official checkpoint. Not exported yet; the native backend and DecisionEngine facade follow in the next PR. Design: doc/decision_engine.md.
Adds the BackendDecision contract, the native path (safetensors reader, ggml head graph with Windows ggml-base twins, private pooling-NONE encoder context, worker messages, router forwarding), LlamaEngine hooks with engine-owned head handles, and the public DecisionEngine facade. Web and LiteRT-LM report unsupported. Includes a local-only E2E against the laya 0.3.5 fixture, the decision-model-smoke runner scenario, and docs.
Move the decision encoder context setup into applyDecisionContextParams with unit tests, state that a call already sent to the backend finishes on the unloaded model, and add local F16 backbone measurements.
This was referenced Sep 23, 2026
…code - Stream head tensors into the upload buffer; one head load peaks about 100 MiB lower. - Allocate head weights with ggml_backend_alloc_ctx_tensors instead of a hand-written layout, dropping six GgmlGraphApi entries and their Windows twins. - Replace the fdlibm erf port with Abramowitz and Stegun 7.1.26 (within 1.4e-7). - Compute the head's last layer only for the CLS and marker rows. - Type BackendDecisionSequence.questionType as DecisionQuestionType. - Parse the head config once, with headLayers on DecisionHeadConfig. - Return the encoder output as a view, reuse the service's device matching, and test the head's device choice. - Tighten tests that could not fail and share the synthetic head and fixture helpers.
- Add ChoiceKey, ScoreKey and NoulKey: typed question handles read back with answerOf, including enum and arbitrary option values, checked against the request that produced the result. The string and JSON forms are unchanged. - Accept structured instructions in the question constructors. - Check empty question ids in DecisionRequest, tokenize each text once per call, clamp temperatures once, and reject non-string JSON keys. - Docs: typed questions guide section, Experimental label, corrected accuracy and Laya-compatibility wording, and one home for the measurements.
- systemOneBatch answers the requests as they were when it started, so changing the list mid-call can no longer mislabel answers. - DecisionEngine.load rejects a model unloaded or replaced during the capability probe instead of returning an engine for the old model. - The real-model smoke fails when the requested GPU backend fell back to CPU, and checks that a config longer than the encoder was trained for is rejected. - The guide says the head falls back to CPU when no device of the model's backend is available.
- Merge feat/decision-native (d9abfc3): typed ChoiceKey, ScoreKey and NoulKey, the Experimental label, and BackendDecisionSequence.questionType as DecisionQuestionType. - Web passes questionType.index to the bridge; drop the int32 range check and its test, which only guarded the int field. - Keep the marker-count check next to #606's validation, without the removed question type check. - Test typed key reads and the question identity check through the Web backend; docs combine the Experimental label with Web support.
# Conflicts: # website/docs/guides/decision-models.md
# Conflicts: # lib/src/backends/webgpu/webgpu_backend.dart
# Conflicts: # CHANGELOG.md # website/docs/changelog/recent-releases.md
Fill the Web check with the numbers measured on the published v0.1.47 assets, name the score tolerance Q8_0 on the WASM CPU also misses, scope the guide's F16/GPU claim to the fixture, and state that Web was checked only in headless Chromium on macOS.
This was referenced Sep 23, 2026
# Conflicts: # doc/webgpu_bridge.md # website/docs/platforms/webgpu-bridge.md
# Conflicts: # CHANGELOG.md # website/docs/changelog/recent-releases.md
# Conflicts: # CHANGELOG.md # website/docs/changelog/recent-releases.md
Contributor
|
Chat app preview removed for |
This was referenced Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
WebGpuLlamaBackendimplementsBackendDecisionthroughWebGpuDecisionHeads(lib/src/backends/webgpu/webgpu_decision.dart), which calls the llama-web-bridge decision API version 1:getDecisionCapabilities,loadDecisionHead,runDecisionandfreeDecisionHead.WebAutoBackendforwards these calls; LiteRT-LM Web reports unsupported.v0.1.44tov0.1.47(release 394986324, tag commitee45e86, manifest9c5e9008…), built from llama-web-bridge64ba825(llama-web-bridge#112). Web and native llama.cpp stay onv0.4.1; native pinv0.4.1-1remains approved, as #630 set up.apiVersionother than 1, is unsupported and names the assets needed. A head reporting another version is freed first.headPathandconfigPathresolve againstdocument.baseURI. The page fetches the config and passes its text to the bridge. URLs in errors drop user info, query and fragment.LlamaStateException,LlamaContextException,LlamaModelExceptionorLlamaInferenceException, as on native; malformed bridge responses throwLlamaDecisionException.validateDecisionSequencesnow rejects more markers than tokens, with the bridge core's message.30.0from30, so an integral double in a JSON-like state or question text reaches the model as30(native writes30.0). The guide says so.doc/decision_engine.md(Web section and Web check), decision guide, support matrix, WebGPU bridge pages, the TTS guide's pin note, README, CHANGELOGUnreleased.Part of #604.
Production-readiness scope
DecisionEngineloads and runs on Web, on WebGPU or the bridge WASM CPU, in worker or main-thread mode.v0.1.47+(decision API 1).WEBGPU_BRIDGE_ASSETS_TAG=v0.1.44):capabilitiesForreports "Web decision models need llama-web-bridge assets v0.1.47+ with the decision API (apiVersion 1); the loaded bridge does not expose it." andDecisionEngine.loadthrowsLlamaUnsupportedException.LlamaUnsupportedException.LlamaUnsupportedException.decision-model-web-smokescenario: #621.laya-Q8_0.ggufon the bridge WASM CPU misses the probability and score tolerances on one fixture row (same top option); a bridge-side drift, documented in the guide anddoc/decision_engine.md.runDecision, and one for{apiVersion: 1}withoutsupported.Completeness checklist
toString().PR type guidance
Test Plan
At head
587e3191c; CI ran on merge18c898e, whose tree equals the head tree.dart run tool/prepare_workspace.dart: exit 0, tree cleandart format --output=none --set-exit-if-changed .: 0 changeddart analyze: no issuesdart test -p vm -j 1 --exclude-tags local-only: CI Linux 3039 passed, 3 skipped; macOS 2964 / 78; Windows 2949 / 89. Local decision/WebGPU/web/tooling/service suites: 1092 passed, 4 skipped.dart test -p chrome --exclude-tags local-only: CI 1370 passed. Local decision/WebGPU/web suites: 469 passed.check_platform_boundariesOK;check_webgpu_bridge_tag --verify-manifest: 17 pins onv0.1.47verify_release_docs_versionspasses;--release-prepfails only on the pending companion bump from #630, identically onmain03bca335fMatrix Evidence
static-format-analyzeroot-vmroot-chromecoverage-libdocs-siterelease-doc-version-consistency--release-prepnote abovewebgpu-bridge-tag-consistency9c5e9008…web-production-artifact-smokev0.1.47587e3191c587e3191c)DecisionEnginein Chromium, 24 fixture rows, 24 typed reads per runlaya-Q8_0.ggufand F16 +laya-head.safetensors; publishedv0.1.47; wasm64plain_text/urgency5, same top option). All equal the documented numbers.587e3191c)v0.1.47andv0.1.44; worker and main threadv0.1.44reports unsupported and load throwsLlamaUnsupportedException4b784051a)web-mock-chat-smoke,web-bridge-smoke,web-real-model-smoke,web-speech-to-text-smoke,web-text-to-speech-smoke,gemma4-webgpu-mem64,webgpu-multimodal-regression(local,4b784051a)v0.1.47v0.1.44vsv0.1.47: identical tokenize, greedy output, state round trip, embed, template, cancel and error texts. Since4b784051ano Web-path or chat-app file changed;lib/changed only in native llama.cpp files and the native pin.decision-model-smokelaya-Q8_0.gguf, CPU and Metalfc8fa4637v0.4.1(nowv0.4.1-1from #630); native change here is one validation branch, pinned by a unit testgemma4-litert-webstructured-output-adversarialhigh-risk-exact-head-independent-qa587e3191c/03bca335fReview Notes
v0.1.44→v0.1.47bridge JS diff, chat-app Web regressions and a bridge A/B; it found the three pre-existing bridge bugs filed above.audit:620) with no part in writing the PR, blocking-only: ACCEPT, P1 0, P2 0, P3 1 (the D06 test gap below).587e3191c165ab725535b35d5c1794b40cce62cc: 23 checks pass, 5 skipped, none failing.High-risk regression review
classify_high_risk_changes.dart:artifactConsumer,backendRuntime)audit:620, operator-owned fresh Claude Opus 5.5 subagent (not Codex, not a human).high_risk_readiness.darton the evidence: exit 2,unverifiedPrerequisites(the repository-local ceiling).587e3191c165ab725535b35d5c1794b40cce62cc/03bca335fe499104a6067ec3515ac06208d048a1(main, the merge base)freeDecisionHeadalone); bridge-identity checks in run and free; handle reuse; head forgetting on clear, model load and lost head; URL redaction and base-URL resolution; state, context and busy error mapping; config forwarding; malformed-head release;questionType/configJsonwire names;WebAutoBackendrouting and LiteRT-LM rejection; the native marker-count check; the fourv0.1.47provenance pins. Survivors:_activeBridgeignoring_usingBridge) differs only inside bridge activation.runDecision) and D06 (a capability response without a booleansupported) are test gaps for inputs the published bridge does not produce. P3, non-blocking.fc8fa4637.laya-Q8_0.ggufand a local F16 conversion withlaya-head.safetensorson the publishedv0.1.47assets at the exact head (table above).configPathran with the localrl_agent_config.json(under<base href>), and a head withoutlaya.configfails with the expectedLlamaModelException.