Tracking issue for DecisionEngine: typed, non-autoregressive decisions with Laya-style models (ModernBERT GGUF backbone in llama.cpp + a safetensors decision head), following Laya's system_one and the Jev request/response shape.
Design: doc/decision_engine.md (added in the first PR).
Stack (each merged only after maintainer approval):
- Pure-Dart core: question/answer types,
json.dumps parity, sequence assembly, decoding, parity fixture from laya 0.3.5.
- Native backend, engine hooks,
DecisionEngine facade, export, local-only E2E, docs.
example/basic_app decision example.
example/laya_tetris Flutter example (base and Tetris-tuned heads).
- Head fine-tuning notebook and dataset tool.
- Web: decision module in
llama-web-bridge, asset publication, WebGpuLlamaBackend wiring.
Needs maintainer decisions before it happens: hosting the Tetris-tuned head; publishing new bridge assets.
Known limits: no Unicode normalization (the Hugging Face tokenizer applies NFC); one encoder pass per question; no cancellation; only the English checkpoint has parity evidence.
Tracking issue for
DecisionEngine: typed, non-autoregressive decisions with Laya-style models (ModernBERT GGUF backbone in llama.cpp + a safetensors decision head), following Laya'ssystem_oneand the Jev request/response shape.Design:
doc/decision_engine.md(added in the first PR).Stack (each merged only after maintainer approval):
json.dumpsparity, sequence assembly, decoding, parity fixture from laya 0.3.5.DecisionEnginefacade, export, local-only E2E, docs.example/basic_appdecision example.example/laya_tetrisFlutter example (base and Tetris-tuned heads).llama-web-bridge, asset publication,WebGpuLlamaBackendwiring.Needs maintainer decisions before it happens: hosting the Tetris-tuned head; publishing new bridge assets.
Known limits: no Unicode normalization (the Hugging Face tokenizer applies NFC); one encoder pass per question; no cancellation; only the English checkpoint has parity evidence.