Skip to content

Replace wordseq detector with Shannon-bit binary CNN ensemble - #13

Closed
staging-devin-ai-integration[bot] wants to merge 1 commit into
mainfrom
devin/1787528731-shannon-bnn
Closed

staging-devin-ai-integration[bot] wants to merge 1 commit into
mainfrom
devin/1787528731-shannon-bnn

Conversation

@staging-devin-ai-integration

Copy link
Copy Markdown
Contributor

Summary

Replaces the wordseq/MSQ1 student model and its tokenizer runtime with a fully-binary CNN detector over Shannon-coded raw bytes:

  • Input: the raw 2048-byte Magika begin/end window (no tokenizer) is entropy-coded with a canonical Shannon code learned from training byte frequencies (add-one smoothing, 12-bit max length, (length, byte) canonical order) into a fixed 16,384-bit stream plus a codeword-boundary bitplane.
  • Model: an ensemble of five binary CNNs — binary weights and activations, XNOR + popcount convolutions (z' = 2*matches - fan_in), fixed-point basis combine (FIX=65536), integer thresholds (folded BatchNorm), OR-pooling, segmented popcount heads, i16 integer classifier with per-class float scale/bias. All heavy compute is bitwise/integer CPU work on packed u64 planes.
  • Artifact: assets/bnn/source-bnn2.bin (packed "BBN2", 1,165,464 bytes) replaces assets/magika/source-student-q4.bin; the old runtime (tokenizer.rs, runtime.rs, layers.rs, embedded.rs, reader.rs, activation.rs) is removed along with the fearless_simd dependency.
  • Accuracy (held-out test split, 27,465 files, exact integer export semantics):
accuracy macro recall
shipped wordseq model 0.934972 0.934863
this PR (BNN ensemble) 0.941635 0.943226
  • Rust inference is bit-exact against the Python integer reference (python_parity test with embedded fixtures).
  • Behavioral compatibility preserved: empty/whitespace/short-input handling, all 48 language fixtures, YAML sequence-of-mappings, Markdown list, and bare dash-list YAML/Markdown uncertainty tests pass.
  • Tradeoff: inference is ~43 ms/window vs ~4.5 ms for the old model (five-model ensemble); recorded in BENCHMARKS.md.
  • Training/eval pipeline added under scripts/bnn/ with scripts/bnn/TRAINING.md.

Verification

  • cargo fmt --check
  • cargo clippy --all-targets --all-features -- -D warnings
  • cargo test and cargo test --release (18 tests + 10 doctests pass)
  • BNN_FIXTURES=.../parity_fixtures2.json cargo test --release python_parity (bit-exact Rust/Python parity)
  • cargo bench --bench detect (benchmark precondition asserts the Rust demo classifies as Rust)
  • python eval_export2.py test — exact integer-semantics evaluation of the exported artifact on the held-out test split (numbers above)

Link to Devin session: https://dioxus.staging.devinenterprise.com/sessions/12e41f6dc16a459e90eae45fc48f336a
Open in Devin Desktop: https://dioxus.staging.devinenterprise.com/desktop/session/12e41f6dc16a459e90eae45fc48f336a?variant=devin-insiders
Requested by: @ealmloff

Co-Authored-By: Staging-Devin AI <166158716+staging-devin-ai-integration[bot]@users.noreply.github.com>
@staging-devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR that start with 'DevinAI' or '@devin'.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@ealmloff

Copy link
Copy Markdown
Member

Compare size of model

@ealmloff ealmloff closed this Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant