An interactive explainer for DataForge 2026 — Pathway track ("Explain the Frontier").
Concept: Synaptic Plasticity as Short-Term Memory (with Linear Attention as the closely-related second topic; both serve one central claim.)
Live artifact: https://puregenius369.github.io/the-synapse-is-the-cache/ — opens without sign-in, no build step, no account. Source: this repository.
Because BDH's attention applies no softmax, it is a Hebbian synaptic memory: one fixed-size matrix
σ, written once per token by an outer product and read once per token by a dot product, computing exactly what all-pairs attention computes — but with a bounded capacity, so it forgets through interference unless its activations stay sparse.
The claim has two halves, and the artifact ships a control that falsifies each one.
| Half of the claim | Control that would falsify it | What actually happens |
|---|---|---|
| "…computing exactly what all-pairs attention computes" | Lab 1 → softmax toggle. If the equality survived a nonlinearity, the stated reason for it would be wrong. | Relative difference jumps from ~4×10⁻¹⁶ to ~2.4×10¹. The claim's mechanism is confirmed by its failure mode. |
| "…forgets through interference unless activations stay sparse" | Lab 2 → raise K (associations stored) at fixed n. If retrieval stayed perfect, σ would be a lossless cache. |
Mean retrieval fidelity falls monotonically; lowering key density restores it. |
If neither control could ever make the sentence false, the claim would be too vague to teach.
Audience. Someone who has implemented or carefully read scaled dot-product attention — a strong final-year undergraduate, a master's student, or a working ML engineer meeting post-Transformer architectures for the first time.
Prerequisites. Matrix multiplication; what an outer product is; what a KV cache is for. Not required: having read the Dragon Hatchling paper, any neuroscience, or PyTorch.
Time to the core insight: under 60 seconds. The page opens with both computations already running and the agreement figure on screen; the softmax toggle is one click away.
After using the artifact, a learner can:
- State why removing softmax is what permits attention to be rewritten as a recurrence.
- Perform the reassociation
Σ_s (q·k_s)v_s = q·(Σ_s k_s ⊗ v_s)and say which algebraic property (linearity of the dot product) it depends on. - Explain why a softmax Transformer therefore needs a growing KV cache.
- Predict the effect of raising the number of stored associations, and of changing key sparsity, on retrieval fidelity — and explain the prediction via cross-talk.
- Distinguish BDH (graph model,
σon edges,n×n) from BDH-GPU (tensorised special case,ρ = Eσ,n×d), and say which one the artifact implements. - Say why BDH is not "an SSM in the Mamba sense", naming at least one mechanical reason.
- Correctly classify BDH's public results by evidence level — in particular, that the Sudoku Extreme figure is not in the arXiv paper.
No framework, no package dependencies, and no network calls at runtime — the single
external request is the Google Fonts stylesheet, which degrades to system fonts offline.
There is one trivial build step (node build.js, no dependencies) whose only job is to
inline the scripts; its output is committed and verified.
src/page.html Structure, design tokens, all prose. The build template.
src/bdh-kernel.js The maths. A faithful shrunken port of the attention kernel in
pathwaycom/bdh. Exports both implementations plus measurement helpers.
Runs unmodified in Node and in the browser.
src/app.js Interaction layer: control wiring, canvas rendering, live readouts.
Contains no model maths of its own beyond the Lab-2 association store.
build.js Inlines the two scripts into src/page.html to produce index.html.
index.html The shipped artifact: self-contained, generated, committed.
train/ Optional offline training run; its JSON result is inlined by build.js.
index.html is self-contained — the JavaScript is inlined — so it opens by double-click
from an unzipped folder with no server and no dependence on how a browser treats relative
scripts over file://. The obvious risk with inlining is that the shipped copy drifts from
the tested source, so node src/verify-sync.js re-runs the build in memory and byte-compares
it against the committed index.html. It runs first in npm test. The code a judge runs in
the browser is provably the code the tests exercise.
| Component | Role |
|---|---|
BDH.buildLayer() |
Builds one shrunken BDH-style layer up to the attention block: embeddings → encoder → ReLU → RoPE. Mirrors the ordering in bdh.py's forward(). |
BDH.quadratic() |
Form A. Materialises the full T×T score matrix with tril(diagonal=-1), then multiplies by V — exactly as released. Optional softmax flag exists only as the falsification control. |
BDH.recurrent() |
Form B. Never allocates the score matrix. Carries σ (n×d), reading before writing, and applies the rank-1 Hebbian update. Written independently of Form A — no shared code path. |
BDH.getFreqs(), BDH.ropeRow() |
RoPE, ported from get_freqs / Attention.rope, including the quantize(t, q=2) frequency pairing. |
BDH.rng() |
Seeded mulberry32 PRNG. Every figure on the page is reproducible from its seed. |
Lab 2 (buildAssoc, readBack, fidelity in app.js) |
The capacity experiment: writes K key–value pairs into one σ by the same Hebbian rule and reads each back. Keys are non-negative (as ReLU outputs are) and sparse. |
| Part | Status |
|---|---|
| Lab 1, both attention forms | Live. float64, recomputed on every control change. |
σ heatmap and token scrubber |
Live. Real snapshots of the accumulating state. |
| Lab 2 interference curve and retrieval bars | Live. Every point recomputed; the curve is 13 independent runs, not a stored array. |
| Model weights in Labs 1–2 | Synthetic — random, seeded, untrained. |
| Lab 2 keys and values | Synthetic — random, seeded. |
| Trained-model sparsity panel | Precomputed, and labelled as such on the page. Produced by train/train_tiny_bdh.py (a real PyTorch training run), shipped as train/trained_stats.json and inlined at build time. Not recomputed in the browser. |
| Animation | None. No CSS or scripted animation depicts model behaviour anywhere on the page. |
The one precomputed element is the trained-model panel, which is marked precomputed in its
own header. Everything else is computed live. Nothing on this page is a rendered video or a
scripted animation, and there is no "Run" button because there is nothing to wait for.
These are stated on the page itself as well as here.
- The interactive toy is untrained, so its
ReLUpasses ~50% of coordinates. Lab 2 exposes density as a manual control precisely because that toy cannot earn it. We did train a separate 229K-parameter BDH to test the sparsity half of the claim: non-zero activations fell from 50.6% to 29.3% with no sparsity penalty in the loss. That reproduces the direction the paper reports but not the magnitude — we do not reach ~5% and do not claim to have reproduced that figure. Our model is ~0.23M parameters against their 10M–1B, and our synthetic corpus is nearly saturated (val loss 0.228), so there is little for a token to be "busy" about — and the paper states sparsity tracks how much work a token requires. Honest reading: the direction reproduces at toy scale; the magnitude is a property of scale and data we cannot test here. - Only the attention block is ported. The surrounding ReLU-lowrank MLP, the
x ⊙ ygating, the LayerNorms and the multi-layer stack are not reproduced — the claim does not need them. Values are the raw embeddings, matchingV = xin the released code. - Scale. Real BDH-GPU runs
nin the tens of thousands. Atn ≤ 256you see the mechanism, not the regime. Sparse superposition improves with width, so this toy understates how wellσholds up. - Lab 2 uses random keys. That is the pessimistic case; a trained model chooses its keys. The lab measures superposition in a fixed-size matrix, not the accuracy of any trained BDH model.
σis not unconditionally cheaper. Below the crossover the page reports (T ≈ √(n·d)), theT×Tscore matrix is genuinely smaller. The win is asymptotic.- This is not an official BDH model and reproduces no published benchmark.
The artifact contains a table classifying every BDH claim it mentions. Two findings from checking primary sources are worth stating here because they are widely gotten wrong:
- The Sudoku Extreme 97.4% figure does not appear in the arXiv paper. The string
"Sudoku" does not occur anywhere in arXiv:2509.26507v1 — verified by full-text search of
the complete HTML rendering (301k characters; see
docs/bdh-paper-fulltext.txt). The figure comes from Pathway's research blog, from an internal implementation that is not the public repository, as that repository's own README notes. - BDH-CQ's ARC result is on ARC-AGI-1's public evaluation set (400 tasks) — not the semi-private set, and not ARC-AGI-2.
Also classified on the page: the "1B to 600B" scaling sentence (asserted, no supporting data in the report) and the Amazon SageMaker HyperPod relationship (a training-infrastructure partnership announcement — not a benchmark, not a deployment, not an independent evaluation). As of writing we found no independent external reproduction of BDH's headline results.
No dependencies. Node ≥ 18.
npm test # runs all three checks below
node src/verify-sync.js # index.html matches a fresh build of src/
node src/verify-equivalence.js # the central claim + the softmax falsification control
node src/verify-interference.js # the capacity/sparsity table
node src/bench.js # compute cost per recomputeThe training run behind the sparsity panel is optional and is the only part that needs a dependency (PyTorch, CPU-only, a few minutes):
pip install torch
python train/train_tiny_bdh.py # rewrites train/trained_stats.json
node build.js # re-inlines the new stats into index.htmlverify-equivalence.js prints the relative difference between the two forms across six
configurations (worst case 4.9×10⁻¹⁶), then switches the softmax on and shows the
divergence (2.4×10¹). verify-interference.js prints retrieval fidelity over a grid of
load K and key density.
To serve the artifact locally:
python -m http.server 8777Then open http://localhost:8777. It also opens directly from the filesystem by
double-clicking index.html, verified with no network available: every figure still
computes and all four canvases render. Without a network the three Google webfonts fall
back to Georgia and the system monospace; nothing else changes.
| Setting | Full Lab-1 recompute |
|---|---|
T=48, n=128, d=16 (page default) |
17.0 ms |
T=96, n=192, d=24 |
53.6 ms |
T=128, n=256, d=32 (slider maximum) |
102.0 ms |
Slider caps were chosen from these measurements to keep the worst case well under one second including a mobile penalty.
Satisfying the ≥3 recent primary papers (2022–2026) requirement, each cited beside the claim it supports on the page:
- Kosowski, Uznański, Chorowski, Stamirowska, Bartoszkiewicz (2025). The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain. arXiv:2509.26507 — the architecture; §6.4 sparsity; §6.3 monosemantic synapses; §4.2 scaling comparison.
- Engdahl, Kosowski, Chorowski, Stamirowska, Uznański et al. (2026). BDH-CQ:
In-Context Learning with Recurrent Latent Reasoning.
arXiv:2608.09888 — §3.2 names linear attention as the
simplest realisation of its contextual-state update, in the additive special case
S_t = S_{t−1} + U(D_t); ARC-AGI-1 results. - Yang, Wang, Shen, Panda, Kim (2023–24). Gated Linear Attention Transformers with Hardware-Efficient Training. arXiv:2312.06635 — forgetting as the standard repair for the interference Lab 2 measures.
- Yang, Wang, Zhang, Kim (2024). Parallelizing Linear Transformers with the Delta Rule over Sequence Length. arXiv:2406.06484 — delta-rule writes outperform pure Hebbian writes on associative recall, i.e. the named fix for the limitation this artifact exposes.
- Arora et al. (2024). Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff. arXiv:2402.18668 — recall capacity scales with state size; the trade-off Lab 2 exhibits directly.
Foundational context (pre-2022, cited for priority, not to satisfy the requirement):
- Katharopoulos, Vyas, Pappas, Fleuret (2020). Transformers are RNNs. arXiv:2006.16236 — the original statement of the reassociation the artifact runs.
- Schlag, Irie, Schmidhuber (2021). Linear Transformers Are Secretly Fast Weight Programmers. arXiv:2102.11174 — names the capacity limit and introduces the delta-rule mitigation.
Code: pathwaycom/bdh — the reference
implementation ported here.
docs/bdh-paper-fulltext.txt is a plain-text extraction of arXiv:2509.26507v1 retained so
the full-text search claims above can be re-run. docs/research-digest-raw.txt is the raw
primary-source research log.
| Item | Source | Licence |
|---|---|---|
src/bdh-kernel.js |
Original code, ported from pathwaycom/bdh bdh.py (structure and arithmetic) |
MIT (this repo); upstream MIT |
src/app.js, index.html |
Original | MIT |
mulberry32 PRNG |
Public-domain algorithm (Tommy Ettinger) | Public domain |
| Box–Muller transform | Standard textbook method | — |
| Instrument Serif, Source Serif 4, JetBrains Mono | Google Fonts | SIL Open Font License 1.1 |
| Prose, diagrams, tables | Original | CC BY 4.0 |
| Model weights / datasets | None used. No third-party weights, checkpoints, or datasets. | — |
| Images, icons, illustrations | None. All visuals are canvas-rendered from live computation. | — |
| JS libraries | None. Zero runtime dependencies. | — |
Code is MIT (see LICENSE); prose and figures are CC BY 4.0.
Not affiliated with or endorsed by Pathway. "Dragon Hatchling", "BDH" and "BDH-CQ" refer to Pathway's published work; this is an independent educational reimplementation.
Per the track's requirement to disclose all AI-generated, reused, or forked work:
- This project is not a fork. It is an original artifact. The only reused work is the
arithmetic of
Attention.forwardfrompathwaycom/bdh(MIT), reimplemented in JavaScript at reduced dimensions and cited in the source. - AI assistance (Claude) was used substantially for: primary-source research and retrieval; drafting the JavaScript implementation, the page copy, and this README; and structuring the argument.
- What the team did and verified independently: chose the concept and the falsifiable
claim; trained the tiny BDH in
train/and reported its result honestly, including the fact that it does not reach the paper's ~5% sparsity; verified the equivalence numerically before any prose was written (src/verify-equivalence.js); verified the interference behaviour (src/verify-interference.js); readbdh.pyand arXiv:2509.26507 directly rather than relying on summaries; full-text-searched the paper to establish that the Sudoku figure is absent from it; confirmed arXiv:2608.09888 resolves to a real paper with the stated abstract; and benchmarked to set the slider caps. - Every claim on the page is either derivable from a cited primary source or reproducible with a script in this repository. No number was taken from an AI summary without checking it against the primary source.
- The team understands and can defend every component, and can trace any figure on the page back to the line of code that computes it.
START-HERE.md orientation for a judge opening the zip
index.html the artifact (self-contained, generated by build.js)
build.js builds index.html from src/
src/page.html page template
src/bdh-kernel.js the maths (shared by page and tests)
src/app.js interaction and rendering
src/verify-sync.js proves index.html matches src/
src/verify-equivalence.js reproduces the central claim
src/verify-interference.js reproduces the capacity result
src/bench.js performance measurements
train/train_tiny_bdh.py trains the tiny BDH behind the sparsity panel
train/trained_stats.json its measured output, inlined at build time
docs/one-page-summary.pdf the required one-page concept summary
docs/blog.pdf the required blog, as PDF
docs/bdh-paper-fulltext.txt text extraction of arXiv:2509.26507v1
docs/research-digest-raw.txt primary-source research log
LICENSE MIT