Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

The Synapse Is the Cache

An interactive explainer for DataForge 2026 — Pathway track ("Explain the Frontier").

Concept: Synaptic Plasticity as Short-Term Memory (with Linear Attention as the closely-related second topic; both serve one central claim.)

Live artifact: https://puregenius369.github.io/the-synapse-is-the-cache/ — opens without sign-in, no build step, no account. Source: this repository.


The one-sentence claim

Because BDH's attention applies no softmax, it is a Hebbian synaptic memory: one fixed-size matrix σ, written once per token by an outer product and read once per token by a dot product, computing exactly what all-pairs attention computes — but with a bounded capacity, so it forgets through interference unless its activations stay sparse.

The claim has two halves, and the artifact ships a control that falsifies each one.

Half of the claim Control that would falsify it What actually happens
"…computing exactly what all-pairs attention computes" Lab 1 → softmax toggle. If the equality survived a nonlinearity, the stated reason for it would be wrong. Relative difference jumps from ~4×10⁻¹⁶ to ~2.4×10¹. The claim's mechanism is confirmed by its failure mode.
"…forgets through interference unless activations stay sparse" Lab 2 → raise K (associations stored) at fixed n. If retrieval stayed perfect, σ would be a lossless cache. Mean retrieval fidelity falls monotonically; lowering key density restores it.

If neither control could ever make the sentence false, the claim would be too vague to teach.


Intended learner and prerequisites

Audience. Someone who has implemented or carefully read scaled dot-product attention — a strong final-year undergraduate, a master's student, or a working ML engineer meeting post-Transformer architectures for the first time.

Prerequisites. Matrix multiplication; what an outer product is; what a KV cache is for. Not required: having read the Dragon Hatchling paper, any neuroscience, or PyTorch.

Time to the core insight: under 60 seconds. The page opens with both computations already running and the agreement figure on screen; the softmax toggle is one click away.

Learning objectives

After using the artifact, a learner can:

  1. State why removing softmax is what permits attention to be rewritten as a recurrence.
  2. Perform the reassociation Σ_s (q·k_s)v_s = q·(Σ_s k_s ⊗ v_s) and say which algebraic property (linearity of the dot product) it depends on.
  3. Explain why a softmax Transformer therefore needs a growing KV cache.
  4. Predict the effect of raising the number of stored associations, and of changing key sparsity, on retrieval fidelity — and explain the prediction via cross-talk.
  5. Distinguish BDH (graph model, σ on edges, n×n) from BDH-GPU (tensorised special case, ρ = Eσ, n×d), and say which one the artifact implements.
  6. Say why BDH is not "an SSM in the Mamba sense", naming at least one mechanical reason.
  7. Correctly classify BDH's public results by evidence level — in particular, that the Sudoku Extreme figure is not in the arXiv paper.

Architecture of the artifact

No framework, no package dependencies, and no network calls at runtime — the single external request is the Google Fonts stylesheet, which degrades to system fonts offline. There is one trivial build step (node build.js, no dependencies) whose only job is to inline the scripts; its output is committed and verified.

src/page.html         Structure, design tokens, all prose. The build template.
src/bdh-kernel.js     The maths. A faithful shrunken port of the attention kernel in
                      pathwaycom/bdh. Exports both implementations plus measurement helpers.
                      Runs unmodified in Node and in the browser.
src/app.js            Interaction layer: control wiring, canvas rendering, live readouts.
                      Contains no model maths of its own beyond the Lab-2 association store.
build.js              Inlines the two scripts into src/page.html to produce index.html.
index.html            The shipped artifact: self-contained, generated, committed.
train/                Optional offline training run; its JSON result is inlined by build.js.

index.html is self-contained — the JavaScript is inlined — so it opens by double-click from an unzipped folder with no server and no dependence on how a browser treats relative scripts over file://. The obvious risk with inlining is that the shipped copy drifts from the tested source, so node src/verify-sync.js re-runs the build in memory and byte-compares it against the committed index.html. It runs first in npm test. The code a judge runs in the browser is provably the code the tests exercise.

Role of every major component

Component Role
BDH.buildLayer() Builds one shrunken BDH-style layer up to the attention block: embeddings → encoder → ReLU → RoPE. Mirrors the ordering in bdh.py's forward().
BDH.quadratic() Form A. Materialises the full T×T score matrix with tril(diagonal=-1), then multiplies by V — exactly as released. Optional softmax flag exists only as the falsification control.
BDH.recurrent() Form B. Never allocates the score matrix. Carries σ (n×d), reading before writing, and applies the rank-1 Hebbian update. Written independently of Form A — no shared code path.
BDH.getFreqs(), BDH.ropeRow() RoPE, ported from get_freqs / Attention.rope, including the quantize(t, q=2) frequency pairing.
BDH.rng() Seeded mulberry32 PRNG. Every figure on the page is reproducible from its seed.
Lab 2 (buildAssoc, readBack, fidelity in app.js) The capacity experiment: writes K key–value pairs into one σ by the same Hebbian rule and reads each back. Keys are non-negative (as ReLU outputs are) and sparse.

Which parts are live, precomputed, synthetic, or animated

Part Status
Lab 1, both attention forms Live. float64, recomputed on every control change.
σ heatmap and token scrubber Live. Real snapshots of the accumulating state.
Lab 2 interference curve and retrieval bars Live. Every point recomputed; the curve is 13 independent runs, not a stored array.
Model weights in Labs 1–2 Synthetic — random, seeded, untrained.
Lab 2 keys and values Synthetic — random, seeded.
Trained-model sparsity panel Precomputed, and labelled as such on the page. Produced by train/train_tiny_bdh.py (a real PyTorch training run), shipped as train/trained_stats.json and inlined at build time. Not recomputed in the browser.
Animation None. No CSS or scripted animation depicts model behaviour anywhere on the page.

The one precomputed element is the trained-model panel, which is marked precomputed in its own header. Everything else is computed live. Nothing on this page is a rendered video or a scripted animation, and there is no "Run" button because there is nothing to wait for.


Honest limitations

These are stated on the page itself as well as here.

  1. The interactive toy is untrained, so its ReLU passes ~50% of coordinates. Lab 2 exposes density as a manual control precisely because that toy cannot earn it. We did train a separate 229K-parameter BDH to test the sparsity half of the claim: non-zero activations fell from 50.6% to 29.3% with no sparsity penalty in the loss. That reproduces the direction the paper reports but not the magnitude — we do not reach ~5% and do not claim to have reproduced that figure. Our model is ~0.23M parameters against their 10M–1B, and our synthetic corpus is nearly saturated (val loss 0.228), so there is little for a token to be "busy" about — and the paper states sparsity tracks how much work a token requires. Honest reading: the direction reproduces at toy scale; the magnitude is a property of scale and data we cannot test here.
  2. Only the attention block is ported. The surrounding ReLU-lowrank MLP, the x ⊙ y gating, the LayerNorms and the multi-layer stack are not reproduced — the claim does not need them. Values are the raw embeddings, matching V = x in the released code.
  3. Scale. Real BDH-GPU runs n in the tens of thousands. At n ≤ 256 you see the mechanism, not the regime. Sparse superposition improves with width, so this toy understates how well σ holds up.
  4. Lab 2 uses random keys. That is the pessimistic case; a trained model chooses its keys. The lab measures superposition in a fixed-size matrix, not the accuracy of any trained BDH model.
  5. σ is not unconditionally cheaper. Below the crossover the page reports (T ≈ √(n·d)), the T×T score matrix is genuinely smaller. The win is asymptotic.
  6. This is not an official BDH model and reproduces no published benchmark.

Evidence discipline

The artifact contains a table classifying every BDH claim it mentions. Two findings from checking primary sources are worth stating here because they are widely gotten wrong:

  • The Sudoku Extreme 97.4% figure does not appear in the arXiv paper. The string "Sudoku" does not occur anywhere in arXiv:2509.26507v1 — verified by full-text search of the complete HTML rendering (301k characters; see docs/bdh-paper-fulltext.txt). The figure comes from Pathway's research blog, from an internal implementation that is not the public repository, as that repository's own README notes.
  • BDH-CQ's ARC result is on ARC-AGI-1's public evaluation set (400 tasks) — not the semi-private set, and not ARC-AGI-2.

Also classified on the page: the "1B to 600B" scaling sentence (asserted, no supporting data in the report) and the Amazon SageMaker HyperPod relationship (a training-infrastructure partnership announcement — not a benchmark, not a deployment, not an independent evaluation). As of writing we found no independent external reproduction of BDH's headline results.


Reproducing every number

No dependencies. Node ≥ 18.

npm test                          # runs all three checks below
node src/verify-sync.js           # index.html matches a fresh build of src/
node src/verify-equivalence.js    # the central claim + the softmax falsification control
node src/verify-interference.js   # the capacity/sparsity table
node src/bench.js                 # compute cost per recompute

The training run behind the sparsity panel is optional and is the only part that needs a dependency (PyTorch, CPU-only, a few minutes):

pip install torch
python train/train_tiny_bdh.py    # rewrites train/trained_stats.json
node build.js                     # re-inlines the new stats into index.html

verify-equivalence.js prints the relative difference between the two forms across six configurations (worst case 4.9×10⁻¹⁶), then switches the softmax on and shows the divergence (2.4×10¹). verify-interference.js prints retrieval fidelity over a grid of load K and key density.

To serve the artifact locally:

python -m http.server 8777

Then open http://localhost:8777. It also opens directly from the filesystem by double-clicking index.html, verified with no network available: every figure still computes and all four canvases render. Without a network the three Google webfonts fall back to Georgia and the system monospace; nothing else changes.

Measured performance (Node/V8, warm)

Setting Full Lab-1 recompute
T=48, n=128, d=16 (page default) 17.0 ms
T=96, n=192, d=24 53.6 ms
T=128, n=256, d=32 (slider maximum) 102.0 ms

Slider caps were chosen from these measurements to keep the worst case well under one second including a mobile penalty.


Primary sources

Satisfying the ≥3 recent primary papers (2022–2026) requirement, each cited beside the claim it supports on the page:

  1. Kosowski, Uznański, Chorowski, Stamirowska, Bartoszkiewicz (2025). The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain. arXiv:2509.26507 — the architecture; §6.4 sparsity; §6.3 monosemantic synapses; §4.2 scaling comparison.
  2. Engdahl, Kosowski, Chorowski, Stamirowska, Uznański et al. (2026). BDH-CQ: In-Context Learning with Recurrent Latent Reasoning. arXiv:2608.09888 — §3.2 names linear attention as the simplest realisation of its contextual-state update, in the additive special case S_t = S_{t−1} + U(D_t); ARC-AGI-1 results.
  3. Yang, Wang, Shen, Panda, Kim (2023–24). Gated Linear Attention Transformers with Hardware-Efficient Training. arXiv:2312.06635 — forgetting as the standard repair for the interference Lab 2 measures.
  4. Yang, Wang, Zhang, Kim (2024). Parallelizing Linear Transformers with the Delta Rule over Sequence Length. arXiv:2406.06484 — delta-rule writes outperform pure Hebbian writes on associative recall, i.e. the named fix for the limitation this artifact exposes.
  5. Arora et al. (2024). Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff. arXiv:2402.18668 — recall capacity scales with state size; the trade-off Lab 2 exhibits directly.

Foundational context (pre-2022, cited for priority, not to satisfy the requirement):

  1. Katharopoulos, Vyas, Pappas, Fleuret (2020). Transformers are RNNs. arXiv:2006.16236 — the original statement of the reassociation the artifact runs.
  2. Schlag, Irie, Schmidhuber (2021). Linear Transformers Are Secretly Fast Weight Programmers. arXiv:2102.11174 — names the capacity limit and introduces the delta-rule mitigation.

Code: pathwaycom/bdh — the reference implementation ported here.

docs/bdh-paper-fulltext.txt is a plain-text extraction of arXiv:2509.26507v1 retained so the full-text search claims above can be re-run. docs/research-digest-raw.txt is the raw primary-source research log.


Source, licence, and asset record

Item Source Licence
src/bdh-kernel.js Original code, ported from pathwaycom/bdh bdh.py (structure and arithmetic) MIT (this repo); upstream MIT
src/app.js, index.html Original MIT
mulberry32 PRNG Public-domain algorithm (Tommy Ettinger) Public domain
Box–Muller transform Standard textbook method —
Instrument Serif, Source Serif 4, JetBrains Mono Google Fonts SIL Open Font License 1.1
Prose, diagrams, tables Original CC BY 4.0
Model weights / datasets None used. No third-party weights, checkpoints, or datasets. —
Images, icons, illustrations None. All visuals are canvas-rendered from live computation. —
JS libraries None. Zero runtime dependencies. —

Code is MIT (see LICENSE); prose and figures are CC BY 4.0.

Not affiliated with or endorsed by Pathway. "Dragon Hatchling", "BDH" and "BDH-CQ" refer to Pathway's published work; this is an independent educational reimplementation.


AI assistance disclosure

Per the track's requirement to disclose all AI-generated, reused, or forked work:

  • This project is not a fork. It is an original artifact. The only reused work is the arithmetic of Attention.forward from pathwaycom/bdh (MIT), reimplemented in JavaScript at reduced dimensions and cited in the source.
  • AI assistance (Claude) was used substantially for: primary-source research and retrieval; drafting the JavaScript implementation, the page copy, and this README; and structuring the argument.
  • What the team did and verified independently: chose the concept and the falsifiable claim; trained the tiny BDH in train/ and reported its result honestly, including the fact that it does not reach the paper's ~5% sparsity; verified the equivalence numerically before any prose was written (src/verify-equivalence.js); verified the interference behaviour (src/verify-interference.js); read bdh.py and arXiv:2509.26507 directly rather than relying on summaries; full-text-searched the paper to establish that the Sudoku figure is absent from it; confirmed arXiv:2608.09888 resolves to a real paper with the stated abstract; and benchmarked to set the slider caps.
  • Every claim on the page is either derivable from a cited primary source or reproducible with a script in this repository. No number was taken from an AI summary without checking it against the primary source.
  • The team understands and can defend every component, and can trace any figure on the page back to the line of code that computes it.

Repository map

START-HERE.md                  orientation for a judge opening the zip
index.html                     the artifact (self-contained, generated by build.js)
build.js                       builds index.html from src/
src/page.html                  page template
src/bdh-kernel.js              the maths (shared by page and tests)
src/app.js                     interaction and rendering
src/verify-sync.js             proves index.html matches src/
src/verify-equivalence.js      reproduces the central claim
src/verify-interference.js     reproduces the capacity result
src/bench.js                   performance measurements
train/train_tiny_bdh.py        trains the tiny BDH behind the sparsity panel
train/trained_stats.json       its measured output, inlined at build time
docs/one-page-summary.pdf      the required one-page concept summary
docs/blog.pdf                  the required blog, as PDF
docs/bdh-paper-fulltext.txt    text extraction of arXiv:2509.26507v1
docs/research-digest-raw.txt   primary-source research log
LICENSE                        MIT

About

Interactive explainer: BDH's attention has no softmax, so it is exactly a Hebbian synaptic memory — one fixed-size matrix written once per token. Run both forms, watch them agree to 1e-16, then break it with one toggle. DataForge 2026, Pathway track.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages