Audience
ML-literate students and practitioners who know Transformer attention but have not studied BDH.
DataForge 2026 · Pathway Track · interactive explainer
A growing attention history can sometimes be algebraically reorganized into an evolving, fixed-shape state. This page lets you inspect that idea in a deterministic toy, verify one exact recurrent-versus-history identity, contrast it with softmax, and then stress a separate associative memory until interference appears.
The one-sentence claim (falsifiable): Attention-like retrieval can be implemented by incremental Hebbian-like outer-product writes to a fixed-shape recurrent state; in BDH-GPU, that evolving state—not a token-by-token KV cache—is the mechanism for within-session memory. This is not a claim of exact equivalence to softmax attention. Lab 2 states and numerically tests the narrower equality that does hold.
live + synthetic The labs are an independent, untrained toy and an abstract associative-memory experiment—not an official Pathway model or a reproduced BDH benchmark. The architectural connection is grounded in the primary BDH paper; deviations are listed in Limitations.
ML-literate students and practitioners who know Transformer attention but have not studied BDH.
Vectors, matrix multiplication, dot-product and softmax attention, causal sequence processing, and the purpose of a KV cache. No neuroscience background is needed.
Identify non-negative activity x, a possibly signed low-rank message, the stored compressed state ρ, and the derived conceptual synapse matrix.
Explain why a recurrent outer-product accumulator and its explicit-history sum match, while softmax is only a contrasting operation.
Recognize that the current query reads the pre-write state, so the current key/value pair is excluded from its own retrieval.
Separate live toy measurements, a generic associative-memory stress test, and author-reported claims from the primary papers.
Each token reads the previous state before writing ρ ← γρ + x ⊗ v*. The stored recurrent state is always 96×12 = 1,152 floats, whatever the sequence length. The left view shows non-negative toy activity x; gold rings mark neurons contributing to the current write. The right view is the derived 96×96 projection ρDy, not another stored array. Because the learned projections are replaced by signed random matrices here, v*, the raw write, ρ, and ρDy may contain positive and negative values. Primary evidence: BDH-GPU definition and state-space form source-grounded
Predict, then test: before moving γ, say what should happen. At γ = 0, each write replaces the old recurrent state; at γ = 1, prior writes accumulate. The toy uses scalar γ as a transparent intervention. It is not a complete implementation of the paper's positional operator. The preset token list still exists in the visualizer to drive and label the demo; the fixed-state claim applies to the toy cell's recurrent state, not to the entire teaching UI.
The green output comes from a fixed 12×8 recurrent accumulator; the blue output recomputes the same causal decayed linear-attention sum from every earlier key/value pair: Στ<t γt−1−τ(qt·kτ)vτ. These are two implementations of the same unnormalized operation, so the displayed maximum component error should be near floating-point roundoff. The orange series separately applies normalized exponential softmax weights. It is a contrast, not an equivalent target, and it need not match either linear-attention output. Primary evidence: BDH-GPU linear-attention analysis source-grounded
Explain the result: recurrence is possible because distributivity lets the history sum be accumulated as outer products. The exact equality tested here is recurrent state output = explicit-history decayed linear-attention output. It is not recurrent state output = softmax output. The query reads the pre-write state, so its own key/value pair is excluded; moving γ below 1 discounts older terms.
This separate controlled toy stores random unit key→value pairs in one outer-product matrix: M = Σ ki ⊗ vi, then reads v̂i = kiTM. A cue is correct only when its intended stored value is the nearest value by cosine similarity. The second curve reports the mean strongest cosine against the wrong stored values. Points are deterministic aggregate results from three seeded trials, recomputed in the browser. This demonstrates cross-talk in this generic linear associative memory; it does not measure or predict BDH or BDH-CQ capacity.
Predict, then test: predict how the green accuracy and red wrong-match curve will change when load rises or dimension increases. Test your prediction, then explain the result using cross-terms (kj·ki)vi. The curve is empirical, seed-dependent evidence about this implemented toy only. There is no universal cliff at load = dimension and no BDH-CQ capacity conclusion here.
Say each prediction aloud before touching the control. A good explanation names the state, the operation, and the evidence boundary—not just what the animation looked like.
Predict: what survives after the next Lab 1 write at γ = 0 versus γ = 1? Test: reset and single-step. Explain: use ρ ← γρ + x⊗v*.
Predict: which Lab 2 series must overlap? Test: move query position and γ. Explain: cite the displayed maximum absolute error and why softmax is exempt.
Predict: how should load and dimension affect wrong matches? Test: change both in Lab 3. Explain: call it a seeded generic-memory result, never a calibrated BDH capacity law.
The paper's conceptual graph formulation uses an n×n synaptic state σ. BDH-GPU carries a low-rank, compressed attention state ρ instead. This toy stores ρ as a 96×12 array in its row-vector convention (the public code field retains the legacy name sigma) and computes ρDy only for display. That displayed projection is an effective interaction matrix; it should not be mislabeled as the full conceptual σ.
The subscript t,l−1 means the same token at the previous layer, not the previous token in the same layer. Cross-token memory enters through ρt−1,l. At inference, trained projections are fixed while activity and recurrent state evolve. The ReLU outputs x and y are non-negative; neither that fact nor sparsity implies that v*, the outer-product write, or ρ is entrywise non-negative. The paper reports sparse positive activity and connectivity observations in trained models; this toy does not reproduce those findings. Primary evidence: equations (4), (5), and (8) author-reported
BDH-CQ is context, not Lab 3's conclusion. Its authors describe demonstration-conditioned recurrent memory and iterative continuous latent reasoning without inference-time parameter updates. They report 29.5% pass@2 for a 150M-parameter configuration on the public ARC-AGI-1 evaluation set at a computed inference cost of $0.00070 per task. This project did not run BDH-CQ, reproduce those numbers, or establish that latent iteration increases the capacity of the Lab 3 matrix. Primary evidence: BDH-CQ paper author-reported
| property | Transformer (softmax attention) | BDH-GPU (synaptic memory) |
|---|---|---|
| sequence-dependent state | the usual KV cache adds per-position keys and values | ρ has a fixed shape for a fixed architecture |
| attention operation | normalized exponential weighting over cached positions | linear-attention terms are accumulated in ρ; not an exact softmax identity |
| what changes during a session | cache contents change; trained weights stay fixed | x, y, and ρ change; trained projections stay fixed |
| per-position inspection | cached K/V remain indexed by position | the aggregate state does not expose a free exact token-by-token map |
| sign | architecture-dependent | x and y are non-negative; v*, writes, and ρ may be signed |
| evidence shown here | conceptual comparison only; no Transformer benchmark runs | toy mechanism plus author-reported BDH properties; no official checkpoint runs |
Single layer, single head, no layer norm, no multi-head factorization, random (untrained) E/Dx/Dy. We inject tokens directly into x with a hand-tuned scale; the full layered equations take xt,l−1 and yt,l−1 from the previous layer at the same token. Expect a mechanism illustration, not quantitatively calibrated behavior.
The Lab 2 reference deliberately keeps explicit history to verify the recurrent result and the token strip keeps labels for teaching. The O(1) statement is only with respect to sequence length for the recurrent state of a fixed architecture; that state still costs O(ND) per configured layer/head.
The toy uses scalar decay γ as a controllable forgetting intervention. It does not implement or validate the paper's complete operator U or its positional variants.
Nearest-neighbor accuracy and mean strongest wrong-value cosine are measured for random associations in the implemented matrix. The curve is seed- and setup-dependent. It establishes neither a universal capacity threshold nor a BDH/BDH-CQ capacity result.
The toy's ReLU activity and readout are non-negative. Signed projections make the low-rank message, outer-product writes, compressed state, and derived effective matrix potentially signed; green/red in the matrix view encodes that distinction.
Everything labeled live is computed in your browser from synthetic inputs. “Source-grounded” labels identify equations or architectural interpretation. “Author-reported” labels identify paper results that this project did not reproduce. No official BDH or BDH-CQ checkpoint runs in these labs.
Built for the DataForge 2026 Pathway Track by Ishpreet Singh and Shikhar Goel. Toy implementation, visualizations and text are original team work with disclosed AI assistance. BDH equations and reported results: Pathway (papers above, MIT-licensed reference code pathwaycom/bdh). Fonts: system stack, no external requests. AI assistance used for drafting and editing, disclosed in README. No cookies, no analytics, no network calls after load.
Step 1 of 5