DataForge 2026 · Pathway Track · interactive explainer

Hebbian writes can build attention-like memory.

A growing attention history can sometimes be algebraically reorganized into an evolving, fixed-shape state. This page lets you inspect that idea in a deterministic toy, verify one exact recurrent-versus-history identity, contrast it with softmax, and then stress a separate associative memory until interference appears.

The one-sentence claim (falsifiable): Attention-like retrieval can be implemented by incremental Hebbian-like outer-product writes to a fixed-shape recurrent state; in BDH-GPU, that evolving state—not a token-by-token KV cache—is the mechanism for within-session memory. This is not a claim of exact equivalence to softmax attention. Lab 2 states and numerically tests the narrower equality that does hold.

live + synthetic The labs are an independent, untrained toy and an abstract associative-memory experiment—not an official Pathway model or a reproduced BDH benchmark. The architectural connection is grounded in the primary BDH paper; deviations are listed in Limitations.

Start here

Who this is for, and what you will learn

Audience

ML-literate students and practitioners who know Transformer attention but have not studied BDH.

Prerequisites

Vectors, matrix multiplication, dot-product and softmax attention, causal sequence processing, and the purpose of a KV cache. No neuroscience background is needed.

1 · Trace the state

Identify non-negative activity x, a possibly signed low-rank message, the stored compressed state ρ, and the derived conceptual synapse matrix.

2 · Verify the algebra

Explain why a recurrent outer-product accumulator and its explicit-history sum match, while softmax is only a contrasting operation.

3 · Respect causality

Recognize that the current query reads the pre-write state, so the current key/value pair is excluded from its own retrieval.

4 · State the boundary

Separate live toy measurements, a generic associative-memory stress test, and author-reported claims from the primary papers.

Lab 1

Watch a fixed-shape state being written

Each token reads the previous state before writing ρ ← γρ + x ⊗ v*. The stored recurrent state is always 96×12 = 1,152 floats, whatever the sequence length. The left view shows non-negative toy activity x; gold rings mark neurons contributing to the current write. The right view is the derived 96×96 projection ρDy, not another stored array. Because the learned projections are replaced by signed random matrices here, v*, the raw write, ρ, and ρDy may contain positive and negative values. Primary evidence: BDH-GPU definition and state-space form source-grounded

Neuron activity x live + synthetic
Live neuron activity visualization. Use the text status below for the current token, step, and active fraction.
Derived effective matrix ρDy live + synthetic
Live signed effective-matrix visualization. The stored recurrent state has 1,152 floats and does not grow with the token count.
token: — · step 0 · — This is measured toy activity. The paper's trained-model sparsity observations are separate, author-reported evidence.

Predict, then test: before moving γ, say what should happen. At γ = 0, each write replaces the old recurrent state; at γ = 1, prior writes accumulate. The toy uses scalar γ as a transparent intervention. It is not a complete implementation of the paper's positional operator. The preset token list still exists in the visualizer to drive and label the demo; the fixed-state claim applies to the toy cell's recurrent state, not to the entire teaching UI.

Lab 2

Test the exact identity—then contrast softmax

The green output comes from a fixed 12×8 recurrent accumulator; the blue output recomputes the same causal decayed linear-attention sum from every earlier key/value pair: Στ<t γt−1−τ(qt·kτ)vτ. These are two implementations of the same unnormalized operation, so the displayed maximum component error should be near floating-point roundoff. The orange series separately applies normalized exponential softmax weights. It is a contrast, not an equivalent target, and it need not match either linear-attention output. Primary evidence: BDH-GPU linear-attention analysis source-grounded

Eight output components: recurrent (green), explicit history (blue), and softmax contrast (orange) live + synthetic
Grouped component bars. The green and blue outputs should overlap; the orange softmax contrast may differ. The numeric error is reported below.
recurrent vs explicit-history check: — The first two outputs should agree up to floating-point rounding. Softmax is a separate contrast and is not expected to match.

Explain the result: recurrence is possible because distributivity lets the history sum be accumulated as outer products. The exact equality tested here is recurrent state output = explicit-history decayed linear-attention output. It is not recurrent state output = softmax output. The query reads the pre-write state, so its own key/value pair is excluded; moving γ below 1 discounts older terms.

Lab 3

Stress a generic associative memory

This separate controlled toy stores random unit key→value pairs in one outer-product matrix: M = Σ ki ⊗ vi, then reads v̂i = kiTM. A cue is correct only when its intended stored value is the nearest value by cosine similarity. The second curve reports the mean strongest cosine against the wrong stored values. Points are deterministic aggregate results from three seeded trials, recomputed in the browser. This demonstrates cross-talk in this generic linear associative memory; it does not measure or predict BDH or BDH-CQ capacity.

Nearest-neighbor accuracy (green) and mean strongest wrong-value cosine (dashed red) live + synthetic
Live two-series chart. Use the text result below for the selected load, accuracy, and mean strongest wrong-value cosine.
Waiting for the live nearest-neighbor cue test.

Predict, then test: predict how the green accuracy and red wrong-match curve will change when load rises or dimension increases. Test your prediction, then explain the result using cross-terms (kj·ki)vi. The curve is empirical, seed-dependent evidence about this implemented toy only. There is no universal cliff at load = dimension and no BDH-CQ capacity conclusion here.

60-second challenge

Predict → test → explain

Say each prediction aloud before touching the control. A good explanation names the state, the operation, and the evidence boundary—not just what the animation looked like.

0–20 s · Decay

Predict: what survives after the next Lab 1 write at γ = 0 versus γ = 1? Test: reset and single-step. Explain: use ρ ← γρ + x⊗v*.

20–40 s · Equality

Predict: which Lab 2 series must overlap? Test: move query position and γ. Explain: cite the displayed maximum absolute error and why softmax is exempt.

40–60 s · Interference

Predict: how should load and dimension affect wrong matches? Test: change both in Lab 3. Explain: call it a seeded generic-memory result, never a calibrated BDH capacity law.

BDH module

Map the toy to BDH-GPU—without collapsing the notation

The paper's conceptual graph formulation uses an n×n synaptic state σ. BDH-GPU carries a low-rank, compressed attention state ρ instead. This toy stores ρ as a 96×12 array in its row-vector convention (the public code field retains the legacy name sigma) and computes ρDy only for display. That displayed projection is an effective interaction matrix; it should not be mislabeled as the full conceptual σ.

v*t,l−1 = LN(E yt,l−1)
xt,l = xt,l−1 + (Dx v*t,l−1)⁺
readt,l = ρt−1,l xt,l  ← pre-write state: τ < t only
yt,l = (Dy LN(readt,l))⁺ ⊙ xt,l
ρt,l = (ρt−1,l + v*t,l−1 xt,lᵀ)U  ← Hebbian-like outer-product write

The subscript t,l−1 means the same token at the previous layer, not the previous token in the same layer. Cross-token memory enters through ρt−1,l. At inference, trained projections are fixed while activity and recurrent state evolve. The ReLU outputs x and y are non-negative; neither that fact nor sparsity implies that v*, the outer-product write, or ρ is entrywise non-negative. The paper reports sparse positive activity and connectivity observations in trained models; this toy does not reproduce those findings. Primary evidence: equations (4), (5), and (8) author-reported

BDH-CQ is context, not Lab 3's conclusion. Its authors describe demonstration-conditioned recurrent memory and iterative continuous latent reasoning without inference-time parameter updates. They report 29.5% pass@2 for a 150M-parameter configuration on the public ARC-AGI-1 evaluation set at a computed inference cost of $0.00070 per task. This project did not run BDH-CQ, reproduce those numbers, or establish that latent iteration increases the capacity of the Lab 3 matrix. Primary evidence: BDH-CQ paper author-reported

BDH-GPU vs a standard cached softmax Transformer at inference — scoped comparison source-based synthesis
propertyTransformer (softmax attention)BDH-GPU (synaptic memory)
sequence-dependent statethe usual KV cache adds per-position keys and valuesρ has a fixed shape for a fixed architecture
attention operationnormalized exponential weighting over cached positionslinear-attention terms are accumulated in ρ; not an exact softmax identity
what changes during a sessioncache contents change; trained weights stay fixedx, y, and ρ change; trained projections stay fixed
per-position inspectioncached K/V remain indexed by positionthe aggregate state does not expose a free exact token-by-token map
signarchitecture-dependentx and y are non-negative; v*, writes, and ρ may be signed
evidence shown hereconceptual comparison only; no Transformer benchmark runstoy mechanism plus author-reported BDH properties; no official checkpoint runs
Honesty

Limitations, deviations, misconceptions

Our toy is not BDH

Single layer, single head, no layer norm, no multi-head factorization, random (untrained) E/Dx/Dy. We inject tokens directly into x with a hand-tuned scale; the full layered equations take xt,l−1 and yt,l−1 from the previous layer at the same token. Expect a mechanism illustration, not quantitatively calibrated behavior.

Common misconception: “fixed-state recurrence is softmax attention in another form.” Lab 2 does not show that. It checks an exact identity for causal decayed linear attention and displays softmax as a deliberately non-equivalent contrast.

Only the cell state is fixed-shape

The Lab 2 reference deliberately keeps explicit history to verify the recurrent result and the token strip keeps labels for teaching. The O(1) statement is only with respect to sequence length for the recurrent state of a fixed architecture; that state still costs O(ND) per configured layer/head.

Positional structure is simplified

The toy uses scalar decay γ as a controllable forgetting intervention. It does not implement or validate the paper's complete operator U or its positional variants.

Lab 3 is a separate abstraction

Nearest-neighbor accuracy and mean strongest wrong-value cosine are measured for random associations in the implemented matrix. The curve is seed- and setup-dependent. It establishes neither a universal capacity threshold nor a BDH/BDH-CQ capacity result.

Signs matter

The toy's ReLU activity and readout are non-negative. Signed projections make the low-rank message, outer-product writes, compressed state, and derived effective matrix potentially signed; green/red in the matrix view encodes that distinction.

Evidence levels

Everything labeled live is computed in your browser from synthetic inputs. “Source-grounded” labels identify equations or architectural interpretation. “Author-reported” labels identify paper results that this project did not reproduce. No official BDH or BDH-CQ checkpoint runs in these labs.

Sources

Primary sources

  1. Kosowski, A., Uznański, P., Chorowski, J., Stamirowska, Z., Bartoszkiewicz, M. The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain. arXiv:2509.26507, 2025. arxiv.org/abs/2509.26507 — BDH equations, Hebbian inference, sparsity and degree-distribution findings. code: github.com/pathwaycom/bdh (MIT)
  2. Pathway. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning. arXiv:2608.09888, 2026. arxiv.org/html/2608.09888v1 — demonstrations accumulated additively into contextual state; ARC-AGI-1 result at 150M params.
  3. Gu, A., Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752, 2023. arxiv.org/abs/2312.00752 — the fixed-state alternative lineage BDH is contrasted with (the PS notes BDH-GPU is not an SSM in the Mamba sense).
  4. Yang, S., Wang, B., Shen, Y., Panda, R., Zhang, Y. Gated Linear Attention Transformers with Hardware-Efficient Training. arXiv:2312.06635, 2023. arxiv.org/abs/2312.06635 — linear attention as an RNN with matrix state: the same broad “state plus outer-product write” pattern, without BDH's neuron interpretation.
  5. Sun, Y., Li, X., Dalal, K., Xu, J., Vikram, A., Zhang, G., Dubois, Y., Lu, X., Al-Shedivat, T., Xiong, W. Learning to (Learn at Test Time): RNNs with Expressive Hidden States. arXiv:2407.04620, 2024. arxiv.org/abs/2407.04620 — hidden state as a machine that trains at test time; a different answer to the same fixed-state/adaptation pressure BDH-CQ addresses.
  6. Behrouz, A., Zharkov, P., Mirhoseini, A. Titans: Learning to Memorize at Test Time. arXiv:2501.00663, 2025. arxiv.org/abs/2501.00663 — test-time memorization with surprise-based plasticity updates, adjacent to BDH's Hebbian writes.
  7. Pathway. From attention to synapses: deriving BDH (BDH explainer series, ch. 2). pathway.com/research/bdh-explainer — the derivation used in Lab 2, including the kernel linearization and k = q = x merge.

Built for the DataForge 2026 Pathway Track by Ishpreet Singh and Shikhar Goel. Toy implementation, visualizations and text are original team work with disclosed AI assistance. BDH equations and reported results: Pathway (papers above, MIT-licensed reference code pathwaycom/bdh). Fonts: system stack, no external requests. AI assistance used for drafting and editing, disclosed in README. No cookies, no analytics, no network calls after load.