Abstract
Kalman Delta Networks reformulate linear attention as a linear-Gaussian state-space model with Kalman-filter updates to track memory uncertainty, yielding efficient scan-compatible approximations that improve language modeling performance.
Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. Delta-rule models learn this strength from the current token embedding but do not track confidence in the memory estimate, preventing each write from adapting to accumulated evidence. To represent this uncertainty explicitly, we reformulate recurrent associative memory as a linear--Gaussian state-space model, for which the Kalman filter is the optimal recursive estimator, and introduce a new family of models, Kalman Delta Networks (KDNs). Within KDNs, the transition propagates both the memory state and its uncertainty, allowing the Kalman gain to weight each residual write by accumulated evidence and observation reliability. Under this formulation, Delta-style updates emerge as a special case that substitutes a token-wise isotropic surrogate for predictive covariance and omits covariance tracking. Exact tracking, however, entails a dense, state-dependent Riccati recursion that is poorly suited to GPU-parallel linear-attention scans. To address this issue, we introduce two scan-compatible KDN approximations. Diagonal KDN projects each one-step posterior onto the diagonal Gaussian family through online mean-field variational inference, whereas Isotropic KDN uses an isotropic approximation with a single uncertainty scalar per head. Their uncertainty recurrences are Mobius maps, enabling associative scans with logarithmic parallel depth. Across controlled pretraining at 750M and 1.3B parameters, KDN variants consistently improve perplexity and mean downstream accuracy over state-of-the-art linear-attention models.
Community
A key question behind this work is: should associative memory treat every new observation with the same level of confidence?
Existing Delta-rule models decide how strongly to update memory from the current token, but they do not explicitly track how certain the model already is about what has been stored. In Kalman Delta Networks (KDNs), we introduce uncertainty into recurrent associative memory through a state-space perspective, allowing each update to adapt to both accumulated evidence and observation reliability.
This connects classical Kalman filtering with modern linear attention, while still leading to efficient, scan-compatible architectures for large-scale language modeling.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MARCH: Scaling Recurrent Memory with Content-Routed State Anchors (2026)
- DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling (2026)
- Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory (2026)
- Maglev: Sliding Recurrent Memory (2026)
- Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones (2026)
- DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization (2026)
- DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.07816 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper