A recap, not a new result
The blog post is dated December 4, 2025, but neither piece of research it describes is new. Titans, the memory module, was posted to arXiv on December 31, 2024, credited to Ali Behrouz, Peilin Zhong and Vahab Mirrokni. MIRAS, the generalizing framework, followed on April 17, 2025, adding Meisam Razaviyayn to the author list. The post is Google Research folding roughly a year of its own work into one explainer, not an announcement of anything that shipped that week. [1][2][3]
Atlas interpretation: That matters for how much weight the post's own framing deserves. Titans is presented as "the tool" and MIRAS as "the blueprint" that retroactively explains it, but MIRAS is the later paper. The framework was built to describe an architecture that already existed, then extended into three further variants, Moneta, Yaad and Memora, that MIRAS's four design axes make possible. [2][3]
Attention remembers exactly; recurrence forgets on purpose
A Transformer's attention layer can look back at any earlier token with no loss of detail, which is why it models context accurately, and why its compute cost grows quadratically with sequence length, the tradeoff behind the architecture Titans is trying to get around. A recurrent model sidesteps that cost by compressing everything it has seen into one fixed-size state, cheap to update, but the state has no more room at token ten thousand than it had at token one, so old information gets overwritten rather than remembered. [2]
Atlas interpretation: This is the same tradeoff LSTM's gates were built to manage inside a fixed-size cell, nearly three decades earlier. Titans keeps the recurrent model's cheap update but replaces the fixed vector with a small deep network, a multilayer perceptron, that has more room to store distinctions than a single vector does. The paper's framing splits memory into attention as short-term and this module as long-term, run alongside each other rather than one replacing the other. [2]
A surprise metric decides what gets written
Titans updates its memory module using the gradient of its own prediction error on the incoming token as a proxy for how surprising that token is: an easily predicted continuation produces a small gradient and barely changes the stored state, while a token that violates what memory already expects produces a large one and gets written in. The paper adds momentum, so a run of surprising tokens keeps influencing updates briefly after the surprise passes, and a learned, per-token decay rate that lets old, low-value content fade rather than pile up forever. [2][1]
The paper reports Titans beating both Transformer and Mamba-2 baselines at 360M and 760M parameter scales on language modeling and commonsense reasoning, and reports results on the BABILong long-context benchmark that surpass much larger models, including GPT-4, on that benchmark specifically. All of these are Google's own reported numbers; no independent replication of either paper is cited on the timeline. [2]
MIRAS names four knobs, Titans is one setting of them
MIRAS reframes both attention and linear recurrent models as instances of the same associative-memory problem, one that varies along four axes: what kind of memory architecture stores associations, from a single vector up to a deep network; what objective, the attentional bias, decides how well a new item fits what is already stored, dot-product similarity or a regression loss or others; what retention gate governs forgetting; and what algorithm, from plain gradient descent to Newton's method, updates memory in response. Titans is one point in that space. Moneta, Yaad and Memora are three others, each swapping in a different attentional bias or retention rule to trade off recall precision against long-context stability. [3]
Atlas interpretation: The framework's real claim is not that any one of these four variants wins outright, but that the design space they came from is larger than "attention or recurrence," and that most existing architectures occupy only a narrow corner of it. As with Titans, the comparisons supporting that claim are the authors' own, and as of this writing MIRAS remains an arXiv preprint with no listed peer-reviewed publication venue, so the generality of the claim has not yet been tested outside the lab that proposed it. [3]
Sources
- Titans + MIRAS: Helping AI have long-term memory
Google · Dec 4, 2025
- Titans: Learning to Memorize at Test Time
arXiv · Dec 31, 2024
- It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
arXiv · Apr 17, 2025