AI Post Transformers

← AI Post Transformerseergisteren

MoM: Mixture-of-Memories Fixes Linear Attention Recall

MoM: Mixture-of-Memories Fixes Linear Attention Recalleergisteren

This episode explores MoM (Mixture-of-Memories), a linear sequence modeling architecture from researchers at Shanghai AI Laboratory and collaborating universities that tackles a core weakness in efficient Transformer alternatives: their tendency to forget information from earlier in a sequence. The discussion traces the trade-off at the heart of the field — Transformers preserve every token via a growing key-value cache at quadratic cost, while linear models like Mamba and RWKV compress everything into a single fixed-size memory state, trading recall precision for constant-time efficiency. MoM's proposed fix draws on two distinct sources: a neuroscience-inspired analogy to how the hippocampus uses separate oscillatory channels to keep simultaneous memories from blending together, and the Mixture-of-Experts routing mechanism, applied here to memory states rather than feed-forward layers. The result is an architecture with multiple independent memory slots plus a shared accumulating memory, with a lightweight router directing each token to the appropriate slot. Listeners interested in efficient sequence modeling, long-context recall, or the cross-pollination between neuroscience and deep learning architecture design will find the mechanics of this capacity-versus-interference problem — and its proposed solution — a compelling deep dive.

Sources:

1. MoM: Linear Sequence Modeling with Mixture-of-Memories — Jusen Du, Weigao Sun, Disen Lan, Jiaxi Hu, Yu Cheng, 2025

http://arxiv.org/abs/2502.13685v4

2. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François Fleuret, 2020

https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention

3. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023 (updated 2024)

https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces

4. Gated Linear Attention Transformers with Hardware-Efficient Training — Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, Yoon Kim, 2023

https://scholar.google.com/scholar?q=Gated+Linear+Attention+Transformers+with+Hardware-Efficient+Training

5. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, Jeff Dean, 2017

https://scholar.google.com/scholar?q=Outrageously+Large+Neural+Networks%3A+The+Sparsely-Gated+Mixture-of-Experts+Layer