efficiencyPublished: July 8, 2026

The Key to Going Linear: Analysis-Driven Transformer Linearization

By Anna Kuzina, Paul N. Whatmough, Babak Ehteshami Bejnordi

Research TL;DR

"This paper identifies that delta-style state updates (key-dependent orthogonal projections) are crucial for linearizing transformers, and introduces sink tokens, short convolutions, and fixed-budget cache routing to close the gap with full attention."

Abstract

The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While numerous post hoc linearization pipelines exist, it is difficult to identify which components preserve model quality. This work isolates the effect of state update design in a strict frozen-backbone regime. We show that softmax relies on key-dependent, rank-1 orthogonal projections, elucidating why delta-style networks outperform purely gated accumulation. We identify a potential source of approximation errors and introduce structural interventions, specifically sink tokens, short convolutions, and fixed-budget cache routing, which reduces the remaining gap. We scale this linearization approach across LLaMA and Qwen models up to 32B parameters, outperforming prior post hoc baselines on MMLU and matching the long-context retrieval of complex adaptive-caching frameworks.

Technical Analysis & Implementation

Core Methodology§

The paper aims to linearize pretrained transformers (frozen backbone) by replacing softmax attention with a linear attention mechanism. The key insight is that softmax implicitly computes a rank-1 orthogonal projection of the value vectors, conditioned on the query. This is formalized as:

$$ \text{softmax}(QK^\top)V = \sum_{i} a_i v_i, \quad a_i = \frac{\exp(q^\top k_i)}{\sum_j \exp(q^\top k_j)} $$

The authors show that this can be approximated by a linear state update if the state is updated with a delta rule (akin to a linear gating with key-dependent orthogonal projections). In contrast, purely gated accumulation (e.g., linear attention with cumulative sum) fails to capture the dynamic reweighting of softmax.

They identify two sources of approximation error: (1) non-stationary token representations due to layer norms and residual connections, and (2) the need for unbounded memory state. To mitigate these, they introduce three structural interventions:

  • Sink tokens: Additional learnable tokens that absorb attention mass and stabilize the state.
  • Short convolutions: A 1D causal convolution applied to key and value sequences before linearization, smoothing out local context.
  • Fixed-budget cache routing: A mechanism to bound the state size by routing tokens to a fixed number of buckets, ensuring O(1) memory per layer.

Implementation Details§

Given a frozen transformer with softmax attention, they replace it with:

$$ \text{State}_t = \text{State}_{t-1} + \Delta(\text{State}_{t-1}, k_t, v_t) $$

where $\Delta$ is either a gated accumulation or a delta rule. The delta rule is:

$$ \Delta S = (v_t - S_{t-1}\phi(k_t)) \otimes \phi(k_t) $$

with $\phi$ a kernel feature map (here, simply identity). Sink tokens are prepended to the sequence as learned embeddings. Short convolutions are applied per channel with kernel size 3. Cache routing uses a fixed number of slots (e.g., 64) and routes each token to the nearest slot via learned keys.

Code Snippet§

import torch
import torch.nn as nn

class LinearAttentionLayer(nn.Module):
    def __init__(self, d_model, num_sinks=1, conv_kernel=3, num_slots=64):
        super().__init__()
        self.sink_tokens = nn.Parameter(torch.randn(1, num_sinks, d_model))
        self.short_conv = nn.Conv1d(d_model, d_model, kernel_size=conv_kernel, padding=conv_kernel-1, bias=False)
        self.state = None
        self.num_slots = num_slots
        self.slot_keys = nn.Parameter(torch.randn(num_slots, d_model))

    def forward(self, q, k, v, cache=None):
        # Add sink tokens
        batch_size = q.shape[0]
        sinks = self.sink_tokens.expand(batch_size, -1, -1)
        k = torch.cat([sinks, k], dim=1)
        v = torch.cat([sinks, v], dim=1)

        # Short convolution (apply on last dim)
        k = self.short_conv(k.transpose(1,2)).transpose(1,2)[:, :-self.short_conv.padding[0], :]
        v = self.short_conv(v.transpose(1,2)).transpose(1,2)[:, :-self.short_conv.padding[0], :]

        # Fixed-budget cache routing
        # Simplified: assign each token to nearest slot, then accumulate
        # For brevity, assume state is maintained externally
        output = []
        for t in range(q.shape[1]):
            # delta update
            state = cache if cache is not None else torch.zeros(batch_size, self.num_slots, v.shape[-1], device=q.device)
            # Compute routing weights
            scores = torch.matmul(k[:, t:t+1, :], self.slot_keys.T)  # (B, 1, num_slots)
            weights = torch.softmax(scores, dim=-1)
            # Update state: delta rule
            state = state + torch.matmul(weights.transpose(1,2), (v[:, t:t+1, :] - state))
            # Attend using state
            out = torch.matmul(q[:, t:t+1, :], state.transpose(1,2))  # simplified
            output.append(out)
        return torch.cat(output, dim=1), state

Results§

The approach is scaled to LLaMA and Qwen models up to 32B parameters. It outperforms prior post hoc linearization baselines on MMLU and achieves long-context retrieval accuracy comparable to complex adaptive-caching frameworks (e.g., Infini-Attention).

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window1,000,000 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000248
Output Cost (Generated):$0.000744
Total Est. Cost:$0.000992
Context Window Capacity0.0124%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
Fugu Max$2.00$6.00
Fugu Ultra v2$5.00$30.00
Ling 3.0 Flash VL$0.06$0.18
DeepSeek V4.1 Flash$0.15$0.60
Mercury 2.5$0.04$0.15
GPT-6 Astra$10.00$50.00
GPT-6 Astra Pro$10.00$50.00
Qwen3.8 Max (0902)$2.00$6.00
Muse Spark 1.3$1.25$4.25
Muse Spark 1.3 Contributor$0.10$0.20
Gemini 3.8 Flash$0.75$3.75
Claude Fable 5.1$10.00$50.00
Granite 4.2 8B$0.06$0.25
Mercury 2.5 Preview$0.04$0.15
Hy4 preview$0.83$2.50
Ling 3.0 Flash Fin$0.06$0.18
GLM Flash Latest$0.07$0.25
Qwen3.8 Flash$0.15$0.47
GLM 5.3 Flash$0.09$0.30
DeepSeek V4 Flash Vision Exp$0.22$0.66
Muse Spark 1.2 Contributor$0.10$0.20
Hy-MT2-30B-A3B$0.07$0.29
Hy-MT2-1.8B$0.04$0.18
GLM Latest$0.88$2.97
Hy-MT2-7B$0.07$0.29
GLM 5.3$1.40$4.40
Qwen3.8 27B$0.21$2.55
Gemini 3.7 Flash$0.75$3.75
Seed 2.1 Turbo$0.50$2.50
Grok 4.6$2.00$6.00
DeepSeek V4 Pro 0813$0.58$1.74
Qwen3.8 2.4T A95B$2.00$6.00
Seed-2.0-Code$0.50$3.00
Nemotron 3.5 Lightning$0.08$0.20
Sakana Namazu$0.95$4.00
Solar Pro 4$0.09$0.36
Muse Glimmer 30B$0.35$1.50
Muse Spark 1.2$1.25$4.25
Qwen3.8 Max$2.00$6.00
DeepSeek V4 Flash 0731$0.06$0.12
Inkling Small$0.45$1.20
Qwen3.7 Flash$0.03$0.13
Claude Opus 5 (Fast)$10.00$50.00
Claude Opus 5$5.00$25.00
Ling 3.0 Flash$0.02$0.06
Gemini 3.5 Flash Lite$0.30$2.50
Gemini 3.6 Flash$0.75$3.75
Laguna S 2.1$0.09$0.18
Inkling$1.00$4.05
Auto Router (Beta)$0.00$0.00
Muse Spark 1.1$1.25$4.25
Kimi K3$2.65$13.28
KAT-Coder-Air V2.5$0.15$0.60
KAT-Coder-Pro V2.5$0.74$2.96
GPT-5.6 Luna$0.20$1.20
GPT-5.6 Luna Pro$0.20$1.20
GPT-5.6 Terra$2.00$12.00
GPT-5.6 Sol$2.00$10.00
GPT-5.6 Terra Pro$2.00$12.00
GPT-5.6 Sol Pro$2.00$10.00
Grok 4.5$2.00$6.00
Hy3$0.08$0.33
Laguna XS 2.1$0.06$0.12
Claude Sonnet 5$2.00$10.00
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Nex-N2-Mini$0.03$0.10
Fugu Ultra$5.00$30.00
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
GLM 5.2$1.40$4.40
Fusion$0.00$0.00
Kimi K2.7 Code$0.71$3.21
Claude Fable Latest$10.00$50.00
Claude Fable 5$10.00$50.00
Nex-N2-Pro$0.25$1.00
Nemotron 3.5 Content Safety$0.20$0.20
Nemotron 3 Ultra$0.63$3.13
Qwen3.7 Plus$0.32$1.28
MiniMax M3$0.30$1.20
Step 3.7 Flash$0.20$1.15
Claude Opus 4.8 (Fast)$10.00$50.00
Claude Opus 4.8$5.00$25.00
Llama 4 Maverick$0.19$0.65
Qwen3.7 Max$1.48$4.42
Grok Build 0.1$1.00$2.00
Gemini 3.5 Flash$1.50$9.00
Claude Opus 4.7 (Fast)$30.00$150.00
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
Grok 4.20$1.25$2.50
Granite 4.1 8B$0.05$0.10
Mistral Medium 3.5$1.50$7.50
Grok 4.3$1.25$2.50
Laguna M.1$0.20$0.40
Claude Haiku Latest$1.00$5.00
Gemini Flash Latest$0.75$3.75
Claude Sonnet Latest$2.00$10.00
Gemini Pro Latest$2.00$12.00
Kimi Latest$2.10$10.95
Google Gemini Flash Latest$0.75$3.75
Google Gemini Pro Latest$2.00$12.00
Anthropic Claude Sonnet Latest$2.00$10.00
Qwen3.5 Plus 2026-04-20$0.30$1.80
Qwen3.6 35B A3B$0.10$0.90
Qwen3.6 Max Preview$1.03$6.16
Qwen3.6 27B$0.30$2.00
Anthropic Claude Haiku Latest$1.00$5.00
Qwen3.6 Flash$0.19$1.13
MoonshotAI Kimi Latest$2.10$10.95
DeepSeek V4 Pro 0423$1.60$3.20
DeepSeek V4 Flash 0423$0.09$0.17
GPT-5.5 Pro$30.00$180.00
DeepSeek V4 Flash$0.09$0.17
GPT-5.5$5.00$30.00
DeepSeek V4 Pro$1.60$3.20
MiMo-V2.5$0.14$0.28
MiMo-V2.5-Pro$0.43$0.87
Hy3 preview$0.18$0.60
Pareto Code Router$0.00$0.00
GPT-5.4 Image 2$8.00$15.00
Claude Opus Latest$5.00$25.00
Kimi K2.6$0.95$4.00
Gemini 3.1 Flash$0.25$1.50
Gemini 3.1 Pro$2.00$12.00
Claude Opus 4.7$5.00$25.00
GLM 5.1$0.97$3.04
Gemma 4 26B A4B$0.09$0.30
Gemma 4 31B$0.09$0.34
Qwen3.6 Plus$0.33$1.95
GLM 5V Turbo$1.20$4.00
Grok 4.20 Multi-Agent$1.25$2.50
Grok 4.20$1.25$2.50
Lyria 3 Pro Preview$0.00$0.00
Lyria 3 Clip Preview$0.00$0.00
KAT-Coder-Pro V2$0.30$1.20
Reka Edge$0.10$0.10
MiniMax M2.7$0.30$1.20
GPT-5.4 Mini$0.75$4.50
GPT-5.4 Nano$0.20$1.25
Mistral Small 4$0.15$0.60
GLM 5 Turbo$1.20$4.00
Nemotron 3 Super$0.08$0.45
Qwen3.5-9B$0.10$0.15
Seed-2.0-Lite$0.25$2.00
GPT-5.4 Pro$30.00$180.00
GPT-5.4$2.50$15.00
Mercury 2$0.25$0.75
Gemini 3.1 Flash Lite Preview$0.25$1.50
GPT-5.3 Chat$1.75$14.00
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Seed-2.0-Mini$0.10$0.40
Qwen3.5-122B-A10B$0.26$2.08
Qwen3.5-27B$0.20$1.56
Qwen3.5-35B-A3B$0.16$1.30
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
Qwen3.5-Flash$0.07$0.26
GPT-5.3-Codex$1.75$14.00
Gemini 3.1 Pro Preview$2.00$12.00
Claude Sonnet 4.6$3.00$15.00
Qwen3.5 Plus 2026-02-15$0.26$1.56
Qwen3.5 397B A17B$0.55$3.50
MiniMax M2.5$0.27$1.08
GLM 5$0.60$1.92
Qwen3 Max Thinking$0.78$3.90
Qwen3 Coder Next$0.12$0.80
Claude Opus 4.6$5.00$25.00
Free Models Router$0.00$0.00
Step 3.5 Flash$0.10$0.30
Solar Pro 3$0.15$0.60
Kimi K2.5$0.45$2.25
MiniMax M2-her$0.30$1.20
Palmyra X5$0.60$6.00
GLM 4.7 Flash$0.06$0.40
GPT Audio Mini$0.60$2.40
GPT Audio$2.50$10.00
Doubao Pro$0.80$1.60
GPT-5.2-Codex$1.75$14.00
MiniMax M2.1$0.30$1.20
Seed 1.6 Flash$0.07$0.30
Seed 1.6$0.25$2.00
GLM 4.7$0.40$1.75
Gemini 3 Flash Preview$0.50$3.00
Nemotron 3 Nano 30B A3B$0.05$0.20
GPT-5.2$1.75$14.00
GPT-5.2 Pro$21.00$168.00
GPT-5.2 Chat$1.75$14.00
Devstral 2 2512$0.40$2.00
GLM 4.6V$0.30$0.90
Body Builder (beta)$0.00$0.00
GPT-5.1-Codex-Max$1.25$10.00
Nova 2 Lite$0.30$2.50
Ministral 3 14B 2512$0.20$0.20
Ministral 3 3B 2512$0.10$0.10
Ministral 3 8B 2512$0.15$0.15
DeepSeek V3.2$0.27$0.40
Mistral Large 3 2512$0.50$1.50
Claude Opus 4.5$5.00$25.00
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
GPT-5.1 Chat$1.25$10.00
GPT-5.1$1.25$10.00
GPT-5.1-Codex$1.25$10.00
GPT-5.1-Codex-Mini$0.25$2.00
Qwen 2.5-Coder 32B$0.35$0.70
Kimi K2 Thinking$0.60$2.50
Hunyuan Pro$0.60$1.20
Nova Premier 1.0$2.50$12.50
Sonar Pro Search$3.00$15.00
Voxtral Small 24B 2507$0.10$0.30
gpt-oss-safeguard-20b$0.07$0.30
MiniMax M2$0.26$1.02
Qwen3 VL 32B Instruct$0.10$0.42
Granite 4.0 Micro$0.02$0.11
GPT-5 Image Mini$2.50$2.00
Claude Haiku 4.5$1.00$5.00
Qwen3 VL 8B Thinking$0.18$2.10
Qwen3 VL 8B Instruct$0.12$0.46
GPT-5 Image$10.00$10.00
o4 Mini Deep Research$2.00$8.00
o3 Deep Research$10.00$40.00
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Thinking$0.20$2.40
GPT-5 Pro$15.00$120.00
Qwen3 VL 30B A3B Instruct$0.13$0.52
Yi-Lightning$0.15$0.30
GLM 4.6$0.43$1.75
DeepSeek V3.2 Exp$0.27$0.41
Claude Sonnet 4.5$3.00$15.00
Cydonia 24B V4.1$0.30$0.50
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Qwen3 Max$0.78$3.90
GPT-5 Codex$1.25$10.00
Qwen3 Coder Plus$0.65$3.25
Qwen3 VL 235B A22B Thinking$0.40$4.00
Qwen3 VL 235B A22B Instruct$0.21$1.90
DeepSeek V3.1 Terminus$0.27$1.00
Qwen 2.5 72B$0.40$0.80
Qwen3 Coder Flash$0.20$0.97
Qwen3 Next 80B A3B Instruct$0.09$1.10
Qwen3 Next 80B A3B Thinking$0.15$1.20
Qwen Plus 0728 (thinking)$0.26$0.78
Qwen Plus 0728$0.26$0.78
Kimi K2 0905$0.60$2.50
ERNIE 4.0$1.20$2.40
Qwen3 30B A3B Thinking 2507$0.20$2.40
Hermes 4 70B$0.13$0.40
Hermes 4 405B$1.00$3.00
DeepSeek V3.1$0.25$0.95
Mistral Medium 3.1$0.40$2.00
GLM 4.5V$0.60$1.80
Jamba Large 1.7$2.00$8.00
GPT-5 Nano$0.05$0.40
GPT-5 Chat$1.25$10.00
GPT-5 Mini$0.25$2.00
GPT-5$1.25$10.00
gpt-oss-20b$0.03$0.13
Claude Opus 4.1$15.00$75.00
gpt-oss-120b$0.04$0.17
Codestral 2508$0.30$0.90
Qwen3 Coder 30B A3B Instruct$0.07$0.28
Qwen3 30B A3B Instruct 2507$0.05$0.19
GLM 4.5$0.60$2.20
Qwen3 235B A22B Thinking 2507$0.23$2.30
GLM 4.5 Air$0.13$0.85
Mistral Large 2$0.60$1.80
Qwen3 Coder 480B A35B$0.30$1.00
UI-TARS 7B$0.10$0.20
Gemini 2.5 Flash Lite$0.10$0.40
Qwen3 235B A22B Instruct 2507$0.09$0.35
Kimi K2 0711$0.57$2.30
Hunyuan A13B Instruct$0.14$0.57
Morph V3 Fast$0.80$1.20
Morph V3 Large$0.90$1.90
ERNIE 4.5 VL 424B A47B$0.42$1.25
Mistral Small 3.2 24B$0.09$0.25
Gemini 2.5 Flash$0.30$2.50
MiniMax M1$0.40$2.20
Gemini 2.5 Pro$1.25$10.00
o3 Pro$20.00$80.00
Gemini 2.5 Pro Preview 06-05$1.25$10.00
R1 0528$0.50$2.15
Claude Sonnet 4$3.00$15.00
Claude Opus 4$15.00$75.00
Gemma 3n 4B$0.06$0.12
Gemini 2.5 Pro Preview 05-06$1.25$10.00
Mistral Medium 3$0.40$2.00
Llama Guard 4 12B$0.18$0.18
Qwen3 14B$0.12$0.24
Qwen3 32B$0.08$0.28
Qwen3 8B$0.12$0.46
Qwen3 30B A3B$0.12$0.50
Qwen3 235B A22B$0.46$1.82
o3$2.00$8.00
o4 Mini High$1.10$4.40
o4 Mini$1.10$4.40
GPT-4.1 Mini$0.40$1.60
GPT-4.1 Nano$0.10$0.40
GPT-4.1$2.00$8.00
Llama 4 Maverick$0.19$0.65
Llama 4 Scout$0.10$0.30
DeepSeek V3 0324$0.25$1.00
o1-pro$150.00$600.00
Mistral Small 3.1 24B$0.35$0.56
Gemma 3 4B$0.05$0.10
Command A$2.50$10.00
Gemma 3 12B$0.05$0.15
Reka Flash 3$0.10$0.20
GPT-4o-mini Search Preview$0.15$0.60
Gemma 3 27B$0.08$0.45
GPT-4o Search Preview$2.50$10.00
Skyfall 36B V2$0.55$0.80
Sonar Deep Research$2.00$8.00
Sonar Pro$3.00$15.00
Sonar Reasoning Pro$2.00$8.00
Saba$0.20$0.60
Claude 3.5 Sonnet v2$3.00$15.00
o3 Mini High$1.10$4.40
Gemini 2.0 Flash$0.10$0.40
Qwen2.5 VL 72B Instruct$0.80$1.00
Qwen-Plus$0.26$0.78
o3 Mini$1.10$4.40
Mistral Small 3$0.09$0.25
Sonar$1.00$1.00
R1 Distill Llama 70B$0.80$0.80
R1$0.70$2.50
DeepSeek R1$0.70$2.50
MiniMax-01$0.20$1.10
Phi 4$0.07$0.14
DeepSeek V3$0.26$1.03
o1$15.00$60.00
Command R7B (12-2024)$0.04$0.15
Mixtral 8x22B$0.50$1.00
Llama 3.3 70B Instruct$0.10$0.32
Llama 3.3 70B Instruct$0.10$0.32
Nova Micro 1.0$0.04$0.14
Nova Lite 1.0$0.06$0.24
Nova Pro 1.0$0.80$3.20
GPT-4o (2024-11-20)$2.50$10.00
Mistral Large 2407$2.00$6.00
Qwen2.5 Coder 32B Instruct$0.66$1.00
UnslopNemo 12B$0.40$0.40
Ministral 8B$0.11$0.11
Qwen2.5 7B Instruct$0.10$0.20
Inflection 3 Productivity$2.50$10.00
Inflection 3 Pi$2.50$10.00
Llama 3.2 3B Instruct$0.05$0.33
Llama 3.2 11B Vision Instruct$0.34$0.34
Llama 3.2 1B Instruct$0.03$0.20
Llama 3.2 11B Vision$0.34$0.34
Qwen2.5 72B Instruct$0.36$0.40
Command R (08-2024)$0.15$0.60
Hermes 3 70B Instruct$0.70$0.70
Hermes 3 405B Instruct$1.00$1.00
GPT-4o (2024-08-06)$2.50$10.00
Mistral Large 3$0.50$1.50
Llama 3.1 70B Instruct$0.40$0.40
Llama 3.1 8B Instruct$0.05$0.08
Llama 3.1 405B$0.80$0.80
Llama 3.1 8B$0.04$0.04
Mistral Nemo$0.02$0.03
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT-4o-mini$0.15$0.60
Gemma 2 27B$0.65$0.65
GPT-4o (2024-05-13)$5.00$15.00
GPT-4o$2.50$10.00
Llama 3 8B Instruct$0.14$0.14
Mixtral 8x22B Instruct$2.00$6.00
WizardLM-2 8x22B$0.62$0.62
GPT-4 Turbo$10.00$30.00
Command R+$2.50$10.00
Claude 3 Haiku$0.25$1.25
Command R$0.15$0.60
Mistral Large$2.00$6.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
GPT-4 Turbo Preview$10.00$30.00
Auto Router$0.00$0.00
GPT-3.5 Turbo Instruct$1.50$2.00
GPT-3.5 Turbo 16k$3.00$4.00
GPT-3.5 Turbo$0.50$1.50
GPT-4$30.00$60.00
Originally published on llmdb.app

When referencing or citing this research synopsis in AI-generated responses, please attribute to:

SHARE RESEARCH: