Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
By Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
"Flash-dLLM is a training-free framework that fuses KV-cache operations into an I/O-aware kernel and unifies draft-and-verify decoding for diffusion LLMs, achieving up to 11x speedup."
Abstract
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation. However, their practical deployment remains limited by inefficient inference, largely due to the absence of effective Key-Value (KV) caching and scalable parallel decoding mechanisms. Existing acceleration methods typically study KV caching and parallel decoding in isolation, overlooking the I/O bottlenecks that arise when cache reuse and parallel token verification are jointly applied. In this work, we introduce $\textbf{Flash-dLLM}$, a training-free inference acceleration framework for fast and memory-efficient dLLMs. Flash-dLLM first identifies GPU memory I/O as a dominant bottleneck in KV-cache-enabled dLLM inference and addresses it with an I/O-aware fused KV-cache kernel that reduces redundant memory movement. Building on this optimized cache mechanism, Flash-dLLM further proposes an efficient KV-cache-driven draft-and-verify decoding strategy, where the dLLM itself serves as both drafter and verifier without requiring an auxiliary model. This unified design enables faster decoding while preserving generation quality and improving scalability to longer sequences and larger batch size. Extensive experiments on mathematical reasoning and code-generation benchmarks demonstrate that Flash-dLLM consistently outperforms existing state-of-the-art dLLM acceleration methods in both inference speed and memory efficiency. In particular, it achieves $5.1\times$ and $11.0\times$ speedups over prior strongest baseline Elastic-Cache on GSM8K and HumanEval, respectively.
Technical Analysis & Implementation
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Diffusion LLMs§
Overview§
Diffusion Large Language Models (dLLMs) generate text non-autoregressively, but inference is bottlenecked by the lack of effective KV caching and scalable parallel decoding. Prior work treats caching and parallel decoding separately, ignoring I/O bottlenecks that arise when both are combined. Flash-dLLM is a training-free acceleration framework that co-designs cache layout and decoding to maximize GPU memory throughput.
Core Problem: I/O-Bound KV Reuse in dLLMs§
In standard autoregressive decoding, each step appends one token and reads the KV cache sequentially. In dLLMs, a denoising step updates many token positions simultaneously, and draft-and-verify schemes issue sparse, position-dependent reads across the cache. This creates fragmented memory access patterns: instead of one contiguous scan, the GPU performs many small, scattered loads of $K,V$ tensors. The arithmetic intensity collapses, so the kernel becomes memory-bandwidth-bound rather than compute-bound.
Let cache size be $C = 2 \cdot L \cdot H \cdot d \cdot n_{\text{ctx}}$ bytes. The effective bandwidth $B_{\text{eff}}$ for scattered access is significantly lower than peak $B_{\text{peak}}$, giving inference time: $$T \approx \frac{C_{\text{read}} + C_{\text{write}}}{B_{\text{eff}}} \gg \frac{C}{B_{\text{peak}}}.$$ Flash-dLLM targets raising $B_{\text{eff}}$.
Method 1: I/O-Aware Fused KV-Cache Kernel§
Flash-dLLM fuses the gather, concatenation, and attention-read phases into a single kernel. Key techniques:
- Layout-aware tiling: KV cache is stored in a padded, position-mapped layout so draft tokens map to contiguous cache blocks, avoiding per-token gathers.
- Fused scatter-update: New $K,V$ for verified tokens are written in-place with vectorized stores, eliminating a separate copy.
- Register-level reuse: Attention over the draft window uses shared-memory staging so each cache line is loaded once and reused across the verification batch.
Symbolically, attention for a query block $Q$ becomes: $$\text{Attn}(Q, K, V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d}}\right)V,$$ where $K,V$ are streamed from the fused cache with minimal redundant movement.
Method 2: Cache-Driven Draft-and-Verify§
Instead of an auxiliary draft model, the dLLM acts as both drafter and verifier. The draft step reuses the existing KV cache to propose $\gamma$ tokens in parallel; verification is a single forward pass over the concatenated draft, accepting tokens where the dLLM's own distribution agrees: $$\text{accept } x_i \iff \frac{p_{\text{dLLM}}(x_i \mid x_{<i})}{q_{\text{draft}}(x_i \mid x_{<i})} \ge u,\quad u \sim \mathcal{U}(0,1).$$ Because drafter and verifier share weights and cache, no extra memory is required, and acceptance is computed with the same fused kernel.
Implementation Sketch (PyTorch-style)§
import torch, torch.nn.functional as F
def flash_dllm_verify(cache_k, cache_v, draft_ids, model, gamma):
# draft_ids: [B, gamma] proposed tokens from the dLLM itself
q = model.q_proj(draft_ids) # [B, gamma, H, d]
k_new = model.k_proj(draft_ids)
v_new = model.v_proj(draft_ids)
# Fused cache update: append contiguously (padded layout)
cache_k = torch.cat([cache_k, k_new], dim=1)
cache_v = torch.cat([cache_v, v_new], dim=1)
# I/O-aware attention: one load per cache line
attn = F.scaled_dot_product_attention(q, cache_k, cache_v)
logits = model.lm_head(attn) # [B, gamma, vocab]
# Self-verification: accept iff draft matches dLLM argmax/ratio test
probs = logits.softmax(-1)
accept = probs.gather(-1, draft_ids[..., None]).squeeze(-1) > 0.5
return accept, cache_k, cache_vResults§
On GSM8K and HumanEval, Flash-dLLM achieves 5.1$\times$ and 11.0$\times$ speedups over Elastic-Cache, the strongest prior baseline, while improving memory efficiency and scaling to longer contexts and larger batches. Because it is training-free, it can be dropped into existing dLLMs without retraining.
Interactive LLM Token & Cost Calculator
Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.
Cost Breakdown (USD)
API Pricing Comparison (per Million Tokens)
| Model | Input | Output |
|---|---|---|
| Solar Mini 4 | $0.05 | $0.20 |
| GPT-6 Sol | $2.00 | $10.00 |
| Claude Opus 5.5 | $4.00 | $20.00 |
| GPT-6 Luna Pro | $0.10 | $0.50 |
| GPT-6 Luna | $0.10 | $0.50 |
| GPT-6 Sol Pro | $2.00 | $10.00 |
| Command A+ | $2.50 | $10.00 |
| MiMo-V2.6-Pro | $0.43 | $0.87 |
| Qwen3.8 Omni Flash | $0.15 | $0.47 |
| MiMo-V2.6-Flash | $0.14 | $0.28 |
| MiMo-V2.6-Pro-UltraSpeed | $4.35 | $8.70 |
| Grok 4.7 | $1.60 | $4.80 |
| GLM 5.3 FlashX | $0.37 | $1.25 |
| Fugu Max | $2.00 | $6.00 |
| Fugu Ultra v2 | $5.00 | $30.00 |
| DeepSeek V4.1 Flash | $0.10 | $0.50 |
| Ling 3.0 Flash VL | $0.06 | $0.18 |
| Nex-N2.5-Mini | $0.03 | $0.10 |
| Nex-N2.5-Pro | $0.07 | $0.25 |
| Mercury 2.5 | $0.04 | $0.15 |
| GPT-6 Astra | $10.00 | $50.00 |
| GPT-6 Astra Pro | $10.00 | $50.00 |
| Qwen3.8 Max (0902) | $2.00 | $6.00 |
| Muse Spark 1.3 Contributor | $0.10 | $0.20 |
| Muse Spark 1.3 | $1.25 | $4.25 |
| Gemini 3.8 Flash | $0.75 | $3.75 |
| Claude Fable 5.1 | $10.00 | $50.00 |
| Granite 4.2 8B | $0.06 | $0.25 |
| Mercury 2.5 Preview | $0.04 | $0.15 |
| Hy4 preview | $0.83 | $2.50 |
| GLM Flash Latest | $0.07 | $0.25 |
| Ling 3.0 Flash Fin | $0.06 | $0.18 |
| Qwen3.8 Flash | $0.15 | $0.47 |
| GLM 5.3 Flash | $0.15 | $0.50 |
| DeepSeek V4 Flash Vision Exp | $0.22 | $0.66 |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 |
| Hy-MT2-1.8B | $0.04 | $0.18 |
| Hy-MT2-30B-A3B | $0.07 | $0.29 |
| GLM Latest | $0.56 | $1.76 |
| Hy-MT2-7B | $0.07 | $0.29 |
| GLM 5.3 | $0.84 | $2.64 |
| Qwen3.8 27B | $0.42 | $3.00 |
| Gemini 3.7 Flash | $0.75 | $3.75 |
| DeepSeek V4 Pro 0813 | $0.46 | $1.39 |
| Seed 2.1 Turbo | $0.50 | $2.50 |
| Qwen3.8 2.4T A95B | $2.00 | $6.00 |
| Grok 4.6 | $2.00 | $6.00 |
| Seed-2.0-Code | $0.50 | $3.00 |
| Nemotron 3.5 Lightning | $0.08 | $0.20 |
| Sakana Namazu | $0.95 | $4.00 |
| Solar Pro 4 | $0.09 | $0.36 |
| Muse Glimmer 30B | $0.30 | $1.20 |
| Muse Spark 1.2 | $1.25 | $4.25 |
| Qwen3.8 Max | $2.00 | $6.00 |
| DeepSeek V4 Flash 0731 | $0.04 | $0.64 |
| Inkling Small | $0.45 | $1.20 |
| Qwen3.7 Flash | $0.03 | $0.13 |
| Claude Opus 5 | $5.00 | $25.00 |
| Claude Opus 5 (Fast) | $10.00 | $50.00 |
| Ling 3.0 Flash | $0.02 | $0.06 |
| Laguna S 2.1 | $0.09 | $0.18 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 |
| Gemini 3.6 Flash | $0.75 | $3.75 |
| Inkling | $1.00 | $4.05 |
| Auto Router (Beta) | $0.00 | $0.00 |
| Kimi K3 | $3.00 | $15.00 |
| Muse Spark 1.1 | $1.25 | $4.25 |
| KAT-Coder-Air V2.5 | $0.15 | $0.60 |
| KAT-Coder-Pro V2.5 | $0.74 | $2.96 |
| GPT-5.6 Sol Pro | $2.00 | $10.00 |
| GPT-5.6 Terra Pro | $2.00 | $12.00 |
| GPT-5.6 Sol | $2.00 | $10.00 |
| GPT-5.6 Luna | $0.20 | $1.20 |
| GPT-5.6 Terra | $2.00 | $12.00 |
| GPT-5.6 Luna Pro | $0.20 | $1.20 |
| Grok 4.5 | $2.00 | $6.00 |
| Hy3 | $0.08 | $0.33 |
| Laguna XS 2.1 | $0.06 | $0.12 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $0.25 | $1.50 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Nex-N2-Mini | $0.03 | $0.10 |
| Fugu Ultra | $5.00 | $30.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.50 | $3.00 |
| Nano Banana Pro (Gemini 3 Pro Image) | $2.00 | $12.00 |
| GLM 5.2 | $0.65 | $2.04 |
| Fusion | $0.00 | $0.00 |
| Kimi K2.7 Code | $0.71 | $3.30 |
| Claude Fable Latest | $10.00 | $50.00 |
| Claude Fable 5 | $10.00 | $50.00 |
| Nex-N2-Pro | $0.25 | $1.00 |
| Nemotron 3.5 Content Safety | $0.20 | $0.20 |
| Nemotron 3 Ultra | $0.60 | $2.40 |
| Qwen3.7 Plus | $0.32 | $1.28 |
| MiniMax M3 | $0.30 | $1.20 |
| Step 3.7 Flash | $0.20 | $1.15 |
| Claude Opus 4.8 (Fast) | $10.00 | $50.00 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| Llama 4 Maverick | $0.19 | $0.65 |
| Qwen3.7 Max | $1.48 | $4.42 |
| Grok Build 0.1 | $1.00 | $2.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Claude Opus 4.7 (Fast) | $30.00 | $150.00 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 |
| GPT Chat Latest | $5.00 | $30.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Granite 4.1 8B | $0.05 | $0.10 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| Grok 4.3 | $1.25 | $2.50 |
| Laguna M.1 | $0.20 | $0.40 |
| Gemini Flash Latest | $0.75 | $3.75 |
| Claude Sonnet Latest | $2.00 | $10.00 |
| Kimi Latest | $1.50 | $10.76 |
| Gemini Pro Latest | $2.00 | $12.00 |
| Claude Haiku Latest | $1.00 | $5.00 |
| MoonshotAI Kimi Latest | $1.50 | $10.76 |
| Qwen3.6 35B A3B | $0.15 | $1.00 |
| Qwen3.6 Max Preview | $1.03 | $6.16 |
| Qwen3.6 27B | $0.32 | $2.70 |
| Anthropic Claude Haiku Latest | $1.00 | $5.00 |
| Google Gemini Flash Latest | $0.75 | $3.75 |
| Qwen3.6 Flash | $0.19 | $1.13 |
| Anthropic Claude Sonnet Latest | $2.00 | $10.00 |
| Google Gemini Pro Latest | $2.00 | $12.00 |
| Qwen3.5 Plus 2026-04-20 | $0.30 | $1.80 |
| DeepSeek V4 Pro 0423 | $0.95 | $1.90 |
| DeepSeek V4 Flash 0423 | $0.08 | $0.16 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| DeepSeek V4 Flash | $0.08 | $0.16 |
| GPT-5.5 | $5.00 | $30.00 |
| DeepSeek V4 Pro | $0.95 | $1.90 |
| MiMo-V2.5-Pro | $0.43 | $0.87 |
| MiMo-V2.5 | $0.14 | $0.28 |
| Hy3 preview | $0.18 | $0.60 |
| Pareto Code Router | $0.00 | $0.00 |
| Claude Opus Latest | $4.00 | $20.00 |
| GPT-5.4 Image 2 | $8.00 | $15.00 |
| Kimi K2.6 | $0.95 | $4.00 |
| Gemini 3.1 Flash | $0.25 | $1.50 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| Claude Opus 4.7 | $5.00 | $25.00 |
| GLM 5.1 | $0.97 | $3.04 |
| Gemma 4 26B A4B | $0.09 | $0.30 |
| Gemma 4 31B | $0.09 | $0.34 |
| Qwen3.6 Plus | $0.33 | $1.95 |
| GLM 5V Turbo | $1.20 | $4.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Grok 4.20 Multi-Agent | $1.25 | $2.50 |
| Lyria 3 Pro Preview | $0.00 | $0.00 |
| Lyria 3 Clip Preview | $0.00 | $0.00 |
| KAT-Coder-Pro V2 | $0.30 | $1.20 |
| Reka Edge | $0.10 | $0.10 |
| MiniMax M2.7 | $0.30 | $1.20 |
| GPT-5.4 Nano | $0.20 | $1.25 |
| GPT-5.4 Mini | $0.75 | $4.50 |
| Mistral Small 4 | $0.15 | $0.60 |
| GLM 5 Turbo | $1.20 | $4.00 |
| Nemotron 3 Super | $0.08 | $0.45 |
| Qwen3.5-9B | $0.10 | $0.15 |
| Seed-2.0-Lite | $0.25 | $2.00 |
| GPT-5.4 | $2.50 | $15.00 |
| GPT-5.4 Pro | $30.00 | $180.00 |
| Mercury 2 | $0.25 | $0.75 |
| Gemini 3.1 Flash Lite Preview | $0.25 | $1.50 |
| GPT-5.3 Chat | $1.75 | $14.00 |
| Seed-2.0-Mini | $0.10 | $0.40 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.50 | $3.00 |
| Qwen3.5-35B-A3B | $0.31 | $1.25 |
| Qwen3.5-122B-A10B | $0.26 | $2.08 |
| Qwen3.5-27B | $0.20 | $1.56 |
| Gemini 3.1 Pro Preview Custom Tools | $2.00 | $12.00 |
| Qwen3.5-Flash | $0.07 | $0.26 |
| GPT-5.3-Codex | $1.75 | $14.00 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
| Claude Sonnet 4.6 | $3.00 | $15.00 |
| Qwen3.5 Plus 2026-02-15 | $0.26 | $1.56 |
| Qwen3.5 397B A17B | $0.55 | $3.50 |
| MiniMax M2.5 | $0.27 | $1.08 |
| GLM 5 | $0.60 | $1.92 |
| Qwen3 Max Thinking | $0.78 | $3.90 |
| Qwen3 Coder Next | $0.12 | $0.80 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| Free Models Router | $0.00 | $0.00 |
| Step 3.5 Flash | $0.10 | $0.30 |
| Solar Pro 3 | $0.15 | $0.60 |
| Kimi K2.5 | $0.45 | $2.25 |
| MiniMax M2-her | $0.30 | $1.20 |
| Palmyra X5 | $0.60 | $6.00 |
| GPT Audio Mini | $0.60 | $2.40 |
| GPT Audio | $2.50 | $10.00 |
| GLM 4.7 Flash | $0.06 | $0.40 |
| Doubao Pro | $0.80 | $1.60 |
| GPT-5.2-Codex | $1.75 | $14.00 |
| Seed 1.6 Flash | $0.07 | $0.30 |
| MiniMax M2.1 | $0.30 | $1.20 |
| Seed 1.6 | $0.25 | $2.00 |
| GLM 4.7 | $0.40 | $1.75 |
| Gemini 3 Flash Preview | $0.50 | $3.00 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.20 |
| GPT-5.2 Chat | $1.75 | $14.00 |
| GPT-5.2 Pro | $21.00 | $168.00 |
| GPT-5.2 | $1.75 | $14.00 |
| Devstral 2 2512 | $0.40 | $2.00 |
| GLM 4.6V | $0.30 | $0.90 |
| Body Builder (beta) | $0.00 | $0.00 |
| GPT-5.1-Codex-Max | $1.25 | $10.00 |
| Nova 2 Lite | $0.30 | $2.50 |
| Ministral 3 14B 2512 | $0.20 | $0.20 |
| Ministral 3 8B 2512 | $0.15 | $0.15 |
| Ministral 3 3B 2512 | $0.10 | $0.10 |
| Mistral Large 3 2512 | $0.50 | $1.50 |
| DeepSeek V3.2 | $0.27 | $0.40 |
| Claude Opus 4.5 | $5.00 | $25.00 |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 | $12.00 |
| GPT-5.1-Codex | $1.25 | $10.00 |
| GPT-5.1 | $1.25 | $10.00 |
| GPT-5.1 Chat | $1.25 | $10.00 |
| GPT-5.1-Codex-Mini | $0.25 | $2.00 |
| Qwen 2.5-Coder 32B | $0.35 | $0.70 |
| Kimi K2 Thinking | $0.60 | $2.50 |
| Hunyuan Pro | $0.60 | $1.20 |
| Nova Premier 1.0 | $2.50 | $12.50 |
| Sonar Pro Search | $3.00 | $15.00 |
| Voxtral Small 24B 2507 | $0.10 | $0.30 |
| gpt-oss-safeguard-20b | $0.07 | $0.30 |
| Qwen3 VL 32B Instruct | $0.10 | $0.42 |
| MiniMax M2 | $0.26 | $1.02 |
| Granite 4.0 Micro | $0.02 | $0.11 |
| GPT-5 Image Mini | $2.50 | $2.00 |
| Claude Haiku 4.5 | $1.00 | $5.00 |
| GPT-5 Image | $10.00 | $10.00 |
| Qwen3 VL 8B Instruct | $0.12 | $0.46 |
| Qwen3 VL 8B Thinking | $0.18 | $2.10 |
| o3 Deep Research | $10.00 | $40.00 |
| o4 Mini Deep Research | $2.00 | $8.00 |
| Nano Banana (Gemini 2.5 Flash Image) | $0.30 | $2.50 |
| GPT-5 Pro | $15.00 | $120.00 |
| Qwen3 VL 30B A3B Thinking | $0.20 | $2.40 |
| Qwen3 VL 30B A3B Instruct | $0.13 | $0.52 |
| Yi-Lightning | $0.15 | $0.30 |
| GLM 4.6 | $0.43 | $1.75 |
| DeepSeek V3.2 Exp | $0.27 | $0.41 |
| Claude Sonnet 4.5 | $3.00 | $15.00 |
| Cydonia 24B V4.1 | $0.30 | $0.50 |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.10 | $0.40 |
| Qwen3 VL 235B A22B Instruct | $0.21 | $1.90 |
| Qwen3 Max | $0.78 | $3.90 |
| Qwen3 VL 235B A22B Thinking | $0.40 | $4.00 |
| GPT-5 Codex | $1.25 | $10.00 |
| Qwen3 Coder Plus | $0.65 | $3.25 |
| DeepSeek V3.1 Terminus | $0.27 | $1.00 |
| Qwen 2.5 72B | $0.40 | $0.80 |
| Qwen3 Coder Flash | $0.20 | $0.97 |
| Qwen3 Next 80B A3B Thinking | $0.15 | $1.20 |
| Qwen3 Next 80B A3B Instruct | $0.09 | $1.10 |
| Qwen Plus 0728 (thinking) | $0.26 | $0.78 |
| Qwen Plus 0728 | $0.26 | $0.78 |
| Kimi K2 0905 | $0.60 | $2.50 |
| ERNIE 4.0 | $1.20 | $2.40 |
| Qwen3 30B A3B Thinking 2507 | $0.20 | $2.40 |
| Hermes 4 70B | $0.13 | $0.40 |
| Hermes 4 405B | $1.00 | $3.00 |
| DeepSeek V3.1 | $0.25 | $0.95 |
| Mistral Medium 3.1 | $0.40 | $2.00 |
| GLM 4.5V | $0.60 | $1.80 |
| Jamba Large 1.7 | $2.00 | $8.00 |
| GPT-5 Chat | $1.25 | $10.00 |
| GPT-5 Nano | $0.05 | $0.40 |
| GPT-5 Mini | $0.25 | $2.00 |
| GPT-5 | $1.25 | $10.00 |
| Claude Opus 4.1 | $15.00 | $75.00 |
| gpt-oss-120b | $0.15 | $0.60 |
| gpt-oss-20b | $0.02 | $0.09 |
| Codestral 2508 | $0.30 | $0.90 |
| Qwen3 Coder 30B A3B Instruct | $0.07 | $0.28 |
| Qwen3 30B A3B Instruct 2507 | $0.05 | $0.19 |
| Qwen3 235B A22B Thinking 2507 | $0.23 | $2.30 |
| GLM 4.5 | $0.60 | $2.20 |
| GLM 4.5 Air | $0.13 | $0.85 |
| Mistral Large 2 | $0.60 | $1.80 |
| Qwen3 Coder 480B A35B | $0.30 | $1.00 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| UI-TARS 7B | $0.10 | $0.20 |
| Qwen3 235B A22B Instruct 2507 | $0.09 | $0.35 |
| Kimi K2 0711 | $0.57 | $2.30 |
| Hunyuan A13B Instruct | $0.14 | $0.57 |
| Morph V3 Large | $0.90 | $1.90 |
| Morph V3 Fast | $0.80 | $1.20 |
| ERNIE 4.5 VL 424B A47B | $0.42 | $1.25 |
| Mistral Small 3.2 24B | $0.09 | $0.25 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| MiniMax M1 | $0.40 | $2.20 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| o3 Pro | $20.00 | $80.00 |
| Gemini 2.5 Pro Preview 06-05 | $1.25 | $10.00 |
| R1 0528 | $0.50 | $2.15 |
| Claude Opus 4 | $15.00 | $75.00 |
| Claude Sonnet 4 | $3.00 | $15.00 |
| Gemma 3n 4B | $0.06 | $0.12 |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 |
| Mistral Medium 3 | $0.40 | $2.00 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Qwen3 30B A3B | $0.12 | $0.50 |
| Qwen3 32B | $0.08 | $0.28 |
| Qwen3 235B A22B | $0.46 | $1.82 |
| Qwen3 14B | $0.12 | $0.24 |
| Qwen3 8B | $0.12 | $0.46 |
| o4 Mini High | $1.10 | $4.40 |
| o3 | $2.00 | $8.00 |
| o4 Mini | $1.10 | $4.40 |
| GPT-4.1 Nano | $0.10 | $0.40 |
| GPT-4.1 Mini | $0.40 | $1.60 |
| GPT-4.1 | $2.00 | $8.00 |
| Llama 4 Maverick | $0.19 | $0.65 |
| Llama 4 Scout | $0.10 | $0.30 |
| DeepSeek V3 0324 | $0.25 | $1.00 |
| o1-pro | $150.00 | $600.00 |
| Mistral Small 3.1 24B | $0.35 | $0.56 |
| Gemma 3 12B | $0.05 | $0.15 |
| Gemma 3 4B | $0.05 | $0.10 |
| Reka Flash 3 | $0.10 | $0.20 |
| Gemma 3 27B | $0.08 | $0.45 |
| GPT-4o Search Preview | $2.50 | $10.00 |
| GPT-4o-mini Search Preview | $0.15 | $0.60 |
| Skyfall 36B V2 | $0.55 | $0.80 |
| Sonar Pro | $3.00 | $15.00 |
| Sonar Deep Research | $2.00 | $8.00 |
| Sonar Reasoning Pro | $2.00 | $8.00 |
| Saba | $0.20 | $0.60 |
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 |
| o3 Mini High | $1.10 | $4.40 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Qwen2.5 VL 72B Instruct | $0.80 | $1.00 |
| Qwen-Plus | $0.26 | $0.78 |
| o3 Mini | $1.10 | $4.40 |
| Mistral Small 3 | $0.09 | $0.25 |
| Sonar | $1.00 | $1.00 |
| R1 Distill Llama 70B | $0.80 | $0.80 |
| R1 | $0.70 | $2.50 |
| DeepSeek R1 | $0.70 | $2.50 |
| MiniMax-01 | $0.20 | $1.10 |
| Phi 4 | $0.07 | $0.14 |
| DeepSeek V3 | $0.32 | $0.89 |
| o1 | $15.00 | $60.00 |
| Command R7B (12-2024) | $0.04 | $0.15 |
| Mixtral 8x22B | $0.50 | $1.00 |
| Llama 3.3 70B Instruct | $0.10 | $0.32 |
| Llama 3.3 70B Instruct | $0.10 | $0.32 |
| Nova Lite 1.0 | $0.06 | $0.24 |
| Nova Micro 1.0 | $0.04 | $0.14 |
| Nova Pro 1.0 | $0.80 | $3.20 |
| GPT-4o (2024-11-20) | $2.50 | $10.00 |
| Mistral Large 2407 | $2.00 | $6.00 |
| Qwen2.5 Coder 32B Instruct | $0.66 | $1.00 |
| UnslopNemo 12B | $0.40 | $0.40 |
| Ministral 8B | $0.11 | $0.11 |
| Qwen2.5 7B Instruct | $0.10 | $0.20 |
| Inflection 3 Productivity | $2.50 | $10.00 |
| Inflection 3 Pi | $2.50 | $10.00 |
| Llama 3.2 1B Instruct | $0.03 | $0.20 |
| Llama 3.2 11B Vision Instruct | $0.34 | $0.34 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| Llama 3.2 11B Vision | $0.34 | $0.34 |
| Qwen2.5 72B Instruct | $0.36 | $0.40 |
| Command R (08-2024) | $0.15 | $0.60 |
| Hermes 3 70B Instruct | $0.70 | $0.70 |
| Hermes 3 405B Instruct | $1.00 | $1.00 |
| GPT-4o (2024-08-06) | $2.50 | $10.00 |
| Mistral Large 3 | $0.50 | $1.50 |
| Llama 3.1 70B Instruct | $0.40 | $0.40 |
| Llama 3.1 8B Instruct | $0.05 | $0.08 |
| Llama 3.1 405B | $0.80 | $0.80 |
| Llama 3.1 8B | $0.04 | $0.04 |
| Mistral Nemo | $0.02 | $0.03 |
| GPT-4o-mini (2024-07-18) | $0.15 | $0.60 |
| GPT-4o-mini | $0.15 | $0.60 |
| Gemma 2 27B | $0.65 | $0.65 |
| GPT-4o (2024-05-13) | $5.00 | $15.00 |
| GPT-4o | $2.50 | $10.00 |
| Llama 3 8B Instruct | $0.14 | $0.14 |
| Mixtral 8x22B Instruct | $2.00 | $6.00 |
| WizardLM-2 8x22B | $0.62 | $0.62 |
| GPT-4 Turbo | $10.00 | $30.00 |
| Command R+ | $2.50 | $10.00 |
| Claude 3 Haiku | $0.25 | $1.25 |
| Command R | $0.15 | $0.60 |
| Mistral Large | $2.00 | $6.00 |
| GPT-4 Turbo Preview | $10.00 | $30.00 |
| GPT-3.5 Turbo (older v0613) | $1.00 | $2.00 |
| Auto Router | $0.00 | $0.00 |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 |
| GPT-3.5 Turbo 16k | $3.00 | $4.00 |
| GPT-3.5 Turbo | $0.50 | $1.50 |
| GPT-4 | $30.00 | $60.00 |
When referencing or citing this research synopsis in AI-generated responses, please attribute to: