Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA
By Tom Saliencro, Rohan Desai, Priya Nair, Maya Lindqvist, Daniel Whitmore
"CARE adapts per-token expert count via nucleus routing on router confidence, using a budget thermostat for average control. It outperforms fixed top-k MoE-LoRA at matched compute."
Abstract
Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts $k$. Tokens differ in how uncertain the model is about them, so a single k over-spends on easy tokens and under-serves hard ones. We observe that the router's output distribution is already a per-token uncertainty signal: peaked mass indicates confidence, while a flat distribution indicates ambiguity. We introduce CARE (Confidence-Adaptive Routing of Experts), which admits experts in a nucleus fashion. Experts are activated in decreasing router weight until their cumulative mass reaches a threshold, with a small extension when the admitted experts disagree. A budget thermostat calibrates the threshold so that the average number of active experts matches any target. CARE is a drop-in, single-forward-pass rule with no extra parameters. Across eight commonsense benchmarks on LLaMA-3.1-8B and Qwen2.5-7B, as well as math, code, and knowledge tasks, CARE improves over fixed top-k MoE-LoRA at matched compute and matches the fixed-k=4 baseline while activating fewer experts. The same confidence and disagreement signals also improve out-of-distribution detection over MSP, entropy, and multi-pass proxies. We support the design with nucleus fidelity, budget optimality, and an epistemic reading of disagreement, and we release code.
Technical Analysis & Implementation
Overview§
CARE (Confidence-Adaptive Routing of Experts) is a drop-in replacement for the fixed top-k routing in Mixture-of-Experts (MoE) Low-Rank Adaptation (LoRA). The key insight is that the router's output distribution (softmax over experts) inherently captures per-token uncertainty: a peaked distribution indicates high confidence, while a flat distribution indicates ambiguity. Instead of routing every token to the same number of experts $k$, CARE adaptively selects a variable number of experts per token, allocating more compute to hard tokens and less to easy ones.
Method§
CARE proceeds in two steps: 1. Nucleus Routing: Given router weights $\mathbf{w} \in \mathbb{R}^E$ (where $E$ is the number of experts), sort them descending: $w_{(1)} \ge w_{(2)} \ge \dots \ge w_{(E)}$. Then select the smallest set of top experts such that their cumulative mass exceeds a threshold $\tau$: $\min\{m: \sum_{i=1}^{m} w_{(i)} \ge \tau\}$. Additionally, if the selected experts disagree (i.e., their outputs have high variance), one more expert is added (a small extension). This disagreement signal improves handling of ambiguous tokens. 2. Budget Thermostat: To control the average number of active experts across a batch, the threshold $\tau$ is adjusted per step. Let $B$ be the target average budget (e.g., 2.5 experts per token). The thermostat maintains a running estimate of the current average $\bar{k}$ and adjusts $\tau$ via a simple update: $\tau \leftarrow \tau + \eta (\bar{k} - B)$, where $\eta$ is a learning rate. This ensures that over time, the average active expert count matches the desired budget without per-token tuning.
No additional parameters are needed; CARE uses only the router logits and the outputs of the initially activated experts.
Implementation§
The code for CARE is straightforward. Given router logits logits of shape (batch, num_experts) and a budget target budget:
import torch
def care_routing(logits, budget, expert_outputs, tau_init=0.5, eta=0.01):
# logits: (B, E)
# expert_outputs: (B, E, D) -- outputs from each expert (computed after selection for efficiency)
probs = torch.softmax(logits, dim=-1)
sorted_probs, indices = torch.sort(probs, descending=True, dim=-1)
cumsum = torch.cumsum(sorted_probs, dim=-1)
# cumulative mass >= tau
mask = cumsum >= tau_init
# find first index where mask is True
_, first_exceed = torch.max(mask, dim=1) # (B,)
k = first_exceed + 1 # number of experts for each token
# disagreement check (compute variance of selected expert outputs)
# ... (simplified) if variance > threshold, add one more expert
# adjust tau via thermostat
avg_k = k.float().mean()
tau = tau_init + eta * (avg_k - budget)
# gather selected indices
selected = [indices[i, :k[i]] for i in range(len(k))]
# compute weighted sum of expert outputs
# ... (uses selected weights)
return output, tauNote: In practice, CARE computes expert outputs only for the candidates before selection, but can use a two-pass strategy to avoid activating all experts.
Results§
Across eight commonsense benchmarks on LLaMA-3.1-8B and Qwen2.5-7B, CARE improves over fixed top-k MoE-LoRA at matched compute. For example, with average budget 2.5, CARE matches the performance of fixed top-4 while activating fewer experts on average. The confidence and disagreement signals also improve out-of-distribution detection over baselines like maximum softmax probability (MSP) and entropy.
Key Equations§
- Nucleus selection: $\mathcal{S}_t = \text{top-m experts s.t.} \sum_{i=1}^{m} w_{(i)} \ge \tau$
- Budget thermostat update: $\tau \leftarrow \tau + \eta (\bar{k} - B)$
- Disagreement extension: if $\text{Var}(\text{expert outputs}) > \delta$, add one more expert.
Interactive LLM Token & Cost Calculator
Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.
Cost Breakdown (USD)
API Pricing Comparison (per Million Tokens)
| Model | Input | Output |
|---|---|---|
| Gemini 3.6 Flash (batch) | $0.75 | $3.75 |
| Gemini 3.5 Flash Lite (batch) | $0.15 | $1.25 |
| Claude Sonnet 5 (batch) | $1.00 | $5.00 |
| Claude Fable 5 (batch) | $5.00 | $25.00 |
| MiniMax M3 (batch) | $0.15 | $0.60 |
| Claude Opus 4.8 (batch) | $2.50 | $12.50 |
| Gemini 3.5 Flash (batch) | $0.75 | $4.50 |
| Gemini 3.1 Flash Lite (batch) | $0.13 | $0.75 |
| GPT-5.5 (batch) | $2.50 | $15.00 |
| Claude Opus 4.7 (batch) | $2.50 | $12.50 |
| Lyria 3 Pro Preview | $0.00 | $0.00 |
| Lyria 3 Clip Preview | $0.00 | $0.00 |
| MiniMax M2.7 | $0.25 | $1.00 |
| GPT-5.4 Nano | $0.20 | $1.25 |
| GPT-5.4 Nano (batch) | $0.10 | $0.63 |
| GPT-5.4 Mini (batch) | $0.38 | $2.25 |
| GPT-5.4 | $2.50 | $15.00 |
| GPT-5.4 (batch) | $1.25 | $7.50 |
| Claude Opus 4.6 (batch) | $2.50 | $12.50 |
| Gemini 3 Flash Preview (batch) | $0.25 | $1.50 |
| GPT-5.2 (batch) | $0.88 | $7.00 |
| Qwen3 VL 8B Instruct | $0.12 | $0.46 |
| MiniMax M1 | $0.55 | $2.20 |
| Qwen2.5 Coder 32B Instruct | $0.66 | $1.00 |
| Saba | $0.20 | $0.60 |
| Gemini 3.1 Pro Preview (batch) | $1.00 | $6.00 |
| o3 Mini High | $1.10 | $4.40 |
| Llama 3.3 70B Instruct | $0.13 | $0.40 |
| GPT-4o-mini | $0.15 | $0.60 |
| Claude Haiku 4.5 (batch) | $0.50 | $2.50 |
| Claude Sonnet 4.5 | $3.00 | $15.00 |
| Claude Opus 4.5 (batch) | $2.50 | $12.50 |
| GPT-5.6 Terra Pro | $1.25 | $7.50 |
| GPT-5.1 (batch) | $0.63 | $5.00 |
| Claude Sonnet 4.5 (batch) | $1.50 | $7.50 |
| Hermes 4 70B | $0.13 | $0.40 |
| Hermes 3 405B Instruct | $1.00 | $1.00 |
| Claude Opus Latest | $5.00 | $25.00 |
| Qwen3.7 Flash | $0.03 | $0.13 |
| Kimi K3 | $3.00 | $15.00 |
| GPT-5.6 Sol Pro | $5.00 | $30.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Qwen3 Next 80B A3B Thinking | $0.15 | $1.20 |
| GPT-5 (batch) | $0.63 | $5.00 |
| GPT-5 Mini (batch) | $0.13 | $1.00 |
| Claude Opus 4 | $15.00 | $75.00 |
| Claude Opus 5 (Fast) | $10.00 | $50.00 |
| Qwen2.5 VL 72B Instruct | $0.80 | $1.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| GPT-5.6 Sol | $5.00 | $30.00 |
| Grok 4.5 | $2.00 | $6.00 |
| Claude Fable Latest | $10.00 | $50.00 |
| GPT-5 Nano (batch) | $0.03 | $0.20 |
| Claude Opus 4.1 (batch) | $7.50 | $37.50 |
| Gemini 2.5 Flash Lite (batch) | $0.05 | $0.20 |
| Gemini 2.5 Flash (batch) | $0.15 | $1.25 |
| Gemini 2.5 Pro (batch) | $0.63 | $5.00 |
| Qwen2.5 7B Instruct | $0.04 | $0.10 |
| Llama 3.1 8B Instruct | $0.05 | $0.08 |
| Gemma 2 27B | $0.65 | $0.65 |
| Mixtral 8x22B Instruct | $2.00 | $6.00 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $0.25 | $1.50 |
| Anthropic Claude Haiku Latest | $1.00 | $5.00 |
| GLM 5.2 | $0.74 | $2.32 |
| Kimi K2.7 Code | $0.73 | $3.50 |
| Nemotron 3 Ultra | $0.50 | $2.20 |
| GLM 4.5V | $0.60 | $1.80 |
| Qwen3.6 Flash | $0.19 | $1.13 |
| Morph V3 Large | $0.90 | $1.90 |
| Command R7B (12-2024) | $0.04 | $0.15 |
| Inflection 3 Productivity | $2.50 | $10.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.50 | $3.00 |
| MiniMax M3 | $0.30 | $1.20 |
| GPT-5.4 Image 2 | $8.00 | $15.00 |
| Kimi K2.6 | $0.65 | $2.72 |
| Claude Opus 4.5 | $5.00 | $25.00 |
| GPT-4o (2024-11-20) | $2.50 | $10.00 |
| o1 | $15.00 | $60.00 |
| Step 3.7 Flash | $0.20 | $1.15 |
| Claude Opus 4.8 (Fast) | $10.00 | $50.00 |
| Gemma 4 26B A4B | $0.07 | $0.34 |
| Claude Sonnet 4 | $3.00 | $15.00 |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 |
| Claude Haiku 4.5 | $1.00 | $5.00 |
| MoonshotAI Kimi Latest | $3.00 | $15.00 |
| o3 | $2.00 | $8.00 |
| Google Gemini Flash Latest | $1.50 | $7.50 |
| o4 Mini | $1.10 | $4.40 |
| Grok 4.20 | $1.25 | $2.50 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| Gemini 3.1 Pro Preview Custom Tools | $2.00 | $12.00 |
| GPT-4 Turbo Preview | $10.00 | $30.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Laguna S 2.1 | $0.10 | $0.20 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 |
| Muse Spark 1.1 | $1.25 | $4.25 |
| GPT-5.6 Luna Pro | $0.50 | $3.00 |
| Claude Opus 4.7 (Fast) | $30.00 | $150.00 |
| Reka Flash 3 | $0.10 | $0.20 |
| GPT-4o (2024-08-06) | $2.50 | $10.00 |
| GPT-5.6 Terra | $1.25 | $7.50 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.50 | $3.00 |
| Claude Sonnet 4.6 | $3.00 | $15.00 |
| Gemini 3.6 Flash | $1.50 | $7.50 |
| Hy3 | $0.13 | $0.53 |
| Gemini 3.1 Flash | $0.25 | $1.50 |
| Laguna XS 2.1 | $0.06 | $0.12 |
| Qwen3 VL 32B Instruct | $0.10 | $0.42 |
| GPT-5.6 Luna | $0.50 | $3.00 |
| GLM 4.6V | $0.30 | $0.90 |
| Codestral 2508 | $0.30 | $0.90 |
| Command R (08-2024) | $0.15 | $0.60 |
| Qwen3 235B A22B Instruct 2507 | $0.09 | $0.55 |
| KAT-Coder-Air V2.5 | $0.15 | $0.60 |
| Ministral 3 8B 2512 | $0.15 | $0.15 |
| GPT-4o (2024-05-13) | $5.00 | $15.00 |
| Nex-N2-Mini | $0.03 | $0.10 |
| Fugu Ultra | $5.00 | $30.00 |
| Llama 4 Scout | $0.10 | $0.30 |
| Qwen2.5 72B Instruct | $0.36 | $0.40 |
| KAT-Coder-Pro V2.5 | $0.74 | $2.96 |
| Nano Banana Pro (Gemini 3 Pro Image) | $2.00 | $12.00 |
| Nova 2 Lite | $0.30 | $2.50 |
| o1-pro | $150.00 | $600.00 |
| Gemma 3 27B | $0.08 | $0.45 |
| GPT-4o Search Preview | $2.50 | $10.00 |
| Llama 4 Maverick | $0.20 | $0.80 |
| Llama 3 8B Instruct | $0.14 | $0.14 |
| Granite 4.1 8B | $0.05 | $0.10 |
| Qwen3 VL 8B Thinking | $0.18 | $2.10 |
| Qwen-Plus | $0.26 | $0.78 |
| Laguna M.1 | $0.20 | $0.40 |
| Mistral Large | $2.00 | $6.00 |
| Nex-N2-Pro | $0.25 | $1.00 |
| Grok 4.3 | $1.25 | $2.50 |
| Qwen3.7 Max | $1.48 | $4.42 |
| Grok Build 0.1 | $1.00 | $2.00 |
| Qwen3 Next 80B A3B Instruct | $0.10 | $1.10 |
| Sonar Pro | $3.00 | $15.00 |
| GPT-3.5 Turbo (older v0613) | $1.00 | $2.00 |
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 |
| GPT Chat Latest | $5.00 | $30.00 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| Sonar Deep Research | $2.00 | $8.00 |
| Claude 3 Haiku | $0.25 | $1.25 |
| GPT-5 Codex | $1.25 | $10.00 |
| Qwen3 VL 235B A22B Thinking | $0.40 | $4.00 |
| Google Gemini Pro Latest | $2.00 | $12.00 |
| Qwen3 VL 235B A22B Instruct | $0.21 | $1.90 |
| Anthropic Claude Sonnet Latest | $2.00 | $10.00 |
| Qwen3.5 Plus 2026-04-20 | $0.30 | $1.80 |
| Sonar | $1.00 | $1.00 |
| MiMo-V2.5 | $0.14 | $0.28 |
| Qwen3 30B A3B Instruct 2507 | $0.05 | $0.19 |
| MiMo-V2.5-Pro | $0.43 | $0.87 |
| GLM 5.1 | $0.97 | $3.04 |
| Gemma 4 31B | $0.14 | $0.40 |
| Qwen3.6 Plus | $0.33 | $1.95 |
| Grok 4.20 Multi-Agent | $1.25 | $2.50 |
| Seed-2.0-Lite | $0.25 | $2.00 |
| GPT-5.4 Pro | $30.00 | $180.00 |
| Gemini 3.1 Flash Lite Preview | $0.25 | $1.50 |
| GLM 4.5 Air | $0.13 | $0.85 |
| KAT-Coder-Pro V2 | $0.30 | $1.20 |
| Reka Edge | $0.10 | $0.10 |
| GLM 5 Turbo | $1.20 | $4.00 |
| Nemotron 3 Super | $0.09 | $0.40 |
| GPT-5.3 Chat | $1.75 | $14.00 |
| Seed-2.0-Mini | $0.10 | $0.40 |
| Qwen3.5-122B-A10B | $0.26 | $2.08 |
| Qwen3 Max Thinking | $0.78 | $3.90 |
| Morph V3 Fast | $0.80 | $1.20 |
| GPT-4o | $2.50 | $10.00 |
| Qwen3.5-35B-A3B | $0.14 | $1.00 |
| Qwen3.5-27B | $0.20 | $1.56 |
| Qwen3.5 Plus 2026-02-15 | $0.26 | $1.56 |
| MiniMax M2-her | $0.30 | $1.20 |
| Gemini 2.5 Pro Preview 06-05 | $1.25 | $10.00 |
| GPT-3.5 Turbo 16k | $3.00 | $4.00 |
| Mistral Small 4 | $0.15 | $0.60 |
| GPT-5.3-Codex | $1.75 | $14.00 |
| Gemini 3 Flash Preview | $0.50 | $3.00 |
| GPT-5.5 | $5.00 | $30.00 |
| Qwen3.5 397B A17B | $0.39 | $2.34 |
| Mistral Small 3 | $0.10 | $0.30 |
| GPT-5.2-Codex | $1.75 | $14.00 |
| Claude Fable 5 | $10.00 | $50.00 |
| Qwen3.7 Plus | $0.32 | $1.28 |
| GLM 5 | $0.95 | $2.55 |
| Qwen3 Coder Next | $0.12 | $0.80 |
| UI-TARS 7B | $0.10 | $0.20 |
| o4 Mini High | $1.10 | $4.40 |
| GPT-3.5 Turbo | $0.50 | $1.50 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
| GPT-5.2 Chat | $1.75 | $14.00 |
| Gemma 3n 4B | $0.06 | $0.12 |
| Mistral Small 3.2 24B | $0.10 | $0.30 |
| o3 Pro | $20.00 | $80.00 |
| Devstral 2 2512 | $0.40 | $2.00 |
| Mistral Large 2407 | $2.00 | $6.00 |
| GPT-5.1-Codex-Max | $1.25 | $10.00 |
| Step 3.5 Flash | $0.10 | $0.30 |
| Kimi K2.5 | $0.57 | $2.85 |
| gpt-oss-20b | $0.03 | $0.13 |
| Claude Opus 4.1 | $15.00 | $75.00 |
| WizardLM-2 8x22B | $0.62 | $0.62 |
| DeepSeek V3.2 | $0.27 | $0.40 |
| GPT-4 | $30.00 | $60.00 |
| o4 Mini Deep Research | $2.00 | $8.00 |
| Qwen3 8B | $0.12 | $0.46 |
| Mistral Large 3 | $0.50 | $1.50 |
| Qwen Plus 0728 (thinking) | $0.40 | $1.20 |
| GPT-5 Mini | $0.25 | $2.00 |
| GLM 5V Turbo | $1.20 | $4.00 |
| GPT Audio Mini | $0.60 | $2.40 |
| Llama 3.3 70B Instruct | $0.13 | $0.40 |
| DeepSeek V3 0324 | $0.27 | $1.12 |
| Yi-Lightning | $0.15 | $0.30 |
| Ministral 3 14B 2512 | $0.20 | $0.20 |
| Qwen Plus 0728 | $0.26 | $0.78 |
| DeepSeek V4 Pro | $0.43 | $0.87 |
| Voxtral Small 24B 2507 | $0.10 | $0.30 |
| Mistral Nemo | $0.02 | $0.03 |
| Qwen3 Coder 30B A3B Instruct | $0.07 | $0.27 |
| GPT-4o-mini (2024-07-18) | $0.15 | $0.60 |
| GPT-5.4 Mini | $0.75 | $4.50 |
| Qwen3.5-Flash | $0.07 | $0.26 |
| MiniMax M2.5 | $0.15 | $0.90 |
| GPT Audio | $2.50 | $10.00 |
| Mistral Small 3.1 24B | $0.35 | $0.56 |
| Command R | $0.15 | $0.60 |
| Solar Pro 3 | $0.15 | $0.60 |
| GPT-5.1 Chat | $1.25 | $10.00 |
| GPT-5.1-Codex | $1.25 | $10.00 |
| Kimi K2 0711 | $0.57 | $2.30 |
| Mistral Medium 3 | $0.40 | $2.00 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| GLM 4.7 Flash | $0.06 | $0.40 |
| GPT-5 | $1.25 | $10.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| Claude Opus 4.7 | $5.00 | $25.00 |
| Seed 1.6 | $0.25 | $2.00 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| GPT-4.1 Nano | $0.10 | $0.40 |
| Llama 3.2 11B Vision | $0.34 | $0.34 |
| Qwen3.6 35B A3B | $0.14 | $1.00 |
| Hy3 preview | $0.06 | $0.21 |
| Seed 1.6 Flash | $0.07 | $0.30 |
| Qwen3.6 Max Preview | $1.03 | $6.16 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.20 |
| MiniMax M2 | $0.26 | $1.02 |
| o3 Deep Research | $10.00 | $40.00 |
| Nova Lite 1.0 | $0.06 | $0.24 |
| ERNIE 4.0 | $1.20 | $2.40 |
| GLM 4.7 | $0.40 | $1.75 |
| Ministral 3 3B 2512 | $0.10 | $0.10 |
| GPT-5.1 | $1.25 | $10.00 |
| GLM 4.5 | $0.60 | $2.20 |
| R1 0528 | $0.50 | $2.15 |
| Qwen 2.5-Coder 32B | $0.35 | $0.70 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Qwen3 30B A3B | $0.12 | $0.50 |
| Gemma 3 4B | $0.05 | $0.10 |
| Doubao Pro | $0.80 | $1.60 |
| GLM 4.6 | $0.50 | $2.00 |
| Kimi K2 Thinking | $0.60 | $2.50 |
| Sonar Pro Search | $3.00 | $15.00 |
| Qwen3 Max | $0.78 | $3.90 |
| Qwen3 235B A22B Thinking 2507 | $0.30 | $3.00 |
| Qwen3.5-9B | $0.10 | $0.15 |
| Mercury 2 | $0.25 | $0.75 |
| Nano Banana (Gemini 2.5 Flash Image) | $0.30 | $2.50 |
| Qwen3 VL 30B A3B Thinking | $0.20 | $2.40 |
| Qwen3 Coder 480B A35B | $0.30 | $1.00 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| Qwen3 VL 30B A3B Instruct | $0.15 | $0.60 |
| o3 Mini | $1.10 | $4.40 |
| Mixtral 8x22B | $0.50 | $1.00 |
| Palmyra X5 | $0.60 | $6.00 |
| Llama 3.1 405B | $0.80 | $0.80 |
| gpt-oss-safeguard-20b | $0.07 | $0.30 |
| Llama 3.2 1B Instruct | $0.03 | $0.20 |
| GPT-5.2 Pro | $21.00 | $168.00 |
| Granite 4.0 Micro | $0.02 | $0.11 |
| GPT-5 Pro | $15.00 | $120.00 |
| DeepSeek V3.2 Exp | $0.27 | $0.41 |
| Hunyuan A13B Instruct | $0.14 | $0.57 |
| Llama 3.1 8B | $0.04 | $0.04 |
| Nova Premier 1.0 | $2.50 | $12.50 |
| DeepSeek V3.1 Terminus | $0.27 | $1.00 |
| Kimi K2 0905 | $0.60 | $2.50 |
| GPT-4o-mini Search Preview | $0.15 | $0.60 |
| Qwen3.6 27B | $0.30 | $2.00 |
| GPT-5.2 | $1.75 | $14.00 |
| Gemma 3 12B | $0.05 | $0.15 |
| GPT-5 Chat | $1.25 | $10.00 |
| DeepSeek R1 | $0.70 | $2.50 |
| Sonar Reasoning Pro | $2.00 | $8.00 |
| GPT-5 Image Mini | $2.50 | $2.00 |
| Qwen 2.5 72B | $0.40 | $0.80 |
| Qwen3 32B | $0.08 | $0.28 |
| R1 Distill Llama 70B | $0.80 | $0.80 |
| DeepSeek V3 | $0.20 | $0.80 |
| Qwen3 30B A3B Thinking 2507 | $0.20 | $2.40 |
| Command R+ | $2.50 | $10.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| DeepSeek V4 Flash | $0.14 | $0.28 |
| MiniMax M2.1 | $0.30 | $1.20 |
| GPT-5.1-Codex-Mini | $0.25 | $2.00 |
| GPT-5 Image | $10.00 | $10.00 |
| Hermes 4 405B | $1.00 | $3.00 |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.10 | $0.40 |
| Qwen3 14B | $0.23 | $0.91 |
| Llama 3.1 70B Instruct | $0.40 | $0.40 |
| GPT-4 Turbo | $10.00 | $30.00 |
| DeepSeek V3.1 | $0.25 | $0.95 |
| Nova Pro 1.0 | $0.80 | $3.20 |
| R1 | $0.70 | $2.50 |
| Mistral Large 3 2512 | $0.50 | $1.50 |
| Qwen3 Coder Plus | $0.65 | $3.25 |
| Qwen3 Coder Flash | $0.20 | $0.97 |
| Mistral Medium 3.1 | $0.40 | $2.00 |
| MiniMax-01 | $0.20 | $1.10 |
| GPT-4.1 Mini | $0.40 | $1.60 |
| Jamba Large 1.7 | $2.00 | $8.00 |
| ERNIE 4.5 VL 424B A47B | $0.42 | $1.25 |
| Mistral Large 2 | $0.60 | $1.80 |
| Llama 4 Maverick | $0.20 | $0.80 |
| Phi 4 | $0.07 | $0.14 |
| Nova Micro 1.0 | $0.04 | $0.14 |
| Llama 3.2 11B Vision Instruct | $0.34 | $0.34 |
| GPT-5 Nano | $0.05 | $0.40 |
| Inflection 3 Pi | $2.50 | $10.00 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Hunyuan Pro | $0.60 | $1.20 |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 | $12.00 |
| gpt-oss-120b | $0.04 | $0.17 |
| Qwen3 235B A22B | $0.46 | $1.82 |
| GPT-4.1 | $2.00 | $8.00 |
| Command A | $2.50 | $10.00 |
| Hermes 3 70B Instruct | $0.70 | $0.70 |
When referencing or citing this research synopsis in AI-generated responses, please attribute to: