Pass the Baton: Trajectory-Relayed On-Policy Distillation
By Haolei Xu, Xiaowen Xu, Haiwen Hong, Zixuan Ni, Hongxing Li, Yiwen Qiu, Weiming Lu, Yongliang Shen
"Relay-OPD detects when a student model deviates during on-policy distillation and hands off to the teacher for a short correction segment, reducing waste and improving performance by over 5%."
Abstract
On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute. We identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction, and convert it into a label-free handoff trigger in Relay On-Policy Distillation (Relay-OPD). During training, Relay-OPD constructs relay trajectories by letting the teacher briefly take over at detected trigger points to produce a teacher leg, after which the student resumes and is optimized on the resulting trajectory. A limited relay budget concentrates intervention on critical early positions while limiting departure from the student policy. With a Qwen3-4B-Instruct-2507 teacher and Qwen3-0.6B/1.7B-Non-Thinking students on eight mathematical reasoning benchmarks, Relay-OPD achieves the best or second-best results on every benchmark, outperforming standard OPD by +5.73% and the strongest baseline FastOPD by +1.49% on average for 1.7B, with consistent gains at 0.6B. Training trajectory length is reduced by over 50%.
Technical Analysis & Implementation
Core Methodology§
Relay On-Policy Distillation (Relay-OPD) addresses the prefix failure problem in on-policy distillation (OPD), where a student's early wrong token choices cause cascading errors. The key insight is teacher-student continuation asymmetry: given a prefix where the student has already committed an error, the teacher tends to redirect (change topic or reasoning step) while the student continues along the wrong path. This asymmetry serves as a label-free trigger for intervention.
Algorithm Overview§
During training: 1. Generate student trajectory: Sample a sequence from the student policy $\pi_s$ until a termination condition (e.g., max length, EOS). 2. Detect trigger points: Compare the student's next token logits at each step to those of the teacher $\pi_t$. A trigger is flagged when the student's top token differs from the teacher's top token and the student's confidence for its top token is high while teacher's confidence is low (empirically measured via normalized probabilities). Formally: $$\mathbb{1}_{\text{trigger}}(t) = \left[\arg\max \pi_s(\cdot|x_{<t}) \neq \arg\max \pi_t(\cdot|x_{<t})\right] \wedge \left[\pi_s(\hat{x}_t|x_{<t}) > \tau_s \right] \wedge \left[\pi_t(\hat{x}_t|x_{<t}) < \tau_t\right]$$ where $\hat{x}_t$ is student's top token, $\tau_s$ and $\tau_t$ are thresholds. 3. Relay intervention: At the first trigger point (or earliest in a limited budget), switch generation to the teacher for a fixed small number of steps $K$ (e.g., 2–4 tokens), producing a teacher leg. Then the student resumes from where the teacher left off, continuing the trajectory. 4. Train on relay trajectory: Use the entire trajectory (student prefix + teacher leg + student continuation) as the on-policy data. The student is trained with token-level KL divergence against the teacher on its own tokens, but only on 'student' portions (not on teacher leg). The loss is: $$\mathcal{L}_{\text{Relay-OPD}} = \sum_{t \in \text{student tokens}} D_{KL}(\pi_t(\cdot|x_{<t}) \,||\, \pi_s(\cdot|x_{<t}))$$
Relay Budget§
To avoid excessive deviation from the student's policy, a relay budget $B$ limits the total number of teacher tokens inserted per trajectory. $B$ is typically small (e.g., 10 tokens total) and spread across early positions to maximize effect. This concentrates intervention where prefix failures are most harmful.
Implementation Details§
- Models: Teacher is Qwen3-4B-Instruct-2507; students are Qwen3-0.6B and 1.7B (non-thinking variant, which lacks explicit chain-of-thought)
- Datasets: 8 mathematical reasoning benchmarks (GSM8K, MATH, etc.)
- Hyperparameters: Thresholds $\tau_s=0.7$, $\tau_t=0.3$; teacher leg length $K=3$; relay budget $B=9$.
- Training: Standard supervised fine-tuning setup with batch size 128, learning rate 5e-5, cosine schedule.
Pseudocode / PyTorch Snippet§
import torch
import torch.nn.functional as F
def relay_opd_step(student, teacher, input_ids, max_len, budget, K, tau_s, tau_t):
# input_ids: initial prompt (batch, seq_len)
student.eval() # for generation, student is trained afterwards
teacher.eval()
batch_size = input_ids.shape[0]
current_ids = input_ids
trigger_count = torch.zeros(batch_size, dtype=torch.long)
teacher_leg_active = torch.zeros(batch_size, dtype=torch.bool)
while current_ids.shape[1] < max_len:
# Forward student and teacher
with torch.no_grad():
s_logits = student(current_ids).logits[:, -1, :] # batch x vocab
t_logits = teacher(current_ids).logits[:, -1, :]
s_probs = F.softmax(s_logits, dim=-1)
t_probs = F.softmax(t_logits, dim=-1)
s_top_token = s_probs.argmax(dim=-1)
t_top_token = t_probs.argmax(dim=-1)
# Determine trigger: student top != teacher top AND student confidence high AND teacher confidence low
mask_match = (s_top_token != t_top_token)
mask_s_high = (s_probs.max(dim=-1).values > tau_s)
mask_t_low = (t_probs.gather(1, s_top_token.unsqueeze(1)).squeeze() < tau_t)
trigger = mask_match & mask_s_high & mask_t_low & ~teacher_leg_active
# For triggered samples, start teacher leg if budget remains
can_trigger = trigger_count < budget
start_leg = trigger & can_trigger
# Update trigger count for those starting leg
trigger_count[start_leg] = trigger_count[start_leg] + K
teacher_leg_active[start_leg] = True
# Sample next token: if teacher leg active, use teacher; else use student
if teacher_leg_active.any():
# For teacher leg, sample from teacher logits
t_probs_for_sample = F.softmax(t_logits / 0.7, dim=-1) # temperature 0.7
next_token = torch.multinomial(t_probs_for_sample, 1).squeeze(1)
# Decrement leg counter (a separate counter for each sample, omitted for brevity)
else:
s_probs_for_sample = F.softmax(s_logits / 1.0, dim=-1)
next_token = torch.multinomial(s_probs_for_sample, 1).squeeze(1)
current_ids = torch.cat([current_ids, next_token.unsqueeze(1)], dim=1)
# Update teacher_leg_active: decrement counter per sample (not shown)
# Now train on all student tokens (not teacher leg) with KL div
# Separate trajectory into student and teacher segments; compute loss only on student tokens.
# ... (omitted for brevity)
return lossResults§
Relay-OPD achieves state-of-the-art results on all 8 benchmarks. For the 1.7B student, average accuracy improves by +5.73% over standard OPD and +1.49% over FastOPD. Training trajectory length reduces by over 50%, demonstrating both efficiency and effectiveness.
The method is particularly effective on harder reasoning tasks where prefix failures are more frequent, and the relay budget concentrates intervention on early tokens, minimizing deviation from the student's policy.
Interactive LLM Token & Cost Calculator
Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.
Cost Breakdown (USD)
API Pricing Comparison (per Million Tokens)
| Model | Input | Output |
|---|---|---|
| Gemini 3.6 Flash (batch) | $0.75 | $3.75 |
| Gemini 3.5 Flash Lite (batch) | $0.15 | $1.25 |
| Claude Sonnet 5 (batch) | $1.00 | $5.00 |
| Claude Fable 5 (batch) | $5.00 | $25.00 |
| MiniMax M3 (batch) | $0.15 | $0.60 |
| Claude Opus 4.8 (batch) | $2.50 | $12.50 |
| Gemini 3.5 Flash (batch) | $0.75 | $4.50 |
| Gemini 3.1 Flash Lite (batch) | $0.13 | $0.75 |
| GPT-5.5 (batch) | $2.50 | $15.00 |
| Claude Opus 4.7 (batch) | $2.50 | $12.50 |
| Lyria 3 Pro Preview | $0.00 | $0.00 |
| Lyria 3 Clip Preview | $0.00 | $0.00 |
| MiniMax M2.7 | $0.25 | $1.00 |
| GPT-5.4 Nano | $0.20 | $1.25 |
| GPT-5.4 Nano (batch) | $0.10 | $0.63 |
| GPT-5.4 Mini (batch) | $0.38 | $2.25 |
| GPT-5.4 | $2.50 | $15.00 |
| GPT-5.4 (batch) | $1.25 | $7.50 |
| Claude Opus 4.6 (batch) | $2.50 | $12.50 |
| Gemini 3 Flash Preview (batch) | $0.25 | $1.50 |
| GPT-5.2 (batch) | $0.88 | $7.00 |
| Qwen3 VL 8B Instruct | $0.12 | $0.46 |
| MiniMax M1 | $0.55 | $2.20 |
| Qwen2.5 Coder 32B Instruct | $0.66 | $1.00 |
| Saba | $0.20 | $0.60 |
| Gemini 3.1 Pro Preview (batch) | $1.00 | $6.00 |
| o3 Mini High | $1.10 | $4.40 |
| Llama 3.3 70B Instruct | $0.13 | $0.40 |
| GPT-4o-mini | $0.15 | $0.60 |
| Claude Haiku 4.5 (batch) | $0.50 | $2.50 |
| Claude Sonnet 4.5 | $3.00 | $15.00 |
| Claude Opus 4.5 (batch) | $2.50 | $12.50 |
| GPT-5.6 Terra Pro | $1.25 | $7.50 |
| GPT-5.1 (batch) | $0.63 | $5.00 |
| Claude Sonnet 4.5 (batch) | $1.50 | $7.50 |
| Hermes 4 70B | $0.13 | $0.40 |
| Hermes 3 405B Instruct | $1.00 | $1.00 |
| Claude Opus Latest | $5.00 | $25.00 |
| Qwen3.7 Flash | $0.03 | $0.13 |
| Kimi K3 | $3.00 | $15.00 |
| GPT-5.6 Sol Pro | $5.00 | $30.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Qwen3 Next 80B A3B Thinking | $0.15 | $1.20 |
| GPT-5 (batch) | $0.63 | $5.00 |
| GPT-5 Mini (batch) | $0.13 | $1.00 |
| Claude Opus 4 | $15.00 | $75.00 |
| Claude Opus 5 (Fast) | $10.00 | $50.00 |
| Qwen2.5 VL 72B Instruct | $0.80 | $1.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| GPT-5.6 Sol | $5.00 | $30.00 |
| Grok 4.5 | $2.00 | $6.00 |
| Claude Fable Latest | $10.00 | $50.00 |
| GPT-5 Nano (batch) | $0.03 | $0.20 |
| Claude Opus 4.1 (batch) | $7.50 | $37.50 |
| Gemini 2.5 Flash Lite (batch) | $0.05 | $0.20 |
| Gemini 2.5 Flash (batch) | $0.15 | $1.25 |
| Gemini 2.5 Pro (batch) | $0.63 | $5.00 |
| Qwen2.5 7B Instruct | $0.04 | $0.10 |
| Llama 3.1 8B Instruct | $0.05 | $0.08 |
| Gemma 2 27B | $0.65 | $0.65 |
| Mixtral 8x22B Instruct | $2.00 | $6.00 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $0.25 | $1.50 |
| Anthropic Claude Haiku Latest | $1.00 | $5.00 |
| GLM 5.2 | $0.74 | $2.32 |
| Kimi K2.7 Code | $0.73 | $3.50 |
| Nemotron 3 Ultra | $0.50 | $2.20 |
| GLM 4.5V | $0.60 | $1.80 |
| Qwen3.6 Flash | $0.19 | $1.13 |
| Morph V3 Large | $0.90 | $1.90 |
| Command R7B (12-2024) | $0.04 | $0.15 |
| Inflection 3 Productivity | $2.50 | $10.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.50 | $3.00 |
| MiniMax M3 | $0.30 | $1.20 |
| GPT-5.4 Image 2 | $8.00 | $15.00 |
| Kimi K2.6 | $0.65 | $2.72 |
| Claude Opus 4.5 | $5.00 | $25.00 |
| GPT-4o (2024-11-20) | $2.50 | $10.00 |
| o1 | $15.00 | $60.00 |
| Step 3.7 Flash | $0.20 | $1.15 |
| Claude Opus 4.8 (Fast) | $10.00 | $50.00 |
| Gemma 4 26B A4B | $0.07 | $0.34 |
| Claude Sonnet 4 | $3.00 | $15.00 |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 |
| Claude Haiku 4.5 | $1.00 | $5.00 |
| MoonshotAI Kimi Latest | $3.00 | $15.00 |
| o3 | $2.00 | $8.00 |
| Google Gemini Flash Latest | $1.50 | $7.50 |
| o4 Mini | $1.10 | $4.40 |
| Grok 4.20 | $1.25 | $2.50 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| Gemini 3.1 Pro Preview Custom Tools | $2.00 | $12.00 |
| GPT-4 Turbo Preview | $10.00 | $30.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Laguna S 2.1 | $0.10 | $0.20 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 |
| Muse Spark 1.1 | $1.25 | $4.25 |
| GPT-5.6 Luna Pro | $0.50 | $3.00 |
| Claude Opus 4.7 (Fast) | $30.00 | $150.00 |
| Reka Flash 3 | $0.10 | $0.20 |
| GPT-4o (2024-08-06) | $2.50 | $10.00 |
| GPT-5.6 Terra | $1.25 | $7.50 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.50 | $3.00 |
| Claude Sonnet 4.6 | $3.00 | $15.00 |
| Gemini 3.6 Flash | $1.50 | $7.50 |
| Hy3 | $0.13 | $0.53 |
| Gemini 3.1 Flash | $0.25 | $1.50 |
| Laguna XS 2.1 | $0.06 | $0.12 |
| Qwen3 VL 32B Instruct | $0.10 | $0.42 |
| GPT-5.6 Luna | $0.50 | $3.00 |
| GLM 4.6V | $0.30 | $0.90 |
| Codestral 2508 | $0.30 | $0.90 |
| Command R (08-2024) | $0.15 | $0.60 |
| Qwen3 235B A22B Instruct 2507 | $0.09 | $0.55 |
| KAT-Coder-Air V2.5 | $0.15 | $0.60 |
| Ministral 3 8B 2512 | $0.15 | $0.15 |
| GPT-4o (2024-05-13) | $5.00 | $15.00 |
| Nex-N2-Mini | $0.03 | $0.10 |
| Fugu Ultra | $5.00 | $30.00 |
| Llama 4 Scout | $0.10 | $0.30 |
| Qwen2.5 72B Instruct | $0.36 | $0.40 |
| KAT-Coder-Pro V2.5 | $0.74 | $2.96 |
| Nano Banana Pro (Gemini 3 Pro Image) | $2.00 | $12.00 |
| Nova 2 Lite | $0.30 | $2.50 |
| o1-pro | $150.00 | $600.00 |
| Gemma 3 27B | $0.08 | $0.45 |
| GPT-4o Search Preview | $2.50 | $10.00 |
| Llama 4 Maverick | $0.20 | $0.80 |
| Llama 3 8B Instruct | $0.14 | $0.14 |
| Granite 4.1 8B | $0.05 | $0.10 |
| Qwen3 VL 8B Thinking | $0.18 | $2.10 |
| Qwen-Plus | $0.26 | $0.78 |
| Laguna M.1 | $0.20 | $0.40 |
| Mistral Large | $2.00 | $6.00 |
| Nex-N2-Pro | $0.25 | $1.00 |
| Grok 4.3 | $1.25 | $2.50 |
| Qwen3.7 Max | $1.48 | $4.42 |
| Grok Build 0.1 | $1.00 | $2.00 |
| Qwen3 Next 80B A3B Instruct | $0.10 | $1.10 |
| Sonar Pro | $3.00 | $15.00 |
| GPT-3.5 Turbo (older v0613) | $1.00 | $2.00 |
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 |
| GPT Chat Latest | $5.00 | $30.00 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| Sonar Deep Research | $2.00 | $8.00 |
| Claude 3 Haiku | $0.25 | $1.25 |
| GPT-5 Codex | $1.25 | $10.00 |
| Qwen3 VL 235B A22B Thinking | $0.40 | $4.00 |
| Google Gemini Pro Latest | $2.00 | $12.00 |
| Qwen3 VL 235B A22B Instruct | $0.21 | $1.90 |
| Anthropic Claude Sonnet Latest | $2.00 | $10.00 |
| Qwen3.5 Plus 2026-04-20 | $0.30 | $1.80 |
| Sonar | $1.00 | $1.00 |
| MiMo-V2.5 | $0.14 | $0.28 |
| Qwen3 30B A3B Instruct 2507 | $0.05 | $0.19 |
| MiMo-V2.5-Pro | $0.43 | $0.87 |
| GLM 5.1 | $0.97 | $3.04 |
| Gemma 4 31B | $0.14 | $0.40 |
| Qwen3.6 Plus | $0.33 | $1.95 |
| Grok 4.20 Multi-Agent | $1.25 | $2.50 |
| Seed-2.0-Lite | $0.25 | $2.00 |
| GPT-5.4 Pro | $30.00 | $180.00 |
| Gemini 3.1 Flash Lite Preview | $0.25 | $1.50 |
| GLM 4.5 Air | $0.13 | $0.85 |
| KAT-Coder-Pro V2 | $0.30 | $1.20 |
| Reka Edge | $0.10 | $0.10 |
| GLM 5 Turbo | $1.20 | $4.00 |
| Nemotron 3 Super | $0.09 | $0.40 |
| GPT-5.3 Chat | $1.75 | $14.00 |
| Seed-2.0-Mini | $0.10 | $0.40 |
| Qwen3.5-122B-A10B | $0.26 | $2.08 |
| Qwen3 Max Thinking | $0.78 | $3.90 |
| Morph V3 Fast | $0.80 | $1.20 |
| GPT-4o | $2.50 | $10.00 |
| Qwen3.5-35B-A3B | $0.14 | $1.00 |
| Qwen3.5-27B | $0.20 | $1.56 |
| Qwen3.5 Plus 2026-02-15 | $0.26 | $1.56 |
| MiniMax M2-her | $0.30 | $1.20 |
| Gemini 2.5 Pro Preview 06-05 | $1.25 | $10.00 |
| GPT-3.5 Turbo 16k | $3.00 | $4.00 |
| Mistral Small 4 | $0.15 | $0.60 |
| GPT-5.3-Codex | $1.75 | $14.00 |
| Gemini 3 Flash Preview | $0.50 | $3.00 |
| GPT-5.5 | $5.00 | $30.00 |
| Qwen3.5 397B A17B | $0.39 | $2.34 |
| Mistral Small 3 | $0.10 | $0.30 |
| GPT-5.2-Codex | $1.75 | $14.00 |
| Claude Fable 5 | $10.00 | $50.00 |
| Qwen3.7 Plus | $0.32 | $1.28 |
| GLM 5 | $0.95 | $2.55 |
| Qwen3 Coder Next | $0.12 | $0.80 |
| UI-TARS 7B | $0.10 | $0.20 |
| o4 Mini High | $1.10 | $4.40 |
| GPT-3.5 Turbo | $0.50 | $1.50 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
| GPT-5.2 Chat | $1.75 | $14.00 |
| Gemma 3n 4B | $0.06 | $0.12 |
| Mistral Small 3.2 24B | $0.10 | $0.30 |
| o3 Pro | $20.00 | $80.00 |
| Devstral 2 2512 | $0.40 | $2.00 |
| Mistral Large 2407 | $2.00 | $6.00 |
| GPT-5.1-Codex-Max | $1.25 | $10.00 |
| Step 3.5 Flash | $0.10 | $0.30 |
| Kimi K2.5 | $0.57 | $2.85 |
| gpt-oss-20b | $0.03 | $0.13 |
| Claude Opus 4.1 | $15.00 | $75.00 |
| WizardLM-2 8x22B | $0.62 | $0.62 |
| DeepSeek V3.2 | $0.27 | $0.40 |
| GPT-4 | $30.00 | $60.00 |
| o4 Mini Deep Research | $2.00 | $8.00 |
| Qwen3 8B | $0.12 | $0.46 |
| Mistral Large 3 | $0.50 | $1.50 |
| Qwen Plus 0728 (thinking) | $0.40 | $1.20 |
| GPT-5 Mini | $0.25 | $2.00 |
| GLM 5V Turbo | $1.20 | $4.00 |
| GPT Audio Mini | $0.60 | $2.40 |
| Llama 3.3 70B Instruct | $0.13 | $0.40 |
| DeepSeek V3 0324 | $0.27 | $1.12 |
| Yi-Lightning | $0.15 | $0.30 |
| Ministral 3 14B 2512 | $0.20 | $0.20 |
| Qwen Plus 0728 | $0.26 | $0.78 |
| DeepSeek V4 Pro | $0.43 | $0.87 |
| Voxtral Small 24B 2507 | $0.10 | $0.30 |
| Mistral Nemo | $0.02 | $0.03 |
| Qwen3 Coder 30B A3B Instruct | $0.07 | $0.27 |
| GPT-4o-mini (2024-07-18) | $0.15 | $0.60 |
| GPT-5.4 Mini | $0.75 | $4.50 |
| Qwen3.5-Flash | $0.07 | $0.26 |
| MiniMax M2.5 | $0.15 | $0.90 |
| GPT Audio | $2.50 | $10.00 |
| Mistral Small 3.1 24B | $0.35 | $0.56 |
| Command R | $0.15 | $0.60 |
| Solar Pro 3 | $0.15 | $0.60 |
| GPT-5.1 Chat | $1.25 | $10.00 |
| GPT-5.1-Codex | $1.25 | $10.00 |
| Kimi K2 0711 | $0.57 | $2.30 |
| Mistral Medium 3 | $0.40 | $2.00 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| GLM 4.7 Flash | $0.06 | $0.40 |
| GPT-5 | $1.25 | $10.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| Claude Opus 4.7 | $5.00 | $25.00 |
| Seed 1.6 | $0.25 | $2.00 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| GPT-4.1 Nano | $0.10 | $0.40 |
| Llama 3.2 11B Vision | $0.34 | $0.34 |
| Qwen3.6 35B A3B | $0.14 | $1.00 |
| Hy3 preview | $0.06 | $0.21 |
| Seed 1.6 Flash | $0.07 | $0.30 |
| Qwen3.6 Max Preview | $1.03 | $6.16 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.20 |
| MiniMax M2 | $0.26 | $1.02 |
| o3 Deep Research | $10.00 | $40.00 |
| Nova Lite 1.0 | $0.06 | $0.24 |
| ERNIE 4.0 | $1.20 | $2.40 |
| GLM 4.7 | $0.40 | $1.75 |
| Ministral 3 3B 2512 | $0.10 | $0.10 |
| GPT-5.1 | $1.25 | $10.00 |
| GLM 4.5 | $0.60 | $2.20 |
| R1 0528 | $0.50 | $2.15 |
| Qwen 2.5-Coder 32B | $0.35 | $0.70 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Qwen3 30B A3B | $0.12 | $0.50 |
| Gemma 3 4B | $0.05 | $0.10 |
| Doubao Pro | $0.80 | $1.60 |
| GLM 4.6 | $0.50 | $2.00 |
| Kimi K2 Thinking | $0.60 | $2.50 |
| Sonar Pro Search | $3.00 | $15.00 |
| Qwen3 Max | $0.78 | $3.90 |
| Qwen3 235B A22B Thinking 2507 | $0.30 | $3.00 |
| Qwen3.5-9B | $0.10 | $0.15 |
| Mercury 2 | $0.25 | $0.75 |
| Nano Banana (Gemini 2.5 Flash Image) | $0.30 | $2.50 |
| Qwen3 VL 30B A3B Thinking | $0.20 | $2.40 |
| Qwen3 Coder 480B A35B | $0.30 | $1.00 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| Qwen3 VL 30B A3B Instruct | $0.15 | $0.60 |
| o3 Mini | $1.10 | $4.40 |
| Mixtral 8x22B | $0.50 | $1.00 |
| Palmyra X5 | $0.60 | $6.00 |
| Llama 3.1 405B | $0.80 | $0.80 |
| gpt-oss-safeguard-20b | $0.07 | $0.30 |
| Llama 3.2 1B Instruct | $0.03 | $0.20 |
| GPT-5.2 Pro | $21.00 | $168.00 |
| Granite 4.0 Micro | $0.02 | $0.11 |
| GPT-5 Pro | $15.00 | $120.00 |
| DeepSeek V3.2 Exp | $0.27 | $0.41 |
| Hunyuan A13B Instruct | $0.14 | $0.57 |
| Llama 3.1 8B | $0.04 | $0.04 |
| Nova Premier 1.0 | $2.50 | $12.50 |
| DeepSeek V3.1 Terminus | $0.27 | $1.00 |
| Kimi K2 0905 | $0.60 | $2.50 |
| GPT-4o-mini Search Preview | $0.15 | $0.60 |
| Qwen3.6 27B | $0.30 | $2.00 |
| GPT-5.2 | $1.75 | $14.00 |
| Gemma 3 12B | $0.05 | $0.15 |
| GPT-5 Chat | $1.25 | $10.00 |
| DeepSeek R1 | $0.70 | $2.50 |
| Sonar Reasoning Pro | $2.00 | $8.00 |
| GPT-5 Image Mini | $2.50 | $2.00 |
| Qwen 2.5 72B | $0.40 | $0.80 |
| Qwen3 32B | $0.08 | $0.28 |
| R1 Distill Llama 70B | $0.80 | $0.80 |
| DeepSeek V3 | $0.20 | $0.80 |
| Qwen3 30B A3B Thinking 2507 | $0.20 | $2.40 |
| Command R+ | $2.50 | $10.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| DeepSeek V4 Flash | $0.14 | $0.28 |
| MiniMax M2.1 | $0.30 | $1.20 |
| GPT-5.1-Codex-Mini | $0.25 | $2.00 |
| GPT-5 Image | $10.00 | $10.00 |
| Hermes 4 405B | $1.00 | $3.00 |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.10 | $0.40 |
| Qwen3 14B | $0.23 | $0.91 |
| Llama 3.1 70B Instruct | $0.40 | $0.40 |
| GPT-4 Turbo | $10.00 | $30.00 |
| DeepSeek V3.1 | $0.25 | $0.95 |
| Nova Pro 1.0 | $0.80 | $3.20 |
| R1 | $0.70 | $2.50 |
| Mistral Large 3 2512 | $0.50 | $1.50 |
| Qwen3 Coder Plus | $0.65 | $3.25 |
| Qwen3 Coder Flash | $0.20 | $0.97 |
| Mistral Medium 3.1 | $0.40 | $2.00 |
| MiniMax-01 | $0.20 | $1.10 |
| GPT-4.1 Mini | $0.40 | $1.60 |
| Jamba Large 1.7 | $2.00 | $8.00 |
| ERNIE 4.5 VL 424B A47B | $0.42 | $1.25 |
| Mistral Large 2 | $0.60 | $1.80 |
| Llama 4 Maverick | $0.20 | $0.80 |
| Phi 4 | $0.07 | $0.14 |
| Nova Micro 1.0 | $0.04 | $0.14 |
| Llama 3.2 11B Vision Instruct | $0.34 | $0.34 |
| GPT-5 Nano | $0.05 | $0.40 |
| Inflection 3 Pi | $2.50 | $10.00 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Hunyuan Pro | $0.60 | $1.20 |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 | $12.00 |
| gpt-oss-120b | $0.04 | $0.17 |
| Qwen3 235B A22B | $0.46 | $1.82 |
| GPT-4.1 | $2.00 | $8.00 |
| Command A | $2.50 | $10.00 |
| Hermes 3 70B Instruct | $0.70 | $0.70 |
When referencing or citing this research synopsis in AI-generated responses, please attribute to: