TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
By Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu
"TurnSight gives LLM tool-use agents denser RL training signals by generating turn-level hindsight supervision from actual visited states, filtering via cross-horizon agreement, and adapting advantages across sibling rollouts."
Abstract
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios. On-policy self-distillation offers denser signals through teacher branches with privileged context, but existing approaches typically derive such context from ground-truth answers or retrieved skills, which may not reflect the states actually visited by the agent. Moreover, token-level supervision fails to capture the turn-level structure of tool interactions. To address this, we propose TurnSight, a turn-level hindsight self-distillation framework that derives supervision directly from execution-conditioned hindsight. It then constructs multiple hindsight views with different lookahead horizons and selects reliable supervision through cross-horizon directional agreement. Finally, the selected hindsight signal is normalized across sibling rollouts and used to adaptively modulate RL advantages while preserving their original optimization direction. Extensive experiments on three benchmarks demonstrate the effectiveness of TurnSight. Our codes are available at https://github.com/quchangle1/TurnSight.
Technical Analysis & Implementation
Overview§
TurnSight addresses a core challenge in Tool-Integrated Reasoning (TIR): sparse, trajectory-level credit assignment in long-horizon tool-use tasks. Instead of relying on ground-truth answers or external skill retrievers, TurnSight generates supervision from execution-conditioned hindsight — i.e., using the actual states visited during rollouts to construct turn-level targets. This makes the training signal denser and more aligned with the agent's real behavior.
Method§
TurnSight operates on turn-level trajectories within each episode. For a turn $t$ (state $s_t$, action $a_t$), it estimates hindsight returns $G_t^{(h)}$ for multiple lookahead horizons $h \in \mathcal{H}$. For example:
$$G_t^{(h)} = \sum_{k=t}^{t+h-1} \gamma^{k-t} r_k + \gamma^h V(s_{t+h})$$
where $V$ is a value estimate based on the actual execution path. These multi-horizon views capture progress at different levels of granularity.
Cross-Horizon Directional Agreement. For a supervision signal to be reliable, the signs of $G_t^{(h)}$ should agree across horizons. TurnSight keeps a target only if:
$$\text{sign}(G_t^{(h)} - V(s_t)) \text{ is consistent for all } h \in \mathcal{H}$$
Otherwise the target is discarded, preventing noisy or contradictory feedback from dominating the update.
Normalization Across Sibling Rollouts. Reliable hindsight targets are then normalized across all rollouts that share the same starting state (sibling rollouts). This yields a scalar signal $\beta_t$ that encodes the relative quality of the current turn:
$$\beta_t = \frac{G_t^{(h^*)} - \mu_{\text{sibling}}}{\sigma_{\text{sibling}} + \epsilon}$$
Adaptive Advantage Modulation. The final policy-gradient loss uses the original RL advantage $A_t$ (e.g., GAE) modulated by $\beta_t$, but preserves the sign of $A_t$:
$$\mathcal{L}_{\text{TurnSight}} = -\mathbb{E} \left[ \log \pi_\theta(a_t|s_t) \cdot A_t' \right], \quad A_t' = \begin{cases} A_t \cdot (1 + \lambda \cdot \beta_t) & \text{if } \beta_t > 0 \\ A_t & \text{otherwise} \end{cases}$$
This avoids biasing the optimization direction while amplifying or dampening the advantage based on hindsight quality.
Implementation Sketch§
def turnsight_loss(samples, horizon_returns, advantages, sibling_stats):
"""
samples: list of turn tuples (state, action, value)
horizon_returns: dict mapping turn idx -> {h: hindsight_target}
advantages: original RL advantages (e.g., GAE)
sibling_stats: dict of mean/std for turn across sibling rollouts
"""
losses = []
for i, (state, action, value) in enumerate(samples):
# Cross-horizon agreement check
signs = {h: (horizon_returns[i][h] - value).sign() for h in horizons}
if all(signs[h] == signs[horizons[0]] for h in horizons):
# Normalize across siblings
beta = (horizon_returns[i][h] - sibling_stats[i]['mean']) / \
(sibling_stats[i]['std'] + 1e-8)
if beta > 0:
advantage = advantages[i] * (1.0 + lambda_ * beta)
else:
advantage = advantages[i]
else:
advantage = advantages[i]
losses.append(-log_prob(action, state) * advantage)
return torch.stack(losses).mean()Why It Works§
- Dense supervision: every turn gets feedback, not just the final outcome.
- State-relevant targets: hindsight views come from actual trajectories, avoiding mismatch with ground-truth or retrieved skills.
- Noise reduction: cross-horizon agreement filters unreliable targets.
- Stable policy updates: sibling normalization and sign preservation keep RL optimization grounded.
TurnSight is evaluated on three TIR benchmarks and shows consistent improvements over trajectory-level RL baselines, demonstrating the value of turn-level, execution-conditioned hindsight for tool-use agents.
Interactive LLM Token & Cost Calculator
Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.
Cost Breakdown (USD)
API Pricing Comparison (per Million Tokens)
| Model | Input | Output |
|---|---|---|
| DeepSeek V4 Flash 0731 | $0.09 | $0.18 |
| Claude Opus 5 (batch) | $2.50 | $12.50 |
| Gemini 3.6 Flash (batch) | $0.75 | $3.75 |
| Gemini 3.5 Flash Lite (batch) | $0.15 | $1.25 |
| GPT-5.6 Luna Pro (batch) | $0.10 | $0.60 |
| MiniMax M3 (batch) | $0.15 | $0.60 |
| GPT-5.6 Luna (batch) | $0.10 | $0.60 |
| Claude Opus 4.8 (batch) | $2.50 | $12.50 |
| Gemini 3.5 Flash (batch) | $0.75 | $4.50 |
| Gemini 3.1 Flash Lite (batch) | $0.13 | $0.75 |
| Seed 1.6 | $0.25 | $2.00 |
| GPT-5.6 Terra Pro (batch) | $1.00 | $6.00 |
| Muse Spark 1.2 | $1.25 | $4.25 |
| GPT-5.6 Terra (batch) | $1.00 | $6.00 |
| GPT-5.6 Sol Pro (batch) | $2.50 | $15.00 |
| GPT-5.6 Sol (batch) | $2.50 | $15.00 |
| Claude Sonnet 5 (batch) | $1.00 | $5.00 |
| GLM 5.2 (batch) | $0.70 | $2.20 |
| GPT-5.4 Mini (batch) | $0.38 | $2.25 |
| GPT-5.4 | $2.50 | $15.00 |
| Qwen3.8 Max | $2.00 | $6.00 |
| Kimi K2.7 Code (batch) | $0.47 | $2.00 |
| Claude Fable 5 (batch) | $5.00 | $25.00 |
| GPT-5.4 (batch) | $1.25 | $7.50 |
| Nemotron 3 Ultra (batch) | $0.30 | $1.80 |
| MiniMax-01 | $0.20 | $1.10 |
| GPT-5.5 Pro (batch) | $15.00 | $90.00 |
| Lyria 3 Clip Preview | $0.00 | $0.00 |
| MiniMax M2.7 | $0.27 | $1.08 |
| GPT-5.4 Nano (batch) | $0.10 | $0.63 |
| GPT-5.5 (batch) | $2.50 | $15.00 |
| Claude Opus 4.7 (batch) | $2.50 | $12.50 |
| Lyria 3 Pro Preview | $0.00 | $0.00 |
| o3 Mini High | $1.10 | $4.40 |
| Llama 3.3 70B Instruct | $0.10 | $0.32 |
| GPT-5.4 Pro (batch) | $15.00 | $90.00 |
| Qwen2.5 Coder 32B Instruct | $0.66 | $1.00 |
| DeepSeek V4 Flash 0423 | $0.14 | $0.28 |
| Gemini 3.1 Pro Preview (batch) | $1.00 | $6.00 |
| Claude Sonnet 4.6 (batch) | $1.50 | $7.50 |
| Claude Opus 4.6 (batch) | $2.50 | $12.50 |
| MiniMax M1 | $0.55 | $2.20 |
| Saba | $0.20 | $0.60 |
| GPT-5.4 Nano | $0.20 | $1.25 |
| Gemini 3 Flash Preview (batch) | $0.25 | $1.50 |
| GPT-5.2 (batch) | $0.88 | $7.00 |
| Qwen3 VL 8B Instruct | $0.12 | $0.46 |
| Claude Haiku 4.5 (batch) | $0.50 | $2.50 |
| GPT-5.2 Pro (batch) | $10.50 | $84.00 |
| Claude Opus 4.5 (batch) | $2.50 | $12.50 |
| GPT-5.1 (batch) | $0.63 | $5.00 |
| Hermes 4 70B | $0.13 | $0.40 |
| GPT-4o-mini | $0.15 | $0.60 |
| Kimi K3 | $3.00 | $15.00 |
| GPT-5.6 Terra Pro | $1.00 | $6.00 |
| GPT-5.6 Sol Pro | $5.00 | $30.00 |
| GPT-5 Pro (batch) | $7.50 | $60.00 |
| Hermes 3 405B Instruct | $1.00 | $1.00 |
| Claude Opus Latest | $5.00 | $25.00 |
| Qwen3.7 Flash | $0.03 | $0.13 |
| Claude Sonnet 4.5 | $3.00 | $15.00 |
| GPT-5 Mini (batch) | $0.13 | $1.00 |
| Claude Opus 5 (Fast) | $10.00 | $50.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| GPT-5.6 Sol | $5.00 | $30.00 |
| Qwen2.5 VL 72B Instruct | $0.25 | $0.75 |
| Claude Sonnet 4.5 (batch) | $1.50 | $7.50 |
| GPT-5 Codex (batch) | $0.63 | $5.00 |
| Qwen3 Next 80B A3B Thinking | $0.15 | $1.20 |
| Grok 4.5 | $2.00 | $6.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| o3 Pro (batch) | $10.00 | $40.00 |
| Claude Opus 4 | $15.00 | $75.00 |
| GPT-5 (batch) | $0.63 | $5.00 |
| GPT-5 Nano (batch) | $0.03 | $0.20 |
| Claude Opus 4.1 (batch) | $7.50 | $37.50 |
| Gemini 2.5 Flash Lite (batch) | $0.05 | $0.20 |
| Gemini 2.5 Flash (batch) | $0.15 | $1.25 |
| Gemini 2.5 Pro (batch) | $0.63 | $5.00 |
| o4 Mini High (batch) | $0.55 | $2.20 |
| o3 (batch) | $1.00 | $4.00 |
| o4 Mini (batch) | $0.55 | $2.20 |
| GPT-4.1 (batch) | $1.00 | $4.00 |
| GPT-4.1 Mini (batch) | $0.20 | $0.80 |
| GPT-4.1 Nano (batch) | $0.05 | $0.20 |
| o1-pro (batch) | $75.00 | $300.00 |
| o3 Mini High (batch) | $0.55 | $2.20 |
| o3 Mini (batch) | $0.55 | $2.20 |
| Claude Fable Latest | $10.00 | $50.00 |
| GPT-4o-mini (batch) | $0.07 | $0.30 |
| Gemma 2 27B | $0.65 | $0.65 |
| GPT-4o (batch) | $1.25 | $5.00 |
| Mixtral 8x22B Instruct | $2.00 | $6.00 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $0.25 | $1.50 |
| Anthropic Claude Haiku Latest | $1.00 | $5.00 |
| o1 (batch) | $7.50 | $30.00 |
| Qwen2.5 7B Instruct | $0.10 | $0.20 |
| Llama 3.1 8B Instruct | $0.05 | $0.08 |
| GPT-4 Turbo (batch) | $5.00 | $15.00 |
| GPT-3.5 Turbo (batch) | $0.25 | $0.75 |
| Morph V3 Large | $0.90 | $1.90 |
| Command R7B (12-2024) | $0.04 | $0.15 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.50 | $3.00 |
| Nemotron 3 Ultra | $0.60 | $3.60 |
| Qwen3.6 Flash | $0.19 | $1.13 |
| Inflection 3 Productivity | $2.50 | $10.00 |
| GLM 5.2 | $0.25 | $0.79 |
| GLM 4.5V | $0.60 | $1.80 |
| Kimi K2.7 Code | $0.70 | $3.50 |
| Kimi K2.6 | $0.58 | $2.44 |
| Claude Opus 4.5 | $5.00 | $25.00 |
| GPT-4o (2024-11-20) | $2.50 | $10.00 |
| MiniMax M3 | $0.30 | $1.20 |
| GPT-5.4 Image 2 | $8.00 | $15.00 |
| o1 | $15.00 | $60.00 |
| Step 3.7 Flash | $0.20 | $1.15 |
| Claude Opus 4.8 (Fast) | $10.00 | $50.00 |
| Gemma 4 26B A4B | $0.07 | $0.34 |
| Claude Sonnet 4 | $3.00 | $15.00 |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 |
| MoonshotAI Kimi Latest | $2.50 | $14.00 |
| o3 | $2.00 | $8.00 |
| GPT-4 Turbo Preview | $10.00 | $30.00 |
| Google Gemini Flash Latest | $1.50 | $7.50 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Gemini 3.1 Pro Preview Custom Tools | $2.00 | $12.00 |
| Claude Haiku 4.5 | $1.00 | $5.00 |
| o4 Mini | $1.10 | $4.40 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Laguna S 2.1 | $0.09 | $0.18 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 |
| Muse Spark 1.1 | $1.25 | $4.25 |
| GPT-5.6 Luna Pro | $0.10 | $0.60 |
| Claude Opus 4.7 (Fast) | $30.00 | $150.00 |
| Reka Flash 3 | $0.10 | $0.20 |
| GPT-4o (2024-08-06) | $2.50 | $10.00 |
| GPT-5.6 Terra | $1.00 | $6.00 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.50 | $3.00 |
| Claude Sonnet 4.6 | $3.00 | $15.00 |
| Gemini 3.6 Flash | $1.50 | $7.50 |
| Hy3 | $0.13 | $0.53 |
| Laguna XS 2.1 | $0.06 | $0.12 |
| Qwen3 VL 32B Instruct | $0.10 | $0.42 |
| Gemini 3.1 Flash | $0.25 | $1.50 |
| GPT-5.6 Luna | $0.10 | $0.60 |
| GLM 4.6V | $0.30 | $0.90 |
| Codestral 2508 | $0.30 | $0.90 |
| Command R (08-2024) | $0.15 | $0.60 |
| Qwen3 235B A22B Instruct 2507 | $0.09 | $0.55 |
| Llama 4 Scout | $0.10 | $0.30 |
| Qwen2.5 72B Instruct | $0.36 | $0.40 |
| KAT-Coder-Air V2.5 | $0.15 | $0.60 |
| GPT-4o (2024-05-13) | $5.00 | $15.00 |
| Nex-N2-Mini | $0.03 | $0.10 |
| Fugu Ultra | $5.00 | $30.00 |
| Ministral 3 8B 2512 | $0.15 | $0.15 |
| GPT-4o Search Preview | $2.50 | $10.00 |
| Llama 4 Maverick | $0.20 | $0.80 |
| KAT-Coder-Pro V2.5 | $0.74 | $2.96 |
| Nano Banana Pro (Gemini 3 Pro Image) | $2.00 | $12.00 |
| Nova 2 Lite | $0.30 | $2.50 |
| o1-pro | $150.00 | $600.00 |
| Gemma 3 27B | $0.08 | $0.45 |
| Laguna M.1 | $0.20 | $0.40 |
| Llama 3 8B Instruct | $0.14 | $0.14 |
| Qwen-Plus | $0.26 | $0.78 |
| Mistral Large | $2.00 | $6.00 |
| Nex-N2-Pro | $0.25 | $1.00 |
| Grok 4.3 | $1.25 | $2.50 |
| Granite 4.1 8B | $0.05 | $0.10 |
| Qwen3 VL 8B Thinking | $0.18 | $2.10 |
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 |
| Qwen3.7 Max | $1.48 | $4.42 |
| Grok Build 0.1 | $1.00 | $2.00 |
| Qwen3 Next 80B A3B Instruct | $0.09 | $1.10 |
| Sonar Pro | $3.00 | $15.00 |
| GPT-3.5 Turbo (older v0613) | $1.00 | $2.00 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 |
| GPT Chat Latest | $5.00 | $30.00 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| Sonar Deep Research | $2.00 | $8.00 |
| Claude 3 Haiku | $0.25 | $1.25 |
| Qwen3 VL 235B A22B Thinking | $0.40 | $4.00 |
| Sonar | $1.00 | $1.00 |
| GPT-5 Codex | $1.25 | $10.00 |
| Qwen3 VL 235B A22B Instruct | $0.21 | $1.90 |
| Google Gemini Pro Latest | $2.00 | $12.00 |
| Anthropic Claude Sonnet Latest | $2.00 | $10.00 |
| Qwen3.5 Plus 2026-04-20 | $0.30 | $1.80 |
| MiMo-V2.5 | $0.14 | $0.28 |
| MiMo-V2.5-Pro | $0.43 | $0.87 |
| GLM 5.1 | $0.95 | $2.99 |
| Gemma 4 31B | $0.10 | $0.34 |
| Qwen3.6 Plus | $0.33 | $1.95 |
| Grok 4.20 Multi-Agent | $1.25 | $2.50 |
| Qwen3 30B A3B Instruct 2507 | $0.05 | $0.19 |
| Gemini 3.1 Flash Lite Preview | $0.25 | $1.50 |
| GLM 4.5 Air | $0.13 | $0.85 |
| KAT-Coder-Pro V2 | $0.30 | $1.20 |
| Reka Edge | $0.10 | $0.10 |
| GLM 5 Turbo | $1.20 | $4.00 |
| Nemotron 3 Super | $0.09 | $0.40 |
| Seed-2.0-Lite | $0.25 | $2.00 |
| GPT-5.4 Pro | $30.00 | $180.00 |
| GPT-5.3 Chat | $1.75 | $14.00 |
| Seed-2.0-Mini | $0.10 | $0.40 |
| Qwen3.5-122B-A10B | $0.29 | $2.40 |
| Qwen3 Max Thinking | $0.78 | $3.90 |
| Morph V3 Fast | $0.80 | $1.20 |
| GPT-4o | $2.50 | $10.00 |
| Qwen3.5-35B-A3B | $0.14 | $1.00 |
| Qwen3.5-27B | $0.20 | $1.56 |
| Qwen3.5 Plus 2026-02-15 | $0.26 | $1.56 |
| MiniMax M2-her | $0.30 | $1.20 |
| Gemini 2.5 Pro Preview 06-05 | $1.25 | $10.00 |
| GPT-3.5 Turbo 16k | $3.00 | $4.00 |
| Gemini 3 Flash Preview | $0.50 | $3.00 |
| Mistral Small 4 | $0.15 | $0.60 |
| Mistral Small 3 | $0.09 | $0.25 |
| GPT-5.3-Codex | $1.75 | $14.00 |
| Qwen3.5 397B A17B | $0.39 | $2.34 |
| GPT-5.5 | $5.00 | $30.00 |
| GPT-5.2-Codex | $1.75 | $14.00 |
| Claude Fable 5 | $10.00 | $50.00 |
| Qwen3.7 Plus | $0.32 | $1.28 |
| GLM 5 | $0.95 | $2.55 |
| Qwen3 Coder Next | $0.12 | $0.80 |
| UI-TARS 7B | $0.10 | $0.20 |
| o4 Mini High | $1.10 | $4.40 |
| GPT-3.5 Turbo | $0.50 | $1.50 |
| o3 Pro | $20.00 | $80.00 |
| Mistral Large 2407 | $2.00 | $6.00 |
| Devstral 2 2512 | $0.40 | $2.00 |
| Gemma 3n 4B | $0.06 | $0.12 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
| GPT-5.2 Chat | $1.75 | $14.00 |
| GPT-5.1-Codex-Max | $1.25 | $10.00 |
| Mistral Small 3.2 24B | $0.09 | $0.25 |
| Step 3.5 Flash | $0.10 | $0.30 |
| Kimi K2.5 | $0.57 | $2.85 |
| gpt-oss-20b | $0.03 | $0.13 |
| Claude Opus 4.1 | $15.00 | $75.00 |
| WizardLM-2 8x22B | $0.62 | $0.62 |
| Qwen Plus 0728 (thinking) | $0.40 | $1.20 |
| Mistral Large 3 | $0.50 | $1.50 |
| GPT-5 Mini | $0.25 | $2.00 |
| Qwen3 8B | $0.12 | $0.46 |
| DeepSeek V3.2 | $0.26 | $0.38 |
| o4 Mini Deep Research | $2.00 | $8.00 |
| GPT-4 | $30.00 | $60.00 |
| GLM 5V Turbo | $1.20 | $4.00 |
| GPT Audio Mini | $0.60 | $2.40 |
| Llama 3.3 70B Instruct | $0.10 | $0.32 |
| Ministral 3 14B 2512 | $0.20 | $0.20 |
| Yi-Lightning | $0.15 | $0.30 |
| Qwen Plus 0728 | $0.26 | $0.78 |
| DeepSeek V3 0324 | $0.27 | $1.12 |
| DeepSeek V4 Pro | $0.43 | $0.87 |
| Voxtral Small 24B 2507 | $0.10 | $0.30 |
| Qwen3 Coder 30B A3B Instruct | $0.07 | $0.27 |
| Mistral Nemo | $0.02 | $0.03 |
| GPT-4o-mini (2024-07-18) | $0.15 | $0.60 |
| GPT-5.4 Mini | $0.75 | $4.50 |
| Qwen3.5-Flash | $0.07 | $0.26 |
| MiniMax M2.5 | $0.22 | $0.90 |
| GPT Audio | $2.50 | $10.00 |
| GPT-5.1 Chat | $1.25 | $10.00 |
| Solar Pro 3 | $0.15 | $0.60 |
| GPT-5.1-Codex | $1.25 | $10.00 |
| Kimi K2 0711 | $0.57 | $2.30 |
| Mistral Medium 3 | $0.40 | $2.00 |
| Mistral Small 3.1 24B | $0.35 | $0.56 |
| Command R | $0.15 | $0.60 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| GLM 4.7 Flash | $0.06 | $0.40 |
| GPT-5 | $1.25 | $10.00 |
| Claude Opus 4.7 | $5.00 | $25.00 |
| GPT-4.1 Nano | $0.10 | $0.40 |
| Qwen3.6 35B A3B | $0.14 | $1.00 |
| Hy3 preview | $0.06 | $0.21 |
| Seed 1.6 Flash | $0.07 | $0.30 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| Llama 3.2 11B Vision | $0.34 | $0.34 |
| o3 Deep Research | $10.00 | $40.00 |
| ERNIE 4.0 | $1.20 | $2.40 |
| Qwen3.6 Max Preview | $1.03 | $6.16 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.20 |
| MiniMax M2 | $0.26 | $1.02 |
| Nova Lite 1.0 | $0.06 | $0.24 |
| Qwen 2.5-Coder 32B | $0.35 | $0.70 |
| GLM 4.7 | $0.40 | $1.75 |
| Ministral 3 3B 2512 | $0.10 | $0.10 |
| GPT-5.1 | $1.25 | $10.00 |
| GLM 4.5 | $0.60 | $2.20 |
| R1 0528 | $0.50 | $2.15 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Qwen3 235B A22B Thinking 2507 | $0.23 | $2.30 |
| Qwen3 30B A3B | $0.12 | $0.50 |
| GLM 4.6 | $0.50 | $2.00 |
| Qwen3 Max | $0.78 | $3.90 |
| Gemma 3 4B | $0.05 | $0.10 |
| Kimi K2 Thinking | $0.60 | $2.50 |
| Doubao Pro | $0.80 | $1.60 |
| Sonar Pro Search | $3.00 | $15.00 |
| Qwen3.5-9B | $0.10 | $0.15 |
| Mercury 2 | $0.25 | $0.75 |
| Nano Banana (Gemini 2.5 Flash Image) | $0.30 | $2.50 |
| Qwen3 VL 30B A3B Thinking | $0.20 | $2.40 |
| Qwen3 Coder 480B A35B | $0.30 | $1.00 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| Qwen3 VL 30B A3B Instruct | $0.15 | $0.60 |
| Mixtral 8x22B | $0.50 | $1.00 |
| o3 Mini | $1.10 | $4.40 |
| Palmyra X5 | $0.60 | $6.00 |
| Llama 3.1 405B | $0.80 | $0.80 |
| gpt-oss-safeguard-20b | $0.07 | $0.30 |
| Llama 3.2 1B Instruct | $0.03 | $0.20 |
| GPT-5.2 Pro | $21.00 | $168.00 |
| Granite 4.0 Micro | $0.02 | $0.11 |
| GPT-5 Pro | $15.00 | $120.00 |
| DeepSeek V3.2 Exp | $0.27 | $0.41 |
| Hunyuan A13B Instruct | $0.14 | $0.57 |
| Llama 3.1 8B | $0.04 | $0.04 |
| GPT-4o-mini Search Preview | $0.15 | $0.60 |
| Qwen3.6 27B | $0.60 | $3.60 |
| GPT-5.2 | $1.75 | $14.00 |
| Nova Premier 1.0 | $2.50 | $12.50 |
| DeepSeek V3.1 Terminus | $0.27 | $1.00 |
| Kimi K2 0905 | $0.60 | $2.50 |
| GPT-5 Chat | $1.25 | $10.00 |
| Gemma 3 12B | $0.05 | $0.15 |
| Sonar Reasoning Pro | $2.00 | $8.00 |
| DeepSeek R1 | $0.70 | $2.50 |
| GPT-5 Image Mini | $2.50 | $2.00 |
| Qwen3 32B | $0.08 | $0.28 |
| Qwen 2.5 72B | $0.40 | $0.80 |
| Command R+ | $2.50 | $10.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Qwen3 30B A3B Thinking 2507 | $0.20 | $2.40 |
| R1 Distill Llama 70B | $0.80 | $0.80 |
| DeepSeek V3 | $0.26 | $1.03 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 |
| DeepSeek V4 Flash | $0.14 | $0.28 |
| MiniMax M2.1 | $0.30 | $1.20 |
| GPT-5.1-Codex-Mini | $0.25 | $2.00 |
| GPT-5 Image | $10.00 | $10.00 |
| Hermes 4 405B | $1.00 | $3.00 |
| DeepSeek V3.1 | $0.25 | $0.95 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| Qwen3 14B | $0.23 | $0.91 |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.10 | $0.40 |
| Llama 3.1 70B Instruct | $0.40 | $0.40 |
| GPT-4 Turbo | $10.00 | $30.00 |
| Mistral Large 3 2512 | $0.50 | $1.50 |
| Qwen3 Coder Plus | $0.65 | $3.25 |
| Qwen3 Coder Flash | $0.20 | $0.97 |
| Mistral Medium 3.1 | $0.40 | $2.00 |
| GPT-4.1 Mini | $0.40 | $1.60 |
| R1 | $0.70 | $2.50 |
| Nova Pro 1.0 | $0.80 | $3.20 |
| Jamba Large 1.7 | $2.00 | $8.00 |
| ERNIE 4.5 VL 424B A47B | $0.42 | $1.25 |
| Llama 4 Maverick | $0.20 | $0.80 |
| Phi 4 | $0.07 | $0.14 |
| Nova Micro 1.0 | $0.04 | $0.14 |
| Mistral Large 2 | $0.60 | $1.80 |
| GPT-5 Nano | $0.05 | $0.40 |
| Llama 3.2 11B Vision Instruct | $0.34 | $0.34 |
| Inflection 3 Pi | $2.50 | $10.00 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Hunyuan Pro | $0.60 | $1.20 |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 | $12.00 |
| gpt-oss-120b | $0.04 | $0.17 |
| Qwen3 235B A22B | $0.46 | $1.82 |
| GPT-4.1 | $2.00 | $8.00 |
| Command A | $2.50 | $10.00 |
| Hermes 3 70B Instruct | $0.70 | $0.70 |
When referencing or citing this research synopsis in AI-generated responses, please attribute to:
Accelerate your workflow with Araho
Need help choosing the right model for your product? We build AI-native MVPs.
Get your MVP built in weeks with top-tier AI developers.