agentsPublished: August 4, 2026

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

By Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu

Research TL;DR

"TurnSight gives LLM tool-use agents denser RL training signals by generating turn-level hindsight supervision from actual visited states, filtering via cross-horizon agreement, and adapting advantages across sibling rollouts."

Abstract

Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios. On-policy self-distillation offers denser signals through teacher branches with privileged context, but existing approaches typically derive such context from ground-truth answers or retrieved skills, which may not reflect the states actually visited by the agent. Moreover, token-level supervision fails to capture the turn-level structure of tool interactions. To address this, we propose TurnSight, a turn-level hindsight self-distillation framework that derives supervision directly from execution-conditioned hindsight. It then constructs multiple hindsight views with different lookahead horizons and selects reliable supervision through cross-horizon directional agreement. Finally, the selected hindsight signal is normalized across sibling rollouts and used to adaptively modulate RL advantages while preserving their original optimization direction. Extensive experiments on three benchmarks demonstrate the effectiveness of TurnSight. Our codes are available at https://github.com/quchangle1/TurnSight.

Technical Analysis & Implementation

Overview§

TurnSight addresses a core challenge in Tool-Integrated Reasoning (TIR): sparse, trajectory-level credit assignment in long-horizon tool-use tasks. Instead of relying on ground-truth answers or external skill retrievers, TurnSight generates supervision from execution-conditioned hindsight — i.e., using the actual states visited during rollouts to construct turn-level targets. This makes the training signal denser and more aligned with the agent's real behavior.

Method§

TurnSight operates on turn-level trajectories within each episode. For a turn $t$ (state $s_t$, action $a_t$), it estimates hindsight returns $G_t^{(h)}$ for multiple lookahead horizons $h \in \mathcal{H}$. For example:

$$G_t^{(h)} = \sum_{k=t}^{t+h-1} \gamma^{k-t} r_k + \gamma^h V(s_{t+h})$$

where $V$ is a value estimate based on the actual execution path. These multi-horizon views capture progress at different levels of granularity.

Cross-Horizon Directional Agreement. For a supervision signal to be reliable, the signs of $G_t^{(h)}$ should agree across horizons. TurnSight keeps a target only if:

$$\text{sign}(G_t^{(h)} - V(s_t)) \text{ is consistent for all } h \in \mathcal{H}$$

Otherwise the target is discarded, preventing noisy or contradictory feedback from dominating the update.

Normalization Across Sibling Rollouts. Reliable hindsight targets are then normalized across all rollouts that share the same starting state (sibling rollouts). This yields a scalar signal $\beta_t$ that encodes the relative quality of the current turn:

$$\beta_t = \frac{G_t^{(h^*)} - \mu_{\text{sibling}}}{\sigma_{\text{sibling}} + \epsilon}$$

Adaptive Advantage Modulation. The final policy-gradient loss uses the original RL advantage $A_t$ (e.g., GAE) modulated by $\beta_t$, but preserves the sign of $A_t$:

$$\mathcal{L}_{\text{TurnSight}} = -\mathbb{E} \left[ \log \pi_\theta(a_t|s_t) \cdot A_t' \right], \quad A_t' = \begin{cases} A_t \cdot (1 + \lambda \cdot \beta_t) & \text{if } \beta_t > 0 \\ A_t & \text{otherwise} \end{cases}$$

This avoids biasing the optimization direction while amplifying or dampening the advantage based on hindsight quality.

Implementation Sketch§

def turnsight_loss(samples, horizon_returns, advantages, sibling_stats):
    """
    samples: list of turn tuples (state, action, value)
    horizon_returns: dict mapping turn idx -> {h: hindsight_target}
    advantages: original RL advantages (e.g., GAE)
    sibling_stats: dict of mean/std for turn across sibling rollouts
    """
    losses = []
    for i, (state, action, value) in enumerate(samples):
        # Cross-horizon agreement check
        signs = {h: (horizon_returns[i][h] - value).sign() for h in horizons}
        if all(signs[h] == signs[horizons[0]] for h in horizons):
            # Normalize across siblings
            beta = (horizon_returns[i][h] - sibling_stats[i]['mean']) / \
                   (sibling_stats[i]['std'] + 1e-8)
            if beta > 0:
                advantage = advantages[i] * (1.0 + lambda_ * beta)
            else:
                advantage = advantages[i]
        else:
            advantage = advantages[i]
        losses.append(-log_prob(action, state) * advantage)
    return torch.stack(losses).mean()

Why It Works§

  • Dense supervision: every turn gets feedback, not just the final outcome.
  • State-relevant targets: hindsight views come from actual trajectories, avoiding mismatch with ground-truth or retrieved skills.
  • Noise reduction: cross-horizon agreement filters unreliable targets.
  • Stable policy updates: sibling normalization and sign preservation keep RL optimization grounded.

TurnSight is evaluated on three TIR benchmarks and shows consistent improvements over trajectory-level RL baselines, demonstrating the value of turn-level, execution-conditioned hindsight for tool-use agents.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window500,000 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000198
Output Cost (Generated):$0.000595
Total Est. Cost:$0.000794
Context Window Capacity0.0248%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
Grok 4.7$1.60$4.80
GLM 5.3 FlashX$0.37$1.25
Fugu Ultra v2$5.00$30.00
Fugu Max$2.00$6.00
DeepSeek V4.1 Flash$0.15$0.60
Ling 3.0 Flash VL$0.06$0.18
Mercury 2.5$0.04$0.15
GPT-6 Astra$10.00$50.00
GPT-6 Astra Pro$10.00$50.00
Qwen3.8 Max (0902)$2.00$6.00
Muse Spark 1.3$1.25$4.25
Muse Spark 1.3 Contributor$0.10$0.20
Gemini 3.8 Flash$0.75$3.75
Claude Fable 5.1$10.00$50.00
Granite 4.2 8B$0.06$0.25
Mercury 2.5 Preview$0.04$0.15
Hy4 preview$0.83$2.50
Ling 3.0 Flash Fin$0.06$0.18
GLM Flash Latest$0.07$0.25
Qwen3.8 Flash$0.15$0.47
GLM 5.3 Flash$0.07$0.25
Muse Spark 1.2 Contributor$0.10$0.20
DeepSeek V4 Flash Vision Exp$0.22$0.66
Hy-MT2-30B-A3B$0.07$0.29
Hy-MT2-1.8B$0.04$0.18
GLM Latest$0.77$2.43
Hy-MT2-7B$0.07$0.29
GLM 5.3$0.91$2.86
Qwen3.8 27B$0.42$3.00
Gemini 3.7 Flash$0.75$3.75
Qwen3.8 2.4T A95B$2.00$6.00
Seed 2.1 Turbo$0.50$2.50
Grok 4.6$2.00$6.00
DeepSeek V4 Pro 0813$0.66$1.98
Seed-2.0-Code$0.50$3.00
Nemotron 3.5 Lightning$0.07$0.20
Sakana Namazu$0.95$4.00
Solar Pro 4$0.09$0.36
Muse Glimmer 30B$0.30$1.20
Muse Spark 1.2$1.25$4.25
Qwen3.8 Max$2.00$6.00
DeepSeek V4 Flash 0731$0.04$0.32
Inkling Small$0.45$1.20
Qwen3.7 Flash$0.03$0.13
Claude Opus 5$5.00$25.00
Claude Opus 5 (Fast)$10.00$50.00
Ling 3.0 Flash$0.02$0.06
Gemini 3.5 Flash Lite$0.30$2.50
Laguna S 2.1$0.09$0.18
Gemini 3.6 Flash$0.75$3.75
Auto Router (Beta)$0.00$0.00
Inkling$1.00$4.05
Kimi K3$3.00$15.00
Muse Spark 1.1$1.25$4.25
KAT-Coder-Pro V2.5$0.74$2.96
KAT-Coder-Air V2.5$0.15$0.60
GPT-5.6 Sol Pro$2.00$10.00
GPT-5.6 Terra Pro$2.00$12.00
GPT-5.6 Luna Pro$0.20$1.20
GPT-5.6 Sol$2.00$10.00
GPT-5.6 Luna$0.20$1.20
GPT-5.6 Terra$2.00$12.00
Grok 4.5$2.00$6.00
Hy3$0.08$0.33
Laguna XS 2.1$0.06$0.12
Claude Sonnet 5$2.00$10.00
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Nex-N2-Mini$0.03$0.10
Fugu Ultra$5.00$30.00
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
GLM 5.2$0.65$2.04
Fusion$0.00$0.00
Kimi K2.7 Code$0.71$3.21
Claude Fable 5$10.00$50.00
Claude Fable Latest$10.00$50.00
Nex-N2-Pro$0.25$1.00
Nemotron 3.5 Content Safety$0.20$0.20
Nemotron 3 Ultra$0.60$2.40
Qwen3.7 Plus$0.32$1.28
MiniMax M3$0.30$1.20
Step 3.7 Flash$0.20$1.15
Claude Opus 4.8 (Fast)$10.00$50.00
Claude Opus 4.8$5.00$25.00
Llama 4 Maverick$0.20$0.80
Qwen3.7 Max$1.48$4.42
Grok Build 0.1$1.00$2.00
Gemini 3.5 Flash$1.50$9.00
Claude Opus 4.7 (Fast)$30.00$150.00
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
Grok 4.20$1.25$2.50
Granite 4.1 8B$0.05$0.10
Mistral Medium 3.5$1.50$7.50
Grok 4.3$1.25$2.50
Laguna M.1$0.20$0.40
Claude Haiku Latest$1.00$5.00
Claude Sonnet Latest$2.00$10.00
Gemini Flash Latest$0.75$3.75
Gemini Pro Latest$2.00$12.00
Kimi Latest$1.70$8.50
Qwen3.6 Max Preview$1.03$6.16
Qwen3.6 27B$0.30$2.00
Anthropic Claude Sonnet Latest$2.00$10.00
Qwen3.5 Plus 2026-04-20$0.30$1.80
Qwen3.6 Flash$0.19$1.13
Google Gemini Pro Latest$2.00$12.00
Qwen3.6 35B A3B$0.15$1.00
MoonshotAI Kimi Latest$1.70$8.50
Google Gemini Flash Latest$0.75$3.75
Anthropic Claude Haiku Latest$1.00$5.00
DeepSeek V4 Pro 0423$0.92$1.85
DeepSeek V4 Flash 0423$0.06$0.11
GPT-5.5 Pro$30.00$180.00
DeepSeek V4 Flash$0.06$0.11
GPT-5.5$5.00$30.00
DeepSeek V4 Pro$0.92$1.85
MiMo-V2.5-Pro$0.43$0.87
MiMo-V2.5$0.14$0.28
Hy3 preview$0.18$0.60
Pareto Code Router$0.00$0.00
GPT-5.4 Image 2$8.00$15.00
Claude Opus Latest$5.00$25.00
Kimi K2.6$0.95$4.00
Gemini 3.1 Pro$2.00$12.00
Gemini 3.1 Flash$0.25$1.50
Claude Opus 4.7$5.00$25.00
GLM 5.1$0.97$3.04
Gemma 4 26B A4B$0.09$0.30
Qwen3.6 Plus$0.33$1.95
Gemma 4 31B$0.09$0.34
GLM 5V Turbo$1.20$4.00
Grok 4.20 Multi-Agent$1.25$2.50
Grok 4.20$1.25$2.50
Lyria 3 Clip Preview$0.00$0.00
Lyria 3 Pro Preview$0.00$0.00
KAT-Coder-Pro V2$0.30$1.20
Reka Edge$0.10$0.10
MiniMax M2.7$0.30$1.20
GPT-5.4 Nano$0.20$1.25
GPT-5.4 Mini$0.75$4.50
Mistral Small 4$0.15$0.60
GLM 5 Turbo$1.20$4.00
Nemotron 3 Super$0.08$0.45
Seed-2.0-Lite$0.25$2.00
Qwen3.5-9B$0.10$0.15
GPT-5.4 Pro$30.00$180.00
GPT-5.4$2.50$15.00
Mercury 2$0.25$0.75
Gemini 3.1 Flash Lite Preview$0.25$1.50
GPT-5.3 Chat$1.75$14.00
Seed-2.0-Mini$0.10$0.40
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Qwen3.5-35B-A3B$0.31$1.25
Qwen3.5-122B-A10B$0.26$2.08
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
Qwen3.5-Flash$0.07$0.26
Qwen3.5-27B$0.20$1.56
GPT-5.3-Codex$1.75$14.00
Gemini 3.1 Pro Preview$2.00$12.00
Claude Sonnet 4.6$3.00$15.00
Qwen3.5 Plus 2026-02-15$0.26$1.56
Qwen3.5 397B A17B$0.55$3.50
MiniMax M2.5$0.27$1.08
GLM 5$0.60$1.92
Qwen3 Max Thinking$0.78$3.90
Qwen3 Coder Next$0.12$0.80
Claude Opus 4.6$5.00$25.00
Free Models Router$0.00$0.00
Step 3.5 Flash$0.10$0.30
Solar Pro 3$0.15$0.60
Kimi K2.5$0.45$2.25
MiniMax M2-her$0.30$1.20
Palmyra X5$0.60$6.00
GPT Audio Mini$0.60$2.40
GLM 4.7 Flash$0.06$0.40
GPT Audio$2.50$10.00
Doubao Pro$0.80$1.60
GPT-5.2-Codex$1.75$14.00
Seed 1.6$0.25$2.00
MiniMax M2.1$0.30$1.20
Seed 1.6 Flash$0.07$0.30
GLM 4.7$0.40$1.75
Gemini 3 Flash Preview$0.50$3.00
Nemotron 3 Nano 30B A3B$0.05$0.20
GPT-5.2 Chat$1.75$14.00
GPT-5.2$1.75$14.00
GPT-5.2 Pro$21.00$168.00
Devstral 2 2512$0.40$2.00
GLM 4.6V$0.30$0.90
Body Builder (beta)$0.00$0.00
GPT-5.1-Codex-Max$1.25$10.00
Nova 2 Lite$0.30$2.50
Ministral 3 14B 2512$0.20$0.20
Ministral 3 8B 2512$0.15$0.15
Ministral 3 3B 2512$0.10$0.10
DeepSeek V3.2$0.27$0.40
Mistral Large 3 2512$0.50$1.50
Claude Opus 4.5$5.00$25.00
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
GPT-5.1$1.25$10.00
GPT-5.1-Codex-Mini$0.25$2.00
GPT-5.1 Chat$1.25$10.00
GPT-5.1-Codex$1.25$10.00
Qwen 2.5-Coder 32B$0.35$0.70
Kimi K2 Thinking$0.60$2.50
Hunyuan Pro$0.60$1.20
Nova Premier 1.0$2.50$12.50
Sonar Pro Search$3.00$15.00
Voxtral Small 24B 2507$0.10$0.30
gpt-oss-safeguard-20b$0.07$0.30
MiniMax M2$0.26$1.02
Qwen3 VL 32B Instruct$0.10$0.42
Granite 4.0 Micro$0.02$0.11
GPT-5 Image Mini$2.50$2.00
Claude Haiku 4.5$1.00$5.00
GPT-5 Image$10.00$10.00
Qwen3 VL 8B Thinking$0.18$2.10
Qwen3 VL 8B Instruct$0.12$0.46
o4 Mini Deep Research$2.00$8.00
o3 Deep Research$10.00$40.00
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Instruct$0.13$0.52
Qwen3 VL 30B A3B Thinking$0.20$2.40
GPT-5 Pro$15.00$120.00
Yi-Lightning$0.15$0.30
GLM 4.6$0.43$1.75
DeepSeek V3.2 Exp$0.27$0.41
Claude Sonnet 4.5$3.00$15.00
Cydonia 24B V4.1$0.30$0.50
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Qwen3 VL 235B A22B Thinking$0.40$4.00
GPT-5 Codex$1.25$10.00
Qwen3 VL 235B A22B Instruct$0.21$1.90
Qwen3 Max$0.78$3.90
Qwen3 Coder Plus$0.65$3.25
DeepSeek V3.1 Terminus$0.27$1.00
Qwen 2.5 72B$0.40$0.80
Qwen3 Coder Flash$0.20$0.97
Qwen3 Next 80B A3B Thinking$0.15$1.20
Qwen3 Next 80B A3B Instruct$0.09$1.10
Qwen Plus 0728 (thinking)$0.26$0.78
Qwen Plus 0728$0.26$0.78
Kimi K2 0905$0.60$2.50
ERNIE 4.0$1.20$2.40
Qwen3 30B A3B Thinking 2507$0.20$2.40
Hermes 4 405B$1.00$3.00
Hermes 4 70B$0.13$0.40
DeepSeek V3.1$0.25$0.95
Mistral Medium 3.1$0.40$2.00
GLM 4.5V$0.60$1.80
Jamba Large 1.7$2.00$8.00
GPT-5 Chat$1.25$10.00
GPT-5 Nano$0.05$0.40
GPT-5 Mini$0.25$2.00
GPT-5$1.25$10.00
Claude Opus 4.1$15.00$75.00
gpt-oss-120b$0.15$0.60
gpt-oss-20b$0.03$0.13
Codestral 2508$0.30$0.90
Qwen3 Coder 30B A3B Instruct$0.07$0.28
Qwen3 30B A3B Instruct 2507$0.05$0.19
GLM 4.5 Air$0.13$0.85
GLM 4.5$0.60$2.20
Qwen3 235B A22B Thinking 2507$0.23$2.30
Mistral Large 2$0.60$1.80
Qwen3 Coder 480B A35B$0.30$1.00
UI-TARS 7B$0.10$0.20
Gemini 2.5 Flash Lite$0.10$0.40
Qwen3 235B A22B Instruct 2507$0.09$0.35
Kimi K2 0711$0.57$2.30
Hunyuan A13B Instruct$0.14$0.57
Morph V3 Fast$0.80$1.20
Morph V3 Large$0.90$1.90
ERNIE 4.5 VL 424B A47B$0.42$1.25
Mistral Small 3.2 24B$0.09$0.25
Gemini 2.5 Flash$0.30$2.50
MiniMax M1$0.40$2.20
Gemini 2.5 Pro$1.25$10.00
o3 Pro$20.00$80.00
Gemini 2.5 Pro Preview 06-05$1.25$10.00
R1 0528$0.50$2.15
Claude Sonnet 4$3.00$15.00
Claude Opus 4$15.00$75.00
Gemma 3n 4B$0.06$0.12
Mistral Medium 3$0.40$2.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
Llama Guard 4 12B$0.18$0.18
Qwen3 235B A22B$0.46$1.82
Qwen3 32B$0.08$0.28
Qwen3 8B$0.12$0.46
Qwen3 14B$0.12$0.24
Qwen3 30B A3B$0.12$0.50
o4 Mini$1.10$4.40
o3$2.00$8.00
o4 Mini High$1.10$4.40
GPT-4.1 Mini$0.40$1.60
GPT-4.1 Nano$0.10$0.40
GPT-4.1$2.00$8.00
Llama 4 Maverick$0.20$0.80
Llama 4 Scout$0.10$0.30
DeepSeek V3 0324$0.25$1.00
o1-pro$150.00$600.00
Mistral Small 3.1 24B$0.35$0.56
Gemma 3 12B$0.05$0.15
Command A$2.50$10.00
Gemma 3 4B$0.05$0.10
Reka Flash 3$0.10$0.20
Gemma 3 27B$0.08$0.45
GPT-4o-mini Search Preview$0.15$0.60
GPT-4o Search Preview$2.50$10.00
Skyfall 36B V2$0.55$0.80
Sonar Reasoning Pro$2.00$8.00
Sonar Pro$3.00$15.00
Sonar Deep Research$2.00$8.00
Saba$0.20$0.60
Claude 3.5 Sonnet v2$3.00$15.00
o3 Mini High$1.10$4.40
Gemini 2.0 Flash$0.10$0.40
Qwen-Plus$0.26$0.78
Qwen2.5 VL 72B Instruct$0.80$1.00
o3 Mini$1.10$4.40
Mistral Small 3$0.09$0.25
Sonar$1.00$1.00
R1 Distill Llama 70B$0.80$0.80
R1$0.70$2.50
DeepSeek R1$0.70$2.50
MiniMax-01$0.20$1.10
Phi 4$0.07$0.14
DeepSeek V3$0.32$0.89
o1$15.00$60.00
Command R7B (12-2024)$0.04$0.15
Mixtral 8x22B$0.50$1.00
Llama 3.3 70B Instruct$0.10$0.32
Llama 3.3 70B Instruct$0.10$0.32
Nova Lite 1.0$0.06$0.24
Nova Pro 1.0$0.80$3.20
Nova Micro 1.0$0.04$0.14
GPT-4o (2024-11-20)$2.50$10.00
Mistral Large 2407$2.00$6.00
Qwen2.5 Coder 32B Instruct$0.66$1.00
UnslopNemo 12B$0.40$0.40
Ministral 8B$0.11$0.11
Qwen2.5 7B Instruct$0.10$0.20
Inflection 3 Productivity$2.50$10.00
Inflection 3 Pi$2.50$10.00
Llama 3.2 3B Instruct$0.05$0.33
Llama 3.2 11B Vision Instruct$0.34$0.34
Llama 3.2 1B Instruct$0.03$0.20
Llama 3.2 11B Vision$0.34$0.34
Qwen2.5 72B Instruct$0.36$0.40
Command R (08-2024)$0.15$0.60
Hermes 3 70B Instruct$0.70$0.70
Hermes 3 405B Instruct$1.00$1.00
GPT-4o (2024-08-06)$2.50$10.00
Mistral Large 3$0.50$1.50
Llama 3.1 70B Instruct$0.40$0.40
Llama 3.1 8B Instruct$0.05$0.08
Llama 3.1 8B$0.04$0.04
Llama 3.1 405B$0.80$0.80
Mistral Nemo$0.02$0.03
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT-4o-mini$0.15$0.60
Gemma 2 27B$0.65$0.65
GPT-4o (2024-05-13)$5.00$15.00
GPT-4o$2.50$10.00
Llama 3 8B Instruct$0.14$0.14
Mixtral 8x22B Instruct$2.00$6.00
WizardLM-2 8x22B$0.62$0.62
GPT-4 Turbo$10.00$30.00
Command R+$2.50$10.00
Claude 3 Haiku$0.25$1.25
Command R$0.15$0.60
Mistral Large$2.00$6.00
GPT-4 Turbo Preview$10.00$30.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
Auto Router$0.00$0.00
GPT-3.5 Turbo Instruct$1.50$2.00
GPT-3.5 Turbo 16k$3.00$4.00
GPT-3.5 Turbo$0.50$1.50
GPT-4$30.00$60.00
SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.