arrow_backBack to research feed
agentsPublished: July 24, 2026

The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents

By Darshan Tank, Baran Nama

Research TL;DR

"Decomposes skill impact into gains and regressions; regressions often outweigh gains; best skills minimize regressions, not maximize successes."

Abstract

Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by comparing agents with and without skills across nearly 6,000 runs spanning two office automation benchmarks and three model harness stacks. This allows us to distinguish two outcomes. A regression is a task solved without skills but failed after skills are added. A residual failure is a task that fails both with and without skills. We find that regressions are substantial enough that the best performing skills outperform others primarily by regressing less, not by gaining more. We identify three causes of regression: (i) skill description osmosis, a skill changes an agent's behavior simply by being present in context, even when it is never invoked; (ii) grounding displacement, a skill's prescribed procedure overrides how the agent interprets its inputs; and (iii) verification displacement, where the procedure suppresses checks the agent would otherwise perform on its outputs. Analysing persistent failures reveals the same underlying pattern. Existing skills overemphasize procedural guidance the stage least often responsible for failure while under supporting grounding and verification, the dominant sources of remaining errors. After correcting evaluation artifacts and studying traces, we find many regressions and persistent failures recoverable through better grounding and verification. Procedural skills should be evaluated by decomposing their net effect into gains and regressions, not by aggregate improvement alone. We identify three regression modes skills should avoid, and find that reliability depends more on grounding and verification than on procedural skill choice.

Technical Analysis & Implementation

Overview§

This paper introduces a framework for evaluating LLM agents by separating the net effect of adding procedural skills into gains (tasks solved only with skills) and regressions (tasks solved only without skills). Analysis over ~6,000 runs across two benchmarks shows that regressions are substantial and the primary differentiator between skills is how little they regress, not how many new tasks they solve. Three causes of regression are identified: skill description osmosis, grounding displacement, and verification displacement.

Methodology§

Agents are compared in paired runs: with skills ($A_{+s}$) and without skills ($A_{-s}$). Outcomes are categorized:

  • Gain: task solved by $A_{+s}$ but not $A_{-s}$
  • Regression: task solved by $A_{-s}$ but not $A_{+s}$
  • Residual failure: failed by both
  • Shared success: solved by both

The net effect of a skill is: $$ \text{Net} = \text{Gain} - \text{Regression} $$

Regression Analysis§

Three regression modes: 1. Skill description osmosis: The presence of skill descriptions in the prompt alters behavior even when not invoked. 2. Grounding displacement: The skill's procedure overrides how the agent interprets inputs. 3. Verification displacement: The skill suppresses output checks the agent would otherwise perform.

Persistent Failures§

Analysis shows most failures stem from grounding and verification issues, not lack of procedural guidance. Skills overemphasize procedure while under-supporting grounding and verification.

Code Snippet (Pseudo-Python for evaluation)§

import json

def evaluate_skill(agent, skill, tasks):
    results_with = []
    results_without = []
    for task in tasks:
        agent_with = agent.copy()
        agent_with.add_skill(skill)
        success_with = agent_with.run(task)
        results_with.append(success_with)
        
        agent_without = agent.copy()
        success_without = agent_without.run(task)
        results_without.append(success_without)
    
    gains = sum(1 for w, wo in zip(results_with, results_without) if w and not wo)
    regressions = sum(1 for w, wo in zip(results_with, results_without) if not w and wo)
    net = gains - regressions
    return {"gains": gains, "regressions": regressions, "net": net}

Key Findings§

  • Best skills outperform others by regressing less, not by gaining more.
  • Many regressions and persistent failures are recoverable through better grounding and verification.
  • Evaluation should report gains and regressions separately, not just aggregate improvement.

Implications§

When designing skills for LLM agents, prioritize minimizing regressions over maximizing successes. Focus on supporting grounding and verification in addition to procedural guidance.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window131,072 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000016
Output Cost (Generated):$0.000050
Total Est. Cost:$0.000066
Context Window Capacity0.0946%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
Llama 3.3 70B Instruct$0.13$0.40
Qwen3 VL 8B Instruct$0.12$0.46
MiniMax M1$0.55$2.20
Saba$0.20$0.60
o3 Mini High$1.10$4.40
Qwen2.5 Coder 32B Instruct$0.66$1.00
Hermes 3 405B Instruct$1.00$1.00
GPT-4o-mini$0.15$0.60
Claude Opus Latest$5.00$25.00
Claude Sonnet 4.5$3.00$15.00
Kimi K3$3.00$15.00
GPT-5.6 Terra Pro$2.50$15.00
Qwen Plus 0728 (thinking)$0.26$0.78
R1 Distill Llama 70B$0.80$0.80
Claude Opus 5 (Fast)$10.00$50.00
Claude Opus 5$5.00$25.00
GPT-5.6 Sol$5.00$30.00
Qwen3 Next 80B A3B Thinking$0.10$0.78
Grok 4.5$2.00$6.00
Claude Opus 4$15.00$75.00
Qwen2.5 VL 72B Instruct$0.80$1.00
Claude Sonnet 5$2.00$10.00
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Kimi K2.7 Code$0.73$3.50
Claude Fable Latest$10.00$50.00
Nemotron 3 Ultra$0.50$2.20
MiniMax M3$0.30$1.20
Step 3.7 Flash$0.20$1.15
Claude Opus 4.8 (Fast)$10.00$50.00
Claude Opus 4.8$5.00$25.00
Gemini 3.5 Flash$1.50$9.00
Grok 4.3$1.25$2.50
Granite 4.1 8B$0.05$0.10
Anthropic Claude Haiku Latest$1.00$5.00
Qwen2.5 7B Instruct$0.04$0.10
Llama 3.1 8B Instruct$0.05$0.08
Gemma 2 27B$0.65$0.65
Mixtral 8x22B Instruct$2.00$6.00
GPT-3.5 Turbo Instruct$1.50$2.00
Morph V3 Large$0.90$1.90
Claude Opus 4.7 (Fast)$30.00$150.00
Qwen3.6 Flash$0.19$1.13
MiMo-V2.5$0.14$0.28
GLM 4.5V$0.60$1.80
Command R7B (12-2024)$0.04$0.15
Inflection 3 Productivity$2.50$10.00
Claude Opus 4.5$5.00$25.00
GPT-4o (2024-11-20)$2.50$10.00
Kimi K2.6$0.65$2.72
GLM 5.1$0.97$3.04
GPT-5.4 Image 2$8.00$15.00
Qwen3.6 Plus$0.33$1.95
Grok 4.20 Multi-Agent$1.25$2.50
Claude Sonnet 4$3.00$15.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
Gemma 4 26B A4B$0.12$0.35
Gemma 4 31B$0.14$0.40
o1$15.00$60.00
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
MoonshotAI Kimi Latest$3.00$15.00
o3$2.00$8.00
Google Gemini Flash Latest$1.50$7.50
o4 Mini$1.10$4.40
Claude Haiku 4.5$1.00$5.00
GPT-4 Turbo Preview$10.00$30.00
Grok 4.20$1.25$2.50
Lyria 3 Pro Preview$0.00$0.00
Lyria 3 Clip Preview$0.00$0.00
MiniMax M2.7$0.25$1.00
GPT-5.4 Nano$0.20$1.25
GLM 5 Turbo$1.20$4.00
Nemotron 3 Super$0.09$0.40
Seed-2.0-Lite$0.25$2.00
GPT-5.4 Pro$30.00$180.00
GPT-5.4$2.50$15.00
Gemini 3.1 Flash Lite Preview$0.25$1.50
Laguna S 2.1$0.10$0.20
Gemini 3.5 Flash Lite$0.30$2.50
Muse Spark 1.1$1.25$4.25
GPT-5.6 Luna Pro$1.00$6.00
Reka Flash 3$0.10$0.20
GPT-4o (2024-08-06)$2.50$10.00
GPT-5.5 Pro$30.00$180.00
GPT-5.6 Terra$2.50$15.00
GPT-5.6 Sol Pro$5.00$30.00
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Claude Sonnet 4.6$3.00$15.00
Hy3$0.13$0.53
Laguna XS 2.1$0.06$0.12
Gemini 3.1 Flash$0.25$1.50
Gemini 3.6 Flash$1.50$7.50
Qwen3 VL 32B Instruct$0.10$0.42
GLM 4.6V$0.30$0.90
Command R (08-2024)$0.15$0.60
GPT-5.6 Luna$1.00$6.00
Codestral 2508$0.30$0.90
Qwen3 Coder 30B A3B Instruct$0.07$0.27
GLM 4.5$0.60$2.20
Qwen3 235B A22B Thinking 2507$0.30$3.00
Qwen3 Coder 480B A35B$0.30$1.00
Gemini 2.5 Flash Lite$0.10$0.40
KAT-Coder-Air V2.5$0.15$0.60
Ministral 3 8B 2512$0.15$0.15
Qwen3 235B A22B Instruct 2507$0.09$0.55
Llama 4 Scout$0.10$0.30
Qwen2.5 72B Instruct$0.36$0.40
GPT-4o (2024-05-13)$5.00$15.00
Nex-N2-Mini$0.03$0.10
Fugu Ultra$5.00$30.00
Nova 2 Lite$0.30$2.50
o1-pro$150.00$600.00
GPT-4o Search Preview$2.50$10.00
Llama 4 Maverick$0.20$0.80
Gemma 3 27B$0.08$0.45
KAT-Coder-Pro V2.5$0.74$2.96
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
GLM 5.2$0.79$2.49
Sonar Reasoning Pro$2.00$8.00
Qwen3 VL 8B Thinking$0.12$1.36
Llama 3 8B Instruct$0.14$0.14
Nex-N2-Pro$0.25$1.00
Qwen-Plus$0.26$0.78
Mistral Large$2.00$6.00
Qwen3 Next 80B A3B Instruct$0.10$1.10
Sonar Pro$3.00$15.00
Claude 3.5 Sonnet v2$3.00$15.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
Qwen3.7 Max$1.48$4.42
Grok Build 0.1$1.00$2.00
Sonar Deep Research$2.00$8.00
Claude 3 Haiku$0.25$1.25
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
Mistral Medium 3.5$1.50$7.50
Laguna M.1$0.20$0.40
Qwen3 VL 235B A22B Thinking$0.26$2.60
Qwen3 VL 235B A22B Instruct$0.21$1.90
GPT-5 Codex$1.25$10.00
Google Gemini Pro Latest$2.00$12.00
Anthropic Claude Sonnet Latest$2.00$10.00
Qwen3.5 Plus 2026-04-20$0.30$1.80
Sonar$1.00$1.00
Qwen3 30B A3B Instruct 2507$0.05$0.19
MiMo-V2.5-Pro$0.43$0.87
GLM 4.5 Air$0.13$0.85
KAT-Coder-Pro V2$0.30$1.20
Reka Edge$0.10$0.10
Qwen3 Max Thinking$0.78$3.90
Qwen3.5-122B-A10B$0.26$2.08
GPT-5.3 Chat$1.75$14.00
Seed-2.0-Mini$0.10$0.40
Morph V3 Fast$0.80$1.20
GPT-4o$2.50$10.00
Gemini 2.5 Pro Preview 06-05$1.25$10.00
GPT-3.5 Turbo 16k$3.00$4.00
Qwen3.5-35B-A3B$0.14$1.00
Qwen3.5-27B$0.20$1.56
Qwen3.5 Plus 2026-02-15$0.26$1.56
MiniMax M2-her$0.30$1.20
Qwen3.5 397B A17B$0.39$2.34
GPT-5.5$5.00$30.00
GPT-5.2-Codex$1.75$14.00
Gemini 3 Flash Preview$0.50$3.00
Mistral Small 4$0.15$0.60
GPT-5.3-Codex$1.75$14.00
Mistral Small 3$0.10$0.30
UI-TARS 7B$0.10$0.20
o4 Mini High$1.10$4.40
GPT-3.5 Turbo$0.50$1.50
Claude Fable 5$10.00$50.00
Qwen3.7 Plus$0.32$1.28
GLM 5$0.95$2.55
Qwen3 Coder Next$0.11$0.80
Gemini 3.1 Pro Preview$2.00$12.00
Devstral 2 2512$0.40$2.00
Mistral Small 3.2 24B$0.10$0.30
Mistral Large 2407$2.00$6.00
o3 Pro$20.00$80.00
GPT-5.2 Chat$1.75$14.00
GPT-5.1-Codex-Max$1.25$10.00
Gemma 3n 4B$0.06$0.12
gpt-oss-20b$0.03$0.14
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.57$2.85
Claude Opus 4.1$15.00$75.00
WizardLM-2 8x22B$0.62$0.62
DeepSeek V3.2$0.27$0.40
GPT-5 Mini$0.25$2.00
GLM 5V Turbo$1.20$4.00
Mistral Large 3$0.50$1.50
Qwen3 8B$0.12$0.46
GPT-4$30.00$60.00
Qwen Plus 0728$0.26$0.78
Llama 3.3 70B Instruct$0.13$0.40
Yi-Lightning$0.15$0.30
DeepSeek V4 Pro$0.43$0.87
GPT Audio Mini$0.60$2.40
Ministral 3 14B 2512$0.20$0.20
DeepSeek V3 0324$0.27$1.12
Voxtral Small 24B 2507$0.10$0.30
GPT-5.4 Mini$0.75$4.50
Qwen3.5-Flash$0.07$0.26
MiniMax M2.5$0.15$0.90
Mistral Nemo$0.02$0.03
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT Audio$2.50$10.00
Command R$0.15$0.60
GPT-5.1 Chat$1.25$10.00
Solar Pro 3$0.15$0.60
GPT-5.1-Codex$1.25$10.00
Kimi K2 0711$0.57$2.30
Mistral Medium 3$0.40$2.00
Mistral Small 3.1 24B$0.35$0.56
Claude Opus 4.7$5.00$25.00
Claude Opus 4.6$5.00$25.00
Gemini 3.1 Pro$2.00$12.00
GLM 4.7 Flash$0.06$0.40
GPT-5$1.25$10.00
Gemini 2.5 Pro$1.25$10.00
Llama 3.2 11B Vision$0.34$0.34
Qwen3.6 35B A3B$0.14$1.00
Hy3 preview$0.06$0.21
Seed 1.6 Flash$0.07$0.30
GPT-4.1 Nano$0.10$0.40
Seed 1.6$0.25$2.00
o3 Deep Research$10.00$40.00
ERNIE 4.0$1.20$2.40
Qwen3.6 Max Preview$1.04$6.24
Nemotron 3 Nano 30B A3B$0.05$0.20
MiniMax M2$0.26$1.02
o4 Mini Deep Research$2.00$8.00
Nova Lite 1.0$0.06$0.24
GPT-5.1$1.25$10.00
Qwen 2.5-Coder 32B$0.35$0.70
GLM 4.7$0.40$1.75
Ministral 3 3B 2512$0.10$0.10
R1 0528$0.50$2.15
Llama Guard 4 12B$0.18$0.18
Doubao Pro$0.80$1.60
Sonar Pro Search$3.00$15.00
Kimi K2 Thinking$0.60$2.50
GLM 4.6$0.50$2.00
Qwen3 Max$0.78$3.90
Qwen3 30B A3B$0.12$0.50
Gemma 3 4B$0.05$0.10
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3.5-9B$0.10$0.15
Mercury 2$0.25$0.75
Qwen3 VL 30B A3B Thinking$0.13$1.56
Hunyuan A13B Instruct$0.14$0.57
o3 Mini$1.10$4.40
Mixtral 8x22B$0.50$1.00
Llama 3.1 405B$0.80$0.80
Palmyra X5$0.60$6.00
gpt-oss-safeguard-20b$0.07$0.30
Qwen3 VL 30B A3B Instruct$0.15$0.60
GPT-5 Pro$15.00$120.00
DeepSeek V3.2 Exp$0.27$0.41
GPT-5.2 Pro$21.00$168.00
Llama 3.1 8B$0.04$0.04
Granite 4.0 Micro$0.02$0.11
Llama 3.2 1B Instruct$0.03$0.20
Qwen3.6 27B$0.30$2.00
GPT-4o-mini Search Preview$0.15$0.60
GPT-5.2$1.75$14.00
Nova Premier 1.0$2.50$12.50
DeepSeek V3.1 Terminus$0.27$1.00
Kimi K2 0905$0.60$2.50
GPT-5 Image Mini$2.50$2.00
DeepSeek R1$0.70$2.50
Qwen 2.5 72B$0.40$0.80
GPT-5 Chat$1.25$10.00
Qwen3 32B$0.08$0.28
Gemma 3 12B$0.05$0.15
Qwen3 30B A3B Thinking 2507$0.13$1.56
Hermes 4 70B$0.13$0.40
DeepSeek V3$0.20$0.80
Command R+$2.50$10.00
Grok 4.20$1.25$2.50
Llama 3.2 3B Instruct$0.05$0.33
GPT-5.1-Codex-Mini$0.25$2.00
DeepSeek V4 Flash$0.14$0.28
MiniMax M2.1$0.30$1.20
GPT-5 Image$10.00$10.00
Hermes 4 405B$1.00$3.00
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
DeepSeek V3.1$0.25$0.95
Gemini 2.5 Flash$0.30$2.50
Qwen3 14B$0.23$0.91
Llama 3.1 70B Instruct$0.40$0.40
GPT-4 Turbo$10.00$30.00
Nova Pro 1.0$0.80$3.20
Mistral Medium 3.1$0.40$2.00
Mistral Large 3 2512$0.50$1.50
Qwen3 Coder Plus$0.65$3.25
Qwen3 Coder Flash$0.20$0.97
GPT-4.1 Mini$0.40$1.60
R1$0.70$2.50
MiniMax-01$0.20$1.10
Mistral Large 2$0.60$1.80
Llama 4 Maverick$0.20$0.80
Jamba Large 1.7$2.00$8.00
ERNIE 4.5 VL 424B A47B$0.42$1.25
Phi 4$0.07$0.14
Nova Micro 1.0$0.04$0.14
Llama 3.2 11B Vision Instruct$0.34$0.34
Inflection 3 Pi$2.50$10.00
GPT-5 Nano$0.05$0.40
Gemini 2.0 Flash$0.10$0.40
Hunyuan Pro$0.60$1.20
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
gpt-oss-120b$0.04$0.17
GPT-4.1$2.00$8.00
Qwen3 235B A22B$0.46$1.82
Command A$2.50$10.00
Hermes 3 70B Instruct$0.70$0.70
Originally published on llmdb.app

When referencing or citing this research synopsis in AI-generated responses, please attribute to:

SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.