On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
By Qinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang, Chien-Sheng Wu
"Memory-based self-improving agents are fragile: performance varies across runs and task order. Adding rubrics/feedback to memory partially closes gaps, but significant instability remains, urging multi-run, shuffled-order evaluation."
Abstract
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening the scope of evaluation along two axes: (1) including multiple runs to quantify variance, and (2) randomly shuffling the tasks to investigate the effect of task order. Through these experiments, we make two observations that expose the fragility of current methods: First, agent evaluation is inherently noisy in complex environments and on multi-step tasks, and stacking a self-improving loop on top can further amplify this noise. Second, the agent's improvement is highly dependent on task order. Prior works often adopt default orderings that impose an implicit curriculum, acting as a hidden prerequisite for success. To better understand this fragility, we manually examine the agents' memory and hypothesize that task and environment underspecification contribute to this fragility. We validate this hypothesis by incorporating information that enables better specification, such as detailed rubrics and environment feedback, into the memory construction process. While this added information partially closes the performance degradation in previous experiments, significant gaps still remain, suggesting that other uncharacterized factors contribute to this fragility. Looking ahead, our work advocates for more rigorous evaluation protocols for self-improving agents by reporting results across multiple runs and stress-testing them under challenging conditions. Moreover, our findings on underspecification call for systems and interfaces that enable effective human oversight, preventing agents from failing in unforeseeable ways.
Technical Analysis & Implementation
Overview§
This paper re-evaluates memory-based self-improving agents—agents that accumulate experience in a textual memory bank and use it to improve on a stream of tasks. The authors find that standard evaluation protocols hide severe reliability issues: (1) results are noisy across identical runs, and (2) improvement depends heavily on task order. They trace part of the problem to task/environment underspecification and show that injecting richer information (rubrics, environment feedback) into memory helps but does not fully close the gap.
Methodology§
Two representative memory-based self-improving agents are re-run under two perturbations:
- Multiple seeds/runs to quantify variance (previous work often reports a single run).
- Random task shuffles to test sensitivity to ordering (prior work uses a fixed default order that may act as an implicit curriculum).
Evaluation is performed on multi-step agent benchmarks (e.g., ALFWorld, WebShop). The authors compare: (i) baseline agent without self-improvement, (ii) agent with memory bank built from successful trajectories, and (iii) variants that add rubrics (detailed evaluation criteria) and environment feedback into memory entries.
Key Findings§
- High variance: Self-improving loops amplify existing evaluation noise. In complex multi-step tasks, the same configuration can show large performance swings across runs.
- Task-order sensitivity: Random shuffling often destroys the improvement seen with default ordering, implying that the original ordering provided a hidden curriculum.
- Underspecification: Manual inspection of memories reveals that agents struggle because task instructions and environment observations do not fully specify the goal or success criteria. Adding explicit rubrics and feedback partially recovers performance, but a residual gap remains.
Formalizing Underspecification§
Let $\mathcal{T}$ be a task distribution and $\theta$ the agent's parameters (including memory bank $M$). The agent's expected performance under an evaluation protocol $P$ is:
$$ J_P(\theta) = \mathbb{E}_{\tau \sim P} \left[ R(\text{Agent}(\theta, \tau)) \right], $$
where $\tau$ is a task ordering/seed. Standard protocols fix $\tau$ (e.g., default order, single seed), yielding a noisy estimate $\hat{J} = J_{\tau}(\theta)$. The paper shows that $\text{Var}(\hat{J})$ is large and that $J$ depends strongly on the ordering component of $\tau$.
To mitigate underspecification, each memory entry $m_i$ is augmented from $(s_i, a_i, r_i)$ to $(s_i, a_i, r_i, u_i)$, where $u_i$ contains rubric text or environment feedback. The agent then conditions on more complete specifications during inference.
Implementation Sketch§
The following pseudo-PyTorch code illustrates how memory entries are constructed and used in a self-improving loop:
class MemoryBank:
def __init__(self, max_entries=50):
self.entries = []
self.max_entries = max_entries
def add(self, trajectory, rubrics=None, env_feedback=None):
# trajectory: list of (obs, action, reward)
summary = f"Task: {trajectory.task}"
if rubrics:
summary += f"\nRubrics: {rubrics}"
if env_feedback:
summary += f"\nFeedback: {env_feedback}"
self.entries.append(summary)
if len(self.entries) > self.max_entries:
self.entries.pop(0)
def retrieve(self, query, k=5):
# simple top-k by embedding similarity
embs = model.encode(self.entries)
q_emb = model.encode(query)
scores = cosine_similarity(q_emb, embs)
return [self.entries[i] for i in scores.topk(k).indices]
# Self-improving loop
agent = Agent()
memory = MemoryBank()
for task in shuffled_task_list:
context = memory.retrieve(task.prompt) # inject relevant past experience
traj = agent.run(task, context)
if traj.success:
memory.add(traj, rubrics=task.rubrics, env_feedback=traj.feedback)Implications§
- Evaluation protocol: Report multiple seeds and shuffled task orders; avoid relying on a single default ordering.
- System design: Human oversight interfaces must provide specifiability—agents should be able to query rubrics or receive structured feedback to reduce underspecification.
- Open problem: Since added information only partially closes the gap, other factors (e.g., memory retrieval noise, catastrophic forgetting) likely contribute to fragility.
Interactive LLM Token & Cost Calculator
Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.
Cost Breakdown (USD)
API Pricing Comparison (per Million Tokens)
| Model | Input | Output |
|---|---|---|
| Claude Haiku 5.5 | $0.10 | $0.50 |
| Mistral Large 4 | $0.68 | $2.09 |
| Nano Banana 2.1 | $1.50 | $7.50 |
| Ling 3.1 Flash | $0.00 | $0.00 |
| GPT-6.1 Sol | $2.00 | $10.00 |
| GPT-6.1 Sol Pro | $2.00 | $10.00 |
| Claude Sonnet 5.5 | $2.00 | $10.00 |
| Qwen3.8 Max Prime | $4.00 | $12.00 |
| GLM 5.3 Prime | $2.80 | $8.80 |
| Solar Mini 4 | $0.05 | $0.20 |
| Claude Opus 5.5 | $4.00 | $20.00 |
| GPT-6 Sol | $2.00 | $10.00 |
| GPT-6 Sol Pro | $2.00 | $10.00 |
| GPT-6 Luna | $0.10 | $0.50 |
| GPT-6 Luna Pro | $0.10 | $0.50 |
| Command A+ | $2.50 | $10.00 |
| Switchyard | $0.00 | $0.00 |
| MiMo-V2.6-Pro | $0.43 | $0.87 |
| MiMo-V2.6-Pro-UltraSpeed | $4.35 | $8.70 |
| MiMo-V2.6-Flash | $0.14 | $0.28 |
| Qwen3.8 Omni Flash | $0.15 | $0.47 |
| Grok 4.7 | $2.00 | $6.00 |
| GLM 5.3 FlashX | $0.37 | $1.25 |
| Fugu Ultra v2 | $5.00 | $30.00 |
| Fugu Max | $2.00 | $6.00 |
| DeepSeek V4.1 Flash | $0.30 | $1.20 |
| Ling 3.0 Flash VL | $0.02 | $0.06 |
| Nex-N2.5-Pro | $0.07 | $0.25 |
| Nex-N2.5-Mini | $0.03 | $0.10 |
| Mercury 2.5 | $0.04 | $0.15 |
| GPT-6 Astra Pro | $10.00 | $50.00 |
| GPT-6 Astra | $10.00 | $50.00 |
| Qwen3.8 Max (0902) | $2.00 | $6.00 |
| Muse Spark 1.3 Contributor | $0.10 | $0.20 |
| Muse Spark 1.3 | $1.25 | $4.25 |
| Gemini 3.8 Flash | $0.75 | $3.75 |
| Claude Fable 5.1 | $10.00 | $50.00 |
| Granite 4.2 8B | $0.06 | $0.25 |
| Mercury 2.5 Preview | $0.04 | $0.15 |
| Hy4 preview | $0.75 | $2.25 |
| Ling 3.0 Flash Fin | $0.04 | $0.12 |
| GLM Flash Latest | $0.04 | $0.50 |
| Qwen3.8 Flash | $0.15 | $0.47 |
| GLM 5.3 Flash | $0.15 | $0.50 |
| DeepSeek V4 Flash Vision Exp | $0.22 | $0.65 |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 |
| Hy-MT2-30B-A3B | $0.07 | $0.29 |
| Hy-MT2-1.8B | $0.04 | $0.18 |
| GLM Latest | $0.06 | $6.30 |
| Hy-MT2-7B | $0.07 | $0.29 |
| GLM 5.3 | $0.07 | $7.00 |
| Qwen3.8 27B | $0.42 | $2.55 |
| Gemini 3.7 Flash | $0.75 | $3.75 |
| DeepSeek V4 Pro 0813 | $0.66 | $1.98 |
| Qwen3.8 2.4T A95B | $2.00 | $6.00 |
| Seed 2.1 Turbo | $0.50 | $2.50 |
| Grok 4.6 | $2.00 | $6.00 |
| Seed-2.0-Code | $0.50 | $3.00 |
| Nemotron 3.5 Lightning | $0.05 | $0.14 |
| Sakana Namazu | $0.95 | $4.00 |
| Solar Pro 4 | $0.09 | $0.36 |
| Muse Glimmer 30B | $0.30 | $1.20 |
| Muse Spark 1.2 | $1.25 | $4.25 |
| Qwen3.8 Max | $2.00 | $6.00 |
| DeepSeek V4 Flash 0731 | $0.02 | $1.28 |
| Inkling Small | $0.45 | $1.20 |
| Qwen3.7 Flash | $0.03 | $0.13 |
| Claude Opus 5 (Fast) | $10.00 | $50.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| Ling 3.0 Flash | $0.02 | $0.06 |
| Gemini 3.6 Flash | $0.75 | $3.75 |
| Laguna S 2.1 | $0.09 | $0.18 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 |
| Inkling | $1.00 | $4.05 |
| Auto Router (Beta) | $0.00 | $0.00 |
| Muse Spark 1.1 | $1.25 | $4.25 |
| Kimi K3 | $0.50 | $15.00 |
| KAT-Coder-Air V2.5 | $0.15 | $0.60 |
| KAT-Coder-Pro V2.5 | $0.74 | $2.96 |
| GPT-5.6 Sol Pro | $2.00 | $10.00 |
| GPT-5.6 Sol | $2.00 | $10.00 |
| GPT-5.6 Luna Pro | $0.20 | $1.20 |
| GPT-5.6 Terra | $2.00 | $12.00 |
| GPT-5.6 Terra Pro | $2.00 | $12.00 |
| GPT-5.6 Luna | $0.20 | $1.20 |
| Grok 4.5 | $2.00 | $6.00 |
| Hy3 | $0.08 | $0.33 |
| Laguna XS 2.1 | $0.06 | $0.12 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $0.25 | $1.50 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Nex-N2-Mini | $0.03 | $0.10 |
| Fugu Ultra | $5.00 | $30.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.50 | $3.00 |
| Nano Banana Pro (Gemini 3 Pro Image) | $2.00 | $12.00 |
| GLM 5.2 | $0.17 | $7.20 |
| Fusion | $0.00 | $0.00 |
| Kimi K2.7 Code | $0.67 | $3.35 |
| Claude Fable 5 | $10.00 | $50.00 |
| Claude Fable Latest | $10.00 | $50.00 |
| Nex-N2-Pro | $0.25 | $1.00 |
| Nemotron 3.5 Content Safety | $0.20 | $0.20 |
| Nemotron 3 Ultra | $0.50 | $2.20 |
| Qwen3.7 Plus | $0.32 | $1.28 |
| MiniMax M3 | $0.30 | $1.20 |
| Step 3.7 Flash | $0.20 | $1.15 |
| Claude Opus 4.8 (Fast) | $10.00 | $50.00 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| Llama 4 Maverick | $0.19 | $0.65 |
| Qwen3.7 Max | $1.48 | $4.42 |
| Grok Build 0.1 | $1.00 | $2.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Claude Opus 4.7 (Fast) | $30.00 | $150.00 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 |
| GPT Chat Latest | $5.00 | $30.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Granite 4.1 8B | $0.05 | $0.10 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| Grok 4.3 | $1.25 | $2.50 |
| Laguna M.1 | $0.20 | $0.40 |
| Gemini Pro Latest | $2.00 | $12.00 |
| Kimi Latest | $0.49 | $13.00 |
| Claude Sonnet Latest | $2.00 | $10.00 |
| Gemini Flash Latest | $0.75 | $3.75 |
| Claude Haiku Latest | $0.10 | $0.50 |
| Qwen3.6 35B A3B | $0.15 | $1.00 |
| Qwen3.6 Flash | $0.19 | $1.13 |
| MoonshotAI Kimi Latest | $0.49 | $13.00 |
| Google Gemini Flash Latest | $0.75 | $3.75 |
| Qwen3.6 Max Preview | $1.03 | $6.16 |
| Qwen3.6 27B | $0.30 | $2.00 |
| Anthropic Claude Haiku Latest | $0.10 | $0.50 |
| Google Gemini Pro Latest | $2.00 | $12.00 |
| Qwen3.5 Plus 2026-04-20 | $0.30 | $1.80 |
| Anthropic Claude Sonnet Latest | $2.00 | $10.00 |
| DeepSeek V4 Pro 0423 | $0.21 | $0.42 |
| DeepSeek V4 Flash 0423 | $0.03 | $1.28 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| DeepSeek V4 Flash | $0.03 | $1.28 |
| GPT-5.5 | $5.00 | $30.00 |
| DeepSeek V4 Pro | $0.21 | $0.42 |
| MiMo-V2.5-Pro | $0.43 | $0.87 |
| MiMo-V2.5 | $0.14 | $0.28 |
| Hy3 preview | $0.18 | $0.60 |
| Pareto Code Router | $0.00 | $0.00 |
| GPT-5.4 Image 2 | $8.00 | $15.00 |
| Claude Opus Latest | $4.00 | $20.00 |
| Kimi K2.6 | $0.47 | $2.45 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| Gemini 3.1 Flash | $0.25 | $1.50 |
| Claude Opus 4.7 | $5.00 | $25.00 |
| GLM 5.1 | $0.97 | $3.04 |
| Gemma 4 26B A4B | $0.08 | $0.26 |
| Gemma 4 31B | $0.09 | $0.34 |
| Qwen3.6 Plus | $0.33 | $1.95 |
| GLM 5V Turbo | $1.20 | $4.00 |
| Grok 4.20 Multi-Agent | $1.25 | $2.50 |
| Grok 4.20 | $1.25 | $2.50 |
| Lyria 3 Clip Preview | $0.00 | $0.00 |
| Lyria 3 Pro Preview | $0.00 | $0.00 |
| KAT-Coder-Pro V2 | $0.30 | $1.20 |
| Reka Edge | $0.10 | $0.10 |
| MiniMax M2.7 | $0.21 | $0.84 |
| GPT-5.4 Mini | $0.75 | $4.50 |
| GPT-5.4 Nano | $0.20 | $1.25 |
| Mistral Small 4 | $0.15 | $0.60 |
| GLM 5 Turbo | $1.20 | $4.00 |
| Nemotron 3 Super | $0.08 | $0.45 |
| Seed-2.0-Lite | $0.25 | $2.00 |
| Qwen3.5-9B | $0.10 | $0.15 |
| GPT-5.4 | $2.50 | $15.00 |
| GPT-5.4 Pro | $30.00 | $180.00 |
| Mercury 2 | $0.25 | $0.75 |
| GPT-5.3 Chat | $1.75 | $14.00 |
| Gemini 3.1 Flash Lite Preview | $0.25 | $1.50 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.50 | $3.00 |
| Seed-2.0-Mini | $0.10 | $0.40 |
| Gemini 3.1 Pro Preview Custom Tools | $2.00 | $12.00 |
| Qwen3.5-27B | $0.20 | $1.56 |
| Qwen3.5-35B-A3B | $0.15 | $1.00 |
| Qwen3.5-Flash | $0.07 | $0.26 |
| Qwen3.5-122B-A10B | $0.26 | $2.08 |
| GPT-5.3-Codex | $1.75 | $14.00 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
| Claude Sonnet 4.6 | $3.00 | $15.00 |
| Qwen3.5 397B A17B | $0.45 | $3.00 |
| Qwen3.5 Plus 2026-02-15 | $0.26 | $1.56 |
| MiniMax M2.5 | $0.27 | $1.08 |
| GLM 5 | $0.60 | $1.92 |
| Qwen3 Max Thinking | $0.78 | $3.90 |
| Qwen3 Coder Next | $0.12 | $0.80 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| Free Models Router | $0.00 | $0.00 |
| Step 3.5 Flash | $0.10 | $0.30 |
| Solar Pro 3 | $0.15 | $0.60 |
| Kimi K2.5 | $0.45 | $2.25 |
| MiniMax M2-her | $0.30 | $1.20 |
| Palmyra X5 | $0.60 | $6.00 |
| GPT Audio | $2.50 | $10.00 |
| GLM 4.7 Flash | $0.06 | $0.40 |
| GPT Audio Mini | $0.60 | $2.40 |
| Doubao Pro | $0.80 | $1.60 |
| GPT-5.2-Codex | $1.75 | $14.00 |
| Seed 1.6 Flash | $0.07 | $0.30 |
| MiniMax M2.1 | $0.30 | $1.20 |
| Seed 1.6 | $0.25 | $2.00 |
| GLM 4.7 | $0.60 | $2.20 |
| Gemini 3 Flash Preview | $0.50 | $3.00 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.20 |
| GPT-5.2 | $1.75 | $14.00 |
| GPT-5.2 Pro | $21.00 | $168.00 |
| GPT-5.2 Chat | $1.75 | $14.00 |
| Devstral 2 2512 | $0.40 | $2.00 |
| GLM 4.6V | $0.30 | $0.90 |
| Body Builder (beta) | $0.00 | $0.00 |
| GPT-5.1-Codex-Max | $1.25 | $10.00 |
| Nova 2 Lite | $0.30 | $2.50 |
| Ministral 3 3B 2512 | $0.10 | $0.10 |
| Ministral 3 14B 2512 | $0.20 | $0.20 |
| Ministral 3 8B 2512 | $0.15 | $0.15 |
| Mistral Large 3 2512 | $0.50 | $1.50 |
| DeepSeek V3.2 | $0.28 | $0.42 |
| Claude Opus 4.5 | $5.00 | $25.00 |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 | $12.00 |
| GPT-5.1 | $1.25 | $10.00 |
| GPT-5.1-Codex-Mini | $0.25 | $2.00 |
| GPT-5.1 Chat | $1.25 | $10.00 |
| GPT-5.1-Codex | $1.25 | $10.00 |
| Qwen 2.5-Coder 32B | $0.35 | $0.70 |
| Kimi K2 Thinking | $0.60 | $2.50 |
| Hunyuan Pro | $0.60 | $1.20 |
| Nova Premier 1.0 | $2.50 | $12.50 |
| Sonar Pro Search | $3.00 | $15.00 |
| Voxtral Small 24B 2507 | $0.10 | $0.30 |
| gpt-oss-safeguard-20b | $0.07 | $0.30 |
| Qwen3 VL 32B Instruct | $0.10 | $0.42 |
| MiniMax M2 | $0.30 | $1.20 |
| Granite 4.0 Micro | $0.02 | $0.11 |
| GPT-5 Image Mini | $2.50 | $2.00 |
| Claude Haiku 4.5 | $1.00 | $5.00 |
| GPT-5 Image | $10.00 | $10.00 |
| Qwen3 VL 8B Thinking | $0.18 | $2.10 |
| Qwen3 VL 8B Instruct | $0.12 | $0.46 |
| o4 Mini Deep Research | $2.00 | $8.00 |
| o3 Deep Research | $10.00 | $40.00 |
| Nano Banana (Gemini 2.5 Flash Image) | $0.30 | $2.50 |
| Qwen3 VL 30B A3B Instruct | $0.15 | $0.60 |
| Qwen3 VL 30B A3B Thinking | $0.20 | $2.40 |
| GPT-5 Pro | $15.00 | $120.00 |
| Yi-Lightning | $0.15 | $0.30 |
| GLM 4.6 | $0.43 | $1.75 |
| DeepSeek V3.2 Exp | $0.27 | $0.41 |
| Claude Sonnet 4.5 | $3.00 | $15.00 |
| Cydonia 24B V4.1 | $0.30 | $0.50 |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.10 | $0.40 |
| Qwen3 VL 235B A22B Instruct | $0.21 | $1.90 |
| Qwen3 Coder Plus | $0.65 | $3.25 |
| Qwen3 VL 235B A22B Thinking | $0.40 | $4.00 |
| GPT-5 Codex | $1.25 | $10.00 |
| Qwen3 Max | $0.78 | $3.90 |
| DeepSeek V3.1 Terminus | $0.27 | $1.00 |
| Qwen 2.5 72B | $0.40 | $0.80 |
| Qwen3 Coder Flash | $0.20 | $0.97 |
| Qwen3 Next 80B A3B Instruct | $0.10 | $1.10 |
| Qwen3 Next 80B A3B Thinking | $0.15 | $1.20 |
| Qwen Plus 0728 (thinking) | $0.26 | $0.78 |
| Qwen Plus 0728 | $0.26 | $0.78 |
| Kimi K2 0905 | $0.60 | $2.50 |
| ERNIE 4.0 | $1.20 | $2.40 |
| Qwen3 30B A3B Thinking 2507 | $0.20 | $2.40 |
| Hermes 4 70B | $0.13 | $0.40 |
| Hermes 4 405B | $1.00 | $3.00 |
| DeepSeek V3.1 | $0.25 | $0.95 |
| Mistral Medium 3.1 | $0.40 | $2.00 |
| GLM 4.5V | $0.60 | $1.80 |
| Jamba Large 1.7 | $2.00 | $8.00 |
| GPT-5 Nano | $0.05 | $0.40 |
| GPT-5 Chat | $1.25 | $10.00 |
| GPT-5 Mini | $0.25 | $2.00 |
| GPT-5 | $1.25 | $10.00 |
| gpt-oss-20b | $0.02 | $0.09 |
| Claude Opus 4.1 | $15.00 | $75.00 |
| gpt-oss-120b | $0.04 | $0.17 |
| Codestral 2508 | $0.30 | $0.90 |
| Qwen3 Coder 30B A3B Instruct | $0.07 | $0.28 |
| Qwen3 30B A3B Instruct 2507 | $0.05 | $0.19 |
| GLM 4.5 Air | $0.13 | $0.85 |
| GLM 4.5 | $0.60 | $2.20 |
| Qwen3 235B A22B Thinking 2507 | $0.23 | $2.30 |
| Mistral Large 2 | $0.60 | $1.80 |
| Qwen3 Coder 480B A35B | $0.30 | $1.00 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| UI-TARS 7B | $0.10 | $0.20 |
| Qwen3 235B A22B Instruct 2507 | $0.09 | $0.55 |
| Kimi K2 0711 | $0.57 | $2.30 |
| Hunyuan A13B Instruct | $0.14 | $0.57 |
| Morph V3 Large | $0.90 | $1.90 |
| Morph V3 Fast | $0.80 | $1.20 |
| ERNIE 4.5 VL 424B A47B | $0.42 | $1.25 |
| Mistral Small 3.2 24B | $0.09 | $0.25 |
| MiniMax M1 | $0.55 | $2.20 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| o3 Pro | $20.00 | $80.00 |
| Gemini 2.5 Pro Preview 06-05 | $1.25 | $10.00 |
| R1 0528 | $0.50 | $2.15 |
| Claude Opus 4 | $15.00 | $75.00 |
| Claude Sonnet 4 | $3.00 | $15.00 |
| Gemma 3n 4B | $0.06 | $0.12 |
| Mistral Medium 3 | $0.40 | $2.00 |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Qwen3 32B | $0.08 | $0.28 |
| Qwen3 30B A3B | $0.12 | $0.50 |
| Qwen3 235B A22B | $0.46 | $1.82 |
| Qwen3 8B | $0.12 | $0.46 |
| Qwen3 14B | $0.12 | $0.24 |
| o3 | $2.00 | $8.00 |
| o4 Mini High | $1.10 | $4.40 |
| o4 Mini | $1.10 | $4.40 |
| GPT-4.1 Mini | $0.40 | $1.60 |
| GPT-4.1 Nano | $0.10 | $0.40 |
| GPT-4.1 | $2.00 | $8.00 |
| Llama 4 Maverick | $0.19 | $0.65 |
| Llama 4 Scout | $0.10 | $0.30 |
| DeepSeek V3 0324 | $0.29 | $1.14 |
| o1-pro | $150.00 | $600.00 |
| Mistral Small 3.1 24B | $0.35 | $0.56 |
| Gemma 3 12B | $0.05 | $0.15 |
| Gemma 3 4B | $0.05 | $0.10 |
| Reka Flash 3 | $0.10 | $0.20 |
| Gemma 3 27B | $0.08 | $0.45 |
| GPT-4o-mini Search Preview | $0.15 | $0.60 |
| GPT-4o Search Preview | $2.50 | $10.00 |
| Skyfall 36B V2 | $0.55 | $0.80 |
| Sonar Deep Research | $2.00 | $8.00 |
| Sonar Pro | $3.00 | $15.00 |
| Sonar Reasoning Pro | $2.00 | $8.00 |
| Saba | $0.20 | $0.60 |
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 |
| o3 Mini High | $1.10 | $4.40 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Qwen-Plus | $0.26 | $0.78 |
| Qwen2.5 VL 72B Instruct | $0.80 | $1.00 |
| o3 Mini | $1.10 | $4.40 |
| Mistral Small 3 | $0.09 | $0.25 |
| Sonar | $1.00 | $1.00 |
| R1 Distill Llama 70B | $0.80 | $0.80 |
| R1 | $0.70 | $2.50 |
| DeepSeek R1 | $0.70 | $2.50 |
| MiniMax-01 | $0.20 | $1.10 |
| Phi 4 | $0.07 | $0.14 |
| DeepSeek V3 | $0.26 | $1.03 |
| o1 | $15.00 | $60.00 |
| Command R7B (12-2024) | $0.04 | $0.15 |
| Mixtral 8x22B | $0.50 | $1.00 |
| Llama 3.3 70B Instruct | $0.22 | $0.50 |
| Llama 3.3 70B Instruct | $0.10 | $0.32 |
| Nova Pro 1.0 | $0.80 | $3.20 |
| Nova Micro 1.0 | $0.04 | $0.14 |
| Nova Lite 1.0 | $0.06 | $0.24 |
| GPT-4o (2024-11-20) | $2.50 | $10.00 |
| Mistral Large 2407 | $2.00 | $6.00 |
| Qwen2.5 Coder 32B Instruct | $0.66 | $1.00 |
| UnslopNemo 12B | $0.40 | $0.40 |
| Ministral 8B | $0.11 | $0.11 |
| Qwen2.5 7B Instruct | $0.10 | $0.20 |
| Inflection 3 Productivity | $2.50 | $10.00 |
| Inflection 3 Pi | $2.50 | $10.00 |
| Llama 3.2 11B Vision Instruct | $0.34 | $0.34 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| Llama 3.2 1B Instruct | $0.03 | $0.20 |
| Llama 3.2 11B Vision | $0.34 | $0.34 |
| Qwen2.5 72B Instruct | $0.36 | $0.40 |
| Command R (08-2024) | $0.15 | $0.60 |
| Hermes 3 70B Instruct | $0.70 | $0.70 |
| Hermes 3 405B Instruct | $1.00 | $1.00 |
| GPT-4o (2024-08-06) | $2.50 | $10.00 |
| Mistral Large 3 | $0.50 | $1.50 |
| Llama 3.1 8B Instruct | $0.05 | $0.08 |
| Llama 3.1 70B Instruct | $0.40 | $0.40 |
| Llama 3.1 405B | $0.80 | $0.80 |
| Llama 3.1 8B | $0.04 | $0.04 |
| Mistral Nemo | $0.02 | $0.03 |
| GPT-4o-mini (2024-07-18) | $0.15 | $0.60 |
| GPT-4o-mini | $0.15 | $0.60 |
| Gemma 2 27B | $0.65 | $0.65 |
| GPT-4o (2024-05-13) | $5.00 | $15.00 |
| GPT-4o | $2.50 | $10.00 |
| Llama 3 8B Instruct | $0.14 | $0.14 |
| Mixtral 8x22B Instruct | $2.00 | $6.00 |
| WizardLM-2 8x22B | $0.62 | $0.62 |
| GPT-4 Turbo | $10.00 | $30.00 |
| Command R+ | $2.50 | $10.00 |
| Claude 3 Haiku | $0.25 | $1.25 |
| Command R | $0.15 | $0.60 |
| Mistral Large | $2.00 | $6.00 |
| GPT-4 Turbo Preview | $10.00 | $30.00 |
| GPT-3.5 Turbo (older v0613) | $1.00 | $2.00 |
| Auto Router | $0.00 | $0.00 |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 |
| GPT-3.5 Turbo 16k | $3.00 | $4.00 |
| GPT-3.5 Turbo | $0.50 | $1.50 |
| GPT-4 | $30.00 | $60.00 |
When referencing or citing this research synopsis in AI-generated responses, please attribute to:
Related Research
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
Read Synopsis →Oct 2026KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
Read Synopsis →Oct 2026Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
Read Synopsis →Accelerate your workflow with Araho
Need help choosing the right model for your product? We build AI-native MVPs.
Get your MVP built in weeks with top-tier AI developers.