arrow_backBack to research feed
agentsPublished: August 18, 2026

On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification

By Qinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang, Chien-Sheng Wu

Research TL;DR

"Memory-based self-improving agents are fragile: performance varies across runs and task order. Adding rubrics/feedback to memory partially closes gaps, but significant instability remains, urging multi-run, shuffled-order evaluation."

Abstract

Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening the scope of evaluation along two axes: (1) including multiple runs to quantify variance, and (2) randomly shuffling the tasks to investigate the effect of task order. Through these experiments, we make two observations that expose the fragility of current methods: First, agent evaluation is inherently noisy in complex environments and on multi-step tasks, and stacking a self-improving loop on top can further amplify this noise. Second, the agent's improvement is highly dependent on task order. Prior works often adopt default orderings that impose an implicit curriculum, acting as a hidden prerequisite for success. To better understand this fragility, we manually examine the agents' memory and hypothesize that task and environment underspecification contribute to this fragility. We validate this hypothesis by incorporating information that enables better specification, such as detailed rubrics and environment feedback, into the memory construction process. While this added information partially closes the performance degradation in previous experiments, significant gaps still remain, suggesting that other uncharacterized factors contribute to this fragility. Looking ahead, our work advocates for more rigorous evaluation protocols for self-improving agents by reporting results across multiple runs and stress-testing them under challenging conditions. Moreover, our findings on underspecification call for systems and interfaces that enable effective human oversight, preventing agents from failing in unforeseeable ways.

Technical Analysis & Implementation

Overview§

This paper re-evaluates memory-based self-improving agents—agents that accumulate experience in a textual memory bank and use it to improve on a stream of tasks. The authors find that standard evaluation protocols hide severe reliability issues: (1) results are noisy across identical runs, and (2) improvement depends heavily on task order. They trace part of the problem to task/environment underspecification and show that injecting richer information (rubrics, environment feedback) into memory helps but does not fully close the gap.

Methodology§

Two representative memory-based self-improving agents are re-run under two perturbations:

  • Multiple seeds/runs to quantify variance (previous work often reports a single run).
  • Random task shuffles to test sensitivity to ordering (prior work uses a fixed default order that may act as an implicit curriculum).

Evaluation is performed on multi-step agent benchmarks (e.g., ALFWorld, WebShop). The authors compare: (i) baseline agent without self-improvement, (ii) agent with memory bank built from successful trajectories, and (iii) variants that add rubrics (detailed evaluation criteria) and environment feedback into memory entries.

Key Findings§

  • High variance: Self-improving loops amplify existing evaluation noise. In complex multi-step tasks, the same configuration can show large performance swings across runs.
  • Task-order sensitivity: Random shuffling often destroys the improvement seen with default ordering, implying that the original ordering provided a hidden curriculum.
  • Underspecification: Manual inspection of memories reveals that agents struggle because task instructions and environment observations do not fully specify the goal or success criteria. Adding explicit rubrics and feedback partially recovers performance, but a residual gap remains.

Formalizing Underspecification§

Let $\mathcal{T}$ be a task distribution and $\theta$ the agent's parameters (including memory bank $M$). The agent's expected performance under an evaluation protocol $P$ is:

$$ J_P(\theta) = \mathbb{E}_{\tau \sim P} \left[ R(\text{Agent}(\theta, \tau)) \right], $$

where $\tau$ is a task ordering/seed. Standard protocols fix $\tau$ (e.g., default order, single seed), yielding a noisy estimate $\hat{J} = J_{\tau}(\theta)$. The paper shows that $\text{Var}(\hat{J})$ is large and that $J$ depends strongly on the ordering component of $\tau$.

To mitigate underspecification, each memory entry $m_i$ is augmented from $(s_i, a_i, r_i)$ to $(s_i, a_i, r_i, u_i)$, where $u_i$ contains rubric text or environment feedback. The agent then conditions on more complete specifications during inference.

Implementation Sketch§

The following pseudo-PyTorch code illustrates how memory entries are constructed and used in a self-improving loop:

class MemoryBank:
    def __init__(self, max_entries=50):
        self.entries = []
        self.max_entries = max_entries

    def add(self, trajectory, rubrics=None, env_feedback=None):
        # trajectory: list of (obs, action, reward)
        summary = f"Task: {trajectory.task}"
        if rubrics:
            summary += f"\nRubrics: {rubrics}"
        if env_feedback:
            summary += f"\nFeedback: {env_feedback}"
        self.entries.append(summary)
        if len(self.entries) > self.max_entries:
            self.entries.pop(0)

    def retrieve(self, query, k=5):
        # simple top-k by embedding similarity
        embs = model.encode(self.entries)
        q_emb = model.encode(query)
        scores = cosine_similarity(q_emb, embs)
        return [self.entries[i] for i in scores.topk(k).indices]

# Self-improving loop
agent = Agent()
memory = MemoryBank()
for task in shuffled_task_list:
    context = memory.retrieve(task.prompt)  # inject relevant past experience
    traj = agent.run(task, context)
    if traj.success:
        memory.add(traj, rubrics=task.rubrics, env_feedback=traj.feedback)

Implications§

  • Evaluation protocol: Report multiple seeds and shuffled task orders; avoid relying on a single default ordering.
  • System design: Human oversight interfaces must provide specifiability—agents should be able to query rubrics or receive structured feedback to reduce underspecification.
  • Open problem: Since added information only partially closes the gap, other factors (e.g., memory retrieval noise, catastrophic forgetting) likely contribute to fragility.
Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window400,000 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000047
Output Cost (Generated):$0.000279
Total Est. Cost:$0.000326
Context Window Capacity0.0310%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
GPT-5.4 Mini (batch)$0.38$2.25
GPT-5.4 Pro (batch)$15.00$90.00
Seed 1.6$0.25$2.00
MiniMax M3 (batch)$0.30$1.20
Claude Opus 4.8 (batch)$2.50$12.50
Gemini 3.5 Flash (batch)$0.75$4.50
Gemini 3.1 Flash Lite (batch)$0.13$0.75
GPT-5.4$2.50$15.00
GPT-5.4 (batch)$1.25$7.50
Muse Spark 1.2 Contributor$0.10$0.20
DeepSeek V4 Flash Vision Exp$0.22$0.66
Hy-MT2-1.8B$0.04$0.18
Hy-MT2-30B-A3B$0.07$0.29
Hy-MT2-7B$0.07$0.29
GLM 5.3$1.40$4.40
Gemini 3.7 Flash$0.38$1.88
Gemini 3.7 Flash (batch)$0.19$0.94
Seed 2.1 Turbo$0.50$2.50
Qwen3.8 2.4T A95B$2.00$6.00
Seed-2.0-Code$0.50$3.00
DeepSeek V4 Pro 0813$1.12$3.37
Grok 4.6$2.00$6.00
DeepSeek V4 Pro$0.40$0.79
Qwen2.5 Coder 32B Instruct$0.66$1.00
Lyria 3 Pro Preview$0.00$0.00
GPT-5.4 Nano (batch)$0.10$0.63
MiniMax M2.7$0.30$1.20
MiniMax-01$0.20$1.10
GLM 5.2 (batch)$1.40$4.40
Kimi K2.7 Code (batch)$0.95$4.00
Claude Fable 5 (batch)$5.00$25.00
Claude Opus 4.7 (batch)$2.50$12.50
Nemotron 3 Ultra (batch)$0.60$3.60
GPT-5.5 Pro (batch)$15.00$90.00
GPT-5.5 (batch)$2.50$15.00
Qwen3.8 27B$0.40$3.00
Nemotron 3.5 Lightning$0.08$0.20
Sakana Namazu$0.95$4.00
Solar Pro 4$0.03$0.12
Muse Glimmer 30B$0.35$1.50
Muse Spark 1.2$1.25$4.25
Qwen3.8 Max$2.00$6.00
DeepSeek V4 Flash 0731$0.08$0.18
Claude Opus 5 (batch)$2.50$12.50
o3 Mini High$1.10$4.40
MiniMax M1$0.55$2.20
Llama 3.3 70B Instruct$0.10$0.32
GPT-5.4 Nano$0.20$1.25
Gemini 3.6 Flash (batch)$0.38$1.88
Gemini 3.5 Flash Lite (batch)$0.15$1.25
Saba$0.20$0.60
GPT-5.2 (batch)$0.88$7.00
Qwen3 VL 8B Instruct$0.12$0.46
DeepSeek V4 Pro 0423$0.40$0.79
DeepSeek V4 Flash 0423$0.05$0.10
Lyria 3 Clip Preview$0.00$0.00
GPT-5.6 Luna Pro (batch)$0.10$0.60
GPT-5.6 Luna (batch)$0.10$0.60
GPT-5.6 Terra Pro (batch)$1.00$6.00
Gemini 3 Flash Preview (batch)$0.25$1.50
GPT-5.6 Terra (batch)$1.00$6.00
GPT-5.6 Sol Pro$2.00$10.00
Hermes 3 405B Instruct$1.00$1.00
GPT-5 Pro (batch)$7.50$60.00
Ministral 8B$0.11$0.11
GPT-4o-mini$0.15$0.60
Claude Opus 4.5 (batch)$2.50$12.50
Qwen3.7 Flash$0.03$0.13
Claude Opus Latest$5.00$25.00
Gemini 3.1 Pro Preview (batch)$1.00$6.00
GPT-5.2 Pro (batch)$10.50$84.00
Claude Sonnet 4.6 (batch)$1.50$7.50
Claude Opus 4.6 (batch)$2.50$12.50
GPT-5.1 (batch)$0.63$5.00
Claude Haiku 4.5 (batch)$0.50$2.50
GPT-5.6 Terra Pro$2.00$12.00
Claude Sonnet 4.5$3.00$15.00
GPT-5.6 Sol Pro (batch)$1.00$5.00
GPT-5.6 Sol (batch)$1.00$5.00
Kimi K3$3.00$15.00
GPT-5 Codex (batch)$0.63$5.00
Qwen3 Next 80B A3B Thinking$0.15$1.20
GPT-5 (batch)$0.63$5.00
GPT-5 Mini (batch)$0.13$1.00
Grok 4.5$2.00$6.00
Claude Sonnet 5$2.00$10.00
o3 Pro (batch)$10.00$40.00
Claude Sonnet 5 (batch)$1.00$5.00
Claude Sonnet 4.5 (batch)$1.50$7.50
Qwen2.5 VL 72B Instruct$0.80$1.00
Claude Opus 4$15.00$75.00
Claude Opus 5 (Fast)$10.00$50.00
Claude Opus 5$5.00$25.00
GPT-5.6 Sol$2.00$10.00
o3 Mini (batch)$0.55$2.20
Claude Fable Latest$10.00$50.00
Hermes 4 70B$0.13$0.40
GPT-5 Nano (batch)$0.03$0.20
Claude Opus 4.1 (batch)$7.50$37.50
Gemini 2.5 Flash Lite (batch)$0.05$0.20
Gemini 2.5 Flash (batch)$0.15$1.25
GPT-4.1 Mini (batch)$0.20$0.80
Gemini 2.5 Pro (batch)$0.63$5.00
o4 Mini High (batch)$0.55$2.20
o3 (batch)$1.00$4.00
o4 Mini (batch)$0.55$2.20
GPT-4.1 (batch)$1.00$4.00
GPT-4.1 Nano (batch)$0.05$0.20
o1-pro (batch)$75.00$300.00
o3 Mini High (batch)$0.55$2.20
Llama 3.1 8B Instruct$0.05$0.08
GPT-4o-mini (batch)$0.07$0.30
GPT-3.5 Turbo (batch)$0.25$0.75
GPT-4o (batch)$1.25$5.00
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Mixtral 8x22B Instruct$2.00$6.00
Gemma 2 27B$0.65$0.65
GPT-4 Turbo (batch)$5.00$15.00
Anthropic Claude Haiku Latest$1.00$5.00
o1 (batch)$7.50$30.00
Qwen2.5 7B Instruct$0.10$0.20
Morph V3 Large$0.90$1.90
Command R7B (12-2024)$0.04$0.15
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Nemotron 3 Ultra$0.60$3.60
Qwen3.6 Flash$0.19$1.13
Inflection 3 Productivity$2.50$10.00
GLM 5.2$0.97$3.04
Kimi K2.7 Code$0.67$3.40
GLM 4.5V$0.60$1.80
Kimi K2.6$0.95$4.00
Claude Opus 4.5$5.00$25.00
GPT-4o (2024-11-20)$2.50$10.00
MiniMax M3$0.30$1.20
GPT-5.4 Image 2$8.00$15.00
o1$15.00$60.00
Step 3.7 Flash$0.20$1.15
Claude Opus 4.8 (Fast)$10.00$50.00
Gemma 4 26B A4B$0.07$0.34
Claude Sonnet 4$3.00$15.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
o3$2.00$8.00
o4 Mini$1.10$4.40
MoonshotAI Kimi Latest$2.60$13.00
Google Gemini Flash Latest$0.38$1.88
Grok 4.20$1.25$2.50
GPT-4 Turbo Preview$10.00$30.00
Claude Opus 4.8$5.00$25.00
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
Claude Haiku 4.5$1.00$5.00
Gemini 3.5 Flash$1.50$9.00
Laguna S 2.1$0.09$0.18
Gemini 3.5 Flash Lite$0.30$2.50
Muse Spark 1.1$1.25$4.25
GPT-5.6 Luna Pro$0.20$1.20
Claude Opus 4.7 (Fast)$30.00$150.00
Reka Flash 3$0.10$0.20
GPT-4o (2024-08-06)$2.50$10.00
GPT-5.5 Pro$30.00$180.00
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Claude Sonnet 4.6$3.00$15.00
GPT-5.6 Terra$2.00$12.00
Gemini 3.6 Flash$0.75$3.75
Hy3$0.13$0.53
Laguna XS 2.1$0.06$0.12
Qwen3 VL 32B Instruct$0.10$0.42
Gemini 3.1 Flash$0.25$1.50
GPT-5.6 Luna$0.20$1.20
GLM 4.6V$0.30$0.90
Codestral 2508$0.30$0.90
Command R (08-2024)$0.15$0.60
Llama 4 Scout$0.10$0.30
Qwen2.5 72B Instruct$0.36$0.40
KAT-Coder-Air V2.5$0.15$0.60
GPT-4o (2024-05-13)$5.00$15.00
Nex-N2-Mini$0.03$0.10
Fugu Ultra$5.00$30.00
Qwen3 235B A22B Instruct 2507$0.09$0.55
Ministral 3 8B 2512$0.15$0.15
Llama 4 Maverick$0.20$0.80
GPT-4o Search Preview$2.50$10.00
Gemma 3 27B$0.08$0.45
KAT-Coder-Pro V2.5$0.74$2.96
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
Nova 2 Lite$0.30$2.50
o1-pro$150.00$600.00
Grok 4.3$1.25$2.50
Granite 4.1 8B$0.05$0.10
Qwen3 VL 8B Thinking$0.18$2.10
Llama 3 8B Instruct$0.14$0.14
Laguna M.1$0.20$0.40
Qwen-Plus$0.26$0.78
Mistral Large$2.00$6.00
Nex-N2-Pro$0.25$1.00
Qwen3.7 Max$1.48$4.42
Grok Build 0.1$1.00$2.00
Qwen3 Next 80B A3B Instruct$0.10$1.10
Sonar Pro$3.00$15.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
Claude 3.5 Sonnet v2$3.00$15.00
Sonar Deep Research$2.00$8.00
Claude 3 Haiku$0.25$1.25
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
Mistral Medium 3.5$1.50$7.50
MiMo-V2.5$0.14$0.28
Qwen3 VL 235B A22B Thinking$0.40$4.00
Qwen3 VL 235B A22B Instruct$0.21$1.90
Sonar$1.00$1.00
GPT-5 Codex$1.25$10.00
Google Gemini Pro Latest$2.00$12.00
Anthropic Claude Sonnet Latest$2.00$10.00
Qwen3.5 Plus 2026-04-20$0.30$1.80
Qwen3.6 Plus$0.33$1.95
Grok 4.20 Multi-Agent$1.25$2.50
Qwen3 30B A3B Instruct 2507$0.05$0.19
MiMo-V2.5-Pro$0.43$0.87
GLM 5.1$0.97$3.04
Gemma 4 31B$0.10$0.34
GPT-5.4 Pro$30.00$180.00
Gemini 3.1 Flash Lite Preview$0.25$1.50
GLM 4.5 Air$0.13$0.85
KAT-Coder-Pro V2$0.30$1.20
Reka Edge$0.10$0.10
GLM 5 Turbo$1.20$4.00
Nemotron 3 Super$0.09$0.40
Seed-2.0-Lite$0.25$2.00
Seed-2.0-Mini$0.10$0.40
Qwen3.5-122B-A10B$0.26$2.08
Qwen3 Max Thinking$0.78$3.90
Morph V3 Fast$0.80$1.20
GPT-4o$2.50$10.00
GPT-5.3 Chat$1.75$14.00
Qwen3.5 Plus 2026-02-15$0.26$1.56
MiniMax M2-her$0.30$1.20
Gemini 2.5 Pro Preview 06-05$1.25$10.00
GPT-3.5 Turbo 16k$3.00$4.00
Qwen3.5-35B-A3B$0.25$1.25
Qwen3.5-27B$0.20$1.56
GPT-5.5$5.00$30.00
GPT-5.2-Codex$1.75$14.00
Mistral Small 4$0.15$0.60
Mistral Small 3$0.07$0.20
GPT-5.3-Codex$1.75$14.00
Qwen3.5 397B A17B$0.39$2.34
Gemini 3 Flash Preview$0.50$3.00
o4 Mini High$1.10$4.40
GPT-3.5 Turbo$0.50$1.50
Claude Fable 5$10.00$50.00
Qwen3.7 Plus$0.32$1.28
GLM 5$0.60$1.92
Qwen3 Coder Next$0.12$0.80
UI-TARS 7B$0.10$0.20
Devstral 2 2512$0.40$2.00
o3 Pro$20.00$80.00
Mistral Small 3.2 24B$0.07$0.20
Gemma 3n 4B$0.06$0.12
Mistral Large 2407$2.00$6.00
Gemini 3.1 Pro Preview$2.00$12.00
GPT-5.2 Chat$1.75$14.00
GPT-5.1-Codex-Max$1.25$10.00
gpt-oss-20b$0.03$0.13
Claude Opus 4.1$15.00$75.00
WizardLM-2 8x22B$0.62$0.62
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.45$2.25
Qwen Plus 0728 (thinking)$0.26$0.78
GPT-5 Mini$0.25$2.00
Mistral Large 3$0.50$1.50
Qwen3 8B$0.12$0.46
GPT-4$30.00$60.00
o4 Mini Deep Research$2.00$8.00
GLM 5V Turbo$1.20$4.00
DeepSeek V3.2$0.26$0.38
Llama 3.3 70B Instruct$0.10$0.32
Yi-Lightning$0.15$0.30
GPT Audio Mini$0.60$2.40
Ministral 3 14B 2512$0.20$0.20
Qwen Plus 0728$0.26$0.78
DeepSeek V3 0324$0.25$1.00
Voxtral Small 24B 2507$0.10$0.30
Qwen3 Coder 30B A3B Instruct$0.07$0.28
Mistral Nemo$0.02$0.03
GPT-5.4 Mini$0.75$4.50
GPT Audio$2.50$10.00
GPT-4o-mini (2024-07-18)$0.15$0.60
Qwen3.5-Flash$0.07$0.26
MiniMax M2.5$0.27$1.08
GPT-5.1 Chat$1.25$10.00
Solar Pro 3$0.15$0.60
GPT-5.1-Codex$1.25$10.00
Kimi K2 0711$0.57$2.30
Mistral Medium 3$0.40$2.00
Mistral Small 3.1 24B$0.35$0.56
Command R$0.15$0.60
Claude Opus 4.6$5.00$25.00
GLM 4.7 Flash$0.06$0.40
GPT-5$1.25$10.00
Claude Opus 4.7$5.00$25.00
Gemini 3.1 Pro$2.00$12.00
GPT-4.1 Nano$0.10$0.40
Llama 3.2 11B Vision$0.34$0.34
Qwen3.6 35B A3B$0.14$1.00
Hy3 preview$0.18$0.60
Seed 1.6 Flash$0.07$0.30
Gemini 2.5 Pro$1.25$10.00
ERNIE 4.0$1.20$2.40
Qwen3.6 Max Preview$1.03$6.16
Nemotron 3 Nano 30B A3B$0.05$0.20
MiniMax M2$0.26$1.02
Nova Lite 1.0$0.06$0.24
o3 Deep Research$10.00$40.00
Qwen 2.5-Coder 32B$0.35$0.70
GLM 4.7$0.40$1.75
Ministral 3 3B 2512$0.10$0.10
GPT-5.1$1.25$10.00
GLM 4.5$0.60$2.20
R1 0528$0.50$2.15
Llama Guard 4 12B$0.18$0.18
Doubao Pro$0.80$1.60
Qwen3 30B A3B$0.12$0.50
GLM 4.6$0.50$2.00
Kimi K2 Thinking$0.60$2.50
Gemma 3 4B$0.05$0.10
Sonar Pro Search$3.00$15.00
Qwen3 Max$0.78$3.90
Qwen3 235B A22B Thinking 2507$0.23$2.30
Qwen3.5-9B$0.10$0.15
Mercury 2$0.25$0.75
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Thinking$0.20$2.40
Qwen3 Coder 480B A35B$0.30$1.00
Gemini 2.5 Flash Lite$0.10$0.40
Qwen3 VL 30B A3B Instruct$0.13$0.52
o3 Mini$1.10$4.40
Llama 3.1 405B$0.80$0.80
Palmyra X5$0.60$6.00
gpt-oss-safeguard-20b$0.07$0.30
Mixtral 8x22B$0.50$1.00
Llama 3.2 1B Instruct$0.03$0.20
GPT-5.2 Pro$21.00$168.00
Granite 4.0 Micro$0.02$0.11
GPT-5 Pro$15.00$120.00
DeepSeek V3.2 Exp$0.27$0.41
Hunyuan A13B Instruct$0.14$0.57
Llama 3.1 8B$0.04$0.04
Nova Premier 1.0$2.50$12.50
DeepSeek V3.1 Terminus$0.27$1.00
Kimi K2 0905$0.60$2.50
GPT-4o-mini Search Preview$0.15$0.60
Qwen3.6 27B$0.60$3.60
GPT-5.2$1.75$14.00
Gemma 3 12B$0.05$0.15
GPT-5 Chat$1.25$10.00
DeepSeek R1$0.70$2.50
Sonar Reasoning Pro$2.00$8.00
GPT-5 Image Mini$2.50$2.00
Qwen3 32B$0.08$0.28
Qwen 2.5 72B$0.40$0.80
Command R+$2.50$10.00
Qwen3 30B A3B Thinking 2507$0.20$2.40
Grok 4.20$1.25$2.50
R1 Distill Llama 70B$0.80$0.80
DeepSeek V3$0.26$1.03
Llama 3.2 3B Instruct$0.05$0.33
GPT-3.5 Turbo Instruct$1.50$2.00
MiniMax M2.1$0.30$1.20
GPT-5.1-Codex-Mini$0.25$2.00
GPT-5 Image$10.00$10.00
Hermes 4 405B$1.00$3.00
DeepSeek V4 Flash$0.05$0.10
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Gemini 2.5 Flash$0.30$2.50
Qwen3 14B$0.12$0.24
Llama 3.1 70B Instruct$0.40$0.40
GPT-4 Turbo$10.00$30.00
DeepSeek V3.1$0.55$1.65
Qwen3 Coder Plus$0.65$3.25
Qwen3 Coder Flash$0.20$0.97
Mistral Medium 3.1$0.40$2.00
GPT-4.1 Mini$0.40$1.60
R1$0.70$2.50
Nova Pro 1.0$0.80$3.20
Mistral Large 3 2512$0.50$1.50
ERNIE 4.5 VL 424B A47B$0.42$1.25
Jamba Large 1.7$2.00$8.00
Llama 4 Maverick$0.20$0.80
Phi 4$0.07$0.14
Nova Micro 1.0$0.04$0.14
Mistral Large 2$0.60$1.80
GPT-5 Nano$0.05$0.40
Llama 3.2 11B Vision Instruct$0.34$0.34
Inflection 3 Pi$2.50$10.00
Gemini 2.0 Flash$0.10$0.40
Hunyuan Pro$0.60$1.20
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
gpt-oss-120b$0.04$0.17
Qwen3 235B A22B$0.46$1.82
GPT-4.1$2.00$8.00
Command A$2.50$10.00
Hermes 3 70B Instruct$0.70$0.70
SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.