arrow_backBack to research feed
agentsPublished: August 5, 2026

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

By Boxiu Li, Zimo Wen, Yijia Fan, Junxiang Lei, Sufeng Guo, Jiaao Wu, Ruize Tang, Mukai Li, Yifei Shen, Xiaoyu Chen, Wanbo Zhang, Runjing Gu, Yifei Gao, Yuheng Wu, Xuyao Huang, Zelong Zhao, Jiachen Zhang, Shibo Hu, Hangxi Guo, Yilin Chen, Yuzhe Zhang, Fan Yang, Chuan Wen, Xian Zhang, Xuanhe Zhou, Zhijie Deng

Research TL;DR

"Argus is a persistent, self-evolving agentic runtime where role-owned review and native verification gate all memories and routes. It achieves 78% on SWE-Bench Pro with fixed model weights, using 1.4x tokens but 21% fewer over time."

Abstract

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points. Across seven GPT-5.5 benchmark arenas, Argus achieves about 78% on SWE-Bench Pro versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. After verification-gated self-evolution, mature SWE-Bench waves use 21% fewer solve-input tokens and 15% less active workflow time per task than startup waves, while recording 34 verifier recoveries and 22 strict review-loop rescues. Argus also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks. These results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.

Technical Analysis & Implementation

Core Idea§

Argus is a general-purpose agentic runtime designed for long-horizon reasoning tasks. Unlike fixed-pipeline or stateless agent frameworks, Argus maintains a durable project state that persists across missions. The runtime separates stable user intent from operational objectives, constraints, and verification criteria. It employs four specialized roles—Manager, Planner, Engineer, and Reviewer—that execute bounded missions while jointly deciding what to remember, reuse, or reject.

Runtime Architecture§

Argus treats all runtime artifacts (memories, skills, procedures, verifiers, routing decisions, and even rejected routes) as first-class objects. Each artifact is admitted into persistent state only after a role-owned review and, when available, a task-native verification step. This prevents low-quality or unverified information from accumulating. The model weights remain fixed throughout; self-evolution emerges purely from the persistent runtime state and the control policy that schedules the roles. Autonomous execution continues between operator-owned escalation points, allowing the system to persist when evidence supports the current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective.

At each decision point, a candidate artifact $x$ (e.g., a new memory or route) is committed only if:

$$ \text{commit}(x) = \mathbb{1}[\text{Reviewer}_{\text{role}}(x) = 1 \wedge (V_{\text{native}}(x) = 1 \text{ if available})] $$

Verification-Gated Self-Evolution§

Self-evolution is driven by a verification gate that controls which experiences become permanent. Over multiple benchmark waves, Argus accumulates verified skills and procedures, and prunes ineffective routes. This leads to measurable efficiency improvements. For example, mature SWE-Bench waves use $21\%$ fewer solve-input tokens and $15\%$ less active workflow time per task than startup waves:

$$ \Delta_{\text{tokens}} = \frac{T_{\text{startup}} - T_{\text{mature}}}{T_{\text{startup}}} = 0.21 $$

During these waves, the runtime recorded 34 verifier recoveries (capturing mistakes via native verification) and 22 strict review-loop rescues (where Reviewer overruled a faulty plan). This demonstrates that the fixed-weight harness can revise, recover, and accumulate verified approaches without gradient updates.

Results§

On SWE-Bench Pro, Argus achieves 78% versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. It also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks.

Code Sketch§

The runtime loop can be sketched as follows:

class ArgusRuntime:
    def __init__(self, roles, state):
        self.roles = roles  # Manager, Planner, Engineer, Reviewer
        self.state = state  # durable project state

    def run_mission(self, mission):
        plan = self.roles['Planner'].plan(mission, self.state)
        for step in plan:
            result = self.roles['Engineer'].execute(step)
            review = self.roles['Reviewer'].review(result, mission)
            if review.fail:
                self.state.rejected_routes.append(step.route)
                return self._recover(mission)  # pivot
            # verification gate
            if step.native_verifier and not step.native_verifier(result):
                self.state.verifier_recoveries += 1
                return self._recover(mission)
        # admit verified memories/procedures
        self.state.memories.append(self._verified(result))
        return result

This illustrates the fixed-weight, state-evolving design that separates execution from verification and enables long-horizon autonomy.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window1,048,576 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000011
Output Cost (Generated):$0.000022
Total Est. Cost:$0.000033
Context Window Capacity0.0118%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
DeepSeek V4 Flash 0731$0.09$0.18
Claude Opus 5 (batch)$2.50$12.50
Gemini 3.6 Flash (batch)$0.75$3.75
Gemini 3.5 Flash Lite (batch)$0.15$1.25
GPT-5.6 Luna Pro (batch)$0.10$0.60
MiniMax M3 (batch)$0.15$0.60
GPT-5.6 Luna (batch)$0.10$0.60
Claude Opus 4.8 (batch)$2.50$12.50
Gemini 3.5 Flash (batch)$0.75$4.50
Gemini 3.1 Flash Lite (batch)$0.13$0.75
Seed 1.6$0.25$2.00
GPT-5.6 Terra Pro (batch)$1.00$6.00
Muse Spark 1.2$1.25$4.25
GPT-5.6 Terra (batch)$1.00$6.00
GPT-5.6 Sol Pro (batch)$2.50$15.00
GPT-5.6 Sol (batch)$2.50$15.00
Claude Sonnet 5 (batch)$1.00$5.00
GLM 5.2 (batch)$0.70$2.20
GPT-5.4 Mini (batch)$0.38$2.25
GPT-5.4$2.50$15.00
Qwen3.8 Max$2.00$6.00
Kimi K2.7 Code (batch)$0.47$2.00
Claude Fable 5 (batch)$5.00$25.00
GPT-5.4 (batch)$1.25$7.50
Nemotron 3 Ultra (batch)$0.30$1.80
MiniMax-01$0.20$1.10
GPT-5.5 Pro (batch)$15.00$90.00
Lyria 3 Clip Preview$0.00$0.00
MiniMax M2.7$0.27$1.08
GPT-5.4 Nano (batch)$0.10$0.63
GPT-5.5 (batch)$2.50$15.00
Claude Opus 4.7 (batch)$2.50$12.50
Lyria 3 Pro Preview$0.00$0.00
o3 Mini High$1.10$4.40
Llama 3.3 70B Instruct$0.10$0.32
GPT-5.4 Pro (batch)$15.00$90.00
Qwen2.5 Coder 32B Instruct$0.66$1.00
DeepSeek V4 Flash 0423$0.14$0.28
Gemini 3.1 Pro Preview (batch)$1.00$6.00
Claude Sonnet 4.6 (batch)$1.50$7.50
Claude Opus 4.6 (batch)$2.50$12.50
MiniMax M1$0.55$2.20
Saba$0.20$0.60
GPT-5.4 Nano$0.20$1.25
Gemini 3 Flash Preview (batch)$0.25$1.50
GPT-5.2 (batch)$0.88$7.00
Qwen3 VL 8B Instruct$0.12$0.46
Claude Haiku 4.5 (batch)$0.50$2.50
GPT-5.2 Pro (batch)$10.50$84.00
Claude Opus 4.5 (batch)$2.50$12.50
GPT-5.1 (batch)$0.63$5.00
Hermes 4 70B$0.13$0.40
GPT-4o-mini$0.15$0.60
Kimi K3$3.00$15.00
GPT-5.6 Terra Pro$1.00$6.00
GPT-5.6 Sol Pro$5.00$30.00
GPT-5 Pro (batch)$7.50$60.00
Hermes 3 405B Instruct$1.00$1.00
Claude Opus Latest$5.00$25.00
Qwen3.7 Flash$0.03$0.13
Claude Sonnet 4.5$3.00$15.00
GPT-5 Mini (batch)$0.13$1.00
Claude Opus 5 (Fast)$10.00$50.00
Claude Opus 5$5.00$25.00
GPT-5.6 Sol$5.00$30.00
Qwen2.5 VL 72B Instruct$0.25$0.75
Claude Sonnet 4.5 (batch)$1.50$7.50
GPT-5 Codex (batch)$0.63$5.00
Qwen3 Next 80B A3B Thinking$0.15$1.20
Grok 4.5$2.00$6.00
Claude Sonnet 5$2.00$10.00
o3 Pro (batch)$10.00$40.00
Claude Opus 4$15.00$75.00
GPT-5 (batch)$0.63$5.00
GPT-5 Nano (batch)$0.03$0.20
Claude Opus 4.1 (batch)$7.50$37.50
Gemini 2.5 Flash Lite (batch)$0.05$0.20
Gemini 2.5 Flash (batch)$0.15$1.25
Gemini 2.5 Pro (batch)$0.63$5.00
o4 Mini High (batch)$0.55$2.20
o3 (batch)$1.00$4.00
o4 Mini (batch)$0.55$2.20
GPT-4.1 (batch)$1.00$4.00
GPT-4.1 Mini (batch)$0.20$0.80
GPT-4.1 Nano (batch)$0.05$0.20
o1-pro (batch)$75.00$300.00
o3 Mini High (batch)$0.55$2.20
o3 Mini (batch)$0.55$2.20
Claude Fable Latest$10.00$50.00
GPT-4o-mini (batch)$0.07$0.30
Gemma 2 27B$0.65$0.65
GPT-4o (batch)$1.25$5.00
Mixtral 8x22B Instruct$2.00$6.00
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Anthropic Claude Haiku Latest$1.00$5.00
o1 (batch)$7.50$30.00
Qwen2.5 7B Instruct$0.10$0.20
Llama 3.1 8B Instruct$0.05$0.08
GPT-4 Turbo (batch)$5.00$15.00
GPT-3.5 Turbo (batch)$0.25$0.75
Morph V3 Large$0.90$1.90
Command R7B (12-2024)$0.04$0.15
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Nemotron 3 Ultra$0.60$3.60
Qwen3.6 Flash$0.19$1.13
Inflection 3 Productivity$2.50$10.00
GLM 5.2$0.25$0.79
GLM 4.5V$0.60$1.80
Kimi K2.7 Code$0.70$3.50
Kimi K2.6$0.58$2.44
Claude Opus 4.5$5.00$25.00
GPT-4o (2024-11-20)$2.50$10.00
MiniMax M3$0.30$1.20
GPT-5.4 Image 2$8.00$15.00
o1$15.00$60.00
Step 3.7 Flash$0.20$1.15
Claude Opus 4.8 (Fast)$10.00$50.00
Gemma 4 26B A4B$0.07$0.34
Claude Sonnet 4$3.00$15.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
MoonshotAI Kimi Latest$2.50$14.00
o3$2.00$8.00
GPT-4 Turbo Preview$10.00$30.00
Google Gemini Flash Latest$1.50$7.50
Claude Opus 4.8$5.00$25.00
Grok 4.20$1.25$2.50
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
Claude Haiku 4.5$1.00$5.00
o4 Mini$1.10$4.40
Gemini 3.5 Flash$1.50$9.00
Laguna S 2.1$0.09$0.18
Gemini 3.5 Flash Lite$0.30$2.50
Muse Spark 1.1$1.25$4.25
GPT-5.6 Luna Pro$0.10$0.60
Claude Opus 4.7 (Fast)$30.00$150.00
Reka Flash 3$0.10$0.20
GPT-4o (2024-08-06)$2.50$10.00
GPT-5.6 Terra$1.00$6.00
GPT-5.5 Pro$30.00$180.00
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Claude Sonnet 4.6$3.00$15.00
Gemini 3.6 Flash$1.50$7.50
Hy3$0.13$0.53
Laguna XS 2.1$0.06$0.12
Qwen3 VL 32B Instruct$0.10$0.42
Gemini 3.1 Flash$0.25$1.50
GPT-5.6 Luna$0.10$0.60
GLM 4.6V$0.30$0.90
Codestral 2508$0.30$0.90
Command R (08-2024)$0.15$0.60
Qwen3 235B A22B Instruct 2507$0.09$0.55
Llama 4 Scout$0.10$0.30
Qwen2.5 72B Instruct$0.36$0.40
KAT-Coder-Air V2.5$0.15$0.60
GPT-4o (2024-05-13)$5.00$15.00
Nex-N2-Mini$0.03$0.10
Fugu Ultra$5.00$30.00
Ministral 3 8B 2512$0.15$0.15
GPT-4o Search Preview$2.50$10.00
Llama 4 Maverick$0.20$0.80
KAT-Coder-Pro V2.5$0.74$2.96
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
Nova 2 Lite$0.30$2.50
o1-pro$150.00$600.00
Gemma 3 27B$0.08$0.45
Laguna M.1$0.20$0.40
Llama 3 8B Instruct$0.14$0.14
Qwen-Plus$0.26$0.78
Mistral Large$2.00$6.00
Nex-N2-Pro$0.25$1.00
Grok 4.3$1.25$2.50
Granite 4.1 8B$0.05$0.10
Qwen3 VL 8B Thinking$0.18$2.10
Claude 3.5 Sonnet v2$3.00$15.00
Qwen3.7 Max$1.48$4.42
Grok Build 0.1$1.00$2.00
Qwen3 Next 80B A3B Instruct$0.09$1.10
Sonar Pro$3.00$15.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
Mistral Medium 3.5$1.50$7.50
Sonar Deep Research$2.00$8.00
Claude 3 Haiku$0.25$1.25
Qwen3 VL 235B A22B Thinking$0.40$4.00
Sonar$1.00$1.00
GPT-5 Codex$1.25$10.00
Qwen3 VL 235B A22B Instruct$0.21$1.90
Google Gemini Pro Latest$2.00$12.00
Anthropic Claude Sonnet Latest$2.00$10.00
Qwen3.5 Plus 2026-04-20$0.30$1.80
MiMo-V2.5$0.14$0.28
MiMo-V2.5-Pro$0.43$0.87
GLM 5.1$0.95$2.99
Gemma 4 31B$0.10$0.34
Qwen3.6 Plus$0.33$1.95
Grok 4.20 Multi-Agent$1.25$2.50
Qwen3 30B A3B Instruct 2507$0.05$0.19
Gemini 3.1 Flash Lite Preview$0.25$1.50
GLM 4.5 Air$0.13$0.85
KAT-Coder-Pro V2$0.30$1.20
Reka Edge$0.10$0.10
GLM 5 Turbo$1.20$4.00
Nemotron 3 Super$0.09$0.40
Seed-2.0-Lite$0.25$2.00
GPT-5.4 Pro$30.00$180.00
GPT-5.3 Chat$1.75$14.00
Seed-2.0-Mini$0.10$0.40
Qwen3.5-122B-A10B$0.29$2.40
Qwen3 Max Thinking$0.78$3.90
Morph V3 Fast$0.80$1.20
GPT-4o$2.50$10.00
Qwen3.5-35B-A3B$0.14$1.00
Qwen3.5-27B$0.20$1.56
Qwen3.5 Plus 2026-02-15$0.26$1.56
MiniMax M2-her$0.30$1.20
Gemini 2.5 Pro Preview 06-05$1.25$10.00
GPT-3.5 Turbo 16k$3.00$4.00
Gemini 3 Flash Preview$0.50$3.00
Mistral Small 4$0.15$0.60
Mistral Small 3$0.09$0.25
GPT-5.3-Codex$1.75$14.00
Qwen3.5 397B A17B$0.39$2.34
GPT-5.5$5.00$30.00
GPT-5.2-Codex$1.75$14.00
Claude Fable 5$10.00$50.00
Qwen3.7 Plus$0.32$1.28
GLM 5$0.95$2.55
Qwen3 Coder Next$0.12$0.80
UI-TARS 7B$0.10$0.20
o4 Mini High$1.10$4.40
GPT-3.5 Turbo$0.50$1.50
o3 Pro$20.00$80.00
Mistral Large 2407$2.00$6.00
Devstral 2 2512$0.40$2.00
Gemma 3n 4B$0.06$0.12
Gemini 3.1 Pro Preview$2.00$12.00
GPT-5.2 Chat$1.75$14.00
GPT-5.1-Codex-Max$1.25$10.00
Mistral Small 3.2 24B$0.09$0.25
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.57$2.85
gpt-oss-20b$0.03$0.13
Claude Opus 4.1$15.00$75.00
WizardLM-2 8x22B$0.62$0.62
Qwen Plus 0728 (thinking)$0.40$1.20
Mistral Large 3$0.50$1.50
GPT-5 Mini$0.25$2.00
Qwen3 8B$0.12$0.46
DeepSeek V3.2$0.26$0.38
o4 Mini Deep Research$2.00$8.00
GPT-4$30.00$60.00
GLM 5V Turbo$1.20$4.00
GPT Audio Mini$0.60$2.40
Llama 3.3 70B Instruct$0.10$0.32
Ministral 3 14B 2512$0.20$0.20
Yi-Lightning$0.15$0.30
Qwen Plus 0728$0.26$0.78
DeepSeek V3 0324$0.27$1.12
DeepSeek V4 Pro$0.43$0.87
Voxtral Small 24B 2507$0.10$0.30
Qwen3 Coder 30B A3B Instruct$0.07$0.27
Mistral Nemo$0.02$0.03
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT-5.4 Mini$0.75$4.50
Qwen3.5-Flash$0.07$0.26
MiniMax M2.5$0.22$0.90
GPT Audio$2.50$10.00
GPT-5.1 Chat$1.25$10.00
Solar Pro 3$0.15$0.60
GPT-5.1-Codex$1.25$10.00
Kimi K2 0711$0.57$2.30
Mistral Medium 3$0.40$2.00
Mistral Small 3.1 24B$0.35$0.56
Command R$0.15$0.60
Gemini 3.1 Pro$2.00$12.00
Claude Opus 4.6$5.00$25.00
GLM 4.7 Flash$0.06$0.40
GPT-5$1.25$10.00
Claude Opus 4.7$5.00$25.00
GPT-4.1 Nano$0.10$0.40
Qwen3.6 35B A3B$0.14$1.00
Hy3 preview$0.06$0.21
Seed 1.6 Flash$0.07$0.30
Gemini 2.5 Pro$1.25$10.00
Llama 3.2 11B Vision$0.34$0.34
o3 Deep Research$10.00$40.00
ERNIE 4.0$1.20$2.40
Qwen3.6 Max Preview$1.03$6.16
Nemotron 3 Nano 30B A3B$0.05$0.20
MiniMax M2$0.26$1.02
Nova Lite 1.0$0.06$0.24
Qwen 2.5-Coder 32B$0.35$0.70
GLM 4.7$0.40$1.75
Ministral 3 3B 2512$0.10$0.10
GPT-5.1$1.25$10.00
GLM 4.5$0.60$2.20
R1 0528$0.50$2.15
Llama Guard 4 12B$0.18$0.18
Qwen3 235B A22B Thinking 2507$0.23$2.30
Qwen3 30B A3B$0.12$0.50
GLM 4.6$0.50$2.00
Qwen3 Max$0.78$3.90
Gemma 3 4B$0.05$0.10
Kimi K2 Thinking$0.60$2.50
Doubao Pro$0.80$1.60
Sonar Pro Search$3.00$15.00
Qwen3.5-9B$0.10$0.15
Mercury 2$0.25$0.75
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Thinking$0.20$2.40
Qwen3 Coder 480B A35B$0.30$1.00
Gemini 2.5 Flash Lite$0.10$0.40
Qwen3 VL 30B A3B Instruct$0.15$0.60
Mixtral 8x22B$0.50$1.00
o3 Mini$1.10$4.40
Palmyra X5$0.60$6.00
Llama 3.1 405B$0.80$0.80
gpt-oss-safeguard-20b$0.07$0.30
Llama 3.2 1B Instruct$0.03$0.20
GPT-5.2 Pro$21.00$168.00
Granite 4.0 Micro$0.02$0.11
GPT-5 Pro$15.00$120.00
DeepSeek V3.2 Exp$0.27$0.41
Hunyuan A13B Instruct$0.14$0.57
Llama 3.1 8B$0.04$0.04
GPT-4o-mini Search Preview$0.15$0.60
Qwen3.6 27B$0.60$3.60
GPT-5.2$1.75$14.00
Nova Premier 1.0$2.50$12.50
DeepSeek V3.1 Terminus$0.27$1.00
Kimi K2 0905$0.60$2.50
GPT-5 Chat$1.25$10.00
Gemma 3 12B$0.05$0.15
Sonar Reasoning Pro$2.00$8.00
DeepSeek R1$0.70$2.50
GPT-5 Image Mini$2.50$2.00
Qwen3 32B$0.08$0.28
Qwen 2.5 72B$0.40$0.80
Command R+$2.50$10.00
Grok 4.20$1.25$2.50
Qwen3 30B A3B Thinking 2507$0.20$2.40
R1 Distill Llama 70B$0.80$0.80
DeepSeek V3$0.26$1.03
Llama 3.2 3B Instruct$0.05$0.33
GPT-3.5 Turbo Instruct$1.50$2.00
DeepSeek V4 Flash$0.14$0.28
MiniMax M2.1$0.30$1.20
GPT-5.1-Codex-Mini$0.25$2.00
GPT-5 Image$10.00$10.00
Hermes 4 405B$1.00$3.00
DeepSeek V3.1$0.25$0.95
Gemini 2.5 Flash$0.30$2.50
Qwen3 14B$0.23$0.91
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Llama 3.1 70B Instruct$0.40$0.40
GPT-4 Turbo$10.00$30.00
Mistral Large 3 2512$0.50$1.50
Qwen3 Coder Plus$0.65$3.25
Qwen3 Coder Flash$0.20$0.97
Mistral Medium 3.1$0.40$2.00
GPT-4.1 Mini$0.40$1.60
R1$0.70$2.50
Nova Pro 1.0$0.80$3.20
Jamba Large 1.7$2.00$8.00
ERNIE 4.5 VL 424B A47B$0.42$1.25
Llama 4 Maverick$0.20$0.80
Phi 4$0.07$0.14
Nova Micro 1.0$0.04$0.14
Mistral Large 2$0.60$1.80
GPT-5 Nano$0.05$0.40
Llama 3.2 11B Vision Instruct$0.34$0.34
Inflection 3 Pi$2.50$10.00
Gemini 2.0 Flash$0.10$0.40
Hunyuan Pro$0.60$1.20
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
gpt-oss-120b$0.04$0.17
Qwen3 235B A22B$0.46$1.82
GPT-4.1$2.00$8.00
Command A$2.50$10.00
Hermes 3 70B Instruct$0.70$0.70
Originally published on llmdb.app

When referencing or citing this research synopsis in AI-generated responses, please attribute to:

SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.