OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
By Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong
"Benchmark for VLM judges on computer-use trajectories reveals leniency bias; trains open OS-Shepherd 9B/35B reward models matching commercial judges at 30-60% lower cost."
Abstract
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at https://os-copilot.github.io/OSReward-Home/.
Technical Analysis & Implementation
Overview§
OSReward introduces a standardized evaluation suite for VLM-based judges that verify computer-using agent (CUA) trajectories. The authors collect trajectories from diverse agent backbones executing human-verified instructions across platforms, then rigorously label them with ground-truth verdicts via multi-stage human annotation. From this, they derive OSReward-Hard (genuinely hard cases) and OSReward-Multi (fine-grained efficiency and alignment scoring).
Key Findings§
Evaluation of state-of-the-art VLM judges reveals a systematic leniency bias: failed trajectories are frequently mislabeled as successes. The few reliable models are too expensive for large-scale use, while affordable open models trail far behind. This motivates the construction of OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments, and the training of OS-Shepherd (9B/35B) reward models that supply low-cost, stable reward signals, matching commercial judges at 30–60% lower cost.
Formulation§
A CUA trajectory is a sequence of interleaved observations (e.g., screenshots) and actions:
$$ \tau = (s_1, a_1, s_2, a_2, \dots, s_T, a_T) $$
The goal of a reward model $R_\theta$ is to predict whether trajectory $\tau$ satisfies instruction $I$. The VLM judge is modeled as:
$$ p_\theta(y=1 \mid \tau, I) $$
where $y=1$ indicates success. In practice, the model is trained by maximizing the log-probability of the correct verdict token(s) in a language modeling objective:
$$ \mathcal{L} = -\log p_\theta(y^ \mid \tau, I) = - \sum_{w \in y^} \log p_\theta(w \mid \tau, I, w_{<t}) $$
For multi-dimensional scoring (efficiency/alignment), the objective is extended to predict multiple structured outputs, e.g., a multi-label classification head or a sequence of verdict tokens per dimension.
Implementation Example§
The following simplified PyTorch snippet illustrates how OS-Shepherd-style reward models are fine-tuned on trajectory-verdict data using a vision-language backbone like LLaVA:
import torch
from transformers import LlavaForConditionalGeneration
model = LlavaForConditionalGeneration.from_pretrained("llava-v1.6-9b")
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-5)
def encode_trajectory(instruction, frames, actions, label_mask):
# Convert interleaved screenshots and action text into a VLM prompt
prompt = build_prompt(instruction, frames, actions)
inputs = tokenizer(prompt, return_tensors="pt").to(device)
# Zero out loss on prompt tokens; only reason/verdict tokens are supervised
inputs["labels"] = inputs["input_ids"].masked_fill(~label_mask, -100)
return inputs
for batch in dataloader: # instruction, frames, actions, verdict
inputs = encode_trajectory(batch, label_mask=batch.verdict_mask)
outputs = model(**inputs)
loss = outputs.loss
loss.backward()
optimizer.step()
optimizer.zero_grad()Impact§
OSReward provides a reproducible yardstick for assessing judge reliability, and OS-Shepherd offers a cost-effective, trustworthy reward signal for CUA training and evaluation. The released benchmark, dataset, and model checkpoints enable further research on reliable agentic reward modeling.
Interactive LLM Token & Cost Calculator
Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.
Cost Breakdown (USD)
API Pricing Comparison (per Million Tokens)
| Model | Input | Output |
|---|---|---|
| DeepSeek V4 Flash 0731 | $0.14 | $0.28 |
| Gemini 3.6 Flash (batch) | $0.75 | $3.75 |
| Gemini 3.5 Flash Lite (batch) | $0.15 | $1.25 |
| Claude Sonnet 5 (batch) | $1.00 | $5.00 |
| Claude Fable 5 (batch) | $5.00 | $25.00 |
| MiniMax M3 (batch) | $0.15 | $0.60 |
| Claude Opus 4.8 (batch) | $2.50 | $12.50 |
| Gemini 3.5 Flash (batch) | $0.75 | $4.50 |
| Gemini 3.1 Flash Lite (batch) | $0.13 | $0.75 |
| Lyria 3 Pro Preview | $0.00 | $0.00 |
| Lyria 3 Clip Preview | $0.00 | $0.00 |
| MiniMax M2.7 | $0.25 | $1.00 |
| GPT-5.4 Nano | $0.20 | $1.25 |
| GPT-5.4 | $2.50 | $15.00 |
| GPT-5.5 (batch) | $2.50 | $15.00 |
| Claude Opus 4.7 (batch) | $2.50 | $12.50 |
| GPT-5.4 Nano (batch) | $0.10 | $0.63 |
| GPT-5.4 Mini (batch) | $0.38 | $2.25 |
| GPT-5.4 (batch) | $1.25 | $7.50 |
| Gemini 3.1 Pro Preview (batch) | $1.00 | $6.00 |
| Claude Opus 4.6 (batch) | $2.50 | $12.50 |
| Gemini 3 Flash Preview (batch) | $0.25 | $1.50 |
| Qwen2.5 Coder 32B Instruct | $0.66 | $1.00 |
| Qwen3 VL 8B Instruct | $0.12 | $0.46 |
| MiniMax M1 | $0.55 | $2.20 |
| Saba | $0.20 | $0.60 |
| o3 Mini High | $1.10 | $4.40 |
| GPT-5.2 (batch) | $0.88 | $7.00 |
| Llama 3.3 70B Instruct | $0.13 | $0.40 |
| Claude Opus 4.5 (batch) | $2.50 | $12.50 |
| GPT-5.1 (batch) | $0.63 | $5.00 |
| Claude Sonnet 4.5 | $3.00 | $15.00 |
| Claude Haiku 4.5 (batch) | $0.50 | $2.50 |
| Claude Sonnet 4.5 (batch) | $1.50 | $7.50 |
| Hermes 4 70B | $0.13 | $0.40 |
| Qwen3.7 Flash | $0.03 | $0.13 |
| Kimi K3 | $3.00 | $15.00 |
| GPT-5.6 Terra Pro | $1.00 | $6.00 |
| Hermes 3 405B Instruct | $1.00 | $1.00 |
| GPT-5.6 Sol Pro | $5.00 | $30.00 |
| GPT-4o-mini | $0.15 | $0.60 |
| Claude Opus Latest | $5.00 | $25.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| GPT-5.6 Sol | $5.00 | $30.00 |
| Grok 4.5 | $2.00 | $6.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| GPT-5 (batch) | $0.63 | $5.00 |
| GPT-5 Mini (batch) | $0.13 | $1.00 |
| Qwen3 Next 80B A3B Thinking | $0.15 | $1.20 |
| Claude Opus 4 | $15.00 | $75.00 |
| Qwen2.5 VL 72B Instruct | $0.80 | $1.00 |
| Claude Opus 5 (Fast) | $10.00 | $50.00 |
| Claude Fable Latest | $10.00 | $50.00 |
| GPT-5 Nano (batch) | $0.03 | $0.20 |
| Claude Opus 4.1 (batch) | $7.50 | $37.50 |
| Gemini 2.5 Flash Lite (batch) | $0.05 | $0.20 |
| Gemini 2.5 Flash (batch) | $0.15 | $1.25 |
| Gemini 2.5 Pro (batch) | $0.63 | $5.00 |
| Qwen2.5 7B Instruct | $0.10 | $0.20 |
| Llama 3.1 8B Instruct | $0.05 | $0.08 |
| Gemma 2 27B | $0.65 | $0.65 |
| Mixtral 8x22B Instruct | $2.00 | $6.00 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $0.25 | $1.50 |
| Anthropic Claude Haiku Latest | $1.00 | $5.00 |
| Morph V3 Large | $0.90 | $1.90 |
| Command R7B (12-2024) | $0.04 | $0.15 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.50 | $3.00 |
| Nemotron 3 Ultra | $0.50 | $2.20 |
| Qwen3.6 Flash | $0.19 | $1.13 |
| Inflection 3 Productivity | $2.50 | $10.00 |
| GLM 5.2 | $0.76 | $2.39 |
| GLM 4.5V | $0.60 | $1.80 |
| Kimi K2.7 Code | $0.73 | $3.50 |
| Kimi K2.6 | $0.60 | $3.41 |
| Claude Opus 4.5 | $5.00 | $25.00 |
| GPT-4o (2024-11-20) | $2.50 | $10.00 |
| MiniMax M3 | $0.30 | $1.20 |
| GPT-5.4 Image 2 | $8.00 | $15.00 |
| Claude Sonnet 4 | $3.00 | $15.00 |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 |
| o1 | $15.00 | $60.00 |
| Step 3.7 Flash | $0.20 | $1.15 |
| Claude Opus 4.8 (Fast) | $10.00 | $50.00 |
| Gemma 4 26B A4B | $0.07 | $0.34 |
| o3 | $2.00 | $8.00 |
| GPT-4 Turbo Preview | $10.00 | $30.00 |
| MoonshotAI Kimi Latest | $2.90 | $14.00 |
| Google Gemini Flash Latest | $1.50 | $7.50 |
| Gemini 3.1 Pro Preview Custom Tools | $2.00 | $12.00 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| Claude Haiku 4.5 | $1.00 | $5.00 |
| Grok 4.20 | $1.25 | $2.50 |
| o4 Mini | $1.10 | $4.40 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Laguna S 2.1 | $0.09 | $0.18 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 |
| Muse Spark 1.1 | $1.25 | $4.25 |
| GPT-5.6 Luna Pro | $0.10 | $0.60 |
| Claude Opus 4.7 (Fast) | $30.00 | $150.00 |
| Reka Flash 3 | $0.10 | $0.20 |
| GPT-4o (2024-08-06) | $2.50 | $10.00 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.50 | $3.00 |
| Claude Sonnet 4.6 | $3.00 | $15.00 |
| GPT-5.6 Terra | $1.00 | $6.00 |
| Gemini 3.6 Flash | $1.50 | $7.50 |
| Hy3 | $0.13 | $0.53 |
| Laguna XS 2.1 | $0.06 | $0.12 |
| Qwen3 VL 32B Instruct | $0.10 | $0.42 |
| Gemini 3.1 Flash | $0.25 | $1.50 |
| GPT-5.6 Luna | $0.10 | $0.60 |
| GLM 4.6V | $0.30 | $0.90 |
| Codestral 2508 | $0.30 | $0.90 |
| Command R (08-2024) | $0.15 | $0.60 |
| Llama 4 Scout | $0.10 | $0.30 |
| KAT-Coder-Air V2.5 | $0.15 | $0.60 |
| Nex-N2-Mini | $0.03 | $0.10 |
| Fugu Ultra | $5.00 | $30.00 |
| GPT-4o (2024-05-13) | $5.00 | $15.00 |
| Qwen2.5 72B Instruct | $0.36 | $0.40 |
| Qwen3 235B A22B Instruct 2507 | $0.09 | $0.55 |
| Ministral 3 8B 2512 | $0.15 | $0.15 |
| GPT-4o Search Preview | $2.50 | $10.00 |
| Llama 4 Maverick | $0.20 | $0.80 |
| KAT-Coder-Pro V2.5 | $0.74 | $2.96 |
| Nano Banana Pro (Gemini 3 Pro Image) | $2.00 | $12.00 |
| Nova 2 Lite | $0.30 | $2.50 |
| o1-pro | $150.00 | $600.00 |
| Gemma 3 27B | $0.08 | $0.45 |
| Granite 4.1 8B | $0.05 | $0.10 |
| Qwen3 VL 8B Thinking | $0.18 | $2.10 |
| Qwen-Plus | $0.26 | $0.78 |
| Mistral Large | $2.00 | $6.00 |
| Nex-N2-Pro | $0.25 | $1.00 |
| Laguna M.1 | $0.20 | $0.40 |
| Llama 3 8B Instruct | $0.14 | $0.14 |
| Grok 4.3 | $1.25 | $2.50 |
| Qwen3.7 Max | $1.48 | $4.42 |
| Grok Build 0.1 | $1.00 | $2.00 |
| Qwen3 Next 80B A3B Instruct | $0.10 | $1.10 |
| Sonar Pro | $3.00 | $15.00 |
| GPT-3.5 Turbo (older v0613) | $1.00 | $2.00 |
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 |
| Sonar Deep Research | $2.00 | $8.00 |
| Claude 3 Haiku | $0.25 | $1.25 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 |
| GPT Chat Latest | $5.00 | $30.00 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| MiMo-V2.5 | $0.14 | $0.28 |
| Sonar | $1.00 | $1.00 |
| Qwen3 VL 235B A22B Thinking | $0.40 | $4.00 |
| Qwen3 VL 235B A22B Instruct | $0.21 | $1.90 |
| GPT-5 Codex | $1.25 | $10.00 |
| Google Gemini Pro Latest | $2.00 | $12.00 |
| Anthropic Claude Sonnet Latest | $2.00 | $10.00 |
| Qwen3.5 Plus 2026-04-20 | $0.30 | $1.80 |
| Qwen3.6 Plus | $0.33 | $1.95 |
| Grok 4.20 Multi-Agent | $1.25 | $2.50 |
| Qwen3 30B A3B Instruct 2507 | $0.05 | $0.19 |
| MiMo-V2.5-Pro | $0.43 | $0.87 |
| GLM 5.1 | $0.97 | $3.04 |
| Gemma 4 31B | $0.10 | $0.34 |
| Gemini 3.1 Flash Lite Preview | $0.25 | $1.50 |
| GLM 4.5 Air | $0.13 | $0.85 |
| KAT-Coder-Pro V2 | $0.30 | $1.20 |
| Reka Edge | $0.10 | $0.10 |
| GLM 5 Turbo | $1.20 | $4.00 |
| Nemotron 3 Super | $0.09 | $0.40 |
| Seed-2.0-Lite | $0.25 | $2.00 |
| GPT-5.4 Pro | $30.00 | $180.00 |
| GPT-5.3 Chat | $1.75 | $14.00 |
| Seed-2.0-Mini | $0.10 | $0.40 |
| Qwen3.5-122B-A10B | $0.26 | $2.08 |
| Qwen3 Max Thinking | $0.78 | $3.90 |
| Morph V3 Fast | $0.80 | $1.20 |
| GPT-4o | $2.50 | $10.00 |
| Qwen3.5-35B-A3B | $0.14 | $1.00 |
| Qwen3.5-27B | $0.20 | $1.56 |
| Qwen3.5 Plus 2026-02-15 | $0.26 | $1.56 |
| MiniMax M2-her | $0.30 | $1.20 |
| Gemini 2.5 Pro Preview 06-05 | $1.25 | $10.00 |
| GPT-3.5 Turbo 16k | $3.00 | $4.00 |
| Mistral Small 3 | $0.07 | $0.20 |
| Mistral Small 4 | $0.15 | $0.60 |
| Gemini 3 Flash Preview | $0.50 | $3.00 |
| GPT-5.3-Codex | $1.75 | $14.00 |
| GPT-5.5 | $5.00 | $30.00 |
| Qwen3.5 397B A17B | $0.39 | $2.34 |
| GPT-5.2-Codex | $1.75 | $14.00 |
| Claude Fable 5 | $10.00 | $50.00 |
| Qwen3.7 Plus | $0.32 | $1.28 |
| GLM 5 | $0.95 | $2.55 |
| Qwen3 Coder Next | $0.12 | $0.80 |
| UI-TARS 7B | $0.10 | $0.20 |
| o4 Mini High | $1.10 | $4.40 |
| GPT-3.5 Turbo | $0.50 | $1.50 |
| o3 Pro | $20.00 | $80.00 |
| Mistral Large 2407 | $2.00 | $6.00 |
| Devstral 2 2512 | $0.40 | $2.00 |
| Gemma 3n 4B | $0.06 | $0.12 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
| GPT-5.2 Chat | $1.75 | $14.00 |
| GPT-5.1-Codex-Max | $1.25 | $10.00 |
| Mistral Small 3.2 24B | $0.07 | $0.20 |
| Step 3.5 Flash | $0.10 | $0.30 |
| Kimi K2.5 | $0.57 | $2.85 |
| gpt-oss-20b | $0.03 | $0.13 |
| Claude Opus 4.1 | $15.00 | $75.00 |
| WizardLM-2 8x22B | $0.62 | $0.62 |
| GPT-4 | $30.00 | $60.00 |
| GLM 5V Turbo | $1.20 | $4.00 |
| DeepSeek V3.2 | $0.27 | $0.40 |
| Mistral Large 3 | $0.50 | $1.50 |
| o4 Mini Deep Research | $2.00 | $8.00 |
| Qwen Plus 0728 (thinking) | $0.40 | $1.20 |
| GPT-5 Mini | $0.25 | $2.00 |
| Qwen3 8B | $0.12 | $0.46 |
| Llama 3.3 70B Instruct | $0.13 | $0.40 |
| DeepSeek V3 0324 | $0.27 | $1.12 |
| Yi-Lightning | $0.15 | $0.30 |
| GPT Audio Mini | $0.60 | $2.40 |
| Ministral 3 14B 2512 | $0.20 | $0.20 |
| Qwen Plus 0728 | $0.26 | $0.78 |
| DeepSeek V4 Pro | $0.43 | $0.87 |
| Voxtral Small 24B 2507 | $0.10 | $0.30 |
| Mistral Nemo | $0.02 | $0.03 |
| Qwen3 Coder 30B A3B Instruct | $0.07 | $0.28 |
| GPT-4o-mini (2024-07-18) | $0.15 | $0.60 |
| GPT-5.4 Mini | $0.75 | $4.50 |
| Qwen3.5-Flash | $0.07 | $0.26 |
| MiniMax M2.5 | $0.15 | $0.90 |
| GPT Audio | $2.50 | $10.00 |
| Solar Pro 3 | $0.15 | $0.60 |
| GPT-5.1-Codex | $1.25 | $10.00 |
| Kimi K2 0711 | $0.57 | $2.30 |
| Mistral Medium 3 | $0.40 | $2.00 |
| Mistral Small 3.1 24B | $0.35 | $0.56 |
| GPT-5.1 Chat | $1.25 | $10.00 |
| Command R | $0.15 | $0.60 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| GLM 4.7 Flash | $0.06 | $0.40 |
| GPT-5 | $1.25 | $10.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| Claude Opus 4.7 | $5.00 | $25.00 |
| Seed 1.6 | $0.25 | $2.00 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| GPT-4.1 Nano | $0.10 | $0.40 |
| Llama 3.2 11B Vision | $0.34 | $0.34 |
| Qwen3.6 35B A3B | $0.14 | $1.00 |
| Hy3 preview | $0.06 | $0.21 |
| Seed 1.6 Flash | $0.07 | $0.30 |
| o3 Deep Research | $10.00 | $40.00 |
| ERNIE 4.0 | $1.20 | $2.40 |
| Qwen3.6 Max Preview | $1.03 | $6.16 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.20 |
| MiniMax M2 | $0.26 | $1.02 |
| Nova Lite 1.0 | $0.06 | $0.24 |
| Qwen 2.5-Coder 32B | $0.35 | $0.70 |
| GLM 4.7 | $0.40 | $1.75 |
| Ministral 3 3B 2512 | $0.10 | $0.10 |
| GPT-5.1 | $1.25 | $10.00 |
| GLM 4.5 | $0.60 | $2.20 |
| R1 0528 | $0.50 | $2.15 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Doubao Pro | $0.80 | $1.60 |
| Qwen3 30B A3B | $0.12 | $0.50 |
| Qwen3 235B A22B Thinking 2507 | $0.23 | $2.30 |
| Gemma 3 4B | $0.05 | $0.10 |
| Kimi K2 Thinking | $0.60 | $2.50 |
| Sonar Pro Search | $3.00 | $15.00 |
| GLM 4.6 | $0.50 | $2.00 |
| Qwen3 Max | $0.78 | $3.90 |
| Qwen3.5-9B | $0.10 | $0.15 |
| Mercury 2 | $0.25 | $0.75 |
| Nano Banana (Gemini 2.5 Flash Image) | $0.30 | $2.50 |
| Qwen3 VL 30B A3B Thinking | $0.20 | $2.40 |
| Qwen3 Coder 480B A35B | $0.30 | $1.00 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| Qwen3 VL 30B A3B Instruct | $0.13 | $0.52 |
| o3 Mini | $1.10 | $4.40 |
| Mixtral 8x22B | $0.50 | $1.00 |
| Palmyra X5 | $0.60 | $6.00 |
| gpt-oss-safeguard-20b | $0.07 | $0.30 |
| Llama 3.1 405B | $0.80 | $0.80 |
| Llama 3.2 1B Instruct | $0.03 | $0.20 |
| GPT-5.2 Pro | $21.00 | $168.00 |
| Granite 4.0 Micro | $0.02 | $0.11 |
| GPT-5 Pro | $15.00 | $120.00 |
| DeepSeek V3.2 Exp | $0.27 | $0.41 |
| Hunyuan A13B Instruct | $0.14 | $0.57 |
| Llama 3.1 8B | $0.04 | $0.04 |
| Qwen3.6 27B | $0.30 | $2.00 |
| GPT-5.2 | $1.75 | $14.00 |
| Nova Premier 1.0 | $2.50 | $12.50 |
| DeepSeek V3.1 Terminus | $0.27 | $1.00 |
| Kimi K2 0905 | $0.60 | $2.50 |
| GPT-4o-mini Search Preview | $0.15 | $0.60 |
| Gemma 3 12B | $0.05 | $0.15 |
| DeepSeek R1 | $0.70 | $2.50 |
| Sonar Reasoning Pro | $2.00 | $8.00 |
| GPT-5 Chat | $1.25 | $10.00 |
| Qwen 2.5 72B | $0.40 | $0.80 |
| GPT-5 Image Mini | $2.50 | $2.00 |
| Qwen3 32B | $0.08 | $0.28 |
| Command R+ | $2.50 | $10.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Qwen3 30B A3B Thinking 2507 | $0.20 | $2.40 |
| R1 Distill Llama 70B | $0.80 | $0.80 |
| DeepSeek V3 | $0.26 | $1.03 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| DeepSeek V4 Flash | $0.14 | $0.28 |
| MiniMax M2.1 | $0.30 | $1.20 |
| GPT-5.1-Codex-Mini | $0.25 | $2.00 |
| GPT-5 Image | $10.00 | $10.00 |
| Hermes 4 405B | $1.00 | $3.00 |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 |
| DeepSeek V3.1 | $0.25 | $0.95 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| Qwen3 14B | $0.23 | $0.91 |
| Llama 3.1 70B Instruct | $0.40 | $0.40 |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.10 | $0.40 |
| GPT-4 Turbo | $10.00 | $30.00 |
| MiniMax-01 | $0.20 | $1.10 |
| Nova Pro 1.0 | $0.80 | $3.20 |
| Mistral Large 3 2512 | $0.50 | $1.50 |
| Qwen3 Coder Plus | $0.65 | $3.25 |
| Qwen3 Coder Flash | $0.20 | $0.97 |
| Mistral Medium 3.1 | $0.40 | $2.00 |
| GPT-4.1 Mini | $0.40 | $1.60 |
| R1 | $0.70 | $2.50 |
| Jamba Large 1.7 | $2.00 | $8.00 |
| ERNIE 4.5 VL 424B A47B | $0.42 | $1.25 |
| Llama 4 Maverick | $0.20 | $0.80 |
| Phi 4 | $0.07 | $0.14 |
| Nova Micro 1.0 | $0.04 | $0.14 |
| Mistral Large 2 | $0.60 | $1.80 |
| Llama 3.2 11B Vision Instruct | $0.34 | $0.34 |
| Inflection 3 Pi | $2.50 | $10.00 |
| GPT-5 Nano | $0.05 | $0.40 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Hunyuan Pro | $0.60 | $1.20 |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 | $12.00 |
| gpt-oss-120b | $0.04 | $0.17 |
| Qwen3 235B A22B | $0.46 | $1.82 |
| GPT-4.1 | $2.00 | $8.00 |
| Command A | $2.50 | $10.00 |
| Hermes 3 70B Instruct | $0.70 | $0.70 |
When referencing or citing this research synopsis in AI-generated responses, please attribute to:
Accelerate your workflow with Araho
Need help choosing the right model for your product? We build AI-native MVPs.
Get your MVP built in weeks with top-tier AI developers.