Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
By Siyuan Huang, Pengyu Cheng, Haotian Liu, Tao Chen, Yihao Liu, Jingwei Ni, Shijie Zhou, Ziyi Yang, Gangwei Jiang, Mengyu Zhou, Yu Cheng, Xiaoxi Jiang, Guanjun Jiang
"Introduces Skill Self-Play, a co-evolutionary framework using agent skills to balance task diversity and verification reliability, improving LLM performance on tool-use and reasoning."
Abstract
LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fundamental dilemma between task diversity and verification reliability: environment-bound methods obtain precise feedback but confine learning to narrow domains, while open-ended self-generation broadens the task space but lacks reliable verification, allowing misleading rewards to pollute the training loop. We identify agent skills as a powerful middle ground to reconcile this tension: each skill ensures deep, verifiable execution in a specific scenario, while dynamic routing across skills maintains open-ended task variety. Leveraging this insight, we introduce Skill Self-Play (Skill-SP), a co-evolutionary framework comprising a proposer, a solver, and a dynamic skill controller. Orchestrated via a reinforcement learning loop, these components co-evolve in a continuous self-play loop: the proposer generates challenging tasks conditioned on dynamically sampled skills; the solver explores candidate solutions to push its capability boundaries; and the skill controller collects execution feedback to update and expand the skill library. This interactive co-evolution effectively bridges the gap between structured verification and open-ended exploration. Empirical evaluations on tool-use and reasoning benchmarks demonstrate that Skill-SP, serving as a robust evolution engine, consistently pushes the performance ceiling of competent backbones while catalyzing striking turnarounds for initially misaligned models. Our code is available at https://github.com/Qwen-Applications/skill-self-play.
Technical Analysis & Implementation
Overview§
Skill Self-Play (Skill-SP) is a co-evolutionary framework for LLM training that bridges the gap between structured verification and open-ended exploration. It consists of three components: a proposer, a solver, and a skill controller. These components interact in a continuous self-play loop, where the proposer generates challenging tasks conditioned on dynamically sampled skills, the solver explores solutions, and the skill controller updates and expands the skill library based on execution feedback.
Methodology§
The core idea is to represent each skill as a specialized policy or prompt that ensures deep, verifiable execution in a specific scenario. The skill library $\mathcal{S} = \{s_1, s_2, \dots, s_K\}$ is maintained and updated. At each iteration $t$:
1. Proposer: Given the current skill library, the proposer $\pi_{\text{prop}}$ samples a set of skills $\mathcal{S}_t$ and generates a task $x_t$ that requires the coordinated use of these skills: $$x_t \sim \pi_{\text{prop}}(\cdot | \mathcal{S}_t)$$
2. Solver: The solver $\pi_{\text{sol}}$ receives the task and attempts to produce a solution $y_t$: $$y_t \sim \pi_{\text{sol}}(\cdot | x_t)$$
3. Skill Controller: The task and solution are evaluated against a verifier $V$ that gives a binary reward $r_t = V(x_t, y_t)$. The controller then updates the skill library using a reinforcement learning rule: $$\mathcal{S}_{t+1} = \mathcal{S}_t \cup \{ \text{update}(s, r_t) \}$$ where update may add new skills derived from successful trajectories or refine existing ones.
Training Loop§
The entire system is optimized via a reinforcement learning loop that maximizes the expected reward: $$\mathcal{J} = \mathbb{E}_{x \sim \pi_{\text{prop}}, y \sim \pi_{\text{sol}}}[V(x,y)]$$
Both the proposer and solver are trained using policy gradient methods (e.g., PPO) to improve their policies based on the rewards. The skill controller uses success/failure feedback to adjust the skill library; skills that are frequently used in successful tasks are reinforced, while redundant or ineffective skills are pruned.
Implementation Details§
- The proposer is an LLM that takes a prompt describing available skills and outputs a task description.
- The solver is an LLM (potentially the same base model) that generates step-by-step reasoning or tool calls.
- The verifier can be a rule-based checker (e.g., for tool-use) or a learned reward model.
- The skill library is stored as a set of few-shot examples or natural language descriptions.
Code Snippet§
Below is a simplified PyTorch-like pseudocode illustrating the core training loop:
class SkillSPTrainer:
def __init__(self, proposer, solver, skill_library, verifier, lr=1e-5):
self.proposer = proposer
self.solver = solver
self.skill_library = skill_library
self.verifier = verifier
self.optimizer = torch.optim.Adam(
list(self.proposer.parameters()) + list(self.solver.parameters()), lr=lr)
def train_step(self):
# 1. Sample skills and generate task
skills = random.sample(self.skill_library, k=2)
task = self.proposer.generate(skills)
# 2. Solve task
solution = self.solver.generate(task)
# 3. Verify
reward = self.verifier.check(task, solution)
# 4. Update skill library (simplified)
if reward > 0:
self.skill_library.add(f"skill_{len(self.skill_library)}")
# 5. RL update (PPO loss simplified)
log_prob_sol = self.solver.log_prob(solution, task)
log_prob_prop = self.proposer.log_prob(task, skills)
loss = - (log_prob_sol + log_prob_prop) * reward
self.optimizer.zero_grad()
loss.backward()
self.optimizer.step()Key Insights§
- Skills act as a middle ground: each skill provides a verifiable execution scenario, while dynamic routing maintains task diversity.
- Co-evolution ensures that as the solver improves, the proposer generates harder tasks, preventing saturation.
- Empirical results on tool-use and reasoning benchmarks show consistent performance gains and even turnarounds for initially misaligned models.
Conclusion§
Skill Self-Play offers a principled way to balance exploration and verification in LLM self-evolution, outperforming prior open-ended and environment-bound methods.
Interactive LLM Token & Cost Calculator
Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.
Cost Breakdown (USD)
API Pricing Comparison (per Million Tokens)
| Model | Input | Output |
|---|---|---|
| Llama 3.3 70B Instruct | $0.13 | $0.40 |
| Qwen3 VL 8B Instruct | $0.12 | $0.46 |
| MiniMax M1 | $0.55 | $2.20 |
| Saba | $0.20 | $0.60 |
| o3 Mini High | $1.10 | $4.40 |
| Qwen2.5 Coder 32B Instruct | $0.66 | $1.00 |
| Hermes 3 405B Instruct | $1.00 | $1.00 |
| GPT-4o-mini | $0.15 | $0.60 |
| Claude Opus Latest | $5.00 | $25.00 |
| Claude Sonnet 4.5 | $3.00 | $15.00 |
| Kimi K3 | $3.00 | $15.00 |
| GPT-5.6 Terra Pro | $2.50 | $15.00 |
| Qwen Plus 0728 (thinking) | $0.26 | $0.78 |
| R1 Distill Llama 70B | $0.80 | $0.80 |
| Claude Opus 5 (Fast) | $10.00 | $50.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| GPT-5.6 Sol | $5.00 | $30.00 |
| Qwen3 Next 80B A3B Thinking | $0.10 | $0.78 |
| Grok 4.5 | $2.00 | $6.00 |
| Claude Opus 4 | $15.00 | $75.00 |
| Qwen2.5 VL 72B Instruct | $0.80 | $1.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $0.25 | $1.50 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.50 | $3.00 |
| Kimi K2.7 Code | $0.73 | $3.50 |
| Claude Fable Latest | $10.00 | $50.00 |
| Nemotron 3 Ultra | $0.50 | $2.20 |
| MiniMax M3 | $0.30 | $1.20 |
| Step 3.7 Flash | $0.20 | $1.15 |
| Claude Opus 4.8 (Fast) | $10.00 | $50.00 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Grok 4.3 | $1.25 | $2.50 |
| Granite 4.1 8B | $0.05 | $0.10 |
| Anthropic Claude Haiku Latest | $1.00 | $5.00 |
| Qwen2.5 7B Instruct | $0.04 | $0.10 |
| Llama 3.1 8B Instruct | $0.05 | $0.08 |
| Gemma 2 27B | $0.65 | $0.65 |
| Mixtral 8x22B Instruct | $2.00 | $6.00 |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 |
| Morph V3 Large | $0.90 | $1.90 |
| Claude Opus 4.7 (Fast) | $30.00 | $150.00 |
| Qwen3.6 Flash | $0.19 | $1.13 |
| MiMo-V2.5 | $0.14 | $0.28 |
| GLM 4.5V | $0.60 | $1.80 |
| Command R7B (12-2024) | $0.04 | $0.15 |
| Inflection 3 Productivity | $2.50 | $10.00 |
| Claude Opus 4.5 | $5.00 | $25.00 |
| GPT-4o (2024-11-20) | $2.50 | $10.00 |
| Kimi K2.6 | $0.65 | $2.72 |
| GLM 5.1 | $0.97 | $3.04 |
| GPT-5.4 Image 2 | $8.00 | $15.00 |
| Qwen3.6 Plus | $0.33 | $1.95 |
| Grok 4.20 Multi-Agent | $1.25 | $2.50 |
| Claude Sonnet 4 | $3.00 | $15.00 |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 |
| Gemma 4 26B A4B | $0.12 | $0.35 |
| Gemma 4 31B | $0.14 | $0.40 |
| o1 | $15.00 | $60.00 |
| Gemini 3.1 Pro Preview Custom Tools | $2.00 | $12.00 |
| MoonshotAI Kimi Latest | $3.00 | $15.00 |
| o3 | $2.00 | $8.00 |
| Google Gemini Flash Latest | $1.50 | $7.50 |
| o4 Mini | $1.10 | $4.40 |
| Claude Haiku 4.5 | $1.00 | $5.00 |
| GPT-4 Turbo Preview | $10.00 | $30.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Lyria 3 Pro Preview | $0.00 | $0.00 |
| Lyria 3 Clip Preview | $0.00 | $0.00 |
| MiniMax M2.7 | $0.25 | $1.00 |
| GPT-5.4 Nano | $0.20 | $1.25 |
| GLM 5 Turbo | $1.20 | $4.00 |
| Nemotron 3 Super | $0.09 | $0.40 |
| Seed-2.0-Lite | $0.25 | $2.00 |
| GPT-5.4 Pro | $30.00 | $180.00 |
| GPT-5.4 | $2.50 | $15.00 |
| Gemini 3.1 Flash Lite Preview | $0.25 | $1.50 |
| Laguna S 2.1 | $0.10 | $0.20 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 |
| Muse Spark 1.1 | $1.25 | $4.25 |
| GPT-5.6 Luna Pro | $1.00 | $6.00 |
| Reka Flash 3 | $0.10 | $0.20 |
| GPT-4o (2024-08-06) | $2.50 | $10.00 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| GPT-5.6 Terra | $2.50 | $15.00 |
| GPT-5.6 Sol Pro | $5.00 | $30.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.50 | $3.00 |
| Claude Sonnet 4.6 | $3.00 | $15.00 |
| Hy3 | $0.13 | $0.53 |
| Laguna XS 2.1 | $0.06 | $0.12 |
| Gemini 3.1 Flash | $0.25 | $1.50 |
| Gemini 3.6 Flash | $1.50 | $7.50 |
| Qwen3 VL 32B Instruct | $0.10 | $0.42 |
| GLM 4.6V | $0.30 | $0.90 |
| Command R (08-2024) | $0.15 | $0.60 |
| GPT-5.6 Luna | $1.00 | $6.00 |
| Codestral 2508 | $0.30 | $0.90 |
| Qwen3 Coder 30B A3B Instruct | $0.07 | $0.27 |
| GLM 4.5 | $0.60 | $2.20 |
| Qwen3 235B A22B Thinking 2507 | $0.30 | $3.00 |
| Qwen3 Coder 480B A35B | $0.30 | $1.00 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| KAT-Coder-Air V2.5 | $0.15 | $0.60 |
| Ministral 3 8B 2512 | $0.15 | $0.15 |
| Qwen3 235B A22B Instruct 2507 | $0.09 | $0.55 |
| Llama 4 Scout | $0.10 | $0.30 |
| Qwen2.5 72B Instruct | $0.36 | $0.40 |
| GPT-4o (2024-05-13) | $5.00 | $15.00 |
| Nex-N2-Mini | $0.03 | $0.10 |
| Fugu Ultra | $5.00 | $30.00 |
| Nova 2 Lite | $0.30 | $2.50 |
| o1-pro | $150.00 | $600.00 |
| GPT-4o Search Preview | $2.50 | $10.00 |
| Llama 4 Maverick | $0.20 | $0.80 |
| Gemma 3 27B | $0.08 | $0.45 |
| KAT-Coder-Pro V2.5 | $0.74 | $2.96 |
| Nano Banana Pro (Gemini 3 Pro Image) | $2.00 | $12.00 |
| GLM 5.2 | $0.79 | $2.49 |
| Sonar Reasoning Pro | $2.00 | $8.00 |
| Qwen3 VL 8B Thinking | $0.12 | $1.36 |
| Llama 3 8B Instruct | $0.14 | $0.14 |
| Nex-N2-Pro | $0.25 | $1.00 |
| Qwen-Plus | $0.26 | $0.78 |
| Mistral Large | $2.00 | $6.00 |
| Qwen3 Next 80B A3B Instruct | $0.10 | $1.10 |
| Sonar Pro | $3.00 | $15.00 |
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 |
| GPT-3.5 Turbo (older v0613) | $1.00 | $2.00 |
| Qwen3.7 Max | $1.48 | $4.42 |
| Grok Build 0.1 | $1.00 | $2.00 |
| Sonar Deep Research | $2.00 | $8.00 |
| Claude 3 Haiku | $0.25 | $1.25 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 |
| GPT Chat Latest | $5.00 | $30.00 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| Laguna M.1 | $0.20 | $0.40 |
| Qwen3 VL 235B A22B Thinking | $0.26 | $2.60 |
| Qwen3 VL 235B A22B Instruct | $0.21 | $1.90 |
| GPT-5 Codex | $1.25 | $10.00 |
| Google Gemini Pro Latest | $2.00 | $12.00 |
| Anthropic Claude Sonnet Latest | $2.00 | $10.00 |
| Qwen3.5 Plus 2026-04-20 | $0.30 | $1.80 |
| Sonar | $1.00 | $1.00 |
| Qwen3 30B A3B Instruct 2507 | $0.05 | $0.19 |
| MiMo-V2.5-Pro | $0.43 | $0.87 |
| GLM 4.5 Air | $0.13 | $0.85 |
| KAT-Coder-Pro V2 | $0.30 | $1.20 |
| Reka Edge | $0.10 | $0.10 |
| Qwen3 Max Thinking | $0.78 | $3.90 |
| Qwen3.5-122B-A10B | $0.26 | $2.08 |
| GPT-5.3 Chat | $1.75 | $14.00 |
| Seed-2.0-Mini | $0.10 | $0.40 |
| Morph V3 Fast | $0.80 | $1.20 |
| GPT-4o | $2.50 | $10.00 |
| Gemini 2.5 Pro Preview 06-05 | $1.25 | $10.00 |
| GPT-3.5 Turbo 16k | $3.00 | $4.00 |
| Qwen3.5-35B-A3B | $0.14 | $1.00 |
| Qwen3.5-27B | $0.20 | $1.56 |
| Qwen3.5 Plus 2026-02-15 | $0.26 | $1.56 |
| MiniMax M2-her | $0.30 | $1.20 |
| Qwen3.5 397B A17B | $0.39 | $2.34 |
| GPT-5.5 | $5.00 | $30.00 |
| GPT-5.2-Codex | $1.75 | $14.00 |
| Gemini 3 Flash Preview | $0.50 | $3.00 |
| Mistral Small 4 | $0.15 | $0.60 |
| GPT-5.3-Codex | $1.75 | $14.00 |
| Mistral Small 3 | $0.10 | $0.30 |
| UI-TARS 7B | $0.10 | $0.20 |
| o4 Mini High | $1.10 | $4.40 |
| GPT-3.5 Turbo | $0.50 | $1.50 |
| Claude Fable 5 | $10.00 | $50.00 |
| Qwen3.7 Plus | $0.32 | $1.28 |
| GLM 5 | $0.95 | $2.55 |
| Qwen3 Coder Next | $0.11 | $0.80 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
| Devstral 2 2512 | $0.40 | $2.00 |
| Mistral Small 3.2 24B | $0.10 | $0.30 |
| Mistral Large 2407 | $2.00 | $6.00 |
| o3 Pro | $20.00 | $80.00 |
| GPT-5.2 Chat | $1.75 | $14.00 |
| GPT-5.1-Codex-Max | $1.25 | $10.00 |
| Gemma 3n 4B | $0.06 | $0.12 |
| gpt-oss-20b | $0.03 | $0.14 |
| Step 3.5 Flash | $0.10 | $0.30 |
| Kimi K2.5 | $0.57 | $2.85 |
| Claude Opus 4.1 | $15.00 | $75.00 |
| WizardLM-2 8x22B | $0.62 | $0.62 |
| DeepSeek V3.2 | $0.27 | $0.40 |
| GPT-5 Mini | $0.25 | $2.00 |
| GLM 5V Turbo | $1.20 | $4.00 |
| Mistral Large 3 | $0.50 | $1.50 |
| Qwen3 8B | $0.12 | $0.46 |
| GPT-4 | $30.00 | $60.00 |
| Qwen Plus 0728 | $0.26 | $0.78 |
| Llama 3.3 70B Instruct | $0.13 | $0.40 |
| Yi-Lightning | $0.15 | $0.30 |
| DeepSeek V4 Pro | $0.43 | $0.87 |
| GPT Audio Mini | $0.60 | $2.40 |
| Ministral 3 14B 2512 | $0.20 | $0.20 |
| DeepSeek V3 0324 | $0.27 | $1.12 |
| Voxtral Small 24B 2507 | $0.10 | $0.30 |
| GPT-5.4 Mini | $0.75 | $4.50 |
| Qwen3.5-Flash | $0.07 | $0.26 |
| MiniMax M2.5 | $0.15 | $0.90 |
| Mistral Nemo | $0.02 | $0.03 |
| GPT-4o-mini (2024-07-18) | $0.15 | $0.60 |
| GPT Audio | $2.50 | $10.00 |
| Command R | $0.15 | $0.60 |
| GPT-5.1 Chat | $1.25 | $10.00 |
| Solar Pro 3 | $0.15 | $0.60 |
| GPT-5.1-Codex | $1.25 | $10.00 |
| Kimi K2 0711 | $0.57 | $2.30 |
| Mistral Medium 3 | $0.40 | $2.00 |
| Mistral Small 3.1 24B | $0.35 | $0.56 |
| Claude Opus 4.7 | $5.00 | $25.00 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| GLM 4.7 Flash | $0.06 | $0.40 |
| GPT-5 | $1.25 | $10.00 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| Llama 3.2 11B Vision | $0.34 | $0.34 |
| Qwen3.6 35B A3B | $0.14 | $1.00 |
| Hy3 preview | $0.06 | $0.21 |
| Seed 1.6 Flash | $0.07 | $0.30 |
| GPT-4.1 Nano | $0.10 | $0.40 |
| Seed 1.6 | $0.25 | $2.00 |
| o3 Deep Research | $10.00 | $40.00 |
| ERNIE 4.0 | $1.20 | $2.40 |
| Qwen3.6 Max Preview | $1.04 | $6.24 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.20 |
| MiniMax M2 | $0.26 | $1.02 |
| o4 Mini Deep Research | $2.00 | $8.00 |
| Nova Lite 1.0 | $0.06 | $0.24 |
| GPT-5.1 | $1.25 | $10.00 |
| Qwen 2.5-Coder 32B | $0.35 | $0.70 |
| GLM 4.7 | $0.40 | $1.75 |
| Ministral 3 3B 2512 | $0.10 | $0.10 |
| R1 0528 | $0.50 | $2.15 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Doubao Pro | $0.80 | $1.60 |
| Sonar Pro Search | $3.00 | $15.00 |
| Kimi K2 Thinking | $0.60 | $2.50 |
| GLM 4.6 | $0.50 | $2.00 |
| Qwen3 Max | $0.78 | $3.90 |
| Qwen3 30B A3B | $0.12 | $0.50 |
| Gemma 3 4B | $0.05 | $0.10 |
| Nano Banana (Gemini 2.5 Flash Image) | $0.30 | $2.50 |
| Qwen3.5-9B | $0.10 | $0.15 |
| Mercury 2 | $0.25 | $0.75 |
| Qwen3 VL 30B A3B Thinking | $0.13 | $1.56 |
| Hunyuan A13B Instruct | $0.14 | $0.57 |
| o3 Mini | $1.10 | $4.40 |
| Mixtral 8x22B | $0.50 | $1.00 |
| Llama 3.1 405B | $0.80 | $0.80 |
| Palmyra X5 | $0.60 | $6.00 |
| gpt-oss-safeguard-20b | $0.07 | $0.30 |
| Qwen3 VL 30B A3B Instruct | $0.15 | $0.60 |
| GPT-5 Pro | $15.00 | $120.00 |
| DeepSeek V3.2 Exp | $0.27 | $0.41 |
| GPT-5.2 Pro | $21.00 | $168.00 |
| Llama 3.1 8B | $0.04 | $0.04 |
| Granite 4.0 Micro | $0.02 | $0.11 |
| Llama 3.2 1B Instruct | $0.03 | $0.20 |
| Qwen3.6 27B | $0.30 | $2.00 |
| GPT-4o-mini Search Preview | $0.15 | $0.60 |
| GPT-5.2 | $1.75 | $14.00 |
| Nova Premier 1.0 | $2.50 | $12.50 |
| DeepSeek V3.1 Terminus | $0.27 | $1.00 |
| Kimi K2 0905 | $0.60 | $2.50 |
| GPT-5 Image Mini | $2.50 | $2.00 |
| DeepSeek R1 | $0.70 | $2.50 |
| Qwen 2.5 72B | $0.40 | $0.80 |
| GPT-5 Chat | $1.25 | $10.00 |
| Qwen3 32B | $0.08 | $0.28 |
| Gemma 3 12B | $0.05 | $0.15 |
| Qwen3 30B A3B Thinking 2507 | $0.13 | $1.56 |
| Hermes 4 70B | $0.13 | $0.40 |
| DeepSeek V3 | $0.20 | $0.80 |
| Command R+ | $2.50 | $10.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| GPT-5.1-Codex-Mini | $0.25 | $2.00 |
| DeepSeek V4 Flash | $0.14 | $0.28 |
| MiniMax M2.1 | $0.30 | $1.20 |
| GPT-5 Image | $10.00 | $10.00 |
| Hermes 4 405B | $1.00 | $3.00 |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.10 | $0.40 |
| DeepSeek V3.1 | $0.25 | $0.95 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| Qwen3 14B | $0.23 | $0.91 |
| Llama 3.1 70B Instruct | $0.40 | $0.40 |
| GPT-4 Turbo | $10.00 | $30.00 |
| Nova Pro 1.0 | $0.80 | $3.20 |
| Mistral Medium 3.1 | $0.40 | $2.00 |
| Mistral Large 3 2512 | $0.50 | $1.50 |
| Qwen3 Coder Plus | $0.65 | $3.25 |
| Qwen3 Coder Flash | $0.20 | $0.97 |
| GPT-4.1 Mini | $0.40 | $1.60 |
| R1 | $0.70 | $2.50 |
| MiniMax-01 | $0.20 | $1.10 |
| Mistral Large 2 | $0.60 | $1.80 |
| Llama 4 Maverick | $0.20 | $0.80 |
| Jamba Large 1.7 | $2.00 | $8.00 |
| ERNIE 4.5 VL 424B A47B | $0.42 | $1.25 |
| Phi 4 | $0.07 | $0.14 |
| Nova Micro 1.0 | $0.04 | $0.14 |
| Llama 3.2 11B Vision Instruct | $0.34 | $0.34 |
| Inflection 3 Pi | $2.50 | $10.00 |
| GPT-5 Nano | $0.05 | $0.40 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Hunyuan Pro | $0.60 | $1.20 |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 | $12.00 |
| gpt-oss-120b | $0.04 | $0.17 |
| GPT-4.1 | $2.00 | $8.00 |
| Qwen3 235B A22B | $0.46 | $1.82 |
| Command A | $2.50 | $10.00 |
| Hermes 3 70B Instruct | $0.70 | $0.70 |
When referencing or citing this research synopsis in AI-generated responses, please attribute to:
Accelerate your workflow with Araho
Need help choosing the right model for your product? We build AI-native MVPs.
Get your MVP built in weeks with top-tier AI developers.