KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
By Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer
"KaliBench: 8,504 NL-to-CLI pairs across 1,642 Kali tools enables runtime-free verifiable rewards; no open-weight model exceeds 42% exact-command accuracy, but RLVR lifts an 8B model to 685B MoE parity."
Abstract
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
Technical Analysis & Implementation
KaliBench: Fine-Grained Cybersecurity Tool-Use Benchmark§
Motivation & Problem Formulation§
LLMs are increasingly tasked with translating analyst intent into concrete command-line invocations. Existing cybersecurity benchmarks evaluate either static knowledge or end-to-end agentic behavior, but neither directly measures the model's ability to emit executable commands for real security tools. The failure mode is brittle: a single misordered flag, wrong value binding, or syntax slip renders the invocation invalid. KaliBench frames this as a structured natural-language → CLI translation problem over the Kali Linux tool ecosystem.
Formally, given a query $q$ in natural language and a tool inventory $\mathcal{T}$, the model must produce a command $c$ such that:
$$\hat{c} = \arg\max_{c} p_\theta(c \mid q), \quad \text{subject to } \text{exec}(c) \neq \emptyset,\ \text{canon}(c) = \text{canon}(c^*)$$
where $c^*$ is the gold command, $\text{exec}$ is sandboxed execution, and $\text{canon}$ is deterministic canonicalization (whitespace, flag ordering, alias resolution).
Dataset Construction & Verification Pipeline§
KaliBench contains 8,504 query–command pairs spanning 1,642 tools, organized along 23 capability dimensions and 5 security phases (recon, scanning, exploitation, post-exploitation, reporting). Three mechanisms ensure quality:
- Manuscript-grounded pipeline — commands are extracted from authoritative tool documentation, then deterministically canonicalized to remove surface-form variance (aliases like
-hvs--help, redundant whitespace, equivalent orderings). - Alias-aware evaluation — scoring credits semantically equivalent invocations, preventing spurious penalties for valid alias choices.
- Multi-stage verification — LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement are chained so every gold pair is both semantically correct and practically runnable.
Runtime-Free Verifiable Rewards§
A key contribution is that the deterministic canonicalization + alias-aware matching yield a reward signal without execution at training time. For a candidate command $c$ predicted from query $q$:
$$R(c, c^) = \alpha \cdot \mathbb{1}[\text{tool}(c) = \text{tool}(c^)] + \beta \cdot \frac{|F(c) \cap F(c^)|}{|F(c) \cup F(c^)|} + \gamma \cdot \mathbb{1}[\text{canon}(c) = \text{canon}(c^*)]$$
Here $F(\cdot)$ extracts the set of flag–value bindings, and the Jaccard term provides a graded partial-credit signal for argument construction when the exact command is not matched. This makes the reward cheap (no sandbox spin-up) yet verifiable and stable for RLVR.
Results & Training§
The authors evaluate 24 configurations of general-purpose and security-focused open-weight models across three evaluation modes (with/without tool hints, restricted flag set). In the unrestricted, no-hint setting, no open-weight model exceeds 42% exact-command accuracy. Supervised fine-tuning and reinforcement learning with the verifiable rewards derived from KaliBench significantly improve an 8B model, matching a 685B MoE baseline — evidence that dense, canonical reward signals substitute effectively for raw scale on this tool-use task.
import torch, torch.nn.functional as F
def verifiable_reward(pred, gold, w=(0.3, 0.3, 0.4)):
# pred/gold: dicts {'tool': str, 'flags': set[(flag,value)]}
tool_match = float(pred['tool'] == gold['tool'])
fp, fg = pred['flags'], gold['flags']
jac = len(fp & fg) / max(len(fp | fg), 1)
exact = float(canon(pred) == canon(gold))
a, b, g = w
return a * tool_match + b * jac + g * exact
def grpo_step(policy, tokenizer, prompts, golds, opt, beta=0.04):
# sample G completions per prompt
G = 8
with torch.no_grad():
eps = torch.randn_like(policy.logits) # simplified policy noise
rewards = torch.tensor([
verifiable_reward(parse(tok.decode(s)), g)
for p, g in zip(prompts, golds) for s in sample_G(policy, p, G)
])
# group-normalized advantages (GRPO)
adv = (rewards - rewards.mean()) / (rewards.std() + 1e-6)
loss = -adv * policy.logprob_sum + beta * kl_to_ref(policy)
loss.backward(); opt.step(); opt.zero_grad()
@torch.no_grad()
def sample_G(policy, prompt, G):
return policy.generate(prompt, num_return_sequences=G, do_sample=True, temperature=0.8)Takeaways§
KaliBench reframes cybersecurity tool use as a deterministic, canonicalizable structured-prediction problem. The conversion of brittle CLI syntax into exact-match-plus-Jaccard rewards unlocks RLVR at low cost, and the empirical result — 8B ≈ 685B — argues that reward density, not parameter count, drives command-level cybersecurity competence. The benchmark is released as both an evaluation suite and a training substrate.
Interactive LLM Token & Cost Calculator
Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.
Cost Breakdown (USD)
API Pricing Comparison (per Million Tokens)
| Model | Input | Output |
|---|---|---|
| Ling 3.1 Flash | $0.00 | $0.00 |
| GPT-6.1 Sol Pro | $2.00 | $10.00 |
| GPT-6.1 Sol | $2.00 | $10.00 |
| Claude Sonnet 5.5 | $2.00 | $10.00 |
| GLM 5.3 Prime | $2.80 | $8.80 |
| Qwen3.8 Max Prime | $4.00 | $12.00 |
| Solar Mini 4 | $0.05 | $0.20 |
| Claude Opus 5.5 | $4.00 | $20.00 |
| GPT-6 Sol | $2.00 | $10.00 |
| GPT-6 Luna | $0.10 | $0.50 |
| GPT-6 Luna Pro | $0.10 | $0.50 |
| GPT-6 Sol Pro | $2.00 | $10.00 |
| Command A+ | $2.50 | $10.00 |
| MiMo-V2.6-Pro-UltraSpeed | $4.35 | $8.70 |
| MiMo-V2.6-Flash | $0.14 | $0.28 |
| Qwen3.8 Omni Flash | $0.15 | $0.47 |
| MiMo-V2.6-Pro | $0.43 | $0.87 |
| Grok 4.7 | $2.00 | $6.00 |
| GLM 5.3 FlashX | $0.37 | $1.25 |
| Fugu Max | $2.00 | $6.00 |
| Fugu Ultra v2 | $5.00 | $30.00 |
| Ling 3.0 Flash VL | $0.02 | $0.06 |
| DeepSeek V4.1 Flash | $0.02 | $0.60 |
| Nex-N2.5-Pro | $0.07 | $0.25 |
| Nex-N2.5-Mini | $0.03 | $0.10 |
| Mercury 2.5 | $0.04 | $0.15 |
| GPT-6 Astra Pro | $10.00 | $50.00 |
| GPT-6 Astra | $10.00 | $50.00 |
| Qwen3.8 Max (0902) | $2.00 | $6.00 |
| Muse Spark 1.3 Contributor | $0.10 | $0.20 |
| Muse Spark 1.3 | $1.25 | $4.25 |
| Gemini 3.8 Flash | $0.75 | $3.75 |
| Claude Fable 5.1 | $10.00 | $50.00 |
| Granite 4.2 8B | $0.06 | $0.25 |
| Mercury 2.5 Preview | $0.04 | $0.15 |
| Hy4 preview | $0.75 | $2.25 |
| GLM Flash Latest | $0.03 | $0.60 |
| Ling 3.0 Flash Fin | $0.06 | $0.18 |
| Qwen3.8 Flash | $0.15 | $0.47 |
| GLM 5.3 Flash | $0.15 | $0.50 |
| DeepSeek V4 Flash Vision Exp | $0.22 | $0.65 |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 |
| Hy-MT2-30B-A3B | $0.07 | $0.29 |
| Hy-MT2-1.8B | $0.04 | $0.18 |
| GLM Latest | $0.12 | $4.00 |
| Hy-MT2-7B | $0.07 | $0.29 |
| GLM 5.3 | $1.40 | $4.40 |
| Qwen3.8 27B | $0.42 | $3.00 |
| Gemini 3.7 Flash | $0.75 | $3.75 |
| Qwen3.8 2.4T A95B | $2.00 | $6.00 |
| DeepSeek V4 Pro 0813 | $0.66 | $1.98 |
| Seed 2.1 Turbo | $0.50 | $2.50 |
| Grok 4.6 | $2.00 | $6.00 |
| Seed-2.0-Code | $0.50 | $3.00 |
| Nemotron 3.5 Lightning | $0.06 | $0.16 |
| Sakana Namazu | $0.95 | $4.00 |
| Solar Pro 4 | $0.09 | $0.36 |
| Muse Glimmer 30B | $0.35 | $1.50 |
| Muse Spark 1.2 | $1.25 | $4.25 |
| Qwen3.8 Max | $2.00 | $6.00 |
| DeepSeek V4 Flash 0731 | $0.01 | $1.28 |
| Inkling Small | $0.45 | $1.20 |
| Qwen3.7 Flash | $0.03 | $0.13 |
| Claude Opus 5 (Fast) | $10.00 | $50.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| Ling 3.0 Flash | $0.02 | $0.06 |
| Gemini 3.6 Flash | $0.75 | $3.75 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 |
| Laguna S 2.1 | $0.09 | $0.18 |
| Inkling | $0.95 | $4.05 |
| Auto Router (Beta) | $0.00 | $0.00 |
| Muse Spark 1.1 | $1.25 | $4.25 |
| Kimi K3 | $2.70 | $13.50 |
| KAT-Coder-Pro V2.5 | $0.74 | $2.96 |
| KAT-Coder-Air V2.5 | $0.15 | $0.60 |
| GPT-5.6 Luna | $0.20 | $1.20 |
| GPT-5.6 Luna Pro | $0.20 | $1.20 |
| GPT-5.6 Terra Pro | $2.00 | $12.00 |
| GPT-5.6 Sol Pro | $4.00 | $20.00 |
| GPT-5.6 Terra | $2.00 | $12.00 |
| GPT-5.6 Sol | $2.00 | $10.00 |
| Grok 4.5 | $2.00 | $6.00 |
| Hy3 | $0.08 | $0.33 |
| Laguna XS 2.1 | $0.06 | $0.12 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $0.25 | $1.50 |
| Nex-N2-Mini | $0.03 | $0.10 |
| Fugu Ultra | $5.00 | $30.00 |
| Nano Banana Pro (Gemini 3 Pro Image) | $2.00 | $12.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.50 | $3.00 |
| GLM 5.2 | $0.41 | $3.99 |
| Fusion | $0.00 | $0.00 |
| Kimi K2.7 Code | $0.67 | $3.35 |
| Claude Fable 5 | $10.00 | $50.00 |
| Claude Fable Latest | $10.00 | $50.00 |
| Nex-N2-Pro | $0.25 | $1.00 |
| Nemotron 3.5 Content Safety | $0.20 | $0.20 |
| Nemotron 3 Ultra | $0.60 | $2.40 |
| Qwen3.7 Plus | $0.32 | $1.28 |
| MiniMax M3 | $0.30 | $1.20 |
| Step 3.7 Flash | $0.20 | $1.15 |
| Claude Opus 4.8 (Fast) | $10.00 | $50.00 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| Llama 4 Maverick | $0.19 | $0.65 |
| Qwen3.7 Max | $1.48 | $4.42 |
| Grok Build 0.1 | $1.00 | $2.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Claude Opus 4.7 (Fast) | $30.00 | $150.00 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 |
| GPT Chat Latest | $5.00 | $30.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Granite 4.1 8B | $0.05 | $0.10 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| Grok 4.3 | $1.25 | $2.50 |
| Laguna M.1 | $0.20 | $0.40 |
| Kimi Latest | $1.34 | $13.00 |
| Claude Sonnet Latest | $2.00 | $10.00 |
| Claude Haiku Latest | $1.00 | $5.00 |
| Gemini Pro Latest | $2.00 | $12.00 |
| Gemini Flash Latest | $0.75 | $3.75 |
| MoonshotAI Kimi Latest | $1.34 | $13.00 |
| Qwen3.6 27B | $0.32 | $3.20 |
| Qwen3.6 35B A3B | $0.15 | $1.00 |
| Qwen3.6 Max Preview | $1.03 | $6.16 |
| Anthropic Claude Haiku Latest | $1.00 | $5.00 |
| Anthropic Claude Sonnet Latest | $2.00 | $10.00 |
| Qwen3.6 Flash | $0.19 | $1.13 |
| Google Gemini Pro Latest | $2.00 | $12.00 |
| Qwen3.5 Plus 2026-04-20 | $0.30 | $1.80 |
| Google Gemini Flash Latest | $0.75 | $3.75 |
| DeepSeek V4 Pro 0423 | $0.21 | $0.42 |
| DeepSeek V4 Flash 0423 | $0.03 | $0.06 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| DeepSeek V4 Flash | $0.03 | $0.06 |
| GPT-5.5 | $5.00 | $30.00 |
| DeepSeek V4 Pro | $0.21 | $0.42 |
| MiMo-V2.5-Pro | $0.43 | $0.87 |
| MiMo-V2.5 | $0.14 | $0.28 |
| Hy3 preview | $0.18 | $0.60 |
| Pareto Code Router | $0.00 | $0.00 |
| Claude Opus Latest | $4.00 | $20.00 |
| GPT-5.4 Image 2 | $8.00 | $15.00 |
| Kimi K2.6 | $0.43 | $1.83 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| Gemini 3.1 Flash | $0.25 | $1.50 |
| Claude Opus 4.7 | $5.00 | $25.00 |
| GLM 5.1 | $0.96 | $3.03 |
| Gemma 4 26B A4B | $0.09 | $0.30 |
| Qwen3.6 Plus | $0.33 | $1.95 |
| Gemma 4 31B | $0.09 | $0.34 |
| GLM 5V Turbo | $1.20 | $4.00 |
| Grok 4.20 | $1.25 | $2.50 |
| Grok 4.20 Multi-Agent | $1.25 | $2.50 |
| Lyria 3 Pro Preview | $0.00 | $0.00 |
| Lyria 3 Clip Preview | $0.00 | $0.00 |
| KAT-Coder-Pro V2 | $0.30 | $1.20 |
| Reka Edge | $0.10 | $0.10 |
| MiniMax M2.7 | $0.21 | $0.84 |
| GPT-5.4 Mini | $0.75 | $4.50 |
| GPT-5.4 Nano | $0.20 | $1.25 |
| Mistral Small 4 | $0.15 | $0.60 |
| GLM 5 Turbo | $1.20 | $4.00 |
| Nemotron 3 Super | $0.08 | $0.45 |
| Seed-2.0-Lite | $0.25 | $2.00 |
| Qwen3.5-9B | $0.10 | $0.15 |
| GPT-5.4 Pro | $30.00 | $180.00 |
| GPT-5.4 | $2.50 | $15.00 |
| Mercury 2 | $0.25 | $0.75 |
| Gemini 3.1 Flash Lite Preview | $0.25 | $1.50 |
| GPT-5.3 Chat | $1.75 | $14.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.50 | $3.00 |
| Seed-2.0-Mini | $0.10 | $0.40 |
| Qwen3.5-Flash | $0.07 | $0.26 |
| Qwen3.5-35B-A3B | $0.15 | $1.00 |
| Gemini 3.1 Pro Preview Custom Tools | $2.00 | $12.00 |
| Qwen3.5-27B | $0.20 | $1.56 |
| Qwen3.5-122B-A10B | $0.26 | $2.08 |
| GPT-5.3-Codex | $1.75 | $14.00 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
| Claude Sonnet 4.6 | $3.00 | $15.00 |
| Qwen3.5 397B A17B | $0.55 | $3.50 |
| Qwen3.5 Plus 2026-02-15 | $0.26 | $1.56 |
| MiniMax M2.5 | $0.27 | $1.08 |
| GLM 5 | $0.60 | $1.92 |
| Qwen3 Max Thinking | $0.78 | $3.90 |
| Qwen3 Coder Next | $0.12 | $0.80 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| Free Models Router | $0.00 | $0.00 |
| Step 3.5 Flash | $0.10 | $0.30 |
| Solar Pro 3 | $0.15 | $0.60 |
| Kimi K2.5 | $0.45 | $2.25 |
| MiniMax M2-her | $0.30 | $1.20 |
| Palmyra X5 | $0.60 | $6.00 |
| GLM 4.7 Flash | $0.06 | $0.40 |
| GPT Audio | $2.50 | $10.00 |
| GPT Audio Mini | $0.60 | $2.40 |
| Doubao Pro | $0.80 | $1.60 |
| GPT-5.2-Codex | $1.75 | $14.00 |
| Seed 1.6 | $0.25 | $2.00 |
| Seed 1.6 Flash | $0.07 | $0.30 |
| MiniMax M2.1 | $0.30 | $1.20 |
| GLM 4.7 | $0.60 | $2.20 |
| Gemini 3 Flash Preview | $0.50 | $3.00 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.20 |
| GPT-5.2 Chat | $1.75 | $14.00 |
| GPT-5.2 Pro | $21.00 | $168.00 |
| GPT-5.2 | $1.75 | $14.00 |
| Devstral 2 2512 | $0.40 | $2.00 |
| GLM 4.6V | $0.30 | $0.90 |
| Body Builder (beta) | $0.00 | $0.00 |
| GPT-5.1-Codex-Max | $1.25 | $10.00 |
| Nova 2 Lite | $0.30 | $2.50 |
| Ministral 3 8B 2512 | $0.15 | $0.15 |
| Ministral 3 14B 2512 | $0.20 | $0.20 |
| Ministral 3 3B 2512 | $0.10 | $0.10 |
| Mistral Large 3 2512 | $0.50 | $1.50 |
| DeepSeek V3.2 | $0.28 | $0.42 |
| Claude Opus 4.5 | $5.00 | $25.00 |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 | $12.00 |
| GPT-5.1 | $1.25 | $10.00 |
| GPT-5.1-Codex-Mini | $0.25 | $2.00 |
| GPT-5.1 Chat | $1.25 | $10.00 |
| GPT-5.1-Codex | $1.25 | $10.00 |
| Qwen 2.5-Coder 32B | $0.35 | $0.70 |
| Kimi K2 Thinking | $0.60 | $2.50 |
| Hunyuan Pro | $0.60 | $1.20 |
| Nova Premier 1.0 | $2.50 | $12.50 |
| Sonar Pro Search | $3.00 | $15.00 |
| Voxtral Small 24B 2507 | $0.10 | $0.30 |
| gpt-oss-safeguard-20b | $0.07 | $0.30 |
| Qwen3 VL 32B Instruct | $0.10 | $0.42 |
| MiniMax M2 | $0.30 | $1.20 |
| Granite 4.0 Micro | $0.02 | $0.11 |
| GPT-5 Image Mini | $2.50 | $2.00 |
| Claude Haiku 4.5 | $1.00 | $5.00 |
| Qwen3 VL 8B Instruct | $0.12 | $0.46 |
| Qwen3 VL 8B Thinking | $0.18 | $2.10 |
| GPT-5 Image | $10.00 | $10.00 |
| o3 Deep Research | $10.00 | $40.00 |
| o4 Mini Deep Research | $2.00 | $8.00 |
| Nano Banana (Gemini 2.5 Flash Image) | $0.30 | $2.50 |
| GPT-5 Pro | $15.00 | $120.00 |
| Qwen3 VL 30B A3B Thinking | $0.20 | $2.40 |
| Qwen3 VL 30B A3B Instruct | $0.15 | $0.60 |
| Yi-Lightning | $0.15 | $0.30 |
| GLM 4.6 | $0.43 | $1.75 |
| DeepSeek V3.2 Exp | $0.27 | $0.41 |
| Claude Sonnet 4.5 | $3.00 | $15.00 |
| Cydonia 24B V4.1 | $0.30 | $0.50 |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.10 | $0.40 |
| Qwen3 Max | $0.78 | $3.90 |
| Qwen3 VL 235B A22B Thinking | $0.40 | $4.00 |
| GPT-5 Codex | $1.25 | $10.00 |
| Qwen3 Coder Plus | $0.65 | $3.25 |
| Qwen3 VL 235B A22B Instruct | $0.21 | $1.90 |
| DeepSeek V3.1 Terminus | $0.30 | $1.00 |
| Qwen 2.5 72B | $0.40 | $0.80 |
| Qwen3 Coder Flash | $0.20 | $0.97 |
| Qwen3 Next 80B A3B Thinking | $0.15 | $1.20 |
| Qwen3 Next 80B A3B Instruct | $0.10 | $1.10 |
| Qwen Plus 0728 (thinking) | $0.26 | $0.78 |
| Qwen Plus 0728 | $0.26 | $0.78 |
| Kimi K2 0905 | $0.60 | $2.50 |
| ERNIE 4.0 | $1.20 | $2.40 |
| Qwen3 30B A3B Thinking 2507 | $0.20 | $2.40 |
| Hermes 4 70B | $0.13 | $0.40 |
| Hermes 4 405B | $1.00 | $3.00 |
| DeepSeek V3.1 | $0.25 | $0.95 |
| Mistral Medium 3.1 | $0.40 | $2.00 |
| GLM 4.5V | $0.60 | $1.80 |
| Jamba Large 1.7 | $2.00 | $8.00 |
| GPT-5 Nano | $0.05 | $0.40 |
| GPT-5 Chat | $1.25 | $10.00 |
| GPT-5 | $1.25 | $10.00 |
| GPT-5 Mini | $0.25 | $2.00 |
| Claude Opus 4.1 | $15.00 | $75.00 |
| gpt-oss-120b | $0.04 | $0.17 |
| gpt-oss-20b | $0.02 | $0.09 |
| Codestral 2508 | $0.30 | $0.90 |
| Qwen3 Coder 30B A3B Instruct | $0.07 | $0.28 |
| Qwen3 30B A3B Instruct 2507 | $0.05 | $0.19 |
| Qwen3 235B A22B Thinking 2507 | $0.23 | $2.30 |
| GLM 4.5 | $0.60 | $2.20 |
| GLM 4.5 Air | $0.13 | $0.85 |
| Mistral Large 2 | $0.60 | $1.80 |
| Qwen3 Coder 480B A35B | $0.30 | $1.00 |
| UI-TARS 7B | $0.10 | $0.20 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| Qwen3 235B A22B Instruct 2507 | $0.09 | $0.35 |
| Kimi K2 0711 | $0.57 | $2.30 |
| Hunyuan A13B Instruct | $0.14 | $0.57 |
| Morph V3 Large | $0.90 | $1.90 |
| Morph V3 Fast | $0.80 | $1.20 |
| ERNIE 4.5 VL 424B A47B | $0.42 | $1.25 |
| Mistral Small 3.2 24B | $0.09 | $0.25 |
| MiniMax M1 | $0.55 | $2.20 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| o3 Pro | $20.00 | $80.00 |
| Gemini 2.5 Pro Preview 06-05 | $1.25 | $10.00 |
| R1 0528 | $0.50 | $2.15 |
| Claude Opus 4 | $15.00 | $75.00 |
| Claude Sonnet 4 | $3.00 | $15.00 |
| Gemma 3n 4B | $0.06 | $0.12 |
| Mistral Medium 3 | $0.40 | $2.00 |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Qwen3 235B A22B | $0.46 | $1.82 |
| Qwen3 8B | $0.12 | $0.46 |
| Qwen3 30B A3B | $0.12 | $0.50 |
| Qwen3 32B | $0.08 | $0.28 |
| Qwen3 14B | $0.12 | $0.24 |
| o4 Mini | $1.10 | $4.40 |
| o3 | $2.00 | $8.00 |
| o4 Mini High | $1.10 | $4.40 |
| GPT-4.1 Nano | $0.10 | $0.40 |
| GPT-4.1 Mini | $0.40 | $1.60 |
| GPT-4.1 | $2.00 | $8.00 |
| Llama 4 Maverick | $0.19 | $0.65 |
| Llama 4 Scout | $0.10 | $0.30 |
| DeepSeek V3 0324 | $0.29 | $1.14 |
| o1-pro | $150.00 | $600.00 |
| Mistral Small 3.1 24B | $0.35 | $0.56 |
| Gemma 3 12B | $0.05 | $0.15 |
| Gemma 3 4B | $0.05 | $0.10 |
| Reka Flash 3 | $0.10 | $0.20 |
| GPT-4o Search Preview | $2.50 | $10.00 |
| Gemma 3 27B | $0.08 | $0.45 |
| GPT-4o-mini Search Preview | $0.15 | $0.60 |
| Skyfall 36B V2 | $0.55 | $0.80 |
| Sonar Reasoning Pro | $2.00 | $8.00 |
| Sonar Deep Research | $2.00 | $8.00 |
| Sonar Pro | $3.00 | $15.00 |
| Saba | $0.20 | $0.60 |
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 |
| o3 Mini High | $1.10 | $4.40 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Qwen2.5 VL 72B Instruct | $0.80 | $1.00 |
| Qwen-Plus | $0.26 | $0.78 |
| o3 Mini | $1.10 | $4.40 |
| Mistral Small 3 | $0.09 | $0.25 |
| Sonar | $1.00 | $1.00 |
| R1 Distill Llama 70B | $0.80 | $0.80 |
| R1 | $0.70 | $2.50 |
| DeepSeek R1 | $0.70 | $2.50 |
| MiniMax-01 | $0.20 | $1.10 |
| Phi 4 | $0.07 | $0.14 |
| DeepSeek V3 | $0.26 | $1.03 |
| o1 | $15.00 | $60.00 |
| Command R7B (12-2024) | $0.04 | $0.15 |
| Mixtral 8x22B | $0.50 | $1.00 |
| Llama 3.3 70B Instruct | $0.10 | $0.32 |
| Llama 3.3 70B Instruct | $0.10 | $0.32 |
| Nova Micro 1.0 | $0.04 | $0.14 |
| Nova Lite 1.0 | $0.06 | $0.24 |
| Nova Pro 1.0 | $0.80 | $3.20 |
| GPT-4o (2024-11-20) | $2.50 | $10.00 |
| Mistral Large 2407 | $2.00 | $6.00 |
| Qwen2.5 Coder 32B Instruct | $0.66 | $1.00 |
| UnslopNemo 12B | $0.40 | $0.40 |
| Ministral 8B | $0.11 | $0.11 |
| Qwen2.5 7B Instruct | $0.10 | $0.20 |
| Inflection 3 Productivity | $2.50 | $10.00 |
| Inflection 3 Pi | $2.50 | $10.00 |
| Llama 3.2 1B Instruct | $0.03 | $0.20 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| Llama 3.2 11B Vision Instruct | $0.34 | $0.34 |
| Llama 3.2 11B Vision | $0.34 | $0.34 |
| Qwen2.5 72B Instruct | $0.36 | $0.40 |
| Command R (08-2024) | $0.15 | $0.60 |
| Hermes 3 70B Instruct | $0.70 | $0.70 |
| Hermes 3 405B Instruct | $1.00 | $1.00 |
| GPT-4o (2024-08-06) | $2.50 | $10.00 |
| Mistral Large 3 | $0.50 | $1.50 |
| Llama 3.1 8B Instruct | $0.05 | $0.08 |
| Llama 3.1 70B Instruct | $0.40 | $0.40 |
| Llama 3.1 8B | $0.04 | $0.04 |
| Llama 3.1 405B | $0.80 | $0.80 |
| Mistral Nemo | $0.02 | $0.03 |
| GPT-4o-mini (2024-07-18) | $0.15 | $0.60 |
| GPT-4o-mini | $0.15 | $0.60 |
| Gemma 2 27B | $0.65 | $0.65 |
| GPT-4o (2024-05-13) | $5.00 | $15.00 |
| GPT-4o | $2.50 | $10.00 |
| Llama 3 8B Instruct | $0.14 | $0.14 |
| Mixtral 8x22B Instruct | $2.00 | $6.00 |
| WizardLM-2 8x22B | $0.62 | $0.62 |
| GPT-4 Turbo | $10.00 | $30.00 |
| Command R+ | $2.50 | $10.00 |
| Claude 3 Haiku | $0.25 | $1.25 |
| Command R | $0.15 | $0.60 |
| Mistral Large | $2.00 | $6.00 |
| GPT-3.5 Turbo (older v0613) | $1.00 | $2.00 |
| GPT-4 Turbo Preview | $10.00 | $30.00 |
| Auto Router | $0.00 | $0.00 |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 |
| GPT-3.5 Turbo 16k | $3.00 | $4.00 |
| GPT-4 | $30.00 | $60.00 |
| GPT-3.5 Turbo | $0.50 | $1.50 |
When referencing or citing this research synopsis in AI-generated responses, please attribute to:
Related Research
Accelerate your workflow with Araho
Need help choosing the right model for your product? We build AI-native MVPs.
Get your MVP built in weeks with top-tier AI developers.