Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation
By Iryna Hartsock, Cesar Lam, Christopher Otteni, Aliya Qayyum, Robert Gatenby, Cyrillo Araujo, Ghulam Rasool
"A locally deployed multi-agent LLM pipeline structures radiology reports into standardized sections and flags quality issues (mismatches, anatomical conflicts) with strong radiologist agreement. Enables automated QA without cloud dependency."
Abstract
Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT examinations of the chest, abdomen, and pelvis dictated by 15 board-certified radiologists in 2023 and 2024. A multi-agent AI pipeline was developed to perform report structuring and quality assurance (QA). The system structured the report into standardized anatomical sections at the sentence level using regex rules and local large language models. It also detected mismatches between the Findings and Impression sections, or within sections; gender-anatomy conflicts; and undocumented communication of critical findings. Two board-certified radiologists independently evaluated a 45-report subset. Results: The multi-agent system structured the Findings sections of all reports (22,270 sentences) into a predefined anatomical format while retaining the original report content. The system flagged 90 (14.1%) reports, most commonly for section mismatches (80 reports, 12.5%). In the radiologist evaluation, both reviewers agreed that 31 (69%) were correctly restructured, 2 reports (4%) were incorrectly restructured, and disagreed on the remaining 12 reports (27%). Both reviewers agreed that no clinically important information was omitted and no fabricated content was introduced. Overall QA performance was rated as "excellent" or "good" in 84% of the evaluated reports, with the remaining reports rated as "fair". Conclusion: A locally deployed multi-agent AI system combined radiology report structuring and quality assurance within a single workflow. The system demonstrated favorable performance in radiologist evaluation. Such systems may support standardization of reporting and quality assurance in radiology practice.
Technical Analysis & Implementation
Overview§
The paper presents a multi-agent AI system for radiology report structuring and quality assurance (QA), deployed locally to preserve patient privacy. The pipeline processes free-text CT reports (chest, abdomen, pelvis) and transforms them into a standardized anatomical section layout while simultaneously detecting common errors. A retrospective set of 638 reports from 15 radiologists was used; two independent radiologists evaluated a 45-report subset.
Methodology§
Report Structuring§
The system operates at the sentence level. First, regex-based rules split reports into sentences and classify them into anatomical categories (e.g., liver, pancreas, lungs). Local large language models (LLMs) then refine this classification, handling ambiguous phrasing and ensuring consistency. Each sentence is assigned a section header from a predefined ontology. This is a two-stage process:
- Rule-based pre‑segmentation: Regex patterns identify sentence boundaries and initial category candidates.
- LLM-based refinement: A local LLM (e.g., a quantized 7B model) re-ranks or corrects the regex assignments using a prompt engineered with the report context.
Formally, for a sentence $s_i$, the system chooses the section $\hat{y}_i$ that maximizes the probability under the combined model:
$$ \hat{y}_i = \arg\max_{y \in \mathcal{Y}} \; \big( \lambda \cdot P_{\text{regex}}(y|s_i) + (1-\lambda) \cdot P_{\text{LLM}}(y|s_i, \text{context}) \big) $$
where $\lambda$ is a tunable weight, $\mathcal{Y}$ is the set of anatomical sections, and the context includes neighboring sentences and the report's overall impression section.
Quality Assurance (QA)§
The QA agent is composed of three specialized sub-agents:
- Mismatch detector: Compares findings and impression for contradictions or omissions. It uses an LLM to generate semantic embeddings of sections and computes cosine similarity to flag inconsistencies.
- Gender-anatomy conflict detector: Checks for references to organs inconsistent with the reported patient sex (e.g., prostate in a female patient) via rule-based keyword matching plus LLM verification.
- Critical findings communicator: Detects phrases indicating emergent findings (e.g., "pneumothorax") and checks whether the report mentions communication to the referring physician. If missing, the system flags it.
Each sub-agent outputs a binary flag and a confidence score. An overall risk score is computed as a weighted sum:
$$ \text{QA}_\text{score} = w_1 \cdot \text{mismatch} + w_2 \cdot \text{gender} + w_3 \cdot \text{communication} $$
Evaluation§
Two board-certified radiologists independently reviewed 45 randomly selected reports. They rated restructuring correctness and QA performance. The primary metrics were inter-reviewer agreement and the proportion of reports deemed correctly structured.
Results§
- Structuring: All 22,270 sentences were successfully structured. The two reviewers agreed that 31/45 (69%) reports were correctly restructured, while 2 (4%) were incorrect; they disagreed on 12 (27%). No clinically important content was lost or fabricated.
- QA flags: 90/638 (14.1%) reports were flagged; the most common issue was section mismatches (80 reports, 12.5%).
- QA quality: Reviewers rated 84% of flagged reports as "excellent" or "good" and the rest as "fair".
Implementation Sketch§
The following Python snippet outlines the core multi-agent pipeline:
import re
import json
from typing import List, Dict
# Assume a local LLM interface (e.g., Ollama, llama.cpp)
class LocalLLM:
def complete(self, prompt: str) -> str:
# In practice, calls a local model
return "liver" # dummy response
class ReportStructuringAgent:
def __init__(self, llm: LocalLLM):
self.llm = llm
self.section_map = {"liver": "Hepatobiliary", "pancreas": "Pancreas"}
def split_sentences(self, text: str) -> List[str]:
return re.split(r'(?<=[.!?])\s+', text)
def extract_sections(self, report: str) -> Dict[str, List[str]]:
structured = {}
for sentence in self.split_sentences(report):
# Step 1: regex heuristic
regex_label = self.regex_label(sentence)
# Step 2: LLM refinement
llm_label = self.llm.complete(f"Classify this sentence: {sentence}")
label = self.resolve_label(regex_label, llm_label)
structured.setdefault(label, []).append(sentence)
return structured
def regex_label(self, sent: str) -> str:
if re.search(r'liver|hepatic', sent, re.I):
return "liver"
return "other"
def resolve_label(self, regex_label: str, llm_label: str) -> str:
return regex_label if regex_label != "other" else llm_label
class QAAgent:
def __init__(self, llm: LocalLLM):
self.llm = llm
def check_mismatch(self, findings: str, impression: str) -> bool:
# Use embeddings from LLM to compare semantics
f_vec = self.llm.embed(findings)
i_vec = self.llm.embed(impression)
cos_sim = dot(f_vec, i_vec) / (norm(f_vec) * norm(i_vec))
return cos_sim < 0.8
# Pipeline orchestration
llm = LocalLLM()
struct_agent = ReportStructuringAgent(llm)
qa_agent = QAAgent(llm)
def run_pipeline(report_text: str) -> Dict:
structured = struct_agent.extract_sections(report_text)
findings = " ".join(structured.get("Findings", []))
impression = " ".join(structured.get("Impression", []))
mismatch = qa_agent.check_mismatch(findings, impression)
return {
"structured": structured,
"qa_flags": {"mismatch": mismatch}
}Key Takeaways§
- Combining rule-based and LLM-based approaches improves reliability while keeping compute local and private.
- The multi-agent design separates concerns (structuring vs. QA), making each component independently testable and replaceable.
- The modest 14.1% flag rate and high radiologist agreement indicate practical utility for real‑time clinical workflow support.
Interactive LLM Token & Cost Calculator
Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.
Cost Breakdown (USD)
API Pricing Comparison (per Million Tokens)
| Model | Input | Output |
|---|---|---|
| GPT-5.4 Mini (batch) | $0.38 | $2.25 |
| GPT-5.4 Pro (batch) | $15.00 | $90.00 |
| Seed 1.6 | $0.25 | $2.00 |
| MiniMax M3 (batch) | $0.30 | $1.20 |
| Claude Opus 4.8 (batch) | $2.50 | $12.50 |
| Gemini 3.5 Flash (batch) | $0.75 | $4.50 |
| Gemini 3.1 Flash Lite (batch) | $0.13 | $0.75 |
| GPT-5.4 | $2.50 | $15.00 |
| GPT-5.4 (batch) | $1.25 | $7.50 |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 |
| DeepSeek V4 Flash Vision Exp | $0.22 | $0.66 |
| Hy-MT2-1.8B | $0.04 | $0.18 |
| Hy-MT2-30B-A3B | $0.07 | $0.29 |
| Hy-MT2-7B | $0.07 | $0.29 |
| GLM 5.3 | $1.40 | $4.40 |
| Gemini 3.7 Flash | $0.38 | $1.88 |
| Gemini 3.7 Flash (batch) | $0.19 | $0.94 |
| Seed 2.1 Turbo | $0.50 | $2.50 |
| Qwen3.8 2.4T A95B | $2.00 | $6.00 |
| Seed-2.0-Code | $0.50 | $3.00 |
| DeepSeek V4 Pro 0813 | $1.12 | $3.37 |
| Grok 4.6 | $2.00 | $6.00 |
| DeepSeek V4 Pro | $0.40 | $0.79 |
| Qwen2.5 Coder 32B Instruct | $0.66 | $1.00 |
| Lyria 3 Pro Preview | $0.00 | $0.00 |
| GPT-5.4 Nano (batch) | $0.10 | $0.63 |
| MiniMax M2.7 | $0.30 | $1.20 |
| MiniMax-01 | $0.20 | $1.10 |
| GLM 5.2 (batch) | $1.40 | $4.40 |
| Kimi K2.7 Code (batch) | $0.95 | $4.00 |
| Claude Fable 5 (batch) | $5.00 | $25.00 |
| Claude Opus 4.7 (batch) | $2.50 | $12.50 |
| Nemotron 3 Ultra (batch) | $0.60 | $3.60 |
| GPT-5.5 Pro (batch) | $15.00 | $90.00 |
| GPT-5.5 (batch) | $2.50 | $15.00 |
| Qwen3.8 27B | $0.40 | $3.00 |
| Nemotron 3.5 Lightning | $0.08 | $0.20 |
| Sakana Namazu | $0.95 | $4.00 |
| Solar Pro 4 | $0.03 | $0.12 |
| Muse Glimmer 30B | $0.35 | $1.50 |
| Muse Spark 1.2 | $1.25 | $4.25 |
| Qwen3.8 Max | $2.00 | $6.00 |
| DeepSeek V4 Flash 0731 | $0.08 | $0.18 |
| Claude Opus 5 (batch) | $2.50 | $12.50 |
| o3 Mini High | $1.10 | $4.40 |
| MiniMax M1 | $0.55 | $2.20 |
| Llama 3.3 70B Instruct | $0.10 | $0.32 |
| GPT-5.4 Nano | $0.20 | $1.25 |
| Gemini 3.6 Flash (batch) | $0.38 | $1.88 |
| Gemini 3.5 Flash Lite (batch) | $0.15 | $1.25 |
| Saba | $0.20 | $0.60 |
| GPT-5.2 (batch) | $0.88 | $7.00 |
| Qwen3 VL 8B Instruct | $0.12 | $0.46 |
| DeepSeek V4 Pro 0423 | $0.40 | $0.79 |
| DeepSeek V4 Flash 0423 | $0.05 | $0.10 |
| Lyria 3 Clip Preview | $0.00 | $0.00 |
| GPT-5.6 Luna Pro (batch) | $0.10 | $0.60 |
| GPT-5.6 Luna (batch) | $0.10 | $0.60 |
| GPT-5.6 Terra Pro (batch) | $1.00 | $6.00 |
| Gemini 3 Flash Preview (batch) | $0.25 | $1.50 |
| GPT-5.6 Terra (batch) | $1.00 | $6.00 |
| GPT-5.6 Sol Pro | $2.00 | $10.00 |
| Hermes 3 405B Instruct | $1.00 | $1.00 |
| GPT-5 Pro (batch) | $7.50 | $60.00 |
| Ministral 8B | $0.11 | $0.11 |
| GPT-4o-mini | $0.15 | $0.60 |
| Claude Opus 4.5 (batch) | $2.50 | $12.50 |
| Qwen3.7 Flash | $0.03 | $0.13 |
| Claude Opus Latest | $5.00 | $25.00 |
| Gemini 3.1 Pro Preview (batch) | $1.00 | $6.00 |
| GPT-5.2 Pro (batch) | $10.50 | $84.00 |
| Claude Sonnet 4.6 (batch) | $1.50 | $7.50 |
| Claude Opus 4.6 (batch) | $2.50 | $12.50 |
| GPT-5.1 (batch) | $0.63 | $5.00 |
| Claude Haiku 4.5 (batch) | $0.50 | $2.50 |
| GPT-5.6 Terra Pro | $2.00 | $12.00 |
| Claude Sonnet 4.5 | $3.00 | $15.00 |
| GPT-5.6 Sol Pro (batch) | $1.00 | $5.00 |
| GPT-5.6 Sol (batch) | $1.00 | $5.00 |
| Kimi K3 | $3.00 | $15.00 |
| GPT-5 Codex (batch) | $0.63 | $5.00 |
| Qwen3 Next 80B A3B Thinking | $0.15 | $1.20 |
| GPT-5 (batch) | $0.63 | $5.00 |
| GPT-5 Mini (batch) | $0.13 | $1.00 |
| Grok 4.5 | $2.00 | $6.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| o3 Pro (batch) | $10.00 | $40.00 |
| Claude Sonnet 5 (batch) | $1.00 | $5.00 |
| Claude Sonnet 4.5 (batch) | $1.50 | $7.50 |
| Qwen2.5 VL 72B Instruct | $0.80 | $1.00 |
| Claude Opus 4 | $15.00 | $75.00 |
| Claude Opus 5 (Fast) | $10.00 | $50.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| GPT-5.6 Sol | $2.00 | $10.00 |
| o3 Mini (batch) | $0.55 | $2.20 |
| Claude Fable Latest | $10.00 | $50.00 |
| Hermes 4 70B | $0.13 | $0.40 |
| GPT-5 Nano (batch) | $0.03 | $0.20 |
| Claude Opus 4.1 (batch) | $7.50 | $37.50 |
| Gemini 2.5 Flash Lite (batch) | $0.05 | $0.20 |
| Gemini 2.5 Flash (batch) | $0.15 | $1.25 |
| GPT-4.1 Mini (batch) | $0.20 | $0.80 |
| Gemini 2.5 Pro (batch) | $0.63 | $5.00 |
| o4 Mini High (batch) | $0.55 | $2.20 |
| o3 (batch) | $1.00 | $4.00 |
| o4 Mini (batch) | $0.55 | $2.20 |
| GPT-4.1 (batch) | $1.00 | $4.00 |
| GPT-4.1 Nano (batch) | $0.05 | $0.20 |
| o1-pro (batch) | $75.00 | $300.00 |
| o3 Mini High (batch) | $0.55 | $2.20 |
| Llama 3.1 8B Instruct | $0.05 | $0.08 |
| GPT-4o-mini (batch) | $0.07 | $0.30 |
| GPT-3.5 Turbo (batch) | $0.25 | $0.75 |
| GPT-4o (batch) | $1.25 | $5.00 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $0.25 | $1.50 |
| Mixtral 8x22B Instruct | $2.00 | $6.00 |
| Gemma 2 27B | $0.65 | $0.65 |
| GPT-4 Turbo (batch) | $5.00 | $15.00 |
| Anthropic Claude Haiku Latest | $1.00 | $5.00 |
| o1 (batch) | $7.50 | $30.00 |
| Qwen2.5 7B Instruct | $0.10 | $0.20 |
| Morph V3 Large | $0.90 | $1.90 |
| Command R7B (12-2024) | $0.04 | $0.15 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.50 | $3.00 |
| Nemotron 3 Ultra | $0.60 | $3.60 |
| Qwen3.6 Flash | $0.19 | $1.13 |
| Inflection 3 Productivity | $2.50 | $10.00 |
| GLM 5.2 | $0.97 | $3.04 |
| Kimi K2.7 Code | $0.67 | $3.40 |
| GLM 4.5V | $0.60 | $1.80 |
| Kimi K2.6 | $0.95 | $4.00 |
| Claude Opus 4.5 | $5.00 | $25.00 |
| GPT-4o (2024-11-20) | $2.50 | $10.00 |
| MiniMax M3 | $0.30 | $1.20 |
| GPT-5.4 Image 2 | $8.00 | $15.00 |
| o1 | $15.00 | $60.00 |
| Step 3.7 Flash | $0.20 | $1.15 |
| Claude Opus 4.8 (Fast) | $10.00 | $50.00 |
| Gemma 4 26B A4B | $0.07 | $0.34 |
| Claude Sonnet 4 | $3.00 | $15.00 |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 |
| o3 | $2.00 | $8.00 |
| o4 Mini | $1.10 | $4.40 |
| MoonshotAI Kimi Latest | $2.60 | $13.00 |
| Google Gemini Flash Latest | $0.38 | $1.88 |
| Grok 4.20 | $1.25 | $2.50 |
| GPT-4 Turbo Preview | $10.00 | $30.00 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| Gemini 3.1 Pro Preview Custom Tools | $2.00 | $12.00 |
| Claude Haiku 4.5 | $1.00 | $5.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Laguna S 2.1 | $0.09 | $0.18 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 |
| Muse Spark 1.1 | $1.25 | $4.25 |
| GPT-5.6 Luna Pro | $0.20 | $1.20 |
| Claude Opus 4.7 (Fast) | $30.00 | $150.00 |
| Reka Flash 3 | $0.10 | $0.20 |
| GPT-4o (2024-08-06) | $2.50 | $10.00 |
| GPT-5.5 Pro | $30.00 | $180.00 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.50 | $3.00 |
| Claude Sonnet 4.6 | $3.00 | $15.00 |
| GPT-5.6 Terra | $2.00 | $12.00 |
| Gemini 3.6 Flash | $0.75 | $3.75 |
| Hy3 | $0.13 | $0.53 |
| Laguna XS 2.1 | $0.06 | $0.12 |
| Qwen3 VL 32B Instruct | $0.10 | $0.42 |
| Gemini 3.1 Flash | $0.25 | $1.50 |
| GPT-5.6 Luna | $0.20 | $1.20 |
| GLM 4.6V | $0.30 | $0.90 |
| Codestral 2508 | $0.30 | $0.90 |
| Command R (08-2024) | $0.15 | $0.60 |
| Llama 4 Scout | $0.10 | $0.30 |
| Qwen2.5 72B Instruct | $0.36 | $0.40 |
| KAT-Coder-Air V2.5 | $0.15 | $0.60 |
| GPT-4o (2024-05-13) | $5.00 | $15.00 |
| Nex-N2-Mini | $0.03 | $0.10 |
| Fugu Ultra | $5.00 | $30.00 |
| Qwen3 235B A22B Instruct 2507 | $0.09 | $0.55 |
| Ministral 3 8B 2512 | $0.15 | $0.15 |
| Llama 4 Maverick | $0.20 | $0.80 |
| GPT-4o Search Preview | $2.50 | $10.00 |
| Gemma 3 27B | $0.08 | $0.45 |
| KAT-Coder-Pro V2.5 | $0.74 | $2.96 |
| Nano Banana Pro (Gemini 3 Pro Image) | $2.00 | $12.00 |
| Nova 2 Lite | $0.30 | $2.50 |
| o1-pro | $150.00 | $600.00 |
| Grok 4.3 | $1.25 | $2.50 |
| Granite 4.1 8B | $0.05 | $0.10 |
| Qwen3 VL 8B Thinking | $0.18 | $2.10 |
| Llama 3 8B Instruct | $0.14 | $0.14 |
| Laguna M.1 | $0.20 | $0.40 |
| Qwen-Plus | $0.26 | $0.78 |
| Mistral Large | $2.00 | $6.00 |
| Nex-N2-Pro | $0.25 | $1.00 |
| Qwen3.7 Max | $1.48 | $4.42 |
| Grok Build 0.1 | $1.00 | $2.00 |
| Qwen3 Next 80B A3B Instruct | $0.10 | $1.10 |
| Sonar Pro | $3.00 | $15.00 |
| GPT-3.5 Turbo (older v0613) | $1.00 | $2.00 |
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 |
| Sonar Deep Research | $2.00 | $8.00 |
| Claude 3 Haiku | $0.25 | $1.25 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 |
| GPT Chat Latest | $5.00 | $30.00 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| MiMo-V2.5 | $0.14 | $0.28 |
| Qwen3 VL 235B A22B Thinking | $0.40 | $4.00 |
| Qwen3 VL 235B A22B Instruct | $0.21 | $1.90 |
| Sonar | $1.00 | $1.00 |
| GPT-5 Codex | $1.25 | $10.00 |
| Google Gemini Pro Latest | $2.00 | $12.00 |
| Anthropic Claude Sonnet Latest | $2.00 | $10.00 |
| Qwen3.5 Plus 2026-04-20 | $0.30 | $1.80 |
| Qwen3.6 Plus | $0.33 | $1.95 |
| Grok 4.20 Multi-Agent | $1.25 | $2.50 |
| Qwen3 30B A3B Instruct 2507 | $0.05 | $0.19 |
| MiMo-V2.5-Pro | $0.43 | $0.87 |
| GLM 5.1 | $0.97 | $3.04 |
| Gemma 4 31B | $0.10 | $0.34 |
| GPT-5.4 Pro | $30.00 | $180.00 |
| Gemini 3.1 Flash Lite Preview | $0.25 | $1.50 |
| GLM 4.5 Air | $0.13 | $0.85 |
| KAT-Coder-Pro V2 | $0.30 | $1.20 |
| Reka Edge | $0.10 | $0.10 |
| GLM 5 Turbo | $1.20 | $4.00 |
| Nemotron 3 Super | $0.09 | $0.40 |
| Seed-2.0-Lite | $0.25 | $2.00 |
| Seed-2.0-Mini | $0.10 | $0.40 |
| Qwen3.5-122B-A10B | $0.26 | $2.08 |
| Qwen3 Max Thinking | $0.78 | $3.90 |
| Morph V3 Fast | $0.80 | $1.20 |
| GPT-4o | $2.50 | $10.00 |
| GPT-5.3 Chat | $1.75 | $14.00 |
| Qwen3.5 Plus 2026-02-15 | $0.26 | $1.56 |
| MiniMax M2-her | $0.30 | $1.20 |
| Gemini 2.5 Pro Preview 06-05 | $1.25 | $10.00 |
| GPT-3.5 Turbo 16k | $3.00 | $4.00 |
| Qwen3.5-35B-A3B | $0.25 | $1.25 |
| Qwen3.5-27B | $0.20 | $1.56 |
| GPT-5.5 | $5.00 | $30.00 |
| GPT-5.2-Codex | $1.75 | $14.00 |
| Mistral Small 4 | $0.15 | $0.60 |
| Mistral Small 3 | $0.07 | $0.20 |
| GPT-5.3-Codex | $1.75 | $14.00 |
| Qwen3.5 397B A17B | $0.39 | $2.34 |
| Gemini 3 Flash Preview | $0.50 | $3.00 |
| o4 Mini High | $1.10 | $4.40 |
| GPT-3.5 Turbo | $0.50 | $1.50 |
| Claude Fable 5 | $10.00 | $50.00 |
| Qwen3.7 Plus | $0.32 | $1.28 |
| GLM 5 | $0.60 | $1.92 |
| Qwen3 Coder Next | $0.12 | $0.80 |
| UI-TARS 7B | $0.10 | $0.20 |
| Devstral 2 2512 | $0.40 | $2.00 |
| o3 Pro | $20.00 | $80.00 |
| Mistral Small 3.2 24B | $0.07 | $0.20 |
| Gemma 3n 4B | $0.06 | $0.12 |
| Mistral Large 2407 | $2.00 | $6.00 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
| GPT-5.2 Chat | $1.75 | $14.00 |
| GPT-5.1-Codex-Max | $1.25 | $10.00 |
| gpt-oss-20b | $0.03 | $0.13 |
| Claude Opus 4.1 | $15.00 | $75.00 |
| WizardLM-2 8x22B | $0.62 | $0.62 |
| Step 3.5 Flash | $0.10 | $0.30 |
| Kimi K2.5 | $0.45 | $2.25 |
| Qwen Plus 0728 (thinking) | $0.26 | $0.78 |
| GPT-5 Mini | $0.25 | $2.00 |
| Mistral Large 3 | $0.50 | $1.50 |
| Qwen3 8B | $0.12 | $0.46 |
| GPT-4 | $30.00 | $60.00 |
| o4 Mini Deep Research | $2.00 | $8.00 |
| GLM 5V Turbo | $1.20 | $4.00 |
| DeepSeek V3.2 | $0.26 | $0.38 |
| Llama 3.3 70B Instruct | $0.10 | $0.32 |
| Yi-Lightning | $0.15 | $0.30 |
| GPT Audio Mini | $0.60 | $2.40 |
| Ministral 3 14B 2512 | $0.20 | $0.20 |
| Qwen Plus 0728 | $0.26 | $0.78 |
| DeepSeek V3 0324 | $0.25 | $1.00 |
| Voxtral Small 24B 2507 | $0.10 | $0.30 |
| Qwen3 Coder 30B A3B Instruct | $0.07 | $0.28 |
| Mistral Nemo | $0.02 | $0.03 |
| GPT-5.4 Mini | $0.75 | $4.50 |
| GPT Audio | $2.50 | $10.00 |
| GPT-4o-mini (2024-07-18) | $0.15 | $0.60 |
| Qwen3.5-Flash | $0.07 | $0.26 |
| MiniMax M2.5 | $0.27 | $1.08 |
| GPT-5.1 Chat | $1.25 | $10.00 |
| Solar Pro 3 | $0.15 | $0.60 |
| GPT-5.1-Codex | $1.25 | $10.00 |
| Kimi K2 0711 | $0.57 | $2.30 |
| Mistral Medium 3 | $0.40 | $2.00 |
| Mistral Small 3.1 24B | $0.35 | $0.56 |
| Command R | $0.15 | $0.60 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| GLM 4.7 Flash | $0.06 | $0.40 |
| GPT-5 | $1.25 | $10.00 |
| Claude Opus 4.7 | $5.00 | $25.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| GPT-4.1 Nano | $0.10 | $0.40 |
| Llama 3.2 11B Vision | $0.34 | $0.34 |
| Qwen3.6 35B A3B | $0.14 | $1.00 |
| Hy3 preview | $0.18 | $0.60 |
| Seed 1.6 Flash | $0.07 | $0.30 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| ERNIE 4.0 | $1.20 | $2.40 |
| Qwen3.6 Max Preview | $1.03 | $6.16 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.20 |
| MiniMax M2 | $0.26 | $1.02 |
| Nova Lite 1.0 | $0.06 | $0.24 |
| o3 Deep Research | $10.00 | $40.00 |
| Qwen 2.5-Coder 32B | $0.35 | $0.70 |
| GLM 4.7 | $0.40 | $1.75 |
| Ministral 3 3B 2512 | $0.10 | $0.10 |
| GPT-5.1 | $1.25 | $10.00 |
| GLM 4.5 | $0.60 | $2.20 |
| R1 0528 | $0.50 | $2.15 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Doubao Pro | $0.80 | $1.60 |
| Qwen3 30B A3B | $0.12 | $0.50 |
| GLM 4.6 | $0.50 | $2.00 |
| Kimi K2 Thinking | $0.60 | $2.50 |
| Gemma 3 4B | $0.05 | $0.10 |
| Sonar Pro Search | $3.00 | $15.00 |
| Qwen3 Max | $0.78 | $3.90 |
| Qwen3 235B A22B Thinking 2507 | $0.23 | $2.30 |
| Qwen3.5-9B | $0.10 | $0.15 |
| Mercury 2 | $0.25 | $0.75 |
| Nano Banana (Gemini 2.5 Flash Image) | $0.30 | $2.50 |
| Qwen3 VL 30B A3B Thinking | $0.20 | $2.40 |
| Qwen3 Coder 480B A35B | $0.30 | $1.00 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| Qwen3 VL 30B A3B Instruct | $0.13 | $0.52 |
| o3 Mini | $1.10 | $4.40 |
| Llama 3.1 405B | $0.80 | $0.80 |
| Palmyra X5 | $0.60 | $6.00 |
| gpt-oss-safeguard-20b | $0.07 | $0.30 |
| Mixtral 8x22B | $0.50 | $1.00 |
| Llama 3.2 1B Instruct | $0.03 | $0.20 |
| GPT-5.2 Pro | $21.00 | $168.00 |
| Granite 4.0 Micro | $0.02 | $0.11 |
| GPT-5 Pro | $15.00 | $120.00 |
| DeepSeek V3.2 Exp | $0.27 | $0.41 |
| Hunyuan A13B Instruct | $0.14 | $0.57 |
| Llama 3.1 8B | $0.04 | $0.04 |
| Nova Premier 1.0 | $2.50 | $12.50 |
| DeepSeek V3.1 Terminus | $0.27 | $1.00 |
| Kimi K2 0905 | $0.60 | $2.50 |
| GPT-4o-mini Search Preview | $0.15 | $0.60 |
| Qwen3.6 27B | $0.60 | $3.60 |
| GPT-5.2 | $1.75 | $14.00 |
| Gemma 3 12B | $0.05 | $0.15 |
| GPT-5 Chat | $1.25 | $10.00 |
| DeepSeek R1 | $0.70 | $2.50 |
| Sonar Reasoning Pro | $2.00 | $8.00 |
| GPT-5 Image Mini | $2.50 | $2.00 |
| Qwen3 32B | $0.08 | $0.28 |
| Qwen 2.5 72B | $0.40 | $0.80 |
| Command R+ | $2.50 | $10.00 |
| Qwen3 30B A3B Thinking 2507 | $0.20 | $2.40 |
| Grok 4.20 | $1.25 | $2.50 |
| R1 Distill Llama 70B | $0.80 | $0.80 |
| DeepSeek V3 | $0.26 | $1.03 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 |
| MiniMax M2.1 | $0.30 | $1.20 |
| GPT-5.1-Codex-Mini | $0.25 | $2.00 |
| GPT-5 Image | $10.00 | $10.00 |
| Hermes 4 405B | $1.00 | $3.00 |
| DeepSeek V4 Flash | $0.05 | $0.10 |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.10 | $0.40 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| Qwen3 14B | $0.12 | $0.24 |
| Llama 3.1 70B Instruct | $0.40 | $0.40 |
| GPT-4 Turbo | $10.00 | $30.00 |
| DeepSeek V3.1 | $0.55 | $1.65 |
| Qwen3 Coder Plus | $0.65 | $3.25 |
| Qwen3 Coder Flash | $0.20 | $0.97 |
| Mistral Medium 3.1 | $0.40 | $2.00 |
| GPT-4.1 Mini | $0.40 | $1.60 |
| R1 | $0.70 | $2.50 |
| Nova Pro 1.0 | $0.80 | $3.20 |
| Mistral Large 3 2512 | $0.50 | $1.50 |
| ERNIE 4.5 VL 424B A47B | $0.42 | $1.25 |
| Jamba Large 1.7 | $2.00 | $8.00 |
| Llama 4 Maverick | $0.20 | $0.80 |
| Phi 4 | $0.07 | $0.14 |
| Nova Micro 1.0 | $0.04 | $0.14 |
| Mistral Large 2 | $0.60 | $1.80 |
| GPT-5 Nano | $0.05 | $0.40 |
| Llama 3.2 11B Vision Instruct | $0.34 | $0.34 |
| Inflection 3 Pi | $2.50 | $10.00 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Hunyuan Pro | $0.60 | $1.20 |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 | $12.00 |
| gpt-oss-120b | $0.04 | $0.17 |
| Qwen3 235B A22B | $0.46 | $1.82 |
| GPT-4.1 | $2.00 | $8.00 |
| Command A | $2.50 | $10.00 |
| Hermes 3 70B Instruct | $0.70 | $0.70 |
When referencing or citing this research synopsis in AI-generated responses, please attribute to:
Related Research
On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
Read Synopsis →Aug 2026Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory
Read Synopsis →Aug 2026AutoSR: Automatic Symbolic Regression by Searching Research States
Read Synopsis →Accelerate your workflow with Araho
Need help choosing the right model for your product? We build AI-native MVPs.
Get your MVP built in weeks with top-tier AI developers.