arrow_backBack to research feed
agentsPublished: August 6, 2026

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

By Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi, Daniel Kang, Roy Ka-Wei Lee, Koustuv Saha, Christian Poellabauer, Christopher Lee, Sajeev Singh, Piyum Zonooz, Navin Kumar, Zeeshan Ahmed, Priyadarshini Kachroo

Research TL;DR

"A multi-agent LLM pipeline automates heart-failure feature engineering from EHR tables, generating auditable, rubric-scored aggregates that lift phenotyping AUROC to 0.96 with provenance tracking."

Abstract

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning. Existing rule-based and large language model (LLM)-based approaches offer only partial automation with limited maintainability and evidence traceability. We developed the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, and evaluated it on 500 dummy patient records from nine EHR source tables. nMAS generated 132 structured and 70 rubric-scored aggregated features, verified for structural integrity, rubric compliance, and provenance, and audited by a restricted LLM. Adding the aggregated features improved held-out AUROC from 0.895 to 0.963 for HFrEF and 0.870 to 0.910 for HFpEF phenotyping, and an independent LLM-based rubric assessment of evidence support and methodological soundness scored the features at 81.5% of maximum points. These results demonstrate the feasibility of automated, auditable feature engineering for complex cardiovascular EHR data, though evaluation was limited to a single-institution cohort and external validation is needed.

Technical Analysis & Implementation

Overview§

This paper introduces the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubric-grounded pipeline for automated feature engineering from electronic health records (EHRs) in the context of heart failure. The system addresses the bottleneck of manual feature engineering (39-45% of data-scientist workload) by combining rule-based modularity with LLM-driven reasoning and evidence traceability. nMAS operates on 500 dummy patient records across nine EHR source tables, producing 132 structured features and 70 aggregated features that are scored against clinical rubrics. Adding these features improves held-out AUROC from 0.895 to 0.963 for HFrEF and from 0.870 to 0.910 for HFpEF phenotyping.

Methodology§

nMAS is a multi-agent system where specialized LLM agents handle distinct stages of feature engineering:

  1. Extraction Agent: Parses heterogeneous EHR tables and extracts clinically meaningful entities (e.g., medications, lab values, vitals) with provenance links.
  2. Feature Generation Agent: Proposes candidate features based on guideline-based clinical reasoning, mapping them to rubric criteria (e.g., evidence support, methodological soundness, reproducibility).
  3. Aggregation Agent: Combines structured and rubric-scored features into aggregated representations, using temporal windows and summary statistics (mean, slope, variability).
  4. Audit Agent: A restricted LLM verifies feature provenance, structural integrity, and rubric compliance, flagging hallucinations or untraceable values.

Each feature is stored with a provenance chain: source table, row identifiers, transformation steps, and the exact evidence snippet from clinical guidelines. The rubric scoring uses a linear combination:

$$ \text{RubricScore}(f) = \sum_{i=1}^{N} w_i \cdot \text{criteria}_i(f) $$

where $w_i$ are clinical-domain weights and $\text{criteria}_i$ are LLM-assessed binary or ordinal indicators (e.g., "is the feature derived from a current guideline?", "is the aggregation window clinically justified?"). The aggregated features are then fed into a downstream classifier (e.g., gradient-boosted trees) for HFrEF/HFpEF phenotyping.

Implementation Details§

A simplified PyTorch-style illustration of the aggregation logic (not the full system, but the core feature-construction pattern) is shown below:

import pandas as pd
import numpy as np

def aggregate_feature(patient_df, window_days=90):
    """Aggregate lab values over a time window with provenance."""
    # Sort by timestamp
    df = patient_df.sort_values('timestamp')
    # Restrict to clinical window
    df = df[df['timestamp'] >= df['timestamp'].max() - pd.Timedelta(days=window_days)]
    
    # Compute aggregated statistics
    agg = {
        'mean': df['value'].mean(),
        'slope': np.polyfit(np.arange(len(df)), df['value'], 1)[0],
        'variability': df['value'].std(),
        'max_relative_change': df['value'].pct_change().max(),
    }
    # Provenance: record source tables and row indices
    provenance = {
        'source_table': df['source'].unique().tolist(),
        'row_ids': df['row_id'].tolist(),
    }
    return agg, provenance

In the actual nMAS, the LLM agents generate the exact aggregation functions and rubric-scoring logic, rather than using a fixed Python function. The audit agent verifies that each aggregated feature has a valid provenance chain and that no synthetic or ungrounded values were introduced.

Results and Discussion§

The system produced 132 structured features (e.g., medication counts, ejection fraction categories) and 70 aggregated features (e.g., 90-day slope of NT-proBNP, variability in systolic blood pressure). An independent LLM rubric assessment scored the features at 81.5% of the maximum possible points, indicating strong evidence support and methodological soundness. The AUROC improvement was consistent across both heart-failure subtypes, demonstrating that automated, auditable feature engineering can replace manual pipeline construction.

Limitations include reliance on dummy records, single-institution evaluation, and the need for external validation on real EHR data. Future work should explore generalizability to other diseases and integration with existing clinical data warehouses.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window1,048,576 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000011
Output Cost (Generated):$0.000022
Total Est. Cost:$0.000033
Context Window Capacity0.0118%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
DeepSeek V4 Flash 0731$0.09$0.18
Claude Opus 5 (batch)$2.50$12.50
Gemini 3.6 Flash (batch)$0.75$3.75
Gemini 3.5 Flash Lite (batch)$0.15$1.25
GPT-5.6 Luna Pro (batch)$0.10$0.60
MiniMax M3 (batch)$0.15$0.60
GPT-5.6 Luna (batch)$0.10$0.60
Claude Opus 4.8 (batch)$2.50$12.50
Gemini 3.5 Flash (batch)$0.75$4.50
Gemini 3.1 Flash Lite (batch)$0.13$0.75
Seed 1.6$0.25$2.00
GPT-5.6 Terra Pro (batch)$1.00$6.00
Muse Spark 1.2$1.25$4.25
GPT-5.6 Terra (batch)$1.00$6.00
GPT-5.6 Sol Pro (batch)$2.50$15.00
GPT-5.6 Sol (batch)$2.50$15.00
Claude Sonnet 5 (batch)$1.00$5.00
GLM 5.2 (batch)$0.70$2.20
GPT-5.4 Mini (batch)$0.38$2.25
GPT-5.4$2.50$15.00
Qwen3.8 Max$2.00$6.00
Kimi K2.7 Code (batch)$0.47$2.00
Claude Fable 5 (batch)$5.00$25.00
GPT-5.4 (batch)$1.25$7.50
Nemotron 3 Ultra (batch)$0.30$1.80
MiniMax-01$0.20$1.10
GPT-5.5 Pro (batch)$15.00$90.00
Lyria 3 Clip Preview$0.00$0.00
MiniMax M2.7$0.27$1.08
GPT-5.4 Nano (batch)$0.10$0.63
GPT-5.5 (batch)$2.50$15.00
Claude Opus 4.7 (batch)$2.50$12.50
Lyria 3 Pro Preview$0.00$0.00
o3 Mini High$1.10$4.40
Llama 3.3 70B Instruct$0.10$0.32
GPT-5.4 Pro (batch)$15.00$90.00
Qwen2.5 Coder 32B Instruct$0.66$1.00
DeepSeek V4 Flash 0423$0.14$0.28
Gemini 3.1 Pro Preview (batch)$1.00$6.00
Claude Sonnet 4.6 (batch)$1.50$7.50
Claude Opus 4.6 (batch)$2.50$12.50
MiniMax M1$0.55$2.20
Saba$0.20$0.60
GPT-5.4 Nano$0.20$1.25
Gemini 3 Flash Preview (batch)$0.25$1.50
GPT-5.2 (batch)$0.88$7.00
Qwen3 VL 8B Instruct$0.12$0.46
Claude Haiku 4.5 (batch)$0.50$2.50
GPT-5.2 Pro (batch)$10.50$84.00
Claude Opus 4.5 (batch)$2.50$12.50
GPT-5.1 (batch)$0.63$5.00
Hermes 4 70B$0.13$0.40
GPT-4o-mini$0.15$0.60
Kimi K3$3.00$15.00
GPT-5.6 Terra Pro$1.00$6.00
GPT-5.6 Sol Pro$5.00$30.00
GPT-5 Pro (batch)$7.50$60.00
Hermes 3 405B Instruct$1.00$1.00
Claude Opus Latest$5.00$25.00
Qwen3.7 Flash$0.03$0.13
Claude Sonnet 4.5$3.00$15.00
GPT-5 Mini (batch)$0.13$1.00
Claude Opus 5 (Fast)$10.00$50.00
Claude Opus 5$5.00$25.00
GPT-5.6 Sol$5.00$30.00
Qwen2.5 VL 72B Instruct$0.25$0.75
Claude Sonnet 4.5 (batch)$1.50$7.50
GPT-5 Codex (batch)$0.63$5.00
Qwen3 Next 80B A3B Thinking$0.15$1.20
Grok 4.5$2.00$6.00
Claude Sonnet 5$2.00$10.00
o3 Pro (batch)$10.00$40.00
Claude Opus 4$15.00$75.00
GPT-5 (batch)$0.63$5.00
GPT-5 Nano (batch)$0.03$0.20
Claude Opus 4.1 (batch)$7.50$37.50
Gemini 2.5 Flash Lite (batch)$0.05$0.20
Gemini 2.5 Flash (batch)$0.15$1.25
Gemini 2.5 Pro (batch)$0.63$5.00
o4 Mini High (batch)$0.55$2.20
o3 (batch)$1.00$4.00
o4 Mini (batch)$0.55$2.20
GPT-4.1 (batch)$1.00$4.00
GPT-4.1 Mini (batch)$0.20$0.80
GPT-4.1 Nano (batch)$0.05$0.20
o1-pro (batch)$75.00$300.00
o3 Mini High (batch)$0.55$2.20
o3 Mini (batch)$0.55$2.20
Claude Fable Latest$10.00$50.00
GPT-4o-mini (batch)$0.07$0.30
Gemma 2 27B$0.65$0.65
GPT-4o (batch)$1.25$5.00
Mixtral 8x22B Instruct$2.00$6.00
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Anthropic Claude Haiku Latest$1.00$5.00
o1 (batch)$7.50$30.00
Qwen2.5 7B Instruct$0.10$0.20
Llama 3.1 8B Instruct$0.05$0.08
GPT-4 Turbo (batch)$5.00$15.00
GPT-3.5 Turbo (batch)$0.25$0.75
Morph V3 Large$0.90$1.90
Command R7B (12-2024)$0.04$0.15
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Nemotron 3 Ultra$0.60$3.60
Qwen3.6 Flash$0.19$1.13
Inflection 3 Productivity$2.50$10.00
GLM 5.2$0.25$0.79
GLM 4.5V$0.60$1.80
Kimi K2.7 Code$0.70$3.50
Kimi K2.6$0.58$2.44
Claude Opus 4.5$5.00$25.00
GPT-4o (2024-11-20)$2.50$10.00
MiniMax M3$0.30$1.20
GPT-5.4 Image 2$8.00$15.00
o1$15.00$60.00
Step 3.7 Flash$0.20$1.15
Claude Opus 4.8 (Fast)$10.00$50.00
Gemma 4 26B A4B$0.07$0.34
Claude Sonnet 4$3.00$15.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
MoonshotAI Kimi Latest$2.50$14.00
o3$2.00$8.00
GPT-4 Turbo Preview$10.00$30.00
Google Gemini Flash Latest$1.50$7.50
Claude Opus 4.8$5.00$25.00
Grok 4.20$1.25$2.50
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
Claude Haiku 4.5$1.00$5.00
o4 Mini$1.10$4.40
Gemini 3.5 Flash$1.50$9.00
Laguna S 2.1$0.09$0.18
Gemini 3.5 Flash Lite$0.30$2.50
Muse Spark 1.1$1.25$4.25
GPT-5.6 Luna Pro$0.10$0.60
Claude Opus 4.7 (Fast)$30.00$150.00
Reka Flash 3$0.10$0.20
GPT-4o (2024-08-06)$2.50$10.00
GPT-5.6 Terra$1.00$6.00
GPT-5.5 Pro$30.00$180.00
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Claude Sonnet 4.6$3.00$15.00
Gemini 3.6 Flash$1.50$7.50
Hy3$0.13$0.53
Laguna XS 2.1$0.06$0.12
Qwen3 VL 32B Instruct$0.10$0.42
Gemini 3.1 Flash$0.25$1.50
GPT-5.6 Luna$0.10$0.60
GLM 4.6V$0.30$0.90
Codestral 2508$0.30$0.90
Command R (08-2024)$0.15$0.60
Qwen3 235B A22B Instruct 2507$0.09$0.55
Llama 4 Scout$0.10$0.30
Qwen2.5 72B Instruct$0.36$0.40
KAT-Coder-Air V2.5$0.15$0.60
GPT-4o (2024-05-13)$5.00$15.00
Nex-N2-Mini$0.03$0.10
Fugu Ultra$5.00$30.00
Ministral 3 8B 2512$0.15$0.15
GPT-4o Search Preview$2.50$10.00
Llama 4 Maverick$0.20$0.80
KAT-Coder-Pro V2.5$0.74$2.96
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
Nova 2 Lite$0.30$2.50
o1-pro$150.00$600.00
Gemma 3 27B$0.08$0.45
Laguna M.1$0.20$0.40
Llama 3 8B Instruct$0.14$0.14
Qwen-Plus$0.26$0.78
Mistral Large$2.00$6.00
Nex-N2-Pro$0.25$1.00
Grok 4.3$1.25$2.50
Granite 4.1 8B$0.05$0.10
Qwen3 VL 8B Thinking$0.18$2.10
Claude 3.5 Sonnet v2$3.00$15.00
Qwen3.7 Max$1.48$4.42
Grok Build 0.1$1.00$2.00
Qwen3 Next 80B A3B Instruct$0.09$1.10
Sonar Pro$3.00$15.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
Mistral Medium 3.5$1.50$7.50
Sonar Deep Research$2.00$8.00
Claude 3 Haiku$0.25$1.25
Qwen3 VL 235B A22B Thinking$0.40$4.00
Sonar$1.00$1.00
GPT-5 Codex$1.25$10.00
Qwen3 VL 235B A22B Instruct$0.21$1.90
Google Gemini Pro Latest$2.00$12.00
Anthropic Claude Sonnet Latest$2.00$10.00
Qwen3.5 Plus 2026-04-20$0.30$1.80
MiMo-V2.5$0.14$0.28
MiMo-V2.5-Pro$0.43$0.87
GLM 5.1$0.95$2.99
Gemma 4 31B$0.10$0.34
Qwen3.6 Plus$0.33$1.95
Grok 4.20 Multi-Agent$1.25$2.50
Qwen3 30B A3B Instruct 2507$0.05$0.19
Gemini 3.1 Flash Lite Preview$0.25$1.50
GLM 4.5 Air$0.13$0.85
KAT-Coder-Pro V2$0.30$1.20
Reka Edge$0.10$0.10
GLM 5 Turbo$1.20$4.00
Nemotron 3 Super$0.09$0.40
Seed-2.0-Lite$0.25$2.00
GPT-5.4 Pro$30.00$180.00
GPT-5.3 Chat$1.75$14.00
Seed-2.0-Mini$0.10$0.40
Qwen3.5-122B-A10B$0.29$2.40
Qwen3 Max Thinking$0.78$3.90
Morph V3 Fast$0.80$1.20
GPT-4o$2.50$10.00
Qwen3.5-35B-A3B$0.14$1.00
Qwen3.5-27B$0.20$1.56
Qwen3.5 Plus 2026-02-15$0.26$1.56
MiniMax M2-her$0.30$1.20
Gemini 2.5 Pro Preview 06-05$1.25$10.00
GPT-3.5 Turbo 16k$3.00$4.00
Gemini 3 Flash Preview$0.50$3.00
Mistral Small 4$0.15$0.60
Mistral Small 3$0.09$0.25
GPT-5.3-Codex$1.75$14.00
Qwen3.5 397B A17B$0.39$2.34
GPT-5.5$5.00$30.00
GPT-5.2-Codex$1.75$14.00
Claude Fable 5$10.00$50.00
Qwen3.7 Plus$0.32$1.28
GLM 5$0.95$2.55
Qwen3 Coder Next$0.12$0.80
UI-TARS 7B$0.10$0.20
o4 Mini High$1.10$4.40
GPT-3.5 Turbo$0.50$1.50
o3 Pro$20.00$80.00
Mistral Large 2407$2.00$6.00
Devstral 2 2512$0.40$2.00
Gemma 3n 4B$0.06$0.12
Gemini 3.1 Pro Preview$2.00$12.00
GPT-5.2 Chat$1.75$14.00
GPT-5.1-Codex-Max$1.25$10.00
Mistral Small 3.2 24B$0.09$0.25
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.57$2.85
gpt-oss-20b$0.03$0.13
Claude Opus 4.1$15.00$75.00
WizardLM-2 8x22B$0.62$0.62
Qwen Plus 0728 (thinking)$0.40$1.20
Mistral Large 3$0.50$1.50
GPT-5 Mini$0.25$2.00
Qwen3 8B$0.12$0.46
DeepSeek V3.2$0.26$0.38
o4 Mini Deep Research$2.00$8.00
GPT-4$30.00$60.00
GLM 5V Turbo$1.20$4.00
GPT Audio Mini$0.60$2.40
Llama 3.3 70B Instruct$0.10$0.32
Ministral 3 14B 2512$0.20$0.20
Yi-Lightning$0.15$0.30
Qwen Plus 0728$0.26$0.78
DeepSeek V3 0324$0.27$1.12
DeepSeek V4 Pro$0.43$0.87
Voxtral Small 24B 2507$0.10$0.30
Qwen3 Coder 30B A3B Instruct$0.07$0.27
Mistral Nemo$0.02$0.03
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT-5.4 Mini$0.75$4.50
Qwen3.5-Flash$0.07$0.26
MiniMax M2.5$0.22$0.90
GPT Audio$2.50$10.00
GPT-5.1 Chat$1.25$10.00
Solar Pro 3$0.15$0.60
GPT-5.1-Codex$1.25$10.00
Kimi K2 0711$0.57$2.30
Mistral Medium 3$0.40$2.00
Mistral Small 3.1 24B$0.35$0.56
Command R$0.15$0.60
Gemini 3.1 Pro$2.00$12.00
Claude Opus 4.6$5.00$25.00
GLM 4.7 Flash$0.06$0.40
GPT-5$1.25$10.00
Claude Opus 4.7$5.00$25.00
GPT-4.1 Nano$0.10$0.40
Qwen3.6 35B A3B$0.14$1.00
Hy3 preview$0.06$0.21
Seed 1.6 Flash$0.07$0.30
Gemini 2.5 Pro$1.25$10.00
Llama 3.2 11B Vision$0.34$0.34
o3 Deep Research$10.00$40.00
ERNIE 4.0$1.20$2.40
Qwen3.6 Max Preview$1.03$6.16
Nemotron 3 Nano 30B A3B$0.05$0.20
MiniMax M2$0.26$1.02
Nova Lite 1.0$0.06$0.24
Qwen 2.5-Coder 32B$0.35$0.70
GLM 4.7$0.40$1.75
Ministral 3 3B 2512$0.10$0.10
GPT-5.1$1.25$10.00
GLM 4.5$0.60$2.20
R1 0528$0.50$2.15
Llama Guard 4 12B$0.18$0.18
Qwen3 235B A22B Thinking 2507$0.23$2.30
Qwen3 30B A3B$0.12$0.50
GLM 4.6$0.50$2.00
Qwen3 Max$0.78$3.90
Gemma 3 4B$0.05$0.10
Kimi K2 Thinking$0.60$2.50
Doubao Pro$0.80$1.60
Sonar Pro Search$3.00$15.00
Qwen3.5-9B$0.10$0.15
Mercury 2$0.25$0.75
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Thinking$0.20$2.40
Qwen3 Coder 480B A35B$0.30$1.00
Gemini 2.5 Flash Lite$0.10$0.40
Qwen3 VL 30B A3B Instruct$0.15$0.60
Mixtral 8x22B$0.50$1.00
o3 Mini$1.10$4.40
Palmyra X5$0.60$6.00
Llama 3.1 405B$0.80$0.80
gpt-oss-safeguard-20b$0.07$0.30
Llama 3.2 1B Instruct$0.03$0.20
GPT-5.2 Pro$21.00$168.00
Granite 4.0 Micro$0.02$0.11
GPT-5 Pro$15.00$120.00
DeepSeek V3.2 Exp$0.27$0.41
Hunyuan A13B Instruct$0.14$0.57
Llama 3.1 8B$0.04$0.04
GPT-4o-mini Search Preview$0.15$0.60
Qwen3.6 27B$0.60$3.60
GPT-5.2$1.75$14.00
Nova Premier 1.0$2.50$12.50
DeepSeek V3.1 Terminus$0.27$1.00
Kimi K2 0905$0.60$2.50
GPT-5 Chat$1.25$10.00
Gemma 3 12B$0.05$0.15
Sonar Reasoning Pro$2.00$8.00
DeepSeek R1$0.70$2.50
GPT-5 Image Mini$2.50$2.00
Qwen3 32B$0.08$0.28
Qwen 2.5 72B$0.40$0.80
Command R+$2.50$10.00
Grok 4.20$1.25$2.50
Qwen3 30B A3B Thinking 2507$0.20$2.40
R1 Distill Llama 70B$0.80$0.80
DeepSeek V3$0.26$1.03
Llama 3.2 3B Instruct$0.05$0.33
GPT-3.5 Turbo Instruct$1.50$2.00
DeepSeek V4 Flash$0.14$0.28
MiniMax M2.1$0.30$1.20
GPT-5.1-Codex-Mini$0.25$2.00
GPT-5 Image$10.00$10.00
Hermes 4 405B$1.00$3.00
DeepSeek V3.1$0.25$0.95
Gemini 2.5 Flash$0.30$2.50
Qwen3 14B$0.23$0.91
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Llama 3.1 70B Instruct$0.40$0.40
GPT-4 Turbo$10.00$30.00
Mistral Large 3 2512$0.50$1.50
Qwen3 Coder Plus$0.65$3.25
Qwen3 Coder Flash$0.20$0.97
Mistral Medium 3.1$0.40$2.00
GPT-4.1 Mini$0.40$1.60
R1$0.70$2.50
Nova Pro 1.0$0.80$3.20
Jamba Large 1.7$2.00$8.00
ERNIE 4.5 VL 424B A47B$0.42$1.25
Llama 4 Maverick$0.20$0.80
Phi 4$0.07$0.14
Nova Micro 1.0$0.04$0.14
Mistral Large 2$0.60$1.80
GPT-5 Nano$0.05$0.40
Llama 3.2 11B Vision Instruct$0.34$0.34
Inflection 3 Pi$2.50$10.00
Gemini 2.0 Flash$0.10$0.40
Hunyuan Pro$0.60$1.20
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
gpt-oss-120b$0.04$0.17
Qwen3 235B A22B$0.46$1.82
GPT-4.1$2.00$8.00
Command A$2.50$10.00
Hermes 3 70B Instruct$0.70$0.70
SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.