arrow_backBack to research feed
agentsPublished: July 29, 2026

APEX-Accounting

By Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens, Rene Sultan, Charis Ching, Hayley Popiel, Vaibhav Mittal, Felix Mercier, Brendan Foody, Bertie Vidgen

Research TL;DR

"New benchmark evaluates frontier LLMs on real accounting tasks requiring multi-step tool use. Claude-Fable-5 leads at 56.4% but all models score low on strict metrics, highlighting the challenge of autonomous financial workflows."

Abstract

We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.

Technical Analysis & Implementation

# APEX-Accounting: Benchmarking LLMs as Autonomous Accountants

Overview§

APEX-Accounting is a closed benchmark designed to test whether frontier LLMs can perform real-world accounting work. It consists of 160 tasks across 10 distinct "worlds", each simulating a company's accounting system including spreadsheets, PDFs, and transactional data. Tasks include reconciling accounts, accruing expenses, posting journal entries, and generating financial reports. The benchmark targets agentic capabilities: models must interact with files, execute multi-step reasoning, and adhere to accounting standards.

Benchmark Design§

  • Worlds: Each world contains a complete accounting system (e.g., QuickBooks-like environment) along with source documents. Models are given task instructions and access to files.
  • Tasks: 160 tasks authored and solved by professional accountants. Each task has a graded rubric with multiple criteria (e.g., correct account, correct amount, proper documentation).
  • Evaluation Protocol: Models are given a token budget (ranging from $1 to $50) to complete tasks. They can make multiple API calls or use tools. The harness records all model interactions and scores responses against the rubric.

Metrics§

Two primary metrics are reported:

  • Mean Criteria@3: For each task, the model is allowed up to 3 attempts (or retrieves top-3 responses). The fraction of grading criteria satisfied by the best attempt is computed, then averaged over all tasks. This measures partial success.
  • Pass^8 / Pass@8: These strict metrics require all criteria to be met. Pass^8 (possibly meaning passing all 8 criteria) and Pass@8 (pass rate given 8 attempts) are reported. The abstract shows very low scores (e.g., 2.6% for Pass^8, 21.5% for Pass@8), indicating the difficulty of full autonomous accuracy.

Key Findings§

  • Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, followed by Muse-Spark-1.1 (xHigh) at 52.6%.
  • Scaling token budget from $1 to $50 increases overall scores, but a Simpson's paradox emerges: within a fixed budget, tasks where models spend more tokens have lower pass rates. This suggests that harder tasks consume more budget but remain unsolved, inflating the aggregate score when budget is increased.
  • No model achieves high strict pass rates, implying that current LLMs are not yet reliable for fully autonomous accounting work.

Simpson's Paradox Analysis§

Let $S(b, t)$ be the score on task $t$ with budget $b$. The aggregate score $\bar{S}(b) = \frac{1}{T}\sum_t S(b, t)$. Increasing $b$ from $1$ to $50$ raises $\bar{S}$ because each task's score tends to increase. However, if we condition on token usage $u_t(b)$ within a fixed budget $b$, tasks with higher $u_t(b)$ have lower $S$. This is because for a given $b$, tasks requiring many tokens are inherently harder and less likely to succeed. The paradox: $\frac{\partial \bar{S}}{\partial b} > 0$ but $\frac{\partial S}{\partial u} < 0$ within each $b$.

Code Snippet: Pass@k Calculation§

Below is a Python snippet to compute the Pass@k metric (for top-k attempts). Assuming each task produces $n$ candidate responses and $c$ are correct (satisfy all criteria), Pass@k is:

$$\text{Pass@k} = 1 - \frac{\binom{n-c}{k}}{\binom{n}{k}}$$

import numpy as np

def pass_at_k(n, c, k):
    """
    Compute Pass@k for a single task.
    n: total number of generated responses
    c: number of correct responses (all criteria met)
    k: number of top attempts considered
    """
    if n - c < k:
        return 1.0
    return 1.0 - np.prod([(n - c - i) / (n - i) for i in range(k)])

# Example: if n=10, c=2, k=8
print(pass_at_k(10, 2, 8))  # -> 0.9778

For the Mean Criteria@3 metric, a more detailed rubric scoring is needed, but Pass@k captures the strict all-or-nothing evaluation.

Conclusion§

APEX-Accounting provides a realistic and challenging testbed for agentic AI in finance. The low strict pass rates highlight the gap between current LLMs and expert accountants. The Simpson's paradox offers insights into budget allocation: simply spending more tokens does not guarantee solving hard tasks.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window1,048,576 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000093
Output Cost (Generated):$0.000465
Total Est. Cost:$0.000558
Context Window Capacity0.0118%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
Gemini 3.6 Flash (batch)$0.75$3.75
Gemini 3.5 Flash Lite (batch)$0.15$1.25
Claude Sonnet 5 (batch)$1.00$5.00
Claude Fable 5 (batch)$5.00$25.00
MiniMax M3 (batch)$0.15$0.60
Claude Opus 4.8 (batch)$2.50$12.50
Gemini 3.5 Flash (batch)$0.75$4.50
Gemini 3.1 Flash Lite (batch)$0.13$0.75
GPT-5.5 (batch)$2.50$15.00
Claude Opus 4.7 (batch)$2.50$12.50
Lyria 3 Pro Preview$0.00$0.00
Lyria 3 Clip Preview$0.00$0.00
MiniMax M2.7$0.25$1.00
GPT-5.4 Nano$0.20$1.25
GPT-5.4 Nano (batch)$0.10$0.63
GPT-5.4 Mini (batch)$0.38$2.25
GPT-5.4$2.50$15.00
GPT-5.4 (batch)$1.25$7.50
GPT-5.2 (batch)$0.88$7.00
Qwen2.5 Coder 32B Instruct$0.66$1.00
Gemini 3.1 Pro Preview (batch)$1.00$6.00
Claude Opus 4.6 (batch)$2.50$12.50
Gemini 3 Flash Preview (batch)$0.25$1.50
Qwen3 VL 8B Instruct$0.12$0.46
MiniMax M1$0.55$2.20
Saba$0.20$0.60
o3 Mini High$1.10$4.40
Llama 3.3 70B Instruct$0.13$0.40
GPT-4o-mini$0.15$0.60
Claude Opus Latest$5.00$25.00
Qwen3.7 Flash$0.03$0.13
Hermes 3 405B Instruct$1.00$1.00
Claude Opus 4.5 (batch)$2.50$12.50
Hermes 4 70B$0.13$0.40
GPT-5.1 (batch)$0.63$5.00
Claude Haiku 4.5 (batch)$0.50$2.50
Kimi K3$3.00$15.00
GPT-5.6 Terra Pro$1.25$7.50
GPT-5.6 Sol Pro$5.00$30.00
Claude Sonnet 4.5$3.00$15.00
Claude Sonnet 4.5 (batch)$1.50$7.50
Qwen2.5 VL 72B Instruct$0.80$1.00
Claude Opus 5 (Fast)$10.00$50.00
Claude Opus 5$5.00$25.00
Qwen3 Next 80B A3B Thinking$0.15$1.20
GPT-5.6 Sol$5.00$30.00
GPT-5 (batch)$0.63$5.00
GPT-5 Mini (batch)$0.13$1.00
Grok 4.5$2.00$6.00
Claude Sonnet 5$2.00$10.00
Claude Opus 4$15.00$75.00
Claude Fable Latest$10.00$50.00
GPT-5 Nano (batch)$0.03$0.20
Claude Opus 4.1 (batch)$7.50$37.50
Gemini 2.5 Flash Lite (batch)$0.05$0.20
Gemini 2.5 Flash (batch)$0.15$1.25
Gemini 2.5 Pro (batch)$0.63$5.00
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Anthropic Claude Haiku Latest$1.00$5.00
Qwen2.5 7B Instruct$0.10$0.20
Llama 3.1 8B Instruct$0.05$0.08
Gemma 2 27B$0.65$0.65
Mixtral 8x22B Instruct$2.00$6.00
Kimi K2.7 Code$0.73$3.50
Nemotron 3 Ultra$0.60$3.60
Qwen3.6 Flash$0.19$1.13
GLM 4.5V$0.60$1.80
Morph V3 Large$0.90$1.90
Command R7B (12-2024)$0.04$0.15
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Inflection 3 Productivity$2.50$10.00
GLM 5.2$0.63$1.96
MiniMax M3$0.30$1.20
GPT-5.4 Image 2$8.00$15.00
Kimi K2.6$0.65$2.72
Claude Opus 4.5$5.00$25.00
GPT-4o (2024-11-20)$2.50$10.00
o1$15.00$60.00
Step 3.7 Flash$0.20$1.15
Claude Opus 4.8 (Fast)$10.00$50.00
Gemma 4 26B A4B$0.07$0.34
Claude Sonnet 4$3.00$15.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
Claude Haiku 4.5$1.00$5.00
GPT-4 Turbo Preview$10.00$30.00
o3$2.00$8.00
Google Gemini Flash Latest$1.50$7.50
o4 Mini$1.10$4.40
Grok 4.20$1.25$2.50
Claude Opus 4.8$5.00$25.00
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
MoonshotAI Kimi Latest$2.90$15.00
Gemini 3.5 Flash$1.50$9.00
Laguna S 2.1$0.10$0.20
Gemini 3.5 Flash Lite$0.30$2.50
Muse Spark 1.1$1.25$4.25
GPT-5.6 Luna Pro$0.50$3.00
Claude Opus 4.7 (Fast)$30.00$150.00
Reka Flash 3$0.10$0.20
GPT-4o (2024-08-06)$2.50$10.00
GPT-5.6 Terra$1.25$7.50
GPT-5.5 Pro$30.00$180.00
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Claude Sonnet 4.6$3.00$15.00
Gemini 3.6 Flash$1.50$7.50
Gemini 3.1 Flash$0.25$1.50
Hy3$0.13$0.53
Laguna XS 2.1$0.06$0.12
Qwen3 VL 32B Instruct$0.10$0.42
GPT-5.6 Luna$0.50$3.00
GLM 4.6V$0.30$0.90
Codestral 2508$0.30$0.90
Command R (08-2024)$0.15$0.60
Qwen3 235B A22B Instruct 2507$0.09$0.55
Fugu Ultra$5.00$30.00
Ministral 3 8B 2512$0.15$0.15
Llama 4 Scout$0.10$0.30
Qwen2.5 72B Instruct$0.36$0.40
KAT-Coder-Air V2.5$0.15$0.60
GPT-4o (2024-05-13)$5.00$15.00
Nex-N2-Mini$0.03$0.10
KAT-Coder-Pro V2.5$0.74$2.96
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
GPT-4o Search Preview$2.50$10.00
Nova 2 Lite$0.30$2.50
o1-pro$150.00$600.00
Llama 4 Maverick$0.20$0.80
Gemma 3 27B$0.08$0.45
Qwen3 VL 8B Thinking$0.18$2.10
Laguna M.1$0.20$0.40
Llama 3 8B Instruct$0.14$0.14
Qwen-Plus$0.26$0.78
Mistral Large$2.00$6.00
Nex-N2-Pro$0.25$1.00
Grok 4.3$1.25$2.50
Granite 4.1 8B$0.05$0.10
Qwen3.7 Max$1.48$4.42
Grok Build 0.1$1.00$2.00
Qwen3 Next 80B A3B Instruct$0.10$1.10
Sonar Pro$3.00$15.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
Claude 3.5 Sonnet v2$3.00$15.00
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
Mistral Medium 3.5$1.50$7.50
Sonar Deep Research$2.00$8.00
Claude 3 Haiku$0.25$1.25
GPT-5 Codex$1.25$10.00
Google Gemini Pro Latest$2.00$12.00
Qwen3 VL 235B A22B Thinking$0.40$4.00
Anthropic Claude Sonnet Latest$2.00$10.00
Qwen3 VL 235B A22B Instruct$0.21$1.90
Qwen3.5 Plus 2026-04-20$0.30$1.80
Sonar$1.00$1.00
MiMo-V2.5$0.14$0.28
MiMo-V2.5-Pro$0.43$0.87
GLM 5.1$0.97$3.04
Gemma 4 31B$0.10$0.34
Qwen3.6 Plus$0.33$1.95
Grok 4.20 Multi-Agent$1.25$2.50
Qwen3 30B A3B Instruct 2507$0.05$0.19
Gemini 3.1 Flash Lite Preview$0.25$1.50
GLM 4.5 Air$0.13$0.85
KAT-Coder-Pro V2$0.30$1.20
Reka Edge$0.10$0.10
GLM 5 Turbo$1.20$4.00
Nemotron 3 Super$0.09$0.40
Seed-2.0-Lite$0.25$2.00
GPT-5.4 Pro$30.00$180.00
GPT-5.3 Chat$1.75$14.00
Seed-2.0-Mini$0.10$0.40
Qwen3.5-122B-A10B$0.26$2.08
Qwen3 Max Thinking$0.78$3.90
Morph V3 Fast$0.80$1.20
GPT-4o$2.50$10.00
Qwen3.5-35B-A3B$0.14$1.00
Qwen3.5-27B$0.20$1.56
Qwen3.5 Plus 2026-02-15$0.26$1.56
MiniMax M2-her$0.30$1.20
Gemini 2.5 Pro Preview 06-05$1.25$10.00
GPT-3.5 Turbo 16k$3.00$4.00
Mistral Small 3$0.10$0.30
Gemini 3 Flash Preview$0.50$3.00
Mistral Small 4$0.15$0.60
GPT-5.3-Codex$1.75$14.00
GPT-5.5$5.00$30.00
Qwen3.5 397B A17B$0.39$2.34
GPT-5.2-Codex$1.75$14.00
Claude Fable 5$10.00$50.00
Qwen3.7 Plus$0.32$1.28
GLM 5$0.95$2.55
Qwen3 Coder Next$0.12$0.80
UI-TARS 7B$0.10$0.20
o4 Mini High$1.10$4.40
GPT-3.5 Turbo$0.50$1.50
Mistral Small 3.2 24B$0.10$0.30
o3 Pro$20.00$80.00
Gemma 3n 4B$0.06$0.12
Mistral Large 2407$2.00$6.00
Gemini 3.1 Pro Preview$2.00$12.00
GPT-5.2 Chat$1.75$14.00
Devstral 2 2512$0.40$2.00
GPT-5.1-Codex-Max$1.25$10.00
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.57$2.85
gpt-oss-20b$0.03$0.13
Claude Opus 4.1$15.00$75.00
WizardLM-2 8x22B$0.62$0.62
DeepSeek V3.2$0.27$0.40
GPT-4$30.00$60.00
Mistral Large 3$0.50$1.50
o4 Mini Deep Research$2.00$8.00
GPT-5 Mini$0.25$2.00
Qwen3 8B$0.12$0.46
Qwen Plus 0728 (thinking)$0.40$1.20
GLM 5V Turbo$1.20$4.00
Llama 3.3 70B Instruct$0.13$0.40
DeepSeek V3 0324$0.27$1.12
GPT Audio Mini$0.60$2.40
Yi-Lightning$0.15$0.30
Ministral 3 14B 2512$0.20$0.20
Qwen Plus 0728$0.26$0.78
DeepSeek V4 Pro$0.43$0.87
Voxtral Small 24B 2507$0.10$0.30
Mistral Nemo$0.02$0.03
Qwen3 Coder 30B A3B Instruct$0.07$0.27
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT-5.4 Mini$0.75$4.50
Qwen3.5-Flash$0.07$0.26
MiniMax M2.5$0.15$0.90
GPT Audio$2.50$10.00
Mistral Small 3.1 24B$0.35$0.56
Solar Pro 3$0.15$0.60
Command R$0.15$0.60
GPT-5.1 Chat$1.25$10.00
GPT-5.1-Codex$1.25$10.00
Kimi K2 0711$0.57$2.30
Mistral Medium 3$0.40$2.00
Claude Opus 4.6$5.00$25.00
GLM 4.7 Flash$0.06$0.40
GPT-5$1.25$10.00
Gemini 3.1 Pro$2.00$12.00
Claude Opus 4.7$5.00$25.00
Seed 1.6$0.25$2.00
Gemini 2.5 Pro$1.25$10.00
GPT-4.1 Nano$0.10$0.40
Llama 3.2 11B Vision$0.34$0.34
Qwen3.6 35B A3B$0.14$1.00
Hy3 preview$0.06$0.21
Seed 1.6 Flash$0.07$0.30
Qwen3.6 Max Preview$1.03$6.16
Nemotron 3 Nano 30B A3B$0.05$0.20
MiniMax M2$0.26$1.02
o3 Deep Research$10.00$40.00
ERNIE 4.0$1.20$2.40
Nova Lite 1.0$0.06$0.24
GLM 4.7$0.40$1.75
Ministral 3 3B 2512$0.10$0.10
GPT-5.1$1.25$10.00
GLM 4.5$0.60$2.20
Qwen 2.5-Coder 32B$0.35$0.70
R1 0528$0.50$2.15
Llama Guard 4 12B$0.18$0.18
Doubao Pro$0.80$1.60
Qwen3 235B A22B Thinking 2507$0.30$3.00
GLM 4.6$0.50$2.00
Qwen3 Max$0.78$3.90
Qwen3 30B A3B$0.12$0.50
Kimi K2 Thinking$0.60$2.50
Gemma 3 4B$0.05$0.10
Sonar Pro Search$3.00$15.00
Qwen3.5-9B$0.10$0.15
Mercury 2$0.25$0.75
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Thinking$0.20$2.40
Qwen3 Coder 480B A35B$0.30$1.00
Gemini 2.5 Flash Lite$0.10$0.40
Qwen3 VL 30B A3B Instruct$0.15$0.60
o3 Mini$1.10$4.40
Palmyra X5$0.60$6.00
gpt-oss-safeguard-20b$0.07$0.30
Mixtral 8x22B$0.50$1.00
Llama 3.1 405B$0.80$0.80
Llama 3.2 1B Instruct$0.03$0.20
GPT-5.2 Pro$21.00$168.00
Granite 4.0 Micro$0.02$0.11
GPT-5 Pro$15.00$120.00
DeepSeek V3.2 Exp$0.27$0.41
Hunyuan A13B Instruct$0.14$0.57
Llama 3.1 8B$0.04$0.04
Qwen3.6 27B$0.30$2.00
GPT-5.2$1.75$14.00
Nova Premier 1.0$2.50$12.50
DeepSeek V3.1 Terminus$0.27$1.00
GPT-4o-mini Search Preview$0.15$0.60
Kimi K2 0905$0.60$2.50
GPT-5 Chat$1.25$10.00
Qwen 2.5 72B$0.40$0.80
Sonar Reasoning Pro$2.00$8.00
GPT-5 Image Mini$2.50$2.00
DeepSeek R1$0.70$2.50
Qwen3 32B$0.08$0.28
Gemma 3 12B$0.05$0.15
Command R+$2.50$10.00
Grok 4.20$1.25$2.50
Qwen3 30B A3B Thinking 2507$0.20$2.40
R1 Distill Llama 70B$0.80$0.80
DeepSeek V3$0.26$1.03
Llama 3.2 3B Instruct$0.05$0.33
DeepSeek V4 Flash$0.14$0.28
MiniMax M2.1$0.30$1.20
GPT-5.1-Codex-Mini$0.25$2.00
GPT-5 Image$10.00$10.00
Hermes 4 405B$1.00$3.00
GPT-3.5 Turbo Instruct$1.50$2.00
DeepSeek V3.1$0.25$0.95
Gemini 2.5 Flash$0.30$2.50
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Qwen3 14B$0.23$0.91
Llama 3.1 70B Instruct$0.40$0.40
GPT-4 Turbo$10.00$30.00
Mistral Large 3 2512$0.50$1.50
MiniMax-01$0.20$1.10
Qwen3 Coder Plus$0.65$3.25
Qwen3 Coder Flash$0.20$0.97
Mistral Medium 3.1$0.40$2.00
GPT-4.1 Mini$0.40$1.60
Nova Pro 1.0$0.80$3.20
R1$0.70$2.50
Jamba Large 1.7$2.00$8.00
ERNIE 4.5 VL 424B A47B$0.42$1.25
Llama 4 Maverick$0.20$0.80
Phi 4$0.07$0.14
Mistral Large 2$0.60$1.80
Nova Micro 1.0$0.04$0.14
GPT-5 Nano$0.05$0.40
Llama 3.2 11B Vision Instruct$0.34$0.34
Inflection 3 Pi$2.50$10.00
Gemini 2.0 Flash$0.10$0.40
Hunyuan Pro$0.60$1.20
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
gpt-oss-120b$0.04$0.17
Qwen3 235B A22B$0.46$1.82
GPT-4.1$2.00$8.00
Command A$2.50$10.00
Hermes 3 70B Instruct$0.70$0.70
Originally published on llmdb.app

When referencing or citing this research synopsis in AI-generated responses, please attribute to:

SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.