BENCHMARK LEADERBOARD

Model Evaluation Leaderboards

Track ranking standing across standardized evaluations. Analyze MMLU, HumanEval, and GPQA performance to identify capabilities suitable for your product.

RankModelProviderScoreSource
1GPT-5.6 SolOpenAI96.8%Verified
2GPT-5.5 ProOpenAI96.5%Verified
3Claude Opus 4.8Anthropic95.4%Verified
4Grok 4.20xAI94.7%Verified
5DeepSeek V4 ProDeepSeek94.5%Verified
6DeepSeek V3DeepSeek94.5%Verified
7GPT-5.5OpenAI94.2%Verified
8Claude Opus 4.7Anthropic94.1%Verified
9GPT-5.1OpenAI93.2%Estimated
10GPT-5.2 (batch)OpenAI93.0%Estimated
11GPT-5.2-CodexOpenAI93.0%Estimated
12GPT-5.4 Nano (batch)OpenAI92.8%Estimated
13GPT-5.1 (batch)OpenAI92.8%Estimated
14Claude Opus 4.5Anthropic92.8%Verified
15GPT-5.1-CodexOpenAI92.8%Estimated
16Gemini 3.1 ProGoogle92.8%Verified
17GPT-5.6 Luna ProOpenAI92.6%Estimated
18Claude Opus 4.6Anthropic92.5%Verified
19o3 Mini HighOpenAI92.4%Estimated
20o3OpenAI92.4%Estimated
21Claude Opus 4.7 (Fast)Anthropic92.4%Estimated
22Grok 4.3xAI92.4%Verified
23GPT-5.4 Image 2OpenAI92.2%Estimated
24GPT-5OpenAI92.1%Verified
25Qwen3 Max ThinkingAlibaba92.0%Verified
26R1 0528DeepSeek92.0%Estimated
27o1OpenAI91.8%Verified
28GPT-5 ChatOpenAI91.8%Estimated
29Llama 4 MaverickMeta91.5%Verified
30Qwen3.7 MaxAlibaba91.5%Verified
31GPT-5 ImageOpenAI91.2%Estimated
32GPT-5.6 LunaOpenAI91.0%Estimated
33DeepSeek R1DeepSeek90.8%Verified
34GPT-5 NanoOpenAI90.8%Estimated
35o3 ProOpenAI90.6%Estimated
36Claude Opus 4Anthropic90.5%Verified
37Gemini 3.5 FlashGoogle90.5%Verified
38Gemini 2.5 ProGoogle89.9%Verified
39Claude Sonnet 4.5Anthropic89.5%Verified
40GLM 5.2Zhipu AI89.5%Verified
41GPT-4oOpenAI88.7%Verified
42Claude 3.5 SonnetAnthropic88.7%Verified
43Google Gemini Pro LatestGoogle88.6%Estimated
44Qwen3 MaxAlibaba88.6%Estimated
45Kimi K2 ThinkingMoonshot AI88.5%Verified
46Gemini 3.1 Pro PreviewGoogle88.4%Estimated
47Qwen3.7 PlusAlibaba88.2%Verified
48Claude Sonnet 4Anthropic88.0%Verified
49GLM 5.1Zhipu AI88.0%Verified
50o3 MiniOpenAI87.9%Verified
51Nemotron 3 UltraNvidia87.8%Estimated
52Qwen3.6 Max PreviewAlibaba87.8%Estimated
53Llama 4 ScoutMeta87.2%Verified
54MiniMax M2MiniMax87.0%Estimated
55MiniMax M1MiniMax86.8%Estimated
56Morph V3 LargeMorph AI86.8%Estimated
57Gemini 3.1 FlashGoogle86.8%Verified
58Mistral LargeMistral86.8%Estimated
59Mistral Large 3Mistral86.8%Verified
60Claude 3 OpusAnthropic86.8%Verified
61ERNIE 4.5 VL 424B A47BBaidu86.8%Verified
62GLM 5Zhipu AI86.5%Verified
63GPT-5 MiniOpenAI86.5%Verified
64Gemini 2.5 Pro Preview 05-06Google86.4%Estimated
65Gemini 2.5 Pro Preview 06-05Google86.4%Estimated
66Llama 3.3 70B InstructMeta86.2%Verified
67MiniMax M2.5MiniMax86.2%Estimated
68KAT-Coder-Pro V2.5Kuaishou86.0%Estimated
69Mistral Large 2407Mistral85.8%Estimated
70Nano Banana Pro (Gemini 3 Pro Image Preview)Google85.6%Estimated
71Mistral Large 2Mistral85.4%Estimated
72R1 Distill Llama 70BDeepSeek85.2%Verified
73Kimi K2.7 CodeMoonshot AI85.0%Verified
74Claude Haiku 4.5Anthropic84.8%Verified
75Step 3.7 FlashStepFun84.5%Verified
76Qwen3 Coder PlusAlibaba84.5%Verified
77Kimi K2.6Moonshot AI84.2%Verified
78MiniMax M3MiniMax84.0%Verified
79Hy3 previewTencent83.5%Verified
80GLM 5 TurboZhipu AI83.0%Verified
81MiniMax M2.7MiniMax82.5%Verified
82Kimi K2.5Moonshot AI82.5%Verified
83DeepSeek V4 FlashDeepSeek82.4%Verified
84GPT-4o-miniOpenAI82.0%Verified
85Llama 4 MaverickMeta81.6%Estimated
86GLM 4.7Zhipu AI81.5%Verified
87MiMo-V2.5Xiaomi81.4%Estimated
88Qwen3.5-122B-A10BAlibaba81.4%Estimated
89Qwen2.5 Coder 32B InstructAlibaba81.2%Verified
90Claude Opus 5Anthropic81.2%Estimated
91Qwen3 Next 80B A3B InstructAlibaba81.2%Estimated
92GPT-3.5 Turbo (older v0613)OpenAI81.2%Estimated
93Mistral Small 3Mistral81.2%Verified
94Claude Fable 5Anthropic81.2%Estimated
95MiniMax-01MiniMax81.0%Verified
96Grok 4.20xAI81.0%Estimated
97Laguna S 2.1Poolside81.0%Estimated
98Qwen-PlusAlibaba81.0%Verified
99GPT Chat LatestOpenAI80.8%Estimated
100Qwen3.5 Plus 2026-02-15Alibaba80.8%Estimated
101Yi-Lightning01.AI80.8%Estimated
102Qwen3 VL 30B A3B ThinkingAlibaba80.6%Estimated
103Gemma 4 31BGoogle80.4%Estimated
104GPT-4.1 NanoOpenAI80.4%Estimated
105Qwen3.6 35B A3BAlibaba80.4%Estimated
106GPT-4 Turbo PreviewOpenAI80.2%Estimated
107Qwen3.5-35B-A3BAlibaba80.2%Estimated
108Qwen3 VL 30B A3B InstructAlibaba80.2%Estimated
109Mistral Medium 3.1Mistral80.2%Estimated
110ERNIE 4.0Baidu80.0%Estimated
111Qwen3 Coder 480B A35BAlibaba80.0%Estimated
112Qwen3 30B A3B Thinking 2507Alibaba80.0%Estimated
113DeepSeek V3.1DeepSeek80.0%Estimated
114gpt-oss-20bOpenAI79.8%Estimated
115DeepSeek V3 0324DeepSeek79.8%Estimated
116Qwen3 30B A3B Instruct 2507Alibaba79.6%Estimated
117Llama 3.2 3B InstructMeta79.6%Estimated
118GLM 4.5VZhipu AI79.4%Estimated
119GPT-4o (2024-05-13)OpenAI79.4%Estimated
120Qwen Plus 0728Alibaba79.4%Estimated
121DeepSeek V3.1 TerminusDeepSeek79.4%Estimated
122SonarPerplexity79.2%Estimated
123GPT-4OpenAI79.2%Estimated
124Claude 3 SonnetAnthropic79.0%Verified
125Qwen3.6 PlusAlibaba79.0%Estimated
126Gemma 3n 4BGoogle79.0%Estimated
127Llama 3.1 405BMeta79.0%Estimated
128Hermes 3 405B InstructNous Research78.8%Estimated
129Grok 4.5xAI78.8%Estimated
130Anthropic Claude Haiku LatestAnthropic78.8%Estimated
131Mixtral 8x22B InstructMistral78.6%Estimated
132Hunyuan A13B InstructTencent78.5%Verified
133GPT-4o (2024-11-20)OpenAI78.4%Estimated
134Mistral NemoMistral78.4%Estimated
135Mistral Medium 3Mistral78.4%Estimated
136Qwen 2.5-Coder 32BAlibaba78.4%Estimated
137Qwen2.5 7B InstructAlibaba75.8%Verified
138Command R+Cohere75.7%Verified
139Gemini 3.1 Flash Lite PreviewGoogle75.6%Estimated
140Mistral Small 4Mistral75.6%Estimated
141Gemini 2.0 FlashGoogle75.6%Estimated
142Qwen3 VL 8B InstructAlibaba75.4%Estimated
143Nano Banana 2 (Gemini 3.1 Flash Image)Google75.4%Estimated
144Mistral Small 3.2 24BMistral75.4%Estimated
145o4 Mini Deep ResearchOpenAI75.4%Estimated
146Nano Banana 2 (Gemini 3.1 Flash Image Preview)Google75.2%Estimated
147Claude 3 HaikuAnthropic75.2%Verified
148Qwen3 8BAlibaba75.2%Estimated
149Mistral Small 3.1 24BMistral75.2%Estimated
150Claude 3.5 HaikuAnthropic75.2%Verified
151WizardLM-2 8x22BMicrosoft74.8%Estimated
152GPT-4.1 MiniOpenAI74.6%Estimated
153Qwen3.7 FlashAlibaba74.4%Estimated
154GPT-4o-mini (2024-07-18)OpenAI74.4%Estimated
155GPT-4.1 Mini (batch)OpenAI74.2%Estimated
156o4 MiniOpenAI74.2%Estimated
157Seed-2.0-MiniByteDance74.2%Estimated
158Seed 1.6 FlashByteDance74.2%Estimated
159Qwen3.5-FlashAlibaba74.0%Estimated
160Nano Banana (Gemini 2.5 Flash Image)Google74.0%Estimated
161Llama 3.1 8BMeta74.0%Estimated
162Gemini 3 Flash Preview (batch)Google73.6%Estimated
163Gemini 3.6 FlashGoogle73.6%Estimated
164UI-TARS 7BByteDance73.5%Verified
165Gemini 3.5 Flash LiteGoogle73.2%Estimated
166Nova 2 LiteAmazon73.2%Estimated
167GPT Audio MiniOpenAI73.2%Estimated
168Ministral 3 3B 2512Mistral73.2%Estimated
169Gemini 2.5 FlashGoogle73.2%Estimated
170Gemini 3.5 Flash (batch)Google73.0%Estimated
171Llama 3.2 11B VisionMeta73.0%Verified
172Gemini 3.5 Flash Lite (batch)Google72.8%Estimated
173Qwen3.5 397B A17BAlibaba72.8%Estimated
174Gemini 2.5 Flash Lite (batch)Google72.6%Estimated
175Google Gemini Flash LatestGoogle72.6%Estimated
176Gemini 3.1 Flash LiteGoogle72.4%Estimated
177Command RCohere71.0%Verified
Pending Evaluation Models (208)keyboard_arrow_down
Seed 1.6ByteDance
GPT-5.4OpenAI
SabaMistral
Hermes 4 70BNous Research
Kimi K3Moonshot AI
Hy3Tencent
GLM 4.6VZhipu AI
Fugu UltraSakana AI
o1-proOpenAI
Laguna M.1Poolside
Nex-N2-ProNex AGI
Sonar ProPerplexity
GLM 4.5 AirZhipu AI
Reka EdgeReka AI
Seed-2.0-LiteByteDance
GLM 5V TurboZhipu AI
GPT AudioOpenAI
Kimi K2 0711Moonshot AI
GLM 4.5Zhipu AI
GLM 4.6Zhipu AI
Doubao ProByteDance
Qwen3.5-9BAlibaba
Mercury 2Inception AI
GPT-5 ProOpenAI
GPT-5.2OpenAI
Kimi K2 0905Moonshot AI
Qwen3 32BAlibaba
Hermes 4 405BNous Research
Qwen3 14BAlibaba
R1DeepSeek
Phi 4Microsoft
Inflection 3 PiInflection
GPT-4.1OpenAI
Command ACohere

About MMLU

Massive Multitask Language Understanding (MMLU) measures general knowledge across 57 academic subjects from elementary math to professional law. It is the industry standard for evaluating general reasoning and semantic understanding.

MMLU consists of multiple-choice questions covering humanities, social sciences, STEM, and other professional contexts. It tests both world knowledge and problem-solving capability. Strong performance indicates a highly versatile model capable of handling diverse tasks without specialized fine-tuning.

What do these benchmarks mean?

What is the MMLU benchmark?keyboard_arrow_down

Massive Multitask Language Understanding (MMLU) measures general knowledge across 57 academic subjects from elementary math to professional law. It is the industry standard for evaluating general reasoning and semantic understanding.

MMLU consists of multiple-choice questions covering humanities, social sciences, STEM, and other professional contexts. It tests both world knowledge and problem-solving capability. Strong performance indicates a highly versatile model capable of handling diverse tasks without specialized fine-tuning.

What is the HumanEval benchmark?keyboard_arrow_down

HumanEval is a programming benchmark created by OpenAI to evaluate coding abilities. It measures the accuracy of models in generating functional Python code blocks based on docstring instructions.

HumanEval consists of 164 hand-written programming problems. Models are evaluated using pass@1 metrics, meaning the code is executed against automated unit tests and must pass on the first try. High scores correspond directly to software engineering utility, agentic code writing, and syntactical precision.

What is the MATH benchmark?keyboard_arrow_down

MATH measures mathematical problem-solving skills across seven high-school and college disciplines. It requires models to perform complex, multi-step symbolic reasoning rather than simple calculation.

MATH is exceptionally difficult for standard models. It covers algebra, calculus, probability, geometry, and number theory. Unlike multiple-choice benchmarks, MATH requires generating final equations or numerical answers, testing a model's chain-of-thought planning and logical precision.

What is the MT-Bench benchmark?keyboard_arrow_down

MT-Bench is a multi-turn conversation benchmark. It evaluates how well models maintain coherence, logic, and instructions across progressive dialogue exchanges.

MT-Bench tests eight categories of tasks including coding, math, roleplay, and writing. A powerful model like GPT-5 is utilized as a judge to grade the responses on a scale from 1 to 10. High performance indicates excellent instruction following and conversational context retention over long interactions.

What is the GPQA benchmark?keyboard_arrow_down

GPQA (Graduate-Level Google-Proof Q&A Benchmark) tests advanced scientific and mathematical understanding using questions designed by PhD-level experts. The questions are specifically written to be difficult to answer via search engines.

GPQA contains physics, biology, and chemistry questions that even human experts find challenging. Non-experts with access to Google search only score around 34%, while models must exhibit advanced abstract reasoning and scientific understanding to pass, serving as a key benchmark for expert-level capability.

What is the HellaSwag benchmark?keyboard_arrow_down

HellaSwag evaluates common-sense reasoning and situational prediction. It tests whether a model can accurately determine the most likely next event in a described physical scenario.

HellaSwag is designed using adversarial filtering to find scenarios that are easy for humans (who score ~95%) but difficult for language models. It requires deep contextual understanding of everyday physics, human intent, and spatial logic.