Model Evaluation Leaderboards
Track ranking standing across standardized evaluations. Analyze MMLU, HumanEval, and GPQA performance to identify capabilities suitable for your product.
| Rank | Model | Provider | Score | Source |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol | OpenAI | 96.8% | Verified |
| 2 | GPT-5.5 Pro | OpenAI | 96.5% | Verified |
| 3 | Claude Opus 4.8 | Anthropic | 95.4% | Verified |
| 4 | Grok 4.20 | xAI | 94.7% | Verified |
| 5 | DeepSeek V4 Pro | DeepSeek | 94.5% | Verified |
| 6 | DeepSeek V3 | DeepSeek | 94.5% | Verified |
| 7 | GPT-5.5 | OpenAI | 94.2% | Verified |
| 8 | Claude Opus 4.7 | Anthropic | 94.1% | Verified |
| 9 | GPT-5.1 | OpenAI | 93.2% | Estimated |
| 10 | GPT-5.2 (batch) | OpenAI | 93.0% | Estimated |
| 11 | GPT-5.2-Codex | OpenAI | 93.0% | Estimated |
| 12 | GPT-5.4 Nano (batch) | OpenAI | 92.8% | Estimated |
| 13 | GPT-5.1 (batch) | OpenAI | 92.8% | Estimated |
| 14 | Claude Opus 4.5 | Anthropic | 92.8% | Verified |
| 15 | GPT-5.1-Codex | OpenAI | 92.8% | Estimated |
| 16 | Gemini 3.1 Pro | 92.8% | Verified | |
| 17 | GPT-5.6 Luna Pro | OpenAI | 92.6% | Estimated |
| 18 | Claude Opus 4.6 | Anthropic | 92.5% | Verified |
| 19 | o3 Mini High | OpenAI | 92.4% | Estimated |
| 20 | o3 | OpenAI | 92.4% | Estimated |
| 21 | Claude Opus 4.7 (Fast) | Anthropic | 92.4% | Estimated |
| 22 | Grok 4.3 | xAI | 92.4% | Verified |
| 23 | GPT-5.4 Image 2 | OpenAI | 92.2% | Estimated |
| 24 | GPT-5 | OpenAI | 92.1% | Verified |
| 25 | Qwen3 Max Thinking | Alibaba | 92.0% | Verified |
| 26 | R1 0528 | DeepSeek | 92.0% | Estimated |
| 27 | o1 | OpenAI | 91.8% | Verified |
| 28 | GPT-5 Chat | OpenAI | 91.8% | Estimated |
| 29 | Llama 4 Maverick | Meta | 91.5% | Verified |
| 30 | Qwen3.7 Max | Alibaba | 91.5% | Verified |
| 31 | GPT-5 Image | OpenAI | 91.2% | Estimated |
| 32 | GPT-5.6 Luna | OpenAI | 91.0% | Estimated |
| 33 | DeepSeek R1 | DeepSeek | 90.8% | Verified |
| 34 | GPT-5 Nano | OpenAI | 90.8% | Estimated |
| 35 | o3 Pro | OpenAI | 90.6% | Estimated |
| 36 | Claude Opus 4 | Anthropic | 90.5% | Verified |
| 37 | Gemini 3.5 Flash | 90.5% | Verified | |
| 38 | Gemini 2.5 Pro | 89.9% | Verified | |
| 39 | Claude Sonnet 4.5 | Anthropic | 89.5% | Verified |
| 40 | GLM 5.2 | Zhipu AI | 89.5% | Verified |
| 41 | GPT-4o | OpenAI | 88.7% | Verified |
| 42 | Claude 3.5 Sonnet | Anthropic | 88.7% | Verified |
| 43 | Google Gemini Pro Latest | 88.6% | Estimated | |
| 44 | Qwen3 Max | Alibaba | 88.6% | Estimated |
| 45 | Kimi K2 Thinking | Moonshot AI | 88.5% | Verified |
| 46 | Gemini 3.1 Pro Preview | 88.4% | Estimated | |
| 47 | Qwen3.7 Plus | Alibaba | 88.2% | Verified |
| 48 | Claude Sonnet 4 | Anthropic | 88.0% | Verified |
| 49 | GLM 5.1 | Zhipu AI | 88.0% | Verified |
| 50 | o3 Mini | OpenAI | 87.9% | Verified |
| 51 | Nemotron 3 Ultra | Nvidia | 87.8% | Estimated |
| 52 | Qwen3.6 Max Preview | Alibaba | 87.8% | Estimated |
| 53 | Llama 4 Scout | Meta | 87.2% | Verified |
| 54 | MiniMax M2 | MiniMax | 87.0% | Estimated |
| 55 | MiniMax M1 | MiniMax | 86.8% | Estimated |
| 56 | Morph V3 Large | Morph AI | 86.8% | Estimated |
| 57 | Gemini 3.1 Flash | 86.8% | Verified | |
| 58 | Mistral Large | Mistral | 86.8% | Estimated |
| 59 | Mistral Large 3 | Mistral | 86.8% | Verified |
| 60 | Claude 3 Opus | Anthropic | 86.8% | Verified |
| 61 | ERNIE 4.5 VL 424B A47B | Baidu | 86.8% | Verified |
| 62 | GLM 5 | Zhipu AI | 86.5% | Verified |
| 63 | GPT-5 Mini | OpenAI | 86.5% | Verified |
| 64 | Gemini 2.5 Pro Preview 05-06 | 86.4% | Estimated | |
| 65 | Gemini 2.5 Pro Preview 06-05 | 86.4% | Estimated | |
| 66 | Llama 3.3 70B Instruct | Meta | 86.2% | Verified |
| 67 | MiniMax M2.5 | MiniMax | 86.2% | Estimated |
| 68 | KAT-Coder-Pro V2.5 | Kuaishou | 86.0% | Estimated |
| 69 | Mistral Large 2407 | Mistral | 85.8% | Estimated |
| 70 | Nano Banana Pro (Gemini 3 Pro Image Preview) | 85.6% | Estimated | |
| 71 | Mistral Large 2 | Mistral | 85.4% | Estimated |
| 72 | R1 Distill Llama 70B | DeepSeek | 85.2% | Verified |
| 73 | Kimi K2.7 Code | Moonshot AI | 85.0% | Verified |
| 74 | Claude Haiku 4.5 | Anthropic | 84.8% | Verified |
| 75 | Step 3.7 Flash | StepFun | 84.5% | Verified |
| 76 | Qwen3 Coder Plus | Alibaba | 84.5% | Verified |
| 77 | Kimi K2.6 | Moonshot AI | 84.2% | Verified |
| 78 | MiniMax M3 | MiniMax | 84.0% | Verified |
| 79 | Hy3 preview | Tencent | 83.5% | Verified |
| 80 | GLM 5 Turbo | Zhipu AI | 83.0% | Verified |
| 81 | MiniMax M2.7 | MiniMax | 82.5% | Verified |
| 82 | Kimi K2.5 | Moonshot AI | 82.5% | Verified |
| 83 | DeepSeek V4 Flash | DeepSeek | 82.4% | Verified |
| 84 | GPT-4o-mini | OpenAI | 82.0% | Verified |
| 85 | Llama 4 Maverick | Meta | 81.6% | Estimated |
| 86 | GLM 4.7 | Zhipu AI | 81.5% | Verified |
| 87 | MiMo-V2.5 | Xiaomi | 81.4% | Estimated |
| 88 | Qwen3.5-122B-A10B | Alibaba | 81.4% | Estimated |
| 89 | Qwen2.5 Coder 32B Instruct | Alibaba | 81.2% | Verified |
| 90 | Claude Opus 5 | Anthropic | 81.2% | Estimated |
| 91 | Qwen3 Next 80B A3B Instruct | Alibaba | 81.2% | Estimated |
| 92 | GPT-3.5 Turbo (older v0613) | OpenAI | 81.2% | Estimated |
| 93 | Mistral Small 3 | Mistral | 81.2% | Verified |
| 94 | Claude Fable 5 | Anthropic | 81.2% | Estimated |
| 95 | MiniMax-01 | MiniMax | 81.0% | Verified |
| 96 | Grok 4.20 | xAI | 81.0% | Estimated |
| 97 | Laguna S 2.1 | Poolside | 81.0% | Estimated |
| 98 | Qwen-Plus | Alibaba | 81.0% | Verified |
| 99 | GPT Chat Latest | OpenAI | 80.8% | Estimated |
| 100 | Qwen3.5 Plus 2026-02-15 | Alibaba | 80.8% | Estimated |
| 101 | Yi-Lightning | 01.AI | 80.8% | Estimated |
| 102 | Qwen3 VL 30B A3B Thinking | Alibaba | 80.6% | Estimated |
| 103 | Gemma 4 31B | 80.4% | Estimated | |
| 104 | GPT-4.1 Nano | OpenAI | 80.4% | Estimated |
| 105 | Qwen3.6 35B A3B | Alibaba | 80.4% | Estimated |
| 106 | GPT-4 Turbo Preview | OpenAI | 80.2% | Estimated |
| 107 | Qwen3.5-35B-A3B | Alibaba | 80.2% | Estimated |
| 108 | Qwen3 VL 30B A3B Instruct | Alibaba | 80.2% | Estimated |
| 109 | Mistral Medium 3.1 | Mistral | 80.2% | Estimated |
| 110 | ERNIE 4.0 | Baidu | 80.0% | Estimated |
| 111 | Qwen3 Coder 480B A35B | Alibaba | 80.0% | Estimated |
| 112 | Qwen3 30B A3B Thinking 2507 | Alibaba | 80.0% | Estimated |
| 113 | DeepSeek V3.1 | DeepSeek | 80.0% | Estimated |
| 114 | gpt-oss-20b | OpenAI | 79.8% | Estimated |
| 115 | DeepSeek V3 0324 | DeepSeek | 79.8% | Estimated |
| 116 | Qwen3 30B A3B Instruct 2507 | Alibaba | 79.6% | Estimated |
| 117 | Llama 3.2 3B Instruct | Meta | 79.6% | Estimated |
| 118 | GLM 4.5V | Zhipu AI | 79.4% | Estimated |
| 119 | GPT-4o (2024-05-13) | OpenAI | 79.4% | Estimated |
| 120 | Qwen Plus 0728 | Alibaba | 79.4% | Estimated |
| 121 | DeepSeek V3.1 Terminus | DeepSeek | 79.4% | Estimated |
| 122 | Sonar | Perplexity | 79.2% | Estimated |
| 123 | GPT-4 | OpenAI | 79.2% | Estimated |
| 124 | Claude 3 Sonnet | Anthropic | 79.0% | Verified |
| 125 | Qwen3.6 Plus | Alibaba | 79.0% | Estimated |
| 126 | Gemma 3n 4B | 79.0% | Estimated | |
| 127 | Llama 3.1 405B | Meta | 79.0% | Estimated |
| 128 | Hermes 3 405B Instruct | Nous Research | 78.8% | Estimated |
| 129 | Grok 4.5 | xAI | 78.8% | Estimated |
| 130 | Anthropic Claude Haiku Latest | Anthropic | 78.8% | Estimated |
| 131 | Mixtral 8x22B Instruct | Mistral | 78.6% | Estimated |
| 132 | Hunyuan A13B Instruct | Tencent | 78.5% | Verified |
| 133 | GPT-4o (2024-11-20) | OpenAI | 78.4% | Estimated |
| 134 | Mistral Nemo | Mistral | 78.4% | Estimated |
| 135 | Mistral Medium 3 | Mistral | 78.4% | Estimated |
| 136 | Qwen 2.5-Coder 32B | Alibaba | 78.4% | Estimated |
| 137 | Qwen2.5 7B Instruct | Alibaba | 75.8% | Verified |
| 138 | Command R+ | Cohere | 75.7% | Verified |
| 139 | Gemini 3.1 Flash Lite Preview | 75.6% | Estimated | |
| 140 | Mistral Small 4 | Mistral | 75.6% | Estimated |
| 141 | Gemini 2.0 Flash | 75.6% | Estimated | |
| 142 | Qwen3 VL 8B Instruct | Alibaba | 75.4% | Estimated |
| 143 | Nano Banana 2 (Gemini 3.1 Flash Image) | 75.4% | Estimated | |
| 144 | Mistral Small 3.2 24B | Mistral | 75.4% | Estimated |
| 145 | o4 Mini Deep Research | OpenAI | 75.4% | Estimated |
| 146 | Nano Banana 2 (Gemini 3.1 Flash Image Preview) | 75.2% | Estimated | |
| 147 | Claude 3 Haiku | Anthropic | 75.2% | Verified |
| 148 | Qwen3 8B | Alibaba | 75.2% | Estimated |
| 149 | Mistral Small 3.1 24B | Mistral | 75.2% | Estimated |
| 150 | Claude 3.5 Haiku | Anthropic | 75.2% | Verified |
| 151 | WizardLM-2 8x22B | Microsoft | 74.8% | Estimated |
| 152 | GPT-4.1 Mini | OpenAI | 74.6% | Estimated |
| 153 | Qwen3.7 Flash | Alibaba | 74.4% | Estimated |
| 154 | GPT-4o-mini (2024-07-18) | OpenAI | 74.4% | Estimated |
| 155 | GPT-4.1 Mini (batch) | OpenAI | 74.2% | Estimated |
| 156 | o4 Mini | OpenAI | 74.2% | Estimated |
| 157 | Seed-2.0-Mini | ByteDance | 74.2% | Estimated |
| 158 | Seed 1.6 Flash | ByteDance | 74.2% | Estimated |
| 159 | Qwen3.5-Flash | Alibaba | 74.0% | Estimated |
| 160 | Nano Banana (Gemini 2.5 Flash Image) | 74.0% | Estimated | |
| 161 | Llama 3.1 8B | Meta | 74.0% | Estimated |
| 162 | Gemini 3 Flash Preview (batch) | 73.6% | Estimated | |
| 163 | Gemini 3.6 Flash | 73.6% | Estimated | |
| 164 | UI-TARS 7B | ByteDance | 73.5% | Verified |
| 165 | Gemini 3.5 Flash Lite | 73.2% | Estimated | |
| 166 | Nova 2 Lite | Amazon | 73.2% | Estimated |
| 167 | GPT Audio Mini | OpenAI | 73.2% | Estimated |
| 168 | Ministral 3 3B 2512 | Mistral | 73.2% | Estimated |
| 169 | Gemini 2.5 Flash | 73.2% | Estimated | |
| 170 | Gemini 3.5 Flash (batch) | 73.0% | Estimated | |
| 171 | Llama 3.2 11B Vision | Meta | 73.0% | Verified |
| 172 | Gemini 3.5 Flash Lite (batch) | 72.8% | Estimated | |
| 173 | Qwen3.5 397B A17B | Alibaba | 72.8% | Estimated |
| 174 | Gemini 2.5 Flash Lite (batch) | 72.6% | Estimated | |
| 175 | Google Gemini Flash Latest | 72.6% | Estimated | |
| 176 | Gemini 3.1 Flash Lite | 72.4% | Estimated | |
| 177 | Command R | Cohere | 71.0% | Verified |
Pending Evaluation Models (208)keyboard_arrow_down
About MMLU
Massive Multitask Language Understanding (MMLU) measures general knowledge across 57 academic subjects from elementary math to professional law. It is the industry standard for evaluating general reasoning and semantic understanding.
What do these benchmarks mean?
What is the MMLU benchmark?keyboard_arrow_down
Massive Multitask Language Understanding (MMLU) measures general knowledge across 57 academic subjects from elementary math to professional law. It is the industry standard for evaluating general reasoning and semantic understanding.
MMLU consists of multiple-choice questions covering humanities, social sciences, STEM, and other professional contexts. It tests both world knowledge and problem-solving capability. Strong performance indicates a highly versatile model capable of handling diverse tasks without specialized fine-tuning.
What is the HumanEval benchmark?keyboard_arrow_down
HumanEval is a programming benchmark created by OpenAI to evaluate coding abilities. It measures the accuracy of models in generating functional Python code blocks based on docstring instructions.
HumanEval consists of 164 hand-written programming problems. Models are evaluated using pass@1 metrics, meaning the code is executed against automated unit tests and must pass on the first try. High scores correspond directly to software engineering utility, agentic code writing, and syntactical precision.
What is the MATH benchmark?keyboard_arrow_down
MATH measures mathematical problem-solving skills across seven high-school and college disciplines. It requires models to perform complex, multi-step symbolic reasoning rather than simple calculation.
MATH is exceptionally difficult for standard models. It covers algebra, calculus, probability, geometry, and number theory. Unlike multiple-choice benchmarks, MATH requires generating final equations or numerical answers, testing a model's chain-of-thought planning and logical precision.
What is the MT-Bench benchmark?keyboard_arrow_down
MT-Bench is a multi-turn conversation benchmark. It evaluates how well models maintain coherence, logic, and instructions across progressive dialogue exchanges.
MT-Bench tests eight categories of tasks including coding, math, roleplay, and writing. A powerful model like GPT-5 is utilized as a judge to grade the responses on a scale from 1 to 10. High performance indicates excellent instruction following and conversational context retention over long interactions.
What is the GPQA benchmark?keyboard_arrow_down
GPQA (Graduate-Level Google-Proof Q&A Benchmark) tests advanced scientific and mathematical understanding using questions designed by PhD-level experts. The questions are specifically written to be difficult to answer via search engines.
GPQA contains physics, biology, and chemistry questions that even human experts find challenging. Non-experts with access to Google search only score around 34%, while models must exhibit advanced abstract reasoning and scientific understanding to pass, serving as a key benchmark for expert-level capability.
What is the HellaSwag benchmark?keyboard_arrow_down
HellaSwag evaluates common-sense reasoning and situational prediction. It tests whether a model can accurately determine the most likely next event in a described physical scenario.
HellaSwag is designed using adversarial filtering to find scenarios that are easy for humans (who score ~95%) but difficult for language models. It requires deep contextual understanding of everyday physics, human intent, and spatial logic.