BENCHMARK LEADERBOARD

Model Evaluation Leaderboards

Track ranking standing across standardized evaluations. Analyze MMLU, HumanEval, and GPQA performance to identify capabilities suitable for your product.

RankModelProviderReleasedScoreSource
1OpenAIJul 202696.8%Verified
2OpenAIApr 202696.5%Verified
3OpenAISep 202696.2%Verified
4AnthropicSep 202695.8%Verified
5OpenAISep 202695.4%Verified
6AnthropicMay 202695.4%Verified
7xAIMay 202694.7%Verified
8DeepSeekApr 202694.5%Verified
9DeepSeekDec 202494.5%Verified
10OpenAIApr 202694.2%Verified
11AnthropicApr 202694.1%Verified
12OpenAINov 202593.6%Estimated
13OpenAIOct 202593.6%Estimated
14OpenAIMar 202593.6%Estimated
15OpenAIJul 202693.4%Estimated
16OpenAIMar 202693.4%Estimated
17OpenAIFeb 202693.4%Verified
18OpenAIDec 202593.4%Estimated
19OpenAIMar 202693.2%Estimated
20OpenAINov 202593.2%Estimated
21OpenAIJan 202693.0%Estimated
22OpenAIOct 202593.0%Estimated
23AlibabaSep 202692.8%Verified
24GoogleApr 202692.8%Verified
25AnthropicNov 202592.8%Verified
26OpenAINov 202592.8%Estimated
27OpenAIJul 202692.6%Estimated
28OpenAIJul 202692.6%Estimated
29AnthropicMay 202692.6%Estimated
30AnthropicFeb 202692.5%Verified
31AnthropicMay 202692.4%Estimated
32xAIApr 202692.4%Verified
33OpenAINov 202592.4%Estimated
34OpenAIApr 202592.4%Estimated
35OpenAIFeb 202592.4%Estimated
36OpenAIApr 202692.2%Estimated
37OpenAIDec 202592.2%Estimated
38OpenAIAug 202592.1%Verified
39OpenAIMar 202692.0%Estimated
40AlibabaFeb 202692.0%Verified
41DeepSeekMay 202592.0%Estimated
42OpenAIAug 202591.8%Estimated
43OpenAIDec 202491.8%Verified
44OpenAIDec 202591.6%Estimated
45MetaMay 202691.5%Verified
46AlibabaMay 202691.5%Verified
47AnthropicAug 202591.4%Estimated
48GoogleSep 202691.2%Verified
49OpenAIOct 202591.2%Estimated
50OpenAIJul 202691.0%Estimated
51OpenAISep 202591.0%Estimated
52OpenAIJul 202690.8%Estimated
53OpenAIOct 202590.8%Estimated
54OpenAIAug 202590.8%Estimated
55DeepSeekJan 202590.8%Verified
56DeepSeekJan 202590.8%Verified
57OpenAIMar 202690.6%Estimated
58OpenAIJun 202590.6%Estimated
59MetaSep 202690.5%Verified
60GoogleMay 202690.5%Verified
61AnthropicMay 202590.5%Verified
62OpenAIMar 202690.4%Estimated
63AnthropicFeb 202690.4%Verified
64OpenAIDec 202590.4%Estimated
65GoogleJun 202589.9%Verified
66Zhipu AIJun 202689.5%Verified
67AnthropicSep 202589.5%Verified
68OpenAIMay 202488.7%Verified
69GoogleApr 202688.6%Estimated
70DeepSeekApr 202688.6%Estimated
71AlibabaSep 202588.6%Estimated
72Moonshot AINov 202588.5%Verified
73AlibabaAug 202688.4%Estimated
74GoogleFeb 202688.4%Estimated
75MiniMaxJan 202688.4%Estimated
76AlibabaJun 202688.2%Verified
77GoogleApr 202688.2%Estimated
78ByteDanceJan 202688.2%Estimated
79Zhipu AIApr 202688.0%Verified
80AnthropicMay 202588.0%Verified
81OpenAIJan 202587.9%Verified
82Sakana AISep 202687.8%Estimated
83NvidiaJun 202687.8%Estimated
84AlibabaApr 202687.8%Estimated
85AI21 LabsAug 202587.8%Estimated
86PerplexityMar 202587.8%Estimated
87MistralDec 202587.4%Estimated
88InflectionOct 202487.4%Estimated
89Sakana AISep 202687.2%Estimated
90MetaApr 202587.2%Verified
91MiniMaxOct 202587.0%Estimated
92MetaDec 202487.0%Estimated
93GoogleApr 202686.8%Verified
94KuaishouMar 202686.8%Estimated
95GoogleFeb 202686.8%Estimated
96PerplexityOct 202586.8%Estimated
97Morph AIJul 202586.8%Estimated
98BaiduJun 202586.8%Verified
99MiniMaxJun 202586.8%Estimated
100MistralJul 202486.8%Verified
101MistralFeb 202486.8%Estimated
102XiaomiApr 202686.6%Estimated
103MetaJul 202486.6%Estimated
104Zhipu AIFeb 202686.5%Verified
105OpenAIAug 202586.5%Verified
106Inception AIAug 202686.4%Verified
107UpstageAug 202686.4%Estimated
108GoogleJun 202586.4%Estimated
109GoogleMay 202586.4%Estimated
110OpenAIApr 202486.4%Verified
111MiniMaxFeb 202686.2%Estimated
112UpstageJan 202686.2%Estimated
113MetaDec 202486.2%Verified
114KuaishouJul 202686.0%Estimated
115GoogleMar 202686.0%Estimated
116TencentNov 202586.0%Estimated
117PerplexityMar 202586.0%Estimated
118AmazonDec 202486.0%Estimated
119Nous ResearchAug 202486.0%Estimated
120DeepSeekAug 202685.8%Estimated
121GoogleJun 202685.8%Estimated
122Nous ResearchAug 202585.8%Estimated
123MistralNov 202485.8%Estimated
124Nex AGIJun 202685.6%Estimated
125GoogleNov 202585.6%Estimated
126Sakana AIJun 202685.4%Estimated
127MiniMaxDec 202585.4%Estimated
128MistralJul 202585.4%Estimated
129AlibabaSep 202485.3%Verified
130DeepSeekJan 202585.2%Verified
131Moonshot AIJun 202685.0%Verified
132AnthropicOct 202584.8%Verified
133StepFunMay 202684.5%Verified
134AlibabaSep 202584.5%Verified
135Moonshot AIApr 202684.2%Verified
136Moonshot AIApr 202684.2%Verified
137MiniMaxMay 202684.0%Verified
138TencentApr 202683.5%Verified
139Zhipu AIMar 202683.0%Verified
140MiniMaxMar 202682.5%Verified
141Moonshot AIJan 202682.5%Verified
142DeepSeekApr 202682.4%Verified
143OpenAIJul 202482.0%Verified
144Zhipu AIAug 202681.6%Estimated
145KuaishouJul 202681.6%Estimated
146AnthropicApr 202681.6%Estimated
147Zhipu AIApr 202681.6%Estimated
148ByteDanceDec 202581.6%Estimated
149AlibabaSep 202581.6%Estimated
150Morph AIJul 202581.6%Estimated
151MetaApr 202581.6%Estimated
152MistralDec 202481.6%Estimated
153Zhipu AIDec 202581.5%Verified
154MetaAug 202681.4%Estimated
155OpenRouterJun 202681.4%Estimated
156NvidiaJun 202681.4%Estimated
157XiaomiApr 202681.4%Estimated
158AlibabaFeb 202681.4%Estimated
159InflectionOct 202481.4%Estimated
160MetaSep 202681.2%Estimated
161AnthropicJul 202681.2%Estimated
162MetaJul 202681.2%Estimated
163PoolsideJul 202681.2%Estimated
164AnthropicJun 202681.2%Estimated
165PoolsideApr 202681.2%Estimated
166AmazonOct 202581.2%Estimated
167AlibabaSep 202581.2%Estimated
168MistralJan 202581.2%Verified
169AlibabaNov 202481.2%Verified
170MetaSep 202481.2%Estimated
171CohereAug 202481.2%Estimated
172OpenAIJan 202481.2%Estimated
173MetaAug 202681.0%Estimated
174AlibabaAug 202681.0%Estimated
175AnthropicJul 202681.0%Estimated
176PoolsideJul 202681.0%Estimated
177MistralApr 202681.0%Estimated
178xAIMar 202681.0%Estimated
179NvidiaDec 202581.0%Estimated
180DeepSeekSep 202581.0%Estimated
181AlibabaSep 202581.0%Estimated
182Moonshot AISep 202581.0%Estimated
183OpenAIApr 202581.0%Estimated
184The DrummerMar 202581.0%Estimated
185AlibabaFeb 202581.0%Verified
186MiniMaxJan 202581.0%Verified
187NvidiaAug 202680.8%Estimated
188Sakana AIAug 202680.8%Estimated
189OpenAIMay 202680.8%Estimated
190xAIMar 202680.8%Estimated
191AlibabaFeb 202680.8%Estimated
192OpenAIOct 202580.8%Estimated
19301.AIOct 202580.8%Estimated
194GoogleMar 202580.8%Estimated
195AlibabaMar 202680.6%Estimated
196AlibabaOct 202580.6%Estimated
197AlibabaSep 202580.6%Estimated
198CohereMar 202580.6%Estimated
199AlibabaApr 202680.4%Estimated
200AlibabaApr 202680.4%Estimated
201GoogleApr 202680.4%Estimated
202AlibabaFeb 202680.4%Estimated
203OpenRouterDec 202580.4%Estimated
204AlibabaJul 202580.4%Estimated
205Zhipu AIJul 202580.4%Estimated
206AlibabaApr 202580.4%Estimated
207OpenAIApr 202580.4%Estimated
208PerplexityMar 202580.4%Estimated
209OpenAIAug 202480.4%Estimated
210AnthropicApr 202680.2%Estimated
211AlibabaFeb 202680.2%Estimated
212DeepSeekDec 202580.2%Estimated
213AlibabaOct 202580.2%Estimated
214AlibabaOct 202580.2%Estimated
215AlibabaSep 202580.2%Estimated
216MistralAug 202580.2%Estimated
217OpenAIJan 202480.2%Estimated
218TencentAug 202680.0%Estimated
219ByteDanceAug 202680.0%Estimated
220WriterJan 202680.0%Estimated
221BaiduAug 202580.0%Estimated
222AlibabaAug 202580.0%Estimated
223DeepSeekAug 202580.0%Estimated
224AlibabaJul 202580.0%Estimated
225AlibabaJul 202580.0%Estimated
226Moonshot AIJul 202580.0%Estimated
227AlibabaApr 202580.0%Estimated
228AlibabaApr 202580.0%Estimated
229AlibabaApr 202580.0%Estimated
230GoogleMar 202580.0%Estimated
231AlibabaFeb 202580.0%Estimated
232OpenRouterApr 202679.8%Estimated
233OpenRouterFeb 202679.8%Estimated
234Zhipu AISep 202579.8%Estimated
235OpenAIAug 202579.8%Estimated
236MistralAug 202579.8%Estimated
237MetaApr 202579.8%Estimated
238DeepSeekMar 202579.8%Estimated
239OpenAIMar 202579.8%Estimated
240OpenAISep 202379.8%Estimated
241xAIMay 202679.6%Estimated
242AnthropicApr 202679.6%Estimated
243OpenAIJan 202679.6%Estimated
244Zhipu AIDec 202579.6%Estimated
245The DrummerSep 202579.6%Estimated
246AlibabaJul 202579.6%Estimated
247Zhipu AIJul 202579.6%Estimated
248MetaSep 202479.6%Estimated
249TencentAug 202679.4%Estimated
250Zhipu AIAug 202679.4%Estimated
251TencentJul 202679.4%Estimated
252GoogleMar 202679.4%Estimated
253MistralDec 202579.4%Estimated
254DeepSeekSep 202579.4%Estimated
255AlibabaSep 202579.4%Estimated
256Zhipu AIAug 202579.4%Estimated
257OpenAIAug 202579.4%Estimated
258OpenAIMay 202479.4%Estimated
259OpenAIAug 202379.4%Estimated
260OpenAIMay 202379.4%Estimated
261xAISep 202679.2%Estimated
262Inception AIMar 202679.2%Estimated
263AnthropicFeb 202579.2%Estimated
264PerplexityJan 202579.2%Estimated
265The DrummerNov 202479.2%Estimated
266MetaSep 202479.2%Estimated
267OpenRouterNov 202379.2%Estimated
268OpenAIMay 202379.2%Estimated
269xAIAug 202679.0%Estimated
270Moonshot AIApr 202679.0%Estimated
271AlibabaApr 202679.0%Estimated
272Reka AIMar 202679.0%Estimated
273GoogleMay 202579.0%Estimated
274MetaJul 202479.0%Estimated
275OpenRouterJul 202678.8%Estimated
276xAIJul 202678.8%Estimated
277AnthropicJun 202678.8%Estimated
278AnthropicApr 202678.8%Estimated
279AnthropicApr 202678.8%Estimated
280NvidiaMar 202678.8%Estimated
281AlibabaJul 202578.8%Estimated
282Nous ResearchAug 202478.8%Estimated
283ByteDanceAug 202678.6%Estimated
284Thinking MachinesJul 202678.6%Estimated
285Moonshot AIJul 202678.6%Estimated
286AnthropicJun 202678.6%Estimated
287AlibabaSep 202578.6%Estimated
288Nous ResearchAug 202578.6%Estimated
289MistralApr 202478.6%Estimated
290TencentJul 202578.5%Verified
291Inception AISep 202678.4%Estimated
292MetaAug 202678.4%Estimated
293GoogleApr 202678.4%Estimated
294AlibabaNov 202578.4%Estimated
295MistralMay 202578.4%Estimated
296MistralFeb 202578.4%Estimated
297OpenAINov 202478.4%Estimated
298MistralJul 202478.4%Estimated
299Zhipu AIJan 202677.2%Verified
300AlibabaOct 202475.8%Verified
301CohereApr 202475.7%Verified
302Zhipu AISep 202675.6%Estimated
303IBMAug 202675.6%Estimated
304GoogleApr 202675.6%Estimated
305MistralMar 202675.6%Estimated
306GoogleMar 202675.6%Estimated
307Reka AIMar 202575.6%Estimated
308GoogleFeb 202575.6%Estimated
309Zhipu AIAug 202675.4%Estimated
310GoogleJun 202675.4%Estimated
311IBMApr 202675.4%Estimated
312AlibabaOct 202575.4%Estimated
313OpenAIOct 202575.4%Estimated
314MistralJun 202575.4%Estimated
315Inclusion AIJul 202675.2%Estimated
316GoogleFeb 202675.2%Estimated
317AlibabaApr 202575.2%Estimated
318MistralMar 202575.2%Estimated
319GoogleMar 202575.2%Estimated
320AnthropicMar 202475.2%Verified
321Inclusion AISep 202675.0%Estimated
322TencentAug 202675.0%Estimated
323GoogleJun 202675.0%Estimated
324GoogleJul 202475.0%Estimated
325MicrosoftApr 202474.8%Estimated
326AlibabaAug 202674.6%Estimated
327AlibabaAug 202674.6%Estimated
328OpenAIApr 202574.6%Estimated
329OpenAIApr 202574.6%Estimated
330MicrosoftJan 202574.6%Estimated
331AmazonDec 202474.6%Estimated
332AlibabaJul 202674.4%Estimated
333ByteDanceMar 202674.4%Estimated
334MistralOct 202574.4%Estimated
335MetaJul 202474.4%Estimated
336OpenAIJul 202474.4%Estimated
337Zhipu AIAug 202674.2%Estimated
338DeepSeekAug 202674.2%Estimated
339AlibabaApr 202674.2%Estimated
340AlibabaApr 202674.2%Estimated
341ByteDanceFeb 202674.2%Estimated
342ByteDanceDec 202574.2%Estimated
343MistralDec 202574.2%Estimated
344OpenAIApr 202574.2%Estimated
345CohereDec 202474.2%Estimated
346Nex AGIJun 202674.0%Estimated
347AlibabaFeb 202674.0%Estimated
348AlibabaFeb 202674.0%Estimated
349GoogleDec 202574.0%Estimated
350GoogleOct 202574.0%Estimated
351MetaJul 202474.0%Estimated
352GoogleAug 202673.8%Estimated
353IBMOct 202573.8%Estimated
354GoogleSep 202573.8%Estimated
355TencentAug 202673.6%Estimated
356GoogleJul 202673.6%Estimated
357MistralOct 202473.6%Estimated
358ByteDanceJul 202573.5%Verified
359Thinking MachinesJul 202673.4%Estimated
360OpenAIMar 202573.4%Estimated
361GoogleJul 202673.2%Estimated
362OpenAIJan 202673.2%Estimated
363AmazonDec 202573.2%Estimated
364MistralDec 202573.2%Estimated
365AlibabaSep 202573.2%Estimated
366GoogleJun 202573.2%Estimated
367DeepSeekJul 202673.0%Estimated
368StepFunJan 202673.0%Estimated
369MistralDec 202573.0%Estimated
370GoogleJul 202573.0%Estimated
371MetaSep 202473.0%Verified
372DeepSeekSep 202672.8%Estimated
373Inclusion AIAug 202672.8%Estimated
374AlibabaFeb 202672.8%Estimated
375GoogleApr 202672.6%Estimated
376DeepSeekApr 202672.6%Estimated
377MetaApr 202472.6%Estimated
378GoogleMay 202672.4%Estimated
379AlibabaOct 202572.4%Estimated
380AmazonDec 202472.4%Estimated
381CohereMar 202471.0%Verified

About MMLU

Massive Multitask Language Understanding (MMLU) measures general knowledge across 57 academic subjects from elementary math to professional law. It is the industry standard for evaluating general reasoning and semantic understanding.

MMLU consists of multiple-choice questions covering humanities, social sciences, STEM, and other professional contexts. It tests both world knowledge and problem-solving capability. Strong performance indicates a highly versatile model capable of handling diverse tasks without specialized fine-tuning.

What do these benchmarks mean?

What is the MMLU benchmark?keyboard_arrow_down

Massive Multitask Language Understanding (MMLU) measures general knowledge across 57 academic subjects from elementary math to professional law. It is the industry standard for evaluating general reasoning and semantic understanding.

MMLU consists of multiple-choice questions covering humanities, social sciences, STEM, and other professional contexts. It tests both world knowledge and problem-solving capability. Strong performance indicates a highly versatile model capable of handling diverse tasks without specialized fine-tuning.

What is the HumanEval benchmark?keyboard_arrow_down

HumanEval is a programming benchmark created by OpenAI to evaluate coding abilities. It measures the accuracy of models in generating functional Python code blocks based on docstring instructions.

HumanEval consists of 164 hand-written programming problems. Models are evaluated using pass@1 metrics, meaning the code is executed against automated unit tests and must pass on the first try. High scores correspond directly to software engineering utility, agentic code writing, and syntactical precision.

What is the MATH benchmark?keyboard_arrow_down

MATH measures mathematical problem-solving skills across seven high-school and college disciplines. It requires models to perform complex, multi-step symbolic reasoning rather than simple calculation.

MATH is exceptionally difficult for standard models. It covers algebra, calculus, probability, geometry, and number theory. Unlike multiple-choice benchmarks, MATH requires generating final equations or numerical answers, testing a model's chain-of-thought planning and logical precision.

What is the MT-Bench benchmark?keyboard_arrow_down

MT-Bench is a multi-turn conversation benchmark. It evaluates how well models maintain coherence, logic, and instructions across progressive dialogue exchanges.

MT-Bench tests eight categories of tasks including coding, math, roleplay, and writing. A powerful model like GPT-5 is utilized as a judge to grade the responses on a scale from 1 to 10. High performance indicates excellent instruction following and conversational context retention over long interactions.

What is the GPQA benchmark?keyboard_arrow_down

GPQA (Graduate-Level Google-Proof Q&A Benchmark) tests advanced scientific and mathematical understanding using questions designed by PhD-level experts. The questions are specifically written to be difficult to answer via search engines.

GPQA contains physics, biology, and chemistry questions that even human experts find challenging. Non-experts with access to Google search only score around 34%, while models must exhibit advanced abstract reasoning and scientific understanding to pass, serving as a key benchmark for expert-level capability.

What is the HellaSwag benchmark?keyboard_arrow_down

HellaSwag evaluates common-sense reasoning and situational prediction. It tests whether a model can accurately determine the most likely next event in a described physical scenario.

HellaSwag is designed using adversarial filtering to find scenarios that are easy for humans (who score ~95%) but difficult for language models. It requires deep contextual understanding of everyday physics, human intent, and spatial logic.