Model Evaluation Leaderboards
Track ranking standing across standardized evaluations. Analyze MMLU, HumanEval, and GPQA performance to identify capabilities suitable for your product.
| Rank | Model | Provider | Released | Score | Source |
|---|---|---|---|---|---|
| 1 | OpenAI | Jul 2026 | 96.8% | Verified | |
| 2 | OpenAI | Apr 2026 | 96.5% | Verified | |
| 3 | OpenAI | Sep 2026 | 96.2% | Verified | |
| 4 | Anthropic | Sep 2026 | 95.8% | Verified | |
| 5 | OpenAI | Sep 2026 | 95.4% | Verified | |
| 6 | Anthropic | May 2026 | 95.4% | Verified | |
| 7 | xAI | May 2026 | 94.7% | Verified | |
| 8 | DeepSeek | Apr 2026 | 94.5% | Verified | |
| 9 | DeepSeek | Dec 2024 | 94.5% | Verified | |
| 10 | OpenAI | Apr 2026 | 94.2% | Verified | |
| 11 | Anthropic | Apr 2026 | 94.1% | Verified | |
| 12 | OpenAI | Nov 2025 | 93.6% | Estimated | |
| 13 | OpenAI | Oct 2025 | 93.6% | Estimated | |
| 14 | OpenAI | Mar 2025 | 93.6% | Estimated | |
| 15 | OpenAI | Jul 2026 | 93.4% | Estimated | |
| 16 | OpenAI | Mar 2026 | 93.4% | Estimated | |
| 17 | OpenAI | Feb 2026 | 93.4% | Verified | |
| 18 | OpenAI | Dec 2025 | 93.4% | Estimated | |
| 19 | OpenAI | Mar 2026 | 93.2% | Estimated | |
| 20 | OpenAI | Nov 2025 | 93.2% | Estimated | |
| 21 | OpenAI | Jan 2026 | 93.0% | Estimated | |
| 22 | OpenAI | Oct 2025 | 93.0% | Estimated | |
| 23 | Alibaba | Sep 2026 | 92.8% | Verified | |
| 24 | Apr 2026 | 92.8% | Verified | ||
| 25 | Anthropic | Nov 2025 | 92.8% | Verified | |
| 26 | OpenAI | Nov 2025 | 92.8% | Estimated | |
| 27 | OpenAI | Jul 2026 | 92.6% | Estimated | |
| 28 | OpenAI | Jul 2026 | 92.6% | Estimated | |
| 29 | Anthropic | May 2026 | 92.6% | Estimated | |
| 30 | Anthropic | Feb 2026 | 92.5% | Verified | |
| 31 | Anthropic | May 2026 | 92.4% | Estimated | |
| 32 | xAI | Apr 2026 | 92.4% | Verified | |
| 33 | OpenAI | Nov 2025 | 92.4% | Estimated | |
| 34 | OpenAI | Apr 2025 | 92.4% | Estimated | |
| 35 | OpenAI | Feb 2025 | 92.4% | Estimated | |
| 36 | OpenAI | Apr 2026 | 92.2% | Estimated | |
| 37 | OpenAI | Dec 2025 | 92.2% | Estimated | |
| 38 | OpenAI | Aug 2025 | 92.1% | Verified | |
| 39 | OpenAI | Mar 2026 | 92.0% | Estimated | |
| 40 | Alibaba | Feb 2026 | 92.0% | Verified | |
| 41 | DeepSeek | May 2025 | 92.0% | Estimated | |
| 42 | OpenAI | Aug 2025 | 91.8% | Estimated | |
| 43 | OpenAI | Dec 2024 | 91.8% | Verified | |
| 44 | OpenAI | Dec 2025 | 91.6% | Estimated | |
| 45 | Meta | May 2026 | 91.5% | Verified | |
| 46 | Alibaba | May 2026 | 91.5% | Verified | |
| 47 | Anthropic | Aug 2025 | 91.4% | Estimated | |
| 48 | Sep 2026 | 91.2% | Verified | ||
| 49 | OpenAI | Oct 2025 | 91.2% | Estimated | |
| 50 | OpenAI | Jul 2026 | 91.0% | Estimated | |
| 51 | OpenAI | Sep 2025 | 91.0% | Estimated | |
| 52 | OpenAI | Jul 2026 | 90.8% | Estimated | |
| 53 | OpenAI | Oct 2025 | 90.8% | Estimated | |
| 54 | OpenAI | Aug 2025 | 90.8% | Estimated | |
| 55 | DeepSeek | Jan 2025 | 90.8% | Verified | |
| 56 | DeepSeek | Jan 2025 | 90.8% | Verified | |
| 57 | OpenAI | Mar 2026 | 90.6% | Estimated | |
| 58 | OpenAI | Jun 2025 | 90.6% | Estimated | |
| 59 | Meta | Sep 2026 | 90.5% | Verified | |
| 60 | May 2026 | 90.5% | Verified | ||
| 61 | Anthropic | May 2025 | 90.5% | Verified | |
| 62 | OpenAI | Mar 2026 | 90.4% | Estimated | |
| 63 | Anthropic | Feb 2026 | 90.4% | Verified | |
| 64 | OpenAI | Dec 2025 | 90.4% | Estimated | |
| 65 | Jun 2025 | 89.9% | Verified | ||
| 66 | Zhipu AI | Jun 2026 | 89.5% | Verified | |
| 67 | Anthropic | Sep 2025 | 89.5% | Verified | |
| 68 | OpenAI | May 2024 | 88.7% | Verified | |
| 69 | Apr 2026 | 88.6% | Estimated | ||
| 70 | DeepSeek | Apr 2026 | 88.6% | Estimated | |
| 71 | Alibaba | Sep 2025 | 88.6% | Estimated | |
| 72 | Moonshot AI | Nov 2025 | 88.5% | Verified | |
| 73 | Alibaba | Aug 2026 | 88.4% | Estimated | |
| 74 | Feb 2026 | 88.4% | Estimated | ||
| 75 | MiniMax | Jan 2026 | 88.4% | Estimated | |
| 76 | Alibaba | Jun 2026 | 88.2% | Verified | |
| 77 | Apr 2026 | 88.2% | Estimated | ||
| 78 | ByteDance | Jan 2026 | 88.2% | Estimated | |
| 79 | Zhipu AI | Apr 2026 | 88.0% | Verified | |
| 80 | Anthropic | May 2025 | 88.0% | Verified | |
| 81 | OpenAI | Jan 2025 | 87.9% | Verified | |
| 82 | Sakana AI | Sep 2026 | 87.8% | Estimated | |
| 83 | Nvidia | Jun 2026 | 87.8% | Estimated | |
| 84 | Alibaba | Apr 2026 | 87.8% | Estimated | |
| 85 | AI21 Labs | Aug 2025 | 87.8% | Estimated | |
| 86 | Perplexity | Mar 2025 | 87.8% | Estimated | |
| 87 | Mistral | Dec 2025 | 87.4% | Estimated | |
| 88 | Inflection | Oct 2024 | 87.4% | Estimated | |
| 89 | Sakana AI | Sep 2026 | 87.2% | Estimated | |
| 90 | Meta | Apr 2025 | 87.2% | Verified | |
| 91 | MiniMax | Oct 2025 | 87.0% | Estimated | |
| 92 | Meta | Dec 2024 | 87.0% | Estimated | |
| 93 | Apr 2026 | 86.8% | Verified | ||
| 94 | Kuaishou | Mar 2026 | 86.8% | Estimated | |
| 95 | Feb 2026 | 86.8% | Estimated | ||
| 96 | Perplexity | Oct 2025 | 86.8% | Estimated | |
| 97 | Morph AI | Jul 2025 | 86.8% | Estimated | |
| 98 | Baidu | Jun 2025 | 86.8% | Verified | |
| 99 | MiniMax | Jun 2025 | 86.8% | Estimated | |
| 100 | Mistral | Jul 2024 | 86.8% | Verified | |
| 101 | Mistral | Feb 2024 | 86.8% | Estimated | |
| 102 | Xiaomi | Apr 2026 | 86.6% | Estimated | |
| 103 | Meta | Jul 2024 | 86.6% | Estimated | |
| 104 | Zhipu AI | Feb 2026 | 86.5% | Verified | |
| 105 | OpenAI | Aug 2025 | 86.5% | Verified | |
| 106 | Inception AI | Aug 2026 | 86.4% | Verified | |
| 107 | Upstage | Aug 2026 | 86.4% | Estimated | |
| 108 | Jun 2025 | 86.4% | Estimated | ||
| 109 | May 2025 | 86.4% | Estimated | ||
| 110 | OpenAI | Apr 2024 | 86.4% | Verified | |
| 111 | MiniMax | Feb 2026 | 86.2% | Estimated | |
| 112 | Upstage | Jan 2026 | 86.2% | Estimated | |
| 113 | Meta | Dec 2024 | 86.2% | Verified | |
| 114 | Kuaishou | Jul 2026 | 86.0% | Estimated | |
| 115 | Mar 2026 | 86.0% | Estimated | ||
| 116 | Tencent | Nov 2025 | 86.0% | Estimated | |
| 117 | Perplexity | Mar 2025 | 86.0% | Estimated | |
| 118 | Amazon | Dec 2024 | 86.0% | Estimated | |
| 119 | Nous Research | Aug 2024 | 86.0% | Estimated | |
| 120 | DeepSeek | Aug 2026 | 85.8% | Estimated | |
| 121 | Jun 2026 | 85.8% | Estimated | ||
| 122 | Nous Research | Aug 2025 | 85.8% | Estimated | |
| 123 | Mistral | Nov 2024 | 85.8% | Estimated | |
| 124 | Nex AGI | Jun 2026 | 85.6% | Estimated | |
| 125 | Nov 2025 | 85.6% | Estimated | ||
| 126 | Sakana AI | Jun 2026 | 85.4% | Estimated | |
| 127 | MiniMax | Dec 2025 | 85.4% | Estimated | |
| 128 | Mistral | Jul 2025 | 85.4% | Estimated | |
| 129 | Alibaba | Sep 2024 | 85.3% | Verified | |
| 130 | DeepSeek | Jan 2025 | 85.2% | Verified | |
| 131 | Moonshot AI | Jun 2026 | 85.0% | Verified | |
| 132 | Anthropic | Oct 2025 | 84.8% | Verified | |
| 133 | StepFun | May 2026 | 84.5% | Verified | |
| 134 | Alibaba | Sep 2025 | 84.5% | Verified | |
| 135 | Moonshot AI | Apr 2026 | 84.2% | Verified | |
| 136 | Moonshot AI | Apr 2026 | 84.2% | Verified | |
| 137 | MiniMax | May 2026 | 84.0% | Verified | |
| 138 | Tencent | Apr 2026 | 83.5% | Verified | |
| 139 | Zhipu AI | Mar 2026 | 83.0% | Verified | |
| 140 | MiniMax | Mar 2026 | 82.5% | Verified | |
| 141 | Moonshot AI | Jan 2026 | 82.5% | Verified | |
| 142 | DeepSeek | Apr 2026 | 82.4% | Verified | |
| 143 | OpenAI | Jul 2024 | 82.0% | Verified | |
| 144 | Zhipu AI | Aug 2026 | 81.6% | Estimated | |
| 145 | Kuaishou | Jul 2026 | 81.6% | Estimated | |
| 146 | Anthropic | Apr 2026 | 81.6% | Estimated | |
| 147 | Zhipu AI | Apr 2026 | 81.6% | Estimated | |
| 148 | ByteDance | Dec 2025 | 81.6% | Estimated | |
| 149 | Alibaba | Sep 2025 | 81.6% | Estimated | |
| 150 | Morph AI | Jul 2025 | 81.6% | Estimated | |
| 151 | Meta | Apr 2025 | 81.6% | Estimated | |
| 152 | Mistral | Dec 2024 | 81.6% | Estimated | |
| 153 | Zhipu AI | Dec 2025 | 81.5% | Verified | |
| 154 | Meta | Aug 2026 | 81.4% | Estimated | |
| 155 | OpenRouter | Jun 2026 | 81.4% | Estimated | |
| 156 | Nvidia | Jun 2026 | 81.4% | Estimated | |
| 157 | Xiaomi | Apr 2026 | 81.4% | Estimated | |
| 158 | Alibaba | Feb 2026 | 81.4% | Estimated | |
| 159 | Inflection | Oct 2024 | 81.4% | Estimated | |
| 160 | Meta | Sep 2026 | 81.2% | Estimated | |
| 161 | Anthropic | Jul 2026 | 81.2% | Estimated | |
| 162 | Meta | Jul 2026 | 81.2% | Estimated | |
| 163 | Poolside | Jul 2026 | 81.2% | Estimated | |
| 164 | Anthropic | Jun 2026 | 81.2% | Estimated | |
| 165 | Poolside | Apr 2026 | 81.2% | Estimated | |
| 166 | Amazon | Oct 2025 | 81.2% | Estimated | |
| 167 | Alibaba | Sep 2025 | 81.2% | Estimated | |
| 168 | Mistral | Jan 2025 | 81.2% | Verified | |
| 169 | Alibaba | Nov 2024 | 81.2% | Verified | |
| 170 | Meta | Sep 2024 | 81.2% | Estimated | |
| 171 | Cohere | Aug 2024 | 81.2% | Estimated | |
| 172 | OpenAI | Jan 2024 | 81.2% | Estimated | |
| 173 | Meta | Aug 2026 | 81.0% | Estimated | |
| 174 | Alibaba | Aug 2026 | 81.0% | Estimated | |
| 175 | Anthropic | Jul 2026 | 81.0% | Estimated | |
| 176 | Poolside | Jul 2026 | 81.0% | Estimated | |
| 177 | Mistral | Apr 2026 | 81.0% | Estimated | |
| 178 | xAI | Mar 2026 | 81.0% | Estimated | |
| 179 | Nvidia | Dec 2025 | 81.0% | Estimated | |
| 180 | DeepSeek | Sep 2025 | 81.0% | Estimated | |
| 181 | Alibaba | Sep 2025 | 81.0% | Estimated | |
| 182 | Moonshot AI | Sep 2025 | 81.0% | Estimated | |
| 183 | OpenAI | Apr 2025 | 81.0% | Estimated | |
| 184 | The Drummer | Mar 2025 | 81.0% | Estimated | |
| 185 | Alibaba | Feb 2025 | 81.0% | Verified | |
| 186 | MiniMax | Jan 2025 | 81.0% | Verified | |
| 187 | Nvidia | Aug 2026 | 80.8% | Estimated | |
| 188 | Sakana AI | Aug 2026 | 80.8% | Estimated | |
| 189 | OpenAI | May 2026 | 80.8% | Estimated | |
| 190 | xAI | Mar 2026 | 80.8% | Estimated | |
| 191 | Alibaba | Feb 2026 | 80.8% | Estimated | |
| 192 | OpenAI | Oct 2025 | 80.8% | Estimated | |
| 193 | 01.AI | Oct 2025 | 80.8% | Estimated | |
| 194 | Mar 2025 | 80.8% | Estimated | ||
| 195 | Alibaba | Mar 2026 | 80.6% | Estimated | |
| 196 | Alibaba | Oct 2025 | 80.6% | Estimated | |
| 197 | Alibaba | Sep 2025 | 80.6% | Estimated | |
| 198 | Cohere | Mar 2025 | 80.6% | Estimated | |
| 199 | Alibaba | Apr 2026 | 80.4% | Estimated | |
| 200 | Alibaba | Apr 2026 | 80.4% | Estimated | |
| 201 | Apr 2026 | 80.4% | Estimated | ||
| 202 | Alibaba | Feb 2026 | 80.4% | Estimated | |
| 203 | OpenRouter | Dec 2025 | 80.4% | Estimated | |
| 204 | Alibaba | Jul 2025 | 80.4% | Estimated | |
| 205 | Zhipu AI | Jul 2025 | 80.4% | Estimated | |
| 206 | Alibaba | Apr 2025 | 80.4% | Estimated | |
| 207 | OpenAI | Apr 2025 | 80.4% | Estimated | |
| 208 | Perplexity | Mar 2025 | 80.4% | Estimated | |
| 209 | OpenAI | Aug 2024 | 80.4% | Estimated | |
| 210 | Anthropic | Apr 2026 | 80.2% | Estimated | |
| 211 | Alibaba | Feb 2026 | 80.2% | Estimated | |
| 212 | DeepSeek | Dec 2025 | 80.2% | Estimated | |
| 213 | Alibaba | Oct 2025 | 80.2% | Estimated | |
| 214 | Alibaba | Oct 2025 | 80.2% | Estimated | |
| 215 | Alibaba | Sep 2025 | 80.2% | Estimated | |
| 216 | Mistral | Aug 2025 | 80.2% | Estimated | |
| 217 | OpenAI | Jan 2024 | 80.2% | Estimated | |
| 218 | Tencent | Aug 2026 | 80.0% | Estimated | |
| 219 | ByteDance | Aug 2026 | 80.0% | Estimated | |
| 220 | Writer | Jan 2026 | 80.0% | Estimated | |
| 221 | Baidu | Aug 2025 | 80.0% | Estimated | |
| 222 | Alibaba | Aug 2025 | 80.0% | Estimated | |
| 223 | DeepSeek | Aug 2025 | 80.0% | Estimated | |
| 224 | Alibaba | Jul 2025 | 80.0% | Estimated | |
| 225 | Alibaba | Jul 2025 | 80.0% | Estimated | |
| 226 | Moonshot AI | Jul 2025 | 80.0% | Estimated | |
| 227 | Alibaba | Apr 2025 | 80.0% | Estimated | |
| 228 | Alibaba | Apr 2025 | 80.0% | Estimated | |
| 229 | Alibaba | Apr 2025 | 80.0% | Estimated | |
| 230 | Mar 2025 | 80.0% | Estimated | ||
| 231 | Alibaba | Feb 2025 | 80.0% | Estimated | |
| 232 | OpenRouter | Apr 2026 | 79.8% | Estimated | |
| 233 | OpenRouter | Feb 2026 | 79.8% | Estimated | |
| 234 | Zhipu AI | Sep 2025 | 79.8% | Estimated | |
| 235 | OpenAI | Aug 2025 | 79.8% | Estimated | |
| 236 | Mistral | Aug 2025 | 79.8% | Estimated | |
| 237 | Meta | Apr 2025 | 79.8% | Estimated | |
| 238 | DeepSeek | Mar 2025 | 79.8% | Estimated | |
| 239 | OpenAI | Mar 2025 | 79.8% | Estimated | |
| 240 | OpenAI | Sep 2023 | 79.8% | Estimated | |
| 241 | xAI | May 2026 | 79.6% | Estimated | |
| 242 | Anthropic | Apr 2026 | 79.6% | Estimated | |
| 243 | OpenAI | Jan 2026 | 79.6% | Estimated | |
| 244 | Zhipu AI | Dec 2025 | 79.6% | Estimated | |
| 245 | The Drummer | Sep 2025 | 79.6% | Estimated | |
| 246 | Alibaba | Jul 2025 | 79.6% | Estimated | |
| 247 | Zhipu AI | Jul 2025 | 79.6% | Estimated | |
| 248 | Meta | Sep 2024 | 79.6% | Estimated | |
| 249 | Tencent | Aug 2026 | 79.4% | Estimated | |
| 250 | Zhipu AI | Aug 2026 | 79.4% | Estimated | |
| 251 | Tencent | Jul 2026 | 79.4% | Estimated | |
| 252 | Mar 2026 | 79.4% | Estimated | ||
| 253 | Mistral | Dec 2025 | 79.4% | Estimated | |
| 254 | DeepSeek | Sep 2025 | 79.4% | Estimated | |
| 255 | Alibaba | Sep 2025 | 79.4% | Estimated | |
| 256 | Zhipu AI | Aug 2025 | 79.4% | Estimated | |
| 257 | OpenAI | Aug 2025 | 79.4% | Estimated | |
| 258 | OpenAI | May 2024 | 79.4% | Estimated | |
| 259 | OpenAI | Aug 2023 | 79.4% | Estimated | |
| 260 | OpenAI | May 2023 | 79.4% | Estimated | |
| 261 | xAI | Sep 2026 | 79.2% | Estimated | |
| 262 | Inception AI | Mar 2026 | 79.2% | Estimated | |
| 263 | Anthropic | Feb 2025 | 79.2% | Estimated | |
| 264 | Perplexity | Jan 2025 | 79.2% | Estimated | |
| 265 | The Drummer | Nov 2024 | 79.2% | Estimated | |
| 266 | Meta | Sep 2024 | 79.2% | Estimated | |
| 267 | OpenRouter | Nov 2023 | 79.2% | Estimated | |
| 268 | OpenAI | May 2023 | 79.2% | Estimated | |
| 269 | xAI | Aug 2026 | 79.0% | Estimated | |
| 270 | Moonshot AI | Apr 2026 | 79.0% | Estimated | |
| 271 | Alibaba | Apr 2026 | 79.0% | Estimated | |
| 272 | Reka AI | Mar 2026 | 79.0% | Estimated | |
| 273 | May 2025 | 79.0% | Estimated | ||
| 274 | Meta | Jul 2024 | 79.0% | Estimated | |
| 275 | OpenRouter | Jul 2026 | 78.8% | Estimated | |
| 276 | xAI | Jul 2026 | 78.8% | Estimated | |
| 277 | Anthropic | Jun 2026 | 78.8% | Estimated | |
| 278 | Anthropic | Apr 2026 | 78.8% | Estimated | |
| 279 | Anthropic | Apr 2026 | 78.8% | Estimated | |
| 280 | Nvidia | Mar 2026 | 78.8% | Estimated | |
| 281 | Alibaba | Jul 2025 | 78.8% | Estimated | |
| 282 | Nous Research | Aug 2024 | 78.8% | Estimated | |
| 283 | ByteDance | Aug 2026 | 78.6% | Estimated | |
| 284 | Thinking Machines | Jul 2026 | 78.6% | Estimated | |
| 285 | Moonshot AI | Jul 2026 | 78.6% | Estimated | |
| 286 | Anthropic | Jun 2026 | 78.6% | Estimated | |
| 287 | Alibaba | Sep 2025 | 78.6% | Estimated | |
| 288 | Nous Research | Aug 2025 | 78.6% | Estimated | |
| 289 | Mistral | Apr 2024 | 78.6% | Estimated | |
| 290 | Tencent | Jul 2025 | 78.5% | Verified | |
| 291 | Inception AI | Sep 2026 | 78.4% | Estimated | |
| 292 | Meta | Aug 2026 | 78.4% | Estimated | |
| 293 | Apr 2026 | 78.4% | Estimated | ||
| 294 | Alibaba | Nov 2025 | 78.4% | Estimated | |
| 295 | Mistral | May 2025 | 78.4% | Estimated | |
| 296 | Mistral | Feb 2025 | 78.4% | Estimated | |
| 297 | OpenAI | Nov 2024 | 78.4% | Estimated | |
| 298 | Mistral | Jul 2024 | 78.4% | Estimated | |
| 299 | Zhipu AI | Jan 2026 | 77.2% | Verified | |
| 300 | Alibaba | Oct 2024 | 75.8% | Verified | |
| 301 | Cohere | Apr 2024 | 75.7% | Verified | |
| 302 | Zhipu AI | Sep 2026 | 75.6% | Estimated | |
| 303 | IBM | Aug 2026 | 75.6% | Estimated | |
| 304 | Apr 2026 | 75.6% | Estimated | ||
| 305 | Mistral | Mar 2026 | 75.6% | Estimated | |
| 306 | Mar 2026 | 75.6% | Estimated | ||
| 307 | Reka AI | Mar 2025 | 75.6% | Estimated | |
| 308 | Feb 2025 | 75.6% | Estimated | ||
| 309 | Zhipu AI | Aug 2026 | 75.4% | Estimated | |
| 310 | Jun 2026 | 75.4% | Estimated | ||
| 311 | IBM | Apr 2026 | 75.4% | Estimated | |
| 312 | Alibaba | Oct 2025 | 75.4% | Estimated | |
| 313 | OpenAI | Oct 2025 | 75.4% | Estimated | |
| 314 | Mistral | Jun 2025 | 75.4% | Estimated | |
| 315 | Inclusion AI | Jul 2026 | 75.2% | Estimated | |
| 316 | Feb 2026 | 75.2% | Estimated | ||
| 317 | Alibaba | Apr 2025 | 75.2% | Estimated | |
| 318 | Mistral | Mar 2025 | 75.2% | Estimated | |
| 319 | Mar 2025 | 75.2% | Estimated | ||
| 320 | Anthropic | Mar 2024 | 75.2% | Verified | |
| 321 | Inclusion AI | Sep 2026 | 75.0% | Estimated | |
| 322 | Tencent | Aug 2026 | 75.0% | Estimated | |
| 323 | Jun 2026 | 75.0% | Estimated | ||
| 324 | Jul 2024 | 75.0% | Estimated | ||
| 325 | Microsoft | Apr 2024 | 74.8% | Estimated | |
| 326 | Alibaba | Aug 2026 | 74.6% | Estimated | |
| 327 | Alibaba | Aug 2026 | 74.6% | Estimated | |
| 328 | OpenAI | Apr 2025 | 74.6% | Estimated | |
| 329 | OpenAI | Apr 2025 | 74.6% | Estimated | |
| 330 | Microsoft | Jan 2025 | 74.6% | Estimated | |
| 331 | Amazon | Dec 2024 | 74.6% | Estimated | |
| 332 | Alibaba | Jul 2026 | 74.4% | Estimated | |
| 333 | ByteDance | Mar 2026 | 74.4% | Estimated | |
| 334 | Mistral | Oct 2025 | 74.4% | Estimated | |
| 335 | Meta | Jul 2024 | 74.4% | Estimated | |
| 336 | OpenAI | Jul 2024 | 74.4% | Estimated | |
| 337 | Zhipu AI | Aug 2026 | 74.2% | Estimated | |
| 338 | DeepSeek | Aug 2026 | 74.2% | Estimated | |
| 339 | Alibaba | Apr 2026 | 74.2% | Estimated | |
| 340 | Alibaba | Apr 2026 | 74.2% | Estimated | |
| 341 | ByteDance | Feb 2026 | 74.2% | Estimated | |
| 342 | ByteDance | Dec 2025 | 74.2% | Estimated | |
| 343 | Mistral | Dec 2025 | 74.2% | Estimated | |
| 344 | OpenAI | Apr 2025 | 74.2% | Estimated | |
| 345 | Cohere | Dec 2024 | 74.2% | Estimated | |
| 346 | Nex AGI | Jun 2026 | 74.0% | Estimated | |
| 347 | Alibaba | Feb 2026 | 74.0% | Estimated | |
| 348 | Alibaba | Feb 2026 | 74.0% | Estimated | |
| 349 | Dec 2025 | 74.0% | Estimated | ||
| 350 | Oct 2025 | 74.0% | Estimated | ||
| 351 | Meta | Jul 2024 | 74.0% | Estimated | |
| 352 | Aug 2026 | 73.8% | Estimated | ||
| 353 | IBM | Oct 2025 | 73.8% | Estimated | |
| 354 | Sep 2025 | 73.8% | Estimated | ||
| 355 | Tencent | Aug 2026 | 73.6% | Estimated | |
| 356 | Jul 2026 | 73.6% | Estimated | ||
| 357 | Mistral | Oct 2024 | 73.6% | Estimated | |
| 358 | ByteDance | Jul 2025 | 73.5% | Verified | |
| 359 | Thinking Machines | Jul 2026 | 73.4% | Estimated | |
| 360 | OpenAI | Mar 2025 | 73.4% | Estimated | |
| 361 | Jul 2026 | 73.2% | Estimated | ||
| 362 | OpenAI | Jan 2026 | 73.2% | Estimated | |
| 363 | Amazon | Dec 2025 | 73.2% | Estimated | |
| 364 | Mistral | Dec 2025 | 73.2% | Estimated | |
| 365 | Alibaba | Sep 2025 | 73.2% | Estimated | |
| 366 | Jun 2025 | 73.2% | Estimated | ||
| 367 | DeepSeek | Jul 2026 | 73.0% | Estimated | |
| 368 | StepFun | Jan 2026 | 73.0% | Estimated | |
| 369 | Mistral | Dec 2025 | 73.0% | Estimated | |
| 370 | Jul 2025 | 73.0% | Estimated | ||
| 371 | Meta | Sep 2024 | 73.0% | Verified | |
| 372 | DeepSeek | Sep 2026 | 72.8% | Estimated | |
| 373 | Inclusion AI | Aug 2026 | 72.8% | Estimated | |
| 374 | Alibaba | Feb 2026 | 72.8% | Estimated | |
| 375 | Apr 2026 | 72.6% | Estimated | ||
| 376 | DeepSeek | Apr 2026 | 72.6% | Estimated | |
| 377 | Meta | Apr 2024 | 72.6% | Estimated | |
| 378 | May 2026 | 72.4% | Estimated | ||
| 379 | Alibaba | Oct 2025 | 72.4% | Estimated | |
| 380 | Amazon | Dec 2024 | 72.4% | Estimated | |
| 381 | Cohere | Mar 2024 | 71.0% | Verified |
About MMLU
Massive Multitask Language Understanding (MMLU) measures general knowledge across 57 academic subjects from elementary math to professional law. It is the industry standard for evaluating general reasoning and semantic understanding.
What do these benchmarks mean?
What is the MMLU benchmark?keyboard_arrow_down
Massive Multitask Language Understanding (MMLU) measures general knowledge across 57 academic subjects from elementary math to professional law. It is the industry standard for evaluating general reasoning and semantic understanding.
MMLU consists of multiple-choice questions covering humanities, social sciences, STEM, and other professional contexts. It tests both world knowledge and problem-solving capability. Strong performance indicates a highly versatile model capable of handling diverse tasks without specialized fine-tuning.
What is the HumanEval benchmark?keyboard_arrow_down
HumanEval is a programming benchmark created by OpenAI to evaluate coding abilities. It measures the accuracy of models in generating functional Python code blocks based on docstring instructions.
HumanEval consists of 164 hand-written programming problems. Models are evaluated using pass@1 metrics, meaning the code is executed against automated unit tests and must pass on the first try. High scores correspond directly to software engineering utility, agentic code writing, and syntactical precision.
What is the MATH benchmark?keyboard_arrow_down
MATH measures mathematical problem-solving skills across seven high-school and college disciplines. It requires models to perform complex, multi-step symbolic reasoning rather than simple calculation.
MATH is exceptionally difficult for standard models. It covers algebra, calculus, probability, geometry, and number theory. Unlike multiple-choice benchmarks, MATH requires generating final equations or numerical answers, testing a model's chain-of-thought planning and logical precision.
What is the MT-Bench benchmark?keyboard_arrow_down
MT-Bench is a multi-turn conversation benchmark. It evaluates how well models maintain coherence, logic, and instructions across progressive dialogue exchanges.
MT-Bench tests eight categories of tasks including coding, math, roleplay, and writing. A powerful model like GPT-5 is utilized as a judge to grade the responses on a scale from 1 to 10. High performance indicates excellent instruction following and conversational context retention over long interactions.
What is the GPQA benchmark?keyboard_arrow_down
GPQA (Graduate-Level Google-Proof Q&A Benchmark) tests advanced scientific and mathematical understanding using questions designed by PhD-level experts. The questions are specifically written to be difficult to answer via search engines.
GPQA contains physics, biology, and chemistry questions that even human experts find challenging. Non-experts with access to Google search only score around 34%, while models must exhibit advanced abstract reasoning and scientific understanding to pass, serving as a key benchmark for expert-level capability.
What is the HellaSwag benchmark?keyboard_arrow_down
HellaSwag evaluates common-sense reasoning and situational prediction. It tests whether a model can accurately determine the most likely next event in a described physical scenario.
HellaSwag is designed using adversarial filtering to find scenarios that are easy for humans (who score ~95%) but difficult for language models. It requires deep contextual understanding of everyday physics, human intent, and spatial logic.