LLM benchmark leaderboards
EvalsAs of August 31, 2026, this index ranks 768 public LLM evals next to live API token prices. Benchmark names on model and comparison pages link here. Official boards own the methodology; these pages add $/M.
Evals with the most published scores
Coverage only — not search demand. The table below lists every public eval.
| Benchmark | Models scored | What it measures |
|---|---|---|
| HLE | 116 | Humanity's Last Exam: 2,500 expert questions designed so search engines fail. |
| SWE Bench | 8 | the SWE Bench score |
| Humanity's Last Exam (No Tools) | 5 | the Humanity's Last Exam (No Tools) score |
| Humanity's Last Exam (With Tools) | 5 | the Humanity's Last Exam (With Tools) score |
| Terminal-Bench 2.0 | 71 | the Terminal-Bench 2.0 score |
| Terminal-Bench | 70 | Terminal-Bench: can the model finish jobs in a real shell. |
| Terminal-Bench 2.1 | 55 | Terminal-Bench: can the model finish jobs in a real shell. |
| Terminal-Bench 2 | 48 | the Terminal-Bench 2 score |
| Terminal-Bench 3.0 | 3 | the Terminal-Bench 3.0 score |
| ARC-AGI-1 Verified | 74 | the ARC-AGI-1 Verified score |
| Chatbot Arena | 16 | Chatbot Arena human votes: people pick the better reply in a blind test. That is a writing and taste rank, not a trivia quiz. |
| ARC-AGI-3 | 10 | the ARC-AGI-3 score |
| Arc Agi | 8 | the Arc Agi score |
| DeepSWE | 35 | DeepSWE: repository-level coding agents on real GitHub work. |
| SWE-bench Verified | 148 | the SWE-bench Verified score |
| SWE-bench Pro | 57 | the SWE-bench Pro score |
| SWE-bench Pro (Public) | 10 | the SWE-bench Pro (Public) score |
| MMLU | 106 | MMLU, a multiple-choice knowledge exam. It does not measure prose. |
| ARC-AGI 2 | 71 | ARC-AGI: novel visual puzzles, not memorized exams. |
| LiveBench | 59 | the LiveBench score |
| OSWorld | 41 | OSWorld: can the model drive a real desktop (click, type, finish the task). |
| OSWorld Verified | 37 | OSWorld-Verified: cleaned desktop computer-use tasks with an execution check. |
| GDPval | 17 | GDPval: professional knowledge-work, closer to office output than trivia. |
| Screenspot | 16 | the Screenspot score |
| OSWorld 2.0 | 14 | the OSWorld 2.0 score |
| ARC-AGI-2 Verified | 2 | the ARC-AGI-2 Verified score |
| ARC-AGI 2 | 1 | the ARC-AGI 2 score |
| SimpleBench | 60 | the SimpleBench score |
| GSM8K | 50 | the GSM8K score |
| Treebench | 2 | the Treebench score |
| LiveCodeBench | 138 | LiveCodeBench: contest programming on problems released after training. |
| LiveCodeBench v6 | 58 | the LiveCodeBench v6 score |
| tau-bench Retail | 48 | tau-bench: multi-turn tool use on customer-support workflows. |
| tau2-bench (Telecom) | 35 | the tau2-bench (Telecom) score |
| tau2-bench Telecom | 34 | the tau2-bench Telecom score |
| Tau Bench | 6 | the Tau Bench score |
| GPQA | 225 | GPQA: graduate-level science questions that resist simple search. |
| BrowseComp | 62 | the BrowseComp score |
| Tvbench | 3 | the Tvbench score |
| Deck Bench | 1 | the Deck Bench score |
| AIME 2025 | 191 | AIME 2025: American invitational math contest problems. |
| GPQA Diamond | 131 | GPQA Diamond: graduate-level science questions written so Google cannot solve them. |
| Aime 2026 | 21 | the Aime 2026 score |
| HealthBench | 9 | HealthBench: medical conversation quality, not multiple-choice science. |
| MMLU-Pro | 194 | MMLU-Pro, a hard multiple-choice knowledge quiz. It is not a writing test. |
| HumanEval | 69 | the HumanEval score |
| MMMU | 111 | MMMU: college-level questions on diagrams, charts, and images. |
| FrontierMath | 54 | FrontierMath: research-level math that stays hard after models ace AIME. |
| CursorBench | 14 | the CursorBench score |
| Big Bench | 3 | the Big Bench score |
| Mle Bench | 2 | the Mle Bench score |
| SimpleQA | 92 | SimpleQA, which checks whether the model makes up facts. |
| IFEval | 64 | IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.' |
| MATH 500 | 50 | MATH-500: textbook and contest math problems. |
| SkillsBench | 30 | SkillsBench: can the agent actually use bundled skills and tools on multi-step tasks, not just chat about them. |
| HellaSwag | 29 | HellaSwag, a sentence-completion quiz, not a writing sample. |
| Mt Bench | 12 | the Mt Bench score |
| Omnidocbench | 2 | the Omnidocbench score |
| AIME 2024 | 47 | the AIME 2024 score |
| Pinchbench | 6 | the Pinchbench score |
| Vending Bench 2 | 4 | the Vending Bench 2 score |
| Eq Bench | 2 | the Eq Bench score |
| MMMU-Pro | 67 | MMMU-Pro: a harder multimodal exam than MMMU. |
| IFBench | 36 | IFBench, a harder instruction-following set than IFEval. |
| ScreenSpot Pro | 24 | the ScreenSpot Pro score |
| Mrcr | 7 | the Mrcr score |
| Paperbench | 3 | the Paperbench score |
| Cybench | 2 | the Cybench score |
| DeepResearch Bench | 20 | the DeepResearch Bench score |
| Mmbench | 9 | the Mmbench score |
| Triviaqa | 9 | the Triviaqa score |
| Gdpval Aa | 6 | the Gdpval Aa score |
| TriviaQA | 2 | the TriviaQA score |
| SciCode | 91 | the SciCode score |
| MathVista | 43 | the MathVista score |
| ToolAthlon | 42 | the ToolAthlon score |
| Ocrbench | 24 | the Ocrbench score |
| Longbench V2 | 17 | the Longbench V2 score |
| Program Bench | 6 | the Program Bench score |
| Livecodebench Pro | 3 | the Livecodebench Pro score |
| Bigcodebench | 2 | the Bigcodebench score |
| Cl Bench | 2 | the Cl Bench score |
| Bixbench | 1 | the Bixbench score |
| MMMLU | 47 | the MMMLU score |
| SuperGPQA | 34 | the SuperGPQA score |
| Mvbench | 19 | the Mvbench score |
| Claw Eval | 14 | the Claw Eval score |
| Posttrainbench | 6 | the Posttrainbench score |
| Exploitbench | 5 | the Exploitbench score |
| Swe Marathon | 4 | the Swe Marathon score |
| Air Bench | 1 | the Air Bench score |
| SWE-bench Multilingual | 42 | the SWE-bench Multilingual score |
| LVBench | 27 | the LVBench score |
| MRCR v2 | 26 | MRCR: can the model still find planted facts inside a long prompt. |
| WinoGrande | 22 | the WinoGrande score |
| Evalplus | 4 | the Evalplus score |
| Longvideobench | 4 | the Longvideobench score |
| FrontierMath Tier 4 | 40 | the FrontierMath Tier 4 score |
| Fiction.liveBench | 37 | the Fiction.liveBench score |
| BBH | 16 | the BBH score |
| Big Bench Hard | 16 | the Big Bench Hard score |
| HealthBench Professional | 10 | the HealthBench Professional score |
| Finance Agent | 8 | the Finance Agent score |
| Zerobench | 8 | the Zerobench score |
| Ocrbench V2 | 7 | the Ocrbench V2 score |
| Multi Swe Bench | 6 | the Multi Swe Bench score |
| Swe Lancer | 4 | the Swe Lancer score |
| Biomysterybench | 3 | the Biomysterybench score |
| Simpleqa Verified | 2 | the Simpleqa Verified score |
| Swe Atlas | 2 | the Swe Atlas score |
| Deep Swe | 1 | the Deep Swe score |
| Olympiadbench | 1 | the Olympiadbench score |
| MMLU-Redux | 45 | the MMLU-Redux score |
| Writingbench | 15 | the Writingbench score |
| Legal Agent Benchmark | 14 | the Legal Agent Benchmark score |
| Mmlongbench Doc | 5 | the Mmlongbench Doc score |
| Tau3 Bench | 5 | the Tau3 Bench score |
| Wildclawbench | 5 | the Wildclawbench score |
| SWE-bench Multimodal | 4 | the SWE-bench Multimodal score |
| Acebench | 2 | the Acebench score |
| Terminal Bench Hard | 2 | the Terminal Bench Hard score |
| Web Bench | 2 | the Web Bench score |
| Omnibench | 1 | the Omnibench score |
| ProofBench | 52 | the ProofBench score |
| CyberBench | 36 | the CyberBench score |
| VideoMMMU | 26 | the VideoMMMU score |
| T2 Bench | 21 | the T2 Bench score |
| C Eval | 18 | the C Eval score |
| AutomationBench | 15 | the AutomationBench score |
| Graphwalks | 3 | the Graphwalks score |
| Lifescibench | 3 | the Lifescibench score |
| Cfeval | 2 | the Cfeval score |
| Genebench | 2 | the Genebench score |
| Worldbench | 2 | the Worldbench score |
| Automation Bench | 1 | the Automation Bench score |
| Big Bench Audio | 1 | the Big Bench Audio score |
| Repobench | 1 | the Repobench score |
| Swt Bench | 1 | the Swt Bench score |
| Webdev Arena | 1 | the Webdev Arena score |
| APEX-Agents | 51 | the APEX-Agents score |
| Browsecomp Zh | 12 | the Browsecomp Zh score |
| HealthBench Hard | 8 | the HealthBench Hard score |
| Job Bench | 6 | the Job Bench score |
| Agieval | 5 | the Agieval score |
| Blueprint Bench 2 | 3 | the Blueprint Bench 2 score |
| Genebench Pro | 3 | the Genebench Pro score |
| Global Mmlu | 3 | the Global Mmlu score |
| Motionbench | 3 | the Motionbench score |
| Ovobench | 2 | the Ovobench score |
| Labbench2 | 1 | the Labbench2 score |
| Osworld G | 1 | the Osworld G score |
| Phybench | 1 | the Phybench score |
| Yc Bench | 1 | the Yc Bench score |
| Mmlu Prox | 30 | the Mmlu Prox score |
| Muirbench | 12 | the Muirbench score |
| Big Bench Extra Hard | 11 | the Big Bench Extra Hard score |
| Ojbench | 9 | the Ojbench score |
| Complexfuncbench | 7 | the Complexfuncbench score |
| Countbench | 7 | the Countbench score |
| Cmmlu | 6 | the Cmmlu score |
| Benchcad | 4 | the Benchcad score |
| Zclawbench | 4 | the Zclawbench score |
| Gdpval Aa V2 | 2 | the Gdpval Aa V2 score |
| Bankertoolbench | 1 | the Bankertoolbench score |
| Posttrain Bench | 1 | the Posttrain Bench score |
| Vibe Code Bench | 59 | the Vibe Code Bench score |
| MCP Atlas | 38 | the MCP Atlas score |
| Imo Answerbench | 20 | the Imo Answerbench score |
| AA-LCR | 18 | long-context retrieval: can the model use a huge window, not just advertise one. |
| Vita Bench | 10 | the Vita Bench score |
| Tir Bench | 4 | the Tir Bench score |
| Sec Bench Pro | 3 | the Sec Bench Pro score |
| Workspace Bench | 3 | the Workspace Bench score |
| Longcodebench | 2 | the Longcodebench score |
| Cmt Benchmark | 1 | the Cmt Benchmark score |
| Frontier-Bench v0.1 | 1 | the Frontier-Bench v0.1 score |
| Frontier Swe | 1 | the Frontier Swe score |
| Gdpval Aa Elo | 1 | the Gdpval Aa Elo score |
| Harvey Lab | 1 | the Harvey Lab score |
| Livesqlbench | 1 | the Livesqlbench score |
| Mle Bench Lite | 1 | the Mle Bench Lite score |
| Profbench | 1 | the Profbench score |
| Qwenclawbench | 1 | the Qwenclawbench score |
| Swe Perf | 1 | the Swe Perf score |
| DeepSWE 1.1 | 31 | DeepSWE 1.1: repository-level coding agents on real GitHub work. |
| Arc C | 28 | the Arc C score |
| Arena-Hard | 23 | the Arena-Hard score |
| GDP-PDF | 23 | the GDP-PDF score |
| Tau Bench Airline | 23 | the Tau Bench Airline score |
| Bfcl V3 | 19 | the Bfcl V3 score |
| Hallusion Bench | 18 | the Hallusion Bench score |
| Video Mme | 17 | the Video Mme score |
| Global Mmlu Lite | 13 | the Global Mmlu Lite score |
| Vibe Eval | 8 | the Vibe Eval score |
| Video-MME | 8 | the Video-MME score |
| AndroidWorld | 7 | the AndroidWorld score |
| Livecodebench V5 | 7 | the Livecodebench V5 score |
| Graphwalks BFS (256K-1M) | 5 | the Graphwalks BFS (256K-1M) score |
| Alignbench | 4 | the Alignbench score |
| FrontierMath (Tiers 1-3) | 4 | the FrontierMath (Tiers 1-3) score |
| Mmt Bench | 4 | the Mmt Bench score |
| Big Finance Bench | 3 | the Big Finance Bench score |
| Mmbench Video | 3 | the Mmbench Video score |
| Lmarena Text | 2 | the Lmarena Text score |
| Multilingual Mmlu | 2 | the Multilingual Mmlu score |
| Presentbench | 2 | the Presentbench score |
| Visualwebbench | 2 | the Visualwebbench score |
| Amo Bench | 1 | the Amo Bench score |
| Browsecomp Vl | 1 | the Browsecomp Vl score |
| Cyberseceval 4 | 1 | the Cyberseceval 4 score |
| Deepswe V1 1 | 1 | the Deepswe V1 1 score |
| Humaneval Plus | 1 | the Humaneval Plus score |
| Kernelbench Hard | 1 | the Kernelbench Hard score |
| Octocodingbench | 1 | the Octocodingbench score |
| Researchclawbench | 1 | the Researchclawbench score |
| Svg Bench | 1 | the Svg Bench score |
| Swe Fficiency | 1 | the Swe Fficiency score |
| HMMT 2025 | 31 | the HMMT 2025 score |
| RealWorldQA | 31 | the RealWorldQA score |
| tau2-bench Retail | 26 | the tau2-bench Retail score |
| Mathvista Mini | 24 | the Mathvista Mini score |
| Multi If | 23 | the Multi If score |
| tau2-bench Airline | 23 | the tau2-bench Airline score |
| Mmbench V1 1 | 20 | the Mmbench V1 1 score |
| Cc Ocr | 18 | the Cc Ocr score |
| Mm Mt Bench | 17 | the Mm Mt Bench score |
| Arena Hard V2 | 16 | the Arena Hard V2 score |
| Bfcl V4 | 15 | the Bfcl V4 score |
| Agents' Last Exam | 14 | the Agents' Last Exam score |
| Graphwalks BFS (0K-128K) | 14 | the Graphwalks BFS (0K-128K) score |
| Livebench 20241125 | 14 | the Livebench 20241125 score |
| Creative Writing V3 | 13 | the Creative Writing V3 score |
| Facts Grounding | 13 | the Facts Grounding score |
| Mmmu Val | 13 | the Mmmu Val score |
| Multipl E | 13 | the Multipl E score |
| Nova 63 | 11 | the Nova 63 score |
| OfficeQA Pro | 10 | the OfficeQA Pro score |
| Deep Planning | 9 | the Deep Planning score |
| The Agent Company | 9 | the The Agent Company score |
| Embspatialbench | 8 | the Embspatialbench score |
| Matharena Apex | 8 | the Matharena Apex score |
| Wild Bench | 8 | the Wild Bench score |
| CoWorkBench | 6 | the CoWorkBench score |
| Refspatialbench | 6 | the Refspatialbench score |
| Seal 0 | 6 | the Seal 0 score |
| Arc E | 4 | the Arc E score |
| HealthBench Consensus | 4 | the HealthBench Consensus score |
| Mmmuval | 4 | the Mmmuval score |
| Mrcr 1m | 4 | the Mrcr 1m score |
| Aa Briefcase | 3 | the Aa Briefcase score |
| Api Bank | 3 | the Api Bank score |
| Artifacts Bench | 3 | the Artifacts Bench score |
| Cnmo 2024 | 3 | the Cnmo 2024 score |
| Frontiermath Tier 4 V2 | 3 | the Frontiermath Tier 4 V2 score |
| Gdpval Mm | 3 | the Gdpval Mm score |
| Gorilla Benchmark Api Bench | 3 | the Gorilla Benchmark Api Bench score |
| Medchembench | 3 | the Medchembench score |
| Mls Bench Lite | 3 | the Mls Bench Lite score |
| Mmlu Cot | 3 | the Mmlu Cot score |
| Natural Questions | 3 | the Natural Questions score |
| Onemillion Bench | 3 | the Onemillion Bench score |
| Posttrainbench Lite | 3 | the Posttrainbench Lite score |
| Spreadsheetbench V1 | 3 | the Spreadsheetbench V1 score |
| Agent Startup Bench | 2 | the Agent Startup Bench score |
| Finance Agent V1 1 | 2 | the Finance Agent V1 1 score |
| Frontier Science Research | 2 | the Frontier Science Research score |
| Humaneval Mul | 2 | the Humaneval Mul score |
| Imo 2025 | 2 | the Imo 2025 score |
| Kimi Code Bench V2 | 2 | the Kimi Code Bench V2 score |
| Livemathematicianbench | 2 | the Livemathematicianbench score |
| Mbpp Evalplus | 2 | the Mbpp Evalplus score |
| Measurebench | 2 | the Measurebench score |
| Mimo Coding Bench | 2 | the Mimo Coding Bench score |
| Mmlu Stem | 2 | the Mmlu Stem score |
| Mmsibench | 2 | the Mmsibench score |
| Openai Mmlu | 2 | the Openai Mmlu score |
| Ovbench | 2 | the Ovbench score |
| Qwen SWE-bench | 2 | the Qwen SWE-bench score |
| Qwenwebbench | 2 | the Qwenwebbench score |
| Qwenworldbench | 2 | the Qwenworldbench score |
| Swe Atlas Codebase Qna | 2 | the Swe Atlas Codebase Qna score |
| Swe Bench Verified Agentic Coding | 2 | the Swe Bench Verified Agentic Coding score |
| Vibe Pro | 2 | the Vibe Pro score |
| Xdailybench | 2 | the Xdailybench score |
| Androidbench | 1 | the Androidbench score |
| Apex Swe | 1 | the Apex Swe score |
| Automationbench Aa | 1 | the Automationbench Aa score |
| Bigcodebench Hard | 1 | the Bigcodebench Hard score |
| Biolp Bench | 1 | the Biolp Bench score |
| Charxiv Reasoning | 1 | the Charxiv Reasoning score |
| Cl Bench Life | 1 | the Cl Bench Life score |
| Cruxeval O | 1 | the Cruxeval O score |
| Frontiercode Diamond | 1 | the Frontiercode Diamond score |
| Gdpval Aa V2 Elo | 1 | the Gdpval Aa V2 Elo score |
| Gdpval Rubrics | 1 | the Gdpval Rubrics score |
| Hle Full | 1 | the Hle Full score |
| Hr Bench 4k | 1 | the Hr Bench 4k score |
| Infinitebench En Mc | 1 | the Infinitebench En Mc score |
| Kernel Bench L3 | 1 | the Kernel Bench L3 score |
| Kimi Claw 24 7 Bench | 1 | the Kimi Claw 24 7 Bench score |
| Mcp Universe | 1 | the Mcp Universe score |
| Miabench | 1 | the Miabench score |
| Mm Clawbench | 1 | the Mm Clawbench score |
| Mmlu Chat | 1 | the Mmlu Chat score |
| Next.js Evals | 1 | the Next.js Evals score |
| Nih Multi Needle | 1 | the Nih Multi Needle score |
| Plawbench | 1 | the Plawbench score |
| Prbench Finance | 1 | the Prbench Finance score |
| Ruler 128k | 1 | the Ruler 128k score |
| Seccodebench | 1 | the Seccodebench score |
| Spreadsheetbench 2 | 1 | the Spreadsheetbench 2 score |
| Swe Atlas Test Writing | 1 | the Swe Atlas Test Writing score |
| SWE-MM | 1 | the SWE-MM score |
| Swe Review | 1 | the Swe Review score |
| Toolathlon Verified | 1 | the Toolathlon Verified score |
| Vibe V2 | 1 | the Vibe V2 score |
| Vladbench | 1 | the Vladbench score |
| WebArena Verified | 1 | the WebArena Verified score |
| MATH | 99 | the MATH score |
| Arena Code Elo | 86 | the Arena Code Elo score |
| CritPt | 84 | the CritPt score |
| Tax Eval v2 | 84 | the Tax Eval v2 score |
| CorpFin v2 | 80 | the CorpFin v2 score |
| MGSM | 63 | the MGSM score |
| MedCode | 60 | the MedCode score |
| MedScribe | 60 | the MedScribe score |
| SAGE | 55 | the SAGE score |
| CharXiv-R | 54 | CharXiv: can the model read scientific charts. |
| MedQA | 47 | the MedQA score |
| Alder Polyglot | 45 | the Alder Polyglot score |
| BFCL | 45 | the BFCL score |
| IOI | 45 | the IOI score |
| Code Migration | 43 | the Code Migration score |
| FinanceAgent v2 | 42 | FinanceAgent: financial analysis tasks, not a generic chat vibe. |
| EMB | 40 | the EMB score |
| FinanceAgent v1.1 | 40 | FinanceAgent: financial analysis tasks, not a generic chat vibe. |
| Arena Vision Elo | 37 | the Arena Vision Elo score |
| MathVision | 35 | the MathVision score |
| Ai2d | 33 | the Ai2d score |
| Mbpp | 32 | the Mbpp score |
| INCLUDE | 29 | the INCLUDE score |
| Lech Mazur Writing | 29 | the Lech Mazur writing eval, which scores prose quality instead of multiple-choice knowledge. |
| Multichallenge | 29 | the Multichallenge score |
| Docvqa | 27 | the Docvqa score |
| Chartqa | 26 | the Chartqa score |
| DROP | 25 | the DROP score |
| ERQA | 25 | the ERQA score |
| Hmmt25 | 24 | the Hmmt25 score |
| Mmstar | 24 | the Mmstar score |
| Arena Search Elo | 23 | the Arena Search Elo score |
| Polymath | 23 | the Polymath score |
| NL2Repo | 21 | the NL2Repo score |
| Wmt24 | 21 | the Wmt24 score |
| Graphwalks Bfs 128k | 18 | the Graphwalks Bfs 128k score |
| OmniDocBench 1.5 | 18 | the OmniDocBench 1.5 score |
| Charxiv D | 17 | the Charxiv D score |
| Truthfulqa | 17 | the Truthfulqa score |
| FrontierCode | 16 | the FrontierCode score |
| FrontierCode 1.1 | 16 | the FrontierCode 1.1 score |
| Frontierswe | 16 | the Frontierswe score |
| Odinw | 16 | the Odinw score |
| Textvqa | 16 | the Textvqa score |
| Blink | 15 | the Blink score |
| Codeforces | 15 | the Codeforces score |
| Graphwalks Parents 128k | 14 | the Graphwalks Parents 128k score |
| Ocrbench V2 En | 14 | the Ocrbench V2 En score |
| Cybergym | 13 | the Cybergym score |
| Global Piqa | 13 | the Global Piqa score |
| Simplevqa | 13 | the Simplevqa score |
| BabyVision | 12 | the BabyVision score |
| Charadessta | 12 | the Charadessta score |
| Hmmt Feb 26 | 12 | the Hmmt Feb 26 score |
| Infovqatest | 12 | the Infovqatest score |
| Docvqatest | 11 | the Docvqatest score |
| Hiddenmath | 11 | the Hiddenmath score |
| Maxife | 11 | the Maxife score |
| Medxpertqa | 11 | the Medxpertqa score |
| Ocrbench V2 Zh | 11 | the Ocrbench V2 Zh score |
| Aider Polyglot Edit | 10 | the Aider Polyglot Edit score |
| Collie | 10 | the Collie score |
| Infovqa | 10 | the Infovqa score |
| Mlvu | 10 | the Mlvu score |
| Videomme W O Sub | 10 | the Videomme W O Sub score |
| Videomme W Sub | 10 | the Videomme W Sub score |
| Widesearch | 10 | the Widesearch score |
| Egoschema | 9 | the Egoschema score |
| Openai Mrcr 2 Needle 128k | 9 | the Openai Mrcr 2 Needle 128k score |
| Refcoco Avg | 9 | the Refcoco Avg score |
| Androidworld Sr | 8 | the Androidworld Sr score |
| Csimpleqa | 8 | the Csimpleqa score |
| Deepsearchqa | 8 | the Deepsearchqa score |
| Mcp Mark | 8 | the Mcp Mark score |
| Mlvu M | 8 | the Mlvu M score |
| MMMU-Pro (With Tools) | 8 | the MMMU-Pro (With Tools) score |
| Natural2code | 8 | the Natural2code score |
| Zebralogic | 8 | the Zebralogic score |
| Arena Agent Elo | 7 | the Arena Agent Elo score |
| Artificial Analysis Index | 7 | the Artificial Analysis Index score |
| Bird Sql Dev | 7 | the Bird Sql Dev score |
| Dynamath | 7 | the Dynamath score |
| Piqa | 7 | the Piqa score |
| Tau3 Banking | 7 | the Tau3 Banking score |
| V Star | 7 | the V Star score |
| Boolq | 6 | the Boolq score |
| ClawEval-MM | 6 | the ClawEval-MM score |
| Eclektic | 6 | the Eclektic score |
| Fleurs | 6 | the Fleurs score |
| GRIND | 6 | the GRIND score |
| Mmvu | 6 | the Mmvu score |
| Swe Lancer Ic Diamond Subset | 6 | the Swe Lancer Ic Diamond Subset score |
| Theoremqa | 6 | the Theoremqa score |
| Beyond Aime | 5 | the Beyond Aime score |
| Bfcl V2 | 5 | the Bfcl V2 score |
| Browsecomp Long 128k | 5 | the Browsecomp Long 128k score |
| Openai Mrcr 2 Needle 1m | 5 | the Openai Mrcr 2 Needle 1m score |
| Openbookqa | 5 | the Openbookqa score |
| ScienceQA | 5 | the ScienceQA score |
| Social Iqa | 5 | the Social Iqa score |
| Squality | 5 | the Squality score |
| Vision2Web | 5 | the Vision2Web score |
| Worldvqa | 5 | the Worldvqa score |
| Zerobench Sub | 5 | the Zerobench Sub score |
| Aa Index | 4 | the Aa Index score |
| Aider | 4 | the Aider score |
| Covost2 | 4 | the Covost2 score |
| Exploitgym | 4 | the Exploitgym score |
| Graphwalks Parents (0K-128K) | 4 | the Graphwalks Parents (0K-128K) score |
| Hypersim | 4 | the Hypersim score |
| Lingoqa | 4 | the Lingoqa score |
| Mme | 4 | the Mme score |
| MMMU-Pro (No Tools) | 4 | the MMMU-Pro (No Tools) score |
| Mmmu Validation | 4 | the Mmmu Validation score |
| Nexus | 4 | the Nexus score |
| OmniDocBench NED | 4 | the OmniDocBench NED score |
| Ruler | 4 | the Ruler score |
| Slakevqa | 4 | the Slakevqa score |
| Sunrgbd | 4 | the Sunrgbd score |
| Vlmsareblind | 4 | the Vlmsareblind score |
| Wmt23 | 4 | the Wmt23 score |
| Xstest | 4 | the Xstest score |
| Aitz Em | 3 | the Aitz Em score |
| Alpacaeval 2 0 | 3 | the Alpacaeval 2 0 score |
| Amc 2022 23 | 3 | the Amc 2022 23 score |
| Android Control High Em | 3 | the Android Control High Em score |
| Android Control Low Em | 3 | the Android Control Low Em score |
| Benchcad With Python Tool | 3 | the Benchcad With Python Tool score |
| Capture The Flag Challenges | 3 | the Capture The Flag Challenges score |
| Cluewsc | 3 | the Cluewsc score |
| Corpusqa 1m | 3 | the Corpusqa 1m score |
| Crag | 3 | the Crag score |
| Cybersecurity Ctfs | 3 | the Cybersecurity Ctfs score |
| Figqa | 3 | the Figqa score |
| Finqa | 3 | the Finqa score |
| Frontierscience Olympiad | 3 | the Frontierscience Olympiad score |
| Frontierscience Research | 3 | the Frontierscience Research score |
| Fullstackbench En | 3 | the Fullstackbench En score |
| Fullstackbench Zh | 3 | the Fullstackbench Zh score |
| Graphwalks Bfs 1m | 3 | the Graphwalks Bfs 1m score |
| Horizonmath | 3 | the Horizonmath score |
| Humanitys Last Exam With Tools Text Only | 3 | the Humanitys Last Exam With Tools Text Only score |
| Kernelgen 1p | 3 | the Kernelgen 1p score |
| Management Consulting Tasks | 3 | the Management Consulting Tasks score |
| Math Cot | 3 | the Math Cot score |
| Mm Mind2web | 3 | the Mm Mind2web score |
| Mmau | 3 | the Mmau score |
| Mobileworld | 3 | the Mobileworld score |
| Mrcr V2 8 Needle 512k 1m | 3 | the Mrcr V2 8 Needle 512k 1m score |
| Multilingual Mgsm Cot | 3 | the Multilingual Mgsm Cot score |
| Multipl E Humaneval | 3 | the Multipl E Humaneval score |
| Multipl E Mbpp | 3 | the Multipl E Mbpp score |
| Nanogpt | 3 | the Nanogpt score |
| Nuscene | 3 | the Nuscene score |
| OfficeQA | 3 | the OfficeQA score |
| Omniscience | 3 | the Omniscience score |
| Openai Connectors | 3 | the Openai Connectors score |
| Openai Search Function Calling | 3 | the Openai Search Function Calling score |
| Pmc Vqa | 3 | the Pmc Vqa score |
| Pope | 3 | the Pope score |
| Qvhighlights | 3 | the Qvhighlights score |
| Realkie Fcc | 3 | the Realkie Fcc score |
| RecreationBench | 3 | the RecreationBench score |
| Rsi Index | 3 | the Rsi Index score |
| Superchem | 3 | the Superchem score |
| Translation En Set1 Comet22 | 3 | the Translation En Set1 Comet22 score |
| Translation En Set1 Spbleu | 3 | the Translation En Set1 Spbleu score |
| Translation Set1 En Comet22 | 3 | the Translation Set1 En Comet22 score |
| Translation Set1 En Spbleu | 3 | the Translation Set1 En Spbleu score |
| Usamo 2026 | 3 | the Usamo 2026 score |
| Videoholmes | 3 | the Videoholmes score |
| Visfactor | 3 | the Visfactor score |
| Visulogic | 3 | the Visulogic score |
| Vqav2 | 3 | the Vqav2 score |
| Vqav2 Val | 3 | the Vqav2 Val score |
| Advancedif | 2 | the Advancedif score |
| Aethercode | 2 | the Aethercode score |
| Apex | 2 | the Apex score |
| Arcagi2 | 2 | the Arcagi2 score |
| Arxivmath | 2 | the Arxivmath score |
| Attaq | 2 | the Attaq score |
| Autologi | 2 | the Autologi score |
| Bfcl V3 Multiturn | 2 | the Bfcl V3 Multiturn score |
| Browsecomp Long 256k | 2 | the Browsecomp Long 256k score |
| Chartography | 2 | the Chartography score |
| Chartqapro | 2 | the Chartqapro score |
| Charxiv Reasoning No Tools | 2 | the Charxiv Reasoning No Tools score |
| Charxiv Reasoning With Tools | 2 | the Charxiv Reasoning With Tools score |
| Codegolf V2 2 | 2 | the Codegolf V2 2 score |
| Contphy | 2 | the Contphy score |
| Creativework | 2 | the Creativework score |
| Crossvid | 2 | the Crossvid score |
| Design2code | 2 | the Design2code score |
| Doubao Multi Turn Bench | 2 | the Doubao Multi Turn Bench score |
| Dsbench Fullstack | 2 | the Dsbench Fullstack score |
| Dsbench Hard | 2 | the Dsbench Hard score |
| Dude | 2 | the Dude score |
| Emma | 2 | the Emma score |
| Factscore | 2 | the Factscore score |
| Frames | 2 | the Frames score |
| Frontiercode Diamond Xhigh | 2 | the Frontiercode Diamond Xhigh score |
| Frontiercs | 2 | the Frontiercs score |
| Functionalmath | 2 | the Functionalmath score |
| Gameworld | 2 | the Gameworld score |
| Gdm Mrcr V2 8needle 128k Average | 2 | the Gdm Mrcr V2 8needle 128k Average score |
| Gdm Mrcr V2 8needle 1m Pointwise | 2 | the Gdm Mrcr V2 8needle 1m Pointwise score |
| Gdp Pdf No Tools | 2 | the Gdp Pdf No Tools score |
| Govreport | 2 | the Govreport score |
| Graphwalks Parents (256K-1M) | 2 | the Graphwalks Parents (256K-1M) score |
| Groundui 1k | 2 | the Groundui 1k score |
| Gsm 8k Cot | 2 | the Gsm 8k Cot score |
| Humanitys Last Exam No Tools Text Only | 2 | the Humanitys Last Exam No Tools Text Only score |
| If | 2 | the If score |
| Image2floorplan | 2 | the Image2floorplan score |
| Intergps | 2 | the Intergps score |
| Kina | 2 | the Kina score |
| Livesports 3k | 2 | the Livesports 3k score |
| Mathverse | 2 | the Mathverse score |
| Mega Mlqa | 2 | the Mega Mlqa score |
| Mega Tydi Qa | 2 | the Mega Tydi Qa score |
| Mega Udpos | 2 | the Mega Udpos score |
| Mega Xcopa | 2 | the Mega Xcopa score |
| Mega Xstorycloze | 2 | the Mega Xstorycloze score |
| Minerva | 2 | the Minerva score |
| Mm If Eval | 2 | the Mm If Eval score |
| Mmlongbench 128k | 2 | the Mmlongbench 128k score |
| Mmsearch Plus | 2 | the Mmsearch Plus score |
| Mmvet | 2 | the Mmvet score |
| Mobileminiwob Sr | 2 | the Mobileminiwob Sr score |
| Mrcr 128k 8 Needle | 2 | the Mrcr 128k 8 Needle score |
| Msqa | 2 | the Msqa score |
| Multilf | 2 | the Multilf score |
| Musr | 2 | the Musr score |
| OpenAI MRCR v2 8-needle (128K-256K) | 2 | the OpenAI MRCR v2 8-needle (128K-256K) score |
| OpenAI MRCR v2 8-needle (128K-512K) | 2 | the OpenAI MRCR v2 8-needle (128K-512K) score |
| OpenAI MRCR v2 8-needle (256K-1M) | 2 | the OpenAI MRCR v2 8-needle (256K-1M) score |
| OpenAI MRCR v2 8-needle (32K-128K) | 2 | the OpenAI MRCR v2 8-needle (32K-128K) score |
| OpenAI MRCR v2 8-needle (4K-8K) | 2 | the OpenAI MRCR v2 8-needle (4K-8K) score |
| OpenAI MRCR v2 8-needle (512K-1M) | 2 | the OpenAI MRCR v2 8-needle (512K-1M) score |
| OpenAI MRCR v2 8-needle (64K-128K) | 2 | the OpenAI MRCR v2 8-needle (64K-128K) score |
| Perceptionbench | 2 | the Perceptionbench score |
| Perceptiontest | 2 | the Perceptiontest score |
| Physicsfinals | 2 | the Physicsfinals score |
| Polymath En | 2 | the Polymath En score |
| Popqa | 2 | the Popqa score |
| Qasper | 2 | the Qasper score |
| Qmsum | 2 | the Qmsum score |
| Qwen Svg | 2 | the Qwen Svg score |
| Repo Env | 2 | the Repo Env score |
| Repoqa | 2 | the Repoqa score |
| Seedclawbench | 2 | the Seedclawbench score |
| Spider | 2 | the Spider score |
| Summscreenfd | 2 | the Summscreenfd score |
| Swe Bench Verified Agentless | 2 | the Swe Bench Verified Agentless score |
| Tempcompass | 2 | the Tempcompass score |
| Tomato | 2 | the Tomato score |
| Trae Code Gen | 2 | the Trae Code Gen score |
| Trae Error Fix | 2 | the Trae Error Fix score |
| Tydiqa | 2 | the Tydiqa score |
| Usamo25 | 2 | the Usamo25 score |
| Vatex | 2 | the Vatex score |
| Videosimpleqa | 2 | the Videosimpleqa score |
| Vlmsarebiased | 2 | the Vlmsarebiased score |
| Voicebench Avg | 2 | the Voicebench Avg score |
| Aa Briefcase Elo | 1 | the Aa Briefcase Elo score |
| Aa Omniscience Index | 1 | the Aa Omniscience Index score |
| Activitynet | 1 | the Activitynet score |
| Advanced Cybersecurity Completion Rate | 1 | the Advanced Cybersecurity Completion Rate score |
| Ai2 Reasoning Challenge Arc | 1 | the Ai2 Reasoning Challenge Arc score |
| Aime | 1 | the Aime score |
| Arc | 1 | the Arc score |
| Arkitscenes | 1 | the Arkitscenes score |
| Babyvision With Python | 1 | the Babyvision With Python score |
| Bc Vl | 1 | the Bc Vl score |
| Beam 128k | 1 | the Beam 128k score |
| Bigcodebench Full | 1 | the Bigcodebench Full score |
| Biomysterybench Hard | 1 | the Biomysterybench Hard score |
| Biomysterybench Human Solved | 1 | the Biomysterybench Human Solved score |
| Cbnsl | 1 | the Cbnsl score |
| Cc Bench V2 Backend | 1 | the Cc Bench V2 Backend score |
| Cc Bench V2 Frontend | 1 | the Cc Bench V2 Frontend score |
| Cc Bench V2 Repo | 1 | the Cc Bench V2 Repo score |
| Chartmuseum | 1 | the Chartmuseum score |
| Charxiv Rq | 1 | the Charxiv Rq score |
| Charxiv Rq With Python | 1 | the Charxiv Rq With Python score |
| Chexpert Cxr | 1 | the Chexpert Cxr score |
| Ci Memories Coverage | 1 | the Ci Memories Coverage score |
| Ci Memories Violation | 1 | the Ci Memories Violation score |
| Cloningscenarios | 1 | the Cloningscenarios score |
| Cohere Agentic Question Answering | 1 | the Cohere Agentic Question Answering score |
| Cohere Data Analysis | 1 | the Cohere Data Analysis score |
| Cohere Memory Usage Quality | 1 | the Cohere Memory Usage Quality score |
| Common Voice 15 | 1 | the Common Voice 15 score |
| Commonsenseqa | 1 | the Commonsenseqa score |
| Corpusqa | 1 | the Corpusqa score |
| Countqa | 1 | the Countqa score |
| Covost2 En Zh | 1 | the Covost2 En Zh score |
| Crperelation | 1 | the Crperelation score |
| Crux O | 1 | the Crux O score |
| Cruxeval Input Cot | 1 | the Cruxeval Input Cot score |
| Cruxeval Output Cot | 1 | the Cruxeval Output Cot score |
| Cursorbench 3 2 | 1 | the Cursorbench 3 2 score |
| Dailyomni | 1 | the Dailyomni score |
| Deepsearchqa F1 | 1 | the Deepsearchqa F1 score |
| Deepswe 1 0 | 1 | the Deepswe 1 0 score |
| Deepswe 1 0 Pass At 1 | 1 | the Deepswe 1 0 Pass At 1 score |
| Deepswe 1 1 Mini Swe Agent | 1 | the Deepswe 1 1 Mini Swe Agent score |
| Dermmcqa | 1 | the Dermmcqa score |
| Draco | 1 | the Draco score |
| Ds Arena Code | 1 | the Ds Arena Code score |
| Ds Fim Eval | 1 | the Ds Fim Eval score |
| Exploitbench Cap Percent | 1 | the Exploitbench Cap Percent score |
| Finsearchcomp T2 T3 | 1 | the Finsearchcomp T2 T3 score |
| Finsearchcomp T3 | 1 | the Finsearchcomp T3 score |
| Flame Vlm Code | 1 | the Flame Vlm Code score |
| French Mmlu | 1 | the French Mmlu score |
| Frontier Science | 1 | the Frontier Science score |
| Frontier Swe Impl | 1 | the Frontier Swe Impl score |
| Frontiercode Main | 1 | the Frontiercode Main score |
| Gaia2 | 1 | the Gaia2 score |
| Gdp Pdf Mean Criteria No Tools | 1 | the Gdp Pdf Mean Criteria No Tools score |
| Gdp Pdf Mean Criteria With Tools | 1 | the Gdp Pdf Mean Criteria With Tools score |
| Gdp Pdf Strict Pass | 1 | the Gdp Pdf Strict Pass score |
| Giantsteps Tempo | 1 | the Giantsteps Tempo score |
| Gpqa Biology | 1 | the Gpqa Biology score |
| Gpqa Chemistry | 1 | the Gpqa Chemistry score |
| Gpqa Physics | 1 | the Gpqa Physics score |
| Gsm8k Chat | 1 | the Gsm8k Chat score |
| Harvey Lab Aa | 1 | the Harvey Lab Aa score |
| Hipho | 1 | the Hipho score |
| Hle Full With Tools | 1 | the Hle Full With Tools score |
| Hle Verified | 1 | the Hle Verified score |
| Humaneval Average | 1 | the Humaneval Average score |
| Humaneval Er | 1 | the Humaneval Er score |
| Humanevalfim Average | 1 | the Humanevalfim Average score |
| Imagemining | 1 | the Imagemining score |
| Imoproof Adv | 1 | the Imoproof Adv score |
| Infinitebench En Qa | 1 | the Infinitebench En Qa score |
| Infographicsqa | 1 | the Infographicsqa score |
| Instruct Humaneval | 1 | the Instruct Humaneval score |
| Ipho 2025 | 1 | the Ipho 2025 score |
| Lbpp V2 | 1 | the Lbpp V2 score |
| Livecodebench 01 09 | 1 | the Livecodebench 01 09 score |
| Livecodebench V5 24 12 25 2 | 1 | the Livecodebench V5 24 12 25 2 score |
| Loca Bench 256k | 1 | the Loca Bench 256k score |
| Longfact | 1 | the Longfact score |
| Longfact Concepts | 1 | the Longfact Concepts score |
| Longfact Objects | 1 | the Longfact Objects score |
| Lsat | 1 | the Lsat score |
| Mask | 1 | the Mask score |
| Mathverse Mini | 1 | the Mathverse Mini score |
| Mathvision With Python | 1 | the Mathvision With Python score |
| Maverix | 1 | the Maverix score |
| Mbpp Base Version | 1 | the Mbpp Base Version score |
| Mbpp Evalplus Base | 1 | the Mbpp Evalplus Base score |
| Mbpp Pass 1 | 1 | the Mbpp Pass 1 score |
| Mbpp Plus | 1 | the Mbpp Plus score |
| Medxpertqa Mm | 1 | the Medxpertqa Mm score |
| Meld | 1 | the Meld score |
| Mewc | 1 | the Mewc score |
| Mimic Cxr | 1 | the Mimic Cxr score |
| Mm Browsercomp | 1 | the Mm Browsercomp score |
| Mmau Music | 1 | the Mmau Music score |
| Mmau Sound | 1 | the Mmau Sound score |
| Mmau Speech | 1 | the Mmau Speech score |
| Mmbc | 1 | the Mmbc score |
| Mme Realworld | 1 | the Mme Realworld score |
| Mmlu Base | 1 | the Mmlu Base score |
| Mmlu French | 1 | the Mmlu French score |
| Mmlu Redux 2 0 | 1 | the Mmlu Redux 2 0 score |
| Mmmu Pro With Python | 1 | the Mmmu Pro With Python score |
| Mmsearch | 1 | the Mmsearch score |
| Mmvetgpt4turbo | 1 | the Mmvetgpt4turbo score |
| Mrcr 128k 2 Needle | 1 | the Mrcr 128k 2 Needle score |
| Mrcr 128k 4 Needle | 1 | the Mrcr 128k 4 Needle score |
| Mrcr 1m Pointwise | 1 | the Mrcr 1m Pointwise score |
| Mrcr 64k 2 Needle | 1 | the Mrcr 64k 2 Needle score |
| Mrcr 64k 4 Needle | 1 | the Mrcr 64k 4 Needle score |
| Mrcr 64k 8 Needle | 1 | the Mrcr 64k 8 Needle score |
| Mt Aime 2025 | 1 | the Mt Aime 2025 score |
| Mtvqa | 1 | the Mtvqa score |
| Musiccaps | 1 | the Musiccaps score |
| Next.js Evals + AGENTS.md | 1 | the Next.js Evals + AGENTS.md score |
| Nmos | 1 | the Nmos score |
| Nolima 128k | 1 | the Nolima 128k score |
| Nolima 32k | 1 | the Nolima 32k score |
| Nolima 64k | 1 | the Nolima 64k score |
| Objectron | 1 | the Objectron score |
| Officeqa Pro Vision | 1 | the Officeqa Pro Vision score |
| Ojbench Cpp | 1 | the Ojbench Cpp score |
| Omnibench Music | 1 | the Omnibench Music score |
| Omnigaia | 1 | the Omnigaia score |
| Omniscience Non Hallucination Rate | 1 | the Omniscience Non Hallucination Rate score |
| Open Rewrite | 1 | the Open Rewrite score |
| Openai Mrcr 2 Needle 256k | 1 | the Openai Mrcr 2 Needle 256k score |
| Openrca | 1 | the Openrca score |
| Osworld Extended | 1 | the Osworld Extended score |
| Osworld Screenshot Only | 1 | the Osworld Screenshot Only score |
| Pathmcqa | 1 | the Pathmcqa score |
| Phibench | 1 | the Phibench score |
| Pointgrounding | 1 | the Pointgrounding score |
| Prbench Legal | 1 | the Prbench Legal score |
| Protocolqa | 1 | the Protocolqa score |
| Qwen Qoder Bench | 1 | the Qwen Qoder Bench score |
| Qwen React Bench | 1 | the Qwen React Bench score |
| Refcocog | 1 | the Refcocog score |
| Robospatialhome | 1 | the Robospatialhome score |
| Robust If | 1 | the Robust If score |
| Ruler 1000k | 1 | the Ruler 1000k score |
| Ruler 2048k | 1 | the Ruler 2048k score |
| Ruler 512k | 1 | the Ruler 512k score |
| Ruler 64k | 1 | the Ruler 64k score |
| Sat Math | 1 | the Sat Math score |
| Scienceqa | 1 | the Scienceqa score |
| Scienceqa Visual | 1 | the Scienceqa Visual score |
| Sifo | 1 | the Sifo score |
| Sifo Multiturn | 1 | the Sifo Multiturn score |
| Siren Agentdojo Attack Success | 1 | the Siren Agentdojo Attack Success score |
| Siren Agentdojo Utility | 1 | the Siren Agentdojo Utility score |
| Stem | 1 | the Stem score |
| Superglue | 1 | the Superglue score |
| Surds | 1 | the Surds score |
| Swe Bench Pro Resolve Rate | 1 | the Swe Bench Pro Resolve Rate score |
| Swe Bench Verified Multiple Attempts | 1 | the Swe Bench Verified Multiple Attempts score |
| Swe Marathon Resolution Rate | 1 | the Swe Marathon Resolution Rate score |
| Tau3 Airline | 1 | the Tau3 Airline score |
| Tau3 Retail | 1 | the Tau3 Retail score |
| Tau3 Telecom | 1 | the Tau3 Telecom score |
| Terminus | 1 | the Terminus score |
| Tldr9 Test | 1 | the Tldr9 Test score |
| Uniform Bar Exam | 1 | the Uniform Bar Exam score |
| Vcr En Easy | 1 | the Vcr En Easy score |
| Vct | 1 | the Vct score |
| Vibe | 1 | the Vibe score |
| Vibe Android | 1 | the Vibe Android score |
| Vibe Backend | 1 | the Vibe Backend score |
| Vibe Ios | 1 | the Vibe Ios score |
| Vibe Simulation | 1 | the Vibe Simulation score |
| Vibe Web | 1 | the Vibe Web score |
| Video Mme Long No Subtitles | 1 | the Video Mme Long No Subtitles score |
| Vocalsound | 1 | the Vocalsound score |
| Vqa Rad | 1 | the Vqa Rad score |
| Vqav2 Test | 1 | the Vqav2 Test score |
| We Math | 1 | the We Math score |
| Webvoyager | 1 | the Webvoyager score |
| Wmdp | 1 | the Wmdp score |
| Worldvqa Forceanswer | 1 | the Worldvqa Forceanswer score |
| Xlsum English | 1 | the Xlsum English score |
| Zerobench Main Pass At 5 | 1 | the Zerobench Main Pass At 5 score |
| Zerobench Main With Python Pass At 5 | 1 | the Zerobench Main With Python Pass At 5 score |
