GSM8K

    ๐Ÿ† Leaderboard

    As of August 31, 2026, MiMo-V2.5-Pro is #1 for GSM8K at 99.6%. Ranked by the GSM8K score 50 models in this index have a published GSM8K score. Methodology: GSM8K (https://github.com/openai/grade-school-math). GSM8K leaderboard: rank models by GSM8K next to live API token prices. Official methodology: GSM8K (https://github.com/openai/grade-school-math).

    Updated August 31, 2026345 models36 providers

    GSM8K vs price

    50 models. Left is cheaper. Up is a higher score. The line is the best score you can buy at each price โ€” a dot under it is a worse deal than something on the line.

    Best score at each price#1 on this board
    MiMo-V2.5-ProOSS
    Xiaomi ยท Open Source
    99.6%$1.30
    Kimi K2 InstructOSS
    Moonshot AI ยท Open Source
    97.3%$1.00
    o1
    OpenAI ยท Proprietary
    97.1%$75.00
    GPT-4.5
    OpenAI ยท Proprietary
    97%$225.00
    Llama 3.1 405B InstructOSS
    Meta ยท Open Source
    96.8%$1.78
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    96.4%$18.00
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    96.4%$18.00
    Qwen2.5 32B InstructOSS
    Qwen ยท Open Source
    95.9%$3.50
    Gemma 3 27BOSS
    Google ยท Open Source
    95.9%$0.30
    Qwen2.5 72B InstructOSS
    Qwen ยท Open Source
    95.8%$0.75
    DeepSeek-V2.5OSS
    DeepSeek ยท Open Source
    95.1%$0.42
    Claude 3 Opus
    Anthropic ยท Proprietary
    95%$90.00
    Qwen2.5 14B InstructOSS
    Qwen ยท Open Source
    94.8%$1.75
    Nova Pro
    Amazon ยท Proprietary
    94.8%$4.00
    Nova Lite
    Amazon ยท Proprietary
    94.5%$0.30
    Gemma 3 12BOSS
    Google ยท Open Source
    94.4%$0.15
    Qwen3 235B A22BOSS
    Qwen ยท Open Source
    94.4%$0.20
    Mistral Large 2OSS
    Mistral ยท Open Source ยท via Mistral AI
    93%$8.00
    Nova Micro
    Amazon ยท Proprietary
    92.3%$0.17
    Claude 3 Sonnet
    Anthropic ยท Proprietary
    92.3%$18.00
    Kimi K2 BaseOSS
    Moonshot AI ยท Open Source ยท via OpenRouter
    92.1%$2.87
    Qwen2.5 7B InstructOSS
    Qwen ยท Open Source
    91.6%$0.60
    Llama 3.1 Nemotron 70B InstructOSS
    NVIDIA ยท Open Source
    91.4%$2.40
    GPT-4o mini
    OpenAI ยท Proprietary
    91.3%$0.75
    Qwen2.5-Coder 32B InstructOSS
    Qwen ยท Open Source
    91.1%$0.18
    Qwen2 72B InstructOSS
    Qwen ยท Open Source
    91.1%$0.75
    Gemini 1.5 Pro
    Google ยท Proprietary
    90.8%$12.50
    Grok-1.5
    xAI ยท Proprietary
    90%$20.00
    GPT-4
    OpenAI ยท Proprietary
    90.0%$90.00
    Gemma 3 4BOSS
    Google ยท Open Source
    89.2%$0.06
    Claude 3 Haiku
    Anthropic ยท Proprietary
    88.9%$1.50
    Qwen2.5-Omni-7BOSS
    Qwen ยท Open Source
    88.7%$0.50
    Phi-3.5-MoE-instructOSS
    Microsoft ยท Open Source
    88.7%$0.21
    Phi 4 MiniOSS
    Microsoft ยท Open Source ยท via OpenRouter
    88.6%$0.43
    Jamba 1.5 LargeOSS
    AI21 Labs ยท Open Source
    87%$10.00
    Phi-3.5-mini-instructOSS
    Microsoft ยท Open Source
    86.2%$0.20
    Gemini 1.5 Flash
    Google ยท Proprietary
    86.2%$0.75
    Qwen2.5-Coder 7B InstructOSS
    Qwen ยท Open Source
    83.9%$0.60
    Llama 3.1 8B InstructOSS
    Meta ยท Open Source
    82.4%$0.06
    Qwen2 7B InstructOSS
    Qwen ยท Open Source
    82.3%$0.60
    Granite 3.3 8B InstructOSS
    IBM ยท Open Source
    80.9%$1.00
    Llama 3.2 3B InstructOSS
    Meta ยท Open Source
    77.7%$0.03lowest
    Jamba 1.5 MiniOSS
    AI21 Labs ยท Open Source
    75.8%$0.60
    Gemma 2 27BOSS
    Google ยท Open Source ยท via OpenRouter
    74%$1.30
    Command R+OSS
    Cohere ยท Open Source
    70.7%$1.25
    IBM Granite 4.0 Tiny PreviewOSS
    IBM ยท Open Source ยท via OpenRouter
    70.1%$0.13
    Gemma 2 9BOSS
    Google ยท Open Source ยท via Groq
    68.6%$0.40
    Gemma 3 1BOSS
    Google ยท Open Source
    62.8%$0.06
    GPT-3.5 Turbo
    OpenAI ยท Proprietary
    57.8%$2.00
    ERNIE 4.5
    Baidu ยท Proprietary
    25.2%$4.40
    Showing 1โ€“50 of 345 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    345 models across 36 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    2 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Liquid AI

    2 models

    Nous Research

    1 models

    OpenBMB

    1 models

    Sakana AI

    1 models

    Sarvam AI

    2 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads GSM8K right now?

    As of August 31, 2026, MiMo-V2.5-Pro by Xiaomi is #1 for GSM8K at 99.6%. Ranked by the GSM8K score This board also tracks GSM8K. Next on the same board: Kimi K2 Instruct and o1. This gsm8k leaderboard ranks models by GSM8K. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: GSM8K (https://github.com/openai/grade-school-math); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for gsm8k. Ranked by the GSM8K score Input and output are dollars per million tokens.
    RankModelGSM8KInput /MOutput /M
    1MiMo-V2.5-Pro99.6%$0.43$0.87
    2Kimi K2 Instruct97.3%$0.50$0.50
    3o197.1%$15.00$60.00
    4GPT-4.597%$75.00$150.00
    5Llama 3.1 405B Instruct96.8%$0.89$0.89
    6Claude 3.5 Sonnet96.4%$3.00$15.00
    7Claude 3.5 Sonnet96.4%$3.00$15.00
    8Gemma 3 27B95.9%$0.10$0.20

    GSM8K FAQ

    Who ranks #1 on the GSM8K leaderboard?

    As of August 31, 2026, MiMo-V2.5-Pro by Xiaomi ranks #1 on GSM8K at 99.6%. API pricing is $0.43/M input and $0.87/M output.

    What are the top models on GSM8K?

    The current GSM8K ranking as of August 31, 2026 is 1. MiMo-V2.5-Pro at 99.6%; 2. Kimi K2 Instruct at 97.3%; 3. o1 at 97.1%.

    Which gsm8k model is the cheapest?

    Llama 3.2 3B Instruct is the cheapest scored model on this gsm8k leaderboard at $0.01/M input and $0.02/M output ($0.03 blended). MiMo-V2.5-Pro still leads GSM8K at 99.6%.

    Should I always pick the #1 GSM8K model?

    Not automatically. MiMo-V2.5-Pro leads GSM8K, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh GSM8K against input/output price, context window, and related evals.

    How often is the GSM8K leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 31, 2026. Treat it as a current index, not a one-off blog post.

    What is GSM8K?

    GSM8K is a public LLM eval (the GSM8K score). This page ranks models that have published a score, next to live API prices. Official methodology: GSM8K (https://github.com/openai/grade-school-math).

    Where is the GSM8K leaderboard?

    This page is the GSM8K leaderboard. Models are sorted by GSM8K, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this GSM8K ranking different from the official board?

    The official GSM8K page owns the methodology. This page keeps the published GSM8K score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://github.com/openai/grade-school-math