MATH 500

    🏆 Leaderboard

    As of August 31, 2026, LongCat-Flash-Thinking is #1 for MATH 500 at 99.2%. Ranked by MATH-500: textbook and contest math problems. 50 models in this index have a published MATH 500 score. MATH 500 leaderboard with live API prices. MATH-500 — a 500-problem subset of MATH covering competition-level mathematical reasoning.

    Updated August 31, 2026345 models36 providers

    MATH 500 vs price

    49 models. Left is cheaper. Up is a higher score. The line is the best score you can buy at each price — a dot under it is a worse deal than something on the line.

    Best score at each price#1 on this board
    LongCat-Flash-ThinkingOSS
    Meituan · Open Source
    99.2%$1.50
    Sarvam-105BOSS
    Sarvam AI · Open Source
    98.6%$1.60
    GLM-4.5OSS
    Z AI · Open Source
    98.2%$2.80
    GLM-4.5-AirOSS
    Z AI · Open Source
    98.1%$1.30
    o3-mini
    OpenAI · Proprietary
    97.9%$5.50
    Nemotron Nano 9B v2OSS
    NVIDIA · Open Source · via OpenRouter
    97.8%
    Kimi K2-Instruct-0905OSS
    Moonshot AI · Open Source · via OpenRouter
    97.4%$3.10
    Kimi K2 InstructOSS
    Moonshot AI · Open Source
    97.4%$1.00
    DeepSeek-R1OSS
    DeepSeek · Open Source
    97.3%$2.74
    Sarvam-30BOSS
    Sarvam AI · Open Source
    97%$0.60
    Llama 3.1 Nemotron Ultra 253B v1OSS
    NVIDIA · Open Source
    97%$2.40
    MiniMax M1 80KOSS
    MiniMax · Open Source
    96.8%$2.75
    LongCat-Flash-LiteOSS
    Meituan · Open Source
    96.8%$0.50
    Llama-3.3 Nemotron Super 49B v1OSS
    NVIDIA · Open Source
    96.6%$0.50
    LongCat-Flash-ChatOSS
    Meituan · Open Source
    96.4%$1.50
    Gemini 3 Pro
    Google · Proprietary
    96.4%$14.00
    Kimi-k1.5
    Moonshot AI · Proprietary
    96.2%$2.65
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    96.2%$18.00
    MiniMax M1 40KOSS
    MiniMax · Open Source · via OpenRouter
    96%$2.75
    GPT-5
    OpenAI · Proprietary
    96%$11.25
    Llama 3.1 Nemotron Nano 8B V1OSS
    NVIDIA · Open Source
    95.4%$0.20
    GPT-5 mini
    OpenAI · Proprietary
    94.8%$2.25
    GPT OSS 120BOSS
    OpenAI · Open Source
    94.8%$0.54
    Qwen3 235B A22BOSS
    Qwen · Open Source
    94.6%$0.20
    o3
    OpenAI · Proprietary
    94.6%$10.00
    DeepSeek R1 Distill Llama 70BOSS
    DeepSeek · Open Source
    94.5%$0.50
    DeepSeek R1 Distill Qwen 32BOSS
    DeepSeek · Open Source
    94.3%$0.30
    o4-mini
    OpenAI · Proprietary
    94.2%$5.50
    GPT OSS 20BOSS
    OpenAI · Open Source
    94.2%$0.25
    DeepSeek-V3 0324OSS
    DeepSeek · Open Source
    94%$1.42
    GPT-5 nano
    OpenAI · Proprietary
    93.8%$0.45
    Claude Opus 4.1
    Anthropic · Proprietary
    93%$90.00
    QwQ-32B-PreviewOSS
    Qwen · Open Source
    90.6%$0.75
    QwQ-32BOSS
    Qwen · Open Source · via Fireworks
    90.6%$1.80
    o1
    OpenAI · Proprietary
    90.4%$75.00
    Claude Opus 4
    Anthropic · Proprietary
    90.4%$90.00
    Claude Sonnet 4
    Anthropic · Proprietary
    90.3%$18.00
    DeepSeek-V3OSS
    DeepSeek · Open Source
    90.2%$1.37
    o1-mini
    OpenAI · Proprietary
    90%$15.00
    Grok-3
    xAI · Proprietary
    89.8%$18.00
    Gemini 2.0 Flash
    Google · Proprietary
    89.7%$0.50
    MiniMax M2.1OSS
    MiniMax · Open Source
    89%$1.50
    GPT-4.1 mini
    OpenAI · Proprietary
    88%$2.00
    GPT-4.1
    OpenAI · Proprietary
    87.2%$10.00
    GPT-4.1 nano
    OpenAI · Proprietary
    80.2%$0.50
    GPT-4o
    OpenAI · Proprietary
    75.2%$12.50
    GPT-4o mini
    OpenAI · Proprietary
    72.6%$0.75
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    72.4%$18.00
    Granite 3.3 8B InstructOSS
    IBM · Open Source
    69.0%$1.00
    Claude 3.5 Haiku
    Anthropic · Proprietary
    64.2%$4.80
    Showing 150 of 345 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    345 models across 36 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    2 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Liquid AI

    2 models

    Nous Research

    1 models

    OpenBMB

    1 models

    Sakana AI

    1 models

    Sarvam AI

    2 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads MATH 500 right now?

    As of August 31, 2026, LongCat-Flash-Thinking by Meituan is #1 for MATH 500 at 99.2%. Ranked by MATH-500: textbook and contest math problems. This board also tracks MATH 500. Next on the same board: Sarvam-105B and GLM-4.5. This math 500 leaderboard ranks models by MATH 500. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for math 500. Ranked by MATH-500: textbook and contest math problems. Input and output are dollars per million tokens.
    RankModelMATH 500Input /MOutput /M
    1LongCat-Flash-Thinking99.2%$0.30$1.20
    2Sarvam-105B98.6%$0.80$0.80
    3GLM-4.598.2%$0.60$2.20
    4GLM-4.5-Air98.1%$0.20$1.10
    5o3-mini97.9%$1.10$4.40
    6Nemotron Nano 9B v297.8%
    7Kimi K2-Instruct-090597.4%$0.60$2.50
    8Kimi K2 Instruct97.4%$0.50$0.50

    MATH 500 FAQ

    Who ranks #1 on the MATH 500 leaderboard?

    As of August 31, 2026, LongCat-Flash-Thinking by Meituan ranks #1 on MATH 500 at 99.2%. API pricing is $0.30/M input and $1.20/M output.

    What are the top models on MATH 500?

    The current MATH 500 ranking as of August 31, 2026 is 1. LongCat-Flash-Thinking at 99.2%; 2. Sarvam-105B at 98.6%; 3. GLM-4.5 at 98.2%.

    Which math 500 model is the cheapest?

    Llama 3.1 Nemotron Nano 8B V1 is the cheapest scored model on this math 500 leaderboard at $0.04/M input and $0.16/M output ($0.20 blended). LongCat-Flash-Thinking still leads MATH 500 at 99.2%.

    Should I always pick the #1 MATH 500 model?

    Not automatically. LongCat-Flash-Thinking leads MATH 500, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh MATH 500 against input/output price, context window, and related evals.

    How often is the MATH 500 leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 31, 2026. Treat it as a current index, not a one-off blog post.

    What is MATH 500?

    MATH-500 — a 500-problem subset of MATH covering competition-level mathematical reasoning. This page ranks models that have published a MATH 500 score, with live API token prices on the same row.

    Where is the MATH 500 leaderboard?

    This page is the MATH 500 leaderboard. Models are sorted by MATH 500, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this MATH 500 ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published MATH 500 score next to live API $/M so you can pick a production SKU, not only a trophy number.