AIME 2024

    ๐Ÿ† Leaderboard

    As of August 31, 2026, Grok-3 Mini is #1 for AIME 2024 at 95.8%. Ranked by the AIME 2024 score 47 models in this index have a published AIME 2024 score. AIME 2024 leaderboard with live API prices. American Invitational Mathematics Examination 2024 โ€” top-tier high school math competition problems.

    Updated August 31, 2026345 models36 providers

    AIME 2024 vs price

    47 models. Left is cheaper. Up is a higher score. The line is the best score you can buy at each price โ€” a dot under it is a worse deal than something on the line.

    Best score at each price#1 on this board
    Grok-3 Mini
    xAI ยท Proprietary
    95.8%$0.80
    Grok-4
    xAI ยท Proprietary
    94%$18.00
    o4-mini
    OpenAI ยท Proprietary
    93.4%$5.50
    LongCat-Flash-ThinkingOSS
    Meituan ยท Open Source
    93.3%$1.50
    Grok-3
    xAI ยท Proprietary
    93.3%$18.00
    Gemini 2.5 Pro
    Google ยท Proprietary
    92%$11.25
    o3
    OpenAI ยท Proprietary
    91.6%$10.00
    DeepSeek-R1-0528OSS
    DeepSeek ยท Open Source
    91.4%$2.74
    GLM-4.5OSS
    Z AI ยท Open Source
    91%$2.80
    Ministral 3 (14B Reasoning 2512)OSS
    Mistral ยท Open Source ยท via Mistral AI
    89.8%$0.40
    GLM-4.5-AirOSS
    Z AI ยท Open Source
    89.4%$1.30
    Gemini 2.5 Flash
    Google ยท Proprietary
    88%$2.80
    o3-mini
    OpenAI ยท Proprietary
    87.3%$5.50
    DeepSeek R1 Distill Llama 70BOSS
    DeepSeek ยท Open Source
    86.7%$0.50
    o1-pro
    OpenAI ยท Proprietary
    86%$750.00
    Ministral 3 (8B Reasoning 2512)OSS
    Mistral ยท Open Source ยท via Mistral AI
    86%$0.30
    MiniMax M1 80KOSS
    MiniMax ยท Open Source
    86%$2.75
    Qwen3 235B A22BOSS
    Qwen ยท Open Source
    85.7%$0.20
    MiniCPM-SALAOSS
    OpenBMB ยท Open Source
    83.8%$0.20
    MiniMax M1 40KOSS
    MiniMax ยท Open Source ยท via OpenRouter
    83.3%$2.75
    DeepSeek R1 Distill Qwen 32BOSS
    DeepSeek ยท Open Source
    83.3%$0.30
    Qwen3 32BOSS
    Qwen ยท Open Source
    81.4%$0.54
    Granite 3.3 8B InstructOSS
    IBM ยท Open Source
    81.2%$1.00
    Qwen3 30B A3BOSS
    Qwen ยท Open Source
    80.4%$0.54
    Claude 3.7 Sonnet
    Anthropic ยท Proprietary
    80%$18.00
    DeepSeek-R1OSS
    DeepSeek ยท Open Source
    79.8%$2.74
    QwQ-32BOSS
    Qwen ยท Open Source ยท via Fireworks
    79.5%$1.80
    Min istral 3 (3B Reasoning 2512)OSS
    Mistral ยท Open Source ยท via Mistral AI
    77.5%$0.20
    Kimi-k1.5
    Moonshot AI ยท Proprietary
    77.5%$2.65
    o1
    OpenAI ยท Proprietary
    74.3%$75.00
    Magistral MediumOSS
    Mistral ยท Open Source ยท via Mistral AI
    73.6%$7.00
    LongCat-Flash-LiteOSS
    Meituan ยท Open Source
    72.2%$0.50
    Kimi K2 0905
    Moonshot AI ยท Proprietary
    72%$3.10
    Magistral Small 2506OSS
    Mistral ยท Open Source ยท via Mistral AI
    70.7%$2.00
    Kimi K2-Instruct-0905OSS
    Moonshot AI ยท Open Source ยท via OpenRouter
    69.6%$3.10
    Kimi K2 InstructOSS
    Moonshot AI ยท Open Source
    69.6%$1.00
    DeepSeek-V3.1OSS
    DeepSeek ยท Open Source
    66.3%$1.27
    o1-mini
    OpenAI ยท Proprietary
    63.6%$15.00
    DeepSeek-V3 0324OSS
    DeepSeek ยท Open Source
    59.4%$1.42
    QwQ-32B-PreviewOSS
    Qwen ยท Open Source
    50%$0.75
    GPT-4.1 mini
    OpenAI ยท Proprietary
    49.6%$2.00
    GPT-4.1
    OpenAI ยท Proprietary
    48.1%$10.00
    o1-preview
    OpenAI ยท Proprietary
    42%$75.00
    DeepSeek-V3OSS
    DeepSeek ยท Open Source
    39.2%$1.37
    GPT-4.5
    OpenAI ยท Proprietary
    36.7%$225.00
    GPT-4.1 nano
    OpenAI ยท Proprietary
    29.4%$0.50
    GPT-4o
    OpenAI ยท Proprietary
    13.1%$12.50
    ChatGPT-4o Latest
    OpenAI ยท Proprietary
    โ€”$12.50
    Claude 3 Haiku
    Anthropic ยท Proprietary
    โ€”$1.50
    Claude 3 Opus
    Anthropic ยท Proprietary
    โ€”$90.00
    Showing 1โ€“50 of 345 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    345 models across 36 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    2 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Liquid AI

    2 models

    Nous Research

    1 models

    OpenBMB

    1 models

    Sakana AI

    1 models

    Sarvam AI

    2 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads AIME 2024 right now?

    As of August 31, 2026, Grok-3 Mini by xAI is #1 for AIME 2024 at 95.8%. Ranked by the AIME 2024 score This board also tracks AIME 2024. Next on the same board: Grok-4 and o4-mini. This aime 2024 leaderboard ranks models by AIME 2024. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for aime 2024. Ranked by the AIME 2024 score Input and output are dollars per million tokens.
    RankModelAIME 2024Input /MOutput /M
    1Grok-3 Mini95.8%$0.30$0.50
    2Grok-494%$3.00$15.00
    3o4-mini93.4%$1.10$4.40
    4LongCat-Flash-Thinking93.3%$0.30$1.20
    5Grok-393.3%$3.00$15.00
    6Gemini 2.5 Pro92%$1.25$10.00
    7o391.6%$2.00$8.00
    8DeepSeek-R1-052891.4%$0.55$2.19

    AIME 2024 FAQ

    Who ranks #1 on the AIME 2024 leaderboard?

    As of August 31, 2026, Grok-3 Mini by xAI ranks #1 on AIME 2024 at 95.8%. API pricing is $0.30/M input and $0.50/M output.

    What are the top models on AIME 2024?

    The current AIME 2024 ranking as of August 31, 2026 is 1. Grok-3 Mini at 95.8%; 2. Grok-4 at 94%; 3. o4-mini at 93.4%.

    Which aime 2024 model is the cheapest?

    Qwen3 235B A22B is the cheapest scored model on this aime 2024 leaderboard at $0.10/M input and $0.10/M output ($0.20 blended). Grok-3 Mini still leads AIME 2024 at 95.8%.

    Should I always pick the #1 AIME 2024 model?

    Not automatically. Grok-3 Mini leads AIME 2024, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh AIME 2024 against input/output price, context window, and related evals.

    How often is the AIME 2024 leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 31, 2026. Treat it as a current index, not a one-off blog post.

    What is AIME 2024?

    American Invitational Mathematics Examination 2024 โ€” top-tier high school math competition problems. This page ranks models that have published a AIME 2024 score, with live API token prices on the same row.

    Where is the AIME 2024 leaderboard?

    This page is the AIME 2024 leaderboard. Models are sorted by AIME 2024, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this AIME 2024 ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published AIME 2024 score next to live API $/M so you can pick a production SKU, not only a trophy number.