APEX-Agents

    ๐Ÿ† Leaderboard

    As of August 31, 2026, Grok 4.6 is #1 for APEX-Agents at 57.5%. Ranked by the APEX-Agents score 51 models in this index have a published APEX-Agents score. APEX-Agents leaderboard with live API prices. APEX-Agents โ€” agent execution benchmark across professional workflows.

    Updated August 31, 2026345 models36 providers

    APEX-Agents vs price

    51 models. Left is cheaper. Up is a higher score. The line is the best score you can buy at each price โ€” a dot under it is a worse deal than something on the line.

    Best score at each price#1 on this board
    Grok 4.6
    xAI ยท Proprietary
    57.5%$8.00
    Claude Fable 5
    Anthropic ยท Proprietary
    45%$60.00
    Claude Opus 5
    Anthropic ยท Proprietary
    43.5%$30.00
    Claude Opus 4.8
    Anthropic ยท Proprietary
    42.5%$30.00
    Muse Spark 1.1
    Meta ยท Proprietary
    41.9%$5.50
    GPT-5.6 Sol
    OpenAI ยท Proprietary
    39.9%$35.00
    GPT-5.5
    OpenAI ยท Proprietary
    38.5%$35.00
    Kimi K3OSS
    Moonshot AI ยท Open Source
    37.6%$18.00
    GPT-5.4
    OpenAI ยท Proprietary
    36%$17.50
    GLM-5.2OSS
    Z AI ยท Open Source
    35.6%$5.80
    GPT-5.2
    OpenAI ยท Proprietary
    34.4%$15.75
    Grok 4.5
    xAI ยท Proprietary
    34.2%$8.00
    Claude Opus 4.7
    Anthropic ยท Proprietary
    33.9%$30.00
    Seed 2.1 Pro
    ByteDance ยท Proprietary
    33.8%$5.34
    Gemini 3.1 Pro
    Google ยท Proprietary
    33.5%$17.50
    Claude Sonnet 5
    Anthropic ยท Proprietary
    32.5%$12.00
    Claude Opus 4.6
    Anthropic ยท Proprietary
    32.4%$30.00
    GPT-5.3 Codex
    OpenAI ยท Proprietary
    31.8%$15.75
    Gemini 3 Pro
    Google ยท Proprietary
    31.5%$14.00
    Seed 2.1 Turbo
    ByteDance ยท Proprietary
    29.2%$3.00
    Kimi K2.6OSS
    Moonshot AI ยท Open Source
    27.9%$4.93
    MiniMax M3OSS
    MiniMax ยท Open Source
    27.7%$1.50
    Kimi K2.7 CodeOSS
    Moonshot AI ยท Open Source
    27.6%$4.93
    GPT-5.2 Codex
    OpenAI ยท Proprietary
    27.6%$15.75
    Hy3OSS
    Tencent ยท Open Source
    25.6%$0.66
    GPT-5.4 mini
    OpenAI ยท Proprietary
    24.6%$5.25
    Gemini 3 Flash
    Google ยท Proprietary
    24%$3.50
    Claude Sonnet 4.6
    Anthropic ยท Proprietary
    23.7%$18.00
    GPT-5.1 Codex
    OpenAI ยท Proprietary
    20.7%$11.25
    Claude Opus 4.5
    Anthropic ยท Proprietary
    20.7%$30.00
    GPT-5 Codex
    OpenAI ยท Proprietary
    20.1%$11.25
    GPT-5
    OpenAI ยท Proprietary
    18.3%$11.25
    GPT-5.1
    OpenAI ยท Proprietary
    17.5%$11.25
    o3
    OpenAI ยท Proprietary
    17.2%$10.00
    GLM-5OSS
    Z AI ยท Open Source
    17.2%$4.20
    GPT-5.4 nano
    OpenAI ยท Proprietary
    16.9%$1.45
    Grok-4
    xAI ยท Proprietary
    15.2%$18.00
    Kimi K2.5OSS
    Moonshot AI ยท Open Source
    14.4%$3.68
    Gemini 3.1 Flash-Lite
    Google ยท Proprietary
    13%$1.75
    Grok-4.1
    xAI ยท Proprietary
    12.8%$18.00
    Claude Sonnet 4
    Anthropic ยท Proprietary
    9.3%$18.00
    Claude Haiku 4.5
    Anthropic ยท Proprietary
    8.9%$6.00
    GLM-4.7OSS
    Z AI ยท Open Source
    8.7%$2.80
    DeepSeek-V3.2OSS
    DeepSeek ยท Open Source ยท via OpenRouter
    7%$0.57
    Gemini 2.5 Pro
    Google ยท Proprietary
    6.6%$11.25
    MiniMax M2.5OSS
    MiniMax ยท Open Source
    6.2%$1.50
    GPT OSS 120BOSS
    OpenAI ยท Open Source
    4.7%$0.54
    GLM-4.6OSS
    Z AI ยท Open Source
    4%$2.80
    Grok-3
    xAI ยท Proprietary
    2.1%$18.00
    Gemini 2.5 Flash
    Google ยท Proprietary
    1.8%$2.80
    Showing 1โ€“50 of 345 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    345 models across 36 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    2 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Liquid AI

    2 models

    Nous Research

    1 models

    OpenBMB

    1 models

    Sakana AI

    1 models

    Sarvam AI

    2 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads APEX-Agents right now?

    As of August 31, 2026, Grok 4.6 by xAI is #1 for APEX-Agents at 57.5%. Ranked by the APEX-Agents score This board also tracks APEX-Agents. Next on the same board: Claude Fable 5 and Claude Opus 5. This apex-agents leaderboard ranks models by APEX-Agents. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for apex-agents. Ranked by the APEX-Agents score Input and output are dollars per million tokens.
    RankModelAPEX-AgentsInput /MOutput /M
    1Grok 4.657.5%$2.00$6.00
    2Claude Fable 545%$10.00$50.00
    3Claude Opus 543.5%$5.00$25.00
    4Claude Opus 4.842.5%$5.00$25.00
    5Muse Spark 1.141.9%$1.25$4.25
    6GPT-5.6 Sol39.9%$5.00$30.00
    7GPT-5.538.5%$5.00$30.00
    8Kimi K337.6%$3.00$15.00

    APEX-Agents FAQ

    Who ranks #1 on the APEX-Agents leaderboard?

    As of August 31, 2026, Grok 4.6 by xAI ranks #1 on APEX-Agents at 57.5%. API pricing is $2.00/M input and $6.00/M output.

    What are the top models on APEX-Agents?

    The current APEX-Agents ranking as of August 31, 2026 is 1. Grok 4.6 at 57.5%; 2. Claude Fable 5 at 45%; 3. Claude Opus 5 at 43.5%.

    Which apex-agents model is the cheapest?

    GPT OSS 120B is the cheapest scored model on this apex-agents leaderboard at $0.09/M input and $0.45/M output ($0.54 blended). Grok 4.6 still leads APEX-Agents at 57.5%.

    Should I always pick the #1 APEX-Agents model?

    Not automatically. Grok 4.6 leads APEX-Agents, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh APEX-Agents against input/output price, context window, and related evals.

    How often is the APEX-Agents leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 31, 2026. Treat it as a current index, not a one-off blog post.

    What is APEX-Agents?

    APEX-Agents โ€” agent execution benchmark across professional workflows. This page ranks models that have published a APEX-Agents score, with live API token prices on the same row.

    Where is the APEX-Agents leaderboard?

    This page is the APEX-Agents leaderboard. Models are sorted by APEX-Agents, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this APEX-Agents ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published APEX-Agents score next to live API $/M so you can pick a production SKU, not only a trophy number.