IFEval

    ๐Ÿ† Leaderboard

    As of August 31, 2026, Qwen3.5-27B is #1 for IFEval at 95%. Ranked by IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.' 64 models in this index have a published IFEval score. Methodology: IFEval (https://github.com/google-research/google-research/tree/master/instruction_following_eval). IFEval leaderboard: rank models by IFEval next to live API token prices. Official methodology: IFEval (https://github.com/google-research/google-research/tree/master/instruction_following_eval).

    Updated August 31, 2026345 models36 providers

    IFEval vs price

    63 models. Left is cheaper. Up is a higher score. The line is the best score you can buy at each price โ€” a dot under it is a worse deal than something on the line.

    Best score at each price#1 on this board
    Qwen3.5-27BOSS
    Qwen ยท Open Source
    95%$2.70
    Qwen3.7-Plus
    Qwen ยท Proprietary
    94.6%$1.60
    Qwen3.7 Max
    Qwen ยท Proprietary
    94.3%$5.00
    Qwen3.6 Plus
    Qwen ยท Proprietary
    94.3%$3.50
    o3-mini
    OpenAI ยท Proprietary
    93.9%$5.50
    Qwen3.5-122B-A10BOSS
    Qwen ยท Open Source
    93.4%$3.60
    Claude 3.7 Sonnet
    Anthropic ยท Proprietary
    93.2%$18.00
    Qwen3.5-397B-A17BOSS
    Qwen ยท Open Source
    92.6%$4.20
    Nova Pro
    Amazon ยท Proprietary
    92.1%$4.00
    Llama 3.3 70B InstructOSS
    Meta ยท Open Source
    92.1%$0.40
    Qwen3.5-35B-A3BOSS
    Qwen ยท Open Source
    91.9%$2.25
    Qwen3.5-9BOSS
    Qwen ยท Open Source ยท via OpenRouter
    91.5%$0.25
    Gemma 3 27BOSS
    Google ยท Open Source
    90.4%$0.30
    Nemotron Nano 9B v2OSS
    NVIDIA ยท Open Source ยท via OpenRouter
    90.3%โ€”
    Gemma 3 4BOSS
    Google ยท Open Source
    90.2%$0.06
    Qwen3.5-4BOSS
    Qwen ยท Open Source
    89.8%$0.20
    Kimi K2-Instruct-0905OSS
    Moonshot AI ยท Open Source ยท via OpenRouter
    89.8%$3.10
    Kimi K2 InstructOSS
    Moonshot AI ยท Open Source
    89.8%$1.00
    Nova Lite
    Amazon ยท Proprietary
    89.7%$0.30
    LongCat-Flash-ChatOSS
    Meituan ยท Open Source
    89.7%$1.50
    Llama 3.1 Nemotron Ultra 253B v1OSS
    NVIDIA ยท Open Source
    89.5%$2.40
    Qwen3-Next-80B-A3B-ThinkingOSS
    Qwen ยท Open Source
    88.9%$1.65
    Gemma 3 12BOSS
    Google ยท Open Source
    88.9%$0.15
    Qwen3-235B-A22B-Instruct-2507OSS
    Qwen ยท Open Source
    88.7%$0.95
    Llama 3.1 405B InstructOSS
    Meta ยท Open Source
    88.6%$1.78
    Qwen3 VL 235B A22B ThinkingOSS
    Qwen ยท Open Source
    88.2%$3.94
    GPT-4.5
    OpenAI ยท Proprietary
    88.2%$225.00
    Qwen3-235B-A22B-Thinking-2507OSS
    Qwen ยท Open Source
    87.8%$3.30
    Qwen3 VL 32B ThinkingOSS
    Qwen ยท Open Source
    87.8%$0.80
    Qwen3 VL 235B A22B InstructOSS
    Qwen ยท Open Source
    87.8%$1.79
    Qwen3-Next-80B-A3B-InstructOSS
    Qwen ยท Open Source
    87.6%$1.65
    Llama 3.1 70B InstructOSS
    Meta ยท Open Source
    87.5%$0.40
    GPT-4.1
    OpenAI ยท Proprietary
    87.4%$10.00
    Nova Micro
    Amazon ยท Proprietary
    87.2%$0.17
    Kimi-k1.5
    Moonshot AI ยท Proprietary
    87.2%$2.65
    DeepSeek-V3OSS
    DeepSeek ยท Open Source
    86.1%$1.37
    Qwen3 VL 30B A3B InstructOSS
    Qwen ยท Open Source
    85.8%$0.90
    Sarvam-105BOSS
    Sarvam AI ยท Open Source
    84.8%$1.60
    Qwen3 VL 32B InstructOSS
    Qwen ยท Open Source ยท via OpenRouter
    84.7%$0.52
    Qwen2.5 72B InstructOSS
    Qwen ยท Open Source
    84.1%$0.75
    GPT-4.1 mini
    OpenAI ยท Proprietary
    84.1%$2.00
    QwQ-32BOSS
    Qwen ยท Open Source ยท via Fireworks
    83.9%$1.80
    Qwen3 VL 8B InstructOSS
    Qwen ยท Open Source
    83.7%$0.58
    Qwen3 VL 8B ThinkingOSS
    Qwen ยท Open Source
    83.2%$2.27
    Mistral Small 3 24B InstructOSS
    Mistral ยท Open Source ยท via Mistral AI
    82.9%$0.21
    Qwen3 VL 4B ThinkingOSS
    Qwen ยท Open Source
    82.6%$1.10
    Qwen3 VL 4B InstructOSS
    Qwen ยท Open Source
    82.3%$0.70
    LFM2.5-VL-3BOSS
    Liquid AI ยท Open Source
    82.3%$0.15
    Qwen3 VL 30B A3B ThinkingOSS
    Qwen ยท Open Source
    81.7%$1.20
    GPT-4o
    OpenAI ยท Proprietary
    81%$12.50
    Showing 1โ€“50 of 345 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    345 models across 36 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    2 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Liquid AI

    2 models

    Nous Research

    1 models

    OpenBMB

    1 models

    Sakana AI

    1 models

    Sarvam AI

    2 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads IFEval right now?

    As of August 31, 2026, Qwen3.5-27B by Qwen is #1 for IFEval at 95%. Ranked by IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.' This board also tracks IFEval. Next on the same board: Qwen3.7-Plus and Qwen3.7 Max. This ifeval leaderboard ranks models by IFEval. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: IFEval (https://github.com/google-research/google-research/tree/master/instruction_following_eval); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for ifeval. Ranked by IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.' Input and output are dollars per million tokens.
    RankModelIFEvalInput /MOutput /M
    1Qwen3.5-27B95%$0.30$2.40
    2Qwen3.7-Plus94.6%$0.32$1.28
    3Qwen3.7 Max94.3%$1.25$3.75
    4Qwen3.6 Plus94.3%$0.50$3.00
    5o3-mini93.9%$1.10$4.40
    6Qwen3.5-122B-A10B93.4%$0.40$3.20
    7Claude 3.7 Sonnet93.2%$3.00$15.00
    8Qwen3.5-397B-A17B92.6%$0.60$3.60

    IFEval FAQ

    Who ranks #1 on the IFEval leaderboard?

    As of August 31, 2026, Qwen3.5-27B by Qwen ranks #1 on IFEval at 95%. API pricing is $0.30/M input and $2.40/M output.

    What are the top models on IFEval?

    The current IFEval ranking as of August 31, 2026 is 1. Qwen3.5-27B at 95%; 2. Qwen3.7-Plus at 94.6%; 3. Qwen3.7 Max at 94.3%.

    Which ifeval model is the cheapest?

    Llama 3.2 3B Instruct is the cheapest scored model on this ifeval leaderboard at $0.01/M input and $0.02/M output ($0.03 blended). Qwen3.5-27B still leads IFEval at 95%.

    Should I always pick the #1 IFEval model?

    Not automatically. Qwen3.5-27B leads IFEval, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh IFEval against input/output price, context window, and related evals.

    How often is the IFEval leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 31, 2026. Treat it as a current index, not a one-off blog post.

    What is IFEval?

    IFEval is a public LLM eval (IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.'). This page ranks models that have published a score, next to live API prices. Official methodology: IFEval (https://github.com/google-research/google-research/tree/master/instruction_following_eval).

    Where is the IFEval leaderboard?

    This page is the IFEval leaderboard. Models are sorted by IFEval, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this IFEval ranking different from the official board?

    The official IFEval page owns the methodology. This page keeps the published IFEval score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://github.com/google-research/google-research/tree/master/instruction_following_eval