BabyVision

    ๐Ÿ† Leaderboard

    As of August 31, 2026, Kimi K3 is #1 for BabyVision at 85.7%. Ranked by the BabyVision score 12 models in this index have a published BabyVision score. BabyVision leaderboard with live API prices. BabyVision โ€” early visual understanding: can the model read simple scenes, objects, and spatial relations.

    Updated August 31, 2026345 models36 providers

    BabyVision vs price

    12 models. Left is cheaper. Up is a higher score. The line is the best score you can buy at each price โ€” a dot under it is a worse deal than something on the line.

    Best score at each price#1 on this board
    Kimi K3OSS
    Moonshot AI ยท Open Source
    85.7%$18.00
    Qwen3.8-27BOSS
    Qwen ยท Open Source
    85.6%$3.65
    Muse Spark 1.1
    Meta ยท Proprietary
    76.3%$5.50
    Seed 2.1 Pro
    ByteDance ยท Proprietary
    73.7%$5.34
    Qwen3.7-Plus
    Qwen ยท Proprietary
    70.4%$1.60
    Kimi K2.6OSS
    Moonshot AI ยท Open Source
    68.5%$4.93
    Seed 2.1 Turbo
    ByteDance ยท Proprietary
    62.9%$3.00
    GLM-5.3-FlashOSS
    Z AI ยท Open Source
    53.4%$0.65
    Qwen3.5-27BOSS
    Qwen ยท Open Source
    44.6%$2.70
    Qwen3.5-122B-A10BOSS
    Qwen ยท Open Source
    40.2%$3.60
    Qwen3.5-35B-A3BOSS
    Qwen ยท Open Source
    38.4%$2.25
    DeepSeek-V4-Flash-Vision-Exp
    DeepSeek ยท Proprietary
    35.1%$0.88
    ChatGPT-4o Latest
    OpenAI ยท Proprietary
    โ€”$12.50
    Claude 3 Haiku
    Anthropic ยท Proprietary
    โ€”$1.50
    Claude 3 Opus
    Anthropic ยท Proprietary
    โ€”$90.00
    Claude 3 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude 3.5 Haiku
    Anthropic ยท Proprietary
    โ€”$4.80
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude 3.7 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude Fable 5
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Haiku 4.5
    Anthropic ยท Proprietary
    โ€”$6.00
    Claude Mythos 5
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Mythos Preview
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Opus 4
    Anthropic ยท Proprietary
    โ€”$90.00
    Claude Opus 4.1
    Anthropic ยท Proprietary
    โ€”$90.00
    Claude Opus 4.5
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Opus 4.6
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Opus 4.7
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Opus 4.8
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Opus 5
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Sonnet 4
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude Sonnet 4.5
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude Sonnet 4.6
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude Sonnet 5
    Anthropic ยท Proprietary
    โ€”$12.00
    Codestral-22BOSS
    Mistral ยท Open Source ยท via Mistral AI
    โ€”$1.20
    Command A+OSS
    Cohere ยท Open Source
    โ€”$12.50
    Command R+OSS
    Cohere ยท Open Source
    โ€”$1.25
    Composer 2
    Cursor ยท Proprietary
    โ€”$3.00
    Composer 2 Fast
    Cursor ยท Proprietary
    โ€”$9.00
    DeepSeek R1 Distill Llama 70BOSS
    DeepSeek ยท Open Source
    โ€”$0.50
    DeepSeek R1 Distill Qwen 32BOSS
    DeepSeek ยท Open Source
    โ€”$0.30
    DeepSeek VL2OSS
    DeepSeek ยท Open Source ยท via Replicate
    โ€”$2.80
    DeepSeek VL2 SmallOSS
    DeepSeek ยท Open Source ยท via Replicate
    โ€”$1.70
    DeepSeek VL2 TinyOSS
    DeepSeek ยท Open Source ยท via Replicate
    โ€”$0.90
    DeepSeek-R1OSS
    DeepSeek ยท Open Source
    โ€”$2.74
    DeepSeek-R1-0528OSS
    DeepSeek ยท Open Source
    โ€”$2.74
    DeepSeek-V2.5OSS
    DeepSeek ยท Open Source
    โ€”$0.42
    DeepSeek-V3OSS
    DeepSeek ยท Open Source
    โ€”$1.37
    DeepSeek-V3 0324OSS
    DeepSeek ยท Open Source
    โ€”$1.42
    Showing 1โ€“50 of 345 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    345 models across 36 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    2 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Liquid AI

    2 models

    Nous Research

    1 models

    OpenBMB

    1 models

    Sakana AI

    1 models

    Sarvam AI

    2 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads BabyVision right now?

    As of August 31, 2026, Kimi K3 by Moonshot AI is #1 for BabyVision at 85.7%. Ranked by the BabyVision score This board also tracks BabyVision. Next on the same board: Qwen3.8-27B and Muse Spark 1.1. This babyvision leaderboard ranks models by BabyVision. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for babyvision. Ranked by the BabyVision score Input and output are dollars per million tokens.
    RankModelBabyVisionInput /MOutput /M
    1Kimi K385.7%$3.00$15.00
    2Qwen3.8-27B85.6%$0.45$3.20
    3Muse Spark 1.176.3%$1.25$4.25
    4Seed 2.1 Pro73.7%$0.89$4.45
    5Qwen3.7-Plus70.4%$0.32$1.28
    6Kimi K2.668.5%$0.96$3.97
    7Seed 2.1 Turbo62.9%$0.50$2.50
    8GLM-5.3-Flash53.4%$0.15$0.50

    BabyVision FAQ

    Who ranks #1 on the BabyVision leaderboard?

    As of August 31, 2026, Kimi K3 by Moonshot AI ranks #1 on BabyVision at 85.7%. API pricing is $3.00/M input and $15.00/M output.

    What are the top models on BabyVision?

    The current BabyVision ranking as of August 31, 2026 is 1. Kimi K3 at 85.7%; 2. Qwen3.8-27B at 85.6%; 3. Muse Spark 1.1 at 76.3%.

    Which babyvision model is the cheapest?

    GLM-5.3-Flash is the cheapest scored model on this babyvision leaderboard at $0.15/M input and $0.50/M output ($0.65 blended). Kimi K3 still leads BabyVision at 85.7%.

    Should I always pick the #1 BabyVision model?

    Not automatically. Kimi K3 leads BabyVision, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh BabyVision against input/output price, context window, and related evals.

    How often is the BabyVision leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 31, 2026. Treat it as a current index, not a one-off blog post.

    What is BabyVision?

    BabyVision โ€” early visual understanding: can the model read simple scenes, objects, and spatial relations. This page ranks models that have published a BabyVision score, with live API token prices on the same row.

    Where is the BabyVision leaderboard?

    This page is the BabyVision leaderboard. Models are sorted by BabyVision, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this BabyVision ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published BabyVision score next to live API $/M so you can pick a production SKU, not only a trophy number.