Robocurve — internal reference

Real-world image quality vs. serving speed — 12 models

DATA AS OF 2026-07-27

Vision Arena ELO (blind head-to-head votes on real user-uploaded photos) against each model's fastest benchmarked output speed. Up and to the right is better; the dashed staircase is the Pareto frontier — models below/left of it are beaten on both axes by something else.

Gemini (closed API) Gemma 4 (open weights) Other flagships  Pareto frontier whiskers = ±95% ELO CI
Model Vision Arena ELO Max output speed Speed source

Frontier: Claude Fable 5 (best image quality), Gemini 3.6 Flash (best quality-speed balance), Gemma 4 31B on Cerebras (fastest by ~8×). Everything else is dominated.

Excluded for missing data: Gemini 3 Pro (ELO 1289) and Meta's Muse Spark (ELO 1295) — Artificial Analysis reports no speed for either; Inkling by Thinking Machines (87.4 t/s) and Kimi K3 (33.3 t/s, 144 s time-to-first-token) — neither is on the Vision Arena leaderboard yet (see the benchmark section below for where they land). Claude Opus 5 / Opus 4.6 / Sonnet variants and mid-board GPT variants are omitted for legibility; all sit at or below 1300 ELO (under Gemini 3.6 Flash's 1301) at sub-240 t/s speeds, so every one is dominated by Gemini 3.6 Flash and none can change the frontier.

Caveats: ELO intervals are 95% CIs — Gemini 3.6 Flash's ±38 is wide (new model, few votes). Speeds are Artificial Analysis medians at each model's benchmarked reasoning setting, taking the fastest tracked provider for open-weights models (Cerebras for Gemma 31B, Cloudflare for 26B A4B) and the first-party API for closed models. Arena setting and speed-benchmark setting may differ slightly (e.g. thinking budgets). Log-scale x-axis.

Where Inkling and Kimi K3 sit

Neither model is on the Vision Arena leaderboard, so they can't appear in the chart above. Exactly two published vision benchmarks cover both — neither is real-world photos. Same frontier view: benchmark score against fastest benchmarked output speed, log-scale x. Scores are without tool augmentation; reasoning effort per model as published.

Reasoning effort as published: Thinking Machines reports Inkling at effort 0.99, Claude Fable 5 at max, Gemini 3.1 Pro at high, and GPT-5.6 Sol at max/xhigh (no effort stated for Kimi K2.5 / K2.6); Moonshot reports all Kimi K3 results at effort max, temperature 1.0.

Cross-vendor caveat: Kimi K3's scores come from Moonshot's own harness; every other score is from Thinking Machines' Inkling evals. The two harnesses agree on GPT-5.6 Sol (83.0 vs 83.0 MMMU-Pro; 84.6 vs 84.7 CharXiv) but differ by ~3 points on Claude Fable 5, so treat score gaps under ~3 points as harness noise — in particular, K3's 84.8 vs Sol's 84.7 on CharXiv is a statistical tie, and the frontier between them there should be read as shared. With tools the picture shifts: K3 reaches 83.4 MMMU-Pro and 91.3 CharXiv with Python; Inkling reaches 82.0 CharXiv with Python.

Speeds follow the top chart's convention — fastest tracked provider for open-weights models (Fireworks serves Kimi K3 at 164.4 t/s, ~5× Moonshot's own API) and first-party APIs for closed ones; Inkling is 87.4 t/s on the Thinking Machines API. Kimi K2.5 / K2.6 use Artificial Analysis's primary (Moonshot) listings — faster third-party hosts may exist but aren't charted.

Sources: Thinking Machines "Introducing Inkling" eval table; Moonshot AI Kimi-K3 model card (multimodal scores averaged over 3 runs); Artificial Analysis for speeds.