Robocurve — internal reference

Real-world image quality vs. serving speed — 15 models

DATA AS OF 2026-07-28

Vision Arena ELO (blind head-to-head votes on real user-uploaded photos) against each model's fastest benchmarked output speed. Up and to the right is better; the dashed staircase is the Pareto frontier — models below/left of it are beaten on both axes by something else.

Gemini (closed API) Gemma 4 (open weights) Other flagships  Pareto frontier whiskers = ±95% ELO CI
Model Vision Arena ELO Max output speed Speed source

Frontier: Claude Fable 5 (best image quality), Gemini 3.6 Flash, Kimi K2.6 on Fireworks, Gemma 4 31B on Cerebras (fastest by ~5×). Everything else is dominated.

Not on the Vision Arena leaderboard (so they can't appear here): GPT-5.6 Sol — OpenAI's newest; GPT-5.5 is its best arena entry — plus Kimi K3 and Inkling by Thinking Machines (see the benchmark section below for where all three land). GLM-5.2 is text-only (no image input), so it has no vision data anywhere on this page; Zhipu's newest vision model, GLM-5V-Turbo, IS on the arena (1230 ± 7, rank 48) but Artificial Analysis reports no output speed for it — excluded for missing data, like Gemini 3 Pro (ELO 1289) and Meta's Muse Spark (1295).

Omitted for legibility: Claude Opus 4.6 / Sonnet variants, the non-thinking Opus 4.7 / 4.8 listings, and mid-board GPT variants — all at or below 1300 ELO at sub-240 t/s, so every one is dominated by Gemini 3.6 Flash and none can change the frontier. Claude Opus 5 is shown at its only arena listing, the high-effort variant (1299 ELO, 56.2 t/s at high effort).

Caveats: ELO intervals are 95% CIs — Gemini 3.6 Flash's ±38 is wide (new model, few votes). Speeds are Artificial Analysis medians sampled 2026-07-28 at each model's benchmarked reasoning setting, taking the fastest tracked non-quantized provider for open-weights models (Cerebras for Gemma 31B, Makora for 26B A4B, Fireworks for Kimi K2.6, Azure for K2.5) and the first-party API for closed models. Quantized endpoints are excluded — e.g. Together serves K2.6 at 528 t/s in FP4, but that isn't the model the arena scored. Arena setting and speed-benchmark setting may differ slightly (e.g. thinking budgets). Log-scale x-axis.

Where Inkling, Kimi K3 and GPT-5.6 Sol sit

None of the three is on the Vision Arena leaderboard, so they can't appear in the chart above. Exactly two published vision benchmarks cover both Inkling and K3 — neither is real-world photos. Every model from the arena chart with a published no-tools score on these benchmarks is charted too, so the same frontier view applies: score against fastest benchmarked output speed, log-scale x. Reasoning effort per model as published; unlabeled points have no stated effort.

Sources (also shown per point in the tooltips): Thinking Machines "Introducing Inkling" eval table (Inkling at effort 0.99, Kimi K2.5/K2.6, Gemini 3.1 Pro at high, Claude Fable 5 at max, GPT-5.6 Sol at max/xhigh); Moonshot's Kimi-K3 card (K3 at effort max / temp 1.0, plus its comparison columns for GPT-5.5 and Claude Opus 4.8 — Moonshot states no effort setting for those); Google's model eval cards (Gemini 3 / 3.5 / 3.6 Flash, Gemma 4 — self-reported); Anthropic's Opus 4.7 system card (CharXiv at max effort); Alibaba's Qwen3.7 Plus launch table (CharXiv). Artificial Analysis for all speeds, sampled 2026-07-28.

Cross-vendor caveat: these numbers mix five sources with different harnesses and protocols. Where sources overlap they agree to ~3 points at best: Fable 5 scores 84.2 (TM) vs 81.2 (Moonshot) on MMMU-Pro and 86.5 (TM) vs 88.9 (Moonshot) on CharXiv; Google averages MMMU-Pro over the Standard-10 and Vision splits while TM/Moonshot use Standard-10; Anthropic's CharXiv uses its "updated settings". So treat gaps under ~3 points as noise — e.g. K3 84.8 vs Sol 84.7 vs Qwen 84.4 vs Gemini 3.5 Flash 84.2 on CharXiv is a four-way statistical tie, and frontier membership within that band is not meaningful.

Speeds follow the top chart's convention — fastest tracked non-quantized provider for open-weights models (Fireworks serves K3 at 162.7 t/s and K2.6 at 424.6 t/s; Azure serves K2.5 at 142.9 t/s; Cerebras serves Gemma 31B at 2,177 t/s; Makora serves 26B A4B at 166.9 t/s), first-party APIs for closed models, and the Thinking Machines API for Inkling (85.0 t/s — DeepInfra is faster at 228.8 t/s but FP8-quantized, so excluded).

Absent for lack of published no-tools scores: Claude Opus 5 (post-dates both comparison tables; its own card dropped CharXiv for Chartography and reports no MMMU-Pro); Grok 4.5 (xAI publishes neither — Google measured its CharXiv at 81.6 on Google's own harness, not charted); Gemini 3.6 Flash has no MMMU-Pro score and Claude Opus 4.7 / the Gemma 4 models no CharXiv score from their vendors, so they appear on one chart only; Qwen3.7 Plus's MMMU-Pro (79.0 floats around unattributed) could not be verified to the vendor table, so only its CharXiv is charted; GLM-5.2 is text-only and GLM-5V-Turbo has no published score on either benchmark. With tools the picture shifts: K3 reaches 83.4 MMMU-Pro and 91.3 CharXiv with Python; Inkling reaches 82.0 CharXiv with Python.