ModelRig

Leaderboard

Probed capability facts, ranked by effective cost of conformance — the effective $ per 1,000 schema-conformant outputs, repairs included. Every rate is a sampled statistic with n and a 95% CI. Column headers link to what each one means; model and cost cells link to what each model costs.

The public face of conformance & cost monitoring: the same probes that rank models in the open watch your own traffic in the console. Conformance & cost monitoring.

Coverage: v0 corpus, finance-weighted (seeded from our first production customer) — schema corpus per fully-probed model: 95 samples — probe-suite n=75 · demo-rig n=20, 25 of them hard-tier. 37 models across 7 providers — enrollment is explicit; frontier additions land with the next probe cycle. Rates here aggregate the whole corpus — treat them as corpus-wide, not domain-general. The probe-kit's headline ask is fixtures from your domain contribute one.
grouped by fixture source (provenance — where fixtures came from); to rank models by capability, use the first dropdown. what makes a fixture hard ▸

Ranked by the corpus-wide effective $ per 1,000 schema-conformant outputs (no per-tier cost is published), shown beside accuracy on the hard-tier fixture subset — a descriptive statistic, not a RigIndex rank. what this measures. what a RigIndex rank is.

Cost-ranked: the cheapest model that passes sits first, whatever its headline accuracy. Rank by best accuracy instead.

modelaccuracy · hard tiermethod▸schema conformancemethod▸value accuracymethod▸effective $ / 1K conformant (aggregate)groundedmethod▸cache hitsmethod▸probed
deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506

fixtures: demo-rig n=20 · probe-suite n=75

60%n=2595%[88%–98%] n=9589%n=95$0.0490%n=150%n=52026-08-18
deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2593%[86%–96%] n=95100%n=95$0.060100%n=150%n=52026-08-18
deepinfra/google/gemma-4-31B-it

fixtures: demo-rig n=20 · probe-suite n=75

100%n=25100%[96%–100%] n=9598%n=95$0.0820%n=52026-08-18
deepseek/deepseek-chat

fixtures: demo-rig n=20 · probe-suite n=75

80%n=2595%[88%–98%] n=9595%n=95$0.084100%n=15100%n=52026-08-19
deepseek/deepseek-v4-flash

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2596%[90%–98%] n=95100%n=95$0.09380%n=15100%n=52026-08-19
deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2583%[74%–89%] n=95100%n=95$0.140100%n=150%n=52026-08-18
deepseek/deepseek-reasoner

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2595%[88%–98%] n=95100%n=95$0.14893%n=15100%n=52026-08-19
openai/gpt-5.4-nano

fixtures: demo-rig n=20 · probe-suite n=75

56%n=25100%[96%–100%] n=9588%n=95$0.15053%n=1560%n=52026-08-19
deepinfra/Qwen/Qwen3-30B-A3B

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2597%[91%–99%] n=9599%n=95$0.17187%n=150%n=52026-08-18
deepinfra/openai/gpt-oss-120b

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2599%[94%–100%] n=95100%n=95$0.1710%n=150%n=52026-08-19
gemini/gemini-3.1-flash-lite

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2595%[88%–98%] n=95100%n=95$0.207100%n=150%n=52026-08-19
deepinfra/Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo

fixtures: demo-rig n=20 · probe-suite n=75

88%n=2595%[88%–98%] n=9596%n=95$0.22773%n=1580%n=52026-08-18
fireworks/gpt-oss-120b

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2599%[94%–100%] n=95100%n=95$0.26620%n=15100%n=52026-08-19
deepseek/deepseek-v4-pro

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2599%[94%–100%] n=9599%n=95$0.4047%n=15100%n=52026-08-19
openai/gpt-5.4-mini

fixtures: demo-rig n=20 · probe-suite n=75

92%n=25100%[96%–100%] n=9598%n=95$0.52560%n=15100%n=52026-08-19
deepinfra/MiniMaxAI/MiniMax-M3

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2599%[94%–100%] n=95100%n=95$0.69787%n=15100%n=52026-08-18
openai/gpt-5-mini

fixtures: demo-rig n=20 · probe-suite n=75

100%n=25100%[96%–100%] n=95100%n=95$1.04753%n=15100%n=52026-08-19
deepinfra/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B

fixtures: demo-rig n=20 · probe-suite n=75

100%n=25100%[96%–100%] n=95100%n=95$1.18733%n=150%n=52026-08-19
gemini/gemini-2.5-flash

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2595%[88%–98%] n=95100%n=95$1.28193%n=1540%n=52026-08-19
openai/gpt-5.2

fixtures: demo-rig n=20 · probe-suite n=75

100%n=25100%[96%–100%] n=95100%n=95$1.5727%n=1580%n=52026-08-19
deepinfra/moonshotai/Kimi-K2.7-Code

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2597%[91%–99%] n=95100%n=95$1.61813%n=15100%n=52026-08-19
anthropic/claude-haiku-4-5

fixtures: demo-rig n=20 · probe-suite n=75

48%n=25100%[96%–100%] n=9586%n=95$1.78687%n=150%n=52026-08-18
openai/gpt-5.4

fixtures: demo-rig n=20 · probe-suite n=75

100%n=25100%[96%–100%] n=95100%n=95$1.80640%n=15100%n=52026-08-19
deepinfra/zai-org/GLM-5.2

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2598%[93%–99%] n=95100%n=95$1.86060%n=15100%n=52026-08-19
gemini/gemini-3-flash-preview

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2598%[93%–99%] n=95100%n=95$2.637100%n=150%n=52026-08-19
fireworks/glm-5p2

fixtures: demo-rig n=20 · probe-suite n=75

100%n=25100%[96%–100%] n=95100%n=95$2.76347%n=15100%n=52026-08-19
deepinfra/moonshotai/Kimi-K2.6

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2599%[94%–100%] n=95100%n=95$3.6460%n=1580%n=52026-08-19
anthropic/claude-sonnet-5

fixtures: demo-rig n=20 · probe-suite n=75

84%n=2592%[84%–96%] n=9594%n=95$4.36293%n=150%n=52026-08-18
openai/gpt-5.5

fixtures: demo-rig n=20 · probe-suite n=75

100%n=25100%[96%–100%] n=95100%n=95$4.6840%n=15100%n=52026-08-19
fireworks/kimi-k3

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2598%[93%–99%] n=95100%n=95$5.09860%n=15100%n=52026-08-19
anthropic/claude-opus-5

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2599%[94%–100%] n=9597%n=95$9.70593%n=152026-08-18
gemini/gemini-3.1-pro-preview

fixtures: demo-rig n=20 · probe-suite n=75

100%n=25100%[96%–100%] n=95100%n=95$11.558100%n=150%n=52026-08-19
anthropic/claude-fable-5

fixtures: demo-rig n=20 · probe-suite n=75

100%n=2592%[84%–96%] n=9594%n=95$19.98093%n=152026-08-18
openai/gpt-5.5-pro

fixtures: demo-rig n=20 · probe-suite n=75

100%n=25100%[96%–100%] n=95100%n=95$52.8270%n=150%n=52026-08-19
grok/grok-4-1-fast

fixtures: probe-suite n=30 — partial coverage

100%[89%–100%] n=30100%n=30$0.10427%n=15100%n=52026-08-02
grok/grok-4.3

fixtures: probe-suite n=30 — partial coverage

100%[89%–100%] n=30100%n=30$0.54420%n=15100%n=52026-08-02
grok/grok-4.5

fixtures: probe-suite n=30 — partial coverage

100%[89%–100%] n=30100%n=30$1.27827%n=15100%n=52026-08-02

data as of 2026-08-19 · vintage 2026-08-19T14:18:42.501Z · live @ 41d1634

Rerun it yourself: every published stat ships with its raw samples.

A reproduction passes when your rerun lands inside the published interval. Disagreements are contributions — file an issue with your result attached. From zero: clone the public repo, run one probe with your own key (the --envelope flag caps the spend in USD), then verify it against the published result.

shell — from zero
git clone https://github.com/modelrig/modelrig && cd modelrig
export OPENAI_API_KEY="your-key"   # or any one provider you hold a key for
npx modelrig-probes run --model openai/gpt-5.4-nano --class schema --out my-results --envelope 1
npx modelrig-probes verify my-results/<your-result>.json --against probes/results/<published>.json

Per-model detail — declared flags, probed rates, raw sample records, and call notes — lives in registry.json in the public repo.

Registry data © ModelRig contributors, licensed CC BY 4.0. Cite freely; link the data.