Leaderboard
Probed capability facts, ranked by effective cost of conformance — the effective $ per 1,000 schema-conformant outputs, repairs included. Every rate is a sampled statistic with n and a 95% CI. Column headers link to what each one means; model and cost cells link to what each model costs.
The public face of conformance & cost monitoring: the same probes that rank models in the open watch your own traffic in the console. Conformance & cost monitoring.
Ranked by measured value accuracy on the hard-tier fixture subset — a descriptive statistic, not a RigIndex rank. what this measures. what a RigIndex rank is.
27 models tie at 100% on this metric; ties are broken by effective cost, so the order you see is the cost order.
Accuracy-ranked: highest measured accuracy first. Back to cheapest conformant.
data as of 2026-08-19 · vintage 2026-08-19T14:18:42.501Z · live @ 41d1634
Rerun it yourself: every published stat ships with its raw samples.
A reproduction passes when your rerun lands inside the published interval. Disagreements are contributions — file an issue with your result attached. From zero: clone the public repo, run one probe with your own key (the --envelope flag caps the spend in USD), then verify it against the published result.
git clone https://github.com/modelrig/modelrig && cd modelrig export OPENAI_API_KEY="your-key" # or any one provider you hold a key for npx modelrig-probes run --model openai/gpt-5.4-nano --class schema --out my-results --envelope 1 npx modelrig-probes verify my-results/<your-result>.json --against probes/results/<published>.json
Per-model detail — declared flags, probed rates, raw sample records, and call notes — lives in registry.json in the public repo.
Registry data © ModelRig contributors, licensed CC BY 4.0. Cite freely; link the data.