The operating layer for AI workstreams
See. Grade. Improve.
One compliance-aware system for your whole AI workstream — observability, A/B route testing, typed fallbacks, and an inspectable artifact store — proving the cheapest model that still passes.
| modelⓘ | schema conformanceⓘmethod▸ | effective $ / 1K conformantⓘ | probedⓘ |
|---|---|---|---|
| deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506 fixtures: demo-rig n=20 · probe-suite n=75 | 95%[88%–98%] n=95 | $0.049 | 2026-08-18 |
| deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct fixtures: demo-rig n=20 · probe-suite n=75 | 93%[86%–96%] n=95 | $0.060 | 2026-08-18 |
| deepinfra/google/gemma-4-31B-it fixtures: demo-rig n=20 · probe-suite n=75 | 100%[96%–100%] n=95 | $0.082 | 2026-08-18 |
What ModelRig does
Most AI products overpay for models because nobody can prove what cheaper ones can do. ModelRig can conduct regular probes of a wide range of models against your actual workload — so you keep the quality while cutting the bill.
Conformance & cost monitoring
Every call metered by task, not by API key. Every capability claim probed, not repeated — and the discrepancies published.
A/B-tested routing
Challengers race your incumbent on replayed production traffic, gated by confidence intervals; the winner serves behind declared candidates, typed fallback, and hard spend stops. The swap is a pull request, never a surprise.
Artifact scoring & storage
early accessYour agents' work products — versioned, hashed, lineaged, scored. “Why did it do that?” gets an answer.
Running on our own production pipeline today — customer zero, as always. Early access opens with the next release. The pitch is one sentence: when someone asks why the agent did that, you pull up the run and show them.
Set up by your coding agent, in an environment where isolation is tested, not asserted.
The margin loop.
Continuous savings, confirmed performance.
01
Observe first — no migration
Point the exporter at us and see your runs before you change anything.
02
Replay your real traffic against challengers
Shadow bake-offs on logged requests; production never sees the test.
03
Swap via a pull request you review
Never silent; revert is one commit.
04
The test re-runs itself
A price or model change triggers a suggested re-test — last quarter's decision doesn't quietly rot.
Models that pass every schema fixture range from $0.082 to $52.827 per 1,000 conformant outputs on the current leaderboard — a 642× spread. Which end is your stack on?
Declared capabilities don't always hold, either: 32 of 37 probed models carry at least one declared-vs-probed discrepancy. See every discrepancy.
01
Point your coding agent at it
Paste one prompt or run modelrig init; your calls become routes in minutes.
Migrate this project to ModelRig routes. … (copy for the full 6-step prompt)
02
Prove it on your own traffic
Replay logged requests against cheaper candidates; zero production risk.
03
The margin keeps widening
Swap behind the stable route, and the corpus keeps working: new models get probed, your routes get re-checked, and every run makes the next decision sharper. You never touch app code again.
The leaderboard is the credential.
Models ranked by effective cost of conformance — the effective $ per 1,000 schema-conformant outputs, repairs included. Sampled statistics with 95% CIs, never single-shot verdicts.
| modelⓘ | schema conformanceⓘmethod▸ | value accuracyⓘmethod▸ | effective $ / 1K conformantⓘ | groundedⓘmethod▸ | cache hitsⓘmethod▸ | probedⓘ |
|---|---|---|---|---|---|---|
| deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506 fixtures: demo-rig n=20 · probe-suite n=75 | 95%[88%–98%] n=95 | 89%n=95 | $0.049 | 0%n=15 | 0%n=5⚠ | 2026-08-18 |
| deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct fixtures: demo-rig n=20 · probe-suite n=75 | 93%[86%–96%] n=95 | 100%n=95 | $0.060 | 100%n=15 | 0%n=5⚠ | 2026-08-18 |
| deepinfra/google/gemma-4-31B-it fixtures: demo-rig n=20 · probe-suite n=75 | 100%[96%–100%] n=95 | 98%n=95 | $0.082 | — | 0%n=5⚠ | 2026-08-18 |
| deepseek/deepseek-chat declared ≠ probed · schema_served_via_json_modedeclared ≠ probed · citations_without_declared_search fixtures: demo-rig n=20 · probe-suite n=75 | 95%[88%–98%] n=95 | 95%n=95 | $0.084 | 100%n=15 | 100%n=5⚠ | 2026-08-19 |
| deepseek/deepseek-v4-flash declared ≠ probed · schema_served_via_json_modedeclared ≠ probed · citations_without_declared_search fixtures: demo-rig n=20 · probe-suite n=75 | 96%[90%–98%] n=95 | 100%n=95 | $0.093 | 80%n=15 | 100%n=5⚠ | 2026-08-19 |
Rerun it yourself: every published stat ships with its raw samples. Rerun it yourself
Coverage: v0 corpus, finance-weighted (seeded from our first production customer) — schema corpus per fully-probed model: 95 samples — probe-suite n=75 · demo-rig n=20, 25 of them hard-tier. Contribute yours
data as of 2026-08-19 · vintage 2026-08-31T07:59:01.582Z · live @ d229489
Your traffic makes the next decision cheaper.
The corpus keeps working between your deploys: new models get probed, your routes get re-checked, and every run makes the next decision sharper.
97%
of the probe matrix is filled today — every registered model against schema conformance, grounding, and caching (108 of 111 cells)
Full — of a deliberately small matrix. The thin part is the corpus, not the models: the fixtures are v0 and finance-weighted, and on a schema nobody has probed yet, your runs would be the first evidence in the cell.
Your first bake-off doesn't start cold. It starts on what everyone else's traffic already proved about schemas shaped like yours.
Your traffic makes the next decision cheaper — for you first. Same flat 2% either way; the corpus is the product, not a coupon. What we keep, and what you get for it.
We built a 16-step financial research pipeline first. Catching a model whose declared schema support was fiction was the easy part — production told us fast. What we could never keep up with was testing the stream of emerging, cheaper models to find the ones that pass every requirement — not without risking production. We built ModelRig because we needed it.
37
models probed across schema conformance, grounding, and caching
4035
probe samples published, raw per-sample records included
45
declared-vs-probed discrepancies published — every stat re-verifiable
counts computed from registry.json as of 2026-08-31
Three steps. No leap of faith.
step 1
Point your coding agent at it
Paste one prompt or run modelrig init; your calls become routes in minutes.
step 2
Prove it on your own traffic
Replay logged requests against cheaper candidates; zero production risk.
step 3
The margin keeps widening
Swap behind the stable route, and the corpus keeps working: new models get probed, your routes get re-checked, and every run makes the next decision sharper. You never touch app code again.
- Never a silent swap.
- Your data never leaves without you.
- Everything we hold, you can see and delete.
- One flat fee — 2% of what you run through us at list prices, your keys or ours, plus the standard Stripe cost on the charge (the processor's, not ours), and on your own keys the first million requests each month are free.
- Your routes live in your git — leave anytime.
All hosted data lives on SOC 2 Type II–audited infrastructure, encrypted in transit and at rest, row-scoped to your organization by the database itself — and everything we hold, you can see, export, and delete. Security →
Every model decision, defensible.
From gambling on model claims to operating on proof — costs governed, conformance enforced, and when the bill comes, you defend it with a chart.
The cheapest model that actually passes — actually is the product.
| modelⓘ | schema conformanceⓘmethod▸ | effective $ / 1K conformantⓘ | probedⓘ |
|---|---|---|---|
| deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506 fixtures: demo-rig n=20 · probe-suite n=75 | 95%[88%–98%] n=95 | $0.049 | 2026-08-18 |
| deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct fixtures: demo-rig n=20 · probe-suite n=75 | 93%[86%–96%] n=95 | $0.060 | 2026-08-18 |
| deepinfra/google/gemma-4-31B-it fixtures: demo-rig n=20 · probe-suite n=75 | 100%[96%–100%] n=95 | $0.082 | 2026-08-18 |
| deepseek/deepseek-chat declared ≠ probed · schema_served_via_json_modedeclared ≠ probed · citations_without_declared_search fixtures: demo-rig n=20 · probe-suite n=75 | 95%[88%–98%] n=95 | $0.084 | 2026-08-19 |
| deepseek/deepseek-v4-flash declared ≠ probed · schema_served_via_json_modedeclared ≠ probed · citations_without_declared_search fixtures: demo-rig n=20 · probe-suite n=75 | 96%[90%–98%] n=95 | $0.093 | 2026-08-19 |
One flat 2%, your keys or ours. On your own keys, your first million requests each month are free.