ModelRig

The operating layer for AI workstreams

See. Grade. Improve.

One compliance-aware system for your whole AI workstream — observability, A/B route testing, typed fallbacks, and an inspectable artifact store — proving the cheapest model that still passes.

modelrig — probed registry, ranked by effective cost
modelschema conformancemethod▸effective $ / 1K conformantprobed
deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506

fixtures: demo-rig n=20 · probe-suite n=75

95%[88%–98%] n=95$0.0492026-08-18
deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct

fixtures: demo-rig n=20 · probe-suite n=75

93%[86%–96%] n=95$0.0602026-08-18
deepinfra/google/gemma-4-31B-it

fixtures: demo-rig n=20 · probe-suite n=75

100%[96%–100%] n=95$0.0822026-08-18
Rendered live from the public probe registry — real probe data, not a mockup.

What ModelRig does

Most AI products overpay for models because nobody can prove what cheaper ones can do. ModelRig can conduct regular probes of a wide range of models against your actual workload — so you keep the quality while cutting the bill.

  • Conformance & cost monitoring

    Every call metered by task, not by API key. Every capability claim probed, not repeated — and the discrepancies published.

  • A/B-tested routing

    Challengers race your incumbent on replayed production traffic, gated by confidence intervals; the winner serves behind declared candidates, typed fallback, and hard spend stops. The swap is a pull request, never a surprise.

  • Artifact scoring & storage

    early access

    Your agents' work products — versioned, hashed, lineaged, scored. “Why did it do that?” gets an answer.

    Running on our own production pipeline today — customer zero, as always. Early access opens with the next release. The pitch is one sentence: when someone asks why the agent did that, you pull up the run and show them.

Set up by your coding agent, in an environment where isolation is tested, not asserted.

The margin loop.

Continuous savings, confirmed performance.

  1. 01

    Observe first — no migration

    Point the exporter at us and see your runs before you change anything.

  2. 02

    Replay your real traffic against challengers

    Shadow bake-offs on logged requests; production never sees the test.

  3. 03

    Swap via a pull request you review

    Never silent; revert is one commit.

  4. 04

    The test re-runs itself

    A price or model change triggers a suggested re-test — last quarter's decision doesn't quietly rot.

Models that pass every schema fixture range from $0.082 to $52.827 per 1,000 conformant outputs on the current leaderboard — a 642× spread. Which end is your stack on?

Declared capabilities don't always hold, either: 32 of 37 probed models carry at least one declared-vs-probed discrepancy. See every discrepancy.

01

Point your coding agent at it

Paste one prompt or run modelrig init; your calls become routes in minutes.

prompt — paste into your coding agent
Migrate this project to ModelRig routes.
… (copy for the full 6-step prompt)

02

Prove it on your own traffic

Replay logged requests against cheaper candidates; zero production risk.

03

The margin keeps widening

Swap behind the stable route, and the corpus keeps working: new models get probed, your routes get re-checked, and every run makes the next decision sharper. You never touch app code again.

The leaderboard is the credential.

Models ranked by effective cost of conformance — the effective $ per 1,000 schema-conformant outputs, repairs included. Sampled statistics with 95% CIs, never single-shot verdicts.

modelschema conformancemethod▸value accuracymethod▸effective $ / 1K conformantgroundedmethod▸cache hitsmethod▸probed
deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506

fixtures: demo-rig n=20 · probe-suite n=75

95%[88%–98%] n=9589%n=95$0.0490%n=150%n=52026-08-18
deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct

fixtures: demo-rig n=20 · probe-suite n=75

93%[86%–96%] n=95100%n=95$0.060100%n=150%n=52026-08-18
deepinfra/google/gemma-4-31B-it

fixtures: demo-rig n=20 · probe-suite n=75

100%[96%–100%] n=9598%n=95$0.0820%n=52026-08-18
deepseek/deepseek-chat

fixtures: demo-rig n=20 · probe-suite n=75

95%[88%–98%] n=9595%n=95$0.084100%n=15100%n=52026-08-19
deepseek/deepseek-v4-flash

fixtures: demo-rig n=20 · probe-suite n=75

96%[90%–98%] n=95100%n=95$0.09380%n=15100%n=52026-08-19

Rerun it yourself: every published stat ships with its raw samples. Rerun it yourself

Coverage: v0 corpus, finance-weighted (seeded from our first production customer) — schema corpus per fully-probed model: 95 samples — probe-suite n=75 · demo-rig n=20, 25 of them hard-tier. Contribute yours

data as of 2026-08-19 · vintage 2026-08-31T07:59:01.582Z · live @ d229489

Your traffic makes the next decision cheaper.

The corpus keeps working between your deploys: new models get probed, your routes get re-checked, and every run makes the next decision sharper.

97%

of the probe matrix is filled today — every registered model against schema conformance, grounding, and caching (108 of 111 cells)

Full — of a deliberately small matrix. The thin part is the corpus, not the models: the fixtures are v0 and finance-weighted, and on a schema nobody has probed yet, your runs would be the first evidence in the cell.

Your first bake-off doesn't start cold. It starts on what everyone else's traffic already proved about schemas shaped like yours.

Your traffic makes the next decision cheaper — for you first. Same flat 2% either way; the corpus is the product, not a coupon. What we keep, and what you get for it.

We built a 16-step financial research pipeline first. Catching a model whose declared schema support was fiction was the easy part — production told us fast. What we could never keep up with was testing the stream of emerging, cheaper models to find the ones that pass every requirement — not without risking production. We built ModelRig because we needed it.

37

models probed across schema conformance, grounding, and caching

4035

probe samples published, raw per-sample records included

45

declared-vs-probed discrepancies published — every stat re-verifiable

counts computed from registry.json as of 2026-08-31

Three steps. No leap of faith.

  1. step 1

    Point your coding agent at it

    Paste one prompt or run modelrig init; your calls become routes in minutes.

  2. step 2

    Prove it on your own traffic

    Replay logged requests against cheaper candidates; zero production risk.

  3. step 3

    The margin keeps widening

    Swap behind the stable route, and the corpus keeps working: new models get probed, your routes get re-checked, and every run makes the next decision sharper. You never touch app code again.

  • Never a silent swap.
  • Your data never leaves without you.
  • Everything we hold, you can see and delete.
  • One flat fee — 2% of what you run through us at list prices, your keys or ours, plus the standard Stripe cost on the charge (the processor's, not ours), and on your own keys the first million requests each month are free.
  • Your routes live in your git — leave anytime.

All hosted data lives on SOC 2 Type II–audited infrastructure, encrypted in transit and at rest, row-scoped to your organization by the database itself — and everything we hold, you can see, export, and delete. Security →

Every model decision, defensible.

From gambling on model claims to operating on proof — costs governed, conformance enforced, and when the bill comes, you defend it with a chart.

The cheapest model that actually passes — actually is the product.

effective cost of conformance — $ / 1K conformant outputs
modelschema conformancemethod▸effective $ / 1K conformantprobed
deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506

fixtures: demo-rig n=20 · probe-suite n=75

95%[88%–98%] n=95$0.0492026-08-18
deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct

fixtures: demo-rig n=20 · probe-suite n=75

93%[86%–96%] n=95$0.0602026-08-18
deepinfra/google/gemma-4-31B-it

fixtures: demo-rig n=20 · probe-suite n=75

100%[96%–100%] n=95$0.0822026-08-18
deepseek/deepseek-chat

fixtures: demo-rig n=20 · probe-suite n=75

95%[88%–98%] n=95$0.0842026-08-19
deepseek/deepseek-v4-flash

fixtures: demo-rig n=20 · probe-suite n=75

96%[90%–98%] n=95$0.0932026-08-19
The effective-cost column, live from the published leaderboard data.

One flat 2%, your keys or ours. On your own keys, your first million requests each month are free.