ModelRig / what ModelRig does
Stop flying blind on AI spend and quality.
You found out about the bill from the invoice. You found out about the regression from a customer. Both facts were knowable the moment they happened — if every call carried its cost, its task, and its conformance verdict. Ours do.
Metered by task
Tokens in, out, and cached; cost; latency; TTFB; failure class — rolled up by route, tag, and project. Metered by task, not by API key.
The cost of conformance
What does a usable answer cost, repairs included? The effective cost of conformance is the number a vendor pricing page can't give you — it depends on whether the model actually passes your schema.
Probed, not repeated
Capability claims are probed against real fixtures, and the declared-vs-probed discrepancies are published. The probe suite is reproducible; the leaderboard is its public face.
When the market moves, you know
A price or capability shift triggers a suggested re-test. You're not crazy — the model changed; the probe is published and dated.
Your observability stack keeps working
OTLP out to Langfuse, Datadog, or any collector you already run. ModelRig is the meter, not a replacement for your traces.
Probed in the open
Every number is reproducible: the live leaderboard and the published discrepancy set are the monitoring engine's public face — rerun any of it yourself.
the live leaderboard →what each model costs →what each column means →rerun it yourself →