ModelRig

ModelRig / what ModelRig does

Stop flying blind on AI spend and quality.

You found out about the bill from the invoice. You found out about the regression from a customer. Both facts were knowable the moment they happened — if every call carried its cost, its task, and its conformance verdict. Ours do.

Metered by task

Tokens in, out, and cached; cost; latency; TTFB; failure class — rolled up by route, tag, and project. Metered by task, not by API key.

The cost of conformance

What does a usable answer cost, repairs included? The effective cost of conformance is the number a vendor pricing page can't give you — it depends on whether the model actually passes your schema.

Probed, not repeated

Capability claims are probed against real fixtures, and the declared-vs-probed discrepancies are published. The probe suite is reproducible; the leaderboard is its public face.

When the market moves, you know

A price or capability shift triggers a suggested re-test. You're not crazy — the model changed; the probe is published and dated.

Your observability stack keeps working

OTLP out to Langfuse, Datadog, or any collector you already run. ModelRig is the meter, not a replacement for your traces.

Probed in the open

Every number is reproducible: the live leaderboard and the published discrepancy set are the monitoring engine's public face — rerun any of it yourself.

the live leaderboard →what each model costs →what each column means →rerun it yourself →