Models & pricing
Provider list prices for every model in the registry — input, output and cached-input per million tokens (Mtok). Generated from the pinned pricing map with the leaderboard, never hand-authored.
List prices as of 2026-08-19, from the pinned pricing map.
The public face of conformance & cost monitoring: the same probes that rank models in the open watch your own traffic in the console. Conformance & cost monitoring.
List price is not your price. These are provider list prices. Our fee is a flat 2% of usage valued at them; a discount you have negotiated with a provider is yours to keep.
Host lanes are distinct. The same open weights served by two hosts are two entries with two prices — the leading segment of a model's name is the host serving it.
Anthropic
The Claude model family, first-party API.
| model | host | input $/Mtok | output $/Mtok | cached input $/Mtok | as of | links |
|---|---|---|---|---|---|---|
| claude-fable-5 | Anthropic | $10.000 | $50.000 | $1.000 | 2026-08-19 | leaderboard · registry |
| claude-haiku-4-5 | Anthropic | $1.000 | $5.000 | $0.100 | 2026-08-19 | leaderboard · registry |
| claude-opus-5 | Anthropic | $5.000 | $25.000 | $0.500 | 2026-08-19 | leaderboard · registry |
| claude-sonnet-5 | Anthropic | $2.000 | $10.000 | $0.200 | 2026-08-19 | leaderboard · registry |
Pricing source: litellm-pricing-map(pinned)+adapter-static
DeepInfra
A GPU host serving open-weight models from several creators.
| model | host | input $/Mtok | output $/Mtok | cached input $/Mtok | as of | links |
|---|---|---|---|---|---|---|
| MiniMaxAI/MiniMax-M3 | DeepInfra | $0.280 | $1.100 | — | 2026-08-19 | leaderboard · registry |
| Qwen/Qwen3-30B-A3B | DeepInfra | $0.080 | $0.290 | — | 2026-08-19 | leaderboard · registry |
| Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo | DeepInfra | $0.290 | $1.200 | — | 2026-08-19 | leaderboard · registry |
| google/gemma-4-31B-it | DeepInfra | $0.130 | $0.380 | — | 2026-08-19 | leaderboard · registry |
| meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 | DeepInfra | $0.150 | $0.600 | — | 2026-08-19 | leaderboard · registry |
| meta-llama/Llama-4-Scout-17B-16E-Instruct | DeepInfra | $0.080 | $0.300 | — | 2026-08-19 | leaderboard · registry |
| mistralai/Mistral-Small-3.2-24B-Instruct-2506 | DeepInfra | $0.075 | $0.200 | — | 2026-08-19 | leaderboard · registry |
| moonshotai/Kimi-K2.6 | DeepInfra | $0.750 | $3.500 | — | 2026-08-19 | leaderboard · registry |
| moonshotai/Kimi-K2.7-Code | DeepInfra | $0.680 | $3.400 | — | 2026-08-19 | leaderboard · registry |
| nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B | DeepInfra | $0.500 | $2.200 | — | 2026-08-19 | leaderboard · registry |
| openai/gpt-oss-120b | DeepInfra | $0.050 | $0.450 | — | 2026-08-19 | leaderboard · registry |
| zai-org/GLM-5.2 | DeepInfra | $0.750 | $2.400 | — | 2026-08-19 | leaderboard · registry |
Pricing source: litellm-pricing-map(pinned)+adapter-static
DeepSeek
The DeepSeek model family, first-party API.
| model | host | input $/Mtok | output $/Mtok | cached input $/Mtok | as of | links |
|---|---|---|---|---|---|---|
| deepseek-chat | DeepSeek | $0.280 | $0.420 | $0.028 | 2026-08-19 | leaderboard · registry |
| deepseek-reasoner | DeepSeek | $0.280 | $0.420 | $0.028 | 2026-08-19 | leaderboard · registry |
| deepseek-v4-flash | DeepSeek | $0.140 | $0.280 | $0.003 | 2026-08-19 | leaderboard · registry |
| deepseek-v4-pro | DeepSeek | $0.435 | $0.870 | $0.004 | 2026-08-19 | leaderboard · registry |
Pricing source: litellm-pricing-map(pinned)+adapter-static
Fireworks AI
A GPU host serving open-weight models from several creators.
| model | host | input $/Mtok | output $/Mtok | cached input $/Mtok | as of | links |
|---|---|---|---|---|---|---|
| glm-5p2 | Fireworks AI | $1.400 | $4.400 | $0.140 | 2026-08-19 | leaderboard · registry |
| gpt-oss-120b | Fireworks AI | $0.150 | $0.600 | $0.015 | 2026-08-19 | leaderboard · registry |
| kimi-k3 | Fireworks AI | $3.000 | $15.000 | $0.300 | 2026-08-19 | leaderboard · registry |
Pricing source: litellm-pricing-map(pinned)+adapter-static
The Gemini model family, via the Gemini API.
| model | host | input $/Mtok | output $/Mtok | cached input $/Mtok | as of | links |
|---|---|---|---|---|---|---|
| gemini-2.5-flash | $0.300 | $2.500 | $0.030 | 2026-08-19 | leaderboard · registry | |
| gemini-3-flash-preview | $0.500 | $3.000 | $0.050 | 2026-08-19 | leaderboard · registry | |
| gemini-3.1-flash-lite | $0.250 | $1.500 | $0.025 | 2026-08-19 | leaderboard · registry | |
| gemini-3.1-pro-preview | $2.000 | $12.000 | $0.200 | 2026-08-19 | leaderboard · registry |
Pricing source: litellm-pricing-map(pinned)+adapter-static
OpenAI
The GPT model family, first-party API.
| model | host | input $/Mtok | output $/Mtok | cached input $/Mtok | as of | links |
|---|---|---|---|---|---|---|
| gpt-5-mini | OpenAI | $0.250 | $2.000 | $0.025 | 2026-08-19 | leaderboard · registry |
| gpt-5.2 | OpenAI | $1.750 | $14.000 | $0.175 | 2026-08-19 | leaderboard · registry |
| gpt-5.4 | OpenAI | $2.500 | $15.000 | $0.250 | 2026-08-19 | leaderboard · registry |
| gpt-5.4-mini | OpenAI | $0.750 | $4.500 | $0.075 | 2026-08-19 | leaderboard · registry |
| gpt-5.4-nano | OpenAI | $0.200 | $1.250 | $0.020 | 2026-08-19 | leaderboard · registry |
| gpt-5.5 | OpenAI | $5.000 | $30.000 | $0.500 | 2026-08-19 | leaderboard · registry |
| gpt-5.5-pro | OpenAI | $30.000 | $180.000 | $3.000 | 2026-08-19 | leaderboard · registry |
Pricing source: litellm-pricing-map(pinned)+adapter-static
xAI
The Grok model family, first-party API.
| model | host | input $/Mtok | output $/Mtok | cached input $/Mtok | as of | links |
|---|---|---|---|---|---|---|
| grok-4-1-fast | xAI | $0.200 | $0.500 | $0.050 | 2026-08-19 | leaderboard · registry |
| grok-4.3 | xAI | $1.250 | $2.500 | $0.200 | 2026-08-19 | leaderboard · registry |
| grok-4.5 | xAI | $2.000 | $6.000 | $0.500 | 2026-08-19 | leaderboard · registry |
Pricing source: litellm-pricing-map(pinned)+adapter-static
data as of 2026-08-19 · vintage 2026-08-19T14:18:42.501Z · live @ 41d1634