Long-context models
Models that take 200,000 tokens or more in a single request — roughly a long book, a full codebase, or a year of support tickets. Big enough that chunking stops being mandatory and the architecture of your product can change.
A big window is a big invoice
The window is a ceiling, not an allowance. Filling a million-token context costs a million tokens of input at that model's rate, every single call — so the question is never "does it fit" but "what does it cost each time, multiplied by your traffic". Run that number before you design around it.
Caching is what makes it affordable
If the large part of your prompt is the same between calls — a manual, a schema, a contract — a model with a cached-input rate bills that part at a fraction after the first request. That single detail decides whether a long-context design is viable at volume, and it is why the cheapest list price is often not the cheapest model.
Long context versus retrieval
Retrieval is still cheaper per call and still wins on very large corpora. Long context wins on simplicity, on questions that need the whole document at once, and on anything where a retrieval miss produces a confidently wrong answer. Most production systems end up using both — retrieval to narrow, a long window to reason over what came back.
Vendor cost is the list price. Your price is that plus your plan margin — the same figures that appear on your invoice.
Every model here, cheapest first
| Model | Vendor | Context | In / 1M | Out / 1M | Cached in | Your price in | |
|---|---|---|---|---|---|---|---|
| GPT-5 nano gpt-5-nano | OpenAI | 400K | $0.05 | $0.4 | $0.005 | $0.054 | Details |
| GPT-4.1 nano gpt-4.1-nano | OpenAI | 1M | $0.1 | $0.4 | $0.025 | $0.108 | Details |
| Ministral 8B ministral-8b-latest | Mistral AI | 262K | $0.15 | $0.15 | — | $0.161 | Details |
| Mistral Small 3 mistral-small-latest | Mistral AI | 262K | $0.15 | $0.6 | — | $0.161 | Details |
| GPT-5.4 nano gpt-5.4-nano | OpenAI | 400K | $0.2 | $1.25 | $0.02 | $0.215 | Details |
| GPT-5.6 Luna gpt-5.6-luna | OpenAI | 400K | $0.2 | $1.20 | $0.02 | $0.215 | Details |
| Gemini 3.1 Flash-Lite gemini-3.1-flash-lite | Google AI | 1.0M | $0.25 | $1.50 | $0.025 | $0.269 | Details |
| GPT-5 mini gpt-5-mini | OpenAI | 400K | $0.25 | $2 | $0.025 | $0.269 | Details |
| Codestral codestral-latest | Mistral AI | 256K | $0.3 | $0.9 | — | $0.323 | Details |
| Gemini 3.5 Flash-Lite gemini-3.5-flash-lite | Google AI | 1.0M | $0.3 | $2.50 | $0.03 | $0.323 | Details |
| GPT-4.1 mini gpt-4.1-mini | OpenAI | 1M | $0.4 | $1.60 | $0.1 | $0.43 | Details |
| Gemini 3 Flash Preview gemini-3-flash-preview | Google AI | 1.0M | $0.5 | $3 | $0.05 | $0.538 | Details |
| Gemini 3.6 Flash gemini-3.6-flash | Google AI | 1.0M | $0.75 | $3.75 | $0.075 | $0.806 | Details |
| Gemini 3.7 Flash gemini-3.7-flash | Google AI | 1.0M | $0.75 | $3.75 | $0.075 | $0.806 | Details |
| GPT-5.4 mini gpt-5.4-mini | OpenAI | 400K | $0.75 | $4.50 | $0.075 | $0.806 | Details |
| Claude Haiku 4.5 claude-haiku-4-5 | Anthropic | 200K | $1 | $5 | $0.1 | $1.08 | Details |
| grok-build-0.1 grok-build-0.1 | xAI | 256K | $1 | $2 | $0.2 | $1.08 | Details |
| o3-mini o3-mini | OpenAI | 200K | $1.10 | $4.40 | $0.55 | $1.18 | Details |
| o4-mini o4-mini | OpenAI | 200K | $1.10 | $4.40 | $0.275 | $1.18 | Details |
| Gemini 2.5 Pro gemini-2.5-pro | Google AI | 1.0M | $1.25 | $10 | $0.125 | $1.34 | Details |
| GPT-5 gpt-5 | OpenAI | 400K | $1.25 | $10 | $0.125 | $1.34 | Details |
| gpt-5.1 gpt-5.1 | OpenAI | 400K | $1.25 | $10 | $0.125 | $1.34 | Details |
| grok-4.20-0309-non-reasoning grok-4.20-0309-non-reasoning | xAI | 1M | $1.25 | $2.50 | $0.2 | $1.34 | Details |
| grok-4.20-0309-reasoning grok-4.20-0309-reasoning | xAI | 1M | $1.25 | $2.50 | $0.2 | $1.34 | Details |
| grok-4.20-multi-agent-0309 grok-4.20-multi-agent-0309 | xAI | 1M | $1.25 | $2.50 | $0.2 | $1.34 | Details |
| grok-4.3 grok-4.3 | xAI | 1M | $1.25 | $2.50 | $0.2 | $1.34 | Details |
| Gemini 3.5 Flash gemini-3.5-flash | Google AI | 1.0M | $1.50 | $9 | $0.15 | $1.61 | Details |
| gpt-5.2 gpt-5.2 | OpenAI | 400K | $1.75 | $14 | $0.175 | $1.88 | Details |
| GPT-5.3 Codex gpt-5.3-codex | OpenAI | 400K | $1.75 | $14 | $0.175 | $1.88 | Details |
| Claude Sonnet 5 claude-sonnet-5 | Anthropic | 1M | $2 | $10 | $0.2 | $2.15 | Details |
| Gemini 3.1 Pro Preview gemini-3.1-pro-preview | Google AI | 1.0M | $2 | $12 | $0.2 | $2.15 | Details |
| GPT-4.1 gpt-4.1 | OpenAI | 1M | $2 | $8 | $0.5 | $2.15 | Details |
| GPT-5.6 Terra gpt-5.6-terra | OpenAI | 400K | $2 | $12 | $0.2 | $2.15 | Details |
| grok-4.5 grok-4.5 | xAI | 500K | $2 | $6 | $0.3 | $2.15 | Details |
| grok-4.6 grok-4.6 | xAI | 500K | $2 | $6 | $0.5 | $2.15 | Details |
| o3 o3 | OpenAI | 200K | $2 | $8 | $0.5 | $2.15 | Details |
| GPT-5.4 gpt-5.4 | OpenAI | 400K | $2.50 | $15 | $0.25 | $2.69 | Details |
| Claude Sonnet 4.5 claude-sonnet-4-5 | Anthropic | 200K | $3 | $15 | $0.3 | $3.23 | Details |
| Claude Sonnet 4.6 claude-sonnet-4-6 | Anthropic | 1M | $3 | $15 | $0.3 | $3.23 | Details |
| Grok 4 grok-4 | xAI | 256K | $3 | $15 | $0.75 | $3.23 | Details |
| Sonar Pro sonar-pro | Perplexity | 200K | $3 | $15 | — | $3.23 | Details |
| GPT-5.6 Sol gpt-5.6-sol | OpenAI | 400K | $4 | $20 | $0.4 | $4.30 | Details |
| ChatGPT (chat-latest) chat-latest | OpenAI | 400K | $5 | $30 | $0.5 | $5.38 | Details |
| Claude Opus 4.5 claude-opus-4-5-20251101 | Anthropic | 200K | $5 | $25 | $0.5 | $5.38 | Details |
| Claude Opus 4.6 claude-opus-4-6 | Anthropic | 1M | $5 | $25 | $0.5 | $5.38 | Details |
| Claude Opus 4.7 claude-opus-4-7 | Anthropic | 1M | $5 | $25 | $0.5 | $5.38 | Details |
| Claude Opus 4.8 claude-opus-4-8 | Anthropic | 1M | $5 | $25 | $0.5 | $5.38 | Details |
| Claude Opus 5 claude-opus-5 | Anthropic | 1M | $5 | $25 | $0.5 | $5.38 | Details |
| GPT-5.5 gpt-5.5 | OpenAI | 400K | $5 | $30 | $0.5 | $5.38 | Details |
| Claude Fable 5 claude-fable-5 | Anthropic | 1M | $10 | $50 | $1 | $10.75 | Details |
| Claude Fable 5.1 claude-fable-5-1 | Anthropic | 1M | $10 | $50 | $0.25 | $10.75 | Details |
| gpt-5-pro gpt-5-pro | OpenAI | 400K | $15 | $120 | — | $16.13 | Details |
| gpt-5.2-pro gpt-5.2-pro | OpenAI | 400K | $21 | $168 | — | $22.58 | Details |
| GPT-5.4 Pro gpt-5.4-pro | OpenAI | 400K | $30 | $180 | — | $32.26 | Details |
| GPT-5.5 Pro gpt-5.5-pro | OpenAI | 400K | $30 | $180 | — | $32.26 | Details |
| Lyria 3 Clip (30s) lyria-3-clip-preview | Google AI | 1.0M | Billed per media unit — see details | Details | |||
| Lyria 3 Pro (full song) lyria-3-pro-preview | Google AI | 1.0M | Billed per media unit — see details | Details | |||
Your price column is the vendor cost +7%.
Narrow it differently
The best price per token of context on the gateway.
The most consistent family on the gateway: every Claude model here reads images, calls tools and carries a cached-input rate.
Models that spend tokens thinking before they answer.
Or read how the gateway picks between them: routing and fallback, and what it costs: plans and margins.