GPT-5
GPT-5 by OpenAI: $1.25 input and $10 output per 1M tokens, 400K context. Call it through the odnoga LLM gateway.
GPT-5 is served by OpenAI and called through odnoga with the same OpenAI-compatible request shape as every other model in the catalog. It is a reasoning model, so it spends extra output tokens working through a problem before answering — budget for higher output cost on hard tasks. It accepts images alongside text. It supports tool and function calling. It can be forced to return structured JSON. Its 400K-token context window sets how much input you can send in one request. Cached input is billed at $0.125 per 1M tokens, so repeated prefixes cost less.
Specification and price
| Vendor | OpenAI |
|---|---|
| Model ID | gpt-5 |
| Context window | 400K tokens |
| Max output | 128K tokens |
| Capabilities | Reasoning, Vision, Tools, JSON mode, Streaming |
| Per 1M tokens | Vendor cost | Your price on Free+7% |
|---|---|---|
| Input | $1.25 | $1.34 |
| Output | $10 | $10.75 |
| Cached input | $0.125 | $0.134 |
Vendor cost is the list price per million tokens as recorded in the odnoga catalog; your price applies your plan margin with the same formula that bills every request — see pricing. Pricing
Call it through odnoga
const res = await fetch("https://api.odnoga.com/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.ODNOGA_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-5",
messages: [{ role: "user", content: "Hello" }],
}),
});
Same request shape for every vendor in the catalog — swap the model id and odnoga handles keys, routing, limits and cost accounting.
How to use GPT-5
Derived from the odnoga catalog record for this model.
Best for
- Multi-step problems where the answer has to be worked out: planning, debugging, data reconciliation, analysis with intermediate steps.
- Work that mixes images with text — screenshots, scanned documents, charts, product photos.
- Agents and workflows that call your own functions, because the model supports tool calling.
- Machine-readable output you can write straight into a database, using enforced JSON.
- Long inputs: a 400K-token window fits whole contracts, codebases or transcripts in one request.
- User-facing chat where partial output should appear while the model is still writing.
Not the right pick when
- Simple, high-frequency calls — reasoning spends extra output tokens, so a non-reasoning model in the same catalog is usually cheaper and faster.
- Anything where a wrong answer is costly without a human check — no model in the catalog removes that requirement.
Practical tips through odnoga
- 01Pin the model id in a managed prompt version, so a model swap is a version change you can compare and roll back, not an edit in application code.
- 02Compare it against 2–8 other models on the same frozen test cases in the evaluation laboratory before you make it the production default.
- 03Keep the stable part of your prompt at the front: cached input is billed at $0.125 per 1M tokens instead of $1.25.
- 04For repeated identical deterministic calls, odnoga answer reuse returns the stored answer and bills no vendor tokens — turn it off for creative output.
- 05Ask for JSON through the response format rather than in the prompt text — the schema is enforced instead of suggested.
- 06Budget for output tokens: reasoning happens on the output side, so a short answer can still be an expensive call.
- 07Set a fallback model on the route so a vendor incident degrades quality instead of returning an error, and a per-tenant budget so one caller cannot spend the month.
What a month costs
1,000 calls a month, 10,000 input tokens and 2,000 output tokens each, at your Free price (vendor cost +7%):
| Input (10M tokens) | $13.44 |
|---|---|
| Output (2M tokens) | $21.51 |
| Your cost per month on Free | $34.95 |
Vendor list cost $32.50 + odnoga margin $2.45 (+7%). Cached input or answer reuse lowers it; odnoga records both numbers per request.
GPT-5 compared
| Model | Context | Input / 1M | Output / 1M | Capabilities |
|---|---|---|---|---|
| GPT-5 | 400K | $1.25 | $10 | Reasoning, Vision, Tools, JSON mode, Streaming |
| ChatGPT (chat-latest) | 400K | $5 | $30 | Vision, Tools, JSON mode, Streaming |
| GPT Image 1 mini | — | $2 | $8 | Vision |
Questions
- How much does GPT-5 cost per 1M tokens?
- OpenAI lists $1.25 / $10 per 1M input / output tokens in the odnoga catalog. Through odnoga you pay that vendor price plus your plan margin, and every request is recorded with both numbers.
- What is GPT-5 best for?
- Multi-step problems where the answer has to be worked out: planning, debugging, data reconciliation, analysis with intermediate steps. Work that mixes images with text — screenshots, scanned documents, charts, product photos. Agents and workflows that call your own functions, because the model supports tool calling.
- Can I switch to GPT-5 without changing my code?
- Yes. odnoga exposes one OpenAI-compatible endpoint, so switching means sending "gpt-5" as the model id — or changing it in the managed prompt version, with no application deploy.
- How large is the GPT-5 context window?
- 400K tokens of input, with up to 128K tokens of output per response.
All models · Pricing · Docs · Compare models in the evaluation lab