What Does an AI API Really Cost? Chat, RAG and Agent Budgets
A price per million tokens does not tell you what an AI product will cost. Conversation history, repeated context, agent steps, search tools, caching and retries can change the bill by an order of magnitude. This guide turns model pricing into a practical budget for chat, RAG and agents.
Key points
- Calculate input, cached input and output separately, then add tool calls and cache storage.
- A long chat history is billed again on every request unless it is compressed or cached.
- One user action in an agent may trigger 5–20 model calls.
- Routing, batch processing and prompt caching often save more than switching providers.
The basic cost formula
Cost = input tokens × input rate + output tokens × output rate + tools + cache storage.
Token rates are usually quoted per one million tokens:
(input_tokens / 1,000,000 × input_price) + (output_tokens / 1,000,000 × output_price)
If a request contains 3,000 input tokens and returns 700 tokens, with hypothetical rates of $1 per million input tokens and $5 per million output tokens, that call costs about $0.0065. At 100,000 calls, the same pattern costs $650 before search, retries or agent steps.
Open the live AI API price comparison and workload calculator →
Why chat gets more expensive over time
The user sees one new message. The API may receive the full conversation, a system prompt, retrieved documents and tool definitions. By message twenty, the input can be ten times longer than it was at the start.
Control this growth by limiting recent turns, summarising older context, caching stable prompts and logging actual token usage for every request.
Three budget patterns
| Product | Cost drivers | Main risk | Optimisation |
|---|---|---|---|
| Support chat | History, knowledge base, answer | Resending long conversations | Summaries, caching, a small classifier |
| Document RAG | Search, passages, generation | Too much irrelevant context | Better retrieval, reranking, passage limits |
| AI agent | Planning, tools, verification | Unbounded loops and retries | Step caps, routing and error budgets |
RAG costs are not just model costs
A retrieval system also pays for embeddings, vector search, reranking, index storage and the passages sent to the model. Returning twenty broad passages instead of five precise ones can cost more and reduce answer quality. The useful metric is therefore cost per correct answer, not cost per token.
An agent call is a chain of calls
An agent can plan, search, call code, inspect a result, recover from an error and then write the final response. Track the average and maximum number of steps, the share of steps that require a frontier model, paid tool calls, retries and the context carried between steps.
Production agents need hard limits for spend, elapsed time and action count. A rare runaway loop can erase a full day of optimisation.
Where the largest savings come from
Route by task
Classification, extraction, moderation and simple summaries rarely need the most expensive model. Reserve stronger models for hard reasoning, ambiguous cases and final verification.
Cache repeated context
System instructions and reference documents may dominate the input. Prompt caching cuts repeated processing, although cache storage can have its own rate. It pays off only when the same context is reused often enough.
Move offline work to Batch
Batch pricing is designed for asynchronous work such as catalogue enrichment, archive classification, evaluation and overnight reports. It is a poor fit for interactive chat but a strong default for workloads that can wait.
Limit the output
Output tokens are often more expensive than input. Clear schemas, length limits and structured JSON reduce both model spend and downstream parsing work.
API or a local model?
An API is attractive for variable demand because there is no hardware, deployment or idle-GPU cost. Local inference becomes more compelling when demand is stable, data cannot leave the device, or control matters more than access to the strongest model.
Compare total cost: hardware, power, engineering, monitoring, redundancy and utilisation. For personal hardware, start with our local model selector by memory and device.
Metrics every AI product needs
- cost per user action;
- input and output tokens by model and feature;
- cached-input share;
- average agent steps;
- cost per successful result;
- retry and error spend;
- spend by customer, team or project.
A monthly cap does not protect you from one expensive request. Use both an overall budget and a per-operation ceiling.
Choose models on real tasks
The cheapest listed model can cost more in production if it needs extra context, retries or human review. A more expensive model can be cheaper when it succeeds on the first attempt.
- Collect 50–200 representative tasks.
- Define what a correct result means.
- Run the same cases across candidate models.
- Measure quality, latency and total cost.
- Select a model or routing rule for each task class.
Sources and freshness
Prices change frequently. AI Feed keeps the tariff table separate from this methodology and checks the official pricing pages from OpenAI, Anthropic, Google Gemini and xAI. The live table shows its update time on the AI API price monitor.
The conclusion: AI product cost is usually an architecture problem before it is a provider problem. Measure calls and tokens first, then improve context, routing and processing mode before comparing model list prices.