How much does an AI chatbot cost?
Monthly estimate for 10,000 conversations averaging 1,200 input and 500 output tokens.
Compare one workload
Rates only become comparable when input, output and request count are identical. The calculation below uses the same workload for every model.
- Input per request
- 1,200
- Output per request
- 500
- Requests
- 10,000
Current model costs
| Model | Input / 1M | Output / 1M | Workload estimate |
|---|---|---|---|
| Mistral: Mistral NemoMistralai | $0.02 | $0.03 | $0.38 |
| OpenAI: gpt-oss-20bOpenai | $0.02 | $0.09 | $0.67 |
| DeepSeek: DeepSeek V4 Flash 0423Deepseek | $0.04 | $0.08 | $0.92 |
| Mistral: Mistral Small 3Mistralai | $0.05 | $0.08 | $1 |
| Meta: Llama 3.1 8B InstructMeta-llama | $0.05 | $0.08 | $1 |
| Google: Gemma 3 4BGoogle | $0.05 | $0.1 | $1.1 |
| Qwen: Qwen3.7 FlashQwen | $0.03 | $0.13 | $1.01 |
| Mistral: Ministral 3 3B 2512Mistralai | $0.1 | $0.1 | $1.7 |
| Google: Gemma 3 12BGoogle | $0.05 | $0.15 | $1.35 |
| OpenAI: gpt-oss-120bOpenai | $0.04 | $0.17 | $1.29 |
Estimate uses published token rates. Caching, reasoning tokens, tools, images, routing and taxes can change the final bill.
What determines an AI chatbot API bill?
Start with the work unit you can measure: one conversation, document, article or agent run. Record both prompt and response length, then multiply by real monthly volume. Output often costs more than input, so a concise answer can matter more than a short prompt.
Price is only one constraint. Check context length, modalities, tool support, latency and output quality before choosing a production model. Test a small representative sample instead of relying on a single benchmark.