Cheapest AI APIs
Low-cost models for classification, extraction, drafts and high-volume requests.
Compare one workload
Rates only become comparable when input, output and request count are identical. The calculation below uses the same workload for every model.
- Input per request
- 2,000
- Output per request
- 1,000
- Requests
- 1,000
Current model costs
| Model | Input / 1M | Output / 1M | Workload estimate |
|---|---|---|---|
| Mistral: Mistral NemoMistralai | $0.02 | $0.03 | $0.07 |
| OpenAI: gpt-oss-20bOpenai | $0.02 | $0.09 | $0.13 |
| DeepSeek: DeepSeek V4 Flash 0423Deepseek | $0.04 | $0.08 | $0.17 |
| Mistral: Mistral Small 3Mistralai | $0.05 | $0.08 | $0.18 |
| Meta: Llama 3.1 8B InstructMeta-llama | $0.05 | $0.08 | $0.18 |
| Google: Gemma 3 4BGoogle | $0.05 | $0.1 | $0.2 |
| Qwen: Qwen3.7 FlashQwen | $0.03 | $0.13 | $0.19 |
| Mistral: Ministral 3 3B 2512Mistralai | $0.1 | $0.1 | $0.3 |
| Google: Gemma 3 12BGoogle | $0.05 | $0.15 | $0.25 |
| OpenAI: gpt-oss-120bOpenai | $0.04 | $0.17 | $0.24 |
Estimate uses published token rates. Caching, reasoning tokens, tools, images, routing and taxes can change the final bill.
How do you choose a low-cost AI API?
Start with the work unit you can measure: one conversation, document, article or agent run. Record both prompt and response length, then multiply by real monthly volume. Output often costs more than input, so a concise answer can matter more than a short prompt.
Price is only one constraint. Check context length, modalities, tool support, latency and output quality before choosing a production model. Test a small representative sample instead of relying on a single benchmark.