LLM API Cost Calculator
Estimate what your AI feature costs per request, per user and per month on Claude, GPT and Gemini APIs. Prices are the providers’ public list prices, checked on 2026-10-07.
Same workload on every model
Monthly cost of the workload above for each model in the table, cheapest first. Token counts are assumed equal across models, which they are not exactly (see FAQ).
| Model | Per request | Per user / month | Per month |
|---|
How the cost is calculated
Prices are quoted per million tokens (MTok). Output tokens include reasoning (“thinking”) tokens on the providers that bill them as output, which is the case on the Gemini API page (“Output price (including thinking tokens)”).
Price list used (checked 2026-10-07)
Standard paid tier, text, USD per 1M tokens. Batch column = batch input / batch output.
| Model | Provider | Input | Cached input | Output | Batch in / out | Note |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | $10.00 | $0.25 | $50.00 | $5.00 / $25.00 | |
| Claude Opus 5.5 | Anthropic | $4.00 | $0.20 | $20.00 | $2.00 / $10.00 | |
| Claude Sonnet 5.5 | Anthropic | $2.00 | $0.20 | $10.00 | $1.00 / $5.00 | |
| Claude Haiku 4.5 | Anthropic | $1.00 | $0.10 | $5.00 | $0.50 / $2.50 | |
| gpt-6-astra | OpenAI | $10.00 | $1.00 | $50.00 | $5.00 / $25.00 | short-context rate |
| gpt-6.1-sol | OpenAI | $2.00 | $0.10 | $10.00 | $1.00 / $5.00 | short-context rate |
| gpt-6-luna | OpenAI | $0.10 | $0.01 | $0.50 | $0.05 / $0.25 | short-context rate |
| Gemini 3.1 Pro Preview | $2.00 | $0.20 | $12.00 | $1.00 / $6.00 | prompts ≤ 200k tokens | |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | $0.38 / $1.88 | price through Dec 31, 2026; $1.50 / $7.50 from Jan 1, 2027 | |
| Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | $0.15 / $1.25 |
Sources: Anthropic pricing (cached input = “cache hits and refreshes”; batch cached = 50% batch discount stacked on the cache-hit price, as stated on that page) · OpenAI API pricing (short-context rates) · Gemini API pricing (paid tier; cached input = “context caching price”). Not included: cache-write premiums and cache storage fees, long-context surcharges, regional/data-residency uplifts, priority or fast tiers, tool calls, images and audio. Prices change: check the provider page before you commit to a budget.
FAQ
How many tokens is a word?
It depends on the tokenizer and the language, and each provider uses its own. Measure with the provider’s token-counting endpoint or the usage numbers returned by real API calls, then type those averages here. Anthropic notes that the tokenizer of Claude 4.7 and later models “produces approximately 30% more tokens for the same text” than earlier Claude models, so the same prompt is not the same token count everywhere.
What does the cached share mean?
If many requests start with the same long prefix (system prompt, documents, tool definitions), the providers can bill the repeated part at a lower “cached input” rate. Enter the percentage of input tokens you expect to be read from cache. The calculator does not add the cost of writing to the cache (Anthropic bills cache writes above the base input price; OpenAI lists separate cache-write prices; Gemini charges storage per hour), so for low hit rates it slightly underestimates.
When should I tick “Batch API prices”?
When requests do not need an immediate answer (nightly jobs, bulk classification, evaluations). All three providers list discounted batch prices; the table above shows them per model.
Why is my real invoice different?
Real traffic varies: some users send far more requests, retries cost tokens, reasoning models can produce many hidden output tokens, and long prompts may fall into a higher long-context price band on some models. Use the result as an order of magnitude and compare it with your provider’s usage dashboard after launch.
How do I turn this into a gross margin for an AI SaaS?
Divide the cost per user by your plan price: the calculator shows it as “LLM cost as % of your price”. In a full plan you also need hosting, payment processing and support costs, plus the fact that token prices tend to change over time, which is what a monthly financial model handles.