Free calculator
Self-hosting an LLM vs paying for an API: break-even calculator
Monthly cost of serving your token volume on rented GPUs versus an API, the volume where they cross, and the throughput self-hosting would need to match the API. API and GPU list prices are linked to their sources and every input is editable.
No paid links on this page: vendor links go straight to the vendor. Affiliate policy.
Self-host vs API break-even calculator
Runs in your browser. API list prices read 7 October 2026 and GPU on-demand prices read 7 October 2026, each linked to its source. Every price is editable.
API list prices behind the calculator
Standard real-time tier, US dollars per million tokens, read from each provider's own pricing page on 7 October 2026. "Check provider" means the page states no price (for example "Contact sales"). Batch, priority and off-peak prices are in the notes. GPU prices are on the GPU price table. Raw data: api-prices.json.
| Provider | Model | Input | Output | Cached input | Open weights | Source |
|---|---|---|---|---|---|---|
| OpenAI | gpt-5.4-mini | $0.75 | $4.50 | $0.075 | No | 2026-10-07 |
| OpenAI | gpt-5.4-nano | $0.20 | $1.25 | $0.02 | No | 2026-10-07 |
| OpenAI | gpt-6-astra | $10.00 | $50.00 | $1.00 | No | 2026-10-07 |
| OpenAI | gpt-6-luna | $0.10 | $0.50 | $0.01 | No | 2026-10-07 |
| OpenAI | gpt-6.1-sol | $2.00 | $10.00 | $0.10 | No | 2026-10-07 |
| Anthropic | Claude Fable 5.1 | $10.00 | $50.00 | $0.25 | No | 2026-10-07 |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 | No | 2026-10-07 |
| Anthropic | Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | No | 2026-10-07 |
| Anthropic | Claude Sonnet 5.5 | $2.00 | $10.00 | $0.20 | No | 2026-10-07 |
| Gemini 2.5 Pro (gemini-2.5-pro) | $1.25 | $10.00 | $0.125 | No | 2026-10-07 | |
| Gemini 3.1 Pro Preview (gemini-3.1-pro-preview) | $2.00 | $12.00 | $0.20 | No | 2026-10-07 | |
| Gemini 3.5 Flash (gemini-3.5-flash) | $1.50 | $9.00 | $0.15 | No | 2026-10-07 | |
| Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) | $0.30 | $2.50 | $0.03 | No | 2026-10-07 | |
| Gemini 3.8 Flash (gemini-3.8-flash) | $0.75 | $3.75 | $0.075 | No | 2026-10-07 | |
| DeepSeek | deepseek-flash (DeepSeek-V4.1-Flash) | $0.30 | $1.20 | $0.006 | Yes | 2026-10-07 |
| DeepSeek | deepseek-v4-pro (DeepSeek-V4-Pro-0813) | $1.32 | $3.96 | $0.044 | Yes | 2026-10-07 |
| Mistral | Mistral Large 3 | $0.50 | $1.50 | $0.05 | Yes | 2026-10-07 |
| Mistral | Mistral Large 4 | $0.68 | $2.09 | $0.07 | No | 2026-10-07 |
| Mistral | Mistral Medium 3.5 | $1.50 | $7.50 | $0.15 | Yes | 2026-10-07 |
| Mistral | Mistral Small 4 | $0.15 | $0.60 | $0.015 | Yes | 2026-10-07 |
| Together AI | DeepSeek V4 Flash 0731 | $0.14 | $0.28 | $0.03 | Yes | 2026-10-07 |
| Together AI | DeepSeek V4 Pro 0813 | $1.32 | $3.96 | $0.13 | Yes | 2026-10-07 |
| Together AI | DeepSeek V4.1 Flash | $0.30 | $1.20 | $0.006 | Yes | 2026-10-07 |
| Together AI | Gemma 4 31B | $0.39 | $0.97 | Not listed | Yes | 2026-10-07 |
| Together AI | GLM-5.3 | $1.40 | $4.40 | $0.26 | Yes | 2026-10-07 |
| Together AI | GLM-5.3-Flash | $0.15 | $0.50 | $0.03 | Yes | 2026-10-07 |
| Together AI | gpt-oss-120B | $0.15 | $0.60 | Not listed | Yes | 2026-10-07 |
| Together AI | Kimi K3 | $2.70 | $13.50 | $0.27 | Yes | 2026-10-07 |
| Together AI | Llama 3 8B Instruct Lite | $0.14 | $0.14 | Not listed | Yes | 2026-10-07 |
| Together AI | Llama 3.3 70B | $1.04 | $1.04 | Not listed | Yes | 2026-10-07 |
| Together AI | MiniMax M2.7 | $0.30 | $1.20 | $0.06 | Yes | 2026-10-07 |
| Together AI | MiniMax M3 | $0.30 | $1.20 | $0.06 | Yes | 2026-10-07 |
| Together AI | Qwen3 235B A22B Instruct 2507 FP8 Throughput | $0.20 | $0.60 | Not listed | Yes | 2026-10-07 |
| Together AI | Qwen3.5-397B-A17B | $0.60 | $3.60 | $0.35 | Yes | 2026-10-07 |
| Together AI | Qwen3.5 9B | $0.17 | $0.25 | Not listed | Yes | 2026-10-07 |
| Together AI | Qwen3.6-Plus | $0.50 | $3.00 | Not listed | No | 2026-10-07 |
| Together AI | Qwen3.8-2.4T-A95B | $2.00 | $6.00 | $0.25 | Yes | 2026-10-07 |
| Together AI | Qwen3.8 Flash | $0.15 | $0.47 | Not listed | Unclear | 2026-10-07 |
| Fireworks AI | Other base models: 4B to 16B parameters | $0.20 | $0.20 | Not listed | Yes | 2026-10-07 |
| Fireworks AI | Other base models: more than 16B parameters | $0.90 | $0.90 | Not listed | Yes | 2026-10-07 |
| Fireworks AI | Other base models: less than 4B parameters | $0.10 | $0.10 | Not listed | Yes | 2026-10-07 |
| Fireworks AI | Other base models: MoE 56.1B to 176B parameters (e.g. DBRX, Mixtral 8x22B) | $1.20 | $1.20 | Not listed | Yes | 2026-10-07 |
| Fireworks AI | Other base models: MoE up to 56B parameters (e.g. Mixtral 8x7B) | $0.50 | $0.50 | Not listed | Yes | 2026-10-07 |
| Fireworks AI | DeepSeek V4.1 Flash | $0.30 | $1.20 | $0.006 | Yes | 2026-10-07 |
| Fireworks AI | GLM 5.3 | $1.40 | $4.40 | $0.26 | Yes | 2026-10-07 |
| Fireworks AI | GLM 5.3 Flash | $0.15 | $0.50 | $0.03 | Yes | 2026-10-07 |
| Fireworks AI | OpenAI GPT OSS 120B | $0.15 | $0.60 | $0.015 | Yes | 2026-10-07 |
| Fireworks AI | Kimi K3 | $3.00 | $15.00 | $0.30 | Yes | 2026-10-07 |
| Fireworks AI | MiniMax M3 | $0.30 | $1.20 | $0.06 | Yes | 2026-10-07 |
| Fireworks AI | Qwen 3.8 Max | $2.00 | $6.00 | $0.25 | Unclear | 2026-10-07 |
| DeepInfra | DeepSeek-V3.2 | $0.26 | $0.38 | $0.13 | Yes | 2026-10-07 |
| DeepInfra | DeepSeek-V4-Flash | $0.09 | $0.18 | $0.018 | Yes | 2026-10-07 |
| DeepInfra | DeepSeek-V4-Flash-0731 | $0.06 | $0.18 | $0.015 | Yes | 2026-10-07 |
| DeepInfra | DeepSeek-V4-Pro | $1.30 | $2.60 | $0.10 | Yes | 2026-10-07 |
| DeepInfra | gemma-3-27b-it | $0.08 | $0.16 | Not listed | Yes | 2026-10-07 |
| DeepInfra | gemma-4-26B-A4B-it | $0.07 | $0.34 | Not listed | Yes | 2026-10-07 |
| DeepInfra | gemma-4-31B-it | $0.20 | $0.40 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Kimi-K2.6 | $0.75 | $3.50 | $0.15 | Yes | 2026-10-07 |
| DeepInfra | Kimi-K3 | $2.85 | $14.25 | $0.285 | Yes | 2026-10-07 |
| DeepInfra | Meta-Llama-3.1-8B-Instruct-Turbo | $0.02 | $0.04 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Llama-3.3-70B-Instruct-Turbo | $0.10 | $0.32 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Mistral-Small-3.2-24B-Instruct-2506 | $0.075 | $0.20 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Qwen3-14B | $0.12 | $0.24 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Qwen3-235B-A22B-Instruct-2507 | $0.09 | $0.55 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Qwen3-30B-A3B | $0.12 | $0.50 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Qwen3-32B | $0.08 | $0.28 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Qwen3.5-27B | $0.26 | $2.60 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Qwen3.5-35B-A3B | $0.14 | $1.00 | $0.05 | Yes | 2026-10-07 |
| DeepInfra | Qwen3.5-397B-A17B | $0.45 | $3.00 | $0.22 | Yes | 2026-10-07 |
| DeepInfra | Qwen3.5-9B | $0.10 | $0.15 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Qwen3.6-27B | $0.32 | $3.20 | Not listed | Yes | 2026-10-07 |
| DeepInfra | Qwen3.6-35B-A3B | $0.10 | $0.95 | $0.10 | Yes | 2026-10-07 |
| DeepInfra | Qwen3.8-Max | $1.65 | $4.951 | $0.206 | Unclear | 2026-10-07 |
| Groq | GPT OSS 120B (openai/gpt-oss-120b) | $0.15 | $0.60 | Not listed | Yes | 2026-10-07 |
| Groq | GPT OSS 20B (openai/gpt-oss-20b) | $0.075 | $0.30 | Not listed | Yes | 2026-10-07 |
| Groq | Llama 3.1 8B (llama-3.1-8b-instant) | Check provider | Check provider | Not listed | Yes | 2026-10-07 |
| Groq | Llama 3.3 70B (llama-3.3-70b-versatile) | Check provider | Check provider | Not listed | Yes | 2026-10-07 |
| Groq | MiniMax M2.7 (minimaxai/minimax-m2.7) | Check provider | Check provider | Not listed | Yes | 2026-10-07 |
| Groq | Qwen/Qwen3.8-27B (qwen/qwen3.8-27b) | $0.80 | $4.00 | Not listed | Yes | 2026-10-07 |
| Hyperstack | OpenAI gpt-oss-120b | $0.10 | $0.40 | Not listed | Yes | 2026-10-07 |
| Hyperstack | Llama 3.1 8B | $0.20 | $0.20 | Not listed | Yes | 2026-10-07 |
| Hyperstack | Llama 3.3 70B | $0.80 | $0.80 | Not listed | Yes | 2026-10-07 |
| Crusoe | Gemma 4 31B-it | $0.14 | $0.40 | $0.14 | Yes | 2026-10-07 |
| Crusoe | GLM 5.3 | $1.40 | $4.40 | $0.26 | Yes | 2026-10-07 |
| Crusoe | GPT-OSS 120B | $0.05 | $0.20 | $0.05 | Yes | 2026-10-07 |
Notes on each price (tiers, promotions, quantization the host states)
- OpenAI: gpt-5.4-mini
- From the 'All models' Standard table (rows embedded in page data, collapsed by default); row has 3 values mapped as input / cached input / output (no cache-write price). Most recent model actually named '-mini' on the page. Batch/Flex $0.375 / $0.0375 / $2.25; Fast $1.50 / $0.15 / $9.00. No long-context tier shown. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
- OpenAI: gpt-5.4-nano
- From the 'All models' Standard table (collapsed by default); 3 values mapped as input / cached input / output. Most recent model actually named '-nano' on the page. Batch/Flex $0.10 / $0.01 / $0.625. No long-context tier shown. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
- OpenAI: gpt-6-astra
- Listed first under 'Flagship models'. Short context (<=272K input tokens) shown; long context (>272K input): $20 in / $2 cached / $75 out. Cache writes $12.50 (short) / $25 (long); per page, input tokens are billed as input, cached input, or cache write (not additive). Batch/Flex $5 / $0.50 / $25. Fast $20 / $2 / $100; Ultrafast $60 / $6 / $300. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
- OpenAI: gpt-6-luna
- Flagship-table small tier (current generation uses astra/sol/luna naming rather than mini/nano). Short context (<=272K) shown; long context (>272K): $0.20 in / $0.02 cached / $0.75 out. Cache writes $0.125 (short) / $0.25 (long). Batch/Flex $0.05 / $0.005 / $0.25. Fast $0.20 / $0.02 / $1.00. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
- OpenAI: gpt-6.1-sol
- Flagship-table mid tier. Short context (<=272K) shown; long context (>272K): $4 in / $0.20 cached / $15 out. Cache writes $2.50 (short) / $5 (long). Batch/Flex $1 / $0.05 / $5. Fast $4 / $0.20 / $20. Also listed: gpt-6-sol at $2 / $0.20 cached / $10. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
- Anthropic: Claude Fable 5.1
- Top row of model table ('For demanding reasoning and long-horizon agentic work'). Cache hit is 0.025x input ($0.25). 5m write $12.50, 1h write $20. Batch $5 / $25. Full 1M context at standard pricing (Claude 4.6+). Page says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text, so compare per-task rather than per-token cost. Batch API 50% off input and output. Prompt caching: 5-min write 1.25x input, 1-hour write 2x input. inference_geo 'us' (data residency) 1.1x on all token categories for Claude 4.6+ models. Requested docs.claude.com URL redirected to platform.claude.com.
- Anthropic: Claude Haiku 4.5
- Current Haiku. Cache hit $0.10 (0.1x). 5m write $1.25, 1h write $2. Batch $0.50 / $2.50. Uses the previous tokenizer (Claude 4.6 and earlier). Batch API 50% off input and output. Prompt caching: 5-min write 1.25x input, 1-hour write 2x input. inference_geo 'us' (data residency) 1.1x on all token categories for Claude 4.6+ models. Requested docs.claude.com URL redirected to platform.claude.com.
- Anthropic: Claude Opus 5.5
- Current Opus. Cache hit is 0.05x input ($0.20). 5m write $5, 1h write $8. Batch $2 / $10. Fast mode (research preview) $8 in / $40 out. Full 1M context at standard pricing (Claude 4.6+). Page says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text, so compare per-task rather than per-token cost. Batch API 50% off input and output. Prompt caching: 5-min write 1.25x input, 1-hour write 2x input. inference_geo 'us' (data residency) 1.1x on all token categories for Claude 4.6+ models. Requested docs.claude.com URL redirected to platform.claude.com.
- Anthropic: Claude Sonnet 5.5
- Current Sonnet. Cache hit $0.20 (0.1x). 5m write $2.50, 1h write $4. Batch $1 / $5. Full 1M context at standard pricing (Claude 4.6+). Page says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text, so compare per-task rather than per-token cost. Batch API 50% off input and output. Prompt caching: 5-min write 1.25x input, 1-hour write 2x input. inference_geo 'us' (data residency) 1.1x on all token categories for Claude 4.6+ models. Requested docs.claude.com URL redirected to platform.claude.com.
- Google: Gemini 2.5 Pro (gemini-2.5-pro)
- Older generally available Pro model, kept for reference. Prompts <=200k shown; >200k: $2.50 in / $15.00 out / $0.25 cached. Batch $0.625 / $5.00 (<=200k). Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
- Google: Gemini 3.1 Pro Preview (gemini-3.1-pro-preview)
- Only Gemini 3.x Pro text model with its own price table (still 'Preview'). Prompts <=200k tokens shown; prompts >200k: $4.00 in / $18.00 out / $0.40 cached. Context-cache storage $4.50 per 1M tokens per hour. Batch/Flex $1.00 / $6.00 (<=200k), $2.00 / $9.00 (>200k). Priority $3.60 / $21.60 (<=200k). Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
- Google: Gemini 3.5 Flash (gemini-3.5-flash)
- Page calls it 'Our earlier Flash model'; it is not on the promotional pricing the 3.6 to 3.8 Flash models have. Cache storage $1.00 per 1M tokens per hour. Batch/Flex $0.75 / $4.50. Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
- Google: Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite)
- Input price applies to text/image/video/audio. Cache storage $1.00 per 1M tokens per hour. Batch $0.15 / $1.25. Older Gemini 3.1 Flash-Lite is cheaper: $0.25 in (text/image/video; $0.50 audio) / $1.50 out / $0.025 cached. Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
- Google: Gemini 3.8 Flash (gemini-3.8-flash)
- Page calls it 'Our most intelligent Flash model'. PROMOTIONAL: $0.75 in / $3.75 out / $0.075 cached through 2026-12-31; from 2027-01-01 $1.50 / $7.50 / $0.15. Cache storage $0.50 per 1M tokens per hour (rises to $1.00). No >200k tier shown. Batch/Flex $0.375 / $1.875. Priority $1.35 / $6.75. Gemini 3.7 Flash and 3.6 Flash carry the same promotional prices. Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
- DeepSeek: deepseek-flash (DeepSeek-V4.1-Flash)
- Cache-miss input $0.30, cache-hit input $0.006, output $1.20 (peak). Off-peak: $0.15 / $0.003 / $0.60. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are served by V4.1-Flash at this price. Weights on HF (MIT) per HF API check. PEAK rates recorded (the higher, undiscounted rate). Off-peak is half price: peak hours are 01:00-04:00 and 06:00-10:00 UTC, Mon-Fri, excluding Chinese public holidays; all other hours (most of the week) are off-peak. Context 1M, max output 384K. Thinking mode is on by default.
- DeepSeek: deepseek-v4-pro (DeepSeek-V4-Pro-0813)
- Cache-miss input $1.32, cache-hit input $0.044, output $3.96 (peak). Off-peak: $0.66 / $0.022 / $1.98. No vision. Weights on HF (MIT) per HF API check. PEAK rates recorded (the higher, undiscounted rate). Off-peak is half price: peak hours are 01:00-04:00 and 06:00-10:00 UTC, Mon-Fri, excluding Chinese public holidays; all other hours (most of the week) are off-peak. Context 1M, max output 384K. Thinking mode is on by default.
- Mistral: Mistral Large 3
- 675B MoE with open weights (Apache-2.0 on HF). The mistral.ai/pricing FAQ example also quotes 'Mistral Large' at $0.5 in / $1.5 out. Standard tier, default (non-regional) view of docs.mistral.ai/inference/pricing. mistral.ai/pricing (the requested URL) shows only plans, no per-model API table. Batch is 50% off per Mistral's pages.
- Mistral: Mistral Large 4
- SALE PRICE recorded (what is charged today). The page shows the original price struck through: $1.36 in / $0.14 cached / $4.18 out. No sale end date is given. No mistralai Large 4 repo found on Hugging Face, so treated as closed weights. Standard tier, default (non-regional) view of docs.mistral.ai/inference/pricing. mistral.ai/pricing (the requested URL) shows only plans, no per-model API table. Batch is 50% off per Mistral's pages.
- Mistral: Mistral Medium 3.5
- The pricing page recommends it for most tasks and coding. HF has mistralai/Mistral-Medium-3.5-128B under a non-Apache license ('other', probably Mistral Research License), so check commercial self-hosting terms. Standard tier, default (non-regional) view of docs.mistral.ai/inference/pricing. mistral.ai/pricing (the requested URL) shows only plans, no per-model API table. Batch is 50% off per Mistral's pages.
- Mistral: Mistral Small 4
- Open weights (Apache-2.0 on HF, 119B). Same page lists Ministral 3 14B $0.2/$0.2, 8B $0.15/$0.15, 3B $0.1/$0.1, and third-party Z.ai GLM 5.3 at $1.4 / $0.14 cached / $4.4. Standard tier, default (non-regional) view of docs.mistral.ai/inference/pricing. mistral.ai/pricing (the requested URL) shows only plans, no per-model API table. Batch is 50% off per Mistral's pages.
- Together AI: DeepSeek V4 Flash 0731
- Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: DeepSeek V4 Pro 0813
- Input and output match DeepSeek's own peak prices; cached is $0.13 here vs $0.044 at DeepSeek. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: DeepSeek V4.1 Flash
- Matches DeepSeek's own peak prices. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Gemma 4 31B
- Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: GLM-5.3
- GLM-5.2 is also listed at the same $1.40 / $0.26 / $4.40. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: GLM-5.3-Flash
- Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: gpt-oss-120B
- No cached price shown. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Kimi K3
- Marked 'PROMO' in the Vision tab of the same table, so this price may be temporary. Kimi K2.x is not listed on Together's page. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Llama 3 8B Instruct Lite
- This is Llama 3 (not 3.1) 8B 'Lite'. Together lists no Llama 3.1 8B chat model. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Llama 3.3 70B
- Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: MiniMax M2.7
- Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: MiniMax M3
- Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Qwen3 235B A22B Instruct 2507 FP8 Throughput
- FP8, 'Throughput' variant (as named). Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Qwen3.5-397B-A17B
- Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Qwen3.5 9B
- Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Qwen3.6-Plus
- No Qwen/Qwen3.6-Plus repo on HF (only Qwen3.6-27B and 35B-A3B), so treated as Alibaba's closed 'Plus' API model resold by Together. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Qwen3.8-2.4T-A95B
- Page URL slug is 'qwen3-8-max', so this is probably what Fireworks and DeepInfra call 'Qwen3.8 Max'. HF license is 'other'. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Together AI: Qwen3.8 Flash
- Open-weights status unclear: HF has Qwen/Qwen3.8-Flash-Next but no 'Qwen3.8-Flash' repo. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
- Fireworks AI: Other base models: 4B to 16B parameters
- Would cover Llama 3.1 8B, Qwen3 8B, Qwen3 14B and Gemma 3 12B if served. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
- Fireworks AI: Other base models: more than 16B parameters
- Would cover dense Llama 3.3 70B, Llama 3.1 405B, Qwen3 32B, Gemma 3 27B and Mistral Small 24B if served. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
- Fireworks AI: Other base models: less than 4B parameters
- Applies to dense models under 4B parameters. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
- Fireworks AI: Other base models: MoE 56.1B to 176B parameters (e.g. DBRX, Mixtral 8x22B)
- No bucket is listed for MoE above 176B (e.g. Qwen3 235B-A22B), so those models have no rule-based price. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
- Fireworks AI: Other base models: MoE up to 56B parameters (e.g. Mixtral 8x7B)
- Would cover Qwen3 30B-A3B and gpt-oss-20b (about 21B MoE) if served. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
- Fireworks AI: DeepSeek V4.1 Flash
- US region $0.45 / $0.009 / $1.80. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
- Fireworks AI: GLM 5.3
- GLM 5.3 Fast and US region both $2.10 / $0.39 / $6.60. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
- Fireworks AI: GLM 5.3 Flash
- Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
- Fireworks AI: OpenAI GPT OSS 120B
- Priority $0.18 / $0.018 / $0.72. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
- Fireworks AI: Kimi K3
- Kimi K3 Fast $4.50 / $0.45 / $22.50; US region $4.50 / $0.45 / $22.50. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
- Fireworks AI: MiniMax M3
- Priority $0.45 / $0.09 / $1.80. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
- Fireworks AI: Qwen 3.8 Max
- Same price as Together's 'Qwen3.8-2.4T-A95B' (slug qwen3-8-max), so probably the open Qwen/Qwen3.8-2.4T-A95B, but this page does not say so. Priority $3.00 / $0.375 / $9.00. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
- DeepInfra: DeepSeek-V3.2
- Context 160k. DeepSeek-V3.1 is also listed: $0.25 / $0.13 cached / $0.95. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: DeepSeek-V4-Flash
- Context 1024k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: DeepSeek-V4-Flash-0731
- Context 1024k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: DeepSeek-V4-Pro
- Context 1024k. This is the original V4-Pro checkpoint; DeepSeek's own API now serves V4-Pro-0813. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: gemma-3-27b-it
- Context 128k. gemma-3-12b-it is also listed at $0.05 / $0.15. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: gemma-4-26B-A4B-it
- Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: gemma-4-31B-it
- Context 256k. Variants also listed: gemma-4-31B-it-turbo $0.09 / $0.05 cached / $0.34, and gemma-4-31B-it-Ultra $0.27 / $0.76 (128k). Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Kimi-K2.6
- Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Kimi-K3
- Context 1024k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Meta-Llama-3.1-8B-Instruct-Turbo
- Turbo variant; context 128k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Llama-3.3-70B-Instruct-Turbo
- Turbo variant; context 128k. Meta-Llama-3.1-70B-Instruct-Turbo is also listed at $0.40 / $0.40. No 405B listed. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Mistral-Small-3.2-24B-Instruct-2506
- Context 125k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3-14B
- Context 40k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3-235B-A22B-Instruct-2507
- Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3-30B-A3B
- Context 40k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3-32B
- Context 40k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3.5-27B
- Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3.5-35B-A3B
- Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3.5-397B-A17B
- Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3.5-9B
- Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3.6-27B
- Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3.6-35B-A3B
- Context 256k. The page shows cached input equal to input ($0.10 / $0.10 cached). Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- DeepInfra: Qwen3.8-Max
- Context 250k. Probably the same model as Together's Qwen3.8-2.4T-A95B, but there is no Qwen/Qwen3.8-Max repo on HF and DeepInfra also resells closed models (e.g. Qwen3-Max, Gemini, Claude), so open-weights status is unconfirmed. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
- Groq: GPT OSS 120B (openai/gpt-oss-120b)
- Production model, about 500 tokens/s. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
- Groq: GPT OSS 20B (openai/gpt-oss-20b)
- Production model, about 1000 tokens/s. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
- Groq: Llama 3.1 8B (llama-3.1-8b-instant)
- Price shown as 'Contact Sales' (tagged Enterprise); no public per-token price. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
- Groq: Llama 3.3 70B (llama-3.3-70b-versatile)
- Price shown as 'Contact Sales' (tagged Enterprise); no public per-token price. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
- Groq: MiniMax M2.7 (minimaxai/minimax-m2.7)
- Preview, tagged Enterprise; price shown as 'Contact Sales'. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
- Groq: Qwen/Qwen3.8-27B (qwen/qwen3.8-27b)
- PREVIEW model (evaluation only, may be discontinued at short notice). groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
- Hyperstack: OpenAI gpt-oss-120b
- From the 'Token-based Pricing' table on Hyperstack's GPU pricing page.
- Hyperstack: Llama 3.1 8B
- From the 'Token-based Pricing' table on Hyperstack's GPU pricing page.
- Hyperstack: Llama 3.3 70B
- From the 'Token-based Pricing' table on Hyperstack's GPU pricing page.
- Crusoe: Gemma 4 31B-it
- From the 'Serverless Inference pricing' table on Crusoe's cloud pricing page.
- Crusoe: GLM 5.3
- From the 'Serverless Inference pricing' table on Crusoe's cloud pricing page.
- Crusoe: GPT-OSS 120B
- From the 'Serverless Inference pricing' table on Crusoe's cloud pricing page.
Pages we could not use as published, and what we used instead:
https://openai.com/api/pricing/: HTTP 403 (bot protection). Used https://developers.openai.com/api/docs/pricing instead (platform.openai.com/docs/pricing redirects there).https://groq.com/pricing: HTTP 308 redirect to the groq.com homepage, which has no prices. Used https://console.groq.com/docs/models (GroqDocs Supported Models, which lists per-model prices).https://mistral.ai/pricing: Lists plans only, with no per-model API price table in the raw HTML or via WebFetch; the only API figure is an FAQ example ('Mistral Large costs $0.5/M in, $1.5/M out'). Used https://docs.mistral.ai/inference/pricing instead.https://fireworks.ai/pricing: No per-token serverless price table (only embeddings, training and GPU prices); it links to https://docs.fireworks.ai/serverless/pricing, which was used.
How the costs are calculated
API cost per month = input tokens x input price + output tokens x output price
(prompt-cached input at its own price, if the API has one)
GPU busy hours per month = input tokens / prefill tokens per second
+ output tokens / decode tokens per second (per replica)
Dedicated = replicas x GPUs per replica x 730 hours x price per GPU-hour
replicas = busy hours / (730 x utilization), rounded up, at least 1
Hourly = busy hours / utilization x GPUs per replica x price per GPU-hour
Break-even volume = the monthly volume, at your input/output mix, where the two costs are equal
A replica is one running copy of the model. If the model needs two GPUs, as Llama 3.3 70B at BF16 does on 80 GB cards, then a replica is two GPUs.
Throughput is the assumption that decides this
Every price on the page is a list price you can check. Throughput is the one number nobody can look up for you. It depends on the model, the precision, the GPU, the engine and its version, how many requests run at once, and how long your prompts and answers are. Change it by a factor of two and the answer can flip. That is why the result always shows the cost at a quarter, half, double and four times your figure.
Two throughputs, not one:
- Prefill processes the prompt, many tokens per forward pass. It is mostly compute-bound and fast per token.
- Decode generates the answer, one token per sequence per forward pass. It is mostly bound by memory bandwidth, and it is where batching matters. A single stream is slow, and many concurrent streams share each read of the weights.
With continuous batching, both run on the same GPUs and compete for the same time, so the calculator adds their busy hours. Use the total decode rate across all concurrent streams at the concurrency you will actually run, not the speed of one stream.
To measure it, run your engine's serving benchmark (vLLM ships one as vllm bench serve) with your real distribution of prompt and output lengths, at the concurrency you plan to serve. Read input and output tokens per second. Measured throughput figures on this site are planned; none are published yet.
Utilization: you pay for idle GPUs
An API bills for tokens. A dedicated GPU bills for hours, busy or not. Traffic is uneven: if your peak hour carries five times your average load and you size for the peak, your replicas average about 20% busy. The utilization input is the share of paid GPU time spent working at the throughput you entered. The busy-time figure in the result shows what your inputs imply.
Hourly mode assumes you can stop paying the moment there is no traffic, for example with serverless GPUs billed per second. It ignores cold starts, minimum billing periods and capacity that is not there when you ask for it, so treat it as the best case.
A worked example
These are the calculator's defaults. The throughput figures are placeholders, so the example shows the method, not the market.
- Volume: 300 million input and 60 million output tokens a month.
- API: Llama 3.3 70B from Together AI at $1.04 per million tokens in and out, which comes to $374 a month.
- Self-hosted: the same model at BF16 on two H100 SXM GPUs from RunPod Secure Cloud at $3.49 per GPU-hour. One replica running all month costs $5,095.
- Busy time: at 20,000 prefill and 2,000 decode tokens per second, the replica is busy 12.5 hours a month, 1.7% of the time it is paid for.
Self-hosting breaks even at about 4.9 billion tokens a month, 13.6 times this volume. Below that, no throughput can help, because one idle replica already costs more than the whole API bill.
At twenty times the volume (7.2 billion tokens a month), the same replica is busy a third of the time and costs $5,095 against an API bill of $7,488. Self-hosting wins, but only just: if real throughput is half the placeholder, it takes two replicas ($10,191) and the API wins again. That is the throughput point in practice.
When self-hosting is worth it anyway
Cost is not the only reason. Self-host when:
- Data must stay in your own account or region.
- You run a fine-tuned model, or one that no API serves.
- You need control of latency, batching or the engine.
- Your volume is steady and high enough that the sensitivity table favors you at half your measured throughput.
Use an API when traffic is spiky or small, when an API serves the same open model at a price no single GPU can match, or when you have nobody to keep a serving stack up at night.
What is not included
- Engineering and on-call time. Put it in "other self-hosting cost". It is often the largest line.
- Storage, egress, and CPU or RAM that some GPU providers bill separately.
- Quality differences between the API model and the one you host. Compare like with like: the price list marks open-weights models, and some hosts serve quantized versions.
- API rate limits, batch-API discounts (several providers list batch at half price; see the notes under the price table), and committed-use discounts on either side.
Renting the GPUs
For testing a model before you commit, by-the-hour providers are the cheapest way to measure real throughput: RunPod and the Vast.ai marketplace for single cards, and Lambda and DigitalOcean for H100, H200 and B200 instances. Compare current prices in the GPU price table, and size the replica with the VRAM calculator.
Also available as Markdown.