Self-host vs API break-even calculator

Runs in your browser. API list prices read 7 October 2026 and GPU on-demand prices read 7 October 2026, each linked to its source. Every price is editable.

Your monthly volume
The API

Self-hosting

The throughput values above are placeholders, not measurements. Replace them with numbers you measured for your model, precision, GPU, engine and traffic shape. They decide the answer.

How you pay for GPUs

API list prices behind the calculator

Standard real-time tier, US dollars per million tokens, read from each provider's own pricing page on 7 October 2026. "Check provider" means the page states no price (for example "Contact sales"). Batch, priority and off-peak prices are in the notes. GPU prices are on the GPU price table. Raw data: api-prices.json.

ProviderModelInputOutputCached inputOpen weightsSource
OpenAI gpt-5.4-mini $0.75 $4.50 $0.075 No 2026-10-07
OpenAI gpt-5.4-nano $0.20 $1.25 $0.02 No 2026-10-07
OpenAI gpt-6-astra $10.00 $50.00 $1.00 No 2026-10-07
OpenAI gpt-6-luna $0.10 $0.50 $0.01 No 2026-10-07
OpenAI gpt-6.1-sol $2.00 $10.00 $0.10 No 2026-10-07
Anthropic Claude Fable 5.1 $10.00 $50.00 $0.25 No 2026-10-07
Anthropic Claude Haiku 4.5 $1.00 $5.00 $0.10 No 2026-10-07
Anthropic Claude Opus 5.5 $4.00 $20.00 $0.20 No 2026-10-07
Anthropic Claude Sonnet 5.5 $2.00 $10.00 $0.20 No 2026-10-07
Google Gemini 2.5 Pro (gemini-2.5-pro) $1.25 $10.00 $0.125 No 2026-10-07
Google Gemini 3.1 Pro Preview (gemini-3.1-pro-preview) $2.00 $12.00 $0.20 No 2026-10-07
Google Gemini 3.5 Flash (gemini-3.5-flash) $1.50 $9.00 $0.15 No 2026-10-07
Google Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) $0.30 $2.50 $0.03 No 2026-10-07
Google Gemini 3.8 Flash (gemini-3.8-flash) $0.75 $3.75 $0.075 No 2026-10-07
DeepSeek deepseek-flash (DeepSeek-V4.1-Flash) $0.30 $1.20 $0.006 Yes 2026-10-07
DeepSeek deepseek-v4-pro (DeepSeek-V4-Pro-0813) $1.32 $3.96 $0.044 Yes 2026-10-07
Mistral Mistral Large 3 $0.50 $1.50 $0.05 Yes 2026-10-07
Mistral Mistral Large 4 $0.68 $2.09 $0.07 No 2026-10-07
Mistral Mistral Medium 3.5 $1.50 $7.50 $0.15 Yes 2026-10-07
Mistral Mistral Small 4 $0.15 $0.60 $0.015 Yes 2026-10-07
Together AI DeepSeek V4 Flash 0731 $0.14 $0.28 $0.03 Yes 2026-10-07
Together AI DeepSeek V4 Pro 0813 $1.32 $3.96 $0.13 Yes 2026-10-07
Together AI DeepSeek V4.1 Flash $0.30 $1.20 $0.006 Yes 2026-10-07
Together AI Gemma 4 31B $0.39 $0.97 Not listed Yes 2026-10-07
Together AI GLM-5.3 $1.40 $4.40 $0.26 Yes 2026-10-07
Together AI GLM-5.3-Flash $0.15 $0.50 $0.03 Yes 2026-10-07
Together AI gpt-oss-120B $0.15 $0.60 Not listed Yes 2026-10-07
Together AI Kimi K3 $2.70 $13.50 $0.27 Yes 2026-10-07
Together AI Llama 3 8B Instruct Lite $0.14 $0.14 Not listed Yes 2026-10-07
Together AI Llama 3.3 70B $1.04 $1.04 Not listed Yes 2026-10-07
Together AI MiniMax M2.7 $0.30 $1.20 $0.06 Yes 2026-10-07
Together AI MiniMax M3 $0.30 $1.20 $0.06 Yes 2026-10-07
Together AI Qwen3 235B A22B Instruct 2507 FP8 Throughput $0.20 $0.60 Not listed Yes 2026-10-07
Together AI Qwen3.5-397B-A17B $0.60 $3.60 $0.35 Yes 2026-10-07
Together AI Qwen3.5 9B $0.17 $0.25 Not listed Yes 2026-10-07
Together AI Qwen3.6-Plus $0.50 $3.00 Not listed No 2026-10-07
Together AI Qwen3.8-2.4T-A95B $2.00 $6.00 $0.25 Yes 2026-10-07
Together AI Qwen3.8 Flash $0.15 $0.47 Not listed Unclear 2026-10-07
Fireworks AI Other base models: 4B to 16B parameters $0.20 $0.20 Not listed Yes 2026-10-07
Fireworks AI Other base models: more than 16B parameters $0.90 $0.90 Not listed Yes 2026-10-07
Fireworks AI Other base models: less than 4B parameters $0.10 $0.10 Not listed Yes 2026-10-07
Fireworks AI Other base models: MoE 56.1B to 176B parameters (e.g. DBRX, Mixtral 8x22B) $1.20 $1.20 Not listed Yes 2026-10-07
Fireworks AI Other base models: MoE up to 56B parameters (e.g. Mixtral 8x7B) $0.50 $0.50 Not listed Yes 2026-10-07
Fireworks AI DeepSeek V4.1 Flash $0.30 $1.20 $0.006 Yes 2026-10-07
Fireworks AI GLM 5.3 $1.40 $4.40 $0.26 Yes 2026-10-07
Fireworks AI GLM 5.3 Flash $0.15 $0.50 $0.03 Yes 2026-10-07
Fireworks AI OpenAI GPT OSS 120B $0.15 $0.60 $0.015 Yes 2026-10-07
Fireworks AI Kimi K3 $3.00 $15.00 $0.30 Yes 2026-10-07
Fireworks AI MiniMax M3 $0.30 $1.20 $0.06 Yes 2026-10-07
Fireworks AI Qwen 3.8 Max $2.00 $6.00 $0.25 Unclear 2026-10-07
DeepInfra DeepSeek-V3.2 $0.26 $0.38 $0.13 Yes 2026-10-07
DeepInfra DeepSeek-V4-Flash $0.09 $0.18 $0.018 Yes 2026-10-07
DeepInfra DeepSeek-V4-Flash-0731 $0.06 $0.18 $0.015 Yes 2026-10-07
DeepInfra DeepSeek-V4-Pro $1.30 $2.60 $0.10 Yes 2026-10-07
DeepInfra gemma-3-27b-it $0.08 $0.16 Not listed Yes 2026-10-07
DeepInfra gemma-4-26B-A4B-it $0.07 $0.34 Not listed Yes 2026-10-07
DeepInfra gemma-4-31B-it $0.20 $0.40 Not listed Yes 2026-10-07
DeepInfra Kimi-K2.6 $0.75 $3.50 $0.15 Yes 2026-10-07
DeepInfra Kimi-K3 $2.85 $14.25 $0.285 Yes 2026-10-07
DeepInfra Meta-Llama-3.1-8B-Instruct-Turbo $0.02 $0.04 Not listed Yes 2026-10-07
DeepInfra Llama-3.3-70B-Instruct-Turbo $0.10 $0.32 Not listed Yes 2026-10-07
DeepInfra Mistral-Small-3.2-24B-Instruct-2506 $0.075 $0.20 Not listed Yes 2026-10-07
DeepInfra Qwen3-14B $0.12 $0.24 Not listed Yes 2026-10-07
DeepInfra Qwen3-235B-A22B-Instruct-2507 $0.09 $0.55 Not listed Yes 2026-10-07
DeepInfra Qwen3-30B-A3B $0.12 $0.50 Not listed Yes 2026-10-07
DeepInfra Qwen3-32B $0.08 $0.28 Not listed Yes 2026-10-07
DeepInfra Qwen3.5-27B $0.26 $2.60 Not listed Yes 2026-10-07
DeepInfra Qwen3.5-35B-A3B $0.14 $1.00 $0.05 Yes 2026-10-07
DeepInfra Qwen3.5-397B-A17B $0.45 $3.00 $0.22 Yes 2026-10-07
DeepInfra Qwen3.5-9B $0.10 $0.15 Not listed Yes 2026-10-07
DeepInfra Qwen3.6-27B $0.32 $3.20 Not listed Yes 2026-10-07
DeepInfra Qwen3.6-35B-A3B $0.10 $0.95 $0.10 Yes 2026-10-07
DeepInfra Qwen3.8-Max $1.65 $4.951 $0.206 Unclear 2026-10-07
Groq GPT OSS 120B (openai/gpt-oss-120b) $0.15 $0.60 Not listed Yes 2026-10-07
Groq GPT OSS 20B (openai/gpt-oss-20b) $0.075 $0.30 Not listed Yes 2026-10-07
Groq Llama 3.1 8B (llama-3.1-8b-instant) Check provider Check provider Not listed Yes 2026-10-07
Groq Llama 3.3 70B (llama-3.3-70b-versatile) Check provider Check provider Not listed Yes 2026-10-07
Groq MiniMax M2.7 (minimaxai/minimax-m2.7) Check provider Check provider Not listed Yes 2026-10-07
Groq Qwen/Qwen3.8-27B (qwen/qwen3.8-27b) $0.80 $4.00 Not listed Yes 2026-10-07
Hyperstack OpenAI gpt-oss-120b $0.10 $0.40 Not listed Yes 2026-10-07
Hyperstack Llama 3.1 8B $0.20 $0.20 Not listed Yes 2026-10-07
Hyperstack Llama 3.3 70B $0.80 $0.80 Not listed Yes 2026-10-07
Crusoe Gemma 4 31B-it $0.14 $0.40 $0.14 Yes 2026-10-07
Crusoe GLM 5.3 $1.40 $4.40 $0.26 Yes 2026-10-07
Crusoe GPT-OSS 120B $0.05 $0.20 $0.05 Yes 2026-10-07
Notes on each price (tiers, promotions, quantization the host states)
OpenAI: gpt-5.4-mini
From the 'All models' Standard table (rows embedded in page data, collapsed by default); row has 3 values mapped as input / cached input / output (no cache-write price). Most recent model actually named '-mini' on the page. Batch/Flex $0.375 / $0.0375 / $2.25; Fast $1.50 / $0.15 / $9.00. No long-context tier shown. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
OpenAI: gpt-5.4-nano
From the 'All models' Standard table (collapsed by default); 3 values mapped as input / cached input / output. Most recent model actually named '-nano' on the page. Batch/Flex $0.10 / $0.01 / $0.625. No long-context tier shown. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
OpenAI: gpt-6-astra
Listed first under 'Flagship models'. Short context (<=272K input tokens) shown; long context (>272K input): $20 in / $2 cached / $75 out. Cache writes $12.50 (short) / $25 (long); per page, input tokens are billed as input, cached input, or cache write (not additive). Batch/Flex $5 / $0.50 / $25. Fast $20 / $2 / $100; Ultrafast $60 / $6 / $300. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
OpenAI: gpt-6-luna
Flagship-table small tier (current generation uses astra/sol/luna naming rather than mini/nano). Short context (<=272K) shown; long context (>272K): $0.20 in / $0.02 cached / $0.75 out. Cache writes $0.125 (short) / $0.25 (long). Batch/Flex $0.05 / $0.005 / $0.25. Fast $0.20 / $0.02 / $1.00. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
OpenAI: gpt-6.1-sol
Flagship-table mid tier. Short context (<=272K) shown; long context (>272K): $4 in / $0.20 cached / $15 out. Cache writes $2.50 (short) / $5 (long). Batch/Flex $1 / $0.05 / $5. Fast $4 / $0.20 / $20. Also listed: gpt-6-sol at $2 / $0.20 cached / $10. Standard tier. Batch and Flex are 50% of Standard. Output price includes reasoning tokens. Regional processing (data residency) and FedRAMP endpoints +10% for models released on/after 2026-03-05. Requested URL platform.openai.com/docs/pricing redirected to developers.openai.com/api/docs/pricing; openai.com/api/pricing/ returned 403.
Anthropic: Claude Fable 5.1
Top row of model table ('For demanding reasoning and long-horizon agentic work'). Cache hit is 0.025x input ($0.25). 5m write $12.50, 1h write $20. Batch $5 / $25. Full 1M context at standard pricing (Claude 4.6+). Page says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text, so compare per-task rather than per-token cost. Batch API 50% off input and output. Prompt caching: 5-min write 1.25x input, 1-hour write 2x input. inference_geo 'us' (data residency) 1.1x on all token categories for Claude 4.6+ models. Requested docs.claude.com URL redirected to platform.claude.com.
Anthropic: Claude Haiku 4.5
Current Haiku. Cache hit $0.10 (0.1x). 5m write $1.25, 1h write $2. Batch $0.50 / $2.50. Uses the previous tokenizer (Claude 4.6 and earlier). Batch API 50% off input and output. Prompt caching: 5-min write 1.25x input, 1-hour write 2x input. inference_geo 'us' (data residency) 1.1x on all token categories for Claude 4.6+ models. Requested docs.claude.com URL redirected to platform.claude.com.
Anthropic: Claude Opus 5.5
Current Opus. Cache hit is 0.05x input ($0.20). 5m write $5, 1h write $8. Batch $2 / $10. Fast mode (research preview) $8 in / $40 out. Full 1M context at standard pricing (Claude 4.6+). Page says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text, so compare per-task rather than per-token cost. Batch API 50% off input and output. Prompt caching: 5-min write 1.25x input, 1-hour write 2x input. inference_geo 'us' (data residency) 1.1x on all token categories for Claude 4.6+ models. Requested docs.claude.com URL redirected to platform.claude.com.
Anthropic: Claude Sonnet 5.5
Current Sonnet. Cache hit $0.20 (0.1x). 5m write $2.50, 1h write $4. Batch $1 / $5. Full 1M context at standard pricing (Claude 4.6+). Page says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text, so compare per-task rather than per-token cost. Batch API 50% off input and output. Prompt caching: 5-min write 1.25x input, 1-hour write 2x input. inference_geo 'us' (data residency) 1.1x on all token categories for Claude 4.6+ models. Requested docs.claude.com URL redirected to platform.claude.com.
Google: Gemini 2.5 Pro (gemini-2.5-pro)
Older generally available Pro model, kept for reference. Prompts <=200k shown; >200k: $2.50 in / $15.00 out / $0.25 cached. Batch $0.625 / $5.00 (<=200k). Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
Google: Gemini 3.1 Pro Preview (gemini-3.1-pro-preview)
Only Gemini 3.x Pro text model with its own price table (still 'Preview'). Prompts <=200k tokens shown; prompts >200k: $4.00 in / $18.00 out / $0.40 cached. Context-cache storage $4.50 per 1M tokens per hour. Batch/Flex $1.00 / $6.00 (<=200k), $2.00 / $9.00 (>200k). Priority $3.60 / $21.60 (<=200k). Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
Google: Gemini 3.5 Flash (gemini-3.5-flash)
Page calls it 'Our earlier Flash model'; it is not on the promotional pricing the 3.6 to 3.8 Flash models have. Cache storage $1.00 per 1M tokens per hour. Batch/Flex $0.75 / $4.50. Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
Google: Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite)
Input price applies to text/image/video/audio. Cache storage $1.00 per 1M tokens per hour. Batch $0.15 / $1.25. Older Gemini 3.1 Flash-Lite is cheaper: $0.25 in (text/image/video; $0.50 audio) / $1.50 out / $0.025 cached. Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
Google: Gemini 3.8 Flash (gemini-3.8-flash)
Page calls it 'Our most intelligent Flash model'. PROMOTIONAL: $0.75 in / $3.75 out / $0.075 cached through 2026-12-31; from 2027-01-01 $1.50 / $7.50 / $0.15. Cache storage $0.50 per 1M tokens per hour (rises to $1.00). No >200k tier shown. Batch/Flex $0.375 / $1.875. Priority $1.35 / $6.75. Gemini 3.7 Flash and 3.6 Flash carry the same promotional prices. Paid tier, Standard. Output price includes thinking tokens. Batch and Flex are 50% of Standard. First plain curl was redirected to a Google sign-in check; the page loaded fine with cookies kept (public, no login).
DeepSeek: deepseek-flash (DeepSeek-V4.1-Flash)
Cache-miss input $0.30, cache-hit input $0.006, output $1.20 (peak). Off-peak: $0.15 / $0.003 / $0.60. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are served by V4.1-Flash at this price. Weights on HF (MIT) per HF API check. PEAK rates recorded (the higher, undiscounted rate). Off-peak is half price: peak hours are 01:00-04:00 and 06:00-10:00 UTC, Mon-Fri, excluding Chinese public holidays; all other hours (most of the week) are off-peak. Context 1M, max output 384K. Thinking mode is on by default.
DeepSeek: deepseek-v4-pro (DeepSeek-V4-Pro-0813)
Cache-miss input $1.32, cache-hit input $0.044, output $3.96 (peak). Off-peak: $0.66 / $0.022 / $1.98. No vision. Weights on HF (MIT) per HF API check. PEAK rates recorded (the higher, undiscounted rate). Off-peak is half price: peak hours are 01:00-04:00 and 06:00-10:00 UTC, Mon-Fri, excluding Chinese public holidays; all other hours (most of the week) are off-peak. Context 1M, max output 384K. Thinking mode is on by default.
Mistral: Mistral Large 3
675B MoE with open weights (Apache-2.0 on HF). The mistral.ai/pricing FAQ example also quotes 'Mistral Large' at $0.5 in / $1.5 out. Standard tier, default (non-regional) view of docs.mistral.ai/inference/pricing. mistral.ai/pricing (the requested URL) shows only plans, no per-model API table. Batch is 50% off per Mistral's pages.
Mistral: Mistral Large 4
SALE PRICE recorded (what is charged today). The page shows the original price struck through: $1.36 in / $0.14 cached / $4.18 out. No sale end date is given. No mistralai Large 4 repo found on Hugging Face, so treated as closed weights. Standard tier, default (non-regional) view of docs.mistral.ai/inference/pricing. mistral.ai/pricing (the requested URL) shows only plans, no per-model API table. Batch is 50% off per Mistral's pages.
Mistral: Mistral Medium 3.5
The pricing page recommends it for most tasks and coding. HF has mistralai/Mistral-Medium-3.5-128B under a non-Apache license ('other', probably Mistral Research License), so check commercial self-hosting terms. Standard tier, default (non-regional) view of docs.mistral.ai/inference/pricing. mistral.ai/pricing (the requested URL) shows only plans, no per-model API table. Batch is 50% off per Mistral's pages.
Mistral: Mistral Small 4
Open weights (Apache-2.0 on HF, 119B). Same page lists Ministral 3 14B $0.2/$0.2, 8B $0.15/$0.15, 3B $0.1/$0.1, and third-party Z.ai GLM 5.3 at $1.4 / $0.14 cached / $4.4. Standard tier, default (non-regional) view of docs.mistral.ai/inference/pricing. mistral.ai/pricing (the requested URL) shows only plans, no per-model API table. Batch is 50% off per Mistral's pages.
Together AI: DeepSeek V4 Flash 0731
Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: DeepSeek V4 Pro 0813
Input and output match DeepSeek's own peak prices; cached is $0.13 here vs $0.044 at DeepSeek. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: DeepSeek V4.1 Flash
Matches DeepSeek's own peak prices. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Gemma 4 31B
Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: GLM-5.3
GLM-5.2 is also listed at the same $1.40 / $0.26 / $4.40. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: GLM-5.3-Flash
Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: gpt-oss-120B
No cached price shown. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Kimi K3
Marked 'PROMO' in the Vision tab of the same table, so this price may be temporary. Kimi K2.x is not listed on Together's page. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Llama 3 8B Instruct Lite
This is Llama 3 (not 3.1) 8B 'Lite'. Together lists no Llama 3.1 8B chat model. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Llama 3.3 70B
Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: MiniMax M2.7
Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: MiniMax M3
Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Qwen3 235B A22B Instruct 2507 FP8 Throughput
FP8, 'Throughput' variant (as named). Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Qwen3.5-397B-A17B
Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Qwen3.5 9B
Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Qwen3.6-Plus
No Qwen/Qwen3.6-Plus repo on HF (only Qwen3.6-27B and 35B-A3B), so treated as Alibaba's closed 'Plus' API model resold by Together. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Qwen3.8-2.4T-A95B
Page URL slug is 'qwen3-8-max', so this is probably what Fireworks and DeepInfra call 'Qwen3.8 Max'. HF license is 'other'. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Together AI: Qwen3.8 Flash
Open-weights status unclear: HF has Qwen/Qwen3.8-Flash-Next but no 'Qwen3.8-Flash' repo. Serverless Chat table. The page has a 'Batch API price' toggle, but batch prices are not in the static HTML (filled in by JavaScript). Quantization not stated unless named.
Fireworks AI: Other base models: 4B to 16B parameters
Would cover Llama 3.1 8B, Qwen3 8B, Qwen3 14B and Gemma 3 12B if served. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
Fireworks AI: Other base models: more than 16B parameters
Would cover dense Llama 3.3 70B, Llama 3.1 405B, Qwen3 32B, Gemma 3 27B and Mistral Small 24B if served. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
Fireworks AI: Other base models: less than 4B parameters
Applies to dense models under 4B parameters. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
Fireworks AI: Other base models: MoE 56.1B to 176B parameters (e.g. DBRX, Mixtral 8x22B)
No bucket is listed for MoE above 176B (e.g. Qwen3 235B-A22B), so those models have no rule-based price. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
Fireworks AI: Other base models: MoE up to 56B parameters (e.g. Mixtral 8x7B)
Would cover Qwen3 30B-A3B and gpt-oss-20b (about 21B MoE) if served. SIZE-BUCKET RULE: 'For any text or vision model not listed individually, pricing is set by parameter count and architecture.' The same price applies to input and output, with no separate cached-input price. This page does NOT confirm which specific models are on serverless; check Fireworks' model library before mapping a model to this bucket. Batch is 50%.
Fireworks AI: DeepSeek V4.1 Flash
US region $0.45 / $0.009 / $1.80. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
Fireworks AI: GLM 5.3
GLM 5.3 Fast and US region both $2.10 / $0.39 / $6.60. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
Fireworks AI: GLM 5.3 Flash
Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
Fireworks AI: OpenAI GPT OSS 120B
Priority $0.18 / $0.018 / $0.72. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
Fireworks AI: Kimi K3
Kimi K3 Fast $4.50 / $0.45 / $22.50; US region $4.50 / $0.45 / $22.50. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
Fireworks AI: MiniMax M3
Priority $0.45 / $0.09 / $1.80. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
Fireworks AI: Qwen 3.8 Max
Same price as Together's 'Qwen3.8-2.4T-A95B' (slug qwen3-8-max), so probably the open Qwen/Qwen3.8-2.4T-A95B, but this page does not say so. Priority $3.00 / $0.375 / $9.00. Standard serverless tier; cells are input / cached input / output. Batch inference is 50% of serverless on input and output. Priority costs about 1.25x (1.5x for some models); '(US)' region variants cost 1.5x. fireworks.ai/pricing has no per-token table and links to this docs page.
DeepInfra: DeepSeek-V3.2
Context 160k. DeepSeek-V3.1 is also listed: $0.25 / $0.13 cached / $0.95. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: DeepSeek-V4-Flash
Context 1024k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: DeepSeek-V4-Flash-0731
Context 1024k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: DeepSeek-V4-Pro
Context 1024k. This is the original V4-Pro checkpoint; DeepSeek's own API now serves V4-Pro-0813. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: gemma-3-27b-it
Context 128k. gemma-3-12b-it is also listed at $0.05 / $0.15. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: gemma-4-26B-A4B-it
Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: gemma-4-31B-it
Context 256k. Variants also listed: gemma-4-31B-it-turbo $0.09 / $0.05 cached / $0.34, and gemma-4-31B-it-Ultra $0.27 / $0.76 (128k). Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Kimi-K2.6
Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Kimi-K3
Context 1024k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Meta-Llama-3.1-8B-Instruct-Turbo
Turbo variant; context 128k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Llama-3.3-70B-Instruct-Turbo
Turbo variant; context 128k. Meta-Llama-3.1-70B-Instruct-Turbo is also listed at $0.40 / $0.40. No 405B listed. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Mistral-Small-3.2-24B-Instruct-2506
Context 125k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3-14B
Context 40k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3-235B-A22B-Instruct-2507
Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3-30B-A3B
Context 40k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3-32B
Context 40k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3.5-27B
Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3.5-35B-A3B
Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3.5-397B-A17B
Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3.5-9B
Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3.6-27B
Context 256k. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3.6-35B-A3B
Context 256k. The page shows cached input equal to input ($0.10 / $0.10 cached). Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
DeepInfra: Qwen3.8-Max
Context 250k. Probably the same model as Together's Qwen3.8-2.4T-A95B, but there is no Qwen/Qwen3.8-Max repo on HF and DeepInfra also resells closed models (e.g. Qwen3-Max, Gemini, Claude), so open-weights status is unconfirmed. Standard service tier (1x). Priority 1.5x, Flex 0.8x. Hint is taken from DeepInfra's model URL path. Quantization is not stated on the pricing page; DeepInfra 'Turbo' variants are often quantized, so check the model page.
Groq: GPT OSS 120B (openai/gpt-oss-120b)
Production model, about 500 tokens/s. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
Groq: GPT OSS 20B (openai/gpt-oss-20b)
Production model, about 1000 tokens/s. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
Groq: Llama 3.1 8B (llama-3.1-8b-instant)
Price shown as 'Contact Sales' (tagged Enterprise); no public per-token price. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
Groq: Llama 3.3 70B (llama-3.3-70b-versatile)
Price shown as 'Contact Sales' (tagged Enterprise); no public per-token price. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
Groq: MiniMax M2.7 (minimaxai/minimax-m2.7)
Preview, tagged Enterprise; price shown as 'Contact Sales'. groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
Groq: Qwen/Qwen3.8-27B (qwen/qwen3.8-27b)
PREVIEW model (evaluation only, may be discontinued at short notice). groq.com/pricing now redirects (308) to the homepage, so prices come from GroqDocs 'Supported Models' (console.groq.com/docs/models), the provider's own public page. No cached-input price shown on that page.
Hyperstack: OpenAI gpt-oss-120b
From the 'Token-based Pricing' table on Hyperstack's GPU pricing page.
Hyperstack: Llama 3.1 8B
From the 'Token-based Pricing' table on Hyperstack's GPU pricing page.
Hyperstack: Llama 3.3 70B
From the 'Token-based Pricing' table on Hyperstack's GPU pricing page.
Crusoe: Gemma 4 31B-it
From the 'Serverless Inference pricing' table on Crusoe's cloud pricing page.
Crusoe: GLM 5.3
From the 'Serverless Inference pricing' table on Crusoe's cloud pricing page.
Crusoe: GPT-OSS 120B
From the 'Serverless Inference pricing' table on Crusoe's cloud pricing page.

Pages we could not use as published, and what we used instead:

  • https://openai.com/api/pricing/: HTTP 403 (bot protection). Used https://developers.openai.com/api/docs/pricing instead (platform.openai.com/docs/pricing redirects there).
  • https://groq.com/pricing: HTTP 308 redirect to the groq.com homepage, which has no prices. Used https://console.groq.com/docs/models (GroqDocs Supported Models, which lists per-model prices).
  • https://mistral.ai/pricing: Lists plans only, with no per-model API price table in the raw HTML or via WebFetch; the only API figure is an FAQ example ('Mistral Large costs $0.5/M in, $1.5/M out'). Used https://docs.mistral.ai/inference/pricing instead.
  • https://fireworks.ai/pricing: No per-token serverless price table (only embeddings, training and GPU prices); it links to https://docs.fireworks.ai/serverless/pricing, which was used.

How the costs are calculated

API cost per month       = input tokens x input price + output tokens x output price
                           (prompt-cached input at its own price, if the API has one)
GPU busy hours per month = input tokens / prefill tokens per second
                         + output tokens / decode tokens per second        (per replica)
Dedicated                = replicas x GPUs per replica x 730 hours x price per GPU-hour
                           replicas = busy hours / (730 x utilization), rounded up, at least 1
Hourly                   = busy hours / utilization x GPUs per replica x price per GPU-hour
Break-even volume        = the monthly volume, at your input/output mix, where the two costs are equal

A replica is one running copy of the model. If the model needs two GPUs, as Llama 3.3 70B at BF16 does on 80 GB cards, then a replica is two GPUs.

Throughput is the assumption that decides this

Every price on the page is a list price you can check. Throughput is the one number nobody can look up for you. It depends on the model, the precision, the GPU, the engine and its version, how many requests run at once, and how long your prompts and answers are. Change it by a factor of two and the answer can flip. That is why the result always shows the cost at a quarter, half, double and four times your figure.

Two throughputs, not one:

With continuous batching, both run on the same GPUs and compete for the same time, so the calculator adds their busy hours. Use the total decode rate across all concurrent streams at the concurrency you will actually run, not the speed of one stream.

To measure it, run your engine's serving benchmark (vLLM ships one as vllm bench serve) with your real distribution of prompt and output lengths, at the concurrency you plan to serve. Read input and output tokens per second. Measured throughput figures on this site are planned; none are published yet.

Utilization: you pay for idle GPUs

An API bills for tokens. A dedicated GPU bills for hours, busy or not. Traffic is uneven: if your peak hour carries five times your average load and you size for the peak, your replicas average about 20% busy. The utilization input is the share of paid GPU time spent working at the throughput you entered. The busy-time figure in the result shows what your inputs imply.

Hourly mode assumes you can stop paying the moment there is no traffic, for example with serverless GPUs billed per second. It ignores cold starts, minimum billing periods and capacity that is not there when you ask for it, so treat it as the best case.

A worked example

These are the calculator's defaults. The throughput figures are placeholders, so the example shows the method, not the market.

Self-hosting breaks even at about 4.9 billion tokens a month, 13.6 times this volume. Below that, no throughput can help, because one idle replica already costs more than the whole API bill.

At twenty times the volume (7.2 billion tokens a month), the same replica is busy a third of the time and costs $5,095 against an API bill of $7,488. Self-hosting wins, but only just: if real throughput is half the placeholder, it takes two replicas ($10,191) and the API wins again. That is the throughput point in practice.

When self-hosting is worth it anyway

Cost is not the only reason. Self-host when:

Use an API when traffic is spiky or small, when an API serves the same open model at a price no single GPU can match, or when you have nobody to keep a serving stack up at night.

What is not included

Renting the GPUs

For testing a model before you commit, by-the-hour providers are the cheapest way to measure real throughput: RunPod and the Vast.ai marketplace for single cards, and Lambda and DigitalOcean for H100, H200 and B200 instances. Compare current prices in the GPU price table, and size the replica with the VRAM calculator.

Also available as Markdown.