Free calculator
LLM VRAM calculator: weights, KV cache, and which GPUs fit
How much GPU memory a model needs at your context length and concurrency. Weights at BF16, FP8, INT8, INT4 or FP4, KV cache per token and in total, and which GPUs fit, alone or with tensor parallelism. The formula is on the page, and model shapes come from each model's config.json.
No paid links on this page: vendor links go straight to the vendor. Affiliate policy.
LLM VRAM calculator
Runs in your browser. Model shapes from each model's config.json on Hugging Face and GPU specs from vendor datasheets, read 7 October 2026. Both tables are below the calculator.
Model presets
Shapes copied from each model's config.json; parameter counts from the safetensors metadata on Hugging Face (what the checkpoint holds, including any vision encoder). Read 7 October 2026. Raw data: models.json.
| Model | Total params | Active | Attention layers (KV heads x head dim) | Max context | Published as | Source |
|---|---|---|---|---|---|---|
| Llama 3.1 8B Instruct | 8.0B | all | 32 full (8 KV x 128) | 131,072 | BF16 | config.json |
| Llama 3.3 70B Instruct | 70.6B | all | 80 full (8 KV x 128) | 131,072 | BF16 | config.json |
| Llama 3.1 405B Instruct | 405.9B | all | 126 full (8 KV x 128) | 131,072 | BF16 | config.json |
| Qwen3 8B | 8.2B | all | 36 full (8 KV x 128) | 40,960 | BF16 | config.json |
| Qwen3 14B | 14.8B | all | 40 full (8 KV x 128) | 40,960 | BF16 | config.json |
| Qwen3 32B | 32.8B | all | 64 full (8 KV x 128) | 40,960 | BF16 | config.json |
| Qwen2.5 7B Instruct | 7.6B | all | 28 full (4 KV x 128) | 32,768 | BF16 | config.json |
| Qwen2.5 72B Instruct | 72.7B | all | 80 full (8 KV x 128) | 32,768 | BF16 | config.json |
| Qwen3.8 27B | 27.8B | all | 16 full (4 KV x 256) + 48 linear (fixed state) | 262,144 | BF16 | config.json |
| Mistral 7B Instruct v0.3 | 7.2B | all | 32 full (8 KV x 128) | 32,768 | BF16 | config.json |
| Mistral Small 3.2 24B Instruct (2506) | 24.0B | all | 40 full (8 KV x 128) | 131,072 | BF16 | config.json |
| Gemma 4 31B IT | 31.3B | all | 50 sliding 1024 (16 KV x 256) + 10 full (4 KV x 512) | 262,144 | BF16 | config.json |
| Qwen3 30B-A3B Instruct (2507) | 30.5B | 3.3B | 48 full (4 KV x 128) | 262,144 | BF16 | config.json |
| Qwen3 235B-A22B Instruct (2507) | 235.1B | 22.0B | 94 full (4 KV x 128) | 262,144 | BF16 | config.json |
| Qwen3.6 35B-A3B | 36.0B | 3.0B | 10 full (2 KV x 256) + 30 linear (fixed state) | 262,144 | BF16 | config.json |
| Qwen3.5 122B-A10B | 125.1B | 10.0B | 12 full (2 KV x 256) + 36 linear (fixed state) | 262,144 | BF16 | config.json |
| Gemma 4 26B-A4B IT | 25.8B | 3.8B | 25 sliding 1024 (8 KV x 256) + 5 full (2 KV x 512) | 262,144 | BF16 | config.json |
| gpt-oss-20b | 20.9B | 3.6B | 12 sliding 128 (8 KV x 64) + 12 full (8 KV x 64) | 131,072 | MXFP4 | config.json |
| gpt-oss-120b | 116.8B | 5.1B | 18 sliding 128 (8 KV x 64) + 18 full (8 KV x 64) | 131,072 | MXFP4 | config.json |
| GLM-4.5-Air | 110.5B | 12.0B | 46 full (8 KV x 128) | 131,072 | BF16 | config.json |
| MiniMax-M2.7 | 228.7B | 11.0B | 62 full (8 KV x 128) | 204,800 | FP8 | config.json |
| DeepSeek-V3.2 | 685.4B | 41.0B | 61 MLA (latent 576) + 61 indexer | 163,840 | FP8 | config.json |
| GLM-5.3 | 753.3B | 41.8B | 78 MLA (latent 576) + 21 indexer | 1,048,576 | FP8 | config.json |
| Kimi K2 Instruct (0905) | 1026.5B | 32.0B | 61 MLA (latent 576) | 262,144 | FP8 | config.json |
Notes on each preset (hybrid attention, gated repositories, what the counts include)
- Llama 3.1 8B Instruct
- meta-llama/Llama-3.1-8B-Instruct is gated on Hugging Face, so its config.json was read from unsloth/Llama-3.1-8B-Instruct, an ungated copy. The parameter count comes from meta-llama/Llama-3.1-8B-Instruct's own public safetensors metadata.
- Llama 3.3 70B Instruct
- meta-llama/Llama-3.3-70B-Instruct is gated on Hugging Face, so its config.json was read from unsloth/Llama-3.3-70B-Instruct, an ungated copy. The parameter count comes from meta-llama/Llama-3.3-70B-Instruct's own public safetensors metadata.
- Llama 3.1 405B Instruct
- Architecture fields read from unsloth's ungated copy of the config (the file also carries a bitsandbytes quantization block, which does not change the shapes). head_dim is not stated in this config; it is hidden_size / num_attention_heads = 128. meta-llama/Llama-3.1-405B-Instruct is gated on Hugging Face, so its config.json was read from unsloth/Meta-Llama-3.1-405B-Instruct-bnb-4bit, an ungated copy. The parameter count comes from meta-llama/Llama-3.1-405B-Instruct's own public safetensors metadata.
- Qwen3 8B
- Standard grouped-query attention on every layer.
- Qwen3 14B
- Standard grouped-query attention on every layer.
- Qwen3 32B
- Standard grouped-query attention on every layer.
- Qwen2.5 7B Instruct
- config.json sets sliding_window but use_sliding_window is false, so every layer is full attention.
- Qwen2.5 72B Instruct
- config.json sets sliding_window but use_sliding_window is false, so every layer is full attention.
- Qwen3.8 27B
- Hybrid attention: 3 of every 4 layers are Gated DeltaNet linear attention with a fixed-size state per sequence; only the full-attention layers keep a KV cache. Parameter count includes the vision encoder.
- Mistral 7B Instruct v0.3
- Standard grouped-query attention on every layer.
- Mistral Small 3.2 24B Instruct (2506)
- Parameter count includes the vision encoder.
- Gemma 4 31B IT
- 5 of every 6 layers use 1,024-token sliding-window attention; the global layers use 4 KV heads of dimension 512 with keys and values from one projection (attention_k_eq_v). We count both K and V as cached, the conservative reading. Parameter count includes the vision encoder.
- Qwen3 30B-A3B Instruct (2507)
- Active parameters: model card: '30.5B in total and 3.3B activated'.
- Qwen3 235B-A22B Instruct (2507)
- Active parameters: model card: '235B in total and 22B activated'.
- Qwen3.6 35B-A3B
- Hybrid attention (Gated DeltaNet + full attention every 4th layer). Parameter count includes the vision encoder. Active parameters: model card: '35B in total and 3B activated'.
- Qwen3.5 122B-A10B
- Hybrid attention (Gated DeltaNet + full attention every 4th layer). Parameter count includes the vision encoder. Active parameters: model card: '122B in total and 10B activated'.
- Gemma 4 26B-A4B IT
- Sliding-window (1,024 tokens) on 5 of every 6 layers; global layers use 2 KV heads of dimension 512 (attention_k_eq_v; both K and V counted). Parameter count includes the vision encoder. Active parameters: model card: 'Active Parameters 3.8B'.
- gpt-oss-20b
- Published with MoE weights in MXFP4; attention, router, embeddings and LM head stay BF16. Alternating 128-token sliding-window and full-attention layers. Active parameters: model card: '21B parameters with 3.6B active parameters'.
- gpt-oss-120b
- Published with MoE weights in MXFP4; attention, router, embeddings and LM head stay BF16. Alternating 128-token sliding-window and full-attention layers. Active parameters: model card: '117B parameters with 5.1B active parameters'.
- GLM-4.5-Air
- Parameter count from the checkpoint (110.5B) includes the multi-token-prediction layer; the model card's 106B does not. Active parameters: model card: '106 billion total parameters and 12 billion active parameters'.
- MiniMax-M2.7
- Published as an FP8 (block-scaled) checkpoint. Active parameters: computed from config.json: total parameters minus the routed experts a token does not use, (experts - experts_per_token) x 3 x hidden_size x moe_intermediate_size per MoE layer; the publisher's card does not state it. Includes embeddings..
- DeepSeek-V3.2
- Multi-head latent attention (MLA) caches one 576-dim latent per token per layer, shared by all heads, plus a small FP8 key for the sparse-attention indexer. Published as FP8. Parameter count includes the 1-layer multi-token-prediction module, which is loaded only for speculative decoding. Active parameters: computed from config.json: total parameters minus the routed experts a token does not use, (experts - experts_per_token) x 3 x hidden_size x moe_intermediate_size per MoE layer; the publisher's card does not state it. Includes embeddings and the MTP layer..
- GLM-5.3
- DeepSeek-style MLA with a sparse-attention indexer; config.json marks 21 of 78 layers as having their own indexer ('full') and 57 as reusing one ('shared'), so the indexer cache is counted on 21 layers. FP8 KV layout assumed to match DeepSeek-V3.2's in vLLM. Parameter count includes the MTP layer. Active parameters: computed from config.json: total parameters minus the routed experts a token does not use, (experts - experts_per_token) x 3 x hidden_size x moe_intermediate_size per MoE layer; the publisher's card does not state it. Includes embeddings and the MTP layer..
- Kimi K2 Instruct (0905)
- DeepSeek-V3 architecture (MLA). Published as FP8. Active parameters: model card: '32 billion activated parameters and a total of 1 trillion parameters'.
GPU specifications
From each vendor's datasheet, product page or architecture whitepaper. Compute is dense tensor-core TFLOPS (no sparsity). "n/a" means the vendor does not publish it or the GPU lacks that data type. Raw data: gpus.json.
| GPU | Memory | Bandwidth | FP16/BF16 | FP8 | FP4 | GPU-to-GPU | Source |
|---|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 4090 | 24 GB GDDR6X | 1.01 TB/s | 165.2 | 330.3 | n/a | PCIe Gen4, no NVLink | NVIDIA |
| NVIDIA GeForce RTX 5090 | 32 GB GDDR7 | 1.79 TB/s | 209.5 | 419 | 1676 | PCIe Gen5, no NVLink | NVIDIA |
| NVIDIA RTX A6000 | 48 GB GDDR6 | 0.77 TB/s | 154.8 | n/a | n/a | NVLink bridge (2 GPUs) 112.5 GB/s bidirectional; PCIe 4.0 x16 | NVIDIA |
| NVIDIA RTX 6000 Ada Generation | 48 GB GDDR6 | 0.96 TB/s | 364 | 728.5 | n/a | PCIe 4.0 x16, no NVLink | NVIDIA |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 96 GB GDDR7 | 1.60 TB/s | n/a | n/a | n/a | PCIe Gen5 x16, no NVLink | NVIDIA |
| NVIDIA L4 | 24 GB GDDR6 | 0.30 TB/s | 121 | 242 | n/a | PCIe Gen4 x16 64 GB/s, no NVLink listed | NVIDIA |
| NVIDIA A10 | 24 GB GDDR6 | 0.60 TB/s | 125 | n/a | n/a | PCIe Gen4 64 GB/s, no NVLink listed | NVIDIA |
| NVIDIA A40 | 48 GB GDDR6 | 0.70 TB/s | 149.7 | n/a | n/a | NVLink bridge (2-way) 112.5 GB/s bidirectional; PCIe Gen4 | NVIDIA |
| NVIDIA L40S | 48 GB GDDR6 | 0.86 TB/s | 362.05 | 733 | n/a | PCIe Gen4 x16 64 GB/s, no NVLink | NVIDIA |
| NVIDIA A100 40GB SXM | 40 GB HBM2 | 1.55 TB/s | 312 | n/a | n/a | NVLink 600 GB/s; PCIe Gen4 64 GB/s | NVIDIA |
| NVIDIA A100 80GB PCIe | 80 GB HBM2e | 1.94 TB/s | 312 | n/a | n/a | NVLink bridge for 2 GPUs 600 GB/s; PCIe Gen4 64 GB/s | NVIDIA |
| NVIDIA A100 80GB SXM | 80 GB HBM2e | 2.04 TB/s | 312 | n/a | n/a | NVLink 600 GB/s; PCIe Gen4 64 GB/s | NVIDIA |
| NVIDIA H100 PCIe | 80 GB HBM2e | 2.00 TB/s | 756 | 1513 | n/a | NVLink bridge (2 GPUs) 600 GB/s; PCIe Gen5 x16 | NVIDIA |
| NVIDIA H100 NVL | 94 GB HBM3 | 3.90 TB/s | 835.5 | 1670.5 | n/a | NVLink bridge 600 GB/s; PCIe Gen5 128 GB/s | NVIDIA |
| NVIDIA H100 SXM | 80 GB HBM3 | 3.35 TB/s | 989.4 | 1978.9 | n/a | NVLink 900 GB/s; PCIe Gen5 128 GB/s | NVIDIA |
| NVIDIA H200 NVL | 141 GB HBM3e | 4.80 TB/s | 835.5 | 1670.5 | n/a | 2- or 4-way NVLink bridge 900 GB/s per GPU; PCIe Gen5 128 GB/s | NVIDIA |
| NVIDIA H200 SXM | 141 GB HBM3e | 4.80 TB/s | 989.5 | 1979 | n/a | NVLink 900 GB/s; PCIe Gen5 128 GB/s | NVIDIA |
| NVIDIA B200 (HGX B200) | 180 GB HBM3E | 8.00 TB/s | 2250 | 4500 | 9000 | NVLink 5 1.8 TB/s GPU-to-GPU (NVLink Switch) | NVIDIA |
| NVIDIA B300 (HGX B300, Blackwell Ultra) | 270 GB HBM3E | 7.70 TB/s | 2250 | 4500 | 14000 | NVLink 5 1.8 TB/s; PCIe Gen6 256 GB/s | NVIDIA |
| AMD Instinct MI300X | 192 GB HBM3 | 5.30 TB/s | 1307.4 | 2614.9 | n/a | AMD Infinity Fabric: 8 links, 128 GB/s peak link bandwidth; PCIe 5.0 x16 | AMD |
| AMD Instinct MI325X | 256 GB HBM3E | 6.00 TB/s | 1307.4 | 2614.9 | n/a | AMD Infinity Fabric: 8 links, 128 GB/s peak link bandwidth; PCIe 5.0 x16 | AMD |
| AMD Instinct MI355X | 288 GB HBM3E | 8.00 TB/s | 2516.6 | 5033.2 | 10066.3 | AMD Infinity Fabric: 7 x 153.6 GB/s scale-up links; PCIe Gen5 x16 128 GB/s | AMD |
How each GPU's figures were read (sparsity, accumulate rate, conflicting sources)
- NVIDIA GeForce RTX 4090
- Product page lists only '1321 AI TOPS', 24 GB GDDR6X, 'Total Graphics Power (W) 450', 'NVLink (SLI-Ready) No'. Ada whitepaper V2.02 Appendix A: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 165.2/330.4' (used) vs 'with FP16 Accumulate 330.3/660.6'; FP8 'with FP32 Accumulate 330.3/660.6' (used, matches FP32-accumulate choice) vs 'with FP16 Accumulate 660.6/1321.2'; second figure = sparsity. Bandwidth '1008 GB/sec' from whitepaper. Second source: https://images.nvidia.com/aem-dam/Solutions/Data-Center/l....
- NVIDIA GeForce RTX 5090
- Product page lists only '3352 AI TOPS', '32 GB GDDR7', 'Total Graphics Power (W) 575', 'NVLink (SLI-Ready) No'. RTX Blackwell whitepaper v1.1 Table 3: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 209.5/419' (used) vs 'with FP16 Accumulate 419/838'; FP8 'with FP32 Accumulate 419/838' (used) vs 'with FP16 Accumulate 838/1676'; 'Peak FP4 Tensor TFLOPS with FP32 Accumulate (FP4 AI TOPS) 1676/3352'; second figure = sparsity. Bandwidth '1792 GB/sec' from whitepaper. Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/black....
- NVIDIA RTX A6000
- Datasheet: 'Tensor performance 309.7 TFLOPS' with footnote 'Effective teraFLOPS (TFLOPS) using the new sparsity feature'. RTX PRO Blackwell whitepaper v1.1 Table 4 gives dense explicitly: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 154.8/309.6' (same as FP16 accumulate for this pro card). FP8 'N/A' (Ampere has no FP8 tensor cores). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/design....
- NVIDIA RTX 6000 Ada Generation
- Datasheet: 'Tensor performance 1457.0 TFLOPS' = 'Effective FP8 teraFLOPS (TFLOPS) using the new sparsity feature'; 'NVIDIA NVLink No'. RTX PRO Blackwell whitepaper v1.1 Table 4: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 364/728' and 'Peak FP8 Tensor TFLOPS with FP32 Accumulate 728.5/1457' (FP16-accumulate rates identical on this pro card). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/design....
- NVIDIA RTX PRO 6000 Blackwell Server Edition
- Compute left null: product page lists 'FP4 Tensor Core 4 PFLOPS', 'FP8 Tensor Core 2 PFLOPS', 'FP16 | BF16 Tensor Core 1 PFLOP' with NO sparsity marker, and the datasheet only gives 'Peak FP4 AI PFLOPS 4 PFLOPS'. NVIDIA's RTX PRO Blackwell whitepaper v1.1 lists the same-GPU Workstation Edition (GB202, 752 Tensor Cores, 126 TFLOPS FP32 vs 120 for Server Edition) at FP16 503.8/1007.6, FP8 1007.6/2015.2, FP4 2015.2/4030.4 TFLOPS (dense/sparse), so the page figures appear to be sparse; halving would give about 500/1000/2000 TFLOPS, but NVIDIA publishes no dense figure for the Server Edition. Product brief: 'Memory type GDDR7', 'Peak memory bandwidth 1,597 GB/s', 600 W, 'NVIDIA NVLink Not supported'. Second source: https://dam-cdn.nvd.orangelogic.com/AssetLink/3km2720jiy7....
- NVIDIA L4
- Product page: 'FP16 Tensor Core 242 teraFLOPS*', 'FP8 Tensor Core 485 teraFLOPs*', '* Shown with sparsity. Specifications are one-half lower without sparsity.' Ada whitepaper V2.02 Table 5 lists dense explicitly: 'FP16 Tensor Core Performance 121 | 242 TFLOPS', 'FP8 ... 242 | 485 TFLOPS', '24GB GDDR6 w/ ECC'. Neither source mentions NVLink; interconnect listed as PCIe only. Second source: https://images.nvidia.com/aem-dam/Solutions/Data-Center/l....
- NVIDIA A10
- Product page: 'FP16 Tensor Core 125 teraFLOPS | 250 teraFLOPS*', '*With Sparsity' (dense explicit). No FP8 tensor cores (Ampere). Interconnect listed only as 'PCIe Gen4 64GB/s'.
- NVIDIA A40
- Datasheet: 'Peak FP16 Tensor TFLOPS with FP16 Accumulate 149.7 | 299.4*' and 'Peak BF16 Tensor TFLOPS with FP32 Accumulate 149.7 | 299.4*', '* Structural sparsity enabled'. No FP8 tensor cores (Ampere). Datasheet lists PCIe Gen4 as 31.5 GB/s, product page as 64GB/s. Second source: https://www.nvidia.com/en-us/data-center/a40/.
- NVIDIA L40S
- Product page full spec table: 'FP16 Tensor Core 362.05 I 733*', 'FP8 Tensor Core 733 I 1,466*', '*With Sparsity' (dense listed explicitly; note 362.05 is not exactly half of 733). '48GB GDDR6 with ECC', 'Max Power Consumption 350W', 'NVIDIA NVLink Support: No'.
- NVIDIA A100 40GB SXM
- A100 datasheet (r4, 2021): 'FP16 Tensor Core 312 TFLOPS | 624 TFLOPS*', '* With sparsity'; A100 40GB SXM column: '40GB HBM2', '1,555GB/s', TDP 400W. No FP8 tensor cores (Ampere). The current A100 product page lists only the 80GB variants.
- NVIDIA A100 80GB PCIe
- Product page: 'FP16 Tensor Core 312 TFLOPS | 624 TFLOPS*', '* With sparsity' (dense explicit). No FP8 tensor cores (Ampere). NVLink only via bridge pairing two cards.
- NVIDIA A100 80GB SXM
- Product page: 'FP16 Tensor Core 312 TFLOPS | 624 TFLOPS*', '* With sparsity' (dense listed explicitly). No FP8 tensor cores (Ampere). '400W TDP for standard configuration. HGX A100-80GB custom thermal solution (CTS) SKU can support TDPs up to 500W'.
- NVIDIA H100 PCIe
- Hopper whitepaper v1.04 ('Includes final GPU / memory clocks and final TFLOPS performance specs') Table 3, H100 PCIe: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 756/1513' and 'Peak FP8 Tensor TFLOPS 1513/3026' (second figure = sparsity). Product brief PB-11133 v02: 'Memory type HBM2e', 'Peak memory bandwidth 2,000 GB/s', 350 W max, 'Total maximum NVLink bandwidth 600 Gbytes per second' (its overview text says 900 GB/s; table value used). Whitepaper lists bandwidth as 2039 GB/sec. NVIDIA's current H100 product page no longer lists H100 PCIe; an older 2022 H100 datasheet listed preliminary rounded figures (1,600 TFLOPS* FP16, 2TB/s), not used. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/gtcs22....
- NVIDIA H100 NVL
- Product page: 'FP16 Tensor Core* 1,671 teraFLOPS', 'FP8 Tensor Core* 3,341 teraFLOPS', '* With sparsity'; dense values are half. TDP '350-400W (configurable)' (400 W recorded). Product brief PB-11773: 'Memory type HBM3', 'Memory size 94 GB', 'Peak memory bandwidth 3,938 GB/s' (page says 3.9TB/s). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-C....
- NVIDIA H100 SXM
- Product page: 'FP16 Tensor Core* 1,979 teraFLOPS', 'FP8 Tensor Core* 3,958 teraFLOPS', '* With sparsity'; Hopper whitepaper v1.04 Table 3 gives dense explicitly: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 989.4/1978.9', FP8 1978.9/3957.8 (sparse after slash), and '80 GB HBM3'. TDP listed as 'Up to 700W (configurable)'. Second source: https://dam-cdn.nvd.orangelogic.com/AssetLink/705n6ur546g....
- NVIDIA H200 NVL
- Product page: 'FP16 Tensor Core² 1,671 TFLOPS', 'FP8 Tensor Core² 3,341 TFLOPS', '² With sparsity'; dense = half. Marked '¹ Preliminary specifications'. TDP 'Up to 600W (configurable)'. HBM3e from page text about H200 memory.
- NVIDIA H200 SXM
- Product page: 'FP16 Tensor Core² 1,979 TFLOPS', 'FP8 Tensor Core² 3,958 TFLOPS', '² With sparsity'; dense = half. Page footnote '¹ Preliminary specifications. May be subject to change.' Page text: '141 gigabytes (GB) of HBM3e memory at 4.8 terabytes per second'. TDP 'Up to 700W (configurable)'.
- NVIDIA B200 (HGX B200)
- No per-GPU B200 datasheet found; compute derived by dividing NVIDIA's 8-GPU HGX B200 totals by 8. HGX page: 'FP4 Tensor Core 144 PFLOPS | 72 PFLOPS' ('Sparse | Dense'), 'FP8/FP6 Tensor Core 72 PFLOPS', 'FP16/BF16 Tensor Core 36 PFLOPS' ('Specification in Sparse. Dense is 1/2 sparse spec shown.'). PCF summary: 'eight NVIDIA Blackwell B200 GPUs, each with 180 GB of HBM3E', 'Per individual GPU: Configurable up to 1000 W', total bandwidth 'Up to 62 TB/s'. Per-GPU bandwidth 'Up to 8TB/s' from https://docs.nvidia.com/enterprise-reference-architectures/hgx-ai-factory/latest/components.html (62/8 = 7.75 TB/s would follow from the PCF total). Second source: https://images.nvidia.com/aem-dam/Solutions/documents/HGX....
- NVIDIA B300 (HGX B300, Blackwell Ultra)
- NVIDIA Blackwell Ultra Datasheet, 'Individual Blackwell Ultra GPU Specifications', HGX B300 column: 'FP4 Tensor Core 18 PFLOPS | 14 PFLOPS' (Sparse | Dense), 'FP8/FP6 9 PFLOPS', 'FP16/BF16 4.5 PFLOPS' ('Specification in sparse. Dense is 1/2 sparse spec shown'), '270 GB HBM3E | 7.7 TB/s', 'Configurable up to 1,100 W' (GB300 NVL72 variant: 279 GB, 8 TB/s, 1,400 W). Conflicts: HGX page 8-GPU FP4 dense 108 PFLOPS (=13.5/GPU) and docs.nvidia.com reference architecture says B300 SXM 288GB, up to 8TB/s; datasheet per-GPU values used. Second source: https://www.nvidia.com/en-us/data-center/hgx/.
- AMD Instinct MI300X
- Product page: 'Peak Half Precision (FP16) Performance 1.3 PFLOPs' and 'with Structured Sparsity 2.61 PFLOPs'; FP8 '2.61 PFLOPs' / sparsity '5.22 PFLOPs'. Page footnote gives precise dense values: '1307.4 TFLOPS peak theoretical half precision (FP16)', '2614.9 TFLOPS peak theoretical 8-bit precision (FP8)'. TBP '750W Peak'. No FP4 listed.
- AMD Instinct MI325X
- Product page: FP16 '1.3 PFLOPs' (sparsity '2.61 PFLOPs'), FP8 '2.61 PFLOPs' (sparsity '5.22 PFLOPs'); footnote MI325-002: '1307.4 TFLOPS peak theoretical half precision (FP16)... 2614.9 TFLOPS peak theoretical 8-bit precision (FP8)'. TBP '1000W Peak'. No FP4 listed.
- AMD Instinct MI355X
- Product page: 'Peak Half Precision Matrix (FP16) Performance 2.5 PFLOPs' (sparsity '5 PFLOPs'), OCP-FP8 '5 PFLOPs' (sparsity '10.1 PFLOPs'), 'MXFP4 Performance 10.1 PFLOPs' (no sparsity figure), TBP '1400W'. Brochure gives precise values: FP16 matrix 2.5166 PFLOPS (5.0332 w/ sparsity), OCP-FP8 5.0332 (10.0664 w/ sparsity), MXFP4 10.0663 (sparsity N/A). FP4 is MXFP4 (microscaling). Second source: https://www.amd.com/content/dam/amd/en/documents/instinct....
The formula
Serving memory has three parts: the weights, the KV cache, and a runtime allowance. Only the KV cache depends on your traffic.
weights = parameters x bytes per parameter
(quantized formats keep embeddings and the LM head at 16-bit)
KV cache per token = 2 x layers x KV heads x head dimension x bytes per value
KV cache in use = KV per token x tokens per sequence x concurrent sequences
fits when weights / TP + allowance + KV cache per GPU <= gpu_memory_utilization x GPU memory
The 2 counts keys and values. TP is the tensor-parallel size, the number of GPUs one copy of the model is split across.
Worked example. Llama 3.1 8B has 32 layers, 8 KV heads and a head dimension of 128, so at BF16 the cache is 2 x 32 x 8 x 128 x 2 bytes = 131,072 bytes (128 KiB) per token. A sequence at 8,192 tokens holds exactly 1 GiB. The weights are 8,030,261,248 parameters x 2 bytes = 14.96 GiB. An RTX 4090 has 24 GB; vLLM's default gpu_memory_utilization of 0.92 makes 22.08 GiB usable. Take away the weights and a 2 GiB allowance and 5.1 GiB is left, which is room for five concurrent 8k sequences. That is arithmetic from the formula, not a measurement.
KV heads, not attention heads
The cache stores keys and values for KV heads, and most current models have far fewer of those than query heads. Llama 3.1 8B has 32 query heads sharing 8 KV heads (grouped-query attention), which makes its cache four times smaller than if every head kept its own keys and values. Llama 3.3 70B also has only 8 KV heads, across 80 layers: 320 KiB per token. Qwen3 32B has 8 KV heads across 64 layers: 256 KiB per token. Calculators that use num_attention_heads overstate the cache by the grouping factor.
What changes the answer
- Paged attention. vLLM allocates the cache in blocks (16 tokens each by default) as sequences grow, so memory follows the tokens actually in use, not the maximum length you allow. That is why the calculator multiplies by your real context length. vLLM still refuses to start if a single sequence at
max_model_lencannot fit. - Prefix caching. Sequences that share a prompt prefix, such as a long system prompt, share its cache blocks. The calculator assumes no sharing, so for chat traffic with a common prefix it is an upper bound.
- KV-cache quantization. An FP8 cache (
--kv-cache-dtype fp8in vLLM) halves cache memory. The effect on output quality depends on the model and the task. Long-context retrieval is a reasonable place to look first, but measure it on your own evaluation before you rely on it. - Weight quantization is not "parameters x bits". Quantized checkpoints usually keep the embedding table and LM head at 16-bit, and group-wise formats store scales. For Llama 3.1 8B, the embeddings and LM head are 1.05 billion of the 8.03 billion parameters (128,256-token vocabulary x 4,096 hidden x 2). INT4 at group size 128 costs 4.16 bits per weight once the scale and zero point are counted. So the 4-bit model is 5.33 GiB, not 4 GB.
- Mixture of experts. Every expert has to be in memory, so a MoE model's memory follows its total parameters. Its decode speed follows its active parameters. Qwen3 30B-A3B needs 61 GB (56.9 GiB) at BF16 but reads only about 3.3 billion parameters per token.
- Sliding-window layers. gpt-oss alternates full-attention layers with 128-token sliding-window layers, and Gemma 4 runs 1,024-token windows on five of every six layers. A sliding-window layer only ever caches its window. Engines with a hybrid KV-cache manager (vLLM V1) allocate it that way, and so does the calculator. At 128k tokens, Gemma 4 31B needs 10.8 GiB per sequence. An engine that cached the full context on every layer would need 110 GiB.
- Linear attention. Qwen3.5, 3.6 and 3.8 use Gated DeltaNet on three of every four layers. Those layers keep a fixed recurrent state per sequence instead of a growing cache. For Qwen3.8 27B, that state is 147 MiB per sequence, stored in FP32 as its config specifies. The full-attention layers add 64 KiB per token. At short context the state dominates; at long context the cache does.
- Multi-head latent attention (MLA). DeepSeek-V3.2, Kimi K2 and GLM-5.3 cache a single 576-value latent per token per layer, shared by all heads. That makes the cache small: 70,272 bytes per token for Kimi K2 at BF16, against 327,680 for Llama 3.3 70B. DeepSeek-V3.2 also caches a 132-byte FP8 key per layer for its sparse-attention indexer.
Tensor parallelism does not always split the cache
With tensor parallelism, vLLM splits KV heads across GPUs, but a GPU never holds less than one head. When the tensor-parallel size reaches the number of KV heads, heads are replicated. Qwen3.5 122B-A10B has only 2 KV heads on its full-attention layers, so going from 2 to 8 GPUs barely shrinks the per-GPU cache (457 MiB to 402 MiB per 32k-token sequence; the part that still shrinks is the linear-attention state).
MLA is the extreme case. The latent is shared by all heads, so every tensor-parallel GPU keeps the full cache. At 32k tokens, a DeepSeek-V3.2 sequence takes 2.4 GiB on each of 8 GPUs. A Llama 3.3 70B sequence of the same length takes 10 GiB in total, which splits to 1.25 GiB per GPU. This is why large MLA deployments use data-parallel attention. The calculator models tensor parallelism only, and says which GPU counts a model's heads allow.
The runtime allowance
When vLLM starts, it loads the weights. Then it runs a profiling forward pass to measure peak activation memory and accounts for memory outside PyTorch's allocator (the CUDA context, NCCL buffers) and for CUDA graphs. Whatever is left of gpu_memory_utilization x GPU memory becomes KV cache. The calculator stands in for the middle part with one number per GPU. The default of 2 GiB is a stated round number, not a measurement. vLLM logs the real figures at startup, so replace it with yours.
GPU memory is the marketed figure treated as GiB. The capacity the driver reports can be a few percent lower, for example with ECC enabled, so leave headroom if you are within a few percent of the limit.
The decode ceiling column
Generating one token means reading every active weight and the sequence's whole cache from GPU memory. Memory bandwidth divided by those bytes is therefore an upper bound on single-stream decode speed. For Llama 3.1 8B at BF16 on an H100 SXM (3,350 GB/s) at 8k context, that is 3,350 GB/s / 17.1 GB = 196 tokens per second. Real engines land below it. Batching reads the weights once for many sequences, so total throughput goes far above one stream. Use the ceiling to compare GPUs, and the break-even calculator with your own measured throughput to compare costs. Measured numbers on this site are planned, not published.
What the calculator does not model
- Pipeline and expert parallelism, and data-parallel attention.
- Speculative decoding. A draft model, or a model's multi-token-prediction layer, adds weights and its own cache.
- LoRA adapters, CPU offload, and GGUF k-quant formats (which mix bit widths per tensor).
- Activations from image or audio inputs. The weights of the multimodal presets do include the vision encoder.
Where to run it
If the model fits a 24 to 32 GB card, hourly rentals of RTX 4090 and RTX 5090 cards are the cheapest way to test it: RunPod lists both, and Vast.ai is a marketplace where individual hosts set the price. For 70B-class models and long contexts you need 80 to 180 GB data-center GPUs from providers such as Lambda or DigitalOcean. The GPU price table compares on-demand prices from eleven GPU clouds and from AWS, Google Cloud and Azure, most of which pay this site nothing.
Before you rent anything, run the break-even calculator. For many open models, an API serving the same model costs less than one idle GPU.
Also available as Markdown.