# LLM VRAM calculator: weights, KV cache, and which GPUs fit

> How much GPU memory a model needs at your context length and concurrency. Weights at BF16, FP8, INT8, INT4 or FP4, KV cache per token and in total, and which GPUs fit, alone or with tensor parallelism. The formula is on the page, and model shapes come from each model's config.json.

By GPUCostLab. Updated 7 October 2026. Canonical URL: https://gpucostlab.com/llm-vram-calculator/

No paid links on this page: vendor links go straight to the vendor.

> **Interactive calculator.** Pick a model, weight and KV-cache precision, context length and concurrent sequences on the HTML version of this page (https://gpucostlab.com/llm-vram-calculator/). It returns weights, KV cache per token and in total, and which GPUs fit, singly or with tensor parallelism. The data it uses is below.

### Model presets (read 2026-10-07)

| Model | Total params | Active params | Layers | KV heads | Head dim | Max context | Published as | config.json |
|---|---|---|---|---|---|---|---|---|
| Llama 3.1 8B Instruct | 8.0B | all | 32 | 8 | 128 | 131072 | BF16 | https://huggingface.co/unsloth/Llama-3.1-8B-Instruct/blob/main/config.json |
| Llama 3.3 70B Instruct | 70.6B | all | 80 | 8 | 128 | 131072 | BF16 | https://huggingface.co/unsloth/Llama-3.3-70B-Instruct/blob/main/config.json |
| Llama 3.1 405B Instruct | 405.9B | all | 126 | 8 | 128 | 131072 | BF16 | https://huggingface.co/unsloth/Meta-Llama-3.1-405B-Instruct-bnb-4bit/blob/main/config.json |
| Qwen3 8B | 8.2B | all | 36 | 8 | 128 | 40960 | BF16 | https://huggingface.co/Qwen/Qwen3-8B/blob/main/config.json |
| Qwen3 14B | 14.8B | all | 40 | 8 | 128 | 40960 | BF16 | https://huggingface.co/Qwen/Qwen3-14B/blob/main/config.json |
| Qwen3 32B | 32.8B | all | 64 | 8 | 128 | 40960 | BF16 | https://huggingface.co/Qwen/Qwen3-32B/blob/main/config.json |
| Qwen2.5 7B Instruct | 7.6B | all | 28 | 4 | 128 | 32768 | BF16 | https://huggingface.co/Qwen/Qwen2.5-7B-Instruct/blob/main/config.json |
| Qwen2.5 72B Instruct | 72.7B | all | 80 | 8 | 128 | 32768 | BF16 | https://huggingface.co/Qwen/Qwen2.5-72B-Instruct/blob/main/config.json |
| Qwen3.8 27B | 27.8B | all | 64 | 4 | 256 | 262144 | BF16 | https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/config.json |
| Mistral 7B Instruct v0.3 | 7.2B | all | 32 | 8 | 128 | 32768 | BF16 | https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3/blob/main/config.json |
| Mistral Small 3.2 24B Instruct (2506) | 24.0B | all | 40 | 8 | 128 | 131072 | BF16 | https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506/blob/main/config.json |
| Gemma 4 31B IT | 31.3B | all | 60 | 16 | 256 | 262144 | BF16 | https://huggingface.co/google/gemma-4-31B-it/blob/main/config.json |
| Qwen3 30B-A3B Instruct (2507) | 30.5B | 3.3B | 48 | 4 | 128 | 262144 | BF16 | https://huggingface.co/Qwen/Qwen3-30B-A3B-Instruct-2507/blob/main/config.json |
| Qwen3 235B-A22B Instruct (2507) | 235.1B | 22.0B | 94 | 4 | 128 | 262144 | BF16 | https://huggingface.co/Qwen/Qwen3-235B-A22B-Instruct-2507/blob/main/config.json |
| Qwen3.6 35B-A3B | 36.0B | 3.0B | 40 | 2 | 256 | 262144 | BF16 | https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/config.json |
| Qwen3.5 122B-A10B | 125.1B | 10.0B | 48 | 2 | 256 | 262144 | BF16 | https://huggingface.co/Qwen/Qwen3.5-122B-A10B/blob/main/config.json |
| Gemma 4 26B-A4B IT | 25.8B | 3.8B | 30 | 8 | 256 | 262144 | BF16 | https://huggingface.co/google/gemma-4-26B-A4B-it/blob/main/config.json |
| gpt-oss-20b | 20.9B | 3.6B | 24 | 8 | 64 | 131072 | MXFP4 | https://huggingface.co/openai/gpt-oss-20b/blob/main/config.json |
| gpt-oss-120b | 116.8B | 5.1B | 36 | 8 | 64 | 131072 | MXFP4 | https://huggingface.co/openai/gpt-oss-120b/blob/main/config.json |
| GLM-4.5-Air | 110.5B | 12.0B | 46 | 8 | 128 | 131072 | BF16 | https://huggingface.co/zai-org/GLM-4.5-Air/blob/main/config.json |
| MiniMax-M2.7 | 228.7B | 11.0B | 62 | 8 | 128 | 204800 | FP8 | https://huggingface.co/MiniMaxAI/MiniMax-M2.7/blob/main/config.json |
| DeepSeek-V3.2 | 685.4B | 41.0B | 61 | 128 | 56 | 163840 | FP8 | https://huggingface.co/deepseek-ai/DeepSeek-V3.2/blob/main/config.json |
| GLM-5.3 | 753.3B | 41.8B | 78 | 64 | 192 | 1048576 | FP8 | https://huggingface.co/zai-org/GLM-5.3/blob/main/config.json |
| Kimi K2 Instruct (0905) | 1026.5B | 32.0B | 61 | 64 | 112 | 262144 | FP8 | https://huggingface.co/moonshotai/Kimi-K2-Instruct-0905/blob/main/config.json |

### GPU specifications (vendor datasheets, dense TFLOPS)

| GPU | Memory (GB) | Bandwidth (GB/s) | FP16/BF16 | FP8 | FP4 | GPU-to-GPU | Source |
|---|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 4090 | 24 | 1008 | 165.2 | 330.3 | n/a | PCIe Gen4, no NVLink | https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/ |
| NVIDIA GeForce RTX 5090 | 32 | 1792 | 209.5 | 419 | 1676 | PCIe Gen5, no NVLink | https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/ |
| NVIDIA RTX A6000 | 48 | 768 | 154.8 | n/a | n/a | NVLink bridge (2 GPUs) 112.5 GB/s bidirectional; PCIe 4.0 x16 | https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/quadro-product-literature/proviz-print-nvidia-rtx-a6000-datasheet-us-nvidia-1454980-r9-web%20(1).pdf |
| NVIDIA RTX 6000 Ada Generation | 48 | 960 | 364 | 728.5 | n/a | PCIe 4.0 x16, no NVLink | https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/rtx-6000/proviz-print-rtx6000-datasheet-web-2504660.pdf |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 96 | 1597 | n/a | n/a | n/a | PCIe Gen5 x16, no NVLink | https://www.nvidia.com/en-us/data-center/rtx-pro-6000-blackwell-server-edition/ |
| NVIDIA L4 | 24 | 300 | 121 | 242 | n/a | PCIe Gen4 x16 64 GB/s, no NVLink listed | https://www.nvidia.com/en-us/data-center/l4/ |
| NVIDIA A10 | 24 | 600 | 125 | n/a | n/a | PCIe Gen4 64 GB/s, no NVLink listed | https://www.nvidia.com/en-us/data-center/products/a10-gpu/ |
| NVIDIA A40 | 48 | 696 | 149.7 | n/a | n/a | NVLink bridge (2-way) 112.5 GB/s bidirectional; PCIe Gen4 | https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a40/proviz-print-nvidia-a40-datasheet-us-nvidia-1469711-r8-web.pdf |
| NVIDIA L40S | 48 | 864 | 362.05 | 733 | n/a | PCIe Gen4 x16 64 GB/s, no NVLink | https://www.nvidia.com/en-us/data-center/l40s/ |
| NVIDIA A100 40GB SXM | 40 | 1555 | 312 | n/a | n/a | NVLink 600 GB/s; PCIe Gen4 64 GB/s | https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/nvidia-a100-datasheet-us-nvidia-1758950-r4-web.pdf |
| NVIDIA A100 80GB PCIe | 80 | 1935 | 312 | n/a | n/a | NVLink bridge for 2 GPUs 600 GB/s; PCIe Gen4 64 GB/s | https://www.nvidia.com/en-us/data-center/a100/ |
| NVIDIA A100 80GB SXM | 80 | 2039 | 312 | n/a | n/a | NVLink 600 GB/s; PCIe Gen4 64 GB/s | https://www.nvidia.com/en-us/data-center/a100/ |
| NVIDIA H100 PCIe | 80 | 2000 | 756 | 1513 | n/a | NVLink bridge (2 GPUs) 600 GB/s; PCIe Gen5 x16 | https://dam-cdn.nvd.orangelogic.com/AssetLink/705n6ur546g0uk43w0117r17n8042d73.pdf |
| NVIDIA H100 NVL | 94 | 3900 | 835.5 | 1670.5 | n/a | NVLink bridge 600 GB/s; PCIe Gen5 128 GB/s | https://www.nvidia.com/en-us/data-center/h100/ |
| NVIDIA H100 SXM | 80 | 3350 | 989.4 | 1978.9 | n/a | NVLink 900 GB/s; PCIe Gen5 128 GB/s | https://www.nvidia.com/en-us/data-center/h100/ |
| NVIDIA H200 NVL | 141 | 4800 | 835.5 | 1670.5 | n/a | 2- or 4-way NVLink bridge 900 GB/s per GPU; PCIe Gen5 128 GB/s | https://www.nvidia.com/en-us/data-center/h200/ |
| NVIDIA H200 SXM | 141 | 4800 | 989.5 | 1979 | n/a | NVLink 900 GB/s; PCIe Gen5 128 GB/s | https://www.nvidia.com/en-us/data-center/h200/ |
| NVIDIA B200 (HGX B200) | 180 | 8000 | 2250 | 4500 | 9000 | NVLink 5 1.8 TB/s GPU-to-GPU (NVLink Switch) | https://www.nvidia.com/en-us/data-center/hgx/ |
| NVIDIA B300 (HGX B300, Blackwell Ultra) | 270 | 7700 | 2250 | 4500 | 14000 | NVLink 5 1.8 TB/s; PCIe Gen6 256 GB/s | https://dam-cdn.nvd.orangelogic.com/AssetLink/1k0p832eq8r5ca0u5383ie5o4tp3bst1.pdf |
| AMD Instinct MI300X | 192 | 5300 | 1307.4 | 2614.9 | n/a | AMD Infinity Fabric: 8 links, 128 GB/s peak link bandwidth; PCIe 5.0 x16 | https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html |
| AMD Instinct MI325X | 256 | 6000 | 1307.4 | 2614.9 | n/a | AMD Infinity Fabric: 8 links, 128 GB/s peak link bandwidth; PCIe 5.0 x16 | https://www.amd.com/en/products/accelerators/instinct/mi300/mi325x.html |
| AMD Instinct MI355X | 288 | 8000 | 2516.6 | 5033.2 | 10066.3 | AMD Infinity Fabric: 7 x 153.6 GB/s scale-up links; PCIe Gen5 x16 128 GB/s | https://www.amd.com/en/products/accelerators/instinct/mi350/mi355x.html |

Raw data: https://gpucostlab.com/tools/llm-vram-calculator/models.json and https://gpucostlab.com/tools/llm-vram-calculator/gpus.json

## The formula

Serving memory has three parts: the weights, the KV cache, and a runtime allowance. Only the KV cache depends on your traffic.

```
weights            = parameters x bytes per parameter
                     (quantized formats keep embeddings and the LM head at 16-bit)
KV cache per token = 2 x layers x KV heads x head dimension x bytes per value
KV cache in use    = KV per token x tokens per sequence x concurrent sequences
fits when          weights / TP + allowance + KV cache per GPU <= gpu_memory_utilization x GPU memory
```

The 2 counts keys and values. TP is the tensor-parallel size, the number of GPUs one copy of the model is split across.

**Worked example.** Llama 3.1 8B has 32 layers, 8 KV heads and a head dimension of 128, so at BF16 the cache is 2 x 32 x 8 x 128 x 2 bytes = 131,072 bytes (128 KiB) per token. A sequence at 8,192 tokens holds exactly 1 GiB. The weights are 8,030,261,248 parameters x 2 bytes = 14.96 GiB. An RTX 4090 has 24 GB; vLLM's default `gpu_memory_utilization` of 0.92 makes 22.08 GiB usable. Take away the weights and a 2 GiB allowance and 5.1 GiB is left, which is room for five concurrent 8k sequences. That is arithmetic from the formula, not a measurement.

## KV heads, not attention heads

The cache stores keys and values for KV heads, and most current models have far fewer of those than query heads. Llama 3.1 8B has 32 query heads sharing 8 KV heads (grouped-query attention), which makes its cache four times smaller than if every head kept its own keys and values. Llama 3.3 70B also has only 8 KV heads, across 80 layers: 320 KiB per token. Qwen3 32B has 8 KV heads across 64 layers: 256 KiB per token. Calculators that use `num_attention_heads` overstate the cache by the grouping factor.

## What changes the answer

- **Paged attention.** vLLM allocates the cache in blocks (16 tokens each by default) as sequences grow, so memory follows the tokens actually in use, not the maximum length you allow. That is why the calculator multiplies by your real context length. vLLM still refuses to start if a single sequence at `max_model_len` cannot fit.
- **Prefix caching.** Sequences that share a prompt prefix, such as a long system prompt, share its cache blocks. The calculator assumes no sharing, so for chat traffic with a common prefix it is an upper bound.
- **KV-cache quantization.** An FP8 cache (`--kv-cache-dtype fp8` in vLLM) halves cache memory. The effect on output quality depends on the model and the task. Long-context retrieval is a reasonable place to look first, but measure it on your own evaluation before you rely on it.
- **Weight quantization is not "parameters x bits".** Quantized checkpoints usually keep the embedding table and LM head at 16-bit, and group-wise formats store scales. For Llama 3.1 8B, the embeddings and LM head are 1.05 billion of the 8.03 billion parameters (128,256-token vocabulary x 4,096 hidden x 2). INT4 at group size 128 costs 4.16 bits per weight once the scale and zero point are counted. So the 4-bit model is 5.33 GiB, not 4 GB.
- **Mixture of experts.** Every expert has to be in memory, so a MoE model's memory follows its total parameters. Its decode speed follows its active parameters. Qwen3 30B-A3B needs 61 GB (56.9 GiB) at BF16 but reads only about 3.3 billion parameters per token.
- **Sliding-window layers.** gpt-oss alternates full-attention layers with 128-token sliding-window layers, and Gemma 4 runs 1,024-token windows on five of every six layers. A sliding-window layer only ever caches its window. Engines with a hybrid KV-cache manager (vLLM V1) allocate it that way, and so does the calculator. At 128k tokens, Gemma 4 31B needs 10.8 GiB per sequence. An engine that cached the full context on every layer would need 110 GiB.
- **Linear attention.** Qwen3.5, 3.6 and 3.8 use Gated DeltaNet on three of every four layers. Those layers keep a fixed recurrent state per sequence instead of a growing cache. For Qwen3.8 27B, that state is 147 MiB per sequence, stored in FP32 as its config specifies. The full-attention layers add 64 KiB per token. At short context the state dominates; at long context the cache does.
- **Multi-head latent attention (MLA).** DeepSeek-V3.2, Kimi K2 and GLM-5.3 cache a single 576-value latent per token per layer, shared by all heads. That makes the cache small: 70,272 bytes per token for Kimi K2 at BF16, against 327,680 for Llama 3.3 70B. DeepSeek-V3.2 also caches a 132-byte FP8 key per layer for its sparse-attention indexer.

## Tensor parallelism does not always split the cache

With tensor parallelism, vLLM splits KV heads across GPUs, but a GPU never holds less than one head. When the tensor-parallel size reaches the number of KV heads, heads are replicated. Qwen3.5 122B-A10B has only 2 KV heads on its full-attention layers, so going from 2 to 8 GPUs barely shrinks the per-GPU cache (457 MiB to 402 MiB per 32k-token sequence; the part that still shrinks is the linear-attention state).

MLA is the extreme case. The latent is shared by all heads, so every tensor-parallel GPU keeps the full cache. At 32k tokens, a DeepSeek-V3.2 sequence takes 2.4 GiB on each of 8 GPUs. A Llama 3.3 70B sequence of the same length takes 10 GiB in total, which splits to 1.25 GiB per GPU. This is why large MLA deployments use data-parallel attention. The calculator models tensor parallelism only, and says which GPU counts a model's heads allow.

## The runtime allowance

When vLLM starts, it loads the weights. Then it runs a profiling forward pass to measure peak activation memory and accounts for memory outside PyTorch's allocator (the CUDA context, NCCL buffers) and for CUDA graphs. Whatever is left of `gpu_memory_utilization` x GPU memory becomes KV cache. The calculator stands in for the middle part with one number per GPU. The default of 2 GiB is a stated round number, not a measurement. vLLM logs the real figures at startup, so replace it with yours.

GPU memory is the marketed figure treated as GiB. The capacity the driver reports can be a few percent lower, for example with ECC enabled, so leave headroom if you are within a few percent of the limit.

## The decode ceiling column

Generating one token means reading every active weight and the sequence's whole cache from GPU memory. Memory bandwidth divided by those bytes is therefore an upper bound on single-stream decode speed. For Llama 3.1 8B at BF16 on an H100 SXM (3,350 GB/s) at 8k context, that is 3,350 GB/s / 17.1 GB = 196 tokens per second. Real engines land below it. Batching reads the weights once for many sequences, so total throughput goes far above one stream. Use the ceiling to compare GPUs, and the [break-even calculator](/self-host-vs-api-cost/) with your own measured throughput to compare costs. Measured numbers on this site are [planned](/methodology/), not published.

## What the calculator does not model

- Pipeline and expert parallelism, and data-parallel attention.
- Speculative decoding. A draft model, or a model's multi-token-prediction layer, adds weights and its own cache.
- LoRA adapters, CPU offload, and GGUF k-quant formats (which mix bit widths per tensor).
- Activations from image or audio inputs. The weights of the multimodal presets do include the vision encoder.

## Where to run it

If the model fits a 24 to 32 GB card, hourly rentals of RTX 4090 and RTX 5090 cards are the cheapest way to test it: [RunPod](https://www.runpod.io/) lists both, and [Vast.ai](https://vast.ai/) is a marketplace where individual hosts set the price. For 70B-class models and long contexts you need 80 to 180 GB data-center GPUs from providers such as [Lambda](https://lambda.ai/) or [DigitalOcean](https://www.digitalocean.com/). The [GPU price table](/gpu-cloud-prices/) compares on-demand prices from eleven GPU clouds and from AWS, Google Cloud and Azure, most of which pay this site nothing.

Before you rent anything, run the [break-even calculator](/self-host-vs-api-cost/). For many open models, an API serving the same model costs less than one idle GPU.
