# LLM fine-tuning cost estimator: LoRA, QLoRA or full fine-tune, memory and GPU-hours

> Memory per GPU for full fine-tuning, LoRA and QLoRA, whether it fits on the GPUs you pick, the GPU-hours your tokens and epochs take at a stated throughput assumption, and the cost at on-demand list prices, including the cheapest listing for that GPU. Formulas and sources are on the page.

By GPUCostLab. Updated 7 October 2026. Canonical URL: https://gpucostlab.com/llm-fine-tuning-cost/

No paid links on this page: vendor links go straight to the vendor.

> **Interactive estimator.** Pick a model, a method (full fine-tune, LoRA or QLoRA), dataset tokens, epochs, sequence length, GPU type and count, and a price on the HTML version of this page (https://gpucostlab.com/llm-fine-tuning-cost/). It returns memory per GPU by component and whether it fits, GPU-hours, wall-clock time, and the cost at your price and at the cheapest listed price for that GPU. Throughput is an assumption (default: 30% MFU of the GPU's peak dense BF16 throughput), never a measurement.

### Formulas

- Full fine-tune memory: 2 bytes weights + 2 bytes gradients + 12 bytes AdamW state (FP32 master copy and two moments) per parameter, 16 in total (ZeRO, https://arxiv.org/abs/1910.02054). 8 bytes of optimizer state without a master copy, 2 bytes with 8-bit AdamW (https://arxiv.org/abs/2110.02861).
- LoRA: frozen weights at 2 bytes per parameter; adapters at 4 + 4 + 8 bytes per trainable parameter (FP32 weights, gradients, AdamW moments). Trainable parameters = rank x (4 x hidden + 2 x heads x head dim + 2 x KV heads x head dim) per layer, for q, k, v and o.
- QLoRA: frozen weights at 4.127 bits per parameter (NF4 plus double-quantized constants, https://arxiv.org/abs/2305.14314), embeddings and LM head at 16-bit; adapters as LoRA.
- Activations (https://arxiv.org/abs/2205.05198): 34 x sequence x micro-batch x hidden bytes per layer, or 2 x sequence x micro-batch x hidden per layer plus one recomputed layer with gradient checkpointing; plus 4 x sequence x micro-batch x vocabulary bytes of FP32 logits. Sharding (FSDP/ZeRO-3) divides weights, gradients and optimizer state by the GPU count.
- Compute (PaLM, https://arxiv.org/abs/2204.02311, appendix B): training FLOPs per token = 6 x parameters used per token + 12 x layers x heads x head dim x sequence length; 4 x parameters for LoRA and QLoRA, which skip weight gradients for the frozen weights.
- Tokens per second per GPU = MFU x peak dense BF16 FLOP/s / FLOPs per token. GPU-hours = dataset tokens x epochs / tokens per second per GPU / 3,600. Cost = GPU-hours x price per GPU-hour.

### GPU compute and the lowest listed on-demand price (prices read 2026-10-07)

| GPU | Memory (GB) | Peak BF16 TFLOPS (dense) | Lowest $/GPU-hour | Where | $ per exaFLOP at 30% MFU |
|---|---|---|---|---|---|
| NVIDIA GeForce RTX 4090 | 24 | 165.2 | 0.34 | RunPod Community Cloud | 1.91 |
| NVIDIA GeForce RTX 5090 | 32 | 209.5 | 0.69 | RunPod Community Cloud | 3.05 |
| NVIDIA RTX A6000 | 48 | 154.8 | 0.33 | RunPod Community Cloud | 1.97 |
| NVIDIA RTX 6000 Ada Generation | 48 | 364 | 0.74 | RunPod Community Cloud | 1.88 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 96 | Not published | 1.69 | RunPod Community Cloud | n/a |
| NVIDIA L4 | 24 | 121 | 0.44 | RunPod Community Cloud | 3.37 |
| NVIDIA A10 | 24 | 125 | 1.1016 | Modal | 8.16 |
| NVIDIA A40 | 48 | 149.7 | 0.35 | RunPod Community Cloud | 2.16 |
| NVIDIA L40S | 48 | 362.05 | 0.79 | RunPod Community Cloud | 2.02 |
| NVIDIA A100 40GB SXM | 40 | 312 | 1.99 | Lambda | 5.91 |
| NVIDIA A100 80GB PCIe | 80 | 312 | 1.19 | RunPod Community Cloud | 3.53 |
| NVIDIA A100 80GB SXM | 80 | 312 | 1.39 | RunPod Community Cloud | 4.13 |
| NVIDIA H100 PCIe | 80 | 756 | 1.99 | RunPod Community Cloud | 2.44 |
| NVIDIA H100 NVL | 94 | 835.5 | 2.59 | RunPod Community Cloud | 2.87 |
| NVIDIA H100 SXM | 80 | 989.4 | starting at 1.99 | Voltage Park | 1.86 |
| NVIDIA H200 NVL | 141 | 835.5 | Not listed | | n/a |
| NVIDIA H200 SXM | 141 | 989.5 | 3.99 | Hyperstack | 3.73 |
| NVIDIA B200 (HGX B200) | 180 | 2250 | 5.98 | RunPod Community Cloud | 2.46 |
| NVIDIA B300 (HGX B300, Blackwell Ultra) | 270 | 2250 | 6.94 | RunPod Community Cloud | 2.86 |
| AMD Instinct MI300X | 192 | 1307.4 | 2.59 | DigitalOcean GPU Droplets | 1.83 |
| AMD Instinct MI325X | 256 | 1307.4 | 3.8 | DigitalOcean GPU Droplets | 2.69 |
| AMD Instinct MI355X | 288 | 2516.6 | Not listed | | n/a |

Price data: https://gpucostlab.com/tools/fine-tuning-cost-estimator/prices.json

## How the estimate is calculated

```
Memory per GPU   = (weights + gradients + optimizer state) / GPUs, when sharded with FSDP or ZeRO-3
                 + activations + logits + framework overhead           (per GPU, never sharded)

Full fine-tune   = 2 + 2 + 12 bytes per parameter                     (BF16 weights and gradients, AdamW
                                                                        with an FP32 master copy and two moments)
LoRA             = 2 bytes per frozen parameter + 16 bytes per adapter parameter
QLoRA            = 4.127 bits per frozen parameter (NF4 + quantization constants) + 16 bytes per adapter parameter
Activations      = 34 x sequence x micro-batch x hidden bytes per layer
                   or, with gradient checkpointing, 2 x sequence x micro-batch x hidden per layer + one layer's 34
Logits           = 4 bytes x sequence x micro-batch x vocabulary

FLOPs per token  = 6 x parameters used per token + 12 x layers x heads x head dim x sequence length
                   (4 x parameters for LoRA and QLoRA: no weight gradients for the frozen weights)
Tokens/s per GPU = MFU x peak dense BF16 FLOP/s / FLOPs per token    (or your measured figure)
GPU-hours        = dataset tokens x epochs / tokens/s per GPU / 3,600
Cost             = GPU-hours x price per GPU-hour
```

Every constant comes from a paper you can check:

- **16 bytes per parameter for full fine-tuning.** Mixed-precision Adam keeps 16-bit weights and gradients (2 + 2 bytes), plus an FP32 master copy of the weights and two FP32 moments (4 + 4 + 4 = 12 bytes). That accounting is from the [ZeRO paper](https://arxiv.org/abs/1910.02054) (Rajbhandari et al., 2020), which writes it as 2Ψ + 2Ψ + KΨ with K = 12. If your optimizer keeps no FP32 master copy, the state is 8 bytes. With 8-bit AdamW it is 2 bytes ([Dettmers et al., 2022](https://arxiv.org/abs/2110.02861)).
- **LoRA adapters.** [LoRA](https://arxiv.org/abs/2106.09685) trains two small matrices of rank r beside each adapted weight, r x (inputs + outputs) parameters per matrix. The estimator puts adapters on the four attention projections (q, k, v and o) of every layer. At rank 16 that is 13.6 million parameters for Llama 3.1 8B, 0.17% of the model. Adapters are kept in FP32, which is PEFT's default, so each costs 4 bytes of weight, 4 of gradient and 8 of AdamW moments. Adding the MLP projections roughly triples the adapter count, which is still small next to the frozen weights. If you know your count, enter it.
- **QLoRA's 4.127 bits.** The [QLoRA paper](https://arxiv.org/abs/2305.14314) stores frozen weights as 4-bit NormalFloat in blocks of 64 and quantizes the block constants again ("double quantization"), which brings them from 0.5 to 0.127 bits per parameter. The usual Transformers and bitsandbytes setup leaves the embeddings and the LM head unquantized, so they stay at 2 bytes, the same treatment as in the [VRAM calculator](/llm-vram-calculator/).
- **Activations.** [Korthikanti et al. (2022)](https://arxiv.org/abs/2205.05198) count 34 x s x b x h bytes per transformer layer for 16-bit activations, plus 5 x a x s² x b for the attention scores. FlashAttention-style kernels do not store the attention-score matrices, so that second term is dropped. Checkpointing every layer ("full activation recomputation") keeps only each layer's 2 x s x b x h input and recomputes the rest, one layer at a time. Their constant assumes a GPT-style layer with a 4x MLP; models with wider MLPs or many experts store somewhat more.
- **FLOPs per token.** From the [PaLM paper's appendix B](https://arxiv.org/abs/2204.02311) (Chowdhery et al., 2022): 2N for the forward pass, 4N for the backward pass, plus 12 x L x H x Q x T for the attention matmuls. The backward pass does two matmuls for each forward one, one for the input gradients and one for the weight gradients. LoRA and QLoRA skip the weight gradients of the frozen weights, so the estimator uses 4N for them. For mixture-of-experts models, N is the parameters active per token, while memory holds all of them.

## Throughput is the assumption that decides the cost

Memory is arithmetic. Throughput is not, and nobody can look it up for your job. It depends on the GPU, the model, sequence length, micro-batch, kernels, the trainer, and how many GPUs talk to each other over what link. The estimator expresses it as **model FLOPs utilization (MFU)**: the share of the GPU's peak dense BF16 throughput that goes into the model's own FLOPs.

The default is 30%, and it is a rule of thumb, not a measurement. For scale: Google reported 46.2% MFU for PaLM 540B, and by the same accounting Megatron-Turing NLG 530B reached 29.7% ([PaLM, section 4.1 and appendix B](https://arxiv.org/abs/2204.02311)). Meta reported 38-43% BF16 MFU for Llama 3 405B pre-training on up to 16,384 H100s ([Llama 3 paper, table 4](https://arxiv.org/abs/2407.21783)). Those are pre-training runs on heavily tuned stacks. A fine-tuning job on one to eight GPUs with off-the-shelf tooling, short sequences and a micro-batch of one usually gets less. Gradient checkpointing costs time too: Korthikanti et al. measured 30-40% extra execution time for full recomputation, which under the MFU definition shows up as a lower MFU.

To replace the assumption, run a few hundred steps of your real job and read the tokens per second from your trainer's log (or samples per second times sequence length). Divide by the number of GPUs and enter it as measured throughput. The result then says "your measurement" instead of "assumed". The sensitivity table shows the cost if the real figure is half or double.

Measured fine-tuning throughput on the GPUs in the price table is [planned](/methodology/#benchmarks-planned-no-results-yet). None is published yet.

## A worked example

These are the estimator's defaults.

- **Job:** Llama 3.1 8B, LoRA at rank 16 on q, k, v and o (13.6 million trainable parameters), 50 million tokens for 3 epochs (150 million training tokens) at 4,096 tokens per sequence, one sequence per step, gradient checkpointing on, AdamW.
- **Memory on one H100 SXM:** frozen BF16 weights 14.96 GiB, adapters with their gradients and optimizer state 0.20 GiB, stored activations 1.00 GiB, one recomputed layer 0.53 GiB, FP32 logits 1.96 GiB (a 128,256-token vocabulary is expensive at the loss), and a 2 GiB overhead allowance: 20.6 GiB of 80. It fits with room to raise the micro-batch.
- **Compute:** 4 x 8.03 billion = 32.1 GFLOP per token for the weights, plus 12 x 32 layers x 32 heads x 128 x 4,096 = 6.4 GFLOP for attention, 38.6 GFLOP in all. At 30% of the H100 SXM's 989.4 dense BF16 TFLOPS, that is about 7,700 tokens per second, so 150 million tokens take **5.4 GPU-hours**.
- **Cost:** at the lowest H100 SXM listing in the price table, Voltage Park's $1.99 per GPU-hour (a "starting at" price), about $11; at Google Cloud's on-demand a3-highgpu-8g rate of $11.06 per GPU, about $60 (prices read 7 October 2026).

The same job as a **full fine-tune** needs 16 bytes for each of 8.03 billion parameters: 125 GiB per GPU on one H100, which does not fit. Sharded across two or more H100s it does (20.4 GiB each across eight). Its compute is 54.6 GFLOP per token, 41% more than LoRA, so 7.7 GPU-hours at the same MFU. For comparison, the LoRA paper reports 43.1 tokens per second per V100 for LoRA against 32.5 for full fine-tuning on GPT-3 175B, which it calls a 25% speedup. That is a smaller gain than the FLOP counts alone predict (about 1.4 to 1.5 times), one more reason to measure your own.

**Llama 3.3 70B with QLoRA** fits on one 80 GB GPU at about 48 GiB, with 37 GiB of that the 4-bit weights. It takes about 44 GPU-hours for the same 150 million tokens at 30% MFU, before the dequantization overhead that QLoRA adds.

All of this is arithmetic from the stated assumptions, not a measurement.

## LoRA, QLoRA or full fine-tuning: what each costs you

- **Full fine-tuning** changes every weight. Memory is 16 bytes per parameter before activations, so anything past a few billion parameters needs several GPUs with sharded optimizer state, and compute is the full 6N per token.
- **LoRA** freezes the model and trains adapters. Memory drops to the 16-bit weights plus a little, and compute drops by about a third. It is the usual choice when the model fits in 16-bit on the GPUs you have.
- **QLoRA** also stores the frozen weights in 4 bits, which is what puts a 70B model on one 80 GB GPU or an 8B model on a 24 GB card. The weights are dequantized for every matrix multiply, so it is slower than LoRA at the same FLOP count; the FLOP formula does not include that.

Whether LoRA matches full fine-tuning on quality depends on the task and the data, and is not something a cost estimator can tell you. Measure it on your own evaluation set.

## Hyperscaler or GPU cloud

The price list includes AWS, Google Cloud and Azure on-demand instance prices alongside the GPU clouds. Per GPU-hour, hyperscaler on-demand prices are usually several times higher, but the comparison is not like for like. A hyperscaler instance includes its CPUs, memory, local storage and fast networking, and most of that capacity is bought with reservations or committed-use discounts that list prices do not show. Some of these instances come only as 8-GPU machines; the estimator says when that means paying for GPUs your job does not use. The [GPU price table](/gpu-cloud-prices/) puts them side by side.

## What is not included

- Failed and repeated runs, hyperparameter sweeps and evaluation. In practice these can cost more than the final run.
- Data preparation, storage, egress and checkpoint saving.
- Padding. If you do not pack sequences, padding tokens cost the same compute as real ones. Count them in the dataset tokens.
- Per-expert activation memory in mixture-of-experts models, attention FLOPs for latent-attention (MLA) and linear-attention layers, and memory spikes from optimizer steps or long-sequence outliers. The estimator says when a model is affected.
- Communication time across GPUs. The estimate assumes MFU holds as you add GPUs. Over PCIe without NVLink it usually does not.

## Renting the GPUs

For a short measurement run to replace the throughput assumption, by-the-hour providers are the cheapest way in: [RunPod](https://www.runpod.io/) and the [Vast.ai](https://vast.ai/) marketplace for single cards, and [Lambda](https://lambda.ai/) and [DigitalOcean](https://www.digitalocean.com/) for H100, H200 and B200 instances. The estimator's price list is the same data as the [GPU price table](/gpu-cloud-prices/), including the providers that pay this site nothing, and it is always sorted by price.
