Fine-tuning cost estimator

Runs in your browser. Model shapes from each model's config.json, GPU peak throughput from vendor datasheets, and on-demand prices read 7 October 2026. Throughput is an assumption you set, not a measurement.

Model and method
Training data
GPUs and price

Throughput (an assumption)

Throughput decides the cost, and nobody can look it up for your setup. The default derives it from the GPU's peak BF16 throughput at 30% MFU, a rule of thumb, not a measurement. Run a few hundred steps and enter the tokens per second your trainer logs.

GPU compute and what it costs at list price

Peak dense BF16 throughput from each vendor's datasheet (7 October 2026) and the lowest on-demand price listed for that exact GPU in the price table (7 October 2026). The last two columns are arithmetic, not measurements: the price of 1018 FLOPs (one exaFLOP) at 100% of peak, and at 30% MFU, the estimator's default. A training job needs about 6 x parameters x tokens FLOPs for a full fine-tune and 4 x parameters x tokens for LoRA, plus attention.

GPUMemoryPeak BF16 TFLOPS (dense)Lowest $/GPU-hourWhere$ per exaFLOP at peak$ per exaFLOP at 30% MFU
NVIDIA GeForce RTX 4090 24 GB 165.2 $0.340 RunPod (Community Cloud) $0.57 $1.91
NVIDIA GeForce RTX 5090 32 GB 209.5 $0.690 RunPod (Community Cloud) $0.91 $3.05
NVIDIA RTX A6000 48 GB 154.8 $0.330 RunPod (Community Cloud) $0.59 $1.97
NVIDIA RTX 6000 Ada Generation 48 GB 364 $0.740 RunPod (Community Cloud) $0.56 $1.88
NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB Not published $1.69 RunPod (Community Cloud) n/an/a
NVIDIA L4 24 GB 121 $0.440 RunPod (Community Cloud) $1.01 $3.37
NVIDIA A10 24 GB 125 $1.10 Modal $2.45 $8.16
NVIDIA A40 48 GB 149.7 $0.350 RunPod (Community Cloud) $0.65 $2.16
NVIDIA L40S 48 GB 362.05 $0.790 RunPod (Community Cloud) $0.61 $2.02
NVIDIA A100 40GB SXM 40 GB 312 $1.99 Lambda $1.77 $5.91
NVIDIA A100 80GB PCIe 80 GB 312 $1.19 RunPod (Community Cloud) $1.06 $3.53
NVIDIA A100 80GB SXM 80 GB 312 $1.39 RunPod (Community Cloud) $1.24 $4.13
NVIDIA H100 PCIe 80 GB 756 $1.99 RunPod (Community Cloud) $0.73 $2.44
NVIDIA H100 NVL 94 GB 835.5 $2.59 RunPod (Community Cloud) $0.86 $2.87
NVIDIA H100 SXM 80 GB 989.4 starting at $1.99 Voltage Park $0.56 $1.86
NVIDIA H200 NVL 141 GB 835.5 Not listedn/an/a
NVIDIA H200 SXM 141 GB 989.5 $3.99 Hyperstack $1.12 $3.73
NVIDIA B200 (HGX B200) 180 GB 2250 $5.98 RunPod (Community Cloud) $0.74 $2.46
NVIDIA B300 (HGX B300, Blackwell Ultra) 270 GB 2250 $6.94 RunPod (Community Cloud) $0.86 $2.86
AMD Instinct MI300X 192 GB 1307.4 $2.59 DigitalOcean GPU Droplets $0.55 $1.83
AMD Instinct MI325X 256 GB 1307.4 $3.80 DigitalOcean GPU Droplets $0.81 $2.69
AMD Instinct MI355X 288 GB 2516.6 Not listedn/an/a

Memory per parameter, by method

MethodWeightsGradientsOptimizer state (AdamW)Source
Full fine-tune, mixed precision2 bytes (BF16)2 bytes12 bytes: FP32 master copy + two FP32 momentsZeRO, section 3.1
Full fine-tune, no master copy2 bytes2 bytes8 bytes: two FP32 moments8-bit optimizers, section 1.1
Full fine-tune, 8-bit AdamW2 bytes2 bytes2 bytes: two 8-bit moments8-bit optimizers
LoRA2 bytes frozen; 4 bytes per adapter parameter (FP32)4 bytes per adapter parameter8 bytes per adapter parameterLoRA
QLoRA0.516 bytes frozen (NF4 + double-quantized constants); embeddings and LM head 2 bytes4 bytes per adapter parameter8 bytes per adapter parameterQLoRA, section 3

How the estimate is calculated

Memory per GPU   = (weights + gradients + optimizer state) / GPUs, when sharded with FSDP or ZeRO-3
                 + activations + logits + framework overhead           (per GPU, never sharded)

Full fine-tune   = 2 + 2 + 12 bytes per parameter                     (BF16 weights and gradients, AdamW
                                                                        with an FP32 master copy and two moments)
LoRA             = 2 bytes per frozen parameter + 16 bytes per adapter parameter
QLoRA            = 4.127 bits per frozen parameter (NF4 + quantization constants) + 16 bytes per adapter parameter
Activations      = 34 x sequence x micro-batch x hidden bytes per layer
                   or, with gradient checkpointing, 2 x sequence x micro-batch x hidden per layer + one layer's 34
Logits           = 4 bytes x sequence x micro-batch x vocabulary

FLOPs per token  = 6 x parameters used per token + 12 x layers x heads x head dim x sequence length
                   (4 x parameters for LoRA and QLoRA: no weight gradients for the frozen weights)
Tokens/s per GPU = MFU x peak dense BF16 FLOP/s / FLOPs per token    (or your measured figure)
GPU-hours        = dataset tokens x epochs / tokens/s per GPU / 3,600
Cost             = GPU-hours x price per GPU-hour

Every constant comes from a paper you can check:

Throughput is the assumption that decides the cost

Memory is arithmetic. Throughput is not, and nobody can look it up for your job. It depends on the GPU, the model, sequence length, micro-batch, kernels, the trainer, and how many GPUs talk to each other over what link. The estimator expresses it as model FLOPs utilization (MFU): the share of the GPU's peak dense BF16 throughput that goes into the model's own FLOPs.

The default is 30%, and it is a rule of thumb, not a measurement. For scale: Google reported 46.2% MFU for PaLM 540B, and by the same accounting Megatron-Turing NLG 530B reached 29.7% (PaLM, section 4.1 and appendix B). Meta reported 38-43% BF16 MFU for Llama 3 405B pre-training on up to 16,384 H100s (Llama 3 paper, table 4). Those are pre-training runs on heavily tuned stacks. A fine-tuning job on one to eight GPUs with off-the-shelf tooling, short sequences and a micro-batch of one usually gets less. Gradient checkpointing costs time too: Korthikanti et al. measured 30-40% extra execution time for full recomputation, which under the MFU definition shows up as a lower MFU.

To replace the assumption, run a few hundred steps of your real job and read the tokens per second from your trainer's log (or samples per second times sequence length). Divide by the number of GPUs and enter it as measured throughput. The result then says "your measurement" instead of "assumed". The sensitivity table shows the cost if the real figure is half or double.

Measured fine-tuning throughput on the GPUs in the price table is planned. None is published yet.

A worked example

These are the estimator's defaults.

The same job as a full fine-tune needs 16 bytes for each of 8.03 billion parameters: 125 GiB per GPU on one H100, which does not fit. Sharded across two or more H100s it does (20.4 GiB each across eight). Its compute is 54.6 GFLOP per token, 41% more than LoRA, so 7.7 GPU-hours at the same MFU. For comparison, the LoRA paper reports 43.1 tokens per second per V100 for LoRA against 32.5 for full fine-tuning on GPT-3 175B, which it calls a 25% speedup. That is a smaller gain than the FLOP counts alone predict (about 1.4 to 1.5 times), one more reason to measure your own.

Llama 3.3 70B with QLoRA fits on one 80 GB GPU at about 48 GiB, with 37 GiB of that the 4-bit weights. It takes about 44 GPU-hours for the same 150 million tokens at 30% MFU, before the dequantization overhead that QLoRA adds.

All of this is arithmetic from the stated assumptions, not a measurement.

LoRA, QLoRA or full fine-tuning: what each costs you

Whether LoRA matches full fine-tuning on quality depends on the task and the data, and is not something a cost estimator can tell you. Measure it on your own evaluation set.

Hyperscaler or GPU cloud

The price list includes AWS, Google Cloud and Azure on-demand instance prices alongside the GPU clouds. Per GPU-hour, hyperscaler on-demand prices are usually several times higher, but the comparison is not like for like. A hyperscaler instance includes its CPUs, memory, local storage and fast networking, and most of that capacity is bought with reservations or committed-use discounts that list prices do not show. Some of these instances come only as 8-GPU machines; the estimator says when that means paying for GPUs your job does not use. The GPU price table puts them side by side.

What is not included

Renting the GPUs

For a short measurement run to replace the throughput assumption, by-the-hour providers are the cheapest way in: RunPod and the Vast.ai marketplace for single cards, and Lambda and DigitalOcean for H100, H200 and B200 instances. The estimator's price list is the same data as the GPU price table, including the providers that pay this site nothing, and it is always sorted by price.

Also available as Markdown.