What a fine-tune actually costs
Choose a model and a method to see an estimate of training memory, term by term. LoRA and QLoRA are sized only for architectures whose trainable modules are modelled; use it to compare settings, not to choose hardware on its own.
Result
Choose a model and a method, then press Calculate. Weights, gradients and optimizer state are exact arithmetic on a verified parameter count; activation memory is not, and it is labelled separately.
The optimizer figures are the ZeRO paper's own. Mixed-precision Adam keeps an fp32 copy of the parameters, momentum and variance at 4 bytes each, which is why 12 bytes per trainable parameter appears above and not a rounder number. Rajbhandari et al., arXiv:1910.02054 ↗
QLoRA's base weights are held at 4.127 bits per parameter: 4-bit NormalFloat plus the 0.127 bits that double quantization leaves for the quantization constants, which the paper derives as 8/64 + 32/(64 × 256). Dettmers et al., arXiv:2305.14314 ↗
Adapter parameters are counted from the real projection shapes — r × (d_in + d_out) per targeted matrix, as the LoRA paper defines it — rather than as a share of the model, which is the rule of thumb that makes other calculators disagree with each other. Hu et al., arXiv:2106.09685 ↗
There is no measured accuracy for these figures. Weights, adapters, gradients and optimizer state are calculated from the parameter count and the stated dtypes, under the listed assumptions; the activation term is a rough model of what a framework keeps. Framework context, temporary buffers, fused-kernel workspaces and distributed training are not included. Until there is a corpus of real runs to score against, this page publishes the arithmetic and names the assumption instead of quoting an error bar it has not earned.
Serving memory, decode speed and the fit verdict are on the main calculator; the method behind those is on the methodology page.