LLMBOTTLENECK.COM
Qwen3-32B at Q4_K_M, 8,192 tokens

NVIDIA GeForce RTX 4090
vs NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090 decodes 1.78× faster here.

Side by side · Qwen3-32B at Q4_K_M, 8,192 tokens

PropertyNVIDIA GeForce RTX 4090NVIDIA GeForce RTX 5090
Runsfits in device memoryfits in device memory
Decode~32 tok/sFaster than you read~57 tok/sFaster than you read
Memory needed22.7 GB of 24.0 GB22.7 GB of 32.0 GB
Spare or short1.29 GB spare9.29 GB spare
Weights19.8 GB19.8 GB
KV cache2.15 GB2.15 GB
Memory bandwidth1008 GB/s1792 GB/s

NVIDIA GeForce RTX 4090

CAPACITY LIMITEDVerdictunvalidated

It fits, but only 1.29 GB remains; one more current-size context needs 2.15 GB.

Open in the calculator →

NVIDIA GeForce RTX 5090

NOTHING TO FIXVerdictestimate

9.29 GB spare and ~57 tok/s — above the 30 tok/s this site treats as faster than reading.

Open in the calculator →

Read this carefully

Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.

Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.

llmbottleneck
catalogue 2026-10-03models 327devices 135