LLMBOTTLENECK.COM
Qwen3-32B at Q4_K_M, 8,192 tokens

NVIDIA GeForce RTX 3090
vs NVIDIA GeForce RTX 4090

NVIDIA GeForce RTX 4090 decodes 1.08× faster here.

Side by side · Qwen3-32B at Q4_K_M, 8,192 tokens

PropertyNVIDIA GeForce RTX 3090NVIDIA GeForce RTX 4090
Runsfits in device memoryfits in device memory
Decode~30 tok/sAbout reading pace~32 tok/sFaster than you read
Memory needed22.7 GB of 24.0 GB22.7 GB of 24.0 GB
Spare or short1.29 GB spare1.29 GB spare
Weights19.8 GB19.8 GB
KV cache2.15 GB2.15 GB
Memory bandwidth936 GB/s1008 GB/s

NVIDIA GeForce RTX 3090

CAPACITY LIMITEDVerdictunvalidated

It fits, but only 1.29 GB remains; one more current-size context needs 2.15 GB.

Open in the calculator →

NVIDIA GeForce RTX 4090

CAPACITY LIMITEDVerdictunvalidated

It fits, but only 1.29 GB remains; one more current-size context needs 2.15 GB.

Open in the calculator →

Read this carefully

Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.

Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.

llmbottleneck
catalogue 2026-10-03models 327devices 135