LLMBOTTLENECK.COM
Qwen3-32B at Q4_K_M, 8,192 tokens

AMD Radeon™ RX 7900 XTX
vs NVIDIA GeForce RTX 4090

AMD Radeon™ RX 7900 XTX decodes 1.09× faster here.

Side by side · Qwen3-32B at Q4_K_M, 8,192 tokens

PropertyAMD Radeon™ RX 7900 XTXNVIDIA GeForce RTX 4090
Runsfits in device memoryfits in device memory
Decode~35 tok/sFaster than you read~32 tok/sFaster than you read
Memory needed22.7 GB of 24.0 GB22.7 GB of 24.0 GB
Spare or short1.29 GB spare1.29 GB spare
Weights19.8 GB19.8 GB
KV cache2.15 GB2.15 GB
Memory bandwidth960 GB/s1008 GB/s

AMD Radeon™ RX 7900 XTX

CAPACITY LIMITEDVerdictunvalidated

It fits, but only 1.29 GB remains; one more current-size context needs 2.15 GB.

Open in the calculator →

NVIDIA GeForce RTX 4090

CAPACITY LIMITEDVerdictunvalidated

It fits, but only 1.29 GB remains; one more current-size context needs 2.15 GB.

Open in the calculator →

Read this carefully

Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.

Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.

llmbottleneck
catalogue 2026-10-03models 327devices 135