Qwen3-32B at Q4_K_M, 8,192 tokens
AMD Radeon™ RX 7900 XTX
vs NVIDIA GeForce RTX 4090
AMD Radeon™ RX 7900 XTX decodes 1.09× faster here.
Side by side · Qwen3-32B at Q4_K_M, 8,192 tokens
| Property | AMD Radeon™ RX 7900 XTX | NVIDIA GeForce RTX 4090 |
|---|---|---|
| Runs | fits in device memory | fits in device memory |
| Decode | ~35 tok/sFaster than you read | ~32 tok/sFaster than you read |
| Memory needed | 22.7 GB of 24.0 GB | 22.7 GB of 24.0 GB |
| Spare or short | 1.29 GB spare | 1.29 GB spare |
| Weights | 19.8 GB | 19.8 GB |
| KV cache | 2.15 GB | 2.15 GB |
| Memory bandwidth | 960 GB/s | 1008 GB/s |
AMD Radeon™ RX 7900 XTX
CAPACITY LIMITEDVerdictunvalidated
It fits, but only 1.29 GB remains; one more current-size context needs 2.15 GB.
Open in the calculator →NVIDIA GeForce RTX 4090
CAPACITY LIMITEDVerdictunvalidated
It fits, but only 1.29 GB remains; one more current-size context needs 2.15 GB.
Open in the calculator →Read this carefully
Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.
Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.