On NVIDIA GeForce RTX 4090 at Q4_K_M, 8,192 tokens
Qwen3-8B
vs phi-4
Qwen3-8B decodes 1.75× faster here.
Side by side · On NVIDIA GeForce RTX 4090 at Q4_K_M, 8,192 tokens
| Property | Qwen3-8B | phi-4 |
|---|---|---|
| Runs | fits in device memory | fits in device memory |
| Decode | ~116 tok/sFaster than you read | ~67 tok/sFaster than you read |
| Memory needed | 7.04 GB of 24.0 GB | 11.4 GB of 24.0 GB |
| Spare or short | 17.0 GB spare | 12.6 GB spare |
| Weights | 5.03 GB | 8.89 GB |
| KV cache | 1.21 GB | 1.68 GB |
| Memory bandwidth | 1008 GB/s | 1008 GB/s |
Qwen3-8B
NOTHING TO FIXVerdictestimate
17.0 GB spare and ~116 tok/s — above the 30 tok/s this site treats as faster than reading.
Open in the calculator →phi-4
NOTHING TO FIXVerdictestimate
12.6 GB spare and ~67 tok/s — above the 30 tok/s this site treats as faster than reading.
Open in the calculator →Read this carefully
Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.
Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.