LLMBOTTLENECK.COM
On NVIDIA GeForce RTX 4090 at Q4_K_M, 8,192 tokens

Qwen3-8B
vs phi-4

Qwen3-8B decodes 1.75× faster here.

Side by side · On NVIDIA GeForce RTX 4090 at Q4_K_M, 8,192 tokens

PropertyQwen3-8Bphi-4
Runsfits in device memoryfits in device memory
Decode~116 tok/sFaster than you read~67 tok/sFaster than you read
Memory needed7.04 GB of 24.0 GB11.4 GB of 24.0 GB
Spare or short17.0 GB spare12.6 GB spare
Weights5.03 GB8.89 GB
KV cache1.21 GB1.68 GB
Memory bandwidth1008 GB/s1008 GB/s

Qwen3-8B

NOTHING TO FIXVerdictestimate

17.0 GB spare and ~116 tok/s — above the 30 tok/s this site treats as faster than reading.

Open in the calculator →

phi-4

NOTHING TO FIXVerdictestimate

12.6 GB spare and ~67 tok/s — above the 30 tok/s this site treats as faster than reading.

Open in the calculator →

Read this carefully

Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.

Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.

llmbottleneck
catalogue 2026-10-03models 327devices 135