LLMBOTTLENECK.COM
On NVIDIA GeForce RTX 4090 at Q4_K_M, 8,192 tokens

Qwen3-32B
vs Qwen3-30B-A3B

Qwen3-30B-A3B decodes 7.88× faster here.

Side by side · On NVIDIA GeForce RTX 4090 at Q4_K_M, 8,192 tokens

PropertyQwen3-32BQwen3-30B-A3B
Runsfits in device memoryfits in device memory
Decode~32 tok/sFaster than you read~251 tok/sFaster than you read
Memory needed22.7 GB of 24.0 GB20.2 GB of 24.0 GB
Spare or short1.29 GB spare3.84 GB spare
Weights19.8 GB18.6 GB
KV cache2.15 GB0.81 GB
Memory bandwidth1008 GB/s1008 GB/s

Qwen3-32B

CAPACITY LIMITEDVerdictunvalidated

It fits, but only 1.29 GB remains; one more current-size context needs 2.15 GB.

Open in the calculator →

Qwen3-30B-A3B

BALANCEDVerdictestimate

The configuration has memory headroom and no evaluated single hardware change improves founded decode speed by at least 20%.

Open in the calculator →

Read this carefully

Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.

Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.

llmbottleneck
catalogue 2026-10-03models 327devices 135