LLMBOTTLENECK.COM
Qwen3-30B-A3B at Q4_K_M, 8,192 tokens

NVIDIA GeForce RTX 4090
vs Apple M4 Max

NVIDIA GeForce RTX 4090 decodes 2.77× faster here.

Side by side · Qwen3-30B-A3B at Q4_K_M, 8,192 tokens

PropertyNVIDIA GeForce RTX 4090Apple M4 Max
Runsfits in device memoryfits in device memory
Decode~251 tok/sFaster than you read~91 tok/sFaster than you read
Memory needed20.2 GB of 24.0 GB20.2 GB of 128.0 GB
Spare or short3.84 GB spare107.8 GB spare
Weights18.6 GB18.6 GB
KV cache0.81 GB0.81 GB
Memory bandwidth1008 GB/s546 GB/s

NVIDIA GeForce RTX 4090

BALANCEDVerdictestimate

The configuration has memory headroom and no evaluated single hardware change improves founded decode speed by at least 20%.

Open in the calculator →

Apple M4 Max

NOTHING TO FIXVerdictestimate

107.8 GB spare and ~91 tok/s — above the 30 tok/s this site treats as faster than reading.

Open in the calculator →

Read this carefully

Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.

Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.

llmbottleneck
catalogue 2026-10-03models 327devices 135