Qwen3-30B-A3B at Q4_K_M, 8,192 tokens
NVIDIA GeForce RTX 4090
vs Apple M4 Max
NVIDIA GeForce RTX 4090 decodes 2.77× faster here.
Side by side · Qwen3-30B-A3B at Q4_K_M, 8,192 tokens
| Property | NVIDIA GeForce RTX 4090 | Apple M4 Max |
|---|---|---|
| Runs | fits in device memory | fits in device memory |
| Decode | ~251 tok/sFaster than you read | ~91 tok/sFaster than you read |
| Memory needed | 20.2 GB of 24.0 GB | 20.2 GB of 128.0 GB |
| Spare or short | 3.84 GB spare | 107.8 GB spare |
| Weights | 18.6 GB | 18.6 GB |
| KV cache | 0.81 GB | 0.81 GB |
| Memory bandwidth | 1008 GB/s | 546 GB/s |
NVIDIA GeForce RTX 4090
BALANCEDVerdictestimate
The configuration has memory headroom and no evaluated single hardware change improves founded decode speed by at least 20%.
Open in the calculator →Apple M4 Max
NOTHING TO FIXVerdictestimate
107.8 GB spare and ~91 tok/s — above the 30 tok/s this site treats as faster than reading.
Open in the calculator →Read this carefully
Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.
Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.