On NVIDIA GeForce RTX 5090 at Q4_K_M, 8,192 tokens
gpt-oss-20b
vs gpt-oss-120b
Side by side · On NVIDIA GeForce RTX 5090 at Q4_K_M, 8,192 tokens
| Property | gpt-oss-20b | gpt-oss-120b |
|---|---|---|
| Runs | fits in device memory | does not run as set |
| Decode | ~487 tok/sFaster than you read | does not run |
| Memory needed | 13.8 GB of 32.0 GB | 71.9 GB of 32.0 GB |
| Spare or short | 18.2 GB spare | short by 39.9 GB |
| Weights | 12.8 GB | 70.8 GB |
| KV cache | 0.20 GB | 0.31 GB |
| Memory bandwidth | 1792 GB/s | 1792 GB/s |
gpt-oss-20b
NOTHING TO FIXVerdictestimate
18.2 GB spare and ~487 tok/s — above the 30 tok/s this site treats as faster than reading.
Open in the calculator →gpt-oss-120b
NO FITVerdictunvalidated
This configuration is 39.9 GB over the available memory. Offloading the overflow needs more system RAM than was declared.
Open in the calculator →Read this carefully
Both columns come from the same engine and the same published inputs, so the comparison is like for like. It is one model at one context length: change either and the answer can invert, which is what the calculator links above are for — each opens its side with the same model, device, format and context.
Decode figures are calibrated estimates from the published calibration, whose measured error is on the accuracy page.