NVIDIA H200 NVL
141 GB decides what fits. 4800 GB/s decides how fast it runs once it does.
- Largest popular model that fitsQwen3.8-Flash-Next180B params · Q4_K_M · needs 111.6 GB~907 tok/s Faster than you readOpen in the calculator →
- Best fast pickStep-3.7-Flash201B params · Q4_K_M · needs 128.2 GB~282 tok/s Faster than you readOpen in the calculator →
- Most downloaded that fitsQwen3.8-27B28B params · Q4_K_M · needs 18.6 GB~220 tok/s Faster than you readOpen in the calculator →
At 8,192 tokens of context with the whole model in device memory. Speeds are estimates from memory bandwidth, not benchmarks run on this card.
249 of 320 fit entirely
249 of the 320 models the engine can size fit entirely in device memory at Q4_K_M where it is published, otherwise the nearest published format, and 8,192 tokens. 10 run with some layers on system memory (32 GB assumed), and 61 do not run at all.
Featured models · Q4_K_M at 8,192 tokens, including offload
| Model | Needs | Verdict | Decode | Calculator |
|---|---|---|---|---|
| Qwen3.5-2B2.3B parameters | 2.32 GB | fits | ~2646 tok/s | Open → |
| NVIDIA-Nemotron-3-Nano-4B-BF164.0B parameters | 3.46 GB | fits | ~1226 tok/s | Open → |
| Qwen3-VL-8B-Instruct8.8B parameters | 7.04 GB | fits | ~555 tok/s | Open → |
| Qwen3.5-9B9.7B parameters | 7.05 GB | fits | ~672 tok/s | Open → |
| gemma-4-12B-it12B parameters | 9.54 GB | fits | ~399 tok/s | Open → |
| gpt-oss-20b21B parameters | 13.8 GB | fits | ~1305 tok/s | Open → |
| Qwen3.8-27B28B parameters | 18.6 GB | fits | ~220 tok/s | Open → |
| Qwen3.6-27B28B parameters | 18.6 GB | fits | ~220 tok/s | Open → |
| Kimi-Linear-48B-A3B-Instruct49B parameters | 30.4 GB | fits | ~1746 tok/s | Open → |
| Qwen3.8-Flash-Next180B parameters | 111.6 GB | fits | ~907 tok/s | Open → |
| DeepSeek-V4-Flash-Vision-Exp305B parameters | 188.1 GB | needs 47.9 GB of system RAM; 32 GB assumed | does not run | Open → |
| GLM-5.3-Flash321B parameters | 198.3 GB | needs 61.4 GB of system RAM; 32 GB assumed | does not run | Open → |
These examples are selected from prominent labs using the catalogue’s latest Hugging Face 30-day downloads and repository-creation freshness signal, with newer releases guaranteed a place. Offloaded rows assume 32 GB of system RAM, and a speed is only shown for a row that runs. Fitting in memory is not the same as loading: whether the runtime and version you have supports each architecture and format on this machine has not been tested here. Each “Open” link carries the same model, format, context and RAM into the calculator.
Fits entirely in NVIDIA H200 NVL memory — 249 of 320 sized models
The most demanding model that fits is Step-3.7-Flash at Q4_K_M: 128.2 GB of the 141 GB, leaving 12.8 GB spare.
Fully resident at 8,192 tokens (or the model’s own maximum, where that is shorter), offload off, at 4800 GB/s. Each row is a run of the engine for this configuration; the rows start with current, prominent releases and “fits” is memory, not a tested runtime. Older or less prominent models remain available through this search and “Show all”.
Showing 40 of 249 models that fit.
| Model | Format | Needs | Spare | Decode | Downloads / 30d | Calculator |
|---|---|---|---|---|---|---|
| Qwen3.8-27BQwen · 28B params | Q4_K_M | 18.6 GB | 122.4 GB | ~220 tok/s | 6.9M | Open → |
| Qwen3.8-Flash-NextNEWQwen · 180B params | Q4_K_M | 111.6 GB | 29.4 GB | ~907 tok/s | 1.4M | Open → |
| gemma-4-26B-A4B-itGoogle · 26B params | Q4_K_M | 16.8 GB | 124.2 GB | ~1226 tok/s | 13M | Open → |
| gemma-4-31B-itGoogle · 31B params | Q4_K_M | 22.2 GB | 118.8 GB | ~159 tok/s | 9.9M | Open → |
| Qwen3.5-9BQwen · 9.7B params | Q4_K_M | 7.05 GB | 134.0 GB | ~672 tok/s | 9M | Open → |
| Qwen3.5-4BQwen · 4.7B params | Q4_K_M | 3.98 GB | 137.0 GB | ~1208 tok/s | 7.8M | Open → |
| Qwen3.6-35B-A3BQwen · 36B params | Q4_K_M | 23.1 GB | 117.9 GB | ~1760 tok/s | 3.3M | Open → |
| gemma-4-12B-itGoogle · 12B params | Q4_K_M | 9.54 GB | 131.5 GB | ~399 tok/s | 1.9M | Open → |
| Qwen3.6-27BQwen · 28B params | Q4_K_M | 18.6 GB | 122.4 GB | ~220 tok/s | 2.5M | Open → |
| Qwen3.5-2BQwen · 2.3B params | Q4_K_M | 2.32 GB | 138.7 GB | ~2646 tok/s | 4.9M | Open → |
| gemma-4-E4B-itGoogle · 8.0B params | Q4_K_M | 5.86 GB | 135.1 GB | ~1079 tok/s | 4.4M | Open → |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16NVIDIA · 32B params | Q4_K_M | 20.3 GB | 120.7 GB | ~167 tok/s | 530K | Open → |
| NVIDIA-Nemotron-3-Nano-4B-BF16NVIDIA · 4.0B params | Q4_K_M | 3.46 GB | 137.5 GB | ~1226 tok/s | 3.4M | Open → |
| gemma-4-E2B-itGoogle · 5.1B params | Q4_K_M | 4.02 GB | 137.0 GB | ~2850 tok/s | 3M | Open → |
| Qwen3.5-0.8BQwen · 873M params | Q4_K_M | 1.46 GB | 139.5 GB | ~5796 tok/s | 2.6M | Open → |
| Qwen3-VL-8B-InstructQwen · 8.8B params | Q4_K_M | 7.04 GB | 134.0 GB | ~555 tok/s | 14.6M | Open → |
| Qwen3.5-27BQwen · 28B params | Q4_K_M | 18.6 GB | 122.4 GB | ~220 tok/s | 1.9M | Open → |
| North-Micro-Vision-InstructCohereLabs · 2.5B params | Q4_K_M | 2.99 GB | 138.0 GB | ~1740 tok/s | 180.7K | Open → |
| Qwen3.5-35B-A3BQwen · 36B params | Q4_K_M | 23.1 GB | 117.9 GB | ~1760 tok/s | 1.6M | Open → |
| NVIDIA-Nemotron-3-Super-120B-A12B-BF16NVIDIA · 124B params | Q4_K_M | 76.9 GB | 64.1 GB | ~43 tok/s | 1.2M | Open → |
| granite-4.2-8bIBM · 8.8B params | Q4_K_M | 7.49 GB | 133.5 GB | ~505 tok/s | 128.4K | Open → |
| GLM-4.7-Flashzai-org · 31B params | Q4_K_M | 20.4 GB | 120.6 GB | ~166 tok/s | 1.8M | Open → |
| granite-4.1-3bIBM · 3.4B params | Q4_K_M | 3.72 GB | 137.3 GB | ~1178 tok/s | 521.8K | Open → |
| LFM2.5-2.6BLiquidAI · 2.7B params | Q4_K_M | 2.61 GB | 138.4 GB | ~1890 tok/s | 108.1K | Open → |
| Qwen3-VL-4B-InstructQwen · 4.4B params | Q4_K_M | 4.72 GB | 136.3 GB | ~881 tok/s | 3.5M | Open → |
| granite-4.1-30bIBM · 29B params | Q4_K_M | 20.4 GB | 120.6 GB | ~166 tok/s | 301.5K | Open → |
| Qwen3.5-122B-A10BQwen · 125B params | Q4_K_M | 77.9 GB | 63.1 GB | ~644 tok/s | 512.6K | Open → |
| Qwen3-VL-2B-InstructQwen · 2.1B params | Q4_K_M | 3.05 GB | 138.0 GB | ~1598 tok/s | 2.8M | Open → |
| granite-4.2-3bIBM · 3.7B params | Q4_K_M | 3.72 GB | 137.3 GB | ~1178 tok/s | 50.6K | Open → |
| LFM2.5-230MLiquidAI · 230M params | Q4_K_M | 1.05 GB | 139.9 GB | ~13174 tok/s | 88.6K | Open → |
| Qwen3-0.6BQwen · 752M params | Q4_K_M | 2.22 GB | 138.8 GB | ~2451 tok/s | 29.7M | Open → |
| Qwen3-Coder-NextQwen · 80B params | Q4_K_M | 50.0 GB | 91.0 GB | ~1171 tok/s | 596.3K | Open → |
| gpt-oss-20bOpenAI · 21B params | Q4_K_M | 13.8 GB | 127.2 GB | ~1305 tok/s | 6.6M | Open → |
| granite-4.2-30bIBM · 29B params | Q4_K_M | 20.7 GB | 120.3 GB | ~166 tok/s | 35.8K | Open → |
| granite-4.1-8bIBM · 8.8B params | Q4_K_M | 7.49 GB | 133.5 GB | ~505 tok/s | 179.3K | Open → |
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16NVIDIA · 32B params | Q4_K_M | 20.3 GB | 120.7 GB | ~167 tok/s | 875.2K | Open → |
| gpt-oss-120bOpenAI · 117B params | Q4_K_M | 71.9 GB | 69.1 GB | ~918 tok/s | 4.5M | Open → |
| LFM2.5-VL-3BLiquidAI · 3.1B params | Q4_K_M | 2.85 GB | 138.1 GB | ~1890 tok/s | 26.7K | Open → |
| MiMo-V2.6-Distill-Qwen-9BNEWXiaomiMiMo · 9.4B params | Q4_K_M | 6.90 GB | 134.1 GB | ~672 tok/s | 14.6K | Open → |
| Qwen3-4B-Instruct-2507Qwen · 4.0B params | Q4_K_M | 4.72 GB | 136.3 GB | ~881 tok/s | 3.7M | Open → |
Sized at the format most people actually download, not at FP16. The catalogue ordered by downloads →
Rent an H200 NVL by the hour
What the NVIDIA H200 NVL costs on Vast.ai and RunPod right now, per machine, from one card to eight. Every price is read from the provider's own API.
On-demand prices for the whole machine. Vast hosts below 98% measured reliability are left out; RunPod’s Community Cloud is vetted third-party hosts and its Secure Cloud is data-centre capacity. Every rentable card compared.
Referral links Vast.ai, RunPod and Novita pay us a share of what you spend if you sign up through these buttons. It costs you nothing, and it never decides an order or a recommendation: both are computed from the live price and the speed, and options that pay us nothing are listed and recommended on the same terms. How we rank
The specification behind every figure
What the manufacturer publishes for this device, and the pages it was read from.
Manufacturer specification
| Memory scope | dedicated |
|---|---|
| Capacity | 141 GB |
| Published options | 141 GB |
| Bandwidth | 4800 GB/s |
| Memory type | HBM3e |
| Bus width | Not published |
| FP32 peak | 60 TFLOPS |
| Dense matrix peak | Not published without sparsity |
| Max thermal design power (TDP), configurable ceiling | 600 W |
Source ledger
Caveats
- NVIDIA marks the H200 table specifications as preliminary and subject to change; memory capacity is per GPU, not the aggregate of an NVL board.
- No dense matrix throughput is catalogued for this device: NVIDIA's H200 page publishes "FP16 Tensor Core 1,671 teraFLOPS" footnoted "With sparsity", and no dense figure. Time to first token is withheld rather than priced against a sparsity figure.
Published capacity is a hardware ceiling, not guaranteed free runtime memory. The calculator shows the runtime reserve separately rather than folding it into a single number.
Run the diagnostic on the NVIDIA H200 NVL →