Architecture
Retrieved 2026-09-01 at pinned commit ec89142b67d7; parameter count from the safetensors index at the same revision.
- Architecture
- falcon
- Layers
- 32
- Hidden size
- 4,544
- Attention heads
- 71
- Vocabulary
- 65,024
LLM//BOTTLENECK
Sized at 8,192 tokens of context with the whole model in device memory, across 18 common devices. Speeds are estimates, not benchmarks.
Find the NVIDIA GeForce RTX 5070, the smallest card listed here that holds it: Amazon ↗ · eBay (new and used) ↗Store links may pay us a commission. They never decide which card is suggested — the memory arithmetic does.
| Format | Size | Range | Basis |
|---|---|---|---|
| FP16 | 14.4 GB | 10.1 GB – 14.6 GB | size range |
| Q8_0 | 7.67 GB | 5.37 GB – 8.30 GB | size range |
| Q6_K | 5.93 GB | 4.14 GB – 6.68 GB | size range |
| Q5_K_M | 5.16 GB | 3.47 GB – 6.68 GB | size range |
| Q5_0 | 5.04 GB | 3.47 GB – 6.68 GB | size range |
| Q4_K_M | 4.43 GB | 2.84 GB – 6.68 GB | size range |
| Q4_0 | 4.20 GB | 2.84 GB – 6.68 GB | size range |
| Q3_K_M | 3.61 GB | 2.17 GB – 6.68 GB | size range |
| Q2_K | 2.86 GB | 1.66 GB – 6.68 GB | size range |
A file size is not the memory a run needs: the KV cache and the runtime reserve come on top, and the calculator adds both.
| Device | Memory | Verdict | Decode | Calculator |
|---|---|---|---|---|
| NVIDIA GeForce RTX 5070 | 5.30 GB of 12 GB | fits | ~101 tok/s | Open → |
| AMD Radeon™ RX 9070 XT | 5.30 GB of 16 GB | fits | ~111 tok/s | Open → |
| NVIDIA GeForce RTX 4080 | 5.30 GB of 16 GB | fits | ~108 tok/s | Open → |
| NVIDIA GeForce RTX 5080 | 5.30 GB of 16 GB | fits | ~145 tok/s | Open → |
| NVIDIA GeForce RTX 5070 Ti | 5.30 GB of 16 GB | fits | ~135 tok/s | Open → |
| NVIDIA GeForce RTX 5060 Ti 16GB | 5.30 GB of 16 GB | fits | ~68 tok/s | Open → |
| NVIDIA GeForce RTX 4090 | 5.30 GB of 24 GB | fits | ~152 tok/s | Open → |
| AMD Radeon™ RX 7900 XTX | 5.30 GB of 24 GB | fits | ~166 tok/s | Open → |
| NVIDIA GeForce RTX 3090 | 5.30 GB of 24 GB | fits | ~141 tok/s | Open → |
| NVIDIA GeForce RTX 5090 | 5.30 GB of 32 GB | fits | ~271 tok/s | Open → |
| NVIDIA RTX 6000 Ada Generation | 5.30 GB of 48 GB | fits | ~145 tok/s | Open → |
| Apple M4 Pro | 5.30 GB of 64 GB | fits | ~43 tok/s | Open → |
| Apple M5 Pro | 5.30 GB of 64 GB | fits | ~47 tok/s | Open → |
| AMD Ryzen AI Max+ 395 with Radeon 8060S | 5.30 GB of 128 GB | fits | ~44 tok/s | Open → |
| Apple M3 Max | 5.30 GB of 128 GB | fits | ~56 tok/s | Open → |
| Apple M4 Max | 5.30 GB of 128 GB | fits | ~69 tok/s | Open → |
| NVIDIA DGX Spark | 5.30 GB of 128 GB | fits | ~41 tok/s | Open → |
| Apple M2 Ultra | 5.30 GB of 192 GB | fits | ~86 tok/s | Open → |
Decode is a calibrated estimate from the published calibration; capacity is the manufacturer's published ceiling, not guaranteed free memory. Fitting in memory is not the same as loading: whether the runtime and version you have supports this architecture and format on that machine has not been tested here. Offloaded rows assume there is enough system RAM for the overflow — the calculator checks that against the RAM you declare. Each “Open” link keeps this model, format and context.
Every rentable machine that holds the whole model, from Vast.ai and RunPod, with this site's speed estimate for each and the provider's own price, read live. Cheapest first; sort by cost per token or speed instead, or switch the format.
| Machine | Speed | Per hour | Per M tokens | Where | Rent |
|---|---|---|---|---|---|
| 1× RTX 306012 GB | ~54 tok/sFaster than you read | $0.036 | $0.18 | Vast.aiMarketplace · 99.8% reliable | Rent on Vast.ai |
| 1× RTX 308010 GB | ~114 tok/sFaster than you read | $0.082 | $0.20 | Vast.aiMarketplace · 98.3% reliable | Rent on Vast.ai |
| 1× RTX 30708 GB | ~67 tok/sFaster than you read | $0.082 | $0.34 | Vast.aiMarketplace · 99.1% reliableRunPod $0.13/h | Rent on Vast.ai |
| 1× RTX 4060 Ti 16GB16 GB | ~43 tok/sFaster than you read | $0.090 | $0.57 | Vast.aiMarketplace · 99.4% reliable | Rent on Vast.ai |
| 1× RTX 407012 GB | ~76 tok/sFaster than you read | $0.096 | $0.35 | Vast.aiMarketplace · 99.8% reliable | Rent on Vast.ai |
| 1× RTX 309024 GB | ~141 tok/sFaster than you read | $0.12 | $0.24 | Vast.aiMarketplace · 99.5% reliableRunPod $0.50/h | Rent on Vast.ai |
| 1× RTX 5060 Ti 16GB16 GB | ~67 tok/sFaster than you read | $0.12 | $0.50 | Vast.aiMarketplace · 99.1% reliable | Rent on Vast.ai |
| 1× RTX 3080 Ti12 GB | ~137 tok/sFaster than you read | $0.14 | $0.27 | Vast.aiMarketplace · 99.9% reliable | Rent on Vast.ai |
No Q4_K_M file of falcon-7b is catalogued here, so there is no exact command to vouch for. Start the machine from the llama.cpp server image ghcr.io/ggml-org/llama.cpp:server-cuda and point -hf at a Q4_K_M build from the publisher or a quantizer you trust — search Hugging Face for one. Check its file size against the memory figure on this page before you rent.
Every machine holds the whole model at Q4_K_M and 8,192 tokens of context, no offload. Speeds are this site’s single-stream decode estimates; prices are what each provider’s own API quoted, on-demand, for the whole machine. Vast hosts below 98% measured reliability are left out.
Referral links Vast.ai, RunPod and Novita pay us a share of what you spend if you sign up through these buttons. It costs you nothing, and it never decides an order or a recommendation: both are computed from the live price and the speed, and options that pay us nothing are listed and recommended on the same terms. How we rank
Buying a card instead, or paying by the token? Compare a year of falcon-7b three ways →
The numbers every figure above is computed from, with the file they came from.
Retrieved 2026-09-01 at pinned commit ec89142b67d7; parameter count from the safetensors index at the same revision.
curl -s https://llmbottleneck.com/v1/analyze \
-H "Authorization: Bearer $LLMB_KEY" -H "content-type: application/json" \
-d '{"model":"tiiuae-falcon-7b","quantization":"Q4_K_M","context":8192,"hardware":"nvidia-geforce-rtx-4090"}'Same engine, same evidence, every field sourced. Free key in one step →
<script src="https://llmbottleneck.com/widget.js" data-model="tiiuae-falcon-7b" data-quantization="Q4_K_M" data-hardware="nvidia-geforce-rtx-4090" data-context="8192"></script>
No key needed. Unbranded, with your own buy button, on Pro and Business →
LLM Bottleneck. “falcon-7b VRAM and hardware requirements.” Architecture from tiiuae/falcon-7b at revision ec89142b67d7, retrieved 2026-09-01. https://llmbottleneck.com/models/tiiuae-falcon-7b
Every figure above is either the published value or a reconstruction whose measured error is on the accuracy page.