LLMBOTTLENECK.COM
Hugging Face repository / answered by Qwen3.8-Flash-Next

beamster/Qwen3.8-Flash-Next-Sushi-4bpw: what it needs to run

beamster/Qwen3.8-Flash-Next-Sushi-4bpw is a quantization of Qwen/Qwen3.8-Flash-Next. Its configuration matches Qwen3.8-Flash-Next's on every field that decides memory, so Qwen3.8-Flash-Next's requirements are this repository's.

Answered by

Geometrypublished data. Both configurations are published files, read at the revisions named below and compared field by field. Not a test that the model loads.

Qwen3.8-Flash-Next (Qwen/Qwen3.8-Flash-Next) — every field that decides memory is equal; what differs is noted below and does not change a memory answer.

Before you use those figures

  • Its context limit is 1,048,576 tokens against the base's 262,144; the base's figures apply up to 262,144.
  • Its safetensors index counts 36,556,177,299 parameters against the base's 179,999,981,459 (79.7% apart). The compared geometry is the same, so the difference is in weights this comparison does not cover, and the base's file sizes are off by about that share for this repository.

Qwen3.8-Flash-Next: weight-file size by format

FormatSizeRangeBasis
FP16360.2 GB252.0 GB – 363.6 GBsize range
Q8_0191.4 GB133.9 GB – 194.4 GBsize range
Q6_K147.8 GB103.4 GB – 150.6 GBsize range
Q5_K_M128.6 GB86.6 GB – 150.6 GBsize range
Q5_0125.6 GB86.6 GB – 150.6 GBsize range
Q4_K_M110.5 GB70.9 GB – 150.6 GBsize range
Q4_0104.7 GB70.9 GB – 150.6 GBsize range
Q3_K_M90.0 GB54.1 GB – 150.6 GBsize range
Q2_K71.3 GB41.3 GB – 150.6 GBsize range

These are Qwen3.8-Flash-Next’s sizes. A model with the same geometry quantizes to the same size, to within any difference in parameter count noted above. Which devices hold each, and how fast they run it, is on Qwen3.8-Flash-Next’s page.

What was read, and where

  1. beamster/Qwen3.8-Flash-Next-Sushi-4bpw at revision f25c0c4acdc4 is tagged by its publisher as a quantization of Qwen/Qwen3.8-Flash-Next.
  2. beamster/Qwen3.8-Flash-Next-Sushi-4bpw/config.json at f25c0c4acdc4 was compared with Qwen/Qwen3.8-Flash-Next/config.json at de4b8e4d43b9, the revision the catalogue pins: layers and their pattern, widths, head counts, experts, latent and recurrent dimensions, vocabulary, tied embeddings and any vision or audio tower.

Differences that do not change a memory answer

  • max_position_embeddings: 1048576 here, 262144 in the base (context limit).
Licence, as tagged
other
Parameters in its safetensors index
37B
Task, as tagged
image-text-to-text
Last changed on Hugging Face
2026-10-07

Publishing this model? A badge for its card

The memory badge for this repository

[![Memory to run this model, from llmbottleneck.com](https://llmbottleneck.com/badge/beamster/Qwen3.8-Flash-Next-Sushi-4bpw)](https://llmbottleneck.com/hf/beamster/Qwen3.8-Flash-Next-Sushi-4bpw)

It states the memory to run the model at Q4_K_M with 8,192 tokens of context — built on this repository’s own Q4_K_M file when it publishes one — and links back to this page, where the evidence is.

Resolve another repository

The same answer as JSON: GET /v1/resolve?repo=beamster/Qwen3.8-Flash-Next-Sushi-4bpw, with a free key from /keys.

llmbottleneck
catalogue 2026-10-03models 327devices 135