LLMBOTTLENECK.COM
jica98 / qwen3_5_text

qwen3.5-4B-super-coder

Parameter count not published.

Architecturepublished data

Architecture available · weight sizes unavailable

Weight sizeno data

This snapshot holds qwen3.5-4B-super-coder’s architecture but no weight size for any format, so the calculator cannot answer for it and it has no hardware table.

Why: Neither a published artifact nor a published parameter count exists for this model at its pinned revision. A size is never inferred from the model’s name.

The coverage page explains what the catalogue does not size, and why.

Under the hood

Architecture, read from the publisher’s file

The numbers every figure above is computed from, with the file they came from.

Architecture

✓ Architecture read from the published config.json

Retrieved 2026-09-01 at pinned commit 1f2362bff662.

Architecture
qwen3_5_text
Layers
32
Hidden size
2,560
Attention heads
16
KV heads
4
Head dimension
256
Feed-forward width
9,216
Vocabulary
248,320
Context ceiling
262,144
RoPE theta
10,000,000

Where the memory goes

tokenembedding× 32 decoder blocksGrouped-query attention16 query · 4 KV headsfull contextFeed-forwardone networkall activeoutputprojectiongrows with contextfixed per token

Grouped-query attention shares each key/value head across 4 query heads, so the KV cache is 25% of what multi-head attention would need at the same context.

This is a multimodal checkpoint (vision); the diagram and memory sizing above cover the text decoder. Encoder/audio/image activations are not included in the KV or weight figures.

Citing this page

LLM Bottleneck. “qwen3.5-4B-super-coder VRAM and hardware requirements.” Architecture from jica98/qwen3.5-4B-super-coder at revision 1f2362bff662, retrieved 2026-09-01. https://llmbottleneck.com/models/jica98-qwen3-5-4b-super-coder

Every figure above is either the published value or a reconstruction whose measured error is on the accuracy page.

llmbottleneck
catalogue 2026-10-03models 327devices 135