Qwen3.8-Flash-Next, INT4/INT6 mixed (AutoRound), for vLLM on 3x to 4x RTX 3090

Article, Article #2

This repository has two mixed-precision AutoRound quantizations of Qwen3.8-Flash-Next (for 3x3090 and 4x3090 users each), and the vLLM patch that runs it on three 24 GB cards with 262,144 tokens of context x 4 requests with text and images.

Measured on 3x RTX 3090 with 125 GiB of RAM:

3x3090 3x3090+MTP
VRAM, weights per GPU (vision tower included) 21.20 / 20.44 / 21.05 GiB 21.20 / 21.71 / 21.18 GiB
Context 262,144 tokens x 4 requests 262,144 tokens x 2 requests
Host RAM ~85 GiB (67.8 GiB pinned K/V pool) 41.2 GiB pinned K/V pool
Disk 95.37 GiB table file on NVMe same
Decode, 1 request, short prompt 95-99 tok/s 117-141 tok/s
Decode, 1 request, long context (8k-160k) 83-96 tok/s 110-116 tok/s
Decode, 2 requests 188 tok/s total 190 tok/s total
Decode, 4 requests 243-245 tok/s total 188 tok/s total
Prefill ~3,700 tok/s at 248k, 5,913-5,947 tok/s at 39k 5,613-5,809 tok/s at 39k

Both builds support prefix caching and image input.

Benchmarks

Benchmark This quant Official (BF16) ± 1 SE
IFBench, prompt-level loose 81.0 81.3 2.3
GPQA Diamond 90.4 91.7 2.1
LiveCodeBench v6, pass@1 92.4 91.9 2.3

The official figures come from the official model card.

  • IFBench: 300 single-turn prompts. Scoring used the official evaluation_lib, applied to the answer after the reasoning was removed. Strict scores were 73.3 (prompt level) and 76.2 (instruction level). Loose instruction level was 83.4.
  • GPQA Diamond: 198 questions with the simple-evals prompt. The answer options were shuffled with a fixed seed, and the grader read the last Answer: X line. Per domain: physics 95.3, chemistry 87.1, biology 84.2.
  • LiveCodeBench v6: 131 problems dated 2025-02-01 to 2025-04-06 (31 easy, 39 medium, 61 hard). The official label is "25.02-25.05", but livecodebench/code_generation_lite has no problems after 2025-04-06. The prompts used the generic lcb_runner template, and the code was taken from the last fenced block. Per difficulty: easy 100, medium 94.9, hard 86.9.

Pick a build

Directory GPUs Context Use it for
vllm-patch/3x3090/ 3 262,144 x 4 Default, no MTP. Good for throughput
vllm-patch/3x3090+MTP/ 3 262,144 x 2 For one user. One stream is 18-43% faster, 4 streams are 22% slower. Good for latency
vllm-patch/4x3090/ 4 For 4x 3090 users. Maybe more throughput than 3x3090. Recommended to use the weights on the 4x3090 branch with this build

Have four RTX 3090s?

Download the 4x3090 branch instead of main, and use vllm-patch/4x3090/. It is a different quantization made for four cards, which have more room for weights. Routed experts are INT4 gs32 (INT8 gs64 in layers 0, 1, 16, 30, 45, 46, 47), and linear_attn, QSA attention and the shared expert are INT8 gs64 instead of INT6. Everything else is the same as main. The numbers on this page are for the main weights on 3x3090.

hf download Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound --revision 4x3090 --local-dir /srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound-4x3090

Requirements

  • 3 or 4 x 24 GB NVIDIA GPUs (tested on RTX 3090).
  • ~85 GB free host RAM at the shipped settings, 67.8 GiB of it a pinned pool (3x3090+MTP: 41.2 GiB, 4x3090: 74.0 GiB). Less RAM works with less context, see "Settings".
  • 96 GB free on a local NVMe drive for the table file. Not a hard disk, not NFS.
  • podman or docker with the NVIDIA container toolkit.
  • The vLLM image vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96 (0.29.1rc1.dev47+gdc36fcce9). The patches may not apply to other nightlies.
  • PCIe lanes that is fast enough, preferably 4.0 x16.

Quick start (3x3090)

cd vllm-patch/3x3090
podman build -t flash-next-vllm:local -f Dockerfile \
  --build-arg BASE=docker.io/vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96 .
MODEL_DIR=/srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound \
PLE_TABLE=/srv/nvme/flash-next/ngram_table.bin \
podman compose up -d
  • MODEL_DIR: this repository.
  • PLE_TABLE: a file path on NVMe. It does not have to exist. The first start writes it (95.37 GiB) and is slow. Later starts are fast.

The server listens on 127.0.0.1:8000 (OpenAI API), model name Qwen3.8-Flash-Next.

For 3x3090+MTP/ and 4x3090/, see vllm-patch/README.md.

Settings

Change these in compose.yaml. The comments there give the numbers.

Setting Shipped If you change it
--kv-cache-memory-bytes 783000000 Sets context and host RAM together. 550000000: 262,144 x 2.8, 47.5 GiB pinned.
--max-num-seqs 4 Requests that run at once. Raise it only with --kv-cache-memory-bytes
--max-model-len 262144 Maximum context per request. Do not go above 262144
--enable-prefix-caching on Off (--no-enable-prefix-caching): 45% more context for the same RAM, but repeated prompts are prefilled again
--limit-mm-per-prompt, --mm-processor-kwargs 4 images, 2 MP Replace both with --language-model-only for a text-only server. Larger images than 2 MP were not tested
--max-num-batched-tokens 512 1568: 470 MiB more VRAM, same speed. 256: prefill takes 32% longer
--host, --port 127.0.0.1, 8000 Where the server listens

Do not remove PYTORCH_CUDA_ALLOC_CONF, VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS, ulimits, or --mamba-cache-mode align. Long prompts fail or the server does not start without them. The environment switches that turn single patches off are in vllm-patch/README.md.

Check that it started correctly

podman logs flash-next | grep "GPU KV cache size"     # 1,048,576 tokens, 4.00x
podman logs flash-next | grep "QSA host KV"           # 244 blocks x 12144 tokens ... 67.8 GiB, 12 times
podman logs flash-next | grep "QSA staging arena"     # no warning. A warning means slow prefill
podman logs flash-next | grep "PLE page prefetch"     # "process_madvise". "per-range madvise": add cap_add SYS_PTRACE
podman logs flash-next | grep "fused MoE decode"      # "enabled"

Setting attention block size to 1568 in the log means the image is not patched correctly.

3x3090+MTP shows 170 blocks x 9776 tokens ... 41.2 GiB 13 times. To see that MTP works, check that spec_decode_num_accepted_tokens_total in curl -s localhost:8000/metrics goes up. Do not sum all spec_decode_* lines: the *_created lines are timestamps.

Things you should know

  • 262,144 tokens should be considered as the limit, because that is max_position_embeddings. You can go up to 1M context, but I don't recommend it unless you really need it.
  • The 3x3090 build doesn't support speculative decoding. Use 3x3090+MTP for MTP.
  • VLLM_QSA_KV_OFFLOAD=1 requires TP=1, and the mmap PLE backend requires ETP=1.
  • TP is not implemented. I tried, but I didn't see any improvement.

What is quantized

Everything below uses compressed-tensors, pack-quantized, and symmetric group quantization.

Group Scheme What it covers
A INT4, gs128 MoE routed experts (512 per layer, top-10). 58.0 GiB, 92% of the body
B INT6, gs64 linear_attn in and out projections, QSA q/k/v/o_proj, shared expert
C INT8, gs64 hyper-connection low-rank mixers
D INT8, gs128 lm_head, embed_tokens, PLE key_proj and value_proj, indexer index_qk_proj
E BF16 router, vision tower, PLE n-gram table, MTP module

3x3090+MTP converts the MTP experts to INT4 in a separate directory (make_mtp_int4.py).

License

The model weights inherit the upstream Qwen Community License 1.0. See LICENSE in this directory.

The patches in vllm-patch/3x3090/, vllm-patch/3x3090+MTP/ and vllm-patch/4x3090/ are derivative works of vLLM, so they stay under the Apache License, Version 2.0, like vLLM itself. vllm-patch/LICENSE holds the full text, and vllm-patch/NOTICE names the vLLM files that each of them changes.

The Dockerfile and compose.yaml in each directory, and make_mtp_int4.py, contain no vLLM code. They are under the MIT license, in vllm-patch/LICENSE.MIT.

If you find these useful, make sure to like my repo.

Downloads last month
739
Safetensors
Model size
140B params
Tensor type
I32
·
BF16
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound

Quantized
(303)
this model

Collection including Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound