Qwen3.8-Flash-Next, INT4/INT6 mixed (AutoRound), for vLLM on 3x to 4x RTX 3090
This repository has two mixed-precision AutoRound quantizations of Qwen3.8-Flash-Next (for 3x3090 and 4x3090 users each), and the vLLM patch that runs it on three 24 GB cards with 262,144 tokens of context x 4 requests with text and images.
Measured on 3x RTX 3090 with 125 GiB of RAM:
3x3090 |
3x3090+MTP |
|
|---|---|---|
| VRAM, weights per GPU (vision tower included) | 21.20 / 20.44 / 21.05 GiB | 21.20 / 21.71 / 21.18 GiB |
| Context | 262,144 tokens x 4 requests | 262,144 tokens x 2 requests |
| Host RAM | ~85 GiB (67.8 GiB pinned K/V pool) | 41.2 GiB pinned K/V pool |
| Disk | 95.37 GiB table file on NVMe | same |
| Decode, 1 request, short prompt | 95-99 tok/s | 117-141 tok/s |
| Decode, 1 request, long context (8k-160k) | 83-96 tok/s | 110-116 tok/s |
| Decode, 2 requests | 188 tok/s total | 190 tok/s total |
| Decode, 4 requests | 243-245 tok/s total | 188 tok/s total |
| Prefill | ~3,700 tok/s at 248k, 5,913-5,947 tok/s at 39k | 5,613-5,809 tok/s at 39k |
Both builds support prefix caching and image input.
Benchmarks
| Benchmark | This quant | Official (BF16) | ± 1 SE |
|---|---|---|---|
| IFBench, prompt-level loose | 81.0 | 81.3 | 2.3 |
| GPQA Diamond | 90.4 | 91.7 | 2.1 |
| LiveCodeBench v6, pass@1 | 92.4 | 91.9 | 2.3 |
The official figures come from the official model card.
- IFBench: 300 single-turn prompts. Scoring used the official
evaluation_lib, applied to the answer after the reasoning was removed. Strict scores were 73.3 (prompt level) and 76.2 (instruction level). Loose instruction level was 83.4. - GPQA Diamond: 198 questions with the simple-evals prompt. The answer options were shuffled with a fixed seed, and the grader read the last
Answer: Xline. Per domain: physics 95.3, chemistry 87.1, biology 84.2. - LiveCodeBench v6: 131 problems dated 2025-02-01 to 2025-04-06 (31 easy, 39 medium, 61 hard). The official label is "25.02-25.05", but
livecodebench/code_generation_litehas no problems after 2025-04-06. The prompts used the genericlcb_runnertemplate, and the code was taken from the last fenced block. Per difficulty: easy 100, medium 94.9, hard 86.9.
Pick a build
| Directory | GPUs | Context | Use it for |
|---|---|---|---|
vllm-patch/3x3090/ |
3 | 262,144 x 4 | Default, no MTP. Good for throughput |
vllm-patch/3x3090+MTP/ |
3 | 262,144 x 2 | For one user. One stream is 18-43% faster, 4 streams are 22% slower. Good for latency |
vllm-patch/4x3090/ |
4 | For 4x 3090 users. Maybe more throughput than 3x3090. Recommended to use the weights on the 4x3090 branch with this build |
Have four RTX 3090s?
Download the 4x3090 branch instead of main, and use vllm-patch/4x3090/. It is a different quantization made for four cards, which have more room for weights. Routed experts are INT4 gs32 (INT8 gs64 in layers 0, 1, 16, 30, 45, 46, 47), and linear_attn, QSA attention and the shared expert are INT8 gs64 instead of INT6. Everything else is the same as main. The numbers on this page are for the main weights on 3x3090.
hf download Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound --revision 4x3090 --local-dir /srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound-4x3090
Requirements
- 3 or 4 x 24 GB NVIDIA GPUs (tested on RTX 3090).
- ~85 GB free host RAM at the shipped settings, 67.8 GiB of it a pinned pool (
3x3090+MTP: 41.2 GiB,4x3090: 74.0 GiB). Less RAM works with less context, see "Settings". - 96 GB free on a local NVMe drive for the table file. Not a hard disk, not NFS.
- podman or docker with the NVIDIA container toolkit.
- The vLLM image
vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96(0.29.1rc1.dev47+gdc36fcce9). The patches may not apply to other nightlies. - PCIe lanes that is fast enough, preferably 4.0 x16.
Quick start (3x3090)
cd vllm-patch/3x3090
podman build -t flash-next-vllm:local -f Dockerfile \
--build-arg BASE=docker.io/vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96 .
MODEL_DIR=/srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound \
PLE_TABLE=/srv/nvme/flash-next/ngram_table.bin \
podman compose up -d
MODEL_DIR: this repository.PLE_TABLE: a file path on NVMe. It does not have to exist. The first start writes it (95.37 GiB) and is slow. Later starts are fast.
The server listens on 127.0.0.1:8000 (OpenAI API), model name Qwen3.8-Flash-Next.
For 3x3090+MTP/ and 4x3090/, see vllm-patch/README.md.
Settings
Change these in compose.yaml. The comments there give the numbers.
| Setting | Shipped | If you change it |
|---|---|---|
--kv-cache-memory-bytes |
783000000 |
Sets context and host RAM together. 550000000: 262,144 x 2.8, 47.5 GiB pinned. |
--max-num-seqs |
4 |
Requests that run at once. Raise it only with --kv-cache-memory-bytes |
--max-model-len |
262144 |
Maximum context per request. Do not go above 262144 |
--enable-prefix-caching |
on | Off (--no-enable-prefix-caching): 45% more context for the same RAM, but repeated prompts are prefilled again |
--limit-mm-per-prompt, --mm-processor-kwargs |
4 images, 2 MP | Replace both with --language-model-only for a text-only server. Larger images than 2 MP were not tested |
--max-num-batched-tokens |
512 |
1568: 470 MiB more VRAM, same speed. 256: prefill takes 32% longer |
--host, --port |
127.0.0.1, 8000 |
Where the server listens |
Do not remove PYTORCH_CUDA_ALLOC_CONF, VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS, ulimits, or --mamba-cache-mode align. Long prompts fail or the server does not start without them. The environment switches that turn single patches off are in vllm-patch/README.md.
Check that it started correctly
podman logs flash-next | grep "GPU KV cache size" # 1,048,576 tokens, 4.00x
podman logs flash-next | grep "QSA host KV" # 244 blocks x 12144 tokens ... 67.8 GiB, 12 times
podman logs flash-next | grep "QSA staging arena" # no warning. A warning means slow prefill
podman logs flash-next | grep "PLE page prefetch" # "process_madvise". "per-range madvise": add cap_add SYS_PTRACE
podman logs flash-next | grep "fused MoE decode" # "enabled"
Setting attention block size to 1568 in the log means the image is not patched correctly.
3x3090+MTP shows 170 blocks x 9776 tokens ... 41.2 GiB 13 times. To see that MTP works, check that spec_decode_num_accepted_tokens_total in curl -s localhost:8000/metrics goes up. Do not sum all spec_decode_* lines: the *_created lines are timestamps.
Things you should know
- 262,144 tokens should be considered as the limit, because that is
max_position_embeddings. You can go up to 1M context, but I don't recommend it unless you really need it. - The
3x3090build doesn't support speculative decoding. Use3x3090+MTPfor MTP. VLLM_QSA_KV_OFFLOAD=1requires TP=1, and the mmap PLE backend requires ETP=1.- TP is not implemented. I tried, but I didn't see any improvement.
What is quantized
Everything below uses compressed-tensors, pack-quantized, and symmetric group quantization.
| Group | Scheme | What it covers |
|---|---|---|
| A | INT4, gs128 | MoE routed experts (512 per layer, top-10). 58.0 GiB, 92% of the body |
| B | INT6, gs64 | linear_attn in and out projections, QSA q/k/v/o_proj, shared expert |
| C | INT8, gs64 | hyper-connection low-rank mixers |
| D | INT8, gs128 | lm_head, embed_tokens, PLE key_proj and value_proj, indexer index_qk_proj |
| E | BF16 | router, vision tower, PLE n-gram table, MTP module |
3x3090+MTP converts the MTP experts to INT4 in a separate directory (make_mtp_int4.py).
License
The model weights inherit the upstream Qwen Community License 1.0. See LICENSE in this directory.
The patches in vllm-patch/3x3090/, vllm-patch/3x3090+MTP/ and vllm-patch/4x3090/ are derivative works of vLLM, so they stay under the Apache License, Version 2.0, like vLLM itself. vllm-patch/LICENSE holds the full text, and vllm-patch/NOTICE names the vLLM files that each of them changes.
The Dockerfile and compose.yaml in each directory, and make_mtp_int4.py, contain no vLLM code. They are under the MIT license, in vllm-patch/LICENSE.MIT.
If you find these useful, make sure to like my repo.
- Downloads last month
- 739
Model tree for Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound
Base model
Qwen/Qwen3.8-Flash-Next