Instructions to use xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8") model = AutoModelForMultimodalLM.from_pretrained("xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8
- SGLang
How to use xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8 with Docker Model Runner:
docker model run hf.co/xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8
Qwen3.6-27B-Fable-Fusion-711 — FP8 (block)
FP8 block-quantized build of DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP, produced with llm-compressor and built/tested on 2× AMD Radeon AI PRO R9700 (gfx1201 / RDNA4).
Inherits everything from the base model, including its uncensored/abliterated behaviour. Read the base model card before using this.
What this is
- Scheme:
FP8_BLOCK— 128×128 weight blocks, dynamic per-token-group (128) activations, e4m3. Data-free; no calibration set. - Format: native fp8 (
quant_method: fp8,weight_block_size: [128, 128],weight_scale_inv) — the same serialization Qwen's own FP8 checkpoints use. See Format for why, and for the compressed-tensors variant. - Size: ~36.5 GB (from 55.6 GB BF16).
- Multi-token prediction: preserved. Speculative decoding works.
- Vision tower: preserved in BF16 (333 tensors). Multimodal input works.
Quantized vs preserved
| Component | Precision | Count |
|---|---|---|
MLP (gate/up/down_proj, 64 layers) |
FP8 | 192 |
Full-attention q/k/v/o_proj (16 layers) |
FP8 | 64 |
| Linear attention / gated delta net (48 layers) | BF16 | — |
| Vision tower | BF16 | 333 |
embed_tokens, lm_head |
BF16 | — |
| MTP module | BF16 | 15 |
Gated DeltaNet is left in BF16 deliberately, matching the RedHatAI recipes for
this architecture family. Those layers carry recurrent state along the sequence,
so quantization error there propagates into every subsequent token rather than
staying local — the opposite of MLP error. A variant with in_proj_qkvz and
out_proj quantized (~31 GB) is measurably faster; this repo prioritises
accuracy.
If you do quantize the linear-attention stack: in_proj_a and in_proj_b
must stay in BF16. They are (48, 5120) on disk and vLLM fuses them into one
in_proj_ba MergedColumnParallelLinear of 96 outputs, which shards to 24 per
rank at TP=2. 128-block quantization requires partitions divisible by 128, so
load fails with:
ValueError: Weight output_partition_size = 24 is not divisible by
weight quantization block_n = 128.
Exclude them with re:.*in_proj_[ab]$ — they are ~0.25M params per layer against
~84M for in_proj_qkvz, so you lose essentially nothing.
Reproduction
from compressed_tensors.utils import save_mtp_tensors_to_checkpoint
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from transformers import AutoProcessor, AutoTokenizer, Qwen3_5ForConditionalGeneration
MODEL_ID = "DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP"
SAVE_DIR = "ff711-FP8-BLOCK"
model = Qwen3_5ForConditionalGeneration.from_pretrained(
MODEL_ID, dtype="auto", device_map="auto_offload", offload_folder="./offload",
)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
processor = AutoProcessor.from_pretrained(MODEL_ID)
recipe = QuantizationModifier(
targets="Linear",
scheme="FP8_BLOCK",
ignore=[
"re:.*lm_head",
"re:.*embed_tokens$",
"re:.*visual.*",
"re:.*model.visual.*",
"re:.*linear_attn.*",
],
)
oneshot(model=model, recipe=recipe)
model.save_pretrained(SAVE_DIR, save_compressed=True)
tokenizer.save_pretrained(SAVE_DIR)
processor.save_pretrained(SAVE_DIR)
save_mtp_tensors_to_checkpoint(source_model=MODEL_ID, dest_dir=SAVE_DIR)
Then converted to native fp8 (rename *.weight_scale to *.weight_scale_inv,
rewrite quantization_config); see Format.
save_mtp_tensors_to_checkpoint is not optional: the MTP tensors are in the
index but are not loaded by Qwen3_5ForConditionalGeneration, so
save_pretrained silently drops them and you lose speculative decoding.
Package versions used:
llmcompressor— 0.12.1.dev92+g8cec0acccompressed-tensors— 0.17.2a20260729transformers— 5.15.0.dev0torch— 2.11.0+rocm7.14accelerate— 1.14.0safetensors— 0.8.0
The llm-compressor and transformers builds are pre-release; the trailing
+g8cec0acc on llmcompressor is the commit it was built from. Stable releases
at the time did not carry qwen3_5 support or
save_mtp_tensors_to_checkpoint.
Deployment
Tested on 2× R9700 (gfx1201) under
stilldeadcode/vllm-radiance
0.5.7 — vLLM 0.26.0, torch 2.11.0+rocm7.14, ROCm 7.2, AITER built for gfx1201:
vllm serve /models/ff711-FP8-BLOCK \
--served-model-name ff711 \
--kv-cache-dtype fp8 \
--tensor-parallel-size 2 \
--attention-backend ROCM_AITER_UNIFIED_ATTN \
--speculative-config '{"method":"mtp","num_speculative_tokens":8,"attention_backend":"ROCM_AITER_UNIFIED_ATTN","disable_padded_drafter_batch":true}'
Pass --quantization fp8 so the native fp8 loader is selected — that is the
faster path on gfx1201 (see Format).
Add --language-model-only to skip the vision tower and free VRAM for KV cache.
Format
This checkpoint was produced by llm-compressor, which emits compressed-tensors
(quant_method: compressed-tensors, weight_scale), and then converted to the
native fp8 serialization (quant_method: fp8, weight_scale_inv).
The two formats are numerically identical. Both store float8_e4m3fn weights
with 128x128 block scales, and — verified empirically against the original BF16
weights — both store the dequantization multiplier, despite the _inv
suffix implying otherwise. Reconstructing weights as w_fp8 * scale gives 2.26%
mean relative error (the FP8 e4m3 floor); as w_fp8 / scale it is off by nine
orders of magnitude. So the conversion is a rename plus a config rewrite, with
no change to any weight value.
It is worth doing because the two formats select different vLLM loaders. compressed-tensors reports:
Selected TritonFp8BlockScaledMMKernel for CompressedTensorsW8A8Fp8
which works correctly but was measurably slower at decode on gfx1201 than the native fp8 path. Prefill was unaffected either way.
If you need the compressed-tensors variant, it is the same weights under the
other convention — rename *.weight_scale_inv back to *.weight_scale and
restore the config_groups form of quantization_config.
Evaluation
The base model's headline claim is ARC-C 0.711 at 8-bit.
| Benchmark | Base (published) | This checkpoint |
|---|---|---|
| ARC-C | 0.711 | TODO |
TODO: note eval harness, seeds, and sampler settings.
Caveats
- Only exercised on gfx1201 with the runtime above. Untested elsewhere.
- Uncensored/abliterated base — safety behaviour differs from stock Qwen3.6.
- Serialized as native fp8. If your stack expects compressed-tensors, see Format — the weights are the same, only the names differ.
- Downloads last month
- 2,584
Model tree for xzanomad/Qwen3.6-27B-Fable-Fusion-711-FP8
Base model
Qwen/Qwen3.6-27B