GLM-4.7-Flash-Hauhau-Aggressive-W6A8
Native compressed-tensors W6A8 checkpoint published by Xananthium. Credit and upstream terms: HauhauCS/GLM-4.7-Flash-Uncensored-HauhauCS-Aggressive. The source model card is preserved in README_ORIGINAL.md. Quantization provenance and retained high-precision components are recorded in conversion-receipt.json.
Six bits on a 3090? Hell yes. >:)
I wanted W6A8, so that is what we built: six-bit stored weights, dynamic eight-bit activations, and vLLM doing the serving. Humming packs the INT6 values across INT32 word boundaries and unpacks them for the 3090's supported INT8 matrix math. The cards do not need a native six-bit Tensor Core instruction. The extra unpacking and scale work is part of the cost, and the speed tests include it.
The six-bit weights use symmetric groups of 128. Embeddings, routers, norms, vision, and designated sensitive components keep higher precision. BF16 remains the surrounding model dtype; eligible quantized linears actually use INT8 activations, with silent activation fallback disabled. TurboQuant Q4 is a separate KV-cache choice when the architecture supports it.
Accuracy matters too. A smaller checkpoint gets us memory back, but the model still has to follow the damn instructions, call tools, and get the security answers right. That is why the results include the knowledge quiz, security-policy checks, long-context recall, real Claude Code tool use, and concurrency. Failed checks stay in the report.
The measured numbers for this checkpoint are below. A six-bit speed gain over an eight-bit baseline is unmeasured unless a matched comparison report is attached.
Measured on two RTX 3090 24 GiB cards, vLLM 0.31.0
| Test | Result |
|---|---|
| CyberMetric-80-v1, primary 2048-token output budget | 76/80 |
| CyberMetric-80 after capped-response retries at 4096 tokens | 78/80 |
| Synthetic security policy checks, first answer | 23/43 |
| Synthetic long-context replay | 53/54 |
| Synthetic exact-answer reasoning, three repeated samples | 18/24 |
| Reasoning after capped-response retries at 4096 tokens | 18/24 |
| Cases correct in every repeat | 6/8 |
| Ordinary workflow recovery after an initial tool failure | 6/6 |
| Extra attempts at a failed tool approach | 4 |
| Short-context decode | 122.5335 tokens/s |
| Long-context decode | 56.2375 tokens/s |
| Four simultaneous requests, aggregate | 297.6825 tokens/s |
Reports in results/ contain the exact per-model profile, budgets, formatting metrics, failed checks and client integration outcomes. Counting successful test programs does not mean all accuracy checks passed. Token throughput includes reasoning tokens; it is not visible-answer throughput. The policy assessment is synthetic, and lengthy reasoning can exhaust the output budget. Eight independent long prefixes were replayed; this was not a fresh eight-turn conversation. CyberMetric is a small public knowledge quiz that may overlap training data, and does not establish practical security competence or a broad model ranking. Questions and private reasoning traces are not redistributed.
vLLM loading
hf download Xananthium/GLM-4.7-Flash-Hauhau-Aggressive-W6A8 --local-dir ./checkpoint
pip install --no-deps ./checkpoint/runtime/humming-compile-compat
LD_LIBRARY_PATH="$VIRTUAL_ENV"/lib/python3.12/site-packages/nvidia/cu13/lib \
LOCAL_HUMMING_COMPILE_COMPAT=1 \
VLLM_HUMMING_INPUT_QUANT_CONFIG='{"dtype":"int8","group_size":0,"allow_fallback":false}' \
vllm serve ./checkpoint --quantization humming --linear-backend humming --moe-backend humming --tensor-parallel-size 2 --dtype bfloat16 --trust-remote-code --enable-auto-tool-choice --tool-call-parser glm47 --reasoning-parser glm47 --kv-cache-dtype bfloat16 --max-model-len 147456
The loading recipe requires the architecture and quantization support in vLLM 0.31.0. Consult results/serving-profile.json for tested sampling and memory settings. Model-side INT8 activations apply only to eligible quantized linears; embeddings, routers, norms and designated sensitive components retain higher precision. KV cache precision is a separate setting. GLM MLA does not use TurboQuant in this version. Retained MTP weights do not imply speculative decoding was tested.
The complete comparison: The 3090 Plan
Seven contenders. Two RTX 3090s. The Council kept the failed checks in the record. The full report, charts, and evaluation scripts compare cybersecurity knowledge, reasoning, consistency, recovery after failed tools, long context, and concurrency. A Hugging Face mirror contains the same report.
The owner selected Nemotron Heretic W8A8 for deployment after reviewing the results. The earlier automated quality ranking remains visible and recommended REDCELL; the deployment choice does not change any measured score.
This checkpoint's complete final-round record is in results/full-round-evaluation.json. The public report explains the output-budget retries, architecture-specific cache settings, and limits of the small local comparison. No Kali workflows were run.
- Downloads last month
- 29