How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "RedHatAI/GLM-5.2-NVFP4-FP8"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "RedHatAI/GLM-5.2-NVFP4-FP8",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker
docker model run hf.co/RedHatAI/GLM-5.2-NVFP4-FP8
Quick Links

RedHatAI/GLM-5.2-NVFP4-FP8

Model Overview

  • Model Architecture: GlmMoeDsaForCausalLM
    • Input: Text
    • Output: Text
  • Model Optimizations:
    • Weight quantization: Mixed (FP8 attention, FP4 MoE)
    • Activation quantization: Mixed (FP8 attention, FP4 MoE)
  • Release Date: 2026-06-29
  • Version: 1.0
  • Model Developers: RedHatAI

This model is a quantized version of zai-org/GLM-5.2, with the MoE experts in FP4 (NVFP4) and the attention weights in FP8, and was evaluated against the unquantized model.

Model Optimizations

This model was obtained by applying mixed-precision quantization to zai-org/GLM-5.2: the MoE expert weights and activations are quantized to FP4 (NVFP4), while the attention weights and activations are quantized to FP8 (block-scaled), ready for inference with vLLM.

This optimization reduces the number of bits per parameter from 16 to an average of roughly 4–5, cutting disk size and GPU memory requirements by approximately 70–75% compared to the BF16 original.

The weights and activations of the linear operators within the attention and MoE blocks are quantized using LLM Compressor.

Deployment

vLLM Serving

This model is intended for deployment with vLLM and requires the following fix: https://github.com/vllm-project/vllm/pull/47780.

vllm serve RedHatAI/GLM-5.2-NVFP4-FP8 \
    --tensor-parallel-size 4 \
    --reasoning-parser glm45 \
    --tool-call-parser glm47 \
    --enable-auto-tool-choice \
    --kv-cache-dtype fp8

For additional serving options (e.g. MTP speculative decoding, larger tensor parallelism, or 1M-context configuration), refer to the vLLM recipe for GLM-5.2.

Creation

This model was created by applying LLM Compressor with calibration samples from UltraChat (HuggingFaceH4/ultrachat_200k), as presented in the code snippet below.

import torch
from compressed_tensors.offload import init_dist
from compressed_tensors.quantization.quant_scheme import (
    FP8_BLOCK,
    NVFP4,
    QuantizationScheme,
)
from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer

from llmcompressor import oneshot
from llmcompressor.datasets.utils import get_rank_partition
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.utils import load_context

# Load the model
init_dist()
model_id = "zai-org/GLM-5.2"
with load_context():
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        device_map="auto_offload",
        max_memory={},
        offload_folder="./offload",
    )
tokenizer = AutoTokenizer.from_pretrained(model_id)

# Select calibration dataset.
DATASET_ID = "HuggingFaceH4/ultrachat_200k"
DATASET_SPLIT = "train_sft"

# Select number of samples. 512 samples is a good place to start.
# Increasing the number of samples can improve accuracy.
NUM_CALIBRATION_SAMPLES = 512
MAX_SEQUENCE_LENGTH = 2048

# Load dataset and preprocess.
ds = load_dataset(
    DATASET_ID, split=get_rank_partition(DATASET_SPLIT, NUM_CALIBRATION_SAMPLES)
)
ds = ds.shuffle(seed=42)


def preprocess(example):
    return {
        "text": tokenizer.apply_chat_template(
            example["messages"],
            tokenize=False,
        )
    }


ds = ds.map(preprocess)


# Tokenize inputs.
def tokenize(sample):
    return tokenizer(
        sample["text"],
        padding=False,
        max_length=MAX_SEQUENCE_LENGTH,
        truncation=True,
        add_special_tokens=False,
    )


ds = ds.map(tokenize, remove_columns=ds.column_names)


# Configure the quantization algorithm to run.
recipe = QuantizationModifier(
    config_groups={
        "attention_shared_experts": QuantizationScheme(
            targets=[r"re:.*self_attn\..*"],
            **FP8_BLOCK,
        ),
        "mlp": QuantizationScheme(
            targets=[r"re:.*mlp\..*"],
            **NVFP4,
        ),
    },
    ignore=[
        r"re:^model\.layers\.[0-2]\..*"
        r"re:.*mlp\.gate.*",  # not technically necessary
        r"re:.*indexer\.weights_proj$",  # sensitive to quantization
        r"lm_head",
    ],
)

# Apply algorithms.
oneshot(
    model=model,
    dataset=ds,
    batch_size=4,
    recipe=recipe,
    shuffle_calibration_samples=False,
)

# Save to disk compressed.
# Note: base checkpoint generation_config needs fixing for newer transformers versions
model.generation_config.top_p = None
SAVE_DIR = model_id.rstrip("/").split("/")[-1] + "-NVFP4-FP8"
model.save_pretrained(SAVE_DIR, save_compressed=True)
tokenizer.save_pretrained(SAVE_DIR)

torch.distributed.destroy_process_group()

Evaluation

This model was evaluated on GPQA Diamond using lighteval, served with vLLM (OpenAI-compatible API).

Accuracy

Category Benchmark zai-org/GLM-5.2 RedHatAI/GLM-5.2-NVFP4-FP8 Recovery
Reasoning GPQA Diamond (0-shot, pass@1) 91.2 89.1 97.7%
Downloads last month
4,930
Safetensors
Model size
425B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/GLM-5.2-NVFP4-FP8

Base model

zai-org/GLM-5.2
Quantized
(149)
this model