Instructions to use RedHatAI/GLM-5.2-NVFP4-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use RedHatAI/GLM-5.2-NVFP4-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="RedHatAI/GLM-5.2-NVFP4-FP8") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("RedHatAI/GLM-5.2-NVFP4-FP8") model = AutoModelForCausalLM.from_pretrained("RedHatAI/GLM-5.2-NVFP4-FP8", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use RedHatAI/GLM-5.2-NVFP4-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "RedHatAI/GLM-5.2-NVFP4-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/GLM-5.2-NVFP4-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/RedHatAI/GLM-5.2-NVFP4-FP8
- SGLang
How to use RedHatAI/GLM-5.2-NVFP4-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "RedHatAI/GLM-5.2-NVFP4-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/GLM-5.2-NVFP4-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "RedHatAI/GLM-5.2-NVFP4-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/GLM-5.2-NVFP4-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use RedHatAI/GLM-5.2-NVFP4-FP8 with Docker Model Runner:
docker model run hf.co/RedHatAI/GLM-5.2-NVFP4-FP8
RedHatAI/GLM-5.2-NVFP4-FP8
Model Overview
- Model Architecture: GlmMoeDsaForCausalLM
- Input: Text
- Output: Text
- Model Optimizations:
- Weight quantization: Mixed (FP8 attention, FP4 MoE)
- Activation quantization: Mixed (FP8 attention, FP4 MoE)
- Release Date: 2026-06-29
- Version: 1.0
- Model Developers: RedHatAI
This model is a quantized version of zai-org/GLM-5.2, with the MoE experts in FP4 (NVFP4) and the attention weights in FP8, and was evaluated against the unquantized model.
Model Optimizations
This model was obtained by applying mixed-precision quantization to zai-org/GLM-5.2: the MoE expert weights and activations are quantized to FP4 (NVFP4), while the attention weights and activations are quantized to FP8 (block-scaled), ready for inference with vLLM.
This optimization reduces the number of bits per parameter from 16 to an average of roughly 4–5, cutting disk size and GPU memory requirements by approximately 70–75% compared to the BF16 original.
The weights and activations of the linear operators within the attention and MoE blocks are quantized using LLM Compressor.
Deployment
vLLM Serving
This model is intended for deployment with vLLM and requires the following fix: https://github.com/vllm-project/vllm/pull/47780.
vllm serve RedHatAI/GLM-5.2-NVFP4-FP8 \
--tensor-parallel-size 4 \
--reasoning-parser glm45 \
--tool-call-parser glm47 \
--enable-auto-tool-choice \
--kv-cache-dtype fp8
For additional serving options (e.g. MTP speculative decoding, larger tensor parallelism, or 1M-context configuration), refer to the vLLM recipe for GLM-5.2.
Creation
This model was created by applying LLM Compressor with calibration samples from UltraChat (HuggingFaceH4/ultrachat_200k), as presented in the code snippet below.
import torch
from compressed_tensors.offload import init_dist
from compressed_tensors.quantization.quant_scheme import (
FP8_BLOCK,
NVFP4,
QuantizationScheme,
)
from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.datasets.utils import get_rank_partition
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.utils import load_context
# Load the model
init_dist()
model_id = "zai-org/GLM-5.2"
with load_context():
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto_offload",
max_memory={},
offload_folder="./offload",
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
# Select calibration dataset.
DATASET_ID = "HuggingFaceH4/ultrachat_200k"
DATASET_SPLIT = "train_sft"
# Select number of samples. 512 samples is a good place to start.
# Increasing the number of samples can improve accuracy.
NUM_CALIBRATION_SAMPLES = 512
MAX_SEQUENCE_LENGTH = 2048
# Load dataset and preprocess.
ds = load_dataset(
DATASET_ID, split=get_rank_partition(DATASET_SPLIT, NUM_CALIBRATION_SAMPLES)
)
ds = ds.shuffle(seed=42)
def preprocess(example):
return {
"text": tokenizer.apply_chat_template(
example["messages"],
tokenize=False,
)
}
ds = ds.map(preprocess)
# Tokenize inputs.
def tokenize(sample):
return tokenizer(
sample["text"],
padding=False,
max_length=MAX_SEQUENCE_LENGTH,
truncation=True,
add_special_tokens=False,
)
ds = ds.map(tokenize, remove_columns=ds.column_names)
# Configure the quantization algorithm to run.
recipe = QuantizationModifier(
config_groups={
"attention_shared_experts": QuantizationScheme(
targets=[r"re:.*self_attn\..*"],
**FP8_BLOCK,
),
"mlp": QuantizationScheme(
targets=[r"re:.*mlp\..*"],
**NVFP4,
),
},
ignore=[
r"re:^model\.layers\.[0-2]\..*"
r"re:.*mlp\.gate.*", # not technically necessary
r"re:.*indexer\.weights_proj$", # sensitive to quantization
r"lm_head",
],
)
# Apply algorithms.
oneshot(
model=model,
dataset=ds,
batch_size=4,
recipe=recipe,
shuffle_calibration_samples=False,
)
# Save to disk compressed.
# Note: base checkpoint generation_config needs fixing for newer transformers versions
model.generation_config.top_p = None
SAVE_DIR = model_id.rstrip("/").split("/")[-1] + "-NVFP4-FP8"
model.save_pretrained(SAVE_DIR, save_compressed=True)
tokenizer.save_pretrained(SAVE_DIR)
torch.distributed.destroy_process_group()
Evaluation
This model was evaluated on GPQA Diamond using lighteval, served with vLLM (OpenAI-compatible API).
Accuracy
| Category | Benchmark | zai-org/GLM-5.2 | RedHatAI/GLM-5.2-NVFP4-FP8 | Recovery |
|---|---|---|---|---|
| Reasoning | GPQA Diamond (0-shot, pass@1) | 91.2 | 89.1 | 97.7% |
- Downloads last month
- 4,930
Model tree for RedHatAI/GLM-5.2-NVFP4-FP8
Base model
zai-org/GLM-5.2
Install from pip and serve model
# Install vLLM from pip: pip install vllm# Start the vLLM server: vllm serve "RedHatAI/GLM-5.2-NVFP4-FP8"# Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/GLM-5.2-NVFP4-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'