FastH3 rank-16, int8 transformer
A complete FastH3 model directory whose transformer is stored in int8 and whose text encoder is stored in NVFP4. It exists so the four-step FastH3 recipe runs on a single consumer card: the released bf16 transformer is 44 GB and cannot be loaded on a 32 GB GPU at all.
| component | this repository | the bf16 it was made from |
|---|---|---|
| transformer | 23.0 GB, int8, rank-16 AdaLN | 44.2 GB, bf16, rank-16 AdaLN |
| text encoder | 25.9 GB, NVFP4 | 63 GB, bf16 |
| video VAE | 10.4 GB, fp32, unchanged | same |
| audio VAE | 0.6 GB, fp32, unchanged | same |
The rank-16 AdaLN reduction is itself a change against the original release, whose transformer is 65.5 GB in bf16.
Measured on one RTX 5090, 864x480, 124 frames, 5 grid points so four transformer forwards: 27.3 s per clip, 25.0 GB peak, transformer resident 1.6 GB with layerwise offload on. The bf16 transformer on the same card fails while reading shard 11 of 15 with 30.6 GB allocated.
Requirements
- FastVideo with the serialized int8 transformer path, https://github.com/hao-ai-lab/FastVideo/pull/1883. It is not on
mainyet. - A GPU with int8 tensor cores for the transformer, which is Turing and newer.
- A Blackwell GPU with
flashinfer-pythonfor the NVFP4 text encoder, which is sm_100 or sm_120 class. To run the transformer on an older card, pointtext_encoder_weightsat a bf16 encoder directory instead. - Single GPU. Packed int8 rows and per-row scales are not split across tensor-parallel ranks, and FSDP inference rejects quantized transformers.
Run it
hf download KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16-int8 --local-dir ./FastH3-r16-int8
FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 \
fastvideo generate --config your_config.yaml --generator.model_path ./FastH3-r16-int8
Do not pass transformer_quant. The checkpoint declares its own scheme in transformer/config.json and the loader refuses a runtime quantization request beside it. A LoRA adapter is also refused: merging into int8 codes would rewrite the row scales this checkpoint ships. Merge an adapter into the bf16 checkpoint and convert that instead.
What is inside
Every main transformer block stores seven linears as two tensors:
| tensor | dtype | shape | meaning |
|---|---|---|---|
weight |
int8 | [out, in] |
round(W / scale) clamped to -127 to 127 |
weight_scale |
float32 | [out] |
amax(abs(W[row])) / 127 |
The seven are attn.to_q, attn.to_k, attn.to_v, attn.to_out, attn.to_gate_compress, ff.fc_in and ff.fc_out, 350 linears over 50 blocks, 42.4 GB of bf16 becoming 21.2 GB. The token refiner, the rank-16 AdaLN factors, the patch, audio, context and time projections and the output head keep their released dtypes. Activations are quantized per row at call time and multiplied with torch._int_mm.
Accuracy
Each linear was probed during conversion against its bf16 product on random rows: worst case 2.7 percent relative error on ff.fc_out, mean 1.2 to 1.7 percent across the seven kinds.
End to end on one GB10 where both precisions fit, same NVFP4 encoder on both sides, same seeds: mean SSIM against the bf16 transformer is 0.573, 0.824 and 0.694 on three seeds, while two bf16 clips at different seeds score 0.226. The sampler stays on a neighbouring trajectory rather than a different one. SSIM against one sample is not a quality ranking, and a four-step distilled sampler moves under a one to three percent weight perturbation.
How it was made
python scripts/checkpoint_conversion/convert_minimax_h3_transformer_int8.py \
--src FastH3-4-step-Preview-v1-r16/transformer \
--dst FastH3-r16-int8/transformer
The source is the rank-16 reduction of the released four-step checkpoint, https://huggingface.co/KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16, produced with the rank-reduced AdaLN script from FastVideo PR 1699. The text encoder is https://huggingface.co/KyleNeverGivesUp/FastH3-text-encoder-nvfp4.
Scope and limits
Text to audio and video only, four transformer forwards. It inherits the scope of the base checkpoint: FL2VA and Ref2VA were not distilled, and difficult motion, fine detail and some audio may remain below base MiniMax H3. int8 is not a speedup: on a unified-memory box where bf16 also fits, it takes about 1.5 times the denoise time and saves 20 GB of resident weights. Its value is that a 32 GB card can run the model at all.
Licensed under the MiniMax H3 Community License, see LICENSE.
Model tree for KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16-int8
Base model
MiniMaxAI/MiniMax-H3