FastH3 rank-16, int8 transformer

A complete FastH3 model directory whose transformer is stored in int8 and whose text encoder is stored in NVFP4. It exists so the four-step FastH3 recipe runs on a single consumer card: the released bf16 transformer is 44 GB and cannot be loaded on a 32 GB GPU at all.

component this repository the bf16 it was made from
transformer 23.0 GB, int8, rank-16 AdaLN 44.2 GB, bf16, rank-16 AdaLN
text encoder 25.9 GB, NVFP4 63 GB, bf16
video VAE 10.4 GB, fp32, unchanged same
audio VAE 0.6 GB, fp32, unchanged same

The rank-16 AdaLN reduction is itself a change against the original release, whose transformer is 65.5 GB in bf16.

Measured on one RTX 5090, 864x480, 124 frames, 5 grid points so four transformer forwards: 27.3 s per clip, 25.0 GB peak, transformer resident 1.6 GB with layerwise offload on. The bf16 transformer on the same card fails while reading shard 11 of 15 with 30.6 GB allocated.

Requirements

  • FastVideo with the serialized int8 transformer path, https://github.com/hao-ai-lab/FastVideo/pull/1883. It is not on main yet.
  • A GPU with int8 tensor cores for the transformer, which is Turing and newer.
  • A Blackwell GPU with flashinfer-python for the NVFP4 text encoder, which is sm_100 or sm_120 class. To run the transformer on an older card, point text_encoder_weights at a bf16 encoder directory instead.
  • Single GPU. Packed int8 rows and per-row scales are not split across tensor-parallel ranks, and FSDP inference rejects quantized transformers.

Run it

hf download KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16-int8 --local-dir ./FastH3-r16-int8
FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 \
  fastvideo generate --config your_config.yaml --generator.model_path ./FastH3-r16-int8

Do not pass transformer_quant. The checkpoint declares its own scheme in transformer/config.json and the loader refuses a runtime quantization request beside it. A LoRA adapter is also refused: merging into int8 codes would rewrite the row scales this checkpoint ships. Merge an adapter into the bf16 checkpoint and convert that instead.

What is inside

Every main transformer block stores seven linears as two tensors:

tensor dtype shape meaning
weight int8 [out, in] round(W / scale) clamped to -127 to 127
weight_scale float32 [out] amax(abs(W[row])) / 127

The seven are attn.to_q, attn.to_k, attn.to_v, attn.to_out, attn.to_gate_compress, ff.fc_in and ff.fc_out, 350 linears over 50 blocks, 42.4 GB of bf16 becoming 21.2 GB. The token refiner, the rank-16 AdaLN factors, the patch, audio, context and time projections and the output head keep their released dtypes. Activations are quantized per row at call time and multiplied with torch._int_mm.

Accuracy

Each linear was probed during conversion against its bf16 product on random rows: worst case 2.7 percent relative error on ff.fc_out, mean 1.2 to 1.7 percent across the seven kinds.

End to end on one GB10 where both precisions fit, same NVFP4 encoder on both sides, same seeds: mean SSIM against the bf16 transformer is 0.573, 0.824 and 0.694 on three seeds, while two bf16 clips at different seeds score 0.226. The sampler stays on a neighbouring trajectory rather than a different one. SSIM against one sample is not a quality ranking, and a four-step distilled sampler moves under a one to three percent weight perturbation.

How it was made

python scripts/checkpoint_conversion/convert_minimax_h3_transformer_int8.py \
    --src FastH3-4-step-Preview-v1-r16/transformer \
    --dst FastH3-r16-int8/transformer

The source is the rank-16 reduction of the released four-step checkpoint, https://huggingface.co/KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16, produced with the rank-reduced AdaLN script from FastVideo PR 1699. The text encoder is https://huggingface.co/KyleNeverGivesUp/FastH3-text-encoder-nvfp4.

Scope and limits

Text to audio and video only, four transformer forwards. It inherits the scope of the base checkpoint: FL2VA and Ref2VA were not distilled, and difficult motion, fine detail and some audio may remain below base MiniMax H3. int8 is not a speedup: on a unified-memory box where bf16 also fits, it takes about 1.5 times the denoise time and saves 20 GB of resident weights. Its value is that a 32 GB card can run the model at all.

Licensed under the MiniMax H3 Community License, see LICENSE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16-int8