Configuration Parsing Warning:In config.json: "num_experts" must be a number

Gemma 4 E4B IT Q4_K_M GGUF Dequantized BF16 for vLLM

This private artifact was converted from unsloth/gemma-4-E4B-it-GGUF file gemma-4-E4B-it-Q4_K_M.gguf into standalone Gemma4 text safetensors for Transformers and vLLM.

The converter streams GGUF tensors, dequantizes them to fp32, maps Gemma4 text/PLE tensor names, then casts once to BF16 before writing sharded safetensors. The output config is rewritten from the multimodal gemma4 wrapper to standalone gemma4_text / Gemma4ForCausalLM.

Validation

  • GGUF dry-run mapping: 720 input tensors; 719 mapped, rope_freqs.weight skipped as runtime metadata.
  • Output index: 719 HF keys, 5 safetensor shards, 15,036,138,068 bytes.
  • Transformers load: passed as Gemma4ForCausalLM with tied embeddings and no missing language weights.
  • Numeric spot-checks: passed with max_abs_vs_bf16=0 for sampled embeddings, PLE projection/norm, layer norms, attention, and layer scalar tensors.
  • Transformers chat-template generation smoke: passed.
  • vLLM chat-completions smoke: passed on vLLM 0.23.0.

vLLM command used

VLLM_USE_FLASHINFER_SAMPLER=0 vllm serve /path/to/model   --served-model-name gemma-4-e4b-it-gguf-q4_k_m-bf16   --dtype bfloat16   --max-model-len 4096   --gpu-memory-utilization 0.5   --trust-remote-code

This is an instruction model; use the chat template or /v1/chat/completions for behavioral smoke tests.

Downloads last month
19
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for exolabs/gemma-4-E4B-it-Q4_K_M-dequant-bf16-vllm

Finetuned
(409)
this model