gemma-4-E2B-it-text-only

Text-only language model extracted from google/gemma-4-E2B-it, a Gemma 4 any-to-any (image+video+audio+text) model.

What was removed

  • The vision tower (Gemma4VisionModel) and its patch embedder/pooler
  • The audio tower (Gemma4AudioModel) and its conformer-style encoder layers
  • The vision/audio multimodal embedders (Gemma4MultimodalEmbedder) that project soft tokens into the language model's embedding space
  • All image/video/audio token routing logic in Gemma4Model.forward

What was kept

  • Token embeddings and the Per-Layer Embedding (PLE) auxiliary embedding table (embed_tokens, embed_tokens_per_layer) - PLE is a text-decoder feature, not a vision/audio one
  • The full decoder stack: sliding-window / full attention layers (with KV-sharing among the last 20 layers), exactly as in the original text backbone (model.language_model in the original checkpoint -> model here)
  • Final norm and LM head (lm_head, tied to embed_tokens since tie_word_embeddings=True)
  • Tokenizer (unchanged - plain text tokenizer, no image/video/audio preprocessor)

Parity verification

Outputs were checked against the original any-to-any model on text-only prompts (no images/audio/video). Greedy generate() decodes were verified to match token-for-token, and logits matched with a max absolute difference of 0.00e+00 (bf16 numerical noise floor).

Loading

This architecture (gemma4_text / Gemma4ForCausalLM) is natively supported in transformers>=5.9. No trust_remote_code is required:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("tuandunghcmut/gemma-4-E2B-it-text-only", dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained("tuandunghcmut/gemma-4-E2B-it-text-only")

inputs = tok("The capital of France is", return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=30)
print(tok.decode(out[0], skip_special_tokens=True))

A standalone reference implementation (modeling_gemma4_text.py + configuration_gemma4_text.py, verified bit-for-bit identical to the native transformers implementation on random weights, including the PLE and KV-sharing code paths actually exercised by this checkpoint) is included in this repo for transparency/portability. It is not required for loading - transformers already ships this architecture natively - but documents exactly what the text-only forward pass does.

Original model

See google/gemma-4-E2B-it for the full any-to-any model, license, and training details. This repo inherits the gemma license from the base model.

Downloads last month
8
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tuandunghcmut/gemma-4-E2B-it-text-only

Finetuned
(385)
this model
Finetunes
1 model

Collection including tuandunghcmut/gemma-4-E2B-it-text-only