gemma-4-E2B-it-text-only
Text-only language model extracted from google/gemma-4-E2B-it, a Gemma 4 any-to-any (image+video+audio+text) model.
What was removed
- The vision tower (
Gemma4VisionModel) and its patch embedder/pooler - The audio tower (
Gemma4AudioModel) and its conformer-style encoder layers - The vision/audio multimodal embedders (
Gemma4MultimodalEmbedder) that project soft tokens into the language model's embedding space - All image/video/audio token routing logic in
Gemma4Model.forward
What was kept
- Token embeddings and the Per-Layer Embedding (PLE) auxiliary embedding table (
embed_tokens,embed_tokens_per_layer) - PLE is a text-decoder feature, not a vision/audio one - The full decoder stack: sliding-window / full attention layers (with KV-sharing among the last
20layers), exactly as in the original text backbone (model.language_modelin the original checkpoint ->modelhere) - Final norm and LM head (
lm_head, tied toembed_tokenssincetie_word_embeddings=True) - Tokenizer (unchanged - plain text tokenizer, no image/video/audio preprocessor)
Parity verification
Outputs were checked against the original any-to-any model on text-only prompts (no images/audio/video).
Greedy generate() decodes were verified to match token-for-token, and logits matched with a max
absolute difference of 0.00e+00 (bf16 numerical noise floor).
Loading
This architecture (gemma4_text / Gemma4ForCausalLM) is natively supported in
transformers>=5.9. No trust_remote_code is required:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("tuandunghcmut/gemma-4-E2B-it-text-only", dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained("tuandunghcmut/gemma-4-E2B-it-text-only")
inputs = tok("The capital of France is", return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=30)
print(tok.decode(out[0], skip_special_tokens=True))
A standalone reference implementation (modeling_gemma4_text.py + configuration_gemma4_text.py,
verified bit-for-bit identical to the native transformers implementation on random weights, including
the PLE and KV-sharing code paths actually exercised by this checkpoint) is included in this repo for
transparency/portability. It is not required for loading - transformers already ships this
architecture natively - but documents exactly what the text-only forward pass does.
Original model
See google/gemma-4-E2B-it for the full any-to-any model, license, and training details. This repo inherits the gemma license from the base model.
- Downloads last month
- 8