Transformers
Telugu
English
tokenizer
telugu
bpe
akshara

Akshara tokenizer v0 (English + Telugu)

Llama 3 spends 13.3 tokens on an average Telugu word. This tokenizer spends 1.69, with English at 1.32 (FLORES-200 devtest), in a 65,536-token vocabulary.

Compare it yourself: Telugu Tokenizer Leaderboard (27 tokenizers, paste your own text). Code: github.com/kuluruvineeth/akshara (training: src/akshara/tokenizer/train.py, measurement: src/akshara/tokenizer/fertility.py).

Results

Tokens per whitespace-separated word on FLORES-200 devtest (1,012 sentences per language; 16,938 Telugu and 21,901 English words), no special tokens added. Lower is better.

Tokenizer Vocabulary Telugu English
Akshara v0 65,536 1.69 1.32
IndicBERT v2 (encoder only) 250,000 1.70 1.24
MuRIL (encoder only) 197,258 1.96 1.26
Sarvam-1 68,096 2.14 1.44
BLOOM 250,680 2.15 1.25
Gemma 3 262,145 2.84 1.24
GPT-4o (o200k) 200,000 3.06 1.23
Llama 4 201,135 4.52 1.24
DeepSeek-V3 128,815 5.97 1.24
Qwen3 151,669 11.41 1.26
Llama 3.2 128,256 13.30 1.24
SmolLM2 49,152 20.76 1.27

It ties the best Indic encoder tokenizer with about a quarter of its vocabulary, and beats every generative model's tokenizer measured. English costs about 6% more tokens than English-centric tokenizers, partly because digits are split one by one. The full table of 27 tokenizers is in the leaderboard Space.

What made the difference: the split rule

Before BPE runs, a regex cuts text into pieces, and merges can never cross a cut. GPT-4-style rules match words with \p{L}+, but Telugu vowel signs are Unicode category M (marks), not L (letters), so \p{L}+ cuts "తెలుగు" into త | ెల | ుగ | ు before learning starts. This tokenizer matches words with [\p{L}\p{M}]+.

On the same 20M-character sample and the same vocabulary size, that one change takes Telugu from 3.98 to 1.71 tokens per word (2.3× fewer) and costs English almost nothing (1.27 → 1.33).

Use

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("kuluruvineeth/akshara-tokenizer-v0")
ids = tokenizer("తెలుగు భాష చాలా అందమైనది.")["input_ids"]
print(len(ids), tokenizer.decode(ids))  # 6 తెలుగు భాష చాలా అందమైనది.

Or with the tokenizers library alone: Tokenizer.from_pretrained("kuluruvineeth/akshara-tokenizer-v0").

How it was trained

  • Model: byte-level BPE, 65,536 tokens including 64 special tokens: <|endoftext|>, <|im_start|>, <|im_end|>, <think>, </think>, <tool_call>, </tool_call>, <tool_response>, </tool_response>, <|image|>, <|audio|> and 53 reserved slots.
  • Pre-tokenization: the split rule above; digits are split individually.
  • Data: 100M characters: 50M of English web text (FineWeb), 10M of Python (codeparrot-clean-valid) and 40M of Telugu web text (FineWeb-2 tel_Telu).
  • Coverage: every assigned character of the Telugu Unicode block is seeded into training, so none falls back to raw bytes.
  • Reproducible: retraining gives a bit-identical file (sha256 2ade5f3e1d8f6ab4…).

Limitations

  • Version 0: it samples the first rows of each source, not the final training mixture. The production tokenizer is retrained on the real mix.
  • FLORES-200 is one domain (Wikipedia-style sentences); colloquial or code-mixed Telugu may tokenize differently.
  • Tokens per word depends on how words are counted; here, whitespace-separated, the same way for every tokenizer.

Part of Akshara

Akshara is a family of English + Telugu language models built end to end: data engine, tokenizer, pretraining, post-training and serving. Collection: అ Akshara.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train kuluruvineeth/akshara-tokenizer-v0

Space using kuluruvineeth/akshara-tokenizer-v0 1

Collection including kuluruvineeth/akshara-tokenizer-v0