Instructions to use kuluruvineeth/akshara-tokenizer-v0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kuluruvineeth/akshara-tokenizer-v0 with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("kuluruvineeth/akshara-tokenizer-v0", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Akshara tokenizer v0 (English + Telugu)
Llama 3 spends 13.3 tokens on an average Telugu word. This tokenizer spends 1.69, with English at 1.32 (FLORES-200 devtest), in a 65,536-token vocabulary.
Compare it yourself: Telugu Tokenizer Leaderboard
(27 tokenizers, paste your own text). Code: github.com/kuluruvineeth/akshara
(training: src/akshara/tokenizer/train.py, measurement: src/akshara/tokenizer/fertility.py).
Results
Tokens per whitespace-separated word on FLORES-200 devtest (1,012 sentences per language; 16,938 Telugu and 21,901 English words), no special tokens added. Lower is better.
| Tokenizer | Vocabulary | Telugu | English |
|---|---|---|---|
| Akshara v0 | 65,536 | 1.69 | 1.32 |
| IndicBERT v2 (encoder only) | 250,000 | 1.70 | 1.24 |
| MuRIL (encoder only) | 197,258 | 1.96 | 1.26 |
| Sarvam-1 | 68,096 | 2.14 | 1.44 |
| BLOOM | 250,680 | 2.15 | 1.25 |
| Gemma 3 | 262,145 | 2.84 | 1.24 |
| GPT-4o (o200k) | 200,000 | 3.06 | 1.23 |
| Llama 4 | 201,135 | 4.52 | 1.24 |
| DeepSeek-V3 | 128,815 | 5.97 | 1.24 |
| Qwen3 | 151,669 | 11.41 | 1.26 |
| Llama 3.2 | 128,256 | 13.30 | 1.24 |
| SmolLM2 | 49,152 | 20.76 | 1.27 |
It ties the best Indic encoder tokenizer with about a quarter of its vocabulary, and beats every generative model's tokenizer measured. English costs about 6% more tokens than English-centric tokenizers, partly because digits are split one by one. The full table of 27 tokenizers is in the leaderboard Space.
What made the difference: the split rule
Before BPE runs, a regex cuts text into pieces, and merges can never cross a cut. GPT-4-style rules match words with
\p{L}+, but Telugu vowel signs are Unicode category M (marks), not L (letters), so \p{L}+ cuts "తెలుగు" into
త | ెల | ుగ | ు before learning starts. This tokenizer matches words with [\p{L}\p{M}]+.
On the same 20M-character sample and the same vocabulary size, that one change takes Telugu from 3.98 to 1.71 tokens per word (2.3× fewer) and costs English almost nothing (1.27 → 1.33).
Use
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("kuluruvineeth/akshara-tokenizer-v0")
ids = tokenizer("తెలుగు భాష చాలా అందమైనది.")["input_ids"]
print(len(ids), tokenizer.decode(ids)) # 6 తెలుగు భాష చాలా అందమైనది.
Or with the tokenizers library alone: Tokenizer.from_pretrained("kuluruvineeth/akshara-tokenizer-v0").
How it was trained
- Model: byte-level BPE, 65,536 tokens including 64 special tokens:
<|endoftext|>,<|im_start|>,<|im_end|>,<think>,</think>,<tool_call>,</tool_call>,<tool_response>,</tool_response>,<|image|>,<|audio|>and 53 reserved slots. - Pre-tokenization: the split rule above; digits are split individually.
- Data: 100M characters: 50M of English web text (FineWeb), 10M of Python (codeparrot-clean-valid) and 40M of
Telugu web text (FineWeb-2
tel_Telu). - Coverage: every assigned character of the Telugu Unicode block is seeded into training, so none falls back to raw bytes.
- Reproducible: retraining gives a bit-identical file (sha256
2ade5f3e1d8f6ab4…).
Limitations
- Version 0: it samples the first rows of each source, not the final training mixture. The production tokenizer is retrained on the real mix.
- FLORES-200 is one domain (Wikipedia-style sentences); colloquial or code-mixed Telugu may tokenize differently.
- Tokens per word depends on how words are counted; here, whitespace-separated, the same way for every tokenizer.
Part of Akshara
Akshara is a family of English + Telugu language models built end to end: data engine, tokenizer, pretraining, post-training and serving. Collection: అ Akshara.