Instructions to use summykai/Qwen3-14B-chem-dyn-tokenizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use summykai/Qwen3-14B-chem-dyn-tokenizer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="summykai/Qwen3-14B-chem-dyn-tokenizer") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("summykai/Qwen3-14B-chem-dyn-tokenizer") model = AutoModelForCausalLM.from_pretrained("summykai/Qwen3-14B-chem-dyn-tokenizer", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use summykai/Qwen3-14B-chem-dyn-tokenizer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "summykai/Qwen3-14B-chem-dyn-tokenizer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summykai/Qwen3-14B-chem-dyn-tokenizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/summykai/Qwen3-14B-chem-dyn-tokenizer
- SGLang
How to use summykai/Qwen3-14B-chem-dyn-tokenizer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "summykai/Qwen3-14B-chem-dyn-tokenizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summykai/Qwen3-14B-chem-dyn-tokenizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "summykai/Qwen3-14B-chem-dyn-tokenizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summykai/Qwen3-14B-chem-dyn-tokenizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use summykai/Qwen3-14B-chem-dyn-tokenizer with Docker Model Runner:
docker model run hf.co/summykai/Qwen3-14B-chem-dyn-tokenizer
Model Card for summykai/Qwen3-14B-chem-dyn-tokenizer
A Qwen3-14B model augmented with the InternS1 dynamic tokenizer for chemistry/biomolecular text. No full fine-tune was performed; we adapted tokenization and warmed the new token rows so the model is ready for downstream SFT/RL on chemistry tasks.
Figure: Tokenization efficiency from the InternS1 technical report — SMILES/IUPAC compression ~2.64× vs ~1.44–1.51× for common LLM tokenizers.
Model Details
- Developed by: Summykai
- License: Apache-2.0
- Finetuned from: Qwen/Qwen3-14B
- Model type: QwenForCausalLM
- Languages: Multilingual (inherits from Qwen3)
What changed vs. base Qwen3-14B?
Dynamic tokenizer integration (InternS1)
We ship tokenization_interns1.py and three SentencePiece models:
tokenizer_SMILES.model(SMILES/SELFIES)tokenizer_IUPAC.model(IUPAC names)tokenizer_FASTA.model(protein sequences)
The tokenizer auto-detects long chemistry/biomolecular spans and switches to the SP model, or you can wrap with explicit tags:
<SMILES> CC(=O)OC1=CC=CC=C1C(=O)O </SMILES>
<IUPAC> 2-acetoxybenzoic acid </IUPAC>
Logical auto-detect tokens (e.g., <SMILES_AUTO_DETECT>…</SMILES_AUTO_DETECT>) do not consume ids.
Effective vocabulary size: 152,971.
Warm-up of new rows (no full finetune):
- Transformer layers frozen; trained only new embedding rows (+ matching
lm_headrows if untied). - Cross-entropy objective; ~1,500 steps; ~3,986,558 token exposures.
- Stability aids: fp32 CE, grad-clip=1.0, small LR.
Initialization strategy (better than random):
New rows were not left random. We projected domain embeddings into Qwen3’s space (ridge-regularized linear map) to give chemistry tokens meaningful initial locations. After merging, test_merge.py verified max |Δ| emb/head = 0.0.
How to Use
Install: RDKit is strongly recommended for high-quality SMILES auto-detection.
pip install rdkit-pypi sentencepiece transformers
# or via conda:
# conda install -c conda-forge rdkit
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "summykai/Qwen3-14B-chem-dyn-tokenizer"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, torch_dtype="auto")
prompt = "Name and SMILES for aspirin: <SMILES>CC(=O)OC1=CC=CC=C1C(=O)O</SMILES>"
out = model.generate(**tok(prompt, return_tensors="pt"), max_new_tokens=64)
print(tok.decode(out[0], skip_special_tokens=True))
Files
- Model:
model-*.safetensors,model.safetensors.index.json,config.json,generation_config.json - Tokenizer:
vocab.json,merges.txt,tokenizer_config.json,special_tokens_map.json,tokenization_interns1.py,tokenizer_SMILES.model,tokenizer_IUPAC.model,tokenizer_FASTA.model - Figures:
figures/compression_interns1.png
Acknowledgements
- Intern / Shanghai AI Lab — InternS1 tokenizer and report
- Qwen Team — Qwen3-14B base model
Contact
Summykai (HF & GitHub): https://github.com/Summykai
- Downloads last month
- 49