Instructions to use kshitijthakkar/deepseek-v4-mini-1B-init with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kshitijthakkar/deepseek-v4-mini-1B-init with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="kshitijthakkar/deepseek-v4-mini-1B-init") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("kshitijthakkar/deepseek-v4-mini-1B-init") model = AutoModelForCausalLM.from_pretrained("kshitijthakkar/deepseek-v4-mini-1B-init", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use kshitijthakkar/deepseek-v4-mini-1B-init with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kshitijthakkar/deepseek-v4-mini-1B-init" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kshitijthakkar/deepseek-v4-mini-1B-init", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kshitijthakkar/deepseek-v4-mini-1B-init
- SGLang
How to use kshitijthakkar/deepseek-v4-mini-1B-init with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kshitijthakkar/deepseek-v4-mini-1B-init" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kshitijthakkar/deepseek-v4-mini-1B-init", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kshitijthakkar/deepseek-v4-mini-1B-init" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kshitijthakkar/deepseek-v4-mini-1B-init", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use kshitijthakkar/deepseek-v4-mini-1B-init with Docker Model Runner:
docker model run hf.co/kshitijthakkar/deepseek-v4-mini-1B-init
The pretraining loss convergence is not ideal.
I downloaded this checkpoint and continued pretraining the model on the FineWeb-Edu 10B dataset. However, midway through training, the loss had already started to plateau, and I found that the model’s text continuation ability was still very poor.
May I ask whether you have tried pretraining this 1B model yourself? What would you recommend I do to train the model more effectively?
The results of loss:
Epoch: 1/1 | Step: 1/870155 | Global Loss: 15.5567 | PPL: 5703910.5530 | LR: 0.000001 | Seconds/Iteration: 3.4731 | Remaining time: 34 days, 23:28:44.965052
Epoch: 1/1 | Step: 51/870155 | Global Loss: 13.7177 | PPL: 906830.1946 | LR: 0.000041 | Seconds/Iteration: 1.2446 | Remaining time: 12 days, 12:49:05.277774
Epoch: 1/1 | Step: 101/870155 | Global Loss: 10.5906 | PPL: 39758.6556 | LR: 0.000081 | Seconds/Iteration: 1.1923 | Remaining time: 12 days, 0:09:42.542777
Deleted old checkpoint: /data3/checkpoints/checkpoint-75000
Checkpoint saved to: /data3/checkpoints/checkpoint-78000
Epoch: 1/1 | Step: 78001/870155 | Global Loss: 6.0688 | PPL: 432.1412 | LR: 0.000388 | Seconds/Iteration: 1.1957 | Remaining time: 10 days, 23:05:40.997470
Epoch: 1/1 | Step: 78051/870155 | Global Loss: 5.6854 | PPL: 294.5441 | LR: 0.000388 | Seconds/Iteration: 1.1956 | Remaining time: 10 days, 23:04:37.693189
Epoch: 1/1 | Step: 78101/870155 | Global Loss: 5.7766 | PPL: 322.6575 | LR: 0.000388 | Seconds/Iteration: 1.1956 | Remaining time: 10 days, 23:03:30.830656
Deleted old checkpoint: /data3/checkpoints/checkpoint-144000
Checkpoint saved to: /data3/checkpoints/checkpoint-147000
Epoch: 1/1 | Step: 147001/870155 | Global Loss: 5.4705 | PPL: 237.5887 | LR: 0.000358 | Seconds/Iteration: 1.1917 | Remaining time: 9 days, 23:22:37.850622
Epoch: 1/1 | Step: 147051/870155 | Global Loss: 5.5180 | PPL: 249.1469 | LR: 0.000358 | Seconds/Iteration: 1.1917 | Remaining time: 9 days, 23:21:34.257241
Epoch: 1/1 | Step: 147101/870155 | Global Loss: 5.3619 | PPL: 213.1350 | LR: 0.000358 | Seconds/Iteration: 1.1917 | Remaining time: 9 days, 23:20:28.865552
The results of the text continuation test are as follows:
python run_inference.py /home/xxx/dl/llm_train/mini-llm-dsv4/checkpoint/checkpoint-131000 "The history of Rome"
The history of Rome, which, the of the a 201 the of the, and, the. But did not a lot of the the state. The name, and, but the the best method of a small (1998-2, the the to the. This is a large, and, and, a person is not to the a number of their health.
“ a lot of the, in a lot that the, the same as well the, and the. It is important to the to the of the same question to the a great power supplies of the, the best time to the to the time of which the in the, the in the a littlerobert. In for a littleo, and. It’s. But did not know to the most common problem, which. 1989 50-6, and the. When the an increase in the a major problems, the of the same as they are used. These include for example, and the a variety, ands the to be and. A company, 5 of the the, a significant impact on a single man the the the an additional.
Hi @qingshuiL — thanks for the detailed report. We reproduced exactly what you saw, and the plateau has a specific, fixable cause.
Root cause — it's the Lightning Indexer, not your data/LR. In the CSA layers the indexer picks which compressed KV blocks each query attends to via a hard top-k. That selection is non-differentiable (integer indices), so the indexer's weights get zero gradient from the language-modeling loss. In the -init checkpoint (and in the from-flash slices) the indexer is randomly initialized — so on every CSA layer the model is permanently forced to attend to randomly-chosen blocks, and no amount of LM-only training fixes it. That's the ~5.x loss wall and the incoherent text.
The fix is the DeepSeek-V3.2 DSA recipe (arXiv:2512.02556, §2.1.1), in three stages:
- Dense recovery — run the CSA layers densely (attend all compressed blocks, indexer bypassed) and train normally. This recovers the backbone + compressor, which are reachable by the LM gradient.
- Indexer warm-up — freeze the backbone and train only the indexer with a KL loss distilling the dense attention into it:
KL(p_main || softmax(index_scores)), summed over heads and L1-normalized. This is the step the LM loss cannot do. - Sparse stage — enable top-k; train the backbone with LM loss while the indexer keeps a selected-set KL, with the indexer input detached so the two losses don't interfere.
Two practical notes:
- The
-from-flashvariants recover far faster than-init—87% of weights are transferred from DeepSeek-V4-Flash, so dense recovery drops the loss steeply. We took2.8**, after which the sparse path matches the dense loss (i.e. the indexer warm-up worked). For your 1B use-case, start fromdeepseek-v4-mini-300M-from-flashfrom the ~5.x wall down to **kshitijthakkar/deepseek-v4-mini-1B-from-flashrather than-init. - If your context is longer than ~1k tokens, scale
index_topkup so top-k covers a reasonable fraction of the compressed blocks.
Recovered checkpoint as a starting point: kshitijthakkar/deepseek-v4-mini-300M-recovered. Happy to share the three-stage trainer script if it helps.
Thank you so much for the detailed explanation and for taking the time to reproduce and analyze the issue.
I’m just starting to experiment with LLM training, so there are still many training-related details that I don’t fully understand yet. I really appreciate your patience in explaining the root cause and the recovery recipe so clearly.
If you could share the three-stage trainer script, that would be extremely helpful for me to better understand and try the recovery process on my 1B model.
Thanks again for all your help!
Happy to help — I've packaged the three-stage trainer as a recovery kit in the model repo:
https://huggingface.co/kshitijthakkar/deepseek-v4-mini-300M-recovered/tree/main/recovery_kit
It contains:
train_indexer.py— the three-phase trainer (Phase 0 dense recovery → Phase A indexer KL warm-up → Phase B sparse).deepseek_v4/— the modeling code with the indexer-training hooks (set_csa_train_mode,raw_scores,_compute_indexer_kl, and the dense-CSA mode). This is the load-bearing part — the stock modeling can't do the warm-up.optimizer.py— the Muon/AdamW split + LR schedulers (WSD by default).README.md— the recipe + a runnable command.
For your 1B, point --model at kshitijthakkar/deepseek-v4-mini-1B-from-flash — the from-flash variant recovers far faster than -init.
The dataset mix we used is public now too: https://huggingface.co/datasets/kshitijthakkar/loggenix_moe_sequelbox_sft_v1 (the trainer re-tokenizes from its text field to the model's vocab=129280).
Two reminders from the README: keep gradient checkpointing off in Phases A/B (the indexer KL needs the live grad pass), and raise --index-topk if you train at >1k context (e.g. 128 at seq 4096). Would love to hear how the 1B recovery goes.
Thanks a lot for packaging and sharing the recovery kit — it’s been very helpful. I’ve started testing it on the 1B deepseek-v4-mini-1B-from-flash model and wanted to share some initial results and observations.
Below are some of my training logs:
=== Before training (sanity sample) ===
PROMPT: The history of Rome
CONT: tritritri步枪 Xml Xml passing passingственные passing� passing� passing�instead psoriasis psoriasis psoriasis psoriasis� psoriasisственные� psoriasis psoriasis psoriasis psoriasasis psoriasis psoriasis passingственныественныественныественныественныественные
------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
=== Phase 0 (mode=dense) | csa_layers=11 | trainable=1020.6M | lr=3.0e-04 | steps=3000 | grad_ckpt=False | world_size=4 ===
[0] step 1/3000 | lm 11.9474 | kl 0.0000 | lr 2.00e-06 | 448 tok/s
[0] step 20/3000 | lm 11.8737 | kl 0.0000 | lr 4.00e-05 | 1951 tok/s
[0] step 40/3000 | lm 11.5871 | kl 0.0000 | lr 8.00e-05 | 2205 tok/s
[0] step 160/3000 | lm 6.4218 | kl 0.0000 | lr 3.00e-04 | 2494 tok/s
[0] step 180/3000 | lm 5.0812 | kl 0.0000 | lr 3.00e-04 | 2500 tok/s
[0] step 200/3000 | lm 5.1837 | kl 0.0000 | lr 3.00e-04 | 2509 tok/s
[0] step 2960/3000 | lm 2.9708 | kl 0.0000 | lr 1.15e-06 | 2527 tok/s
[0] step 2980/3000 | lm 3.0316 | kl 0.0000 | lr 1.04e-06 | 2527 tok/s
[0] step 3000/3000 | lm 3.2964 | kl 0.0000 | lr 1.00e-06 | 2527 tok/s
=== Phase A (mode=warmup) | csa_layers=11 | trainable=5.1M | lr=1.0e-03 | steps=500 | grad_ckpt=False ===
[A] step 1/500 | lm 0.0000 | kl 0.1740 | lr 4.00e-05 | 1714 tok/s
[A] step 20/500 | lm 0.0000 | kl 0.1267 | lr 8.00e-04 | 4987 tok/s
[A] step 40/500 | lm 0.0000 | kl 0.0645 | lr 9.98e-04 | 5324 tok/s
[A] step 460/500 | lm 0.0000 | kl 0.0104 | lr 1.84e-05 | 5717 tok/s
[A] step 480/500 | lm 0.0000 | kl 0.0101 | lr 5.36e-06 | 5717 tok/s
[A] step 500/500 | lm 0.0000 | kl 0.0115 | lr 1.00e-06 | 5716 tok/s
PROMPT: The history of Rome
CONT: . no two public public access. public public public public public public public public public public public public public public public public public public public public public public public public public public public public public
=== Phase B (mode=sparse) | csa_layers=11 | trainable=1020.6M | lr=1.0e-04 | steps=2000 | grad_ckpt=False ===
[B] step 1/2000 | lm 2.1049 | kl 0.0111 | lr 1.00e-06 | 617 tok/s
[B] step 20/2000 | lm 1.8056 | kl 0.0104 | lr 2.00e-05 | 1436 tok/s
[B] step 40/2000 | lm 2.0077 | kl 0.0106 | lr 4.00e-05 | 1542 tok/s
[B] step 1960/2000 | lm 1.6480 | kl 0.0114 | lr 1.11e-06 | 1662 tok/s
[B] step 1980/2000 | lm 1.7601 | kl 0.0117 | lr 1.03e-06 | 1662 tok/s
[B] step 2000/2000 | lm 1.9777 | kl 0.0126 | lr 1.00e-06 | 1662 tok/s
PROMPT: The history of Rome
CONT:
I) reasoning is a set of public keys. Let's recall the structure of a set of rules for a specific type of security group (like a set of rules) and a set of
------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
From the logs, the loss does seem to converge, and the continuation quality has improved somewhat. However, the model is still not really usable for text continuation.
My current understanding is that this may be due to insufficient training data, and also because the CSA/HCA/Lightning Index components in DeepSeek-V4 are relatively hard to recover/train. Do you think there could be other major factors, such as tokenizer/template mismatch, insufficient dense recovery, index-topk settings, or the model still being undertrained overall?
This is great — your numbers show the recipe worked on the 1B: Phase 0 11.94→3.30, Phase A KL 0.174→0.0115 (the indexer warmed up beautifully), and Phase B 2.10→1.98 (even lower than our 300M). So the CSA/HCA/indexer components recovered fine — those aren't your problem, and neither is index_topk (a KL ~0.01 means the indexer is well-calibrated).
The "not usable for continuation" gap is two things, in order of impact:
1. You're evaluating a chat model as a base model. The dataset (loggenix) is ChatML-formatted (system/user/assistant), so the model learned to operate in chat format, not raw continuation. Prompt it with the template:
msgs = [{"role": "user", "content": "Explain the history of Rome."}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")
Your sample ("...reasoning is a set of public keys... structure of a set of rules...") is the model echoing loggenix's reasoning/agentic style — i.e. it learned that distribution well.
2. Data breadth, not architecture. loggenix is a narrow synthetic SFT mix (~277M tokens), so the model is specialized to it. For general usability, run Phase 0 (dense recovery) on broad text — your FineWeb-Edu is ideal — to build general LM ability, then SFT on task data. Loss 1.98 means it fits loggenix; it doesn't mean it's a general model.
Two smaller levers:
- Push Phase 0 further — 3.30 at 3k steps was likely still dropping (our 300M reached ~2.8). A stronger dense backbone gives Phase B more to work with.
- It's a 1B, so it needs more total tokens than our 300M for the same fluency — "undertrained" is fair; the fix is more/broader Phase-0 data, not architecture changes.
tl;dr: the recovery worked (your sparse loss is excellent). The remaining gap is (a) eval it as a chat model, and (b) widen the Phase-0 data beyond loggenix.
Thanks a lot for your detailed feedback and guidance — this is very helpful.
I’ll go back and try evaluating it with the proper chat template, and also experiment with broader Phase-0 data as you suggested. I’ll report back once I have new results.
Thanks again for your help!