Instructions to use yosefw/Qwen3-0.6B-DSpark-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yosefw/Qwen3-0.6B-DSpark-v2 with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("yosefw/Qwen3-0.6B-DSpark-v2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Qwen3-0.6B-DSpark-v2
A DSpark draft model for speculative decoding with Qwen/Qwen3-0.6B as the verifier, trained with speculators. The drafter proposes 4 tokens at a time and the verifier checks them in one forward pass, so output is identical to running the verifier alone — a lossless speedup.
Same architecture as yosefw/Qwen3-0.6B-DSpark, trained in offline mode: the verifier's hidden states were extracted to disk up front, freeing both GPUs for training.
Training code: rasyosef/train-dspark-draft-models.
Trained on 1,200 samples as a pipeline demonstration, not a deployment-ready drafter.
Usage
vLLM loads the verifier automatically from the config — don't pass it separately.
vllm serve yosefw/Qwen3-0.6B-DSpark-v2 --port 8000 --gpu-memory-utilization 0.75
Then query the OpenAI-compatible endpoint at http://localhost:8000/v1.
Details
5 Qwen3 layers (hidden size 1024, sliding-window attention), ~0.3B params, bfloat16. Block size 4, draft vocabulary reduced to 32,000, aux hidden-state layers 2/14/25, confidence head with Markov (rank 256).
Trained for 5 epochs at lr 3e-4 on 1,200 Magpie prompts regenerated by the verifier itself, with a {"ce": 0.1, "tv": 0.9} loss. Offline mode: a vLLM server in extract_hidden_states mode wrote activations to disk in a separate pass, then shut down before training began — so training ran data-parallel across both GPUs rather than sharing a card with the verifier. 2× T4, 5 h 53 min end to end, including the extraction pass.
Offline vs. online
The two modes differ only in where the verifier's hidden states come from, and produce equivalent drafters. Offline costs one extra extraction pass and a lot of disk, but frees the GPU for training — the better choice on a single card. Online avoids the disk overhead but keeps the verifier resident for the whole run.
| v2 (offline) | v1 (online) | |
|---|---|---|
| Samples | 1,200 | 1,600 |
| Training GPUs | 2 | 1 (verifier holds the other) |
| Wall clock (2× T4) | 5 h 53 min | 8 h 54 min |
Evaluation
Acceptance benchmarks have been published for v1 only, where the weighted acceptance length across the nine RedHatAI/speculator_benchmarks subsets was 1.806. Expect this checkpoint to land somewhat lower given the smaller training set. To measure it yourself, serve the model and run evaluate.py throughput from speculators.
Limitations
Trained on only 1,200 samples, so acceptance would improve considerably with more data. Works only with Qwen/Qwen3-0.6B and is not usable as a standalone model. Real-world speedup depends on your traffic mix — acceptance is highest where the verifier's next token is most predictable, such as code and math, and lowest on translation and summarization. Because verification is lossless, the verifier's own behavior and biases carry through unchanged.
License
Apache-2.0, matching both speculators and the verifier.
- Downloads last month
- 40