Update README.md to add vLLM snippet

#2
Files changed (1) hide show
  1. README.md +18 -0
README.md CHANGED
@@ -217,6 +217,24 @@ tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-
217
  tokens = tokenizer.encode(prompt)
218
  ```
219
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
220
  ## How to Run Locally
221
 
222
  Please refer to the [inference](inference/README.md) folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.
 
217
  tokens = tokenizer.encode(prompt)
218
  ```
219
 
220
+ ## How to Run with vLLM
221
+
222
+ DSpark speculative decoding is enabled with a single flag — add --speculative-config with method: dspark to your vLLM launch command:
223
+
224
+ `--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'`
225
+
226
+ For example, the command below serves the model with vLLM on a single 4×GB300 node.
227
+ See the [vLLM recipe](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash?hardware=b300&features=tool_calling,reasoning) for detailed instructions and other hardware configurations.
228
+
229
+ ```bash
230
+ vllm serve deepseek-ai/DeepSeek-V4-Flash-DSpark \
231
+ --trust-remote-code --kv-cache-dtype fp8 --block-size 256 \
232
+ --data-parallel-size 4 --enable-expert-parallel \
233
+ --moe-backend deep_gemm_mega_moe \
234
+ --attention-config '{"use_fp4_indexer_cache": true}' \
235
+ --speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'
236
+ ```
237
+
238
  ## How to Run Locally
239
 
240
  Please refer to the [inference](inference/README.md) folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.