Question for the DeepSeek team

#52
by madeby561 - opened

Why does the DeepSeek-V4-Flash-0731 model card recommend greedy num_speculative_tokens: 7 when its released DSpark head appears optimized for five positions and underfit for high/max reasoning?

DeepSpec's released Qwen/Gemma block-7 heads were trained on non-thinking outputs, and its README recommends refitting for thinking mode. In my full GSM8K run (1,319 prompts, C32, greedy), DeepSeek-V4-Flash-0731 K7 reached 4.98 inclusive acceptance versus K5's 4.52; under high/max reasoning, K7 fell to 4.11/3.94. Meanwhile, DeepSpec's public Qwen3-4B block-7 head reaches the expected 5-6 range in the same vLLM runtime.

Could you clarify the DeepSeek-V4-Flash-0731 head's training block size, reasoning-mode distribution, and evidence behind the K7 recommendation - or provide a K7 head aligned to the official high/max prompt contract?

Can You create a version Flash for https://github.com/giannisanni/pulsar or https://github.com/FareedKhan-dev/kimi-k3-in-c
for people with small hardware. Slow version but for agents and long answered. For grok build, claude cli or pi agents https://pi.dev/ terminal

Please tell me how big is my language in deepseek? I use Polish Language (meybe other too)
How many data in my language is inside DeepSeek

Sign up or log in to comment