DeepSeek open-sources DeepSpec, delivers up to 85% faster inference for V4
DeepSeek, in collaboration with Peking University, has published a technical report on DSpark — a speculative decoding acceleration framework — and open-sourced the full-stack codebase under the name DeepSpec. DSpark is now deployed in production for DeepSeek-V4. Without any output quality degradation, DSpark boosts single-user generation speed by 60% to 85% for the Flash version and 57% to 78% for the Pro version. It outperforms the previous single-token multi-branch prediction (MTP-1) baseline and significantly raises overall system throughput under strict latency constraints.
Previously, multi-token speculative decoding was difficult to deploy in production. Autoregressive draft models were too slow, while parallel draft models suffered from low acceptance rates in the latter half of long sequences because each position predicted independently. Blindly verifying multi-token drafts under high concurrency caused the large model to waste massive compute on tokens destined for rejection, crashing system throughput. That's why the industry largely stuck with single-token prediction (MTP-1) online.
DSpark overcomes the throughput degradation bottleneck under high concurrency. It first uses the DFlash parallel backbone to generate hidden states, then adds an extremely lightweight Markov head. The Markov head injects neighbor-token dependencies at trivial cost via a lookup table and a single matrix multiplication. The system also integrates a confidence prediction head and a posterior calibration algorithm. To achieve zero-overhead scheduling in production and prevent future information leakage, the scheduler uses an asynchronous mechanism that dynamically decides the candidate token pruning length based on predictions from two steps earlier. This prevents the large model from validating high-risk tail tokens under heavy load.
Beyond DSpark, the open-sourced DeepSpec codebase natively supports open-source models like Qwen3 and Gemma. It provides a complete Python toolchain from downloading prompts and rebuilding model caches to training draft models and running benchmarks. Developers can use the included scripts to customize and deploy dedicated acceleration modules for different open-source models locally.