DeepSeek open-sources DeepSpec, claims 85% inference speed boost
Today, DeepSeek, in collaboration with Peking University, released a technical report on DSpark, a speculative sampling acceleration framework, and open-sourced the full-stack codebase under the name DeepSpec. DSpark is already deployed in DeepSeek-V4's production system. According to the team, it boosts single-user generation speed by 60 to 85 percent for the Flash edition and 57 to 78 percent for the Pro edition — all without any loss in output quality. DSpark outperforms the previous single-token multi-branch prediction (MTP-1) baseline and significantly improves overall system throughput under strict latency constraints.
Previously, multi-token speculative sampling struggled to land in production environments. Autoregressive draft models were too slow, while parallel draft models — which predict each token independently — suffered from extremely low acceptance rates in the latter half of long sequences. Blindly verifying multi-token drafts under high concurrency caused large models to waste massive compute on tokens that were certain to be rejected, crashing overall system throughput. As a result, the industry largely stuck with single-token prediction (MTP-1) in online settings.
DSpark overcomes this throughput degradation under high concurrency. It first uses DFlash, a parallel backbone network, to generate hidden states, then adds an extremely lightweight Markov head. The Markov head uses a lookup table and a single matrix multiplication to inject correlations between adjacent tokens at minimal cost. The system also integrates a confidence prediction head and a posterior calibration algorithm. To ensure zero-overhead scheduling in production and prevent future information leakage, the scheduler uses an asynchronous mechanism that dynamically determines candidate token pruning length based on predictions from two steps prior. This fundamentally prevents the large model from verifying high-risk tail tokens under heavy load.
Beyond DSpark, the open-source DeepSpec codebase natively supports other open-source large models like Qwen3 and Gemma. It provides a complete Python toolchain — from downloading prompts, rebuilding large model caches, training draft models, to benchmark evaluation. Developers can use the provided scripts to customize and deploy dedicated acceleration modules for different open-source models locally.