5 related articles

Deep dive into DFlash 2's parallel draft decoding technology, explaining how its Keep Drafting Parallel mechanism breaks autoregressive bottlenecks for lossless LLM inference acceleration.

In-depth review of Poolside's Laguna S 2.1 open-source coding model: MoE architecture, RL training, DGX Spark local deployment, and real-world agentic coding tests with 8B active parameters.

DeepSeek partners with Peking University to open-source DSpark, an inference acceleration tech boosting single-user speed by 57%-85% under high concurrency. Learn its three core designs and the DSpec framework.

DeepSeek and Peking University open-source DSpark, an inference acceleration tech boosting single-user generation speed by 57%-85% under high concurrency. Learn its 3 core designs and the DSpec framework.
Product ReviewsReal-world test of Qwen 3.6 Multi-Token Prediction (MTP): boost inference speed from 34.2 to 41 tokens/s with just three parameters in ik_llama.cpp — zero quality loss, zero extra models.