AI圈报
产品发布 / 更新精选

SGLang 集成 DSpark 推测解码:置信度驱动的可变长度验证

信息来源:LMSYS:Blog(Chatbot Arena 团队)·
原始标题:Blog DSpark in SGLang: Speculative Decoding with Confidence-Driven, Variable-Length Verification Speculative decoding trades extra compute for fewer decode steps, and the trade sours as load grows: at batch size B with K speculative tokens the target verifies B K tokens every step, and past a po… SGLang Team

内容摘要

SGLang 团队将 DSpark 推测解码算法集成到开源推理引擎中。该算法采用半自回归块起草器一次生成一组 token,并利用置信度头与顺序温度缩放(STS)为每个请求动态分配可变验证长度,从而在高负载下裁剪无效验证成本。SGLang 支持密集模型(如 Qwen3)和稀疏模型(如 DeepSeek-V4),通过全 CUDA 图处理不规则的每请求验证长度。提供三种验证模式:`static`(全长)、`compact`(生产路径)和 `cap-accept`(接受上限测量)。还引入了零开销调度、基于离线成本表的在线调度器、融合 Triton 核等优化。在 H200 上使用 DeepSeek-V4-Flash 的测试中,DSpark 在整个并发扫描范围内比 MTP 和非推测基线实现了更优的吞吐量-延迟权衡。
内容层级精选情报
发布时间(北京时间)
本站收录时间(北京时间)
信息来源LMSYS:Blog(Chatbot Arena 团队)
站内情报编号intel-2bf15066c56008020984c419