Unverified60% confidenceBenchmarkExact time
该测试中模型短上下文场景输出速度约20 token/秒,200K+上下文降至约13 token/秒
2
Sources
60%
Confidence
Long-term
Relevance
9/11/2026
First Seen
Sources
Related Entities
Related Claims
Unverified在输出(Decode)速度测试中,2 万 token 以内可达约 45 tokens/秒,3 万 token 以上骤降至约 20 tokens/秒79% similarUnverified本地模型推理速度约1-10 tokens/秒,而云端API可达50-150 tokens/秒77% similarUnverified启用MTP后推理速度波动显著,有时超过40 tokens/秒,有时降至25 tokens/秒73% similarUnverified采样频率越高轨迹越精细但续航越短,1秒采样通常比10秒采样续航缩短30-50%,长时间活动需在精度与续航间取舍72% similarUnverified微短剧将付费触点压缩到每集90-180秒69% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/897649API
curl https://kongchang.com/api/v1/knowledge/claims/897649MCP
get_claim(id=897649)