Verified70% confidenceBenchmarkExact time
在输出(Decode)速度测试中,2 万 token 以内可达约 45 tokens/秒,3 万 token 以上骤降至约 20 tokens/秒
4
Sources
70%
Confidence
Medium-term (~90 days)
Relevance
7/22/2026
First Seen
Valid until: 10/20/2026
Sources
Related Claims
Unverified该测试中模型短上下文场景输出速度约20 token/秒,200K+上下文降至约13 token/秒79% similarUnverified通常认为单并发 100 tokens/秒以上即可实现流畅的打字机效果,约 200 tokens/秒生成 2000 字长文本仅需约 10 秒79% similarUnverified启用MTP后推理速度波动显著,有时超过40 tokens/秒,有时降至25 tokens/秒79% similarUnverified本地模型推理速度约1-10 tokens/秒,而云端API可达50-150 tokens/秒77% similarUnverifiedA typical 30-minute coding session can consume 118,000 Tokens on terminal command output alone69% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/585468API
curl https://kongchang.com/api/v1/knowledge/claims/585468MCP
get_claim(id=585468)