Unverified50% confidenceOpinionExact time
作者认为12-30 token/s是舒适区,超过30 token/s算好;48-64GB的M5 Max约能跑到20 token/s
1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
9/10/2026
First Seen
Valid until: 12/9/2026
Sources
Related Entities
Related Claims
Verified以LLaMA-2 70B为例,处理4096 token上下文、批量大小32时,KV Cache可消耗超过80GB显存71% similarUnverified200K token的上下文请求相比8K token请求,需要约25倍的KV-Cache显存69% similarUnverified实际工程中chunk_size在256至1024 token之间、chunk_overlap在20至100 token之间是较为常见的经验范围69% similarUnverified1 个 Token 约等于 0.75 个英文单词,现代模型已将上下文窗口上限扩展至 128K Token67% similarUnverified以 Llama-3 70B 为例,在 A100 80GB GPU 上处理 32K token 的预填充阶段约需 800-1200ms,而处理 2K token 仅需约 50ms67% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/887806API
curl https://kongchang.com/api/v1/knowledge/claims/887806MCP
get_claim(id=887806)