[KongchangAI]
Unverified50% confidenceBenchmarkExact time

以 Llama-3 70B 为例,在 A100 80GB GPU 上处理 32K token 的预填充阶段约需 800-1200ms,而处理 2K token 仅需约 50ms

1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
7/6/2026
First Seen
Valid until: 10/4/2026

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/115482
API
curl https://kongchang.com/api/v1/knowledge/claims/115482
MCP
get_claim(id=115482)