[KongchangAI]
Unverified50% confidenceBenchmarkExact time

在显存充足、模型完整驻留GPU时,RTX 4090跑Q4量化的27B模型通常可达15~30 tokens/s;触发CPU卸载后速度可能骤降至1~2 tokens/s甚至更低

1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
7/12/2026
First Seen
Valid until: 10/10/2026

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/489738
API
curl https://kongchang.com/api/v1/knowledge/claims/489738
MCP
get_claim(id=489738)