Unverified50% confidenceBenchmarkExact time
官方跑分显示 Qwen3.6 27B 稠密模型采用 NVFP4 量化时,单并发可达 202 token/s,双并发 400 token/s,8 并发冲到 1147 token/s
1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
9/16/2026
First Seen
Valid until: 12/15/2026
Sources
Related Entities
Related Claims
Unverifiedclaude-code-local 宣称支持 Qwen 3.5 122B(约 65 tokens/秒)、Llama 3.3 70B 和 Gemma 4 31B 等模型75% similarUnverified主流开源多模态模型上下文长度通常在4K到32K之间,LLaVA-1.5为4096 token,Qwen-VL为8192 token,InternVL2扩展到32K75% similarUnverified典型7B模型(32层、32头、128维度)在FP16精度下处理2048 token序列、批大小8时,KV Cache约消耗4GB显存75% similarUnverified据 FreeToken 技术论文,在 RTX 5090 系统上 Qwen3-235B-A3B 可达 77-83 tokens/s,DeepSeek 相关模型在 agent 轨迹上可达 22-25 tokens/s73% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/924349API
curl https://kongchang.com/api/v1/knowledge/claims/924349MCP
get_claim(id=924349)