待验证50% 置信基准精确时间
在SWE Bench Pro等智能体编程基准中,Qwen 3.8据称在部分项目上超越了Opus 4.6 Max
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/9/16
首次发现
有效期至:2026/12/15
来源
涉及实体
相关事实
待验证Qwen3 27B在多项编程和智能体基准测试中的得分超越了Claude Opus 4.676% 相似待验证Qwen 3.6's performance in general programming capability, development skills, multi-turn agent-enhanced capability, and agent task testing leads Qwen 3.5 in many aspects and approaches or exceeds Claude 4.5 Opus on certain metrics.75% 相似待验证阿里官方给出Qwen3.8 Flash的SWE-Bench Pro得62.5,领先Claude Opus 4.6约9分;Cowork Bench 73.9,DeepSeek为45.1;Job Bench 55.7,超出Opus 4.6约20分75% 相似待验证微调后的Qwen3.5-4B模型在BU Bench V1基准上准确率从15%提升至53%74% 相似待验证测评者认为Qwen 3.8 Max在同一滑板游戏prompt下的表现与Fable 5.1接近73% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/925177API
curl https://kongchang.com/api/v1/knowledge/claims/925177MCP
get_claim(id=925177)