待验证50% 置信事实精确时间
Claude Opus 4.8在智能体实战跑分中拿下1890分,比上一代高出137分,甩开GPT 5.5达121分
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
Claude Opus 4.8代码能力登顶,却暴露AI训练致命隐患
bilibili智视界2026/6/23
相关事实
待验证在SABench基准测试中,Claude Opus 4.7得分约33分,而人类高级工程师得分80-90分76% 相似待验证Claude Opus 4.8 reached 85.9% on the GraphWalks 256K subset (up from 76.9% for 4.7) and 68.1% on the full 1-million-token version, nearly double Opus 4.7's score of 40.3%75% 相似部分验证Claude Opus 4.8在SWE-bench基准测试中得分69.2%,GPT 5.5为58.6%,Google Gemini为54.2%75% 相似待验证Opus 4在SWE-bench Verified上的通过率超过72%,显著领先于其他竞争模型74% 相似待验证Claude Opus 4.8 scored 1890 ELO on GDPVal AA, which is 137 points higher than Opus 4.7 and 121 points above GPT 5.574% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/40946API
curl https://kongchang.com/api/v1/knowledge/claims/40946MCP
get_claim(id=40946)