待验证50% 置信基准精确时间
Opus 5.5 并非全线第一,在 Automation Bench 上以 40 略低于 Astra 的 41.4,在 Terminal Bench Science 上落后于 GPT-6 Astra
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/9/25
首次发现
有效期至:2026/12/24
来源
涉及实体
相关事实
待验证在 Terminal Bench 4.0 上,Opus 5.5 得分 66.4%,Astra 为 57.9%,Fable 5.1 为 55.8%,前代 Opus 5 为 52.3%75% 相似待验证Opus 5.5在token效率上不如Opus 5:Opus 5的max档每任务用73k token,Fable用78k,Opus 5.5的max档接近12万token,而GPT-6 Astra只需27k72% 相似待验证Terminal Bench 2.1测试中GPT-5.6 Ultra得分接近92%,Fable 5为88%,考虑误差范围基本持平65% 相似待验证在 Automation Bench 测试中,GPT-6 Sol 的 extra high 档拿下 33% 的第一名,而 5.6 的 max effort 得分为 28.865% 相似待验证Opus 5在Frontier-Bench上的表现超过Opus 4.8的两倍63% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/949381API
curl https://kongchang.com/api/v1/knowledge/claims/949381MCP
get_claim(id=949381)