Unverified75% confidenceFactTime unknown
GPT 5.5在TerminalBench 2.0基准测试中以87%的成绩大幅领先,超过了Anthropic尚未发布的Mythos模型
1
Sources
75%
Confidence
Medium-term (~90 days)
Relevance
5/31/2026
First Seen
Valid until: 8/29/2026
Sources
GPT 5.5 vs Claude Code vs DeepSeek V4:三大编码模型实测对比
bilibili量子菠萝_
Related Entities
Related Claims
UnverifiedGPT 5.5 leads on TerminalBench 2.0 with a score of 87%, surpassing even Anthropic's unreleased Mythos model84% similarUnverified在DeepSWE基准测试中,GPT 5.5以约70%的通过率高居榜首,比Opus 4.7领先超过15个百分点80% similarUnverifiedGPT-5在HumanEval、SWE-bench等主流编程基准测试上相比前代模型表现大幅提升80% similarVerifiedGPT-5.4在SWE-Bench Pro测试中拿下57.7%准确率,超越了GPT-5.3-Codex的56.8%79% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/2464API
curl https://kongchang.com/api/v1/knowledge/claims/2464MCP
get_claim(id=2464)