Unverified50% confidenceFactTime unknown
GPT 5.5 leads on TerminalBench 2.0 with a score of 87%, surpassing even Anthropic's unreleased Mythos model
1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
7/2/2026
First Seen
Valid until: 9/30/2026
Sources
Related Claims
UnverifiedGPT 5.5在TerminalBench 2.0基准测试中以87%的成绩大幅领先,超过了Anthropic尚未发布的Mythos模型84% similarUnverifiedGPT-5.5 与 Claude Mythos 5 在 Terminal Bench 2.1 上均停留在 88% 左右81% similarPartially VerifiedOpenAI 称 GPT-5.6 在 Terminal Bench 2.1 上达到当前最强水平80% similarUnverifiedGPT-5.2 Codex scores slightly higher than GPT-5.2 on SWBench Pro and about 2 percentage points higher on Terminal Bench.76% similarVerifiedGPT-5.4在SWE-Bench Pro测试中拿下57.7%准确率,超越了GPT-5.3-Codex的56.8%76% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/57265API
curl https://kongchang.com/api/v1/knowledge/claims/57265MCP
get_claim(id=57265)