Unverified50% confidenceFactExact time
在DeepSWE基准测试中,GPT 5.5以约70%的通过率高居榜首,比Opus 4.7领先超过15个百分点
1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
7/2/2026
First Seen
Valid until: 9/30/2026
Sources
DeepSWE基准测试揭示真相:GPT 5.5大幅领先Opus 4.7
bilibili7k的每日搬运5/29/2026
Related Claims
UnverifiedGPT-5.6编程能力在关键基准上达到91.9%,安全评测CyberGym得分85.6%,均超越竞品Nexus 581% similarUnverifiedGPT 5.5在TerminalBench 2.0基准测试中以87%的成绩大幅领先,超过了Anthropic尚未发布的Mythos模型80% similarVerifiedGPT-5.4在SWE-Bench Pro测试中拿下57.7%准确率,超越了GPT-5.3-Codex的56.8%79% similarPartially Verified在同等任务上,GPT-5.5使用的输出token比Opus少约72%78% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/47581API
curl https://kongchang.com/api/v1/knowledge/claims/47581MCP
get_claim(id=47581)