Expired75% confidenceFactTime unknown
两款模型的最终测试总比分为平局
1
Sources
75%
Confidence
Medium-term (~90 days)
Relevance
5/31/2026
First Seen
Valid until: 8/29/2026(expired)
Sources
Gemini 3.1 Pro vs Opus 4.6:前端编程能力实测对比
bilibili伊莱文思帕
Related Entities
Related Claims
UnverifiedThe tester assessed that O3 and Gemini 2.5 Pro performed the best in the first round of the snake battle game test.58% similarUnverified在Agent Test: Last Exam测试中,Soul得分53.6分,Anthropic的Opus 4.5得分40.5分54% similarUnverifiedtop-1%准确率是评估量化模型与原始全精度模型输出一致性的指标,关注模型置信度最高的1%的token是否与全精度版本一致53% similarUnverifiedDeepSeek V4 Pro在Humanity's Last Exam基准测试中得分37.753% similarUnverifiedLFM2.5在多项基准测试中展现出与体量高达其4倍模型相当的竞争力53% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/9566API
curl https://kongchang.com/api/v1/knowledge/claims/9566MCP
get_claim(id=9566)