待验证50% 置信事实精确时间
The four models tested in the AI coding comparison were ChatGPT 5.4, Gemini 3.1, DeepSeek V4 Pro, and Kimi 5.1
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
相关事实
待验证The five AI models tested were DeepSeek R1, Claude Sonnet 3.7, ChatGPT o3 Mini, Grok 3, and Qwen 2.5 Max.76% 相似待验证The study tested AI coding performance across models including Sonnet 4.5, GPT 5.2, GPT 5.1 Mini, and Qwen 376% 相似待验证Supabase团队在4个模型(Claude Code/Opus 4.6、Sonnet 4.6、GPT 5.4、GPT 5.4 Mini)上进行了6个场景的对比测试,Skill+MCP组合在每个模型上都优于其他条件76% 相似已验证GPT-4、Claude 3.5 Sonnet、Gemini 1.5 Pro等新一代大语言模型在HumanEval、SWE-bench等代码能力基准测试上的得分已大幅超越此前版本74% 相似待验证AI aggregator platforms support access to models including GPT 5.5 Thinking, Grok 4.2, Claude 4.7, Gemini 3.1 Pro, and DeepSeek V4.73% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/48439API
curl https://kongchang.com/api/v1/knowledge/claims/48439MCP
get_claim(id=48439)