待验证50% 置信事实精确时间
Gemini 2.5 Pro在全栈网站任务中视觉得分仅为11.7,功能分仅为22.6
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
清华智谱发布全栈Web开发基准:顶级AI模型集体翻车
bilibili跟我学一辈子AI吧2026/5/22
相关事实
待验证Gemini 2.5 Pro's visual score drops to 11.7 and functionality score drops to 22.6 on full-stack web development tasks.83% 相似待验证Gemini 2.5 Pro从静态网页到全栈任务的得分降幅超过60%78% 相似待验证Gemini 2.5 Pro scores approximately 63 on static web page generation tasks in the benchmark.76% 相似待验证Gemini 2.5 Pro experienced a cliff-like score drop from approximately 63 on static pages to 11.7 visual score on full-stack tasks, representing a drop of over 80%.73% 相似待验证Gemini 2.5 Pro在Humanities Last Exam基准测试中得分18.8%,远超O3 Mini的14%和其他模型的10%以下72% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/51768API
curl https://kongchang.com/api/v1/knowledge/claims/51768MCP
get_claim(id=51768)