待验证50% 置信事实精确时间
Gemini 2.5 Pro experienced a cliff-like score drop from approximately 63 on static pages to 11.7 visual score on full-stack tasks, representing a drop of over 80%.
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
Tsinghua and Zhipu Release Full-Stack Web Dev Benchmark: Top AI Models Fail Spectacularly
bilibili跟我学一辈子AI吧2026/5/22
相关事实
待验证Gemini 2.5 Pro's visual score drops to 11.7 and functionality score drops to 22.6 on full-stack web development tasks.84% 相似待验证Gemini 2.5 Pro从静态网页到全栈任务的得分降幅超过60%78% 相似待验证Gemini 2.5 Pro在全栈网站任务中视觉得分仅为11.7,功能分仅为22.673% 相似待验证Gemini 2.5 Pro scores approximately 63 on static web page generation tasks in the benchmark.67% 相似待验证Gemini 2在Google的SWE Bench基准测试中得分80.6%67% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/50421API
curl https://kongchang.com/api/v1/knowledge/claims/50421MCP
get_claim(id=50421)