待验证50% 置信事实精确时间
Simon Willison长期使用'鹈鹕骑自行车'作为固定测试用例来横向对比不同AI模型在不同时期的表现
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/9/9
首次发现
有效期至:2026/12/8
来源
涉及实体
相关事实
待验证「画一只骑自行车的鹈鹕」测试最初由开发者Simon Willison提出,用来评估大语言模型和图像生成模型的能力边界78% 相似待验证Simon Willison使用'鹈鹕骑自行车'的SVG生成作为非正式的模型能力基准测试73% 相似待验证测试设计方法论(等价类、边界值、状态迁移、因果图等)和质量风险评估能力是AI时代测试工程师不可替代的核心竞争力59% 相似待验证A Bilibili content creator named Jiuling conducted a head-to-head comparison test of five AI models for game development.58% 相似待验证Industry estimates suggest that over 95% of testing miles in autonomous driving development rely on simulation environments58% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/884107API
curl https://kongchang.com/api/v1/knowledge/claims/884107MCP
get_claim(id=884107)