Unverified50% confidenceBenchmarkExact time
Anthropic于2024年10月率先发布了Claude的Computer Use功能,其论文披露在OSWorld基准上成功率约22%,显著低于人类的72%
1
Sources
50%
Confidence
Long-term
Relevance
7/20/2026
First Seen
Sources
GPT-5.6与OpenAI超级App深度实测:Loop工程替代Prompt工程的新范式
bilibiliGoldenSpiderAI7/11/2026
Related Claims
UnverifiedAnthropic的Computer Use在OSWorld基准测试上的操作准确率约为14.9%(2024年10月数据)68% similarUnverifiedClaude第一次推出计算机使用能力时,OS World得分还不到30%66% similarUnverifiedAnthropic在2023年的论文中系统测量了奉承偏差,发现即便在明显错误前提下多数主流模型仍有相当比例的回复选择顺从用户57% similarUnverified在OS World电脑操作测试中,Sol超过了Claude的Opus系列,消耗算力少85%56% similarUnverifiedAnthropic在推出Claude 4.5 Sonnet的同时,大幅削减了200美元Max套餐用户使用Opus模型的额度,使用限制骤降约20倍56% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/566790API
curl https://kongchang.com/api/v1/knowledge/claims/566790MCP
get_claim(id=566790)