基准WebArena Benchmark
WebArena
面向AI智能体任务评估的基准测试环境,用于标准化智能体在网页交互等真实世界任务中的评估框架
时间轴 (近 90 天)
8月4日
Stanford University's WebArena benchmark, released in 2023, showed that even the most advanced AI models achieve only about a 14% success rate on complex operational tasks on real web pages
待验证80%