[KongchangAI]
Benchmark

BrowseComp-Plus

深度研究能力评估基准,用于测试AI智能体在复杂信息检索与综合分析任务上的表现,是BrowseComp的扩展版本

Timeline (last 90 days)

Oct 5

在BrowseComp-Plus基准上,CLM配合Qwen3 27B取得59.4%,比最好的摘要方案高约11%,计算量少约五分之一

Unverified50%

All Facts (1)

Source Articles