[KongchangAI]
BenchmarkDeepSeek Open Source Benchmark for Unified Evaluation / DeepStop UE

DeepStop UE

DeepSeek发布的开源编程智能体评估基准,涵盖代码生成、调试、多步骤任务执行等场景,核心特点是同时评估模型与工具链(harness)的整体协同表现,而非孤立测试模型语言能力

Timeline (last 90 days)

Sep 18

在使用同一checkpoint、同为max effort设置时,DeepSeek精简版harness在DeepStop UE上得分72.6,而OpenCode仅为65.5

Unverified50%
Sep 18

DeepStop UE(DeepSeek Open Source Benchmark for Unified Evaluation)是DeepSeek用于评估编程智能体能力的基准测试集,考察智能体与harness协同运作的整体表现

Unverified50%

All Facts (2)

Source Articles