待验证50% 置信事实精确时间
Terminal Bench专注衡量agentic编程能力,即模型自主在终端环境中规划、执行并验证多步骤任务的能力
1
来源数
50%
置信度
长期有效
时效性
2026/9/25
首次发现
来源
涉及实体
相关事实
待验证Terminal-Bench 属于交互式Agent基准测试家族,与SWE-bench、WebArena等评测共同构成AI Agent能力的多维评估体系75% 相似待验证Terminal Agents can understand user intent, decompose it into multiple steps, execute them sequentially, and adjust the next step's strategy based on the previous step's output73% 相似待验证Terminal Agents introduce an LLM reasoning layer that enables them to adjust subsequent strategies based on real-time command output, a capability known in engineering as 'Reactive Planning'68% 相似待验证The agent uses a tool called 'get_terminal_full_text' to read terminal output and make decisions based on current state67% 相似待验证Trae 的 Agent 能力可自主拆解任务、生成代码、执行终端命令并调试错误,形成规划→编码→验证的完整闭环67% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/949438API
curl https://kongchang.com/api/v1/knowledge/claims/949438MCP
get_claim(id=949438)