Unverified50% confidenceFactExact time
Terminal Bench专注衡量agentic编程能力,即模型自主在终端环境中规划、执行并验证多步骤任务的能力
1
Sources
50%
Confidence
Long-term
Relevance
9/25/2026
First Seen
Sources
Related Entities
Related Claims
UnverifiedTerminal-Bench 属于交互式Agent基准测试家族,与SWE-bench、WebArena等评测共同构成AI Agent能力的多维评估体系75% similarUnverifiedTerminal Agents can understand user intent, decompose it into multiple steps, execute them sequentially, and adjust the next step's strategy based on the previous step's output73% similarUnverifiedTerminal Agents introduce an LLM reasoning layer that enables them to adjust subsequent strategies based on real-time command output, a capability known in engineering as 'Reactive Planning'68% similarUnverifiedThe agent uses a tool called 'get_terminal_full_text' to read terminal output and make decisions based on current state67% similarUnverifiedTrae 的 Agent 能力可自主拆解任务、生成代码、执行终端命令并调试错误,形成规划→编码→验证的完整闭环67% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/949438API
curl https://kongchang.com/api/v1/knowledge/claims/949438MCP
get_claim(id=949438)