待验证50% 置信事实精确时间
Computer-Use Agent 通过屏幕截图识别界面元素、模拟鼠标点击和键盘输入直接操控图形用户界面,模拟人类的视觉-操作行为
1
来源数
50%
置信度
长期有效
时效性
2026/9/11
首次发现
来源
涉及实体
相关事实
待验证Computer Use通过将屏幕截图作为视觉输入,模型识别UI元素后输出鼠标点击或键盘操作指令,与依赖DOM结构的Selenium、Playwright有本质区别75% 相似待验证屏幕视觉识别+模拟输入方案从操作系统层面看与人类玩家行为几乎无异,是个人开发者最常用的游戏自动化路线74% 相似待验证Anthropic's Computer Use employs a 'screenshot-reason-execute' loop where AI captures the current screen, feeds the image into the Claude model for visual understanding, and outputs the next action including click coordinates, keyboard input, and scroll commands.72% 相似待验证Adept AI 的 ACT-1(Action Transformer)采用像素到动作的端到端学习范式,直接从屏幕截图预测鼠标点击坐标和键盘输入71% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/891652API
curl https://kongchang.com/api/v1/knowledge/claims/891652MCP
get_claim(id=891652)