待验证50% 置信观点精确时间
开发者声称Ox Alpha在同类任务上的表现比Qwen3系列和Anthropic的Opus 4.5好上1000倍
1
来源数
50%
置信度
短期 (~14 天)
时效性
2026/9/12
首次发现
有效期至:2026/9/26
来源
涉及实体
相关事实
待验证在部分测试中Qwen3中型版本面对当前前沿模型能站稳脚跟,与一年前耗资数十亿美元的系统相比表现更胜一筹54% 相似待验证在HLE科学类基准上,Agent A1 据实测数据超过了Qwen 3.6和Step 3.5 Flash53% 相似待验证Qwen 3.6's performance in general programming capability, development skills, multi-turn agent-enhanced capability, and agent task testing leads Qwen 3.5 in many aspects and approaches or exceeds Claude 4.5 Opus on certain metrics.52% 相似待验证Qwen 3 Max在长程智能体任务(构建熊猫数据集、微调Gemma 2B、搭建本地Web界面)中拿下满分10分52% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/914887API
curl https://kongchang.com/api/v1/knowledge/claims/914887MCP
get_claim(id=914887)