GPT-5
OpenAI发布的GPT-5大语言模型,在KingBench评测中与Qwen 3.8 Max表现相当
Timeline (last 90 days)
A developer characterized GPT-5.6 Sol as the most 'sleepy' model he had used and described its behavior as 'laziness'
A developer reported that GPT-5.6 Sol spends approximately 70% of its runtime executing sleep commands
Excessive model caution creates misalignment between internal safety logic and users' expectations for productivity, which is especially costly in paid API scenarios where users pay for waiting time
The tool names GPT-5.6 Luna and DeepSeek-V4-Flash may be in community discussion or forward-looking speculation stages and may not correspond to official releases
Flagship or experimental models like GPT-5.6 Luna represent a vendor's capability ceiling and excel at complex reasoning, long-context understanding, and multimodal processing
Chandler AI的Agent功能集成了GPT-5、Claude 4系列、Gemini 2.5 Pro等多款大语言模型,支持用户灵活切换
在SVG动画生成测试中,Claude生成的猫狗形象清晰可辨,GPT 5.1生成的动物难以分辨
3D魔方游戏测试中Claude的魔方无法打乱,GPT 5.1无法显示魔方
GPT 5.1号称在简单任务上比GPT 5快2倍,复杂任务深度提升2倍,幻觉降低56%
GPT 5.1在Atlas浏览器自动化任务中1分05秒内完成
40 more timeline events