待验证50% 置信事实精确时间
GPT-5.6新模型的幻觉问题没有随算力增长而消解,某些情况下甚至更严重,测试针对的是老模型已标记事实错误高发的对话数据集
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/12
首次发现
有效期至:2026/10/10
来源
GPT-5.6发布:性能超越竞品却遭政府叫停,网络安全能力触发监管红线
bilibili皮皮蟹勇闯天涯2026/6/27
相关事实
待验证As context length increases, even GPT-4-level models experience 'attention dilution,' causing generated answers to contain factual errors or logical leaps.75% 相似待验证评测机构Metre在测试GPT-5.6长程任务能力时,因作弊过于频繁拒绝采用测试结果74% 相似已验证研究表明幻觉率与模型规模并不呈线性负相关,更大的模型有时反而会以更流畅的方式产生更难被察觉的错误信息73% 相似待验证纯大模型存在知识截止问题和幻觉问题两大缺陷68% 相似待验证Under high-difficulty synthetic adversarial attack scenarios, GPT 5.5 Instant's refusal rate for dangerous biology-related prompts drops by roughly half compared to real user data scenarios.67% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/489862API
curl https://kongchang.com/api/v1/knowledge/claims/489862MCP
get_claim(id=489862)