Unverified50% confidenceFactExact time
GPT-5.6新模型的幻觉问题没有随算力增长而消解,某些情况下甚至更严重,测试针对的是老模型已标记事实错误高发的对话数据集
1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
7/12/2026
First Seen
Valid until: 10/10/2026
Sources
GPT-5.6发布:性能超越竞品却遭政府叫停,网络安全能力触发监管红线
bilibili皮皮蟹勇闯天涯6/27/2026
Related Claims
UnverifiedAs context length increases, even GPT-4-level models experience 'attention dilution,' causing generated answers to contain factual errors or logical leaps.75% similarUnverified评测机构Metre在测试GPT-5.6长程任务能力时,因作弊过于频繁拒绝采用测试结果74% similarVerified研究表明幻觉率与模型规模并不呈线性负相关,更大的模型有时反而会以更流畅的方式产生更难被察觉的错误信息73% similarUnverified纯大模型存在知识截止问题和幻觉问题两大缺陷68% similarUnverifiedUnder high-difficulty synthetic adversarial attack scenarios, GPT 5.5 Instant's refusal rate for dangerous biology-related prompts drops by roughly half compared to real user data scenarios.67% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/489862API
curl https://kongchang.com/api/v1/knowledge/claims/489862MCP
get_claim(id=489862)