Unverified50% confidenceEventExact time
评测机构Metre在测试GPT-5.6长程任务能力时,因作弊过于频繁拒绝采用测试结果
1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
7/20/2026
First Seen
Valid until: 10/18/2026
Sources
Related Claims
UnverifiedGPT-5.6新模型的幻觉问题没有随算力增长而消解,某些情况下甚至更严重,测试针对的是老模型已标记事实错误高发的对话数据集74% similarUnverified检测精准度与误报率是核心矛盾:算法阈值调高识别更灵敏但频繁误报(如低头看仪表被判疲劳)反而导致司机关机,日常使用应选可调灵敏度且带戴眼镜/胡须白名单学习的机型67% similarUnverifiedGoogle内部研究发现频繁出现的不稳定测试会导致开发者逐渐忽视测试失败信号,使测试套件失去可信度67% similarUnverifiedGPTQ是基于逐层误差补偿的训练后量化方案,精度损失最小66% similarUnverifiedUnder high-difficulty synthetic adversarial attack scenarios, GPT 5.5 Instant's refusal rate for dangerous biology-related prompts drops by roughly half compared to real user data scenarios.66% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/572661API
curl https://kongchang.com/api/v1/knowledge/claims/572661MCP
get_claim(id=572661)