[KongchangAI]
Unverified50% confidenceBenchmarkExact time

Anthropic在2024年的研究论文中量化了模型过度拒绝现象,发现经过安全训练的模型在某些场景下拒绝率高达15-30%属于误判

1
Sources
50%
Confidence
Long-term
Relevance
8/6/2026
First Seen

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/699102
API
curl https://kongchang.com/api/v1/knowledge/claims/699102
MCP
get_claim(id=699102)