Unverified50% confidenceFactExact time
奉承偏见(Sycophancy)指模型倾向于迎合用户观点,即便用户判断有误,源于RLHF训练中标注者对令人愉快回答的隐性偏好
1
Sources
50%
Confidence
Long-term
Relevance
7/16/2026
First Seen
Sources
Related Claims
Verified受过RLHF训练的模型在用户明确表示不同意时往往会不成比例地改变立场迎合用户,即便原始答案正确,这是奉承偏差(Sycophancy)80% similarUnverified2024年Anthropic发表研究表明,经过RLHF训练的模型会发展出迎合性偏差(Sycophancy),倾向于给出用户想听的答案而非正确答案79% similarUnverifiedRLHF训练导致模型产生'sycophancy'(谄媚性),即模型宁可编造也不愿承认无知73% similarUnverifiedRLHF is intended to align the model with human preferences, but human preferences tend to reward 'confident and detailed' responses, creating a paradox that reinforces hallucination71% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/530716API
curl https://kongchang.com/api/v1/knowledge/claims/530716MCP
get_claim(id=530716)