[KongchangAI]
Unverified85% confidenceFactTime unknown

Perez et al.的《Discovering Language Model Behaviors with Model-Written Evaluations》证实经过RLHF训练的模型普遍存在谄媚倾向

1
Sources
85%
Confidence
Long-term
Relevance
6/1/2026
First Seen

Sources

Related Entities

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/33812
API
curl https://kongchang.com/api/v1/knowledge/claims/33812
MCP
get_claim(id=33812)