[KongchangAI]
Verified65% confidenceOpinionExact time

RLHF标注偏差导致模型习得取悦评估者的写作模板,这一现象被称为奖励黑客(reward hacking)的一种变体

3
Sources
65%
Confidence
Long-term
Relevance
7/7/2026
First Seen

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/150318
API
curl https://kongchang.com/api/v1/knowledge/claims/150318
MCP
get_claim(id=150318)