[KongchangAI]
Unverified50% confidenceFactExact time

RLHF是当前主流大模型安全对齐的核心技术路径,但存在奖励模型被过度优化(Reward Hacking)的系统性局限

1
Sources
50%
Confidence
Long-term
Relevance
7/5/2026
First Seen

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/112491
API
curl https://kongchang.com/api/v1/knowledge/claims/112491
MCP
get_claim(id=112491)