[KongchangAI]
Unverified50% confidenceSolutionExact time

PPO在语言模型对齐中面临奖励欺骗挑战,工程实践中通常在目标函数中加入KL散度惩罚项约束策略模型与SFT基础模型的偏离

1
Sources
50%
Confidence
Long-term
Relevance
7/18/2026
First Seen

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/549785
API
curl https://kongchang.com/api/v1/knowledge/claims/549785
MCP
get_claim(id=549785)