[KongchangAI]
Unverified50% confidenceSolutionExact time

GRPO算法通过组内相对奖励机制直接比较同一问题的一组答案优劣,省去了PPO所需的独立评论家模型

1
Sources
50%
Confidence
Long-term
Relevance
7/16/2026
First Seen

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/533829
API
curl https://kongchang.com/api/v1/knowledge/claims/533829
MCP
get_claim(id=533829)