[KongchangAI]
Verified65% confidenceFactExact time

GRPO通过对同一问题采样一组输出以组内平均奖励作为基线,消除了对独立critic网络的依赖,将显存需求从PPO的约4倍压缩至约2倍

3
Sources
65%
Confidence
Long-term
Relevance
7/17/2026
First Seen

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/542331
API
curl https://kongchang.com/api/v1/knowledge/claims/542331
MCP
get_claim(id=542331)