[KongchangAI]
Verified75% confidenceTradeoffExact time

在大语言模型训练场景中,PPO需要同时维护Actor、Critic、Reward、Reference四个模型副本,单次训练显存占用往往是基础模型的4倍以上

3
Sources
75%
Confidence
Long-term
Relevance
7/10/2026
First Seen

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/429048
API
curl https://kongchang.com/api/v1/knowledge/claims/429048
MCP
get_claim(id=429048)