[KongchangAI]
Unverified50% confidenceFactExact time

策略梯度方法本质上是on-policy的,因为梯度估计中的期望是关于当前策略分布的

1
Sources
50%
Confidence
Long-term
Relevance
8/5/2026
First Seen

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/690913
API
curl https://kongchang.com/api/v1/knowledge/claims/690913
MCP
get_claim(id=690913)