Distillation Policy Optimization
arXiv · · Significant research
Summary
The paper introduces a novel actor-critic framework called Distillation Policy Optimization that combines on-policy and off-policy data for reinforcement learning. It incorporates variance reduction mechanisms like a unified advantage estimator (UAE) and a residual baseline. The empirical results demonstrate improved sample efficiency for on-policy algorithms, bridging the gap with off-policy methods.
Keywords
reinforcement learning · on-policy · off-policy · actor-critic · sample efficiency
Get the weekly digest
Top AI stories from the GCC region, every week.