ARPPO vs PPO: When Does Average Reward RL Win?
Proximal Policy Optimization (PPO) has become the default choice for deep reinforcement learning: reliable, well-understood, and competitive across many benchmarks. But PPO uses discounted reward, which creates a fundamental mismatch with continuing operational tasks. Our ARPPO algorithm [Schneckenreither, 2020; Schneckenreither & Moser, 2025] addresses this directly. Algorithm Comparison Property PPO ARPPO Objective Discounted cumulative reward […]
ARPPO vs PPO: When Does Average Reward RL Win? Read More »
