When Does Your RL Agent Need Memory? Recurrent Policies in PPO and ARPPO
Most RL policies see only the current observation. When sensors drop, signals are noisy, or context spans many steps, a feedforward network is generally insufficient. This post benchmarks LSTM-augmented PPO and ARPPO against feedforward baselines across six partial-observability tasks — and finds a consistent asymmetry between the two algorithms.
When Does Your RL Agent Need Memory? Recurrent Policies in PPO and ARPPO Read More »
