When Does Your RL Agent Need Memory? Recurrent Policies in PPO and ARPPO

Most RL policies see only the current observation. When sensors drop, signals are noisy, or context spans many steps, a feedforward network is generally insufficient. This post benchmarks LSTM-augmented PPO and ARPPO against feedforward baselines across six partial-observability tasks — and finds a consistent asymmetry between the two algorithms.

When Does Your RL Agent Need Memory? Recurrent Policies in PPO and ARPPO Read More »

Reward Normalization Done Right: How ARPPO Uses Welford Statistics for Stable Training

Standard PPO fails on real-world tasks because raw reward signals destabilize value learning. ARPPO solves this with Welford online statistics — numerically stable, constant-memory reward normalization that adapts continuously across training. Combined with an annealed value shrink curriculum, this two-layer stability system makes ARPPO dramatically more reliable than PPO for enterprise deployments.

Reward Normalization Done Right: How ARPPO Uses Welford Statistics for Stable Training Read More »

ARPPO vs PPO: When Does Average Reward RL Win?

Proximal Policy Optimization (PPO) has become the default choice for deep reinforcement learning: reliable, well-understood, and competitive across many benchmarks. But PPO uses discounted reward, which creates a fundamental mismatch with continuing operational tasks. Our ARPPO algorithm [Schneckenreither, 2020; Schneckenreither & Moser, 2025] addresses this directly. Algorithm Comparison Property PPO ARPPO Objective Discounted cumulative reward

ARPPO vs PPO: When Does Average Reward RL Win? Read More »

How Reinforcement Learning Solves Real-World Inventory Optimization

Inventory optimization is one of the oldest problems in operations research. Classical solutions such as EOQ, the newsvendor model, and (s,S) policies assume simplified demand distributions and static cost structures. In the real world, demand is seasonal, trends shift, and supplier lead times vary. Reinforcement learning handles all of this naturally . The RL Formulation

How Reinforcement Learning Solves Real-World Inventory Optimization Read More »

Average Reward vs. Discounted Reward: Why It Matters for Business Optimization

When building AI systems for business operations, one of the most consequential design choices is often overlooked: how do you measure success? Most reinforcement learning textbooks and libraries default to discounted reward . For games and simulated environments, this works fine. For real operational problems, it creates a systematic bias that undermines performance. The Discounting

Average Reward vs. Discounted Reward: Why It Matters for Business Optimization Read More »

x

ARPPO vs Reorder-Point Policies: A Simulation Study of Adaptive Inventory Control

x Read More »

Scroll to Top