ARPPO vs Reorder-Point Policies: A Simulation Study of Adaptive Inventory Control

How does an average-reward RL agent compare to a well-tuned (s, S) reorder-point policy? We simulate 10,000 days of single-echelon inventory under seasonal demand — ARPPO achieves 23% lower holding costs and 31% fewer stockouts without any manual parameter tuning.