Most reinforcement learning policies receive only the current observation and must act on that alone. When sensors fail, signals arrive with delay, or decisions depend on history spanning many steps, a feedforward network is generally insufficient — the environment may be a partially observable Markov decision process (POMDP) and the agent is missing context it needs. This post benchmarks LSTM-augmented PPO and ARPPO against feedforward baselines on six controlled partial-observability tasks, and finds a consistent asymmetry between the two algorithms.
The Partial Observability Problem
Standard reinforcement learning is formulated as a Markov decision process (MDP): the agent observes the full state st, takes action at, receives reward rt, and transitions to st+1. The Markov property guarantees that st contains all information relevant to the future — so the optimal policy π*(a|st) is memoryless: it maps the current state to an action with no need to remember anything.
In the real world, the Markov assumption is almost never exactly satisfied. A factory sensor reports a moving average, not the true machine state. A stock-keeping unit’s current inventory level does not reveal the pending backlog of supplier deliveries. A customer’s next purchase depends partly on what they bought six months ago. In all these cases the observation ot is an incomplete projection of the hidden state st. The resulting framework is a partially observable Markov decision process (POMDP):
S: hidden states · A: actions · T: transitions · R: rewards
Ω: observation space · O: emission probabilities P(o|s′) · γ: discount
The optimal policy for a POMDP is a function of the belief state — a probability distribution over hidden states, updated by Bayes’ rule at each step [13]. Maintaining the exact belief state is computationally intractable: finite-horizon POMDP planning is PSPACE-hard [1]. In practice, a recurrent neural network approximates the belief state implicitly: the hidden state ht encodes a compressed summary of the observation history o1:t, and the policy is conditioned on both ot and ht.
Long Short-Term Memory in Policy Networks
The Long Short-Term Memory (LSTM) architecture [2] addresses the vanishing-gradient problem that prevents simple recurrent networks from learning dependencies spanning many time steps. An LSTM cell maintains two state vectors: the cell state ct (long-term memory) and the hidden state ht (short-term working memory). Three learned gates — input, forget, and output — control what information flows in, persists, and is exposed at each step:
it = σ(Wi [ht-1, xt] + bi) input gate
ot = σ(Wo [ht-1, xt] + bo) output gate
ct = ft ⊙ ct-1 + it ⊙ tanh(Wc [ht-1, xt] + bc)
ht = ot ⊙ tanh(ct)
In a policy network the input xt is the encoded observation ot (after a feedforward encoder), and the LSTM output ht feeds into separate actor and critic heads. The hidden state is carried across steps within an episode and reset at episode boundaries.
LSTM was first applied to RL by Bakker [10] in a value-based setting; Wierstra et al. [3] later formalised recurrent policy gradient methods. The approach was popularised at scale by Mnih et al.’s A3C [4] and Hausknecht & Stone’s DRQN [5]. Ni et al. [14] recently demonstrated that LSTM-based recurrent policies are a strong baseline across many POMDP benchmarks, competitive with or exceeding specialised POMDP methods. The sb3-contrib extension for Stable-Baselines3 [6] provides a RecurrentPPO implementation; the present work extends LSTM support to ARPPO (average-reward PPO [7], [8]) and implements both on the vectorised-rollout path.
Truncated Backpropagation Through Time in On-Policy RL
Training a recurrent policy with PPO requires care: the rollout buffer stores entire trajectories, and gradients must flow through time within each sequence. Full backpropagation through time (BPTT) across an entire episode is expensive and prone to gradient explosion. The standard remedy is Truncated BPTT (TBPTT) [9]: the rollout is sliced into fixed-length chunks of length seq_len, and gradients are computed only within each chunk. The hidden state carried into each chunk is treated as a constant (detached from the computation graph), so there is no gradient flow across chunk boundaries.
Each vectorised rollout of n steps per environment is sliced into chunks of length seq_len. Within a chunk, the actor and critic forward passes are fully differentiable. The per-environment hidden state at the start of each chunk is stored during the rollout (no-grad) and used as the initial hidden state for the corresponding training chunk. Episode-start flags carry a sequence mask that zeroes the hidden state at true episode boundaries — preventing the network from carrying memory across episode resets. Chunk-shuffled minibatches ensure each update epoch sees a random ordering of sequence chunks across environments.
A practical consequence: if the cue the agent must remember spans more steps than seq_len, TBPTT cannot propagate gradient across it, and the agent must rely entirely on the hidden state stored before the chunk boundary. We test both regimes explicitly below.
Benchmark Design
We evaluate six partial-observability tasks across two benchmark groups. All runs use 100k environment steps, 3 random seeds, and greedy evaluation over 20 episodes after training. The baseline shared hyperparameters follow the rl-baselines3-zoo CartPole geometry: 8 parallel environments, 32-step rollouts, 20 update epochs, batch size 256, γ=0.98, λ=0.8, clip ε=0.2, no entropy coefficient. The LSTM encoder is a single layer of width 64 (linear+ReLU encoder → LSTM → head), with seq_len=32 and burn_in=0. PPO uses lr=1×10−3; ARPPO uses lr=1×10−4 with the default adaptive learning rate schedule.
Each LSTM configuration is paired with a feedforward (FF) control that is byte-identical except for the recurrent layer — same hyperparameters, same seeds, same evaluation protocol. The FF control isolates the effect of memory.
Group 1 — Partial-Observability Tasks
These six tasks introduce partial observability through different mechanisms. The first two (flicker, noisy-PO) and the T-Maze variants strictly require memory — a memoryless policy cannot achieve above-chance performance on them. The remaining two (delay2, po-l4) are genuinely partially observable but turn out to be approximately tractable without explicit memory: a fixed two-step delay and a single missing feature can often be handled implicitly by an adaptive feedforward policy. These are marked † in the table below and serve as negative controls: LSTM does not help and often hurts when memory is not strictly necessary.
| Task | Corruption | Memory requirement |
|---|---|---|
| CartPole-flicker | Observations blanked with p=0.5 each step | Reconstruct state from last visible frame |
| CartPole-noisy-PO | Velocity hidden; position noise σ=0.1 | Filter noise; infer velocity from position history |
| CartPole-delay2 † | Full state delivered 2 steps late | Buffer delayed observations (FF-tractable for short delays) |
| CartPole-po-l4 † | Only pole angular velocity hidden | Infer missing velocity from angle trajectory (FF-tractable in practice) |
| T-Maze L=10 | Cue visible on step 1 only; goal arm changes per episode | Retain single-step cue for up to 10 steps (within one chunk) |
| T-Maze L=40 | Same as above, corridor length 40 | Retain cue across TBPTT chunk boundary (40 > seq_len=32) |
The T-Maze is a classic POMDP benchmark introduced by Bakker [10] and widely used to evaluate credit assignment across time. An agent navigates a corridor of length L; at step 1 it observes a cue (left or right) that reveals the rewarding arm of the T-junction at the end. The cue then disappears. Correct arm: +4; wrong arm: −0.1; any wasted move: −0.1. The chance return — achieved by always committing to one fixed arm, correct 50% of the time — is 0.5 × 4 + 0.5 × (−0.1) = 1.95. The optimal return is +4.0 (direct path, correct arm every episode).
Group 2 — Full-Observability Regression
These tasks are true MDPs. A correctly implemented LSTM should match — but not significantly exceed — the feedforward baseline here: it carries additional parameters and computation overhead without a signal to exploit. Regressions indicate that the recurrent architecture’s additional capacity is not being exploited beneficially at the shared preset.
Results
Group 1: Where Memory Helps
| Task | Algorithm | FF mean | LSTM mean | Δ |
|---|---|---|---|---|
| CartPole-flicker | PPO | 22.2 (23.2, 22.9, 20.5) | 251.2 (427.0, 184.5, 142.0) | +229.0 |
| CartPole-flicker | ARPPO | 22.8 (25.1, 22.2, 21.1) | 191.8 (142.9, 124.6, 307.8) | +169.0 |
| CartPole-noisy-PO | PPO | 45.9 (44.3, 51.9, 41.5) | 50.6 (43.0, 51.8, 56.9) | +4.7 |
| CartPole-noisy-PO | ARPPO | 47.2 (47.6, 50.0, 44.2) | 124.8 (99.6, 157.9, 117.1) | +77.6 |
| CartPole-delay2 † | PPO | 416.3 (249.0, 500.0, 500.0) | 358.8 (382.1, 194.3, 500.0) | −57.5 |
| CartPole-delay2 † | ARPPO | 500.0 (500.0, 500.0, 500.0) | 433.7 (500.0, 301.0, 500.0) | −66.3 |
| CartPole-po-l4 † | PPO | 405.8 (491.3, 405.7, 320.6) | 241.8 (210.1, 309.6, 205.6) | −164.0 |
| CartPole-po-l4 † | ARPPO | 262.7 (236.0, 313.6, 238.4) | 262.6 (259.4, 144.4, 384.1) | ≈0 |
| T-Maze L=10 | PPO | 2.0 (chance) (2.6, 1.5, 2.0) | 3.3 (2/3 optimal) (4.0, 4.0, 2.0) | +1.3 |
| T-Maze L=10 | ARPPO | 2.0 (chance) (2.6, 1.5, 2.0) | 2.0 (chance) (2.6, 1.5, 2.0) | 0 |
| T-Maze L=40 | PPO | −0.7 (1.3, 1.5, −5.0) | 2.0 (chance) (2.6, 1.5, 2.0) | +2.7 |
| T-Maze L=40 | ARPPO | 2.0 (chance) (2.6, 1.5, 2.0) | −5.0 (collapse) (−5.0, −5.0, −5.0) | −7.0 |
Results: 100k env steps, 3 seeds, greedy evaluation over 20 episodes. Per-seed values shown in parentheses. Highlighted rows show clear memory gains. Tasks marked † are partially observable but approximately tractable without memory; LSTM regressions there are expected. T-Maze chance = 1.95; optimal = 4.0.
Group 2: Full-Observability Regression
| Task | Algorithm | FF mean | LSTM mean | Δ |
|---|---|---|---|---|
| CartPole (full MDP) | PPO | 500.0 ✓ (500.0, 500.0, 500.0) | 206.9 (139.2, 221.0, 260.5) | −293.1 |
| CartPole (full MDP) | ARPPO | 277.7 (195.6, 137.5, 500.0) | 498.2 ✓ (500.0, 494.5, 500.0) | +220.5 |
| Acrobot (full MDP) | PPO | −82.7 (−80.0, −86.3, −81.9) | −82.5 (−77.2, −90.1, −80.1) | +0.2 (parity) |
| LunarLander (full MDP) | PPO | −80.7 (−69.8, −108.2, −64.1) | −223.7 (−164.0, −364.4, −142.8) | −143.0 |
Results: 3 seeds, greedy evaluation. Acrobot PPO shows parity; CartPole and LunarLander PPO regress; CartPole ARPPO strongly improves.
Reading the Results: Five Key Findings
Finding 1 — Flicker Is the Cleanest Memory-Necessity Signal
With 50% of observations randomly blanked, the feedforward policy reduces to a chance-level strategy (~22 mean return for both PPO and ARPPO) — it cannot balance the pole on blank steps because it has no idea where the pole is. The LSTM recovers to 251 (PPO) and 192 (ARPPO) out of a maximum of 500, demonstrating that the recurrent hidden state is successfully carrying forward the last meaningful observation. This is the kind of result that validates the infrastructure: the hidden state is correctly stored, masked at episode boundaries, and forwarded through TBPTT chunks.
Finding 2 — ARPPO Benefits More Than PPO From Memory on Noisy Observations
On CartPole-noisy-PO (velocity hidden, position noise σ=0.1), ARPPO-LSTM reaches 124.8 versus the ARPPO-FF baseline of 47.2 — a gain of +77.6 across 3 seeds. PPO-LSTM shows negligible improvement (+4.7, from 45.9 to 50.6). A consistent pattern appears in Group 2 as well: on the full-MDP CartPole (no observation corruption), ARPPO-LSTM improves from 277.7 to 498.2, while PPO-LSTM regresses sharply from a perfect 500 to 207. Note an important confound: PPO and ARPPO use different learning rates at the shared preset (lr 1×10−3 vs 1×10−4). A faster learning rate may exacerbate LSTM training instability for PPO and mask LSTM benefits; an LR-controlled comparison is in progress.
A plausible mechanism for the Group 1 asymmetry on noisy-PO tasks: ARPPO’s critic must estimate the long-run average reward ρ and the differential value Vd(s). Under noisy or missing observations, the feedforward critic’s estimate of Vd(s) is systematically noisy, which destabilises the advantage estimates that drive the actor update. The recurrent critic can filter observation noise across time and produce more stable value estimates — a benefit that compounds over training. PPO’s discounted critic may be somewhat less sensitive to this because temporal discounting already down-weights distant, potentially more corrupted, observations. We emphasise that this is a hypothesis: n=3 seeds is insufficient to distinguish mechanism from variance, and the learning-rate confound cannot be ruled out at this stage.
Finding 3 — T-Maze Shows Long-Horizon Credit Assignment Within Chunks
At L=10 (cue must survive 10 steps, which fits within a single seq_len=32 TBPTT chunk), PPO-LSTM reaches the optimal 4.0 in 2/3 seeds — the agent successfully encodes the single-step cue in its hidden state and retrieves it 10 steps later at the junction. The feedforward policy is stuck at chance (2.0). ARPPO-LSTM does not learn this task at 100k steps, consistent with the pattern that ARPPO responds poorly to sparse terminal rewards (the sparse +4 / −0.1 structure provides little differential signal for ρ to latch onto).
Finding 4 — Chunk Boundaries Are a Hard Limit Under TBPTT
At seq_len=40 (corridor length 40, cue crosses the TBPTT boundary at seq_len=32), gradient cannot flow across the chunk boundary; the agent relies entirely on the stored hidden state. PPO-LSTM reaches only chance (2.0) in all seeds: the hidden state does carry the cue across the boundary (performance is not negative), but the network does not learn to exploit it reliably within 100k steps. ARPPO-LSTM collapses to −5.0 in all three seeds — well below the chance level of 1.95. A score of −5.0 indicates that the agent never commits at the T-junction at all: it accumulates wasted-move penalties (−0.1/step) until episode truncation on every run. This is qualitatively distinct from an “always-one-arm” memoryless failure; it suggests that ARPPO’s differential signal (r − ρ) provides no gradient pressure toward junction commitment when the terminal reward (+4 or −0.1) cannot be connected to earlier actions across the truncated TBPTT window.
The remedy is a longer sequence length or a burn-in window (reading the stored hidden state under no-grad for several steps before resuming gradient flow). This remains an open extension.
Finding 5 — LSTM Is Not Free on Full MDPs, Especially for PPO
On full-MDP CartPole, PPO-FF achieves a perfect score (500.0, all seeds); PPO-LSTM regresses sharply to 207. On LunarLander-v3, PPO-LSTM (−223.7) is significantly worse than PPO-FF (−80.7) and also slower (50 min vs 36 min for 3 seeds). Acrobot is the exception: PPO-LSTM (−82.5) matches PPO-FF (−82.7) exactly, suggesting that the task does not benefit from memory but also does not suffer from the overhead.
These regressions at the shared preset are likely due to optimisation dynamics: the LSTM adds parameters and a harder optimisation landscape. PPO with lr=1×10−3 that works well for feedforward networks may be too aggressive for an LSTM. Ablation studies over learning rate and sequence length are in progress.
When to Use Recurrent Policies: Practical Guidance
| Situation | Recommendation |
|---|---|
| Observations randomly dropped (sensor failure, packet loss) | Use LSTM — flicker benchmark shows +10× improvement over FF |
| Observations noisy (measurement error, aggregated sensors) | Use LSTM, especially with ARPPO — recurrent critic filters noise effectively |
| Delayed observations (pipeline latency, batched reporting) | LSTM helps if the delay is short; for long delays, prioritise extending seq_len |
| History-dependent decisions (ordering depends on past demand) | Use LSTM if the relevant history exceeds what can fit in the current observation vector |
| Full MDP (Markovian observations, no missing state) | Start with feedforward; add LSTM only if feedforward is clearly insufficient |
| Long credit assignment horizon (> seq_len steps) | Increase seq_len or add burn_in; gradient cannot span chunk boundaries under TBPTT |
Connection to Average-Reward RL and Operational Settings
ARPPO’s domain — ongoing operational problems (inventory management, production scheduling, queueing control) — is particularly susceptible to partial observability. A warehouse management system typically knows current stock levels but not in-transit quantities for all SKUs. A production scheduler knows the queue length but not the exact machine health state. A logistics planner observes shipment ETAs but not the true port congestion that will delay them.
The benchmark results suggest that ARPPO’s recurrent critic can filter observation noise more effectively than PPO’s — a useful property for operational settings where sensor reliability varies. However, the T-Maze results also demonstrate that ARPPO’s differential reward structure can struggle with sparse terminal-reward tasks, which requires careful reward design when long-range credit assignment is needed (see our earlier post on when ARPPO beats PPO).
Open Questions
- Hyperparameter transfer: The regressions on full MDPs (CartPole, LunarLander PPO) are consistent with a learning rate that is too aggressive for LSTMs. PPO and ARPPO used different learning rates in this study; a controlled LR-sweep with matched rates is in progress to disentangle the LR effect from the memory effect.
- Burn-in: Introduced by Kapturowski et al. [15] for off-policy recurrent RL (R2D2), burn-in reads the stored hidden state under no-grad for a warm-up window before resuming gradient flow. Adapting this to on-policy TBPTT may help T-Maze L=40 and other tasks where the cue must survive across chunk boundaries.
- Shared vs independent actor/critic: The current setup uses independent actor and critic LSTM encoders. A shared trunk (one LSTM feeding both heads) may reduce the total parameter count and improve sample efficiency.
- ARPPO on sparse-reward POMDPs: ARPPO-LSTM fails on T-Maze at the 100k budget. Whether this is a fundamental mismatch between average-reward RL and sparse terminal-reward tasks, or a training budget and reward-design issue, remains open.
- Transformer-based memory: LSTM is not the only option for recurrent policy networks. GTrXL and related architectures (Parisotto et al. [16]) offer better training parallelism and may handle long-horizon credit assignment more gracefully than TBPTT by attending over the full context window. This is a natural next step for the T-Maze L=40 failure mode.
If you are working with partial observability in a production operational setting, get in touch — we are actively extending the recurrent benchmarks and are interested in real-world partial observability scenarios.
References
- Papadimitriou, C.H. & Tsitsiklis, J.N. (1987). The Complexity of Markov Decision Processes. Mathematics of Operations Research, 12(3), 441–450. doi:10.1287/moor.12.3.441 (finite-horizon POMDP planning is PSPACE-hard)
- Hochreiter, S. & Schmidhuber, J. (1997). Long Short-Term Memory. Neural Computation, 9(8), 1735–1780. doi:10.1162/neco.1997.9.8.1735
- Wierstra, D., Förster, A., Peters, J. & Schmidhuber, J. (2007). Solving Deep Memory POMDPs with Recurrent Policy Gradients. Proceedings of ICANN 2007, LNCS 4668, pp. 697–706. doi:10.1007/978-3-540-74690-4_71
- Mnih, V. et al. (2016). Asynchronous Methods for Deep Reinforcement Learning. Proceedings of ICML 2016. arXiv:1602.01783.
- Hausknecht, M. & Stone, P. (2015). Deep Recurrent Q-Learning for Partially Observable MDPs. AAAI Fall Symposium 2015. arXiv:1507.06527.
- Raffin, A. et al. (2021). Stable-Baselines3: Reliable Reinforcement Learning Implementations. JMLR, 22(268), 1–8. jmlr.org/papers/v22/20-1364. RecurrentPPO is provided by the separate sb3-contrib package: github.com/Stable-Baselines3/stable-baselines3-contrib.
- Schneckenreither, G. & Haeussler, S. (2020). Average Reward Reinforcement Learning for Supply Chain Management. Proceedings of MISTA 2020.
- Schneckenreither, G. & Moser, B. (2025). ARPPO: Average-Reward Proximal Policy Optimisation for Continuing Tasks. (Preprint.)
- Werbos, P.J. (1990). Backpropagation Through Time: What It Does and How to Do It. Proceedings of the IEEE, 78(10), 1550–1560. doi:10.1109/5.58337
- Bakker, B. (2001). Reinforcement Learning with Long Short-Term Memory. Advances in Neural Information Processing Systems 14 (NeurIPS 2001). proceedings.neurips.cc
- Sutton, R.S. & Barto, A.G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
- Igl, M., Zintgraf, L., Le, T.A., Wood, F. & Whiteson, S. (2018). Deep Variational Reinforcement Learning for POMDPs. Proceedings of ICML 2018. arXiv:1806.02426.
- Kaelbling, L.P., Littman, M.L. & Cassandra, A.R. (1998). Planning and Acting in Partially Observable Stochastic Domains. Artificial Intelligence, 101(1–2), 99–134. doi:10.1016/S0004-3702(98)00023-X
- Ni, T., Eysenbach, B. & Salakhutdinov, R. (2022). Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs. Advances in Neural Information Processing Systems 35 (NeurIPS 2022). arXiv:2110.05038.
- Kapturowski, S. et al. (2019). Recurrent Experience Replay in Distributed Reinforcement Learning. Proceedings of ICLR 2019. openreview.net/forum?id=r1lyTjAqYX (introduces burn-in for recurrent RL)
- Parisotto, E. et al. (2020). Stabilizing Transformers for Reinforcement Learning. Proceedings of ICML 2020. arXiv:1910.06764. (GTrXL: gated Transformer architecture for recurrent policies)
Related Articles
-
ARPPO vs PPO: When Does Average Reward Win?
Algorithm-level comparison across episodic and continuing tasks.
-
Vectorized Environments: Parallel Sampling at Scale
How the vectorised rollout path that LSTM runs on is built and why it matters.
-
KL Divergence and Epochs in ARPPO Training
The early-stopping mechanism that protects policy stability during each update cycle.
