Gradient-Free versus Gradient-Based Risk Constraints in Reinforcement Learning: A Reservoir Governance Case Study
Abstract
We ask whether the choice of reinforcement-learning optimizer matters for solving a CVaR-constrained reservoir-governance problem, and why. Training CEM (population-based), SAC, and PPO (gradient-based, with a Lagrangian expected-cost safety critic) on an identical simulator and evaluating all three on the full training objective, we find a sharp, reproducible pattern across three of four physical regimes: CEM achieves strictly lower realized tail risk than every SAC/PPO variant even its least conservative policy is safer than their most conservative ones – while SAC and PPO achieve strictly lower cost. A fourth regime, where management capacity saturates, is an instructive exception where CEM is both safest and cheapest. We explain this mechanistically: CEM evaluates the true, non-differentiable CVaR of the episode-maximum deficit directly, while gradient-based methods must substitute a differentiable expected-discounted-cost surrogate. We prove this substitution yields a discount-induced blindness result a per-step safety guarantee that decays geometrically with the time index and confirm it empirically via a discount-factor ablation and an extreme-budget stress test. An independent five-seed replication, using a scale-invariant dual-ascent correction, sharpens the risk ordering to statistical significance; a genuinely trained quantile/distributional safety critic, however, does not close the gap and in fact increases realized risk in every seed tested, pointing to Monte-Carlo quantile-estimation noise rather than La-grangian under-tuning as the remaining obstacle. Closed-loop existence, local-stability, and probabilistic practical-stability theorems complete the analysis, applicable regardless of optimizer.
Related articles
Related articles are currently not available for this article.