Multi-step Proximal Policy Improvement in Offline Reinforcement Learning
The geometric interpretation of offline policy updates as manifold gradient flows is intellectually satisfying, but the practical advantage of multi-step composition over single-step methods isn't demonstrated clearly in this excerpt. Worth reading if you're tuning offline RL agents; skim otherwise.