Dynamic Programming
Overview
Overview
If you already know exactly how the environment works — every transition probability, every reward — you don't need to act, explore, or learn from experience at all. You can compute the optimal policy directly, by repeatedly applying the Bellman equations from Value Functions and Bellman Equations until they converge.