Markov Decision Processes
The Reinforcement Learning Problem described the agent-environment loop informally. The Markov Decision Process (MDP) is the formalism that makes that loop mathematically precise — and precise enough that every algorithm in this section can be stated and analysed against it.