Reinforcement Learning

machine-learning

Handwritten notes on reinforcement learning: Markov decision processes, Q-learning, policy gradients, reward shaping, and exploration vs exploitation.

These notes cover the fundamentals of reinforcement learning — agents learning through interaction:

  • Markov Decision Processes: States, actions, rewards, transitions, discount factors, and the Bellman equation.
  • Value-Based Methods: Q-learning, SARSA, value iteration, and temporal difference learning.
  • Policy-Based Methods: Policy gradients, REINFORCE, actor-critic architectures, and advantage functions.
  • Exploration vs Exploitation: Epsilon-greedy, UCB, Thompson sampling, and the exploration-exploitation tradeoff.