Course Outline
1. Introduction to Deep Reinforcement Learning
- Defining Reinforcement Learning
- Distinguishing between Supervised, Unsupervised, and Reinforcement Learning
- DRL applications in 2025 (robotics, healthcare, finance, logistics)
- Comprehending the agent-environment interaction cycle
2. Core Concepts of Reinforcement Learning
- Markov Decision Processes (MDP)
- States, Actions, Rewards, Policies, and Value functions
- Balancing Exploration and Exploitation
- Monte Carlo techniques and Temporal-Difference (TD) learning
3. Developing Basic RL Algorithms
- Tabular approaches: Dynamic Programming, Policy Evaluation, and Iteration
- Q-Learning and SARSA
- Epsilon-greedy exploration and decay strategies
- Creating RL environments using OpenAI Gymnasium
4. Moving into Deep Reinforcement Learning
- Limitations of tabular methods
- Utilizing neural networks for function approximation
- Deep Q-Network (DQN) structure and workflow
- Experience replay and target networks
5. Advanced DRL Algorithms
- Double DQN, Dueling DQN, and Prioritized Experience Replay
- Policy Gradient Methods: the REINFORCE algorithm
- Actor-Critic architectures (A2C, A3C)
- Proximal Policy Optimization (PPO)
- Soft Actor-Critic (SAC)
6. Managing Continuous Action Spaces
- Challenges in continuous control
- Applying DDPG (Deep Deterministic Policy Gradient)
- Twin Delayed DDPG (TD3)
7. Practical Tools and Frameworks
- Leveraging Stable-Baselines3 and Ray RLlib
- Logging and monitoring via TensorBoard
- Tuning hyperparameters for DRL models
8. Reward Engineering and Environment Design
- Reward shaping and penalty balancing
- Sim-to-real transfer learning concepts
- Creating custom environments in Gymnasium
9. Partially Observable Environments and Generalization
- Addressing incomplete state information (POMDPs)
- Memory-based approaches using LSTMs and RNNs
- Enhancing agent robustness and generalization
10. Game Theory and Multi-Agent Reinforcement Learning
- Overview of multi-agent environments
- Cooperation versus competition
- Use cases in adversarial training and strategy optimization
11. Case Studies and Real-World Applications
- Autonomous driving simulations
- Dynamic pricing and financial trading strategies
- Robotics and industrial automation
12. Troubleshooting and Optimization
- Diagnosing unstable training processes
- Addressing reward sparsity and overfitting
- Scaling DRL models on GPUs and distributed systems
13. Summary and Next Steps
- Recap of DRL architecture and key algorithms
- Industry trends and research directions (e.g., RLHF, hybrid models)
- Further resources and reading materials
Requirements
- Strong command of Python programming
- Solid grasp of Calculus and Linear Algebra
- Foundational knowledge of Probability and Statistics
- Experience developing machine learning models with Python, NumPy, or TensorFlow/PyTorch
Target Audience
- Developers eager to explore AI and intelligent systems
- Data Scientists investigating reinforcement learning frameworks
- Machine Learning Engineers focusing on autonomous systems
Testimonials (2)
Getting people that never used AI some repetition in prompting and people that do use AI to consider different methods to using it.
Matthew Gay - Tarsus Pharmaceuticals
Course - Artificial Intelligence (AI) Overview
Working from first principles in a focused way, and moving to applying case studies within the same day