Build notes · Goh Kun Ming
Pendulum Learning Lab
AI · Notebook experiment
Jul 2025 – Aug 2025
Project overview
Learning control through trial, reward and evaluation.
A reinforcement-learning lab comparing Standard, Double, Dueling and Rainbow DQN on Pendulum-v1. Continuous torque is mapped to discrete actions, with reusable replay, training, evaluation and tuning components. Local smoke execution is kept distinct from full studies; matched multi-seed reward superiority and control stability are not claimed.
Inside this build
- Mapped continuous torque to discrete action bins for value-based learning.
- Configured seeded experiments with replay buffers, warm-up and target-network updates.
- Added prioritised replay and distributional outputs for Rainbow DQN.
- Provided reward summaries, learning-curve analysis, greedy rollouts and comparison views.
- Separated bounded smoke runs from the full 500-episode study configuration.
- Validated experiment tooling and notebook integrity while documenting missing matched multi-seed runs.












