Learning which way to go.
Reinforcement Learning
Grid-world experiments in value iteration and Q-learning, with rewards, obstacles, and animated state histories.
- Python
- NumPy
- Matplotlib
Public source

Two routes to a policy
Value iteration repeatedly updates a grid of values. The Q-learning study combines random exploration with greedy movement toward promising neighboring states. Both record intermediate states for animation.
Seeing the reward propagate
The interesting artifact is the sequence rather than just the final route. A positive destination and negative obstacles gradually shape behavior, making discounting and exploration easier to reason about.
STATE → CHOOSE ACTION → OBSERVE REWARD ↑ │ └──────── UPDATE VALUE ←────┘
Two update rules, one observable grid world.
- 01 / configuration
Grid world
Start / goal / random mines
Rewards + valid actions
- initial values → 2. ValueIteration
- states / actions → 3. QLearning.choose_action
- 02 / planning loop
ValueIteration
Evaluate neighbour values
Repeated grid sweeps
- snapshot → 5. records + opt_pol
- 03 / exploration
QLearning.choose_action
Random or greedy neighbour
Move to a valid next cell
- next state → 4. Scalar cell update
- 04 / learning loop
Scalar cell update
Reward + discounted difference
Update visits and cell value
- continue episode → 3. QLearning.choose_action
- snapshot → 5. records + opt_pol
- 05 / history
records + opt_pol
Copy intermediate value grids
Collect policy states
- history → 6. Animate.generateAnimat
- 06 / presentation
Animate.generateAnimat
Matplotlib frames
Values + mine / goal markers
- 1Grid world ValueIterationinitial values
- 2Grid world QLearning.choose_actionstates / actions
- 3QLearning.choose_action Scalar cell updatenext state
- 4Scalar cell update QLearning.choose_actioncontinue episode
- 5ValueIteration records + opt_polsnapshot
- 6Scalar cell update records + opt_polsnapshot
- 7records + opt_pol Animate.generateAnimathistory
Read from the implementation
QLearning.pyValueIteration.pyAnimate.py
Read the implementation on its own terms
This historical version uses a scalar value per grid cell in its Q-learning study, rather than a complete state-action table. Its README terminology is broader than the implementation; it is presented here as an early experiment in learning behavior.