← Back to the collection
LEARNING / 2021 / LEARNING STUDY

Learning which way to go.

Reinforcement Learning

Grid-world experiments in value iteration and Q-learning, with rewards, obstacles, and animated state histories.

MEDIUM
  • Python
BUILT WITH
  • NumPy
  • Matplotlib
ACCESS

Public source

Grid-world reinforcement-learning visualization from the original project
Repository animation · historical Q-learning state history, shortened for playback.Open film ↗
01

Two routes to a policy

Value iteration repeatedly updates a grid of values. The Q-learning study combines random exploration with greedy movement toward promising neighboring states. Both record intermediate states for animation.

02

Seeing the reward propagate

The interesting artifact is the sequence rather than just the final route. A positive destination and negative obstacles gradually shape behavior, making discounting and exploration easier to reason about.

SYSTEM SKETCH / CONCEPTUAL OVERVIEW
STATE → CHOOSE ACTION → OBSERVE REWARD
  ↑                           │
  └──────── UPDATE VALUE ←────┘
A reading of the architecture, not an application screenshot.
SYSTEM MAP / ALGORITHM FLOW

Two update rules, one observable grid world.

Grid world → ValueIteration: initial values. Grid world → QLearning.choose_action: states / actions. QLearning.choose_action → Scalar cell update: next state. Scalar cell update → QLearning.choose_action: continue episode. ValueIteration → records + opt_pol: snapshot. Scalar cell update → records + opt_pol: snapshot. records + opt_pol → Animate.generateAnimat: history.123456701 / CONFIGURATIONGrid worldStart / goal / random minesRewards + valid actions02 / PLANNING LOOPValueIterationEvaluate neighbour valuesRepeated grid sweeps03 / EXPLORATIONQLearning.choose_actionRandom or greedy neighbourMove to a valid next cell04 / LEARNING LOOPScalar cell updateReward + discounted differenceUpdate visits and cell value05 / HISTORYrecords + opt_polCopy intermediate value gridsCollect policy states06 / PRESENTATIONAnimate.generateAnimatMatplotlib framesValues + mine / goal markers
  1. 01 / configuration

    Grid world

    Start / goal / random mines

    Rewards + valid actions

    • initial values → 2. ValueIteration
    • states / actions → 3. QLearning.choose_action
  2. 02 / planning loop

    ValueIteration

    Evaluate neighbour values

    Repeated grid sweeps

    • snapshot → 5. records + opt_pol
  3. 03 / exploration

    QLearning.choose_action

    Random or greedy neighbour

    Move to a valid next cell

    • next state → 4. Scalar cell update
  4. 04 / learning loop

    Scalar cell update

    Reward + discounted difference

    Update visits and cell value

    • continue episode → 3. QLearning.choose_action
    • snapshot → 5. records + opt_pol
  5. 05 / history

    records + opt_pol

    Copy intermediate value grids

    Collect policy states

    • history → 6. Animate.generateAnimat
  6. 06 / presentation

    Animate.generateAnimat

    Matplotlib frames

    Values + mine / goal markers

    1. 1Grid world ValueIterationinitial values
    2. 2Grid world QLearning.choose_actionstates / actions
    3. 3QLearning.choose_action Scalar cell updatenext state
    4. 4Scalar cell update QLearning.choose_actioncontinue episode
    5. 5ValueIteration records + opt_polsnapshot
    6. 6Scalar cell update records + opt_polsnapshot
    7. 7records + opt_pol Animate.generateAnimathistory
    The historical “Q-learning” code stores one value per cell. Its policy output is collected through a set, so it is not an ordered shortest-path certificate.
    Read from the implementation
    • QLearning.py
    • ValueIteration.py
    • Animate.py
    03

    Read the implementation on its own terms

    This historical version uses a scalar value per grid cell in its Q-learning study, rather than a complete state-action table. Its README terminology is broader than the implementation; it is presented here as an early experiment in learning behavior.

    CONTINUE EXPLORING

    ANN Visualizer →

    A browser experiment that connects neural-network training to a spatial view of inputs, hidden layers, and predictions.

    2023
    • JavaScript