Skip to main content

Exploring a Grid World

In reinforcement learning, an agent learns by trial and error: it acts, sees what happens, and is rewarded when something good happens. But it can only learn from rewards it has actually found, so first it has to explore.

Here the agent is a robot vacuum. Each step it moves one tile up, down, left or right (driving into a wall leaves it where it is), and it earns +1 only when it reaches the dust in the far corner. It is not told where the dust is.

Drive the robot yourself with the arrow keys or the on-screen pad. Then press Random walk to watch an agent that picks a direction at random every step, which is how a learning agent explores before it knows anything, and compare its step count with Go to the dust, the shortest path.

Empty room​

The shortest path is 26 steps; a random walk needs about 1,350 on average.

Four rooms​

A larger 19×19 floor split into four rooms joined by narrow doorways. The shortest path is 36 steps, but a random walk needs about 4,300 on average, because it has to stumble through the doorways by chance.

Random exploration slows down quickly as spaces grow and gain bottlenecks, which is why exploring efficiently is one of the central problems in reinforcement learning.