Artificial intelligence · Reinforcement learning

An AI that learns to play on its own.

No one programmed the rules into it. It watches the screen the way a person would, detects the obstacles, and discovers —through trial and reward— the exact moment to jump in Chrome's dinosaur game. Every attempt makes it better.

PPO
Reinforcement learning
Vision
Detects obstacles on screen
Self
Improves on its own each attempt
What it does

It learns to play like a person does: by watching

No tricks, no access to the game's code. Just the screen, the eyes (computer vision) and experience.

It sees the screen

It captures the game in real time and, with computer vision (OpenCV), detects where the dinosaur is and the obstacles coming toward it.

It decides in milliseconds

At every instant it chooses an action —jump or keep going— trying to dodge what's coming. It reacts faster than you blink.

It learns through reward

Each time it survives a little longer, it gets a reward. The PPO algorithm tunes its "instinct" to repeat what works and avoid what kills it.

It improves with every attempt

It starts clumsy and ends up expert. Episode after episode its score climbs — the learning curve proves it.

Two approaches

A trained neural network that learns on its own, and a fixed-rules version. Perfect for comparing the AI against a traditional script.

It visualizes what it "sees"

A debug mode shows exactly the region of the game the AI is analyzing and the obstacles it detects, drawn on screen. Full transparency into its reasoning.

How it works

The learning loop

The same loop used by AI systems that learn to drive or to play: observe, act, reward, repeat.

Observe

It captures the screen and detects the dino and obstacles with OpenCV.

Act

The neural network (PPO) chooses to jump or not, based on what it sees.

Reward

Surviving adds points; crashing subtracts them.

Learn

It tunes its decisions to maximize future reward.

Under the hood

Technology

Reinforcement learning

PPO with Stable-Baselines3 and a custom Gymnasium environment (observation, actions and reward).

Computer vision

OpenCV + screen capture to detect obstacles from the image, without reading the game's memory.

Neural networks

PyTorch trains the policy that decides each action; the model is saved and reused.

Watch the AI play on its own

All the code is open: the environment, the training, and the script that loads the model and plays autonomously.