A reinforcement learning agent that learns to play Street Fighter Alpha: Warriors' Dreams (Game Boy Color) by playing it — improving its moves and strategy through repeated matches.
This project trains an AI to fight in Street Fighter Alpha by hooking a Game Boy emulator up to a reinforcement learning loop. The emulator runs the actual ROM, the agent reads the game's memory to understand what's happening (health, positions, timer), chooses which buttons to press, and is rewarded or penalized based on the outcome. Over thousands of matches it learns which actions tend to win.
The game runs on the PyBoy emulator, the environment is wrapped in a standard OpenAI Gym interface, and the agent is trained with PPO from Stable-Baselines3.
Observations — Each step, the agent reads values straight out of the game's RAM:
| Value | Memory address |
|---|---|
| Player health | 0xC3E2 |
| Opponent health | 0xC3E4 |
| Player X position | 0xC421 |
| Opponent X position | 0xC621 |
| Round timer | 0xCBB3 |
From these it also derives the distance between the two fighters. The observation vector is these stats plus a rolling history of the agent's last 30 actions.
Actions — A MultiDiscrete([4, 2, 2]) action space:
- A direction: down / left / up / right
- Button A pressed or not
- Button B pressed or not
Rewards — The agent is shaped to fight aggressively and decisively:
- Keeps its own health high (+) and drives the opponent's health down (−)
- Penalized for staying far from the opponent (encourages closing in)
- +100 for a KO win, −100 for losing or running out the clock, −500 for a timeout with the opponent still standing
The game is run with speed-up enabled during training so matches play out as fast as possible.
- PyBoy — Game Boy / GBC emulator with a Python bot API and memory access
- OpenAI Gym — RL environment interface (
SFenv) - Stable-Baselines3 — PPO implementation
- TensorBoard — training metrics & reward curves
- Python 3
StreetFighter-AI/
├── app.py # Minimal PyBoy launcher (boots the ROM)
├── apps.py / envi.py # Early environment experiments
├── game/
│ ├── SFenv.py # The Gym environment (observations, actions, rewards)
│ ├── SFenv(product).py # Refined environment variant
│ ├── check_SFenv.py # Validates the env with SB3's check_env
│ ├── doublecheck_SFenv.py # Runs random-action episodes to sanity-check the env
│ ├── game.py # Scripted gameplay / memory-reading experiments
│ ├── ROM/ # Game ROM + save states (not included — see below)
│ ├── models/ # Saved PPO checkpoints (per training run)
│ └── logs/ # TensorBoard training logs
⚠️ You must supply your own ROM. This repository does not distribute the Street Fighter Alpha: Warriors' Dreams ROM. You need a legally obtained.gbcfile and should place it undergame/ROM/.
pip install pyboy gym stable-baselines3 numpy pillowSeveral scripts currently reference absolute Windows paths to the ROM. Update those paths (in game/SFenv.py, game/game.py, etc.) to your local .gbc file before running.
cd game
python check_SFenv.py # confirms the env matches the Gym API
python doublecheck_SFenv.py # runs random actions for a few episodesLoad SFenv into a Stable-Baselines3 PPO model and call .learn(), saving checkpoints into game/models/ and logging to game/logs/ for TensorBoard:
tensorboard --logdir game/logs- Reward shaping and memory addresses are specific to this particular ROM; other Street Fighter titles use different memory layouts.
- File paths are currently hard-coded — making them relative/configurable is a good first improvement.
- This is a personal learning project exploring game-playing reinforcement learning end to end: emulation, memory reading, environment design, reward shaping, and training.
This project is for educational and research purposes. Street Fighter and Street Fighter Alpha: Warriors' Dreams are trademarks of Capcom. No game ROM is included or distributed here.