MuZero was the DeepMind system that learned to play games without being given their rules, using only observations and rewards.
Announced in 2020, MuZero learned an internal model of the aspects of an environment that mattered for planning. It was tested on Atari games, Go, chess, and shogi. Unlike AlphaZero, which was supplied with the rules of chess, Go, and shogi, MuZero was not given those game rules in advance.
MuZero’s significance lies in combining model-based planning with reinforcement learning. The system did not need to reconstruct every detail of a game world; instead, it learned representations useful for predicting rewards and selecting actions. This distinction is why MuZero is often confused with AlphaZero. AlphaGo specialized in Go, while AlphaZero learned several board games from their known rules. MuZero extended the idea toward settings where an agent must discover useful dynamics from interaction rather than receiving a complete simulator specification.