Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model

doi:10.1038/S41586-020-03051-4

Open AccessJournal Article10.1038/S41586-020-03051-4

Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model

Julian Schrittwieser, +11 more

- 19 Nov 2019

- arXiv: Learning

1.4K

TL;DR: The MuZero algorithm is presented, which, by combining a tree-based search with a learned model, achieves superhuman performance in a range of challenging and visually complex domains, without any knowledge of their underlying dynamics.

Abstract: Constructing agents with planning capabilities has long been one of the main challenges in the pursuit of artificial intelligence. Tree-based planning methods have enjoyed huge success in challenging domains, such as chess and Go, where a perfect simulator is available. However, in real-world problems the dynamics governing the environment are often complex and unknown. In this work we present the MuZero algorithm which, by combining a tree-based search with a learned model, achieves superhuman performance in a range of challenging and visually complex domains, without any knowledge of their underlying dynamics. MuZero learns a model that, when applied iteratively, predicts the quantities most directly relevant to planning: the reward, the action-selection policy, and the value function. When evaluated on 57 different Atari games - the canonical video game environment for testing AI techniques, in which model-based planning approaches have historically struggled - our new algorithm achieved a new state of the art. When evaluated on Go, chess and shogi, without any knowledge of the game rules, MuZero matched the superhuman performance of the AlphaZero algorithm that was supplied with the game rules.

Chat with Paper

AI Agents for this Paper

Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps

Citations

Towards biologically plausible Dreaming and Planning in recurrent spiking networks

Cristiano Capone, +1 more

- 20 May 2022

TL;DR: The authors proposed a two-module (agent and model) spiking neural network in which "living new experiences in a model-based simulated environment" significantly boosts learning and also explore "planning", an online alternative to dreaming, that shows comparable performances.

...read moreread less

E-MCTS: Deep Exploration in Model-Based Reinforcement Learning by Planning with Epistemic Uncertainty

Yaniv Oren, +2 more

- 21 Oct 2022

TL;DR: In this paper , the authors develop a methodology to propagate epistemic uncertainty in Monte-Carlo Tree Search (MCTS) and utilize the propagated uncertainty for a novel deep exploration algorithm by explicitly planning to explore.

...read moreread less

Book Chapter•10.1007/978-3-030-78811-7_34

Value-Based Continuous Control Without Concrete State-Action Value Function.

Jin Zhu, +2 more

- 17 Jul 2021

TL;DR: In this article, the actor-critic method is proposed to implement value-based continuous control in an effective but compromise way, where actions with higher expected return (state-action value, also as Q) will be selected as the action decision.

...read moreread less

Proceedings Article

Expert Initialized Hybrid Model-Based and Model-Free Reinforcement Learning

Jeppe Langaa, +1 more

- 13 Jun 2023

TL;DR: In this paper , a reinforcement learning algorithm that enables fast learning of control policies based on a limited amount of training data, by leveraging the attributes of both model-based and model-free algorithms, is presented.

...read moreread less

Detection of disadvantageous individual decisions for a game with fantastic elements

B.S. Muller

TL;DR: In this paper , the authors proposed a method to automatically detect mistakes made in games based on artificial intelligence, which is going to help players gain more insight into how individual game events affect the outcome of the game and notify them about what they could have done differently to improve the probability of their team to win the game.

...read moreread less

...

Expand

References

•Journal Article•10.1145/3065386

ImageNet classification with deep convolutional neural networks

Alex Krizhevsky, +2 more

- 24 May 2017

- Communications of The ACM

TL;DR: A large, deep convolutional neural network was trained to classify the 1.2 million high-resolution images in the ImageNet LSVRC-2010 contest into the 1000 different classes and employed a recently developed regularization method called "dropout" that proved to be very effective.

...read moreread less

98.2K

•Proceedings Article

ImageNet Classification with Deep Convolutional Neural Networks

Alex Krizhevsky, +2 more

- 03 Dec 2012

TL;DR: The state-of-the-art performance of CNNs was achieved by Deep Convolutional Neural Networks (DCNNs) as discussed by the authors, which consists of five convolutional layers, some of which are followed by max-pooling layers, and three fully-connected layers with a final 1000-way softmax.

...read moreread less

88.4K

•Book

Reinforcement Learning: An Introduction

Richard S. Sutton, +1 more

- 01 Jan 1988

TL;DR: This book provides a clear and simple account of the key ideas and algorithms of reinforcement learning, which ranges from the history of the field's intellectual foundations to the most recent developments and applications.

...read moreread less

39.7K

...

Expand

Related Papers (5)

Reinforcement Learning: An Introduction

[...]

Richard S. Sutton, +1 more

- 01 Jan 1988

Playing Atari with Deep Reinforcement Learning

[...]

Volodymyr Mnih, +6 more

- 19 Dec 2013

- arXiv: Learning

Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model

Chat with Paper

AI Agents for this Paper

Citations

Towards biologically plausible Dreaming and Planning in recurrent spiking networks

E-MCTS: Deep Exploration in Model-Based Reinforcement Learning by Planning with Epistemic Uncertainty

Value-Based Continuous Control Without Concrete State-Action Value Function.

Expert Initialized Hybrid Model-Based and Model-Free Reinforcement Learning

Detection of disadvantageous individual decisions for a game with fantastic elements

References

ImageNet classification with deep convolutional neural networks

ImageNet Classification with Deep Convolutional Neural Networks

Reinforcement Learning: An Introduction

Human-level control through deep reinforcement learning

Mastering the game of Go with deep neural networks and tree search

Related Papers (5)

Mastering the game of Go without human knowledge

Human-level control through deep reinforcement learning

Mastering the game of Go with deep neural networks and tree search

Reinforcement Learning: An Introduction

Playing Atari with Deep Reinforcement Learning