AlphaZero: Chess, Shogi & Go Self-Play Results

DeepMind trained one self-play algorithm from only the rules, then evaluated it against Stockfish, Elmo and AlphaGo under disputed match conditions.

One algorithm, three games

AlphaZero replaced the handcrafted evaluation functions and move-ordering heuristics that chess and shogi engines had relied on for decades with a single neural network trained from nothing but the rules, the same recipe AlphaGo Zero had already used for Go. Training ran for 700,000 steps using 5,000 first-generation TPUs to generate self-play games and 64 second-generation TPUs to train the network, a separate instance for each game. [1]

The paper reports AlphaZero overtook the reference engine on an Elo scale at different points in training for each game: it passed Stockfish in chess after about 4 hours (300,000 steps), passed Elmo in shogi after under 2 hours (110,000 steps), and passed AlphaGo Lee in Go after 8 hours (165,000 steps). DeepMind's own summary rounds this to "within 24 hours" for all three games. [1]

The paper's headline evaluation was a 100-game match against each reference engine at one minute per move, with AlphaZero and the previous AlphaGo Zero running on a single machine with four TPUs while Stockfish and Elmo played at their strongest settings on 64 threads with a 1GB hash table. AlphaZero won 28 games and drew 72 against Stockfish, without losing once; against Elmo it won 90, drew 2 and lost 8. [1]

Atlas interpretation: The result was not won by brute force. The paper reports AlphaZero searched roughly 80,000 positions per second in chess and 40,000 in shogi, against Stockfish's 70 million and Elmo's 35 million, a search a thousand times smaller. What replaced the missing search was a network deciding which few lines were worth looking at, closer to a strong human's selectivity than to the exhaustive alpha-beta search every prior top engine used. [1]

A year arguing about a match

Chess grandmasters Hikaru Nakamura and Larry Kaufman objected to the terms within a day of the paper's release. Stockfish had played without an opening book it normally relies on, and Nakamura called the match "dishonest," while Kaufman argued AlphaZero had effectively built its own opening book during training and so was not facing an equivalent opponent. [3]

A year later, DeepMind published the full study in Science with a revised evaluation: a 1,000-game match in which Stockfish was given an opening book and DeepMind reported a result of 155 wins, 6 losses and 839 draws for AlphaZero, alongside a 91.2 percent win rate against Elmo. Training time to reach the reported peak strength was given as roughly 9 hours for chess, 12 for shogi and 13 days for Go. [2]

Atlas interpretation: The objection and the revision are the same argument playing out on two different clocks. Nakamura and Kaufman were right that the 2017 conditions handed Stockfish a specific, correctable disadvantage, and DeepMind's own fix a year later, adding the book back and widening the match to 1,000 games, is close to an admission of the point. It did not change the conclusion, since AlphaZero still lost only six games with the book restored, but it is worth noticing that the number the field settled on to cite is the 2018 figure, not the one in this paper. [3][2]

Sources

  1. Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

    arXiv · Dec 5, 2017

  2. AlphaZero: Shedding new light on chess, shogi, and Go

    Google DeepMind · Dec 6, 2018

  3. Google's AlphaZero Destroys Stockfish In 100-Game Match

    Chess.com · Dec 6, 2017