Trained on nothing but the rules
DeepMind announced AlphaGo Zero on October 18, 2017. Unlike earlier AlphaGo versions, which trained on a database of human expert games before self-play, this version started from random play with only the rules of Go as input. It learned entirely through games against itself. Its predecessor had beaten Lee Sedol in 2016 after training partly on human game records; AlphaGo Zero skipped that step. [1]
The Nature paper describes a single neural network that predicts both a move and a position's likely winner, trained through a search procedure that plays out possible move sequences and adjusts the network toward the outcomes those searches favor. The paper reports the network improved over roughly 3 days and 4.9 million self-play games on this reduced pipeline. [2]
The reported results
DeepMind reported that AlphaGo Zero beat AlphaGo Lee, the version that had defeated Lee Sedol, by 100 games to 0 in an internal evaluation match. Against AlphaGo Master, a stronger, previously undisclosed version that had beaten top human professionals online, the paper reports a 89 to 11 result over 100 games. [2][1]
Atlas interpretation: These figures come from matches DeepMind ran and reported itself, not an independently refereed tournament. That does not make them unreliable, but it means the numbers describe DeepMind's own evaluation setup and hardware rather than a result a third party reproduced. [2]
What the result argued for
Atlas interpretation: The result was notable less for the win margin than for what training used. Prior AlphaGo systems leaned on a curated set of human games as a starting point, which suggested a ceiling tied to the quality of that data. AlphaGo Zero's self-play approach, reaching a stronger level without any human game records, argued that for a domain with clear rules and a fast simulator, human demonstrations were a scaffold the system could discard once it existed. [2]
The paper also reports that AlphaGo Zero used a single neural network and a simpler search than earlier versions, and ran on far less specialized hardware: four tensor processing units, compared with the larger distributed setups used previously. DeepMind described the algorithm as a general reinforcement-learning method rather than one built for Go specifically, and later reused the same approach for chess and shogi. [1][2]
Sources
- AlphaGo Zero: Starting from scratch
Google DeepMind · Oct 18, 2017
- Mastering the game of Go without human knowledge
Nature · Oct 19, 2017