AlphaGo vs Lee Sedol: Match Result & Move 37 Explained

DeepMind's AlphaGo beat Lee Sedol four games to one. The page explains its move-selection and value networks, Monte Carlo tree search and surprising Move 37.

Deep Blue picked chess moves by evaluating positions exhaustively. Its 32 processors and dedicated chess chips searched roughly 200 million positions a second, extending lines deep enough that a hand-built evaluation function, sharpened with advice from grandmasters, could rank what survived. The search was brute force: width and depth stood in for judgment, not a trained sense of which moves mattered. [5]

AlphaGo replaced the hand-built evaluation function with two trained networks. A policy network proposed a short list of promising moves; a value network scored a resulting position without playing it out. Both were trained first on roughly 30 million positions from human expert games, then refined by reinforcement learning from self-play against earlier versions of themselves. [2]

Monte Carlo tree search narrows the tree instead of widening it

Atlas interpretation: Deep Blue's search explored the legal replies at every ply and pruned branches once they looked lost, so the machine still had to touch most of the board before discarding it. AlphaGo's Monte Carlo tree search worked the other way: the policy network's move probabilities decided where simulations were even run, and the value network, combined with rollout outcomes, updated those branches' estimated worth as play continued. Depth came from repeatedly walking the most promising lines further, not from covering the board more completely. [2][5]

Move 37, mechanically

AlphaGo won the five-game match four games to one, finishing in Seoul on March 15, 2016. In game two, it placed a stone on the fifth line near the edge, a move commentators first read as an error. DeepMind's own account puts the move's probability of appearing in a human game at roughly 1 in 10,000. Lee Sedol left the room; when he returned, he said he had taken AlphaGo for pure probability calculation until he saw it. [3][1]

The supervised-learning policy network, trained only on human games, assigned that move almost no chance of being played, exactly because no human precedent favored it. The value network, shaped instead by self-play, rated the positions it led to highly. Tree search accumulated enough simulated strength along that line to promote a move its own human-imitating half had all but ruled out. [4]

What the comparison actually shows

Atlas interpretation: Deep Blue could not have produced an equivalent moment. Its evaluation function was fixed by its engineers before the match, so a deeper search only confirmed judgments already built in; searching wider finds a better example of what the program already values, not a different notion of value. AlphaGo's networks derived their sense of a good position from self-played games rather than a programmer's rule, so search could surface a move a human-authored evaluation function would never have scored well. The two programs differed less in how far they searched than in where the judgment guiding that search came from. [2][5]

Sources

  1. AlphaGo

    Google DeepMind

  2. Mastering the game of Go with deep neural networks and tree search

    Nature · Jan 27, 2016

  3. Google AI beats Go world champion again to complete historic 4-1 series victory

    TechCrunch · Mar 15, 2016

  4. Was AlphaGo's Move 37 Inevitable?

    Katherine Bailey · Jan 23, 2017

  5. Deep Blue

    IBM · Sep 9, 2026