← New search

Other meanings of AlphaGo

Artificial Intelligence

AlphaGo

AlphaGo is a computer program developed by DeepMind Technologies that plays the board game Go, and in 2016 it became the first artificial intelligence to defeat a professional human Go player without handicap. Its victory over Lee Sedol in a five-game match marked a watershed moment in artificial intelligence research, demonstrating that deep neural networks and reinforcement learning could master a game of immense combinatorial complexity.

2016
Year of historic match vs Lee Sedol
4–1
Match score vs Lee Sedol
10^170
Approximate number of possible board positions
37%
Win rate of AlphaGo's policy network vs. supervised learning baseline
1

Development and architecture

AlphaGo was developed by DeepMind, a London-based AI company acquired by Google in 2014. The program combined two deep neural networks: a policy network that selects the next move and a value network that evaluates board positions. These networks were trained using a three-stage process: supervised learning from human expert games, reinforcement learning through self-play, and Monte Carlo tree search (MCTS) to simulate future positions.1

The policy network was first trained on 30 million positions from human games, achieving 57% accuracy in predicting human moves. Then, through reinforcement learning, the network played against itself millions of times, improving its strategy. The value network learned to predict the winner from any position, reducing the need for exhaustive search. This architecture allowed AlphaGo to evaluate positions with human-like intuition rather than brute-force calculation.2

2

Historic matches

In October 2015, AlphaGo defeated Fan Hui, the European Go champion, 5–0 in a closed match, becoming the first computer program to beat a professional Go player without handicap. This result was published in Nature in January 2016.1

The most famous match occurred in March 2016, when AlphaGo faced Lee Sedol, one of the greatest Go players of all time, in Seoul. AlphaGo won four of five games, with Lee Sedol winning game four with a brilliant move that was later dubbed “God’s hand.” The match was watched by over 200 million people worldwide and sparked intense discussion about the capabilities of artificial intelligence.3

3

Technical innovations

AlphaGo introduced several innovations that were later adopted in other AI systems. One was the use of two separate networks, which allowed the program to balance exploration and exploitation. Another was the use of a “tree search” that combined the policy and value networks to guide simulations, reducing the search space dramatically.2

Notably, AlphaGo’s move 37 in game two against Lee Sedol was a shoulder hit in the upper right corner, a move that human experts initially considered a mistake but later recognized as a creative and powerful play. This move was praised for its originality and has been studied by Go professionals since.3

4

Impact and legacy

AlphaGo’s success had a profound impact on both the game of Go and the field of artificial intelligence. It demonstrated that deep reinforcement learning could solve problems previously thought to be decades away from solution. The techniques developed for AlphaGo were later applied to other domains, including protein folding (AlphaFold) and quantum chemistry.4

In the Go community, AlphaGo’s play has changed the way professionals approach the game, leading to new opening strategies and a deeper understanding of the game’s complexity. The match also raised public awareness of AI, prompting discussions about the future of human-machine collaboration and the ethical implications of advanced AI systems.

5

Lesser-known aspects

AlphaGo’s development involved a team of about 20 researchers, including David Silver, Aja Huang, and Demis Hassabis. Aja Huang, a former Go amateur, was the one who physically placed the stones on the board during the matches against Lee Sedol.

After the Lee Sedol match, DeepMind released a paper detailing AlphaGo’s algorithms, but the full code was not open-sourced. However, a later version, AlphaGo Zero, was described in Nature in 2017, which learned entirely from self-play without any human data, starting from random play and achieving superhuman performance in 40 days.4

AlphaGo’s victory also had a cultural impact in South Korea, where Go is a national pastime. The match was broadcast live on television, and Lee Sedol’s defeat was met with both shock and admiration. In 2019, Lee Sedol retired from professional Go, citing the rise of AI as a factor in his decision.

Glossary

Deep neural network
A machine learning model with multiple layers that can learn complex patterns from data.
Reinforcement learning
A type of machine learning where an agent learns to make decisions by receiving rewards or penalties.
Monte Carlo tree search
A heuristic search algorithm used in decision-making processes, particularly in games.
Policy network
A neural network that maps board positions to probabilities of selecting each possible move.
Value network
A neural network that estimates the expected outcome (win/loss) from a given board position.

AlphaGo's victory over Lee Sedol is often compared to IBM's Deep Blue defeating Garry Kasparov in chess in 1997, but Go's complexity made the achievement far more significant.