Scinovex
articleTop 10% cited

TD-Gammon, a Self-Teaching Backgammon Program, Achieves Master-Level Play

Neural Computation · 1994 · Vol. 6(2) · pp. 215–219
Gerald Tesauro

Abstract

TD-Gammon is a neural network that is able to teach itself to play backgammon solely by playing against itself and learning from the results, based on the TD(λ) reinforcement learning algorithm (Sutton 1988). Despite starting from random initial weights (and hence random initial strategy), TD-Gammon achieves a surprisingly strong level of play. With zero knowledge built in at the start of learning (i.e., given only a “raw” description of the board state), the network learns to play at a strong intermediate level. Furthermore, when a set of hand-crafted features is added to the network's input representation, the result is a truly staggering level of performance: the latest version of TD-Gammon is now estimated to play at a strong master level that is extremely close to the world's best human players.

Artificial Intelligence in GamesReinforcement Learning in RoboticsEvolutionary Algorithms and ApplicationsReinforcement learningSet (abstract data type)Representation (politics)Computer scienceArtificial intelligenceArtificial neural network
Citations
789
FWCI
11.88
field-weighted impact
References
5
Percentile
99%
vs. same field & year
Citations per year
Cited by
Predictive Reward Signal of Dopamine Neurons
Journal of Neurophysiology · 1998 · 4,543 citations
Reinforcement Learning in Continuous Time and Space
Neural Computation · 2000 · 979 citations
Online learning control by association and reinforcement
IEEE Transactions on Neural Networks · 2001 · 774 citations
Action understanding as inverse planning
Cognition · 2009 · 933 citations
Getting Formal with Dopamine and Reward
Neuron · 2002 · 2,583 citations
Dopamine neurons and their role in reward mechanisms
Current Opinion in Neurobiology · 1997 · 765 citations
Metalearning and neuromodulation
Neural Networks · 2002 · 694 citations
References
Learning to Predict by the Methods of Temporal Differences
Machine Learning · 1988 · 3,908 citations
Practical issues in temporal difference learning
Machine Learning · 1992 · 795 citations
Learning to predict by the methods of temporal differences
Machine Learning · 1988 · 2,774 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.