Gerald Tesauro of IBM's Thomas J. Watson Research Center created TD-Gammon in 1992, a neural network trained by temporal-difference learning on games it played against itself. Version 2.1, trained on 1.5 million self-play games, played just below the level of the top human backgammon players, and its unconventional opening choices changed expert practice — a precursor of AlphaGo Zero.
The question, scope, and sources behind this Registry record.
Gerald Tesauro of IBM's Thomas J. Watson Research Center created TD-Gammon in 1992, a neural network trained by temporal-difference learning on games it played against itself. Version 2.1, trained on 1.5 million self-play games, played just below the level of the top human backgammon players, and its unconventional opening choices changed expert practice — a precursor of AlphaGo Zero.
The year Gerald Tesauro's TD-Gammon showed that temporal-difference learning from self-play could approach world-class play.
Change a parameter to stress-test whether a proposed result is still inside the published specification. This is an audit aid, not a proof checker.
This record has no editable parameters. Read the formal question and assumptions before challenging it.
Current frontiers derived from accepted Claims.
An observed or demonstrated result; no opposing bound is implied.
The frontier is not sacred
Most progress starts with a disagreement that survives contact with evidence. If you can push the known lower bound up or pull the upper bound down, show us the work.
≥ when you have shown that at least this value is achievable.≤ when you have shown that anything above this value is impossible.No vibes. State the value, define the scope, and link the paper, proof, code, or reproduction that lets another person check it. Editors review every challenge before the public record changes.
Challenge this recordAssertions tied to evidence, attribution, and review.
The frontier as it changed over time.
Only accepted Claims matching the current specification contribute to the displayed bounds. Strict inequalities remain open; contradictory Claims require editorial review.
1 accepted Claim, with 1 linked evidence records.
Permanent ID limitsregistry.com/limits/LR-TD-GAMMON-SELF-TAUGHT-BACKGAMMON
No active verified bounties are linked to this Limit.
View verified bounty tracker ↗No accepted machine-checked reproductions are recorded for this Limit.
Limits Registry. LR-TD-GAMMON-SELF-TAUGHT-BACKGAMMON. First neural network to reach near-champion play by self-play (TD-Gammon). 2026.