iGaming Journalist & Crypto Casino Analyst
Poker was, for two decades, the benchmark that artificial intelligence could not clear. Chess fell in 1997. Go fell in 2016. Heads-up no-limit hold'em held out until 2017, and six-handed no-limit — the format most online players actually play — held out until 2019. The delay was not a matter of computing power. It was structural: poker is a game of imperfect information, and the techniques that beat chess and Go do not transfer to it.
This is a complete reference timeline of poker AI research, with the published results, win rates, sample sizes and compute costs for every major milestone. Every figure below comes from peer-reviewed publications or the research groups that produced the systems.
Key Findings
Cepheus essentially solved heads-up limit hold'em in 2015, with an exploitability of 0.000986 big blinds per game. DeepStack and Libratus both beat professionals at heads-up no-limit in 2017, Libratus by 14.7 big blinds per 100 hands over 120,000 hands at 99.98% statistical significance. Pluribus beat 15 professionals at six-handed no-limit in 2019, winning 48 milli-big-blinds per game — and cost roughly $144 to train.
Why Poker Was Harder Than Chess or Go
In chess, both players see the entire board. The game has what researchers call perfect information, and it can be decomposed: the best move in a position can be computed by looking only at that position and what follows it. Search algorithms exploit this ruthlessly.
Poker cannot be decomposed that way. Your correct strategy on the river depends on the range of hands you would have arrived there with, which depends on how you played earlier streets, which depends on what your opponent believes about you. A "position" in poker is not a state of the cards; it is a state of the cards plus a probability distribution over everything hidden. Change the strategy in one part of the game tree and the value of every other part changes with it.
Three properties made no-limit hold'em specifically difficult:
- Hidden information. Optimal play requires randomisation. A strategy that never bluffs is exploitable; so is one that bluffs too often. The solution concept is a Nash equilibrium over mixed strategies, not a single best move.
- Continuous action space. No-limit betting allows any wager up to the stack. Even after abstraction, the game tree for heads-up no-limit hold'em is estimated at roughly 10161 decision points.
- Multiplayer dynamics. With three or more players, the guarantees attached to two-player zero-sum equilibria disappear entirely. Nothing in the theory promises that playing an equilibrium strategy at a six-handed table wins money.
The Complete Timeline
| Year | System | Institution | Game | Result |
|---|---|---|---|---|
| 1997-2003 | Loki / Poki | University of Alberta | Limit hold'em | First credible computer opponents; beat weak humans |
| 2003 | PsOpti | University of Alberta | Heads-up limit | First game-theoretic approach via abstraction |
| 2007 | Polaris | University of Alberta | Heads-up limit | Narrowly lost the first Man-vs-Machine match |
| 2008 | Polaris 2 | University of Alberta | Heads-up limit | Narrowly won the rematch — first AI win over pros |
| 2015 | Cepheus | University of Alberta | Heads-up limit | Game essentially solved; published in Science |
| 2015 | Claudico | Carnegie Mellon | Heads-up no-limit | Lost $732,713 over 80,000 hands (~9 bb/100) |
| 2017 | DeepStack | University of Alberta | Heads-up no-limit | Beat 33 pros by ~490 mbb/game |
| 2017 | Libratus | Carnegie Mellon | Heads-up no-limit | Beat 4 elite pros by 14.7 bb/100 over 120,000 hands |
| 2019 | Pluribus | Facebook AI / CMU | Six-handed no-limit | Beat 15 pros at 48 mbb/game; ~$144 training cost |
| 2020 | ReBeL | Facebook AI | Heads-up no-limit | General RL+search framework across imperfect-information games |
1997-2008: The Alberta Era
The University of Alberta's Computer Poker Research Group built the field. Loki and its successor Poki were rule-and-simulation systems that could beat weak human players but were badly exploitable. PsOpti in 2003 was the conceptual breakthrough: rather than trying to model opponents, it approximated a game-theoretic equilibrium on a simplified version of the game, then played that strategy back in the full game.
Polaris took the approach to a public test. In 2007 it lost narrowly to human professionals in the first Man-vs-Machine Poker Championship. In 2008 the rematch went the other way, making Polaris the first poker program to beat a team of professionals in a meaningful match — in heads-up limit hold'em, the simplest competitive variant.
2015: Cepheus Solves a Game
In January 2015, Alberta published "Heads-up limit hold'em poker is solved" in Science. Cepheus was not merely strong; it was, in the paper's terms, essentially unbeatable. The best possible counter-strategy could win only 0.000986 big blinds per game against it in expectation — a figure so small that a human playing a lifetime of hands could not distinguish it from zero.
This was the first competitively played imperfect-information game to be essentially solved. The caveat is important for citation: heads-up limit hold'em has a fixed bet size and two players. It is many orders of magnitude smaller than the no-limit games humans mostly play.
2015: Claudico Loses — Usefully
Four months later, Carnegie Mellon's Claudico faced Doug Polk, Bjorn Li, Dong Kim and Jason Les in the first Brains vs. Artificial Intelligence match. Over 80,000 hands, the humans finished ahead by $732,713 in virtual chips, roughly 9 big blinds per 100 hands. CMU characterised the margin as within statistical noise; the pros disagreed, and the pros were closer to right about the direction.
Claudico's weaknesses were instructive: erratic bet sizing, including tiny bets and enormous overbets that human opponents learned to attack, and no ability to refine its strategy mid-hand. Both failures were addressed directly in the next generation.
2017: The Year No-Limit Fell Twice
DeepStack, published in Science in early 2017, played 44,852 hands against 33 professionals recruited through the International Federation of Poker. It won at roughly 490 milli-big-blinds per game — around 49 big blinds per 100 hands — with the result statistically significant. Its innovation was continual re-solving: rather than precomputing a full strategy, DeepStack solved each decision as it arose, using deep neural networks to estimate the value of future states.
Weeks later, Libratus faced four elite heads-up specialists — Jason Les, Dong Kim, Daniel McAulay and Jimmy Chou — over 120,000 hands at the Rivers Casino in Pittsburgh. It won by 14.7 big blinds per 100 hands, roughly $1.77 million in chips, at 99.98% statistical significance. Libratus combined three components: a precomputed blueprint strategy, nested subgame solving that recomputed strategy in real time as hands progressed, and a self-improvement module that analysed its own most-exploited bet sizes overnight and patched them before the next day's play.
That overnight patching is the detail most often left out of summaries, and it was decisive. The pros found leaks; Libratus closed them faster than they could find new ones.
2019: Pluribus and the Multiplayer Barrier
Pluribus, from Facebook AI Research and Carnegie Mellon, addressed the format that mattered commercially: six-handed no-limit hold'em. It played two experiment formats — one AI against five humans, and one human against five copies of the AI — against a pool of 15 professionals including Chris Ferguson, Darren Elias, Jimmy Chou and Greg Merson.
In the five-humans-plus-one-AI format, Pluribus won 48 milli-big-blinds per game with a standard error of 25, across 10,000 hands played over 12 days. That is roughly 5 big blinds per 100 hands — a rate no human sustains against opposition of that quality.
Pluribus also broke the assumption that superhuman poker required a supercomputer. Its blueprint strategy was computed in eight days on a 64-core machine using under 512GB of RAM, at a market cloud cost of approximately $144. Libratus, by contrast, consumed millions of core-hours on the Pittsburgh Supercomputing Center's Bridges system. The cost curve for a given capability collapsed by orders of magnitude in two years.
Compute Economics: The $144 Superhuman
| System | Year | Training hardware | Approximate training cost |
|---|---|---|---|
| Libratus | 2017 | Bridges supercomputer, millions of core-hours | Institutional supercomputer allocation |
| Pluribus | 2019 | 64-core server, <512GB RAM, 8 days | ~$144 at cloud market rates |
The reason Pluribus was so cheap is the most transferable lesson in this literature. Pluribus did not compute a better full strategy; it computed a coarse blueprint and then spent its intelligence at decision time, searching only the part of the game tree in front of it. Noam Brown, who worked on both Libratus and Pluribus, has since argued that this trade — less pre-computation, more search when the decision is actually faced — generalises well beyond poker. It is the same principle behind the test-time reasoning approaches now central to large language model research.
From the Lab to the Tables
Published poker AI research was never released as a playable product, but the techniques diffused. Commercial solvers built on counterfactual regret minimisation put near-equilibrium strategy in the hands of ordinary players, changing how the game is studied. Understanding the difference between equilibrium play and adjusting to opponents is now foundational, as covered in our guide to GTO versus exploitative strategy.
The same diffusion created an enforcement problem. Real-time assistance software and bots now use descendants of these methods against unwitting opponents, and operators have responded with machine-learning detection trained on hand-history data. Reported enforcement actions in 2025 and 2026 include PokerStars flagging over 3,000 suspicious accounts with 890 resulting in bans after review, and GGPoker's Poker Integrity Council banning 42 accounts in a single investigation involving roughly $1.2 million in seized funds.
The asymmetry is worth stating: the research that proved poker is beatable by machines is public, peer-reviewed and freely readable, while the detection systems protecting online games are proprietary and unpublished. Players evaluating where to play can compare integrity policies via our poker site reference, and those building a study routine will find the fundamentals in our beginner learning path.
What Poker AI Contributed to the Wider Field
Three contributions outlived the poker results themselves. Counterfactual regret minimisation and its Monte Carlo variants became standard tools for imperfect-information games generally. Depth-limited and nested subgame solving showed that search could be made sound in games where a subgame has no well-defined value in isolation. And ReBeL, published in 2020, unified reinforcement learning with search into a single framework applicable to a broad class of imperfect-information games, with poker as the demonstration rather than the destination.
Poker's lasting role in AI research is as the domain where the field learned to handle hidden information, uncertainty and deliberate deception — conditions that describe negotiation, auctions, security and most real-world strategic interaction far better than chess does.
Methodology
Every result in this timeline is taken from the primary publication or the announcing research institution: Science for Cepheus, DeepStack and Pluribus; Carnegie Mellon University releases and contemporaneous reporting for Claudico and Libratus; and University of Alberta Computer Poker Research Group publications for the earlier systems.
Win rates are reported in the units used by the original papers. Milli-big-blinds per game (mbb/game) and big blinds per 100 hands (bb/100) are converted approximately for comparison — 100 mbb/game is roughly 10 bb/100 — but direct cross-match comparison is unsound. Match conditions varied substantially in opponent strength, sample size, stakes and whether variance-reduction techniques such as duplicate hands were used. DeepStack's higher headline win rate than Libratus reflects a different and generally weaker opponent pool, not a stronger system.
Compute cost figures are as reported by the research teams at the cloud market rates prevailing at the time and are not adjusted for subsequent price changes. No figures have been estimated by DeucesCracked.
Frequently Asked Questions
Has poker been solved by computers?
Heads-up limit hold'em has been essentially solved — Cepheus achieved an exploitability of 0.000986 big blinds per game in 2015. No-limit hold'em, in either heads-up or six-handed form, has not been solved. AI systems beat top humans at it, but that is superhuman performance, not a solution.
How much did Libratus beat the pros by?
Libratus won 14.7 big blinds per 100 hands across 120,000 hands against Jason Les, Dong Kim, Daniel McAulay and Jimmy Chou in 2017, equivalent to about $1.77 million in chips, with 99.98% statistical significance.
Why was six-player poker harder than heads-up?
Two-player zero-sum games have a theoretical guarantee: playing a Nash equilibrium cannot lose in the long run. With three or more players that guarantee vanishes. Pluribus succeeded by using a blueprint strategy plus real-time search rather than by relying on equilibrium guarantees.
Did Pluribus really cost only $144 to train?
Its blueprint strategy was computed in eight days on a 64-core server using under 512GB of RAM, which the researchers valued at roughly $144 at then-current cloud rates. That figure covers blueprint training, not the full research programme behind it.
Can I play against these poker AIs?
No. None of Cepheus, Libratus, DeepStack or Pluribus was released as a public product, and the Pluribus code was withheld specifically because of the risk to online poker ecosystems. Commercial solvers use related techniques for off-table study, but using any such tool during live play violates the terms of every licensed operator.
Sources
- Science — Superhuman AI for multiplayer poker (Pluribus)
- Science — DeepStack: Expert-level AI in heads-up no-limit poker
- Communications of the ACM — Heads-Up Limit Hold'em Poker Is Solved
- Carnegie Mellon University — Libratus match results
- PokerNews — Brains vs. AI: Claudico results
Cite This Article
If you use data from this article, please link back to https://www.deucescracked.com/blog/poker-ai-research-timeline-data. Related reference material is collected in our poker strategy hub.
Related Guides
Join the Conversation
Be respectful. No spam. Strategy discussion welcome.