Blackjack card counting: learned strategy tables
A benchmark of blackjack card counting, for 5 Las Vegas games. Two kinds of agents are compared:
- Counting systems: tabular Q-learning agents, one per system (Hi-Lo, Hi-Opt I, Hi-Opt II, Omega II, Zen Count, Wong Halves, Knock-Out). Each sees the shoe through its count and learns the best play at every value of it.
- DQN agents with no predefined count (DQN linear, DQN 2Γ64): they learn from the full composition of the shoe, how many of each rank are left, and approach the optimal gain. They are the yardstick: the gap between a count and the DQN agents measures what a single count leaves on the table. DQN 2Γ64 adds a small neural network to DQN linear, and gains almost nothing over it, so DQN linear is close to the optimum.
No basic strategy or index table was given: every play, index play and insurance decision was learned by self-play in a simulated shoe, with every player at the table playing the same agent.
Code, method and validation: https://github.com/apicquot/blackjack. Interactive results: https://huggingface.co/spaces/apicquot/blackjack-counting-results.
Results at a glance
Best counting system per game (the DQN agents are in each game's page and CSV). Money is in units won per 100 rounds dealt (1 unit = the minimum bet), except the last column: units won per 100 hands actually played, since 0β3 sits out the rounds without a player edge. Betting: 3 units at counts with a player edge, otherwise 1 (1β3) or nothing (0β3).
| game | best count | no count, flat bet | best count, flat bet | bet 1β3, play every round | bet 0β3, sit out bad counts | 0β3: units won per 100 hands played |
|---|---|---|---|---|---|---|
| 1 deck, H17, DAS, 6:5 | Omega II | β1.59 | β0.88 | β0.05 | +1.25 | +4.89 |
| 2 decks, H17, DAS, 3:2 | Zen Count | β0.49 | β0.16 | +0.98 | +1.71 | +5.41 |
| 6 decks, H17, DAS, surrender, 3:2 | Wong Halves | β0.54 | β0.40 | +0.27 | +1.01 | +2.79 |
| 6 decks, S17, DAS, surrender, RSA, 3:2 (high limit) | Wong Halves | β0.27 | β0.13 | +0.72 | +1.27 | +3.50 |
| 8 decks, H17, DAS, surrender, 3:2 | Wong Halves | β0.56 | β0.47 | +0.03 | +0.74 | +2.12 |
Files
For each game, a folder <game>/ with:
config.yaml: the rules (decks, penetration, soft 17, doubling, splits, surrender, payout, players)q_<system>.npz: the trained table of one count systemdqn_<model>.npz: the benchmark agents without a predefined count, DQNs on the full shoe composition:dqn_linear(linear model) and, for single deck,dqn_2x64(plus a network with two hidden layers of 64).W: advantage weights per decision,mlp: the network part,Wr: predicted edge before the deal; load withblackjack.dqn.ModelAgent.loadsummary.csv,by_seat.csv: the results by count system and by seat
Each .npz holds:
| array | shape | meaning |
|---|---|---|
Q |
(count bucket, dealer up β 1, hand, two cards, action) | expected gain in initial bets of each action, then optimal play |
N |
same | visits; 0 means the action was never taken (illegal or unreached) |
Qi, Ni |
(count bucket,) | value of insurance, and visits |
buckets |
(lo, width, n) | count bucket k covers [lo + kΒ·width, lo + (k+1)Β·width); the ends are open |
count |
counting system; rule_* fields: the rules |
- Hands: index 0β17 is hard 4β21, 18β27 is soft 12β21, and 28β37 are splittable pairs A,A, 2,2 β¦ 10,10.
- Actions: stand, hit, double, split, surrender.
- Advantage:
Q β max Q, what each action costs against the best one.
Use
from huggingface_hub import hf_hub_download
from blackjack.qlearn import QTable # pip install git+https://github.com/apicquot/blackjack
qt = QTable.load(hf_hub_download("apicquot/blackjack-counting", "6deck_h17_das_ls/q_hi_lo.npz"))
qt.state(count=2, up=10, hand="H16") # Q, visits and advantage of each action
print(qt.chart(count=2)) # strategy chart at true count +2