Blackjack card counting: learned strategy tables

A benchmark of blackjack card counting, for 5 Las Vegas games. Two kinds of agents are compared:

  • Counting systems: tabular Q-learning agents, one per system (Hi-Lo, Hi-Opt I, Hi-Opt II, Omega II, Zen Count, Wong Halves, Knock-Out). Each sees the shoe through its count and learns the best play at every value of it.
  • DQN agents with no predefined count (DQN linear, DQN 2Γ—64): they learn from the full composition of the shoe, how many of each rank are left, and approach the optimal gain. They are the yardstick: the gap between a count and the DQN agents measures what a single count leaves on the table. DQN 2Γ—64 adds a small neural network to DQN linear, and gains almost nothing over it, so DQN linear is close to the optimum.

No basic strategy or index table was given: every play, index play and insurance decision was learned by self-play in a simulated shoe, with every player at the table playing the same agent.

Code, method and validation: https://github.com/apicquot/blackjack. Interactive results: https://huggingface.co/spaces/apicquot/blackjack-counting-results.

Results at a glance

Best counting system per game (the DQN agents are in each game's page and CSV). Money is in units won per 100 rounds dealt (1 unit = the minimum bet), except the last column: units won per 100 hands actually played, since 0–3 sits out the rounds without a player edge. Betting: 3 units at counts with a player edge, otherwise 1 (1–3) or nothing (0–3).

game best count no count, flat bet best count, flat bet bet 1–3, play every round bet 0–3, sit out bad counts 0–3: units won per 100 hands played
1 deck, H17, DAS, 6:5 Omega II βˆ’1.59 βˆ’0.88 βˆ’0.05 +1.25 +4.89
2 decks, H17, DAS, 3:2 Zen Count βˆ’0.49 βˆ’0.16 +0.98 +1.71 +5.41
6 decks, H17, DAS, surrender, 3:2 Wong Halves βˆ’0.54 βˆ’0.40 +0.27 +1.01 +2.79
6 decks, S17, DAS, surrender, RSA, 3:2 (high limit) Wong Halves βˆ’0.27 βˆ’0.13 +0.72 +1.27 +3.50
8 decks, H17, DAS, surrender, 3:2 Wong Halves βˆ’0.56 βˆ’0.47 +0.03 +0.74 +2.12

Files

For each game, a folder <game>/ with:

  • config.yaml: the rules (decks, penetration, soft 17, doubling, splits, surrender, payout, players)
  • q_<system>.npz: the trained table of one count system
  • dqn_<model>.npz: the benchmark agents without a predefined count, DQNs on the full shoe composition: dqn_linear (linear model) and, for single deck, dqn_2x64 (plus a network with two hidden layers of 64). W: advantage weights per decision, mlp: the network part, Wr: predicted edge before the deal; load with blackjack.dqn.ModelAgent.load
  • summary.csv, by_seat.csv: the results by count system and by seat

Each .npz holds:

array shape meaning
Q (count bucket, dealer up βˆ’ 1, hand, two cards, action) expected gain in initial bets of each action, then optimal play
N same visits; 0 means the action was never taken (illegal or unreached)
Qi, Ni (count bucket,) value of insurance, and visits
buckets (lo, width, n) count bucket k covers [lo + kΒ·width, lo + (k+1)Β·width); the ends are open
count counting system; rule_* fields: the rules
  • Hands: index 0–17 is hard 4–21, 18–27 is soft 12–21, and 28–37 are splittable pairs A,A, 2,2 … 10,10.
  • Actions: stand, hit, double, split, surrender.
  • Advantage: Q βˆ’ max Q, what each action costs against the best one.

Use

from huggingface_hub import hf_hub_download
from blackjack.qlearn import QTable        # pip install git+https://github.com/apicquot/blackjack

qt = QTable.load(hf_hub_download("apicquot/blackjack-counting", "6deck_h17_das_ls/q_hi_lo.npz"))
qt.state(count=2, up=10, hand="H16")       # Q, visits and advantage of each action
print(qt.chart(count=2))                   # strategy chart at true count +2
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Space using apicquot/blackjack-counting 1