TanyaoDojo β Mahjong AI checkpoints
JAX/Flax checkpoints from TanyaoDojo, a
from-scratch mahjong AI stack (vectorized env + behavior cloning + self-play RL).
Every number below comes from the same duplicate 1v3 protocol: the challenger
plays all four seats over identical walls against three copies of a strong
open-source baseline (Mortal v4), seed_key=20260711, seeds from 10000,
placement points [90, 45, 0, -135].
avg_pt is the challenger's mean placement points relative to the baseline;
0 would mean parity. Higher is better.
The headline checkpoint has a 100k-game milestone measurement
(-4.66 +/- 0.535, 16.0M decisions, zero fallbacks) β not a small-sample estimate.
Status, 2026-09-13. The project's strongest model is no longer in this repository's lineage. It is an online value-regression fine-tune on the Mortal stack with a Q-value anchor (loss weight 0.2) that we had shut down in July after reading only a small, unpaired internal test-play metric. Re-measured under the protocol above: -0.75 +/- 1.31 at 12k games, paired +1.75 (z = 2.58) over the previous best checkpoint (-2.50 on the same 12k games, -1.96 at 100k). Those weights are not published here; the anchor patch is in the GitHub repo (
patches/anchor_train.patch). The JAX line below has been paused after its gap was located (see "Where the remaining gap is").Update, 2026-09-14 β what you anchor to decides whether continued training holds. Resuming that model with its anchor still pointing at the earlier checkpoint pulled it back to the earlier checkpoint's play within two hours: riichi choice on 20,691 identical decision points went 41.1% -> 44.4% (the baseline itself: 39.0%), and at ~40k steps the 12k paired difference was -1.26 (z = -1.81). Re-anchoring to the good checkpoint itself held it: riichi choice 39.9%, 12k -0.56, paired +0.22 (z = 0.36). The two anchors compared directly at matched steps: +1.48 +/- 1.42 (z = 2.05). The 4k strength screens during the regression were all inside noise; the per-decision riichi rate saw it immediately.
Update, 2026-09-14 (afternoon) β the play changed, the strength did not, and the gap is in placement. Halving the anchor weight (0.2 -> 0.1) for 40k steps gave 12k -0.53 +/- 1.32, paired +0.25 (z = 0.38) vs the July checkpoint and +0.03 (z = 0.04) vs the end of the 0.2 phase. The style slid along an iso-strength line: deal-in rate fell below the baseline's for the first time (defense gap -28 -> -4 points/hand) while win rate and call rate dropped with it. With point flow nearly level, a placement decomposition located the remaining
0.5: every snapshot enters South 4 ahead of the baseline, games that reach South 4 roughly break even, and the loss sits in games that end early by someone busting (22% of games; -0.56 to -1.63 pt, same sign in 8 of 8 readings; bust rate +0.2 to +1.3 pp, 4th place in early-ended games +1.7 to +5.3 pp). Root-cause candidate: every training config used rank rewards[6, 4, 2, 0](gaps 2/2/2) while the protocol scores[90, 45, 0, -135](gaps 45/45/135) β training priced 4th place at a third of what we measure. A branch with rank rewards rescaled to the protocol's ratios is running now. Separately, the trainer had been bottlenecked since July by single-process data loading (GPU ~5% busy); fixing it cut a 4,000-step cycle from ~49 to ~17 minutes.Update, 2026-09-16 β matching the training reward to the metric moved every mechanism we predicted (on 8k games; see the 09-17 correction). Branching from the previous checkpoint and changing only the rank rewards from
[6, 4, 2, 0]to the protocol's ratios ([2.390, 1.195, 0, -3.585], same standard deviation), 40k steps later: 12k -0.20 +/- 1.33 β the best point estimate this project has recorded (previous best -0.53, the July checkpoint -0.78 on this same GPU-fp32 basis) β but paired differences of +0.33 (z = 0.46) vs the branch point and +0.59 (z = 0.83) vs the July checkpoint are not significant, so no record is claimed. The pre-registered mechanism predictions all landed (8k games, vs the baseline seats in the same games): bust rate 6.53% vs 6.89% before, now level with the baseline's 6.49%; 4th place in games that end early +2.0 -> +0.9 pp; 4th-place rate +1.30 -> +0.25 pp; 1st-place rate +1.15 -> +0.48 pp; and the placement decomposition's "converting position in the final hand" term -1.02 -> -0.00. Riichi rate is now below the baseline's (-0.77 pp) and call rate rose to 30.77%: a high-variance style traded for a steady one. Two diagnostics added for this:diag_place.py(splits avg_pt into position entering the final hand vs conversion, with games that end early by a bust broken out) anddiag_bust.py.Update, 2026-09-17 β the 100k measurement did not confirm it. Forty thousand steps further on (448k), the 12k estimate reached +0.56 +/- 1.30, the first positive 12k reading here. Following the rule set in advance (a 100k point estimate >= +0.01 counts as parity), it was extended to 100k: -0.93 +/- 0.46 (95% CI [-1.39, -0.47]) β significantly below the baseline. The two 12k segments happened to be its two best of five. Paired on the same 28k games it is +0.34 (z = 0.73) over the July checkpoint and indistinguishable from the 408k snapshot (-0.13). At 100k the placement decomposition is unambiguous: almost the entire gap (-0.92 +/- 0.26 of -0.93) sits in games that end early by a bust, while the model still enters South 4 ahead of the baseline. Between 408k and 448k the play kept sliding toward passivity (call rate 30.7% -> 27.3%, win rate -0.72 pp), and the bust gap came back. Lesson recorded: every remaining effect here is under 1 point, and a 12k estimate (+/- 1.3) has no discriminating power near zero β decisions now go to 100k or same-game pairing. The value line's previous 100k reading was -1.96 (rl1_best), so -0.93 is still its best 100k, about 1 point higher.
Correction, 2026-09-17 β the mechanism "hits" above were mostly 8k noise. The 408k snapshot at 100k: -0.64 +/- 0.46 (95% CI [-1.10, -0.18]); paired with 448k on the same 100k games, 448k - 408k = -0.30 +/- 0.49 (z = -1.17). The bust-rate gap, 4th-place gap and final-hand conversion that read +0.04 pp / +0.25 pp / -0.00 on 8k games read +0.28 pp / +0.68 pp / -1.13 on 100k β better than the July checkpoint in direction, but the loss in bust-ended games is not closed, and the claim that all six predictions landed is withdrawn. Whether the rank-reward change helped at all is being measured by extending its branch point to 100k. A second train/eval mismatch is being tested meanwhile: only 16.8% of self-play games were against the evaluation baseline; the rest were against two aggressive earlier checkpoints (riichi choice ~43.8% vs the baseline's 39.0%). A branch from 408k now trains against the baseline in every game, with everything else unchanged, to be paired with 448k at 100k.
Update, 2026-09-18 β the rank-reward change, measured properly: +0.25 +/- 0.48 (z = 1.01), not significant. The branch point (anchor 0.1, old rank rewards) extended to 100k scored -0.89 +/- 0.46. Paired on the same 100k games: 40k steps after the change +0.25 +/- 0.48, 80k steps after -0.05 +/- 0.49. At 100k the change did make placements less extreme (1st/4th-place rates vs the baseline went from +0.81/+1.35 pp to -0.23/+0.68 pp), but the loss in bust-ended games barely moved (final-hand conversion -1.19 -> -1.13), and 40k further steps undid the placement shift. It is kept as the default (it matches the metric and does no harm) but is no longer counted as a validated lever. The opponent-pool branch reaches its 100k checkpoint next.
Update, 2026-09-20 β starting from the baseline itself: +0.10 +/- 0.34 over 100k games. With every training knob on our own base coming out flat, the base was swapped for the champion's own weights (256x54, 23.7M) and fine-tuned with the same anchored value regression (anchor = the baseline itself, weight 0.2, all self-play against it). Here the challenger's
avg_ptis the net gain over the baseline, and the zero point was calibrated first: the baseline playing itself scores -0.056 +/- 0.091 with seat counts [1000, 998, 1001, 1001]. After 40k steps, 100k games give +0.102 +/- 0.337 (95% CI [-0.236, +0.439]) β above zero as a point estimate, statistically indistinguishable from it. Screens along the way (-0.08, -0.69, -0.69, -0.61, -0.14, then +0.59, -0.77) show no trend. Separating +0.10 from zero would need roughly 1.1M games (~55 GPU-hours), an order of magnitude beyond what this project can measure. The mechanism result is the useful part: starting from the champion's weights, the excess bust rate disappears (6.78% vs 6.76%, against +0.28 pp for our own base) and average deal-in cost lands at 5,147 vs the baseline's 5,174 β confirming that "busts more, deals into bigger hands" was a property of our own 192x40 lineage, not something the RL caused or can fix. Our from-scratch base still tops out at 100k -0.64; this line is "baseline + our RL", and is labelled that way.Update, 2026-09-30 β the +0.10 made significant, the remaining gap priced per decision, model size ruled out. Significance. The 40k-step fine-tune from the baseline was re-measured with a pre-registered sequential design (futility-only looks every 200k games, one success test at the end): +0.123 +/- 0.101 over 1.1M games (95% CI [+0.022, +0.225]). Fine-tuning the baseline with our RL does beat it, by a small margin. That is "baseline + our RL"; our own pipeline's best remains 100k -0.64 +/- 0.46. Where our own lineage loses. The arena derives every wall from the seed recorded in the log, so any logged decision can be replayed on the true wall with the baseline's action and played on; a telescoping identity then splits "our result - the baseline playing our seat" exactly into per-decision terms (tooling: riichi-eval, AGPL-3.0). On 100k games / 1.32M disagreements: pushing where the baseline folds against a riichi -0.15 +/- 0.08, riichi judgment about -0.18, ordinary tile choice (72% of disagreements) +0.04 +/- 0.34, i.e. nothing; total -0.43 +/- 0.40 vs the measured -0.64 +/- 0.46. Across the lineage the extra pushing appears with online RL (the offline base pushes and folds symmetrically), and the from-baseline branch shows none of it. Raw-point accounting points the wrong way on riichi decisions. Two targeted fixes that did not work. Flipping to fold wherever our own Q margin is small recovers nothing at any threshold, although folding once at the spots where the baseline folds gains +0.16 +/- 0.14: our Q cannot find those spots. Fine-tuning on 1.44M rollout-priced action pairs: independent hold-out +0.015 +/- 1.00 (single-rollout labels have SD ~15 pt against a ~0.4 pt signal). Model size ruled out. Our offline base retrained at the baseline's size (256x54 instead of 192x40), recipe unchanged: 100k -2.68 +/- 0.46, paired with the 192x40 base -0.18 +/- 0.62 (z = -0.58). More data, width plus hand-crafted danger features, and now capacity all come out flat; the offline recipe tops out near -2.5. The value target is not it either. Mortal-style training values each hand's end with a rank predictor trained on human games. On our tables it is off in two places (entering the final hand in 4th is 5-7 pt worse than predicted, dropping below 8,000 in the East round ~5 pt worse), yet on the same rollouts valued at game end vs at hand end, the per-decision difference is -0.000 +/- 0.045 pt (interim, 47k games). Next. The online value regression only learns from the action actually taken, with 0.5% exploration, so the untaken side of a push/fold decision keeps a stale value; that hypothesis is checked first, then targeted exploration at riichi-exposed decisions.
Checkpoints
| File | Arch | Obs | Training | avg_pt vs baseline |
|---|---|---|---|---|
bc_v2_g186.pkl |
256ch x 10blk (12.4M) | v2 (34x36 + 32) | BC on 10y logs + LR-1e-4 refine | -4.66 +/- 0.535 (100k games) |
bc_lean_g402.pkl |
256ch x 10blk (12.4M) | lean (34x20 + 26) | BC on 14y logs + LR-1e-4 refine | -5.07 +/- 1.54 (12k games) |
bc_lean_w192_ep2.pkl |
192ch x 8blk (5.3M) | lean | BC on 6y logs, 2 epochs | -8.87 +/- 2.70 (4k games) |
rl_oracle_800m.pkl |
256ch x 10blk | lean (actor) | oracle-critic PPO league, 0.8B steps | -8.38 +/- 2.65 (4k) β negative result, wrong objective (see below) |
The last row is published deliberately, and its interpretation was corrected on 2026-08-23 β the correction is more useful than the checkpoint.
An asymmetric actor-critic whose critic sees all four hands fit the value function ~100x better (v_loss 0.102 -> 0.001) yet lost 3.2pt of external strength. It was the fourth consecutive negative RL result here. We originally read this as "oracle critics don't transfer". That reading was wrong.
The real cause: the training objective was not the evaluation objective. All four
RL runs used round_mode="single", which ends the episode after a single hand, and
the reward is mahjong's raw point transfer. The arena scores an entire hanchan by
placement points [90, 45, 0, -135]. Measured directly, not inferred from code: 600
steps x 256 envs produced 1541 episodes (one hand each); rewards ranged [-120, 130]
with rows like [+30, -10, -10, -10] summing to ~0 (point transfers); and mahjax writes
order_points only into the final score, never into rewards, in any round_mode.
A single-hand point maximizer has no concept of 4th place β falling to last costs it nothing extra, so it should be maximally aggressive. This predicts that more RL steps make the arena result worse (observed), and it explains why a better critic hurt more: it converged faster onto the wrong objective.
So this checkpoint is not evidence against oracle critics. It is evidence that we
optimized the wrong function for ~10^10 environment steps. The fix β terminal
placement-aligned reward over full hanchan β is in the repo as
jax_rl/reward_placement.py, and all four RL results above are being redone on it.
First result after the fix (1B steps from the -4.66 base, 11 arena evals of 1600
games each along the way): -4.88 +/- 0.81 β statistically flat versus the base.
The previous four runs lost 3.2 to 10 points each and got worse the longer they ran.
The degradation is gone; the gains are not there yet, and the metrics say why β
approx_kl about 4e-5 per update and only 0.011 KL of drift from the base after a
billion steps, i.e. a tight trust region plus roughly 0.75 hanchan terminals per
rollout left the policy almost unmoved. The next step is therefore signal density,
not a different algorithm: GRP potential shaping (a 7.5M-sample "position -> expected
final placement points" potential, reward becomes phi(s') - phi(s) with the terminal
remainder), which is policy-invariant by the shaping theorem β verified numerically β
and raises the fraction of non-zero-reward steps from 0.1% to 1.10%.
If you are porting an RL recipe (ours or anyone's): check that its episode boundary and reward definition are the same function your benchmark scores.
Shaped run, first segment closed at 1.48B steps (2026-09-04). Drift from the base reached 0.023 KL versus 0.011 at 1B in the unshaped run β the policy moved about twice as far. Two checkpoints have been adjudicated at the full 12k-game protocol: -3.79 +/- 1.54 at 0.54B and -4.04 +/- 1.54 at 1.48B, i.e. +0.87 and +0.62 against the -4.66 base, at 1.05 and 0.74 sigma. Neither is significant. Pooling everything measured on this run (~36k games, six checkpoints, inverse-variance) gives -3.78 +/- 0.85, or +0.88 +/- 1.00 over the base β z = 1.73. Suggestive, not proven, and we are not going to call it more than that.
Both adjudications tell the same story about small samples: the 4k readings that preceded them (-3.50 and -3.57, and one wild -1.29 at 1600 games) were all optimistic, and all pulled back. That is now five for five here. Single-segment readings are for screening only; nothing gets claimed without the 12k double-segment protocol.
We also audited our own confidence intervals, and they were wrong β conservatively.
The eval harness computes its interval as 1.96 * std(per-game points) / sqrt(n) with n counting every game. But duplicate evaluation plays the same
deal four times with rotated seats, so those four games are not independent samples β
that dependence is the entire point of the design. Recomputing over 2,000 deals from a
real 8k-game segment gives an intra-class correlation of -0.079 and a deal-level
interval 0.873x the per-game one. The rotation really does cancel seat luck, so our
published intervals have been about 13% too wide: the true resolution is +/-2.36 at 4k,
+/-1.34 at 12k and +/-0.467 at 100k. Every significance call we have made was therefore
conservative, and none of them flips.
Getting that number required fixing a second bug in the audit script itself: mjai's
reach_accepted event carries no deltas, so summing only hora/ryukyoku deltas silently
omits every 1000-point riichi deposit and misranks close games. The tell was a
self-check we had built in β the recomputed segment average has to reproduce the
harness's own printed value, and it came out 0.47pt off until the deposits were
accounted for. Instrument: jax_rl/mjai_bot/group_ci.py.
The cleanest comparison we can make, and what it cost us to learn. With training paused we replayed the base g186 over the same 2,000 deals the 1.48B-step shaped checkpoint had just played, and compared them deal by deal. Pairing on identical deals correlates the two models at r = +0.307 and removes 30.7% of the variance. The answer: +0.636 pt in favour of the RL checkpoint, 95% CI +/-1.90, z = 0.66. The effect is on the positive side in the most direct measurement available, and 8,000 games cannot resolve it.
That result is worth more as a budgeting lesson than as a score. Inverting the same standard error: proving a +0.64pt effect at 2 sigma needs about 71,500 games per side, roughly 86 hours of CPU. Proving a +2pt effect needs 7,200 games per side, about 9 hours. Making the effect bigger is an order of magnitude cheaper than measuring the small one. So the next move is not more evaluation games β it is loosening the trust region (clip 0.02 -> 0.05) so the policy can move far enough to produce an effect our instruments can actually see. The intra-class correlation also replicated across the two models (-0.079 and -0.093), which is why we now trust the 0.87 correction factor.
And then we checked the premise behind that plan, and it did not hold. The argument for loosening the trust region rested on a KL number: drift from the anchor was only 0.023 nats after 1.48B steps, so "the policy has barely moved." We had never translated that into behaviour. Doing so: on identical states the RL checkpoint and the base disagree on their top-1 action 3.10% of the time β about 30 decisions per hanchan. That is not "barely moved."
The calibration that settles it: the earlier oracle-critic run, the one trained on the wrong objective that lost 3.3pt, disagrees with its base on 5.30% of decisions. So a run we describe as having drifted badly moved only 1.7x further than the current one. Distance from the human-imitation base does not predict strength; direction does. We are recording this as a correction to our own reasoning rather than quietly dropping it, and the loosening experiment is no longer justified by "the anchor is the bottleneck" β if we run it, it will be as an honest coin-flip.
The instrument has a second use that is worth more than the finding: policy_shift.py
reads out how far a config change actually moves the policy in a couple of minutes,
where the arena evaluation that answers the same question takes nine hours.
The Sichuan line cleared its first RL gate
The second line in this project trains θ‘ζε°εΊ (Sichuan bloody mahjong) from scratch β no human data, no behaviour-cloned base, no anchor. Its phase-2 gate was "beat the hand-written rule ladder's L1 within 300M steps." Measured at 105M steps, duplicate 1v3 with the challenger rotating through all four seats on identical deals:
| Challenger | vs 3x L0, 300 deals |
|---|---|
| Hand-written L1 (greedy shanten) | +3.949 +/- 0.355 |
| From-scratch RL, 105M steps | +6.872 +/- 0.500 |
| Paired difference, same deals | +2.922 +/- 0.565, z = 10.1 |
Cleared at 35% of the step budget. Note how cheap this measurement was compared to the riichi line: 300 deals sufficed because the effect is +2.9pt. That is the budgeting lesson from the other line, seen from the good side.
Two things kept us honest here. The evaluation bridges a JAX shadow state alongside the Python reference implementation the rule bots are written against β and if that shadow ever desynchronises, the network is choosing moves for the wrong position while the scores still come out looking fine. So every network decision re-checks the legal-action sets: 58,132 decisions, 0 mismatches, 0 fallbacks. And the baseline was re-measured through the identical code path on the identical deals (+3.949 against the +3.867 on record) rather than quoted from a document.
The result also comes with a defect we are not burying: the policy's entropy has collapsed to 0.048 nats, about 1.05 effective actions, with per-state max ratio reaching 3-6. Every averaged metric looks healthy; only the tail instruments show it. The phase-2 gate does not test entropy, but a near-deterministic policy is unlikely to survive the adapting opponent pool that phase 3 requires.
One claim from our own diagnosis, refuted by measurement. A survey we were working
from asserted that a single hand correlates with final placement at rho < 0.15, and
concluded that group-relative advantage at hand granularity is hopeless. Measured over
708k hanchan (27.2M hand-by-player samples): rho = 0.235, about 57% higher than
claimed. The number is wrong; the direction survives β R^2 = 0.055 means one hand
explains 5.5% of placement variance, so group-relative routes are inefficient rather
than impossible. Script: scripts/rho_kyoku.py in the repo. We publish this because a
diagnosis document is not evidence until you check it.
A second claim, and an instrument that failed first. The same document flagged its own uncertainty about which way the entropy coefficient should move. Two facts settle it as a non-lever here. Our arena play is already greedy (masked argmax), so the entropy bonus never spends evaluated points. And under the policy's own centralized Q-critic, sweeping the sampling temperature over a 200x range (tau 0.01 to 2.0) moves expected value by at most 0.0004 placement points per decision, with every confidence interval covering zero β the policy's support is value-flat, so its stochasticity is close to free. Entropy itself has sat at 0.50 +/- 0.01 nats for 1.3B steps.
Worth recording how we nearly got this wrong: the first version of the instrument used
max_a Q(s,a) as the baseline and reported an "entropy tax" of 0.836 pt per decision.
That number was estimation bias, not entropy β the max over ~9 noisy action values sits
about 1.5 sigma above the mean by construction. The tell was that it did not move at
all with temperature. Contrasting two expectations removes the max operator and the
bias with it. Instrument: jax_rl/entropy_tax.py.
A third claim: that single-player-solver features are worth +1.0 to +1.5 pt. You cannot falsify a counterfactual training gain without running it, but you can bound where it could come from. The main thing such a solver tells you is which discard keeps the hand fastest, so we measured exactly that: over ~6,000 discard decisions of self play, for every legal discard we computed the resulting shanten and the exact ukeire (tile-count acceptance), and compared the policy's choice against the optimum.
The RL policy at 1.34B steps picks a shanten-optimal discard 96.42% of the time and gives up 3.85% of the available ukeire. The BC base β trained only to imitate strong human play β scores 96.36% and 3.72%. The two are the same policy on this axis, and 1.34B steps of RL moved it by 0.06pp. Read that the right way: the deviations from maximum efficiency are not errors, they are trade-offs against safety, yaku value and dora that strong humans make at the identical rate. So the cheap story for +1.0-1.5pt ("the net cannot compute efficiency, hand it the answer") is ruled out; any real gain would have to come from sharpening the trade-off itself, which is the hard part and not what a shanten/ukeire solver gives you. We rate the claim unsupported at the stated magnitude and deprioritized it.
Caveat we will not paper over: our ukeire counts only exclude tiles in the actor's own
hand, not those visible in rivers and melds, and a real solver is wall- and turn-aware
and outputs expected value rather than acceptance. This bounds the static-efficiency
component only. Instrument: jax_rl/ukeire_headroom.py.
A fourth claim, and a correction to our own README. Both the survey and our own notes said the 8-arm league scan carried about 1.9pt of winner's-curse bias, so the best arm's -9.70 should really be read as roughly -11.6. The arithmetic behind that number is exact β the expected maximum of 8 i.i.d. standard normals is 1.4236, and a 4k-game reading has a standard error of 1.38pt, so 1.96pt. It just does not describe what we actually did. Only 2 of the 8 arms were ever evaluated externally; the other 6 were filtered out beforehand by an internal training metric. Selecting the best of 2 external readings carries 0.78pt of bias, not 1.9pt, and the internal pre-filter only inflates that toward 1.9pt to the extent that it predicts external strength β which this project's central finding is that it does not. The debiased reading is about -10.5, with -11.7 as a worst case. Either way it sits far below its own -8.87 base, so nothing about the "self-family league gains nothing externally" conclusion changes.
The forward-looking half matters more than the correction: we now publish every checkpoint's reading and pool them inverse-variance instead of reporting the best one, which is immune to this bias by construction.
Two silent defects we found by measuring instead of reading (2026-09-07)
Both were found on a day spent auditing rather than training. Neither crashed, neither tripped an assertion, and both had been corrupting results for weeks.
1. The environment's wall RNG was a global constant after the first hand
We pin Mahjax at 3fa2826. In that revision
red_mahjong/env.py::_init never seeds RoundState.rng_key β it keeps the dataclass
default PRNGKey(0) β while every subsequent hand draws its wall from
split(round_state.rng_key), and Env.step deletes the key it is handed on its
first line. The consequence: hands 2 through 9 of every hanchan use one fixed set
of 8 decks, identical across seeds, across parallel environments, across episodes
and across runs.
round_mode="single" is unaffected. Under round_mode="half" we measured that
88.9% of training steps land on constant walls. Upstream fixed this in v0.1.3
(commit fade6de, issue #71, 2026-08-08), which describes it as an implicit
full-information oracle for RL.
What makes this worth writing down is the shape of the failure. Tile conservation,
zero-sum scoring, legal-action masks and end-of-hand settlement all pass exactly as
before. Our evaluation runs through libriichi.arena and never touches Mahjax, so it
stayed clean β which is precisely the asymmetry that produces the symptom we had been
staring at for weeks: internal metrics healthy, external strength flat. We switched
round_mode from single to half on 2026-08-23 as part of fixing an unrelated
objective mismatch, and in doing so turned this latent bug on ourselves. About 2.98B
steps of RL (~131 GPU-hours) were trained on ~89% duplicated deals.
Every instrument we had was on the policy side β clip fraction, entropy, KL, per-state max ratio. Not one was on the data side. One line counting distinct decks per batch would have caught it immediately.
2. The per-seat GAE reset was off by one
Our turn-based four-player credit assignment chains GAE per seat, since only
current_player acts but all four seats can receive reward at any step. The reverse
scan zeroed the carry at is_new[t] β the first step of an episode β where it
should zero at the last. So an episode's opening decision was treated as terminal
(bootstrap forced to zero), while its true terminal decision inherited gae/next_value
from the following episode.
Replicated and fixed on the Sichuan trainer, then verified on identical trajectories: per-seat reward conservation error 1.000 -> 0, reward mass dropped at episode boundaries 22.0 -> 0 (1.03% of total), cross-episode value contamination 31 -> 0 (0.095% of decision points), with the valid-sample fraction unchanged at 87.50%. Real game score was never lost (with shaping disabled the leak is exactly 0); what leaked was the shaping term. On the riichi line this matters more than the raw percentage suggests, because placement reward is paid only on the terminal step β the contaminated transitions are exactly the ones carrying the entire objective signal.
Did fixing them explain the plateau? No.
The story fit too well to accept without a test. We re-ran from the same untouched base (g186, -4.66) with both defects fixed and every hyperparameter identical to the contaminated run, and adjudicated at the 12k protocol:
| Checkpoint | Steps | 12k avg_pt | vs base |
|---|---|---|---|
| clean b64 | 0.51B | -3.03 | +1.63 (z = 1.96) |
| clean b384 | 0.85B | -3.38 +/- 1.30 | +1.29 (z = 1.77) |
| clean6 | 1.40B | -3.95 +/- 1.33 | +0.71 |
| contaminated b1408 | 1.48B | -4.04 +/- 1.54 | +0.62 |
At matched compute, clean minus contaminated is +0.09 +/- 1.91. Both bugs were real and are fixed, but neither was the reason the line stopped improving β a negative result that removed "not enough deal diversity" from the list of explanations.
Where the remaining gap is
Before spending more compute, we asked where the missing points actually are. Parsing the 12k duplicate game logs directly (no inference; the challenger rotates through all four seats on the same deals, so seat and tile luck cancel) for clean6 versus the three baseline seats on the same games:
| Per hand | Agent | Baseline | Diff |
|---|---|---|---|
| Win rate | 20.98% | 21.72% | -0.75pp |
| Deal-in rate | 12.99% | 12.83% | +0.16pp |
| Riichi rate | 16.62% | 19.03% | -2.41pp |
| Call rate | 29.36% | 30.63% | -1.27pp |
| Mean win value | 6,479 | 6,695 | -216 |
Net per hand: -109 points, of which offense -95 and defense -13. Defense is already at parity; the gap is almost entirely offensive, and its shape is an agent that is too passive. Tile efficiency had been measured earlier and has no headroom (shanten-optimal discards 96.4%, same as the behaviour-cloned base).
The same instrument on the Mortal-stack value line shows the mirror image: its first RL peak riichis 2.03pp more than the baseline and deals in 0.86pp more; the anchored fine-tune that became the project's best pulled that back to +0.81pp and +0.15pp. Two lines, two opposite errors on the same axis β push-or-fold and riichi judgement β and strength tracks closeness to the baseline on that axis. Mechanism readouts like these now go next to every strength number we report; they explain what an RL change fixed, which an avg_pt with a +/-2.4 interval cannot.
Measuring against a weak opponent saturates, and then lies
The Sichuan section above reports its gate on a scale of "average points versus three copies of L0". That scale turns out to be saturated, which invalidates every margin quoted on it β including the follow-up reading in our repo, which had the 298M-step checkpoint scoring 0.528 worse than the 105M one and concluded that learning had stopped before 100M steps. Both of those were artifacts of the measuring stick, and correcting them is more interesting than the original claim.
L0 is a uniform-random policy, and there is a hard ceiling on what can be extracted from random opponents. We measured that ceiling at about +6.4, and inside the saturated band the ordering inverts:
| vs 3x L0 | head-to-head | |
|---|---|---|
| b192 (25M steps) | +6.775 | -1.422 +/- 0.411 |
| b2272 (298M steps) | +6.343 | 0 (anchor) |
L0 ranks b192 above b2272. Played against each other, b2272 wins by +1.59, agreeing
in both directions (z = +7.9 and -6.8). A positive control β b800 vs 3x b32 = +5.614 +/- 0.465 β rules out the alternative explanation that the head-to-head instrument
simply cannot see anything.
Rebuilt with each checkpoint played directly against the strongest one, on identical deals (the anchor scores exactly 0 by the A(pi,pi) = 0 identity):
| checkpoint | steps | vs 3x b2272 | paired step-over-step |
|---|---|---|---|
| b32 | 4.2M | -5.216 +/- 0.344 | β |
| b96 | 12.6M | -5.124 +/- 0.365 | +0.092, z=0.45 |
| b192 | 25.2M | -1.422 +/- 0.411 | +3.703, z=15.5 |
| b384 | 50.3M | -0.245 +/- 0.369 | +1.177, z=4.51 |
| b800 | 105M | -0.058 +/- 0.323 | +0.187, z=0.91 |
| b1408 | 185M | -0.085 +/- 0.333 | -0.027, z=-0.13 |
| b2272 | 298M | 0 | +0.085, z=0.50 |
| hand-written L1 | β | -2.062 +/- 0.336 | β |
Learning finishes at 50.3M steps, 17% of the budget; the remaining 83% buys +0.245 +/- 0.369 (z=1.30, not significant). The correct reading for b2272 - b800 is the head-to-head +0.035 +/- 0.270, not -0.528: "no stronger" survives, "nominally worse" does not. And the L0 scale had silently swallowed a real +1.18 gain between 25M and 50M steps.
The phase-2 gate itself still stands β b2272 beats L1 by +2.06 on the unsaturated scale, the same order as the +2.4 to +2.9 measured through L0. What does not stand is using any of those margins as a progress metric.
One more correction in the same spirit: the collapsed entropy we reported as 0.048 nats
is an average over all decision points, and 72.9% of Sichuan decision points have
exactly one legal action. Bucketed by legal-action count, the >= 5 bucket fell from
0.495 nats at 4M steps to 0.173 at 298M β a 65% collapse, where the average shows only
14%. The collapse is real (our falsification threshold was 0.5 nats in that bucket, and
0.17 misses it), but the number we had been watching understates it badly, and any
closed-loop controller regulating the average would be regulating a quantity that is
73% structurally frozen.
The generalisable lesson, and the reason this sits next to our objective-mismatch writeup: a measuring instrument can reach its limit before the thing being measured does, and it will not tell you. It just returns a flat β or inverted β curve. Before concluding "no progress", show that the instrument can still resolve a known difference (positive control) and that it has not saturated (re-anchor and re-measure).
Format
Plain pickle of a Flax parameter pytree for LeanACNet(channels, blocks)
(see jax_rl/net_lean.py in the repo). Load and run:
import pickle, jax
from net_lean import LeanACNet # from the TanyaoDojo repo
from obs_v2 import observe_v2 # or obs_lean.observe_lean
params = pickle.load(open("bc_v2_g186.pkl", "rb"))
net = LeanACNet(channels=256, blocks=10)
logits, value = net.apply(params, observe_v2(state)) # state: Mahjax red_mahjong State
Observation must match the checkpoint: bc_v2_* needs obs_v2 (36 planes),
everything else needs obs_lean (20 planes). Planes are stored as
uint8 * scale in datasets (scale 24 for v2, 4 for lean) and divided back at
train/eval time.
Evaluation harness: jax_rl/mjai_bot/run_eval.py in the repo
(--obs v2 for the v2 checkpoint).
Training data β not distributed
These models were trained on Tenhou houou-level game logs. Neither the logs nor
the derived datasets are redistributed here, per Tenhou's terms. The repo ships
the full builder (jax_rl/data_bridge/make_bc_dataset.py) so you can rebuild
equivalent datasets from logs you obtain yourself.
License
MIT for these weights and the core training code. Note that the repo's evaluation
bridge (jax_rl/mjai_bot/) links libriichi and is AGPL-3.0; see the repo's
LICENSING.md.