Your count is right on my side too, your conversion works out of sample, one number in it is my error, and your question has a measured answer. Then I have to correct letter 59, and give you what we did while you were away.
Your numbers re-derive
I re-solved the three rows in u = log(r3/r4) over [-200, 200], as you did.
fold-1e-5 5 roots X3 = 0.000 9.035 9.292 43.585 49.329 (r3 min 6.13e-22, r4 min 2.47e-21)
fold+1e-5 3 roots X3 = 0.000 43.585 49.328
Same numbers to every digit you printed. The slopes of the node, d(d3)/d(delta), from my own closed form by centred differences:
fold-1e-5 fold-1e-6 fold-1e-7 fold-1e-8 fold-1e-9 fold-3e-10 fold-1e-10
slope 16.07 53.88 173.48 551.71 1747.76 3191.82 5527.32
you 16 54 552
Your three conversions reproduce: +3.0e-7 gives fold - 5.57e-9 at eps 1e-7, -5.8e-7 gives fold + 1.08e-8 at 1e-8, -2.3e-6 gives fold + 4.27e-8 at 1e-10.
One number is wrong, and the error is mine. You wrote that my -8e-6 at fold-1e-8 reads as delta - 2.0e-8. With your own slope, -8.034e-6 / 551 = -1.46e-8. The 2e-8 comes from my own notebook line ("a node at delta - 2e-8") which I wrote without dividing by the right slope. The other two entries are right: -4.95e-8 and -4.20e-8.
Your question, measured
Full network, wall 77777 k=3 (referents 3/4), eps 1e-10, six phases per point (warm-up 20 000 steps at fold-1e-6, then 20 000 + k with k = 89, 144, 233, 377, 610, 987, then delta), 40 000 steps traced, statistics on the last 30 000. None of the 42 runs escaped. s = (mean d3 - closed-form node) / local slope.
offset from the exact fold slope s (mean - node)/slope (mean - median)/slope local exponent p
fold-1e-5 16.07 -4.977e-8 +- 4.5e-10 -4.978e-8
fold-1e-6 53.88 -4.235e-8 +- 2.7e-10 -4.236e-8 0.07
fold-1e-7 173.48 -2.779e-8 +- 1.3e-10 -2.780e-8 0.18
fold-1e-8 551.71 -1.458e-8 +- 6.7e-11 -1.376e-8 0.28
fold-1e-9 1747.76 -6.376e-9 +- 7.7e-11 -4.565e-9 0.36
fold-3e-10 3191.82 -3.856e-9 +- 3.0e-11 -2.503e-9 0.42
fold-1e-10 5527.32 -2.350e-9 +- 1.8e-11 -1.445e-9 0.45
p is d ln|s| / d ln(delta_c - delta) between successive rows. The shift keeps falling, so "still shrinking as you approach" holds. It does not stop at 6e-10 to 1e-9: at fold-1e-9 it is 6.4e-9 (4.6e-9 with mean minus median), six to ten times your figure. Precommitted windows were |s(-1e-7)| in [1.5e-8, 4e-8], |s(-1e-8)| in [1e-8, 2e-8], |s(-1e-9)| in [4e-9, 1.2e-8]. All three landed, and so did the two I added afterwards at -3e-10 and -1e-10.
The local exponent climbs 0.07, 0.18, 0.28, 0.36, 0.42, 0.45. It is heading to 1/2, which is the exponent of the node-to-twin distance. So the converted shift goes to zero at the fold, and the reading "threshold = fold - s" cannot give a threshold at eps 1e-10 in that limit. At fold-1e-9 it lands near the region where I see escapes (fold+5e-9 to +7e-9). At fold-1e-10 it says 2.4e-9. I take the landing at -1e-9 as a coincidence of the evaluation point, not a prediction.
To see whether p really reaches 1/2 I went closer with the two-line reduction of the network (sender rows 3 and 4, receiver logits l3 and l4, the 25-referent tail frozen; 16 phases, 2 000 000 recorded steps each, none escaped). First I checked it against the full network on the seven rows above: s agrees to 0.1 to 0.6 % (for example -6.410e-9 against -6.376e-9 at fold-1e-9). The node is solved in 40-digit arithmetic, because my root grid stops resolving the node-twin gap below about 1e-11.
offset slope s (two-line reduction) local p d3 bias = s x slope node-twin gap in d3 |bias| / gap
fold-1e-8 551.7 -1.459e-8 0.282 -8.05e-6 2.21e-5 0.36
fold-1e-9 1747.8 -6.410e-9 0.357 -1.12e-5 7.00e-6 1.60
fold-1e-10 5530.1 -2.355e-9 0.451 -1.30e-5 2.21e-6 5.89
fold-3e-11 10097.7 -1.334e-9 0.472 -1.35e-5 1.21e-6 11.1
fold-1e-11 17490.7 -7.844e-10 0.484 -1.37e-5 7.00e-7 19.6
fold-3e-12 31934.8 -4.343e-10 0.491 -1.39e-5 3.83e-7 36.2
fold-1e-12 55313.7 -2.523e-10 0.494 -1.40e-5 2.21e-7 63.1
Yes, p reaches 1/2: 0.451, 0.472, 0.484, 0.491, 0.494, and s tends to 2.5e-4 sqrt(delta_c - delta) (2.36e-4, 2.48e-4, 2.52e-4 at fold-1e-10, -1e-11, -1e-12). The reason is arithmetic. The d3 bias saturates near 1.4e-5, which is the size of the bursts themselves (sd of d3 is 1.27e-5 at eps 1e-10), while the slope you gave diverges like 1/sqrt(delta_c - delta). The node-twin gap in d3 shrinks like sqrt and falls below the bias at about fold-3e-9. So the reading "the mean is the node of a shifted delta" is a linear response that holds while the bias is small against the gap: at fold-1e-6, where your conversions at eps 2e-8, 5e-8 and 1e-7 sit, the ratio is 0.01. At eps 1e-10 near the fold the bursts exceed the gap by a factor 6 to 63, and the conversion into delta units no longer has that meaning.
Your conversion out of sample
In the deterministic regime it works. I measured delta_c' by the ghost law, t = kappa / sqrt(delta - delta_c') with t0 near 0 and R2 = 1.0000 on every eps, then compared with your conversion of the mean-minus-median skew at fold-1e-6 (three phases for the two new eps):
eps your conversion measured delta_c' - fold difference
1e-7 -5.57e-9 -5.5e-9 (repeated to 1 step on a second warm state) 1 %
5e-8 -3.99e-9 -3.66e-9 +9 %
3e-8 +1.2e-10 +3.2e-10 both ~0
2e-8 +4.04e-9 +3.85e-9 +5 %
1e-8 +1.08e-8 +9.1e-9 (3-point inversion; direct bracket (6e-9, 1e-8]) +18 % / 8 % past the bracket
The 2e-8 and 5e-8 rows were predicted before I ran them (25 % tolerance). The zero of the shift sits at eps close to 3.1e-8, where the skew also vanishes. That coincidence is an identity, since both are the same displacement of the mean.
What I have to correct in letter 59
The "(6e-10, 1e-9)" I gave you for the default eps was the transient of a fresh Adam. Started from the tie with empty moments, the state is kicked over the twin in 27 steps (r4 excursion 144 times the node-twin gap), and that is what the no-warm-up bisection measured. On a settled state (40 000 steps at fold-1e-6, then delta) the state holds at fold+6e-9 and breaks at fold+7e-9 on a 60 000-step budget, on the wall and on collision 6/14 in both directions. That was already not a threshold, and the stronger statement is below.
At the default eps there is no threshold. Escape times at a fixed delta are spread over more than a decade (1 110 to 29 287 steps at fold+7.0e-9, ten phases), the outcome is not monotone in delta at 1e-10 resolution (breaks at fold+6.6e-9, holds at 6.7e-9 and 6.8e-9, breaks at 6.9e-9), and the same run breaks at step 465 or 1 110 depending on rounding. A two-line reduction of the network (sender rows 3 and 4, receiver logits l3 and l4, the 25-referent tail frozen) reproduces the full-network counts at delta >= 6.0e-9 (54 escapes observed, 48.8 expected from its rates, ratio 1.11; that pooled count is the second review's, and I re-derived its five first-crossing times, 490, 574, 593, 614 and 665 steps, with the two-line code myself) and gives the rate over five decades (the second review's table):
delta - fold (1e-9) 5.0 5.25 5.5 5.75 6.0 6.1 6.3 6.5 6.8 7.0 7.5 8.0 10
escape rate / step 1.7e-9 9.3e-9 4.0e-8 2.0e-7 5.9e-7 1.2e-6 2.8e-6 6.6e-6 2.0e-5 3.7e-5 1.1e-4 1.7e-4 4.5e-4
I re-sampled four rows myself with another seed and got 4.4e-5 (7.0), 6.8e-6 (6.5), 8.5e-7 (6.0) and 3.7e-8 (5.5, nine escapes in 2.45e8 steps). The waiting times are close to exponential up to about 6.8e-9 (sd/mean 0.75 to 0.97, bootstrap Lilliefors p mostly above 0.05), not above (sd/mean 0.45 to 0.74 from 7.0e-9 upward, p <= 0.001 nearly everywhere).
So "the threshold" is a function of the budget. Taking the delta where the rate equals 1/T:
budget T (steps) 1e4 6e4 1e5 3e5 1e6 1e7 1e8 1e9
fold + (1e-9) 7.45 6.75 6.61 6.34 6.07 5.64 5.26 5.00
Every hold-or-break bisection in the notebook at eps 1e-10 is one point of this curve.
What we did while you were away
Two reviews written in your style ran as internal checks. Neither is you. Every number they gave, I re-derived before using it, and I retracted some of my own statements on the way.
- The critical mode. I solved the Hessian of the full objective (2 x 729 parameters) at the node: the soft mode (eigenvalue -7.29e-8) is 99.1 % on the sender's row 3 (95.5 % the message logit, 3.7 % the 26 competitors), the stiff mode (-2.27e-4) is 99.1 % on the receiver. My one-variable Kapitza estimate had used u, the receiver variable, as the slow one. So my "97 % cancellation with Adam's weighting" was an artefact of that choice, and I retract it.
- Rate of passage. Away from the default eps, t = kappa / sqrt(delta - delta_c') is exact (R2 = 1.0000, 3 to 9 points per eps). At eps >> sqrt(v) Adam reduces to (lr/eps) times the gradient, and I recomputed the slope of kappa/eps on all 1458 parameters from the fold's own ingredients: alpha = n . d(grad J)/d delta = 1.538e-3 and D3J[n,n,n] = 1.903e-6 give kappa/eps = pi / (lr sqrt(alpha beta)) = 1.6422e6, with beta = D3J/2. Measured at eps 3e-6: 13 846, 9 343 and 5 421 steps at fold + 1e-7, 2e-7, 5e-7, i.e. 1.617e6. My earlier linear fit kappa = 0.0711 + 1.389e6 eps had the wrong asymptote (12 % low at eps 3e-6).
- The sign of the shift, on the full network. Putting eps = 1e-7 on the sender's row 3 only moves the threshold from the 60 000-step cliff near +6.5e-9 (all eps 1e-10) to -1.36e-8; with the receiver alone it stays a positive cliff between +2e-9 and +4e-9. The message logit alone gives a threshold in (+9e-9, +1.2e-8] and the 26 competitors alone (+6e-9, +9e-9], both positive: only the whole row flips the sign. With all eps at 1e-7: -5.43e-9 on the wall, -5.70e-9 on collision 6/14 (seed 12345), -5.67e-9 on collision 5/20 (seed 77777, k=1), and the row alone -1.36, -1.31 and -1.45e-8. Break times at -3e-9 / -1e-9 / +1e-9 are 4107 / 3065 / 2559, 4099 / 3071 / 2546 and 4100 / 3076 / 2556. The eps-dependent shift is a property of the object at fixed (N, beta, lr), not of the code around it.
- Learning rate. At lr 0.035 the thresholds are +3.25e-9 at eps 1e-8 and -2.10e-9 at eps 1e-7 (predicted +3.3e-9 and -2.1e-9 before the run; at lr 0.05 they are +9.1e-9 and -5.5e-9).
- Two regimes. The dispersion (sd/mean) of the break time over ten well-separated phases at fold+1.5e-8:
eps 1e-10 1e-9 3e-9 1e-8 3e-8 1e-7
lr 0.05 0.773 0.712 0.452 0.087 0.011 0.004
lr 0.035 - 0.381 0.040 0.009 - -
lr 0.02 0.112 0.041 0.009 0.015 - -
At lr 0.05 the passage is deterministic (ghost law) from eps 1e-8 and stochastic at 3e-9 and below. The boundary moves by a factor of 5 or more when lr goes from 0.05 to 0.02, much faster than sqrt(v) proportional to lr predicts, and I withdraw that reading. The project's lr 0.05 sits in the stochastic regime at the default eps.
- Retractions of mine from this week: the power law lambda = c (x - x*)^gamma with x* = fold + 5.7e-9 and gamma = 3.5 (its parameters move from gamma 2.3 to 12.2 with the fitting window, and the reduction has 3.7e-8 at 5.5e-9 where the law says zero); a background floor of 1.6e-6 per step (no floor in the reduction); my row at 1e-8 in the rate table (phases k <= 55 all break in 409 to 490 steps, one deterministic passage; k >= 89 give 423 to 3 309 steps, so the rate was inflated about 5 times); the "zero escapes in 20 runs at 5e-9" as a test (the reduction expected 0.003 events); my 2.0e-8 above.
What I have not tested
The full network below delta = fold + 6.0e-9. The reduction is checked against it only above; below, its fidelity is inferred. I am not running the day of compute that a direct check needs.
The functional form of the rate. A power law (gamma 12 on [5, 7]e-9), a Gaussian tail and a Kramers form with exponent 3/2 (x_c = 7.4e-9) all fit that window without rejection, and I have no mechanism for any of them. The second review reports two positive Lyapunov exponents (8.3e-3 and 2.6e-3 per step); I have not recomputed them.
Your conversion on r4. I did d3 only, as you did. The validity limit of the conversion, |bias| / gap of order 1, is read off the table above and not tested finely.
Dependence on lr beyond 0.035 for the eps-dependent shift, and a fourth code.
One question back
The d3 bias saturates at 1.4e-5 while the bursts have sd(d3) = 1.27e-5, whatever delta near the fold. Would you say the bursts carry the mean a fixed distance toward the twin, so that in the stochastic regime the right unit for the shift is a d3 offset (or a fraction of the burst amplitude) and not a delta offset: yes or no?
Scripts: verifier_tour60_conversion_delta.py (roots, slopes, your conversions), verifier_tour60_biais_converti.py and lancer_tour60_traces.sh (the full-network tables above), verifier_tour60_exposant_reduction.py (the reduction rows down to fold-1e-12), verifier_tour59_masque_eps.py and adam_eps_masque.py (per-group eps), verifier_tour59_hessien_reseau_complet.py, verifier_tour59_kappa_fermee_reseau_complet.py, verifier_tour59_dispersion_phase.py, hasard_reduction_tour59.py with reduction_numba_tour59.py and reduction_deux_lignes_tour59.py, verifier_tour59_fenetre_puissance_table.py, verifier_tour59_exponentialite_attentes.py.
Notebook: section "VRAIE CRITIQUE DE DIPANKARSARKAR, 29/09/2026 (tour 60, PAS simulée)", and the sub-section "Après la lettre (28/09/2026)" of the tour-59 section.