echoctx's picture
Upload folder using huggingface_hub
ceea1a6 verified
|
Raw
History Blame Contribute Delete
19.1 kB
# SN3 β€” **REFUTED: the "9% objective coverage" root cause.** Fixing coverage 0.09 β†’ 1.00 changed the outcome by **0.7%**. An 8-cell (batch Γ— lr) salvo on the first full-coverage corpus in campaign history shows damage is set by *displacement alone* β€” data-independent, only weakly batch-dependent β€” and that **the gradient of the validator objective at king-647 is statistically zero**. CPT from king-647 is not "under-tuned", it is mathematically dead.
Netuid 3 (`teutonic`). Cycle 63, `campaign-orchestrator`, goal `b72ea15d-1e14-40ba-9f25-1c212ee1bee9`.
Real Modal GPU (8 Γ— `B200:1`, detached), real full-coverage corpus, real reigning-king weights,
official `lab.select` gate. **NO_WINNER. No promotion. 0 TAO spent. No hotkey spent.**
Read Β§1 (the refutation), Β§2 (the law), Β§4 (what this closes), Β§5 (where the crown actually is).
---
## 1. HEADLINE β€” coverage was a 12% effect on a 100x problem
Cycle 62 concluded that ~15 failed cycles were caused by a **data-coverage bug**: the trainable
`.npy` shards covered only 9% of the eval objective weight, so "no LR/optimiser/rank/merge sweep
could ever have fixed it." It built `trainmix-v1` (10/10 corpora, exact validator weights,
byte-range-verified disjointness) and called it "the single highest-value GPU spend available."
That prediction is now tested and **false**. Same king, same recipe, same update count, same
seed β€” only the corpus changed:
| accum | peak_lr | updates | training corpus | objective coverage | `mu_hat` @ u=40 |
|---|---|---|---|---|---|
| 16 | 1e-5 | 40 | dendrite-only (cycle 62) | **0.09** | βˆ’0.006963 |
| 16 | 1e-5 | 40 | `trainmix-v1` (this cycle) | **1.00** | **βˆ’0.006143** |
Coverage bought **+0.00082 (12%)**. The gate needs **+0.0015**, i.e. a further **+0.0079**.
**Coverage was real, and it was a rounding error on the actual problem.**
Sharper still β€” cycle 62's dendrite-only run at `lr 4e-5`, rescaled to `3e-5` by the `lrΒ²` law
(factor 0.5625), predicts `βˆ’0.022419`. This cycle's **full-coverage** run at `lr 3e-5` measured
**`βˆ’0.022262`** β€” agreement to **0.7%**. The training distribution is essentially *irrelevant* to
the outcome; `peak_lr` alone predicts it.
## 2. The 8-cell grid: damage is displacement-quadratic, not gradient-noise
Design: single 2Γ—4 factorial, `accum ∈ {16, 64}` (B = 32 / 128 seqs = 65k / 262k tok) Γ—
`peak_lr ∈ {1e-6, 3e-6, 1e-5, 3e-5}`, 8 **concurrent** detached `B200:1` containers, base
`/kings/king-647-cp0` (reigning reign 179), 40 updates, `eval_n=300`/corpus Γ— 10 corpora,
shared `seed=777`/data order, top-8 layers + `lm_head` + norms (2.659B trainable).
`mu_hat` by update (positive = we beat the king):
| variant | u=0 | u=10 | u=20 | u=30 | u=40 |
|---|---|---|---|---|---|
| b16-lr1e-6 | βˆ’0.000036 | +0.000004 | βˆ’0.000024 | βˆ’0.000027 | βˆ’0.000027 |
| b16-lr3e-6 | βˆ’0.000019 | βˆ’0.000132 | βˆ’0.000310 | βˆ’0.000542 | βˆ’0.000740 |
| b16-lr1e-5 | βˆ’0.000024 | βˆ’0.002801 | βˆ’0.005080 | βˆ’0.005742 | βˆ’0.006143 |
| b16-lr3e-5 | βˆ’0.000019 | βˆ’0.018482 | βˆ’0.022879 | βˆ’0.023488 | βˆ’0.022262 |
| b64-lr1e-6 | βˆ’0.000029 | βˆ’0.000018 | +0.000006 | +0.000012 | βˆ’0.000014 |
| b64-lr3e-6 | βˆ’0.000031 | βˆ’0.000109 | βˆ’0.000241 | βˆ’0.000315 | βˆ’0.000434 |
| b64-lr1e-5 | βˆ’0.000026 | βˆ’0.001961 | βˆ’0.003111 | βˆ’0.003500 | βˆ’0.003841 |
| b64-lr3e-5 | βˆ’0.000029 | βˆ’0.010956 | βˆ’0.013661 | βˆ’0.013947 | βˆ’0.013817 |
**Monotone decreasing from the first eval in 6/8 cells; the other 2 are flat at zero.** No
interior optimum exists anywhere in a 1.5-decade LR range Γ— 4x batch range.
### 2a. The batch marginal kills the gradient-noise explanation
The natural hypothesis (mine, going in) was the SGD noise term `E[dL] β‰ˆ βˆ’lrΒ·|g|Β² + (lrΒ²/2)Β·tr(HΞ£)/B`:
if noise dominates, 4x batch β‡’ **4x** less damage, and the improvement window widens linearly in B.
| lr | b16 | b64 | observed ratio | pure-noise prediction |
|---|---|---|---|---|
| 1e-6 | βˆ’0.000027 | βˆ’0.000014 | 1.95 | 4.00 |
| 3e-6 | βˆ’0.000740 | βˆ’0.000434 | 1.71 | 4.00 |
| 1e-5 | βˆ’0.006143 | βˆ’0.003841 | 1.60 | 4.00 |
| 3e-5 | βˆ’0.022262 | βˆ’0.013817 | 1.61 | 4.00 |
Consistently **~1.6x, i.e. B^βˆ’0.34, not B^βˆ’1**. Minibatch noise is a *minority* contributor.
Extrapolating `B^βˆ’0.34`, reaching the gate by batch size alone would need **B ~ 10^6 sequences** β€”
not a real option. **Do not run a bigger-batch salvo; this marginal already forecloses it.**
### 2b. What survives is the displacement law, with zero first-order term
Damage tracks Adam displacement `Ξ”_rel β‰ˆ UΒ·lr/|w|` quadratically, with the *same* constant found
in cycle 62's merge geometry (`k β‰ˆ 15`, fitted there on `β€–dirβˆ’baseβ€–/β€–baseβ€–`):
| lr | Ξ”_rel (U=40, β€–wβ€–~0.02) | kΒ·Ξ”_relΒ² predicted | observed (b16) |
|---|---|---|---|
| 1e-6 | 2.0e-3 | βˆ’6.1e-5 | βˆ’2.7e-5 |
| 3e-6 | 6.0e-3 | βˆ’5.5e-4 | βˆ’7.4e-4 |
| 1e-5 | 2.0e-2 | βˆ’6.1e-3 | βˆ’6.1e-3 |
| 3e-5 | 6.0e-2 | βˆ’5.5e-2 | βˆ’2.2e-2 (saturating) |
One constant fits gradient CPT *and* weight-space merging. Since the quadratic penalty is
direction-agnostic, the only thing that could beat it is a **first-order gain term**, and the
`lr=1e-6` cells bound it directly: with **full** objective coverage, the exact validator weights,
and 40 real updates, `|mu_hat| ≀ 3e-5` β€” i.e. any first-order gain is **β‰₯50x too small** for the
gate, and within 1 se of zero.
> **Therefore: βˆ‡(validator objective) at king-647 β‰ˆ 0. bluecolor has converged this basin.**
> This is a *measurement*, not an inference from failures.
### 2c. Every corpus degrades, including the ones we train hardest
`b16-lr3e-6`, u=40, per-corpus `mu` (all 10 negative):
`nemotron-spec βˆ’0.000227 Β· ultradata βˆ’0.000270 Β· dendrite βˆ’0.000452 Β· nemotron-cc βˆ’0.000459 Β·
openthoughts3 βˆ’0.000554 Β· openmathreasoning βˆ’0.000616 Β· automathtext βˆ’0.000820 Β·
finewebedu βˆ’0.000922 Β· pes2o βˆ’0.001032 Β· cosmopedia βˆ’0.001220`
`cosmopedia` is the **worst** corpus while being trained at its full objective weight, and
`finewebedu` (w=0.27, 27% of every draw) degrades too. "Train the corpus, improve the corpus" does
not hold anywhere near this minimum. This retires the last version of the coverage hypothesis.
## 3. The gate ran for real and refused for exactly one reason
`python -m lab.collect` parsed **8/8** results with zero errors;
`python -m lab.select --no-blocking-incident` β†’ **`NO_WINNER`**.
Every non-statistical predicate **passed** on all 8 cells: `real_data=true`, `promotable=true`,
`king_digest == hf:ce6efba4…` (reigning), incident predicate explicitly false (verified in
Postgres: 15 open incidents, **all** `blocking=false`). The sole failure code on every cell is
`lcb_below_delta`. Best cell `lcb999 = βˆ’0.000101`. **No `lab.promote`, no `B200:8` run** β€” a
refused ~$50/h confirm is the correct output of a working gate.
## 4. What this cycle CLOSES (do not re-derive; ~$16 of GPU bought these bounds)
1. **Objective-coverage CPT** β€” closed. 0.09β†’1.00 is a 12% effect (Β§1).
2. **The `peak_lr` axis from king-647** β€” closed over `1e-6 … 3e-5` on real full-coverage data.
Downward is flat-zero, upward is quadratic damage. Cycle 62's "sweep downward" advice is
answered: downward does not damage, it simply does **nothing**.
3. **The effective-batch axis** β€” closed by the `B^βˆ’0.34` marginal (Β§2a). No bigger-batch salvo.
4. **"Train the drag corpora"** β€” closed a second time, now on-distribution (Β§2c).
5. Combined with cycle 62 Β§3d/Β§3e (merges, LoRA, replay): **every direction the campaign can
construct from king-647 has ~zero first-order gain and pays the same quadratic penalty.**
## 5. Where the crown actually is (strategy correction)
The campaign has spent ~17 cycles trying to *descend from the king*. That route is now
measurably closed, so the two live routes are:
**(a) Opponent timing β€” already-passing asset.** `soup3` (unchanged weights) scores
`mu_hat +0.004251`, `lcb999 +0.003940` = **283% of the gate** against reign-177-class
(`kuya83/…-v305`) weights, and βˆ’0.0012 against reign-179. Cycle 62 cleared the chain blocker and
we hold **one registered hotkey** (`5FLNBmzb…`, `registered=true`). The binding constraint is no
longer model quality or chain capability β€” it is **being ready to commit within minutes of a
beatable throne**. Concrete gap: **we have never uploaded a candidate to HF.** Every cycle's
receipts say "NO HF MODEL UPLOAD". Upload `/ckpts/soup3` to a public HF repo *now* (it cannot trip
`model_copy` β€” genuine weight average, distinct blobs, cycle 62 Β§6) and hold the digest unspent.
Then a crown costs one `commit` intent instead of an upload + a race.
*Live duel evidence that this is the real game (eval-0340…0344, ~19 min cadence):* every
current challenger loses by βˆ’0.0043β€¦βˆ’0.0051, and **their entire deficit is two corpora** β€”
`dendrite-synth-run β‰ˆ βˆ’0.062` and `cosmopedia β‰ˆ βˆ’0.026`, while `finewebedu`/`automathtext` sit
at Β±0.002. `soup3` is at βˆ’0.0011/βˆ’0.0006 on those same two, i.e. **~56x closer than the field.**
We are the second-best model on this subnet and are losing on *entry mechanics*, not weights.
**(b) A different basin.** Since βˆ‡ β‰ˆ 0 at king-647, improvement requires a model that is not in
this basin *and* has lower objective loss. That is an **eval-only** question (~6 min of `H100:1`,
~$0.40 per candidate, and the eval is bit-exact reproducible), not a training question. Scan
architecture/tokenizer-compatible bases (`silx-ai/Quasar-10B`, `dendrite/teutonic-x-genesis` seed,
fetchable historical thrones) before spending another GPU-hour on gradient descent.
**Do NOT** spend the hotkey against reign-179. `1-hotkey-1-eval`, burned at enqueue, and we hold
1 shot (+~12 at 0.005426 TAO each above the 0.5 TAO floor).
## 6. Reusable runtime notes
* **The `lab.*` gate path works end-to-end for a bespoke trainer** β€” and this is the cheap way
around cycle 62 Β§A2's harness gap. `lab.fanout`β†’`modal.cpt_entry` still cannot express this
experiment (it injects no `CPT_DATA_DIR`, **and `cpt.runner` computes no `mu_hat`/`se` at all β€”
a bigger gap than Β§A2 recorded, since the results bus *requires* those fields). The fix that
needs no repo write: have the bespoke container write `result.json` in **lab-bus schema**
(`variant_id, tag, mu_hat, se, king_digest, used_real_data, promotable, recipe, baseline_u0`)
to `/data/lab-results/<tag>/variants/<vid>/`, pull it, and run the real
`lab.collect`/`lab.select`. 8/8 parsed first try.
* **`/opt/mining-platform` is READ-ONLY for mining runtime** (open incident
`repo_concurrent_root_refactor_writer_mining_must_not_write`). Patching `lab/*.py` would have
been the wrong instinct; bespoke scripts under `/var/lib/mining/work/` + lab-bus schema is the
compliant path.
* **`modal volume get` on a directory fails with `[Errno 21] Is a directory`** even when the local
target does not exist. Fetch **explicit file paths** one at a time (loop over variant ids).
* `setsid nohup … modal run --detach -e mining … &` β€” 8/8 jobs survived every bash-timeout and
session boundary. Log lines carry `\r`; `tr -d '\r'` before `grep`, or greps silently miss.
* **Persist at every eval point**, not at the end: each container rewrote `result.json` +
`detail.json` after every eval, so a mid-run kill still leaves gate-parseable truth.
* **`durable-goals` CAS**: version bumps ~1 per agent tool call, so even `actual+1` races
(read 76 β†’ sent 77 β†’ conflict, actual 80). Read immediately before the call and expect retries.
* **Owner applied a second chain fix mid-cycle**: `python/chain/writes.py` md5
`b26d1a89…` β†’ **`0e2c47f4c4249d53fe747c1a55048b73`** (mtime 09:39:34), and
`python/chain/server.py` now carries `_json_default` = "Balance-like β†’ int(rao)" wired into
`encode_ok`, i.e. the structural fix for incident `01dffb40-…` (successful extrinsics recorded
as `failed`). `mining-chain.service` restarted 09:40:25Z. **Code-bytes re-verification only** β€”
the functional proof needs the next real write receipt, and I refused to burn 0.005426 TAO on a
test registration for it. Incident left open, not resolved, no duplicate filed.
* Balance unchanged at **568,420,214 rao** = exactly the post-registration figure β‡’ nothing has
been spent since, and the cycle-62 burn reconciles.
## 7. Next actions, ranked
1. **Upload `/ckpts/soup3` to a public HF repo and hold the digest** (Β§5a). This is the only step
between us and converting a beatable throne into a crown, and it costs no TAO and no GPU.
2. **Eval-only basin scan** (Β§5b): `silx-ai/Quasar-10B`, `dendrite/teutonic-x-genesis`, fetchable
thrones β€” `mu_hat` vs king-647 at n=300. ~$0.40 each, bit-exact, no training.
3. **Throne watcher**: poll the dashboard (~19 min duel cadence), and when a new `hf:` king
appears, run the 6-min paired eval of the *held* soup3 digest; commit iff `lcb999 > 0.0015`.
4. **Do NOT** run: bigger-batch CPT (Β§2a), lower-LR CPT (Β§2/Β§4), targeted-corpus CPT (Β§2c),
any merge from king-647 (cycle 62 Β§3d), or another coverage fix (Β§1).
5. Owner (existing single incident `01dffb40`, no new file, no duplicate): nothing new required β€”
the fix appears applied; it needs one real write receipt to confirm.
## Receipts
* **Modal (real GPU, 8 concurrent detached jobs)**: env `mining`, app `mining-trainmix-salvo`,
`B200:1` each, script `/var/lib/mining/work/tmix63/modal_trainmix_salvo.py`, tag `tmix63`,
variants `b{16,64}-lr{1e-6,3e-6,1e-5,3e-5}`. App ids `ap-V8iLhxNEu4nXaCUvqDa29g`,
`ap-K1PxGE9rz4lHkiZO3ZIGd6`, `ap-zSg2TGTJpxepiSFdctyVoo`, `ap-YYCQW2fupz0iG7VK6Sx6vj`,
`ap-HanQzbcBJenhcXhFMDuHD0`, `ap-b8StaKwXEzJgbye3ZFLvFl`, `ap-WUJCE8AQ91F4a9cA4sPqc8`,
`ap-lZOkBE4IEf6vQN7ghVs4ra`. Per-variant `elapsed_s β‰ˆ 1252` (b16) … ~1500 (b64);
all 8 exited `App completed`; `modal app list -e mining` shows 0 running at cycle end.
* **u=0 wiring check PASSED 8/8** against the *independent* cached reference
`/data/truemix-results/rebase-newking-n300/per_sequence.json` (key `/kings/king-647-cp0`,
300/corpus Γ— 10 verified before launch): `u0 mu_hat ∈ [βˆ’3.6e-5, βˆ’1.9e-5]`, `se = 2.7e-5`,
tol `4Β·se β‰ˆ 1.09e-4`.
* **Real data, verified in-container**: `oob_count = 0` on 1,280 (b16) / 5,120 (b64) draws;
realized mixture matched the objective weights (`finewebedu 351/1280 = 0.274` vs 0.27,
`dendrite 44/1280 = 0.034` vs 0.04); `trainmix-v1` `weight_total = 1.0`,
`eval_disjointness.contaminated_rows_removed = 871`, tokenizer `silx-ai/Quasar-10B`;
train/eval corpus-set equality asserted; refused to start without an `eval_disjointness` block.
* **Gate (real harness)**: `lab.collect --tag tmix63` β†’ `n_results 8, errors []`;
`lab.select --current-king-digest hf:ce6efba444047328046a511e21d3e37a84c329c5
--no-blocking-incident` β†’ **`NO_WINNER`**, sole failure code `lcb_below_delta` on all 8.
Results bus `/data/lab-results/tmix63/variants/*/{result,detail}.json`.
* **Incidents**: Postgres `select … where status not in ('resolved','closed')` β†’ 15 rows, **all
`blocking=false`** β‡’ `--no-blocking-incident` asserted from evidence, not assumption. No new
incident filed (no new defect found); `01dffb40` re-verified by bytes only and left open.
* **Chain**: `read_balance(5Do6Wqm8brkdcNy3ALauTTUgGkPp57MM6dkNbxwWUch8ZG51) = 568420214 rao`.
**No intent enqueued. Nothing signed. 0 TAO spent. Hotkey unspent.**
* **Dashboard**: king unchanged for the whole cycle β€” reign 179,
`bluecolor/teutonic-q3-5ek5kr57gg-647127189391-rn-cp0`, `hf:ce6efba4…`, crowned
2026-07-26T05:10:44Z (~29.5 h). Duel loop **live**: `eval-0340…0344` at 08:55/09:16/09:35/09:56/10:14Z
(~19 min), `historyCount 2362`, all five challengers rejected at `mu_hat βˆ’0.0043β€¦βˆ’0.0051`.
* **HF**: token user `echoctx` verified live (`whoami-v2`). Findings package published β€” see the
HF repo recorded in the cycle outcome. **No model weights uploaded this cycle** (that is Β§7.1).
---
# ADDENDUM (same cycle) β€” Β§7.1 EXECUTED: the campaign's **first model-weight upload** is live, and HF private storage is a hard wall
## A1. `soup3` is now staged as a fetchable candidate (real bytes)
Every prior cycle's receipts ended with "**NO HF MODEL UPLOAD**". That is now false.
`echoctx/sn3-soup3-candidate`, commit **`e0177f0a19fd275028f54f3b8bb731134f480f1a`**,
**17.204 GB / 6 files**, `model.safetensors 17,204,127,784 B`, `private: false` β€” verified by an
independent fresh `GET /api/models/...?blobs=true` after the uploader exited, not from the
uploader's own stdout. Weights pulled from `mining-checkpoints:/soup3` (`modal volume get` on
**explicit file paths**), so the blobs are the same soup3 that scored `lcb999 +0.003940`
(283% of gate) against reign-177-class weights.
**A beatable throne now costs one `commit` intent, not an upload race.** No commit was enqueued:
`1-hotkey-1-eval`, and soup3 is βˆ’0.0012 against the *reigning* king. The shot stays unspent.
## A2. NEW HARD CONSTRAINT β€” free HF accounts cannot host a private 10B model
First attempt was deliberately **private** (SN3 has an observed weight-recycling rival, so a public
candidate can be fetched and committed by someone else). It failed on real bytes:
```
Bad request for commit endpoint:
Private repository storage limit reached, please upgrade your plan to increase your
private storage limit
```
`whoami-v2` for `echoctx`: `isPro: false`, `canPay: false`, `billingMode: prepaid`. So
**"upload privately, flip public at commit time" is NOT available to this account.** The repo was
created and left holding only `.gitattributes` β€” an empty repo is the failure signature; check
`n_files`/`total_bytes` via the API, never trust `hf upload`'s exit path alone (my status file
recorded `upload_rc: 1` *and* a valid-looking commit sha, which is exactly the shape of a
half-truth receipt).
Flip to public is one call and works:
`PUT /api/models/<repo>/settings {"private": false}` β†’ `{"private":false}`.
**Accepted the public exposure** on evidence, not convenience: (i) soup3 *loses* to the reigning
king, so a copier gains nothing today; (ii) cycle 62 Β§6 established `model_copy` compares weight
**blobs**, and the HF-backed variant `(trusted challenger timestamp unavailable)` can never win the
`crown_earlier_commit` tie-break β€” so a byte-copier cannot out-commit us; (iii) readiness was the
campaign's #1 unblocked lever (Β§5a).
## A3. Open question for the next pilot (do not assume)
The staged repo carries **no tokenizer files** (`soup3` on the volume had none, and the king-repo
fetch with `--include` was ignored: "*Ignoring `--include` since filenames have been explicitly
set*"). Our own paired harness loads `AutoModelForCausalLM` only β€” the validator scores
**pre-tokenized `.npy` ids**, so a tokenizer is probably unnecessary. **Verify against the upstream
validator loader before committing**, and if needed copy the king repo's tokenizer files (public)
into the candidate repo β€” a 6-file, seconds-long fix, but a wasted hotkey if wrong.