File size: 6,013 Bytes
dfb775d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 | # dcoach β prove a CPU-trained model recalls its training
`dcoach` is the decentralized-aware extension of the [Coach](coach.md). It closes
mindXtrain's founding loop: **author a dataset β imprint a persona on a tiny model
(CPU) β prove the model recalls the training β let governance rule on it β feed the
verdict back into autotune.** It is also the on-ramp to the 2026 decentralized-training
landscape (see [the deep dive](decentralized-training-deep-dive-2026.md)).
Open it at **`/coach/dcoach`** (linked from the Coach header).
## The proof loop
```
persona + skills ββΊ script.jsonl ββΊ imprint-train (trl_local, CPU)
β
βββββββββββββββββββββββββββββββ
βΌ
probe recall βββΊ classroom (before vs after) βββΊ boardroom (rule) βββΊ feedback
(base vs adapter) recall β? persona kept? approve / reject tune next run
```
1. **Author** β a persona (e.g. `codephreak`) plus optional skills (software engineer,
platform architect, bash, solidity) is composed into chat rows
(`data/scripts.py::build_script_rows`). Each row carries the persona **system prompt**
+ a userβassistant turn.
2. **Imprint-train** β a tiny actor (default `HuggingFaceTB/SmolLM2-135M`) is LoRA-trained
on the script on the CPU lane (`train/backend_trl_cpu.py::run_trl_local`). The autotune
plan is frozen AOT β no JIT autotune in the loop.
3. **Probe recall** β `eval/imprint.py::probe_recall` generates the actor's answer to each
inquiry **before** (base model) and **after** (base + adapter). The probe prepends the
*same persona system prompt the adapter trained under*, so the comparison measures what
the imprint actually learned rather than penalising a missing conditioning turn.
4. **Classroom** β `governance/classroom.py::evaluate_classroom` scores before vs after
against the persona baseline (clean-room [llama-style evaluators](#clean-room-eval-tools)):
recall up? persona maintained? `passed = persona_maintained and pairwise β₯ 0.5`.
5. **Boardroom** β the classroom graduation becomes a motion; a board (any-N, preset or
model-backed) rules **approve / reject**. A disputed board is settled by a prime-sized
**dojo**.
6. **Feedback** β `autotune/feedback.py` records `(run_id, params, classroom_score,
outcome)` to an append-only ledger and `suggest_next_params` nudges the next run: a weak
or rejected imprint trains harder (more epochs, `grad_accum=1`); a clean pass holds.
`suggest_from_history` feeds the nudge back into `derive_training_params`.
The whole chain is `governance/proof_loop.py::run_proof_loop`, streamed phase-by-phase to
the UI via **`POST /coach/api/dcoach/run`** (SSE). It is heavy (real CPU training +
generation) β expect a few minutes per run.
## Clean-room eval tools
`eval/llama_evals.py` reimplements the *behaviour* of LlamaIndex's evaluators (MIT) from
their public contract β never copied. Each returns an `EvalScore{scoreβ[0,1], passing,
reasoning, method}`:
| Evaluator | What it measures | Backed by |
|-----------|------------------|-----------|
| `SemanticSimilarityEvaluator` | embedding/lexical closeness of two texts | `eval/imprint.py::_voice_similarity` |
| `CorrectnessEvaluator` | response vs reference (LLM judge, 1β5 β [0,1]) | `governance/panel.chat_once` |
| `PairwiseEvaluator` | after-utterance better than before toward the persona | judge (A/B/TIE) |
| `GuidelineEvaluator` | rubric/agenda compliance | LLM judge |
Endpoints: `POST /coach/api/classroom/evaluate`, `POST /coach/api/eval/prompt`,
`POST /coach/api/autotune/feedback`.
## Prompt tools β test cheap, promote if it wins
**`/coach/prompts`** treats prompting as the cheapest pseudo-training: craft a system
prompt + few-shot demonstrations, run them against a base model (streaming, **no
training**), evaluate the outcome with the eval tools, and only if it's advantageous
**make it permanent** by baking the prompt + demonstrations into an Ollama Modelfile
(`POST /coach/api/modelfile/create`). Non-permanent experiment β promote on results.
## How mindXtrain fits decentralized training
The dcoach page renders a read-only panel (`GET /coach/api/decentralized`) mapping each
2026 network to where mindXtrain plugs in. mindXtrain **does not mine** on any of them β
every one is CUDA-first / hardware-gated. Instead it exposes a *verifiable, payable*
training surface compatible with their verification primitives:
| mindXtrain primitive | Maps to |
|----------------------|---------|
| AOT-only autotune plan (bit-reproducible run) | Gensyn **Verde + RepOps** training verification |
| BLAKE3 verifiable receipt (`mindxtrain receipt`) | TOPLOC / checkpoint-hash verification; Templar **Gauntlet** auditing |
| x402-metered training job | Per-job crypto metering β unbuilt territory across all networks |
| AgenticPlace / ERC-8004 registration | Pluralis unextractable-model ownership / on-chain attribution |
Networks covered: **Prime Intellect** (open stack, RL post-training), **Templar Β· Bittensor
SN3** (Covenant-72B, the only live incentivized training market), **Nous Β· Psyche**
(DisTrO on Solana), **Gensyn** (verification-first, Verde β the closest match), **Pluralis Β·
Node0** (model-parallel over WAN, unextractable models). Full analysis in
[decentralized-training-deep-dive-2026.md](decentralized-training-deep-dive-2026.md) and
[mindxtrain-llm-training-landscape-2026.md](mindxtrain-llm-training-landscape-2026.md).
## Why this matters
This is mindXtrain's **first-run proof**: that a model trained on the CPU lane actually
*recalls* what it was trained on β measured, ruled on, and fed back, not asserted. It is
also the bridge to the [mindX self-training loop](../README.md): the same loop that imprints
`codephreak` here consumes the `machine.dream` corpus to produce the small model mindX falls
back to.
|