File size: 12,539 Bytes
dfb775d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
# Actualization status

A per-module map of what's real Python vs. what gracefully degrades to an
install hint or runtime requirement. Reflects the state after the
"actualize stubs" pass; counts and labels track the canonical layout from
[`blueprints/mindxtrain2.md`](blueprints/mindxtrain2.md) Β§Part 4.

## v1.0.0 production-readiness (objective audit, 2026-06-11)

Honest classification for the v1.0.0 release. **CPU is active**; the GPU path is
code-complete but needs real ROCm hardware to execute.

**Production-ready on CPU now (run today, no GPU):**
- Training: `trl_cpu` + `trl_local` (real checkpoints, in-process TRL).
- Data: `hf`, `local` (JSONL), `mindx_dreams` sources; dedupe/filter/tokenize/pack.
- Provenance: BLAKE3 manifest + `emit_receipt_for_run` + `verify` + `mindxtrain receipt`.
- Persona/imprint: script authoring, `imprint` recall before/after, ollama push (LoRA merge).
- Operator/Coach: recipes, bench dry-run, compile, cost, live-training SSE, receipt card,
  create-dataset, MEI, training-jobs API, `/v1/chat/completions`.
- Provenance chain helpers: x402 invoice/settlement, ERC-8004 encode/broadcast (need `--extra chain`).

**GPU-ready, hardware-pending (code complete; needs MI300X/ROCm to run):**
- Subprocess training backends `axolotl` / `unsloth` / `torchtune` / `primus` (command-built + unit-tested; not executed e2e here).
- Real autotune probes (attention/GEMM timing) β€” dry-run reference on CPU.
- Quark FP8/MXFP4 quantize; vLLM / SGLang serve launchers (commands built, serving needs GPU).

**Stubs β€” NOT claimed as working in 1.0.0 (roadmap):**
- `/v1/agentic` mindX MASTERMIND dispatch β†’ `501` (`operator/app.py`).
- Cloud provisioners `akash` / `ionet` / `bacalhau` / `tensorwave` β†’ `NotImplementedError` (`budget/providers/*`).
- `lighthouse` as a *data source* and `storage/lighthouse.py:get_dir` β†’ redirect to IPFS.

## Headline numbers

- **99** Python modules under `mindxtrain/`.
- **38 actualized** (real implementations using stdlib / already-installed deps + lazy imports).
- **5 cloud-provider stubs** preserved in `budget/providers/*` (post-hackathon).
- **2 deliberate-redirects** that raise with a pointer to a sibling module
  (`storage/lighthouse.py:get_dir` β†’ use `storage.ipfs`).
- **566 tests** pass on a CPU-only laptop (`uv run pytest -q`).
- **0** OLD-namespace imports anywhere (`from xtrain.`, `from automindx.`,
  `from custmodel` are all gone).

## What `uv sync` (no extras) gives you

Every module is *importable*. Anything that doesn't need a heavyweight
runtime works directly:

| Surface | Status |
|---|---|
| `mindxtrain --help` / `--version` / `init` / `init --list` | works |
| `mindxtrain bench --dry-run` | works (synthetic plan) |
| `mindxtrain receipt <manifest.json>` | works (BLAKE3 verify) |
| `mindxtrain.operator.app` (FastAPI, no chat backend) | boots; `/coach/` UI live |
| `mindxtrain.deploy.{registry,hot_swap,ab_test}` | atomic JSON-backed registry |
| `mindxtrain.operator.{tool_router,agent_loop,context,trajectory,approval}` | bounded ReAct, ContextManager, etc. |
| `mindxtrain.provenance.{manifest,hashing,verify}` | BLAKE3 manifest round-trip |
| `mindxtrain.storage.local_fs` | working |
| `mindxtrain.train.distributed` (FSDP/DeepSpeed config builders) | works |
| `mindxtrain.budget.{pricing,resource}` | works (psutil if installed) |

## What the optional-dep groups unlock

Install with `uv sync --extra <group>` (multiple `--extra` flags allowed,
or `--all-extras`):

| Group | Adds | Unlocks |
|---|---|---|
| `ml` | `trl`, `transformers`, `peft`, `accelerate`, `datasets` | `mindxtrain train`, `mindxtrain dataset prep`, `mindxtrain.train.{sft,dpo,grpo,rlhf,tool_use}`, `mindxtrain.data.{curate,tokenize}`, `mindxtrain.train.callbacks` |
| `eval` | `lm-eval`, `lighteval`, `inspect-ai`, `jinja2` | `mindxtrain eval`, `mindxtrain.eval.{harness,lighteval_adapter,inspect_ai_adapter,bfcl,tau_bench,card}` |
| `data` | `datasketch`, `sentence-transformers`, `faiss-cpu`, `pyarrow` | `mindxtrain.data.{dedupe,filter}` semantic paths, `mindxtrain.eval.persona_regression` |
| `serve` | `vllm` | in-process vLLM (the operator FastAPI app proxies via httpx by default) |
| `chain` | `web3`, `py-algorand-sdk`, `huggingface-hub` | `mindxtrain.provenance.{erc8004.broadcast_attestation,x402.validate_settlement,algorand}`, `mindxtrain.storage.hf_hub` |
| `obs` | `opentelemetry-sdk`, `prometheus-client`, `psutil` | `mindxtrain.operator.telemetry.*`, `mindxtrain.budget.resource.detect` |

The `all` extra installs everything except `amd-quark` (which ships with the
rocm/primus container β€” see [HANDOFF.md](HANDOFF.md) Β§3).

## Per-subpackage status

### `mindxtrain.cli`

`main.py` β€” **real**. All 9 verbs (`init`, `bench`, `train`, `eval`,
`quantize`, `serve`, `publish`, `receipt`, `dataset prep`) dispatch into
canonical modules. Exit codes: `0` = ok, `1` = bad input / missing file,
`3` = optional dep missing.

### `mindxtrain.config`

`schema.py` (Pydantic 10-section `XTrainConfig`) and `loader.py`
(YAML render + load) β€” **real, frozen**. Three runtime-defaults JSON files
(`train_default.json`, `eval_default.json`, `deploy_default.json`) ship as
`${ENV}`-interpolated templates per mindxtrain2.md ml-intern style.

### `mindxtrain.data`

| Module | Status | Dep group |
|---|---|---|
| `curate.py` | real (HF datasets streaming) | `--extra ml` |
| `dedupe.py` | real MinHash + SemDeDup | `--extra data` |
| `filter.py` | real (length/repeat/alpha heuristics + optional KenLM) | none (stdlib) |
| `pack.py` | real (greedy first-fit + tar shards) | none (stdlib) |
| `synth.py` | real (httpx β†’ vLLM teacher endpoint) | needs reachable `MINDXTRAIN_TEACHER_BASE_URL` |
| `tokenize.py` | real (AutoTokenizer wrap) | `--extra ml` |
| `verify.py` | real (BLAKE3 walk vs manifest) | none |

### `mindxtrain.models`

| Module | Status |
|---|---|
| `registry.py` | real (Backend ABC + ModelRegistry + preset registry) |
| `chat_template.py` | real (Hermes/Qwen3-Coder/Qwen3-Reasoning parsers) |
| `glm51.py`, `qwen35.py`, `deepseek_v32.py`, `mistral3.py`, `phi4_mini.py` | real Pydantic presets, auto-register on import |

### `mindxtrain.train`

| Module | Status | Dep group |
|---|---|---|
| `dispatch.py` | real 4-way switch | none |
| `axolotl_compile.py` | real (XTrainConfig β†’ Axolotl YAML) | none |
| `sft.py` | real subprocess wrap of `accelerate launch -m axolotl.cli.train` | `--extra ml` + axolotl on PATH |
| `dpo.py`, `grpo.py`, `rlhf.py`, `tool_use.py` | real TRL trainer wraps | `--extra ml` |
| `distributed.py` | real (FSDP / DeepSpeed dict builders, 1- or 8-GPU only) | none |
| `callbacks.py` | real `EvalDuringTraining` + `BestCheckpointKeeper` | `--extra ml` (lazy) |
| `backend_unsloth.py`, `backend_torchtune.py`, `backend_primus.py` | real subprocess wraps | each backend's own install |

### `mindxtrain.eval`

| Module | Status | Dep group |
|---|---|---|
| `harness.py` | real `lm_eval` subprocess + JSON parser | `--extra eval` |
| `lighteval_adapter.py` | real `lighteval accelerate` wrap | `--extra eval` |
| `inspect_ai_adapter.py` | real `inspect eval` wrap | `--extra eval` |
| `bfcl.py` | real `bfcl evaluate` wrap | external (BFCL harness) |
| `tau_bench.py` | real subprocess wrap | external |
| `persona_regression.py` | real (sentence-transformer cosine vs baseline) | `--extra data` |
| `agenda_regression.py` | real (keyword overlap + optional LLM judge) | none + optional `MINDXTRAIN_TEACHER_BASE_URL` |
| `card.py` | real (Jinja2 with stdlib `string.Template` fallback) | optional `--extra eval` |

### `mindxtrain.autotune`

| Module | Status |
|---|---|
| `benchmark.py`, `plan.py`, `gemm_probe.py`, `rccl_probe.py` | real |
| `attention_probe.py` | real (CK vs Triton SDPA timing if torch+ROCm available; CPU fallback returns canonical default) |

### `mindxtrain.operator`

| Module | Status |
|---|---|
| `app.py` (FastAPI), `coach/api.py`, `coach/static/*` | real |
| `tool_router.py` | real (typed `ToolSpec` + dispatch) |
| `agent_loop.py` | real (bounded ReAct + doom-loop detector) |
| `context.py` | real (170k-token compaction + summarize fallback) |
| `trajectory.py` | real (JSONL append-only writer) |
| `approval.py` | real (CLI / Web / Slack transports) |
| `backends/{vllm,openai_compat}.py` | real (httpx clients to OpenAI-compat endpoints) |
| `telemetry/{energy,otel_hooks,prometheus_exporter}.py` | real, gracefully no-op if optional deps missing |
| `prompts/{system_v1,codephreak}.yaml` | real prompt-as-data |

### `mindxtrain.storage`

| Module | Status | Dep group |
|---|---|---|
| `provider.py` | real ABC | none |
| `local_fs.py` | real | none |
| `hf_hub.py` | real (huggingface_hub upload_folder) | `--extra chain` |
| `lighthouse.py` | real httpx POST to Lighthouse REST API; falls back to stub-CID without `LIGHTHOUSE_API_KEY` | none |
| `ipfs.py` | real httpx to kubo `/api/v0/add` | needs running kubo |

### `mindxtrain.provenance`

| Module | Status | Dep group |
|---|---|---|
| `manifest.py` | real (`Manifest` + `emit_receipt`) | none |
| `hashing.py` | real (BLAKE3 file/dir) | none |
| `verify.py` | real (re-hash on-disk artifacts) | none |
| `x402.py` | real httpx invoice + Algorand verify | `--extra chain` |
| `erc8004.py` | real ABI encode + web3 broadcast | `--extra chain` |
| `algorand.py` | real BANKON ENS allocator + ASA info | `--extra chain` |

### `mindxtrain.deploy`

| Module | Status | Dep group |
|---|---|---|
| `registry.py` | real atomic JSON-backed registry | none |
| `hot_swap.py` | real canary-promote + rollback | none |
| `ab_test.py` | real deterministic Splitter | none |
| `api_client.py` | real httpx β†’ mindx.pythai.net + agenticplace.pythai.net | needs deployed services |
| `vllm_launcher.py`, `sglang_rocm.py` | real argv builders | none |
| `quark.py` | real subprocess wrap of `python -m amd_quark.quantize` | rocm/primus container |
| `gptq_rocm.py` | real subprocess wrap | `auto-gptq` ROCm wheel |

### `mindxtrain.budget`

| Module | Status | Dep group |
|---|---|---|
| `pricing.py` | real | none |
| `resource.py` | real (psutil + rocm-smi probes; falls back to defaults) | optional `--extra obs` |
| `providers/{akash,amd_dev_cloud,bacalhau,ionet,tensorwave}.py` | **stubs** (post-hackathon) | each provider's SDK |

## What stays as `NotImplementedError`

7 residual `NotImplementedError` raises across the package:

- `budget/providers/akash.py`, `amd_dev_cloud.py`, `bacalhau.py`, `ionet.py`,
  `tensorwave.py` β€” cloud-burst provisioning. Out of hackathon scope.
- `storage/lighthouse.py:LighthouseProvider.get_dir` β€” deliberately
  redirects to `mindxtrain.storage.ipfs.IpfsProvider.get_dir`.
- `train/dispatch.py` β€” string match in a docstring, not an actual raise.

Run `grep -r "raise NotImplementedError" mindxtrain` to confirm.

## Test coverage

```
tests/
β”œβ”€β”€ test_ab_test.py                   # canary splitter distribution
β”œβ”€β”€ test_agent_loop.py                # bounded ReAct + doom-loop
β”œβ”€β”€ test_autotune_plan.py             # AutotunePlan invariants
β”œβ”€β”€ test_axolotl_compile.py           # XTrainConfig β†’ Axolotl YAML
β”œβ”€β”€ test_cli_smoke.py                 # all 9 verbs reachable
β”œβ”€β”€ test_coach_api.py                 # /coach/api/* endpoints
β”œβ”€β”€ test_config_schema.py             # 10-section schema, recipe round-trip
β”œβ”€β”€ test_context_manager.py           # ContextManager compaction
β”œβ”€β”€ test_data_pipeline.py             # filter / synth / verify
β”œβ”€β”€ test_deploy_registry.py           # registry + hot-swap atomicity
β”œβ”€β”€ test_distributed.py               # FSDP/DeepSpeed builders, xGMI invariant
β”œβ”€β”€ test_manifest.py                  # Manifest + BLAKE3 round-trip
β”œβ”€β”€ test_models_registry.py           # preset + chat-template lookup
β”œβ”€β”€ test_pack.py                      # greedy first-fit packer + tar shards
β”œβ”€β”€ test_parsers.py                   # chat templates
β”œβ”€β”€ test_pricing.py                   # MI300X $/hr math
β”œβ”€β”€ test_provenance_verify.py         # tamper detection
β”œβ”€β”€ test_tool_router.py               # ToolSpec dispatch
└── test_vllm_launcher.py             # vLLM cmd builder
```

`uv run pytest -q` β†’ **564 passed**.

## See also

- [HANDOFF.md](HANDOFF.md) β€” ordered checklist for taking the project from "code is done" to "demo is live."
- [development.md](development.md) β€” toolchain, lazy-import pattern, how to add features.
- [architecture.md](architecture.md) β€” canonical layout + 5-layer architecture.