File size: 12,539 Bytes
dfb775d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 | # Actualization status
A per-module map of what's real Python vs. what gracefully degrades to an
install hint or runtime requirement. Reflects the state after the
"actualize stubs" pass; counts and labels track the canonical layout from
[`blueprints/mindxtrain2.md`](blueprints/mindxtrain2.md) Β§Part 4.
## v1.0.0 production-readiness (objective audit, 2026-06-11)
Honest classification for the v1.0.0 release. **CPU is active**; the GPU path is
code-complete but needs real ROCm hardware to execute.
**Production-ready on CPU now (run today, no GPU):**
- Training: `trl_cpu` + `trl_local` (real checkpoints, in-process TRL).
- Data: `hf`, `local` (JSONL), `mindx_dreams` sources; dedupe/filter/tokenize/pack.
- Provenance: BLAKE3 manifest + `emit_receipt_for_run` + `verify` + `mindxtrain receipt`.
- Persona/imprint: script authoring, `imprint` recall before/after, ollama push (LoRA merge).
- Operator/Coach: recipes, bench dry-run, compile, cost, live-training SSE, receipt card,
create-dataset, MEI, training-jobs API, `/v1/chat/completions`.
- Provenance chain helpers: x402 invoice/settlement, ERC-8004 encode/broadcast (need `--extra chain`).
**GPU-ready, hardware-pending (code complete; needs MI300X/ROCm to run):**
- Subprocess training backends `axolotl` / `unsloth` / `torchtune` / `primus` (command-built + unit-tested; not executed e2e here).
- Real autotune probes (attention/GEMM timing) β dry-run reference on CPU.
- Quark FP8/MXFP4 quantize; vLLM / SGLang serve launchers (commands built, serving needs GPU).
**Stubs β NOT claimed as working in 1.0.0 (roadmap):**
- `/v1/agentic` mindX MASTERMIND dispatch β `501` (`operator/app.py`).
- Cloud provisioners `akash` / `ionet` / `bacalhau` / `tensorwave` β `NotImplementedError` (`budget/providers/*`).
- `lighthouse` as a *data source* and `storage/lighthouse.py:get_dir` β redirect to IPFS.
## Headline numbers
- **99** Python modules under `mindxtrain/`.
- **38 actualized** (real implementations using stdlib / already-installed deps + lazy imports).
- **5 cloud-provider stubs** preserved in `budget/providers/*` (post-hackathon).
- **2 deliberate-redirects** that raise with a pointer to a sibling module
(`storage/lighthouse.py:get_dir` β use `storage.ipfs`).
- **566 tests** pass on a CPU-only laptop (`uv run pytest -q`).
- **0** OLD-namespace imports anywhere (`from xtrain.`, `from automindx.`,
`from custmodel` are all gone).
## What `uv sync` (no extras) gives you
Every module is *importable*. Anything that doesn't need a heavyweight
runtime works directly:
| Surface | Status |
|---|---|
| `mindxtrain --help` / `--version` / `init` / `init --list` | works |
| `mindxtrain bench --dry-run` | works (synthetic plan) |
| `mindxtrain receipt <manifest.json>` | works (BLAKE3 verify) |
| `mindxtrain.operator.app` (FastAPI, no chat backend) | boots; `/coach/` UI live |
| `mindxtrain.deploy.{registry,hot_swap,ab_test}` | atomic JSON-backed registry |
| `mindxtrain.operator.{tool_router,agent_loop,context,trajectory,approval}` | bounded ReAct, ContextManager, etc. |
| `mindxtrain.provenance.{manifest,hashing,verify}` | BLAKE3 manifest round-trip |
| `mindxtrain.storage.local_fs` | working |
| `mindxtrain.train.distributed` (FSDP/DeepSpeed config builders) | works |
| `mindxtrain.budget.{pricing,resource}` | works (psutil if installed) |
## What the optional-dep groups unlock
Install with `uv sync --extra <group>` (multiple `--extra` flags allowed,
or `--all-extras`):
| Group | Adds | Unlocks |
|---|---|---|
| `ml` | `trl`, `transformers`, `peft`, `accelerate`, `datasets` | `mindxtrain train`, `mindxtrain dataset prep`, `mindxtrain.train.{sft,dpo,grpo,rlhf,tool_use}`, `mindxtrain.data.{curate,tokenize}`, `mindxtrain.train.callbacks` |
| `eval` | `lm-eval`, `lighteval`, `inspect-ai`, `jinja2` | `mindxtrain eval`, `mindxtrain.eval.{harness,lighteval_adapter,inspect_ai_adapter,bfcl,tau_bench,card}` |
| `data` | `datasketch`, `sentence-transformers`, `faiss-cpu`, `pyarrow` | `mindxtrain.data.{dedupe,filter}` semantic paths, `mindxtrain.eval.persona_regression` |
| `serve` | `vllm` | in-process vLLM (the operator FastAPI app proxies via httpx by default) |
| `chain` | `web3`, `py-algorand-sdk`, `huggingface-hub` | `mindxtrain.provenance.{erc8004.broadcast_attestation,x402.validate_settlement,algorand}`, `mindxtrain.storage.hf_hub` |
| `obs` | `opentelemetry-sdk`, `prometheus-client`, `psutil` | `mindxtrain.operator.telemetry.*`, `mindxtrain.budget.resource.detect` |
The `all` extra installs everything except `amd-quark` (which ships with the
rocm/primus container β see [HANDOFF.md](HANDOFF.md) Β§3).
## Per-subpackage status
### `mindxtrain.cli`
`main.py` β **real**. All 9 verbs (`init`, `bench`, `train`, `eval`,
`quantize`, `serve`, `publish`, `receipt`, `dataset prep`) dispatch into
canonical modules. Exit codes: `0` = ok, `1` = bad input / missing file,
`3` = optional dep missing.
### `mindxtrain.config`
`schema.py` (Pydantic 10-section `XTrainConfig`) and `loader.py`
(YAML render + load) β **real, frozen**. Three runtime-defaults JSON files
(`train_default.json`, `eval_default.json`, `deploy_default.json`) ship as
`${ENV}`-interpolated templates per mindxtrain2.md ml-intern style.
### `mindxtrain.data`
| Module | Status | Dep group |
|---|---|---|
| `curate.py` | real (HF datasets streaming) | `--extra ml` |
| `dedupe.py` | real MinHash + SemDeDup | `--extra data` |
| `filter.py` | real (length/repeat/alpha heuristics + optional KenLM) | none (stdlib) |
| `pack.py` | real (greedy first-fit + tar shards) | none (stdlib) |
| `synth.py` | real (httpx β vLLM teacher endpoint) | needs reachable `MINDXTRAIN_TEACHER_BASE_URL` |
| `tokenize.py` | real (AutoTokenizer wrap) | `--extra ml` |
| `verify.py` | real (BLAKE3 walk vs manifest) | none |
### `mindxtrain.models`
| Module | Status |
|---|---|
| `registry.py` | real (Backend ABC + ModelRegistry + preset registry) |
| `chat_template.py` | real (Hermes/Qwen3-Coder/Qwen3-Reasoning parsers) |
| `glm51.py`, `qwen35.py`, `deepseek_v32.py`, `mistral3.py`, `phi4_mini.py` | real Pydantic presets, auto-register on import |
### `mindxtrain.train`
| Module | Status | Dep group |
|---|---|---|
| `dispatch.py` | real 4-way switch | none |
| `axolotl_compile.py` | real (XTrainConfig β Axolotl YAML) | none |
| `sft.py` | real subprocess wrap of `accelerate launch -m axolotl.cli.train` | `--extra ml` + axolotl on PATH |
| `dpo.py`, `grpo.py`, `rlhf.py`, `tool_use.py` | real TRL trainer wraps | `--extra ml` |
| `distributed.py` | real (FSDP / DeepSpeed dict builders, 1- or 8-GPU only) | none |
| `callbacks.py` | real `EvalDuringTraining` + `BestCheckpointKeeper` | `--extra ml` (lazy) |
| `backend_unsloth.py`, `backend_torchtune.py`, `backend_primus.py` | real subprocess wraps | each backend's own install |
### `mindxtrain.eval`
| Module | Status | Dep group |
|---|---|---|
| `harness.py` | real `lm_eval` subprocess + JSON parser | `--extra eval` |
| `lighteval_adapter.py` | real `lighteval accelerate` wrap | `--extra eval` |
| `inspect_ai_adapter.py` | real `inspect eval` wrap | `--extra eval` |
| `bfcl.py` | real `bfcl evaluate` wrap | external (BFCL harness) |
| `tau_bench.py` | real subprocess wrap | external |
| `persona_regression.py` | real (sentence-transformer cosine vs baseline) | `--extra data` |
| `agenda_regression.py` | real (keyword overlap + optional LLM judge) | none + optional `MINDXTRAIN_TEACHER_BASE_URL` |
| `card.py` | real (Jinja2 with stdlib `string.Template` fallback) | optional `--extra eval` |
### `mindxtrain.autotune`
| Module | Status |
|---|---|
| `benchmark.py`, `plan.py`, `gemm_probe.py`, `rccl_probe.py` | real |
| `attention_probe.py` | real (CK vs Triton SDPA timing if torch+ROCm available; CPU fallback returns canonical default) |
### `mindxtrain.operator`
| Module | Status |
|---|---|
| `app.py` (FastAPI), `coach/api.py`, `coach/static/*` | real |
| `tool_router.py` | real (typed `ToolSpec` + dispatch) |
| `agent_loop.py` | real (bounded ReAct + doom-loop detector) |
| `context.py` | real (170k-token compaction + summarize fallback) |
| `trajectory.py` | real (JSONL append-only writer) |
| `approval.py` | real (CLI / Web / Slack transports) |
| `backends/{vllm,openai_compat}.py` | real (httpx clients to OpenAI-compat endpoints) |
| `telemetry/{energy,otel_hooks,prometheus_exporter}.py` | real, gracefully no-op if optional deps missing |
| `prompts/{system_v1,codephreak}.yaml` | real prompt-as-data |
### `mindxtrain.storage`
| Module | Status | Dep group |
|---|---|---|
| `provider.py` | real ABC | none |
| `local_fs.py` | real | none |
| `hf_hub.py` | real (huggingface_hub upload_folder) | `--extra chain` |
| `lighthouse.py` | real httpx POST to Lighthouse REST API; falls back to stub-CID without `LIGHTHOUSE_API_KEY` | none |
| `ipfs.py` | real httpx to kubo `/api/v0/add` | needs running kubo |
### `mindxtrain.provenance`
| Module | Status | Dep group |
|---|---|---|
| `manifest.py` | real (`Manifest` + `emit_receipt`) | none |
| `hashing.py` | real (BLAKE3 file/dir) | none |
| `verify.py` | real (re-hash on-disk artifacts) | none |
| `x402.py` | real httpx invoice + Algorand verify | `--extra chain` |
| `erc8004.py` | real ABI encode + web3 broadcast | `--extra chain` |
| `algorand.py` | real BANKON ENS allocator + ASA info | `--extra chain` |
### `mindxtrain.deploy`
| Module | Status | Dep group |
|---|---|---|
| `registry.py` | real atomic JSON-backed registry | none |
| `hot_swap.py` | real canary-promote + rollback | none |
| `ab_test.py` | real deterministic Splitter | none |
| `api_client.py` | real httpx β mindx.pythai.net + agenticplace.pythai.net | needs deployed services |
| `vllm_launcher.py`, `sglang_rocm.py` | real argv builders | none |
| `quark.py` | real subprocess wrap of `python -m amd_quark.quantize` | rocm/primus container |
| `gptq_rocm.py` | real subprocess wrap | `auto-gptq` ROCm wheel |
### `mindxtrain.budget`
| Module | Status | Dep group |
|---|---|---|
| `pricing.py` | real | none |
| `resource.py` | real (psutil + rocm-smi probes; falls back to defaults) | optional `--extra obs` |
| `providers/{akash,amd_dev_cloud,bacalhau,ionet,tensorwave}.py` | **stubs** (post-hackathon) | each provider's SDK |
## What stays as `NotImplementedError`
7 residual `NotImplementedError` raises across the package:
- `budget/providers/akash.py`, `amd_dev_cloud.py`, `bacalhau.py`, `ionet.py`,
`tensorwave.py` β cloud-burst provisioning. Out of hackathon scope.
- `storage/lighthouse.py:LighthouseProvider.get_dir` β deliberately
redirects to `mindxtrain.storage.ipfs.IpfsProvider.get_dir`.
- `train/dispatch.py` β string match in a docstring, not an actual raise.
Run `grep -r "raise NotImplementedError" mindxtrain` to confirm.
## Test coverage
```
tests/
βββ test_ab_test.py # canary splitter distribution
βββ test_agent_loop.py # bounded ReAct + doom-loop
βββ test_autotune_plan.py # AutotunePlan invariants
βββ test_axolotl_compile.py # XTrainConfig β Axolotl YAML
βββ test_cli_smoke.py # all 9 verbs reachable
βββ test_coach_api.py # /coach/api/* endpoints
βββ test_config_schema.py # 10-section schema, recipe round-trip
βββ test_context_manager.py # ContextManager compaction
βββ test_data_pipeline.py # filter / synth / verify
βββ test_deploy_registry.py # registry + hot-swap atomicity
βββ test_distributed.py # FSDP/DeepSpeed builders, xGMI invariant
βββ test_manifest.py # Manifest + BLAKE3 round-trip
βββ test_models_registry.py # preset + chat-template lookup
βββ test_pack.py # greedy first-fit packer + tar shards
βββ test_parsers.py # chat templates
βββ test_pricing.py # MI300X $/hr math
βββ test_provenance_verify.py # tamper detection
βββ test_tool_router.py # ToolSpec dispatch
βββ test_vllm_launcher.py # vLLM cmd builder
```
`uv run pytest -q` β **564 passed**.
## See also
- [HANDOFF.md](HANDOFF.md) β ordered checklist for taking the project from "code is done" to "demo is live."
- [development.md](development.md) β toolchain, lazy-import pattern, how to add features.
- [architecture.md](architecture.md) β canonical layout + 5-layer architecture.
|