mindXtrain / docs /actualization_status.md
Gregory-L's picture
fork mindXtrain from GitHub (Professor-Codephreak/mindXtrain@661bd41) as the mindX-specific line
dfb775d verified
|
Raw History Blame Contribute Delete
12.5 kB
# Actualization status
A per-module map of what's real Python vs. what gracefully degrades to an
install hint or runtime requirement. Reflects the state after the
"actualize stubs" pass; counts and labels track the canonical layout from
[`blueprints/mindxtrain2.md`](blueprints/mindxtrain2.md) Β§Part 4.
## v1.0.0 production-readiness (objective audit, 2026-06-11)
Honest classification for the v1.0.0 release. **CPU is active**; the GPU path is
code-complete but needs real ROCm hardware to execute.
**Production-ready on CPU now (run today, no GPU):**
- Training: `trl_cpu` + `trl_local` (real checkpoints, in-process TRL).
- Data: `hf`, `local` (JSONL), `mindx_dreams` sources; dedupe/filter/tokenize/pack.
- Provenance: BLAKE3 manifest + `emit_receipt_for_run` + `verify` + `mindxtrain receipt`.
- Persona/imprint: script authoring, `imprint` recall before/after, ollama push (LoRA merge).
- Operator/Coach: recipes, bench dry-run, compile, cost, live-training SSE, receipt card,
create-dataset, MEI, training-jobs API, `/v1/chat/completions`.
- Provenance chain helpers: x402 invoice/settlement, ERC-8004 encode/broadcast (need `--extra chain`).
**GPU-ready, hardware-pending (code complete; needs MI300X/ROCm to run):**
- Subprocess training backends `axolotl` / `unsloth` / `torchtune` / `primus` (command-built + unit-tested; not executed e2e here).
- Real autotune probes (attention/GEMM timing) β€” dry-run reference on CPU.
- Quark FP8/MXFP4 quantize; vLLM / SGLang serve launchers (commands built, serving needs GPU).
**Stubs β€” NOT claimed as working in 1.0.0 (roadmap):**
- `/v1/agentic` mindX MASTERMIND dispatch β†’ `501` (`operator/app.py`).
- Cloud provisioners `akash` / `ionet` / `bacalhau` / `tensorwave` β†’ `NotImplementedError` (`budget/providers/*`).
- `lighthouse` as a *data source* and `storage/lighthouse.py:get_dir` β†’ redirect to IPFS.
## Headline numbers
- **99** Python modules under `mindxtrain/`.
- **38 actualized** (real implementations using stdlib / already-installed deps + lazy imports).
- **5 cloud-provider stubs** preserved in `budget/providers/*` (post-hackathon).
- **2 deliberate-redirects** that raise with a pointer to a sibling module
(`storage/lighthouse.py:get_dir` β†’ use `storage.ipfs`).
- **566 tests** pass on a CPU-only laptop (`uv run pytest -q`).
- **0** OLD-namespace imports anywhere (`from xtrain.`, `from automindx.`,
`from custmodel` are all gone).
## What `uv sync` (no extras) gives you
Every module is *importable*. Anything that doesn't need a heavyweight
runtime works directly:
| Surface | Status |
|---|---|
| `mindxtrain --help` / `--version` / `init` / `init --list` | works |
| `mindxtrain bench --dry-run` | works (synthetic plan) |
| `mindxtrain receipt <manifest.json>` | works (BLAKE3 verify) |
| `mindxtrain.operator.app` (FastAPI, no chat backend) | boots; `/coach/` UI live |
| `mindxtrain.deploy.{registry,hot_swap,ab_test}` | atomic JSON-backed registry |
| `mindxtrain.operator.{tool_router,agent_loop,context,trajectory,approval}` | bounded ReAct, ContextManager, etc. |
| `mindxtrain.provenance.{manifest,hashing,verify}` | BLAKE3 manifest round-trip |
| `mindxtrain.storage.local_fs` | working |
| `mindxtrain.train.distributed` (FSDP/DeepSpeed config builders) | works |
| `mindxtrain.budget.{pricing,resource}` | works (psutil if installed) |
## What the optional-dep groups unlock
Install with `uv sync --extra <group>` (multiple `--extra` flags allowed,
or `--all-extras`):
| Group | Adds | Unlocks |
|---|---|---|
| `ml` | `trl`, `transformers`, `peft`, `accelerate`, `datasets` | `mindxtrain train`, `mindxtrain dataset prep`, `mindxtrain.train.{sft,dpo,grpo,rlhf,tool_use}`, `mindxtrain.data.{curate,tokenize}`, `mindxtrain.train.callbacks` |
| `eval` | `lm-eval`, `lighteval`, `inspect-ai`, `jinja2` | `mindxtrain eval`, `mindxtrain.eval.{harness,lighteval_adapter,inspect_ai_adapter,bfcl,tau_bench,card}` |
| `data` | `datasketch`, `sentence-transformers`, `faiss-cpu`, `pyarrow` | `mindxtrain.data.{dedupe,filter}` semantic paths, `mindxtrain.eval.persona_regression` |
| `serve` | `vllm` | in-process vLLM (the operator FastAPI app proxies via httpx by default) |
| `chain` | `web3`, `py-algorand-sdk`, `huggingface-hub` | `mindxtrain.provenance.{erc8004.broadcast_attestation,x402.validate_settlement,algorand}`, `mindxtrain.storage.hf_hub` |
| `obs` | `opentelemetry-sdk`, `prometheus-client`, `psutil` | `mindxtrain.operator.telemetry.*`, `mindxtrain.budget.resource.detect` |
The `all` extra installs everything except `amd-quark` (which ships with the
rocm/primus container β€” see [HANDOFF.md](HANDOFF.md) Β§3).
## Per-subpackage status
### `mindxtrain.cli`
`main.py` β€” **real**. All 9 verbs (`init`, `bench`, `train`, `eval`,
`quantize`, `serve`, `publish`, `receipt`, `dataset prep`) dispatch into
canonical modules. Exit codes: `0` = ok, `1` = bad input / missing file,
`3` = optional dep missing.
### `mindxtrain.config`
`schema.py` (Pydantic 10-section `XTrainConfig`) and `loader.py`
(YAML render + load) β€” **real, frozen**. Three runtime-defaults JSON files
(`train_default.json`, `eval_default.json`, `deploy_default.json`) ship as
`${ENV}`-interpolated templates per mindxtrain2.md ml-intern style.
### `mindxtrain.data`
| Module | Status | Dep group |
|---|---|---|
| `curate.py` | real (HF datasets streaming) | `--extra ml` |
| `dedupe.py` | real MinHash + SemDeDup | `--extra data` |
| `filter.py` | real (length/repeat/alpha heuristics + optional KenLM) | none (stdlib) |
| `pack.py` | real (greedy first-fit + tar shards) | none (stdlib) |
| `synth.py` | real (httpx β†’ vLLM teacher endpoint) | needs reachable `MINDXTRAIN_TEACHER_BASE_URL` |
| `tokenize.py` | real (AutoTokenizer wrap) | `--extra ml` |
| `verify.py` | real (BLAKE3 walk vs manifest) | none |
### `mindxtrain.models`
| Module | Status |
|---|---|
| `registry.py` | real (Backend ABC + ModelRegistry + preset registry) |
| `chat_template.py` | real (Hermes/Qwen3-Coder/Qwen3-Reasoning parsers) |
| `glm51.py`, `qwen35.py`, `deepseek_v32.py`, `mistral3.py`, `phi4_mini.py` | real Pydantic presets, auto-register on import |
### `mindxtrain.train`
| Module | Status | Dep group |
|---|---|---|
| `dispatch.py` | real 4-way switch | none |
| `axolotl_compile.py` | real (XTrainConfig β†’ Axolotl YAML) | none |
| `sft.py` | real subprocess wrap of `accelerate launch -m axolotl.cli.train` | `--extra ml` + axolotl on PATH |
| `dpo.py`, `grpo.py`, `rlhf.py`, `tool_use.py` | real TRL trainer wraps | `--extra ml` |
| `distributed.py` | real (FSDP / DeepSpeed dict builders, 1- or 8-GPU only) | none |
| `callbacks.py` | real `EvalDuringTraining` + `BestCheckpointKeeper` | `--extra ml` (lazy) |
| `backend_unsloth.py`, `backend_torchtune.py`, `backend_primus.py` | real subprocess wraps | each backend's own install |
### `mindxtrain.eval`
| Module | Status | Dep group |
|---|---|---|
| `harness.py` | real `lm_eval` subprocess + JSON parser | `--extra eval` |
| `lighteval_adapter.py` | real `lighteval accelerate` wrap | `--extra eval` |
| `inspect_ai_adapter.py` | real `inspect eval` wrap | `--extra eval` |
| `bfcl.py` | real `bfcl evaluate` wrap | external (BFCL harness) |
| `tau_bench.py` | real subprocess wrap | external |
| `persona_regression.py` | real (sentence-transformer cosine vs baseline) | `--extra data` |
| `agenda_regression.py` | real (keyword overlap + optional LLM judge) | none + optional `MINDXTRAIN_TEACHER_BASE_URL` |
| `card.py` | real (Jinja2 with stdlib `string.Template` fallback) | optional `--extra eval` |
### `mindxtrain.autotune`
| Module | Status |
|---|---|
| `benchmark.py`, `plan.py`, `gemm_probe.py`, `rccl_probe.py` | real |
| `attention_probe.py` | real (CK vs Triton SDPA timing if torch+ROCm available; CPU fallback returns canonical default) |
### `mindxtrain.operator`
| Module | Status |
|---|---|
| `app.py` (FastAPI), `coach/api.py`, `coach/static/*` | real |
| `tool_router.py` | real (typed `ToolSpec` + dispatch) |
| `agent_loop.py` | real (bounded ReAct + doom-loop detector) |
| `context.py` | real (170k-token compaction + summarize fallback) |
| `trajectory.py` | real (JSONL append-only writer) |
| `approval.py` | real (CLI / Web / Slack transports) |
| `backends/{vllm,openai_compat}.py` | real (httpx clients to OpenAI-compat endpoints) |
| `telemetry/{energy,otel_hooks,prometheus_exporter}.py` | real, gracefully no-op if optional deps missing |
| `prompts/{system_v1,codephreak}.yaml` | real prompt-as-data |
### `mindxtrain.storage`
| Module | Status | Dep group |
|---|---|---|
| `provider.py` | real ABC | none |
| `local_fs.py` | real | none |
| `hf_hub.py` | real (huggingface_hub upload_folder) | `--extra chain` |
| `lighthouse.py` | real httpx POST to Lighthouse REST API; falls back to stub-CID without `LIGHTHOUSE_API_KEY` | none |
| `ipfs.py` | real httpx to kubo `/api/v0/add` | needs running kubo |
### `mindxtrain.provenance`
| Module | Status | Dep group |
|---|---|---|
| `manifest.py` | real (`Manifest` + `emit_receipt`) | none |
| `hashing.py` | real (BLAKE3 file/dir) | none |
| `verify.py` | real (re-hash on-disk artifacts) | none |
| `x402.py` | real httpx invoice + Algorand verify | `--extra chain` |
| `erc8004.py` | real ABI encode + web3 broadcast | `--extra chain` |
| `algorand.py` | real BANKON ENS allocator + ASA info | `--extra chain` |
### `mindxtrain.deploy`
| Module | Status | Dep group |
|---|---|---|
| `registry.py` | real atomic JSON-backed registry | none |
| `hot_swap.py` | real canary-promote + rollback | none |
| `ab_test.py` | real deterministic Splitter | none |
| `api_client.py` | real httpx β†’ mindx.pythai.net + agenticplace.pythai.net | needs deployed services |
| `vllm_launcher.py`, `sglang_rocm.py` | real argv builders | none |
| `quark.py` | real subprocess wrap of `python -m amd_quark.quantize` | rocm/primus container |
| `gptq_rocm.py` | real subprocess wrap | `auto-gptq` ROCm wheel |
### `mindxtrain.budget`
| Module | Status | Dep group |
|---|---|---|
| `pricing.py` | real | none |
| `resource.py` | real (psutil + rocm-smi probes; falls back to defaults) | optional `--extra obs` |
| `providers/{akash,amd_dev_cloud,bacalhau,ionet,tensorwave}.py` | **stubs** (post-hackathon) | each provider's SDK |
## What stays as `NotImplementedError`
7 residual `NotImplementedError` raises across the package:
- `budget/providers/akash.py`, `amd_dev_cloud.py`, `bacalhau.py`, `ionet.py`,
`tensorwave.py` β€” cloud-burst provisioning. Out of hackathon scope.
- `storage/lighthouse.py:LighthouseProvider.get_dir` β€” deliberately
redirects to `mindxtrain.storage.ipfs.IpfsProvider.get_dir`.
- `train/dispatch.py` β€” string match in a docstring, not an actual raise.
Run `grep -r "raise NotImplementedError" mindxtrain` to confirm.
## Test coverage
```
tests/
β”œβ”€β”€ test_ab_test.py # canary splitter distribution
β”œβ”€β”€ test_agent_loop.py # bounded ReAct + doom-loop
β”œβ”€β”€ test_autotune_plan.py # AutotunePlan invariants
β”œβ”€β”€ test_axolotl_compile.py # XTrainConfig β†’ Axolotl YAML
β”œβ”€β”€ test_cli_smoke.py # all 9 verbs reachable
β”œβ”€β”€ test_coach_api.py # /coach/api/* endpoints
β”œβ”€β”€ test_config_schema.py # 10-section schema, recipe round-trip
β”œβ”€β”€ test_context_manager.py # ContextManager compaction
β”œβ”€β”€ test_data_pipeline.py # filter / synth / verify
β”œβ”€β”€ test_deploy_registry.py # registry + hot-swap atomicity
β”œβ”€β”€ test_distributed.py # FSDP/DeepSpeed builders, xGMI invariant
β”œβ”€β”€ test_manifest.py # Manifest + BLAKE3 round-trip
β”œβ”€β”€ test_models_registry.py # preset + chat-template lookup
β”œβ”€β”€ test_pack.py # greedy first-fit packer + tar shards
β”œβ”€β”€ test_parsers.py # chat templates
β”œβ”€β”€ test_pricing.py # MI300X $/hr math
β”œβ”€β”€ test_provenance_verify.py # tamper detection
β”œβ”€β”€ test_tool_router.py # ToolSpec dispatch
└── test_vllm_launcher.py # vLLM cmd builder
```
`uv run pytest -q` β†’ **564 passed**.
## See also
- [HANDOFF.md](HANDOFF.md) β€” ordered checklist for taking the project from "code is done" to "demo is live."
- [development.md](development.md) β€” toolchain, lazy-import pattern, how to add features.
- [architecture.md](architecture.md) β€” canonical layout + 5-layer architecture.