mindXtrain / CLAUDE.md
Gregory-L's picture
bankml: the verified CPU engine as backend, serve target and imprint probe (#1)
730c5bb
|
Raw History Blame Contribute Delete
9.72 kB
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## What this is
`mindxtrain` is a single-package Python training framework for fine-tuning open-weight LLMs on AMD MI300X and serving them through an OpenAI-compatible API. The architectural differentiator is a **60-second AOT autotune probe** (`mindxtrain bench`) that fixes attention backend (CK vs Triton), GEMM heuristic, and RCCL config at training start β€” **JIT autotune is forbidden in the production training loop**.
The base install is CPU-only and runs the CLI, Coach UI, `bench --dry-run`, manifest verify, and the operator FastAPI; heavyweight paths gate on opt-in dep groups.
## Commands
The repo uses `uv` with Python 3.12 (pinned `>=3.12,<3.13`). All commands run from the repo root.
```bash
uv sync # base install (CPU-only; 564 tests pass)
uv sync --extra ml --extra eval --extra data # opt into heavyweight groups
uv sync --all-extras # everything except amd-quark (ships in container)
# Standard local cycle β€” CI runs the same:
uv run ruff check . # lint
uv run mypy mindxtrain/config mindxtrain/provenance # mypy --strict (only these two)
uv run pytest -q # β†’ 564 passed
uv run pytest tests/test_config_schema.py -q # single test file
uv run pytest tests/test_config_schema.py::test_xgmi_2gpu_rejected # single test
# CLI entry point (typer; 9 verbs):
uv run mindxtrain --help
uv run mindxtrain init --list # list 12 built-in YAML recipes
uv run mindxtrain init --template qwen3_8b_sft_lora --out run.yaml
uv run mindxtrain bench --dry-run --out plan.json # CPU-safe (real probe needs MI300X)
uv run mindxtrain receipt ./out/runs/<name>/manifest.json --config run.yaml
# Operator FastAPI + Coach UI (no GPU required):
uv run uvicorn mindxtrain.operator.app:app --host 0.0.0.0 --port 8080
# β†’ http://localhost:8080/coach/
```
GPU verbs (`bench` without `--dry-run`, `train`, `quantize`, `serve`) require an AMD MI300X with ROCm 7.2.1 inside `rocm/primus:v26.2`. The full operator path is in `docs/HANDOFF.md`.
Solidity contracts live in `contracts/` (Foundry, solc 0.8.26): `forge install && forge test` from inside `contracts/`.
## Architecture (concentric layers)
The codebase is organized so each inner layer is consumed by the next, never the reverse:
1. **CLI** (`mindxtrain/cli/main.py`, typer) β€” `init | bench | train | dataset prep | eval | quantize | serve | publish | receipt`. Never reaches into the training backend; consumes a Pydantic-validated config + `AutotunePlan` and dispatches downward.
2. **Autotune** (`mindxtrain/autotune/`) β€” the differentiator. Emits `AutotunePlan` JSON, AOT-only.
3. **Dataset** (`mindxtrain/data/`) β€” curate β†’ dedupe (MinHash + SemDeDup) β†’ filter β†’ tokenize β†’ pack β†’ synth β†’ verify.
4. **Training** (`mindxtrain/train/`) β€” backend dispatch (`dispatch.py` β†’ axolotl / unsloth / torchtune / primus for MI300X subprocess; `trl_cpu` for CPU and `trl_local` for device-aware consumer-GPU/CPU-fallback, both in-process TRL). Methods: SFT, DPO, ORPO, GRPO, GSPO, RLHF, tool-use, CPT.
5. **Artifact + Integration** (`mindxtrain/{eval,deploy,storage,provenance,operator}`) β€” Quark FP8/MXFP4 β†’ lm-eval-harness β†’ HF Hub push β†’ Lighthouse pin β†’ mindX register β†’ AgenticPlace β†’ BANKON ENS β†’ x402 metering β†’ ERC-8004 attestation.
Key end-to-end flow: `XTrainConfig` (Pydantic) + `AutotunePlan` β†’ `dispatch_training()` β†’ `checkpoint_dir/` β†’ `eval.json` β†’ `quantized/` β†’ `manifest.json` (BLAKE3 of YAML+dataset+ckpt+eval, plus HF/Lighthouse/INFT/ASA pointers) β†’ operator serves on `/v1/chat/completions`. `mindxtrain receipt` re-hashes and verifies the manifest round-trip.
Recipes live as YAML at `mindxtrain/train/recipes/<name>.yaml` and are auto-picked up by `mindxtrain init --list` and validated by `tests/test_config_schema.py::test_all_recipes_validate`.
## Non-negotiable invariants
These are encoded in the schema/recipes; violating them is a deployment bug, not a style issue.
1. **AOT-only.** No `torch.compile(mode="max-autotune")` in production paths. No JIT autotune in vLLM (`VLLM_USE_TRITON_FLASH_ATTN=0` if needed). `autotune.policy: aot_only` is the YAML contract.
2. **`hardware.gpus: Literal[1, 8]` only.** 2/4-GPU MI300X FSDP groups hit asymmetric xGMI bandwidth β€” schema rejects them at parse time. (Tested in `test_config_schema.py::test_xgmi_2gpu_rejected`, `test_distributed.py`.)
3. **Seven MI300X env vars** are defaults in every recipe's `train.env` (autotune plan may override values, never remove keys): `PYTORCH_ROCM_ARCH=gfx942`, `HSA_NO_SCRATCH_RECLAIM=1`, `HIP_FORCE_DEV_KERNARG=1`, `GPU_MAX_HW_QUEUES=1`, `NVTE_CK_USES_BWD_V3=1`, `NVTE_CK_IS_V3_ATOMIC_FP32=1`, `PRIMUS_TURBO_ATTN_V3_ATOMIC_FP32=1`, `NCCL_MIN_NCHANNELS=112`.
4. **`extra: forbid` + `frozen: true`** on every Pydantic model β€” unknown YAML keys raise `ValidationError`; loaded configs are immutable.
5. **Solidity contracts are write-once.** No proxies, no `Ownable`, no admin keys, no setters in `contracts/src/{mindxtrain_registry,x402_receiver}.sol`. Rotating any parameter requires a fresh deploy.
6. **Numpy pinned `<2.0`** against `torch==2.9.1+rocm7.2.1.lw`.
7. **Container is `rocm/primus:v26.2`**; SHA256 digest snapshot in `ops/containerfiles/digest.lock`.
## Lazy-import pattern (mandatory for optional deps)
Optional dep groups: `ml` (trl, transformers, peft, accelerate, datasets), `eval` (lm-eval, lighteval, inspect-ai, jinja2), `data` (datasketch, sentence-transformers, faiss-cpu, pyarrow), `serve` (vllm), `chain` (web3, py-algorand-sdk, huggingface-hub), `obs` (opentelemetry-sdk, prometheus-client, psutil).
Every module that wants an optional dep guards the import inside the function that needs it. `import mindxtrain.eval.harness` must always succeed even without `--extra eval`. Error messages must include the exact `uv sync --extra <group>` to run. New modules taking optional deps must follow this pattern.
## Clean-room policy (non-negotiable)
mindXtrain is a **clean-room** codebase: functionality that originates in another
project (mindX, external repos, reference implementations) is **reimplemented or
adapted locally from observed behavior or a spec β€” never copied byte-for-byte**. The
in-tree code is owned by mindXtrain and untainted by foreign source.
When you need something from another codebase:
1. **Load it at runtime** via a documented env var / file path (e.g. the Codephreak
persona via `MINDXTRAIN_PERSONA_PATH`), or
2. **Reimplement the behavior locally** in mindXtrain style, citing the source as a
*reference*, not pasting it.
This applies to the whole product, including the Coach: **mindXtrain and Coach train
models**, and the training/dataset/persona machinery is mindXtrain-native β€” it consumes
mindX artifacts (dream corpus, persona) through boundaries, it does not vendor mindX code.
## Reuse boundaries
- **From `/home/hacker/mindX/`** (production codebase): Codephreak persona JSON loaded at runtime via `MINDXTRAIN_PERSONA_PATH`. Do not copy file bytes β€” load via env var (clean-room).
- **Not** from `/home/hacker/aglm/` β€” broken per its own README.
## Adding things
- **New recipe** β†’ drop YAML at `mindxtrain/train/recipes/<name>.yaml`; `test_all_recipes_validate` picks it up.
- **New training backend** β†’ add `mindxtrain/train/backend_<name>.py` exposing `run_<name>(cfg, plan, out_dir) -> Path`; wire into `train/dispatch.py`; add to `TrainingBackend` literal in `config/schema.py`.
- **New operator backend** β†’ subclass `Backend` in `mindxtrain/operator/backends/<name>.py` decorated `@register_backend("<name>")`; side-effect import from `models/registry.py`; add its base-URL branch to `operator/app.py::backend_kwargs` (shared by the operator route and the Coach chat stream) and, if it can be probed, to `backend_reachable` / `backend_first_model`. Registered backends: `openai_compat`, `ollama`, `vllm`, `bankml` ([docs/bankml.md](docs/bankml.md) β€” the verified CPU engine, reached over HTTP/CLI only; receipts on `ChatResponse.receipt`, HTTP 400 β†’ `BankmlRefusal`, never retried).
- **New training method** β†’ add `_MethodBase` subclass in `config/schema.py` with `kind: Literal["<name>"]`; extend `TrainMethod` discriminated union; add `train/<name>.py` runner; update dispatch; add a recipe.
## Documentation hub
| Doc | What it covers |
|-----|----------------|
| `docs/NAV.md` | **Docs index** β€” start here; one line per doc, grouped. |
| `docs/HANDOFF.md` | 11-step operator checklist (local β†’ MI300X droplet β†’ submission). |
| `docs/architecture.md` | 5-layer architecture + MI300X invariants + data flow. |
| `docs/development.md` | Toolchain, optional-deps, lazy-import pattern, debugging table. |
| `docs/actualization_status.md` | Per-module map of what's real vs. requires extras. |
| `docs/autotune.md` | The 60-second AOT probe β€” the differentiator. |
| `docs/cli.md` | Every verb with synopsis, options, exit codes. |
| `docs/yaml_schema.md` | Every field of the 10-section `XTrainConfig`. |
| `docs/coach.md` | Interactive `/coach/` web UI bundled in the operator. |
| `docs/governance.md` | classroom / boardroom (any-N consensus) / dojo (prime-N dispute settlement). |
| `docs/bankml.md` | bankml ([github.com/cryptoAGI/bankml](https://github.com/cryptoAGI/bankml)): `serve --to bankml`, the `bankml` backend, `imprint-bankml` (not comparable with the canonical gate). |
| `docs/blueprints/` | Frozen source design briefs (the spec the project was built against). |