Download CLAUDE.md from PYTHAI/mindXtrain: direct link, hf CLI and curl.
- Browser
- Download file 9.72 kB
-
https://huggingface.co/PYTHAI/mindXtrain/resolve/main/CLAUDE.md
- Command line
-
hf download hf://PYTHAI/mindXtrain/CLAUDE.md
-
curl -L -o CLAUDE.md https://huggingface.co/PYTHAI/mindXtrain/resolve/main/CLAUDE.md
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
What this is
mindxtrain is a single-package Python training framework for fine-tuning open-weight LLMs on AMD MI300X and serving them through an OpenAI-compatible API. The architectural differentiator is a 60-second AOT autotune probe (mindxtrain bench) that fixes attention backend (CK vs Triton), GEMM heuristic, and RCCL config at training start β JIT autotune is forbidden in the production training loop.
The base install is CPU-only and runs the CLI, Coach UI, bench --dry-run, manifest verify, and the operator FastAPI; heavyweight paths gate on opt-in dep groups.
Commands
The repo uses uv with Python 3.12 (pinned >=3.12,<3.13). All commands run from the repo root.
uv sync # base install (CPU-only; 564 tests pass)
uv sync --extra ml --extra eval --extra data # opt into heavyweight groups
uv sync --all-extras # everything except amd-quark (ships in container)
# Standard local cycle β CI runs the same:
uv run ruff check . # lint
uv run mypy mindxtrain/config mindxtrain/provenance # mypy --strict (only these two)
uv run pytest -q # β 564 passed
uv run pytest tests/test_config_schema.py -q # single test file
uv run pytest tests/test_config_schema.py::test_xgmi_2gpu_rejected # single test
# CLI entry point (typer; 9 verbs):
uv run mindxtrain --help
uv run mindxtrain init --list # list 12 built-in YAML recipes
uv run mindxtrain init --template qwen3_8b_sft_lora --out run.yaml
uv run mindxtrain bench --dry-run --out plan.json # CPU-safe (real probe needs MI300X)
uv run mindxtrain receipt ./out/runs/<name>/manifest.json --config run.yaml
# Operator FastAPI + Coach UI (no GPU required):
uv run uvicorn mindxtrain.operator.app:app --host 0.0.0.0 --port 8080
# β http://localhost:8080/coach/
GPU verbs (bench without --dry-run, train, quantize, serve) require an AMD MI300X with ROCm 7.2.1 inside rocm/primus:v26.2. The full operator path is in docs/HANDOFF.md.
Solidity contracts live in contracts/ (Foundry, solc 0.8.26): forge install && forge test from inside contracts/.
Architecture (concentric layers)
The codebase is organized so each inner layer is consumed by the next, never the reverse:
- CLI (
mindxtrain/cli/main.py, typer) βinit | bench | train | dataset prep | eval | quantize | serve | publish | receipt. Never reaches into the training backend; consumes a Pydantic-validated config +AutotunePlanand dispatches downward. - Autotune (
mindxtrain/autotune/) β the differentiator. EmitsAutotunePlanJSON, AOT-only. - Dataset (
mindxtrain/data/) β curate β dedupe (MinHash + SemDeDup) β filter β tokenize β pack β synth β verify. - Training (
mindxtrain/train/) β backend dispatch (dispatch.pyβ axolotl / unsloth / torchtune / primus for MI300X subprocess;trl_cpufor CPU andtrl_localfor device-aware consumer-GPU/CPU-fallback, both in-process TRL). Methods: SFT, DPO, ORPO, GRPO, GSPO, RLHF, tool-use, CPT. - Artifact + Integration (
mindxtrain/{eval,deploy,storage,provenance,operator}) β Quark FP8/MXFP4 β lm-eval-harness β HF Hub push β Lighthouse pin β mindX register β AgenticPlace β BANKON ENS β x402 metering β ERC-8004 attestation.
Key end-to-end flow: XTrainConfig (Pydantic) + AutotunePlan β dispatch_training() β checkpoint_dir/ β eval.json β quantized/ β manifest.json (BLAKE3 of YAML+dataset+ckpt+eval, plus HF/Lighthouse/INFT/ASA pointers) β operator serves on /v1/chat/completions. mindxtrain receipt re-hashes and verifies the manifest round-trip.
Recipes live as YAML at mindxtrain/train/recipes/<name>.yaml and are auto-picked up by mindxtrain init --list and validated by tests/test_config_schema.py::test_all_recipes_validate.
Non-negotiable invariants
These are encoded in the schema/recipes; violating them is a deployment bug, not a style issue.
- AOT-only. No
torch.compile(mode="max-autotune")in production paths. No JIT autotune in vLLM (VLLM_USE_TRITON_FLASH_ATTN=0if needed).autotune.policy: aot_onlyis the YAML contract. hardware.gpus: Literal[1, 8]only. 2/4-GPU MI300X FSDP groups hit asymmetric xGMI bandwidth β schema rejects them at parse time. (Tested intest_config_schema.py::test_xgmi_2gpu_rejected,test_distributed.py.)- Seven MI300X env vars are defaults in every recipe's
train.env(autotune plan may override values, never remove keys):PYTORCH_ROCM_ARCH=gfx942,HSA_NO_SCRATCH_RECLAIM=1,HIP_FORCE_DEV_KERNARG=1,GPU_MAX_HW_QUEUES=1,NVTE_CK_USES_BWD_V3=1,NVTE_CK_IS_V3_ATOMIC_FP32=1,PRIMUS_TURBO_ATTN_V3_ATOMIC_FP32=1,NCCL_MIN_NCHANNELS=112. extra: forbid+frozen: trueon every Pydantic model β unknown YAML keys raiseValidationError; loaded configs are immutable.- Solidity contracts are write-once. No proxies, no
Ownable, no admin keys, no setters incontracts/src/{mindxtrain_registry,x402_receiver}.sol. Rotating any parameter requires a fresh deploy. - Numpy pinned
<2.0againsttorch==2.9.1+rocm7.2.1.lw. - Container is
rocm/primus:v26.2; SHA256 digest snapshot inops/containerfiles/digest.lock.
Lazy-import pattern (mandatory for optional deps)
Optional dep groups: ml (trl, transformers, peft, accelerate, datasets), eval (lm-eval, lighteval, inspect-ai, jinja2), data (datasketch, sentence-transformers, faiss-cpu, pyarrow), serve (vllm), chain (web3, py-algorand-sdk, huggingface-hub), obs (opentelemetry-sdk, prometheus-client, psutil).
Every module that wants an optional dep guards the import inside the function that needs it. import mindxtrain.eval.harness must always succeed even without --extra eval. Error messages must include the exact uv sync --extra <group> to run. New modules taking optional deps must follow this pattern.
Clean-room policy (non-negotiable)
mindXtrain is a clean-room codebase: functionality that originates in another project (mindX, external repos, reference implementations) is reimplemented or adapted locally from observed behavior or a spec β never copied byte-for-byte. The in-tree code is owned by mindXtrain and untainted by foreign source.
When you need something from another codebase:
- Load it at runtime via a documented env var / file path (e.g. the Codephreak
persona via
MINDXTRAIN_PERSONA_PATH), or - Reimplement the behavior locally in mindXtrain style, citing the source as a reference, not pasting it.
This applies to the whole product, including the Coach: mindXtrain and Coach train models, and the training/dataset/persona machinery is mindXtrain-native β it consumes mindX artifacts (dream corpus, persona) through boundaries, it does not vendor mindX code.
Reuse boundaries
- From
/home/hacker/mindX/(production codebase): Codephreak persona JSON loaded at runtime viaMINDXTRAIN_PERSONA_PATH. Do not copy file bytes β load via env var (clean-room). - Not from
/home/hacker/aglm/β broken per its own README.
Adding things
- New recipe β drop YAML at
mindxtrain/train/recipes/<name>.yaml;test_all_recipes_validatepicks it up. - New training backend β add
mindxtrain/train/backend_<name>.pyexposingrun_<name>(cfg, plan, out_dir) -> Path; wire intotrain/dispatch.py; add toTrainingBackendliteral inconfig/schema.py. - New operator backend β subclass
Backendinmindxtrain/operator/backends/<name>.pydecorated@register_backend("<name>"); side-effect import frommodels/registry.py; add its base-URL branch tooperator/app.py::backend_kwargs(shared by the operator route and the Coach chat stream) and, if it can be probed, tobackend_reachable/backend_first_model. Registered backends:openai_compat,ollama,vllm,bankml(docs/bankml.md β the verified CPU engine, reached over HTTP/CLI only; receipts onChatResponse.receipt, HTTP 400 βBankmlRefusal, never retried). - New training method β add
_MethodBasesubclass inconfig/schema.pywithkind: Literal["<name>"]; extendTrainMethoddiscriminated union; addtrain/<name>.pyrunner; update dispatch; add a recipe.
Documentation hub
| Doc | What it covers |
|---|---|
docs/NAV.md |
Docs index β start here; one line per doc, grouped. |
docs/HANDOFF.md |
11-step operator checklist (local β MI300X droplet β submission). |
docs/architecture.md |
5-layer architecture + MI300X invariants + data flow. |
docs/development.md |
Toolchain, optional-deps, lazy-import pattern, debugging table. |
docs/actualization_status.md |
Per-module map of what's real vs. requires extras. |
docs/autotune.md |
The 60-second AOT probe β the differentiator. |
docs/cli.md |
Every verb with synopsis, options, exit codes. |
docs/yaml_schema.md |
Every field of the 10-section XTrainConfig. |
docs/coach.md |
Interactive /coach/ web UI bundled in the operator. |
docs/governance.md |
classroom / boardroom (any-N consensus) / dojo (prime-N dispute settlement). |
docs/bankml.md |
bankml (github.com/cryptoAGI/bankml): serve --to bankml, the bankml backend, imprint-bankml (not comparable with the canonical gate). |
docs/blueprints/ |
Frozen source design briefs (the spec the project was built against). |