mindXtrain / CLAUDE.md
Gregory-L's picture
bankml: the verified CPU engine as backend, serve target and imprint probe (#1)
730c5bb
|
Raw History Blame Contribute Delete
9.72 kB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

What this is

mindxtrain is a single-package Python training framework for fine-tuning open-weight LLMs on AMD MI300X and serving them through an OpenAI-compatible API. The architectural differentiator is a 60-second AOT autotune probe (mindxtrain bench) that fixes attention backend (CK vs Triton), GEMM heuristic, and RCCL config at training start β€” JIT autotune is forbidden in the production training loop.

The base install is CPU-only and runs the CLI, Coach UI, bench --dry-run, manifest verify, and the operator FastAPI; heavyweight paths gate on opt-in dep groups.

Commands

The repo uses uv with Python 3.12 (pinned >=3.12,<3.13). All commands run from the repo root.

uv sync                         # base install (CPU-only; 564 tests pass)
uv sync --extra ml --extra eval --extra data    # opt into heavyweight groups
uv sync --all-extras            # everything except amd-quark (ships in container)

# Standard local cycle β€” CI runs the same:
uv run ruff check .                                       # lint
uv run mypy mindxtrain/config mindxtrain/provenance       # mypy --strict (only these two)
uv run pytest -q                                          # β†’ 564 passed
uv run pytest tests/test_config_schema.py -q              # single test file
uv run pytest tests/test_config_schema.py::test_xgmi_2gpu_rejected   # single test

# CLI entry point (typer; 9 verbs):
uv run mindxtrain --help
uv run mindxtrain init --list                             # list 12 built-in YAML recipes
uv run mindxtrain init --template qwen3_8b_sft_lora --out run.yaml
uv run mindxtrain bench --dry-run --out plan.json         # CPU-safe (real probe needs MI300X)
uv run mindxtrain receipt ./out/runs/<name>/manifest.json --config run.yaml

# Operator FastAPI + Coach UI (no GPU required):
uv run uvicorn mindxtrain.operator.app:app --host 0.0.0.0 --port 8080
# β†’ http://localhost:8080/coach/

GPU verbs (bench without --dry-run, train, quantize, serve) require an AMD MI300X with ROCm 7.2.1 inside rocm/primus:v26.2. The full operator path is in docs/HANDOFF.md.

Solidity contracts live in contracts/ (Foundry, solc 0.8.26): forge install && forge test from inside contracts/.

Architecture (concentric layers)

The codebase is organized so each inner layer is consumed by the next, never the reverse:

  1. CLI (mindxtrain/cli/main.py, typer) β€” init | bench | train | dataset prep | eval | quantize | serve | publish | receipt. Never reaches into the training backend; consumes a Pydantic-validated config + AutotunePlan and dispatches downward.
  2. Autotune (mindxtrain/autotune/) β€” the differentiator. Emits AutotunePlan JSON, AOT-only.
  3. Dataset (mindxtrain/data/) β€” curate β†’ dedupe (MinHash + SemDeDup) β†’ filter β†’ tokenize β†’ pack β†’ synth β†’ verify.
  4. Training (mindxtrain/train/) β€” backend dispatch (dispatch.py β†’ axolotl / unsloth / torchtune / primus for MI300X subprocess; trl_cpu for CPU and trl_local for device-aware consumer-GPU/CPU-fallback, both in-process TRL). Methods: SFT, DPO, ORPO, GRPO, GSPO, RLHF, tool-use, CPT.
  5. Artifact + Integration (mindxtrain/{eval,deploy,storage,provenance,operator}) β€” Quark FP8/MXFP4 β†’ lm-eval-harness β†’ HF Hub push β†’ Lighthouse pin β†’ mindX register β†’ AgenticPlace β†’ BANKON ENS β†’ x402 metering β†’ ERC-8004 attestation.

Key end-to-end flow: XTrainConfig (Pydantic) + AutotunePlan β†’ dispatch_training() β†’ checkpoint_dir/ β†’ eval.json β†’ quantized/ β†’ manifest.json (BLAKE3 of YAML+dataset+ckpt+eval, plus HF/Lighthouse/INFT/ASA pointers) β†’ operator serves on /v1/chat/completions. mindxtrain receipt re-hashes and verifies the manifest round-trip.

Recipes live as YAML at mindxtrain/train/recipes/<name>.yaml and are auto-picked up by mindxtrain init --list and validated by tests/test_config_schema.py::test_all_recipes_validate.

Non-negotiable invariants

These are encoded in the schema/recipes; violating them is a deployment bug, not a style issue.

  1. AOT-only. No torch.compile(mode="max-autotune") in production paths. No JIT autotune in vLLM (VLLM_USE_TRITON_FLASH_ATTN=0 if needed). autotune.policy: aot_only is the YAML contract.
  2. hardware.gpus: Literal[1, 8] only. 2/4-GPU MI300X FSDP groups hit asymmetric xGMI bandwidth β€” schema rejects them at parse time. (Tested in test_config_schema.py::test_xgmi_2gpu_rejected, test_distributed.py.)
  3. Seven MI300X env vars are defaults in every recipe's train.env (autotune plan may override values, never remove keys): PYTORCH_ROCM_ARCH=gfx942, HSA_NO_SCRATCH_RECLAIM=1, HIP_FORCE_DEV_KERNARG=1, GPU_MAX_HW_QUEUES=1, NVTE_CK_USES_BWD_V3=1, NVTE_CK_IS_V3_ATOMIC_FP32=1, PRIMUS_TURBO_ATTN_V3_ATOMIC_FP32=1, NCCL_MIN_NCHANNELS=112.
  4. extra: forbid + frozen: true on every Pydantic model β€” unknown YAML keys raise ValidationError; loaded configs are immutable.
  5. Solidity contracts are write-once. No proxies, no Ownable, no admin keys, no setters in contracts/src/{mindxtrain_registry,x402_receiver}.sol. Rotating any parameter requires a fresh deploy.
  6. Numpy pinned <2.0 against torch==2.9.1+rocm7.2.1.lw.
  7. Container is rocm/primus:v26.2; SHA256 digest snapshot in ops/containerfiles/digest.lock.

Lazy-import pattern (mandatory for optional deps)

Optional dep groups: ml (trl, transformers, peft, accelerate, datasets), eval (lm-eval, lighteval, inspect-ai, jinja2), data (datasketch, sentence-transformers, faiss-cpu, pyarrow), serve (vllm), chain (web3, py-algorand-sdk, huggingface-hub), obs (opentelemetry-sdk, prometheus-client, psutil).

Every module that wants an optional dep guards the import inside the function that needs it. import mindxtrain.eval.harness must always succeed even without --extra eval. Error messages must include the exact uv sync --extra <group> to run. New modules taking optional deps must follow this pattern.

Clean-room policy (non-negotiable)

mindXtrain is a clean-room codebase: functionality that originates in another project (mindX, external repos, reference implementations) is reimplemented or adapted locally from observed behavior or a spec β€” never copied byte-for-byte. The in-tree code is owned by mindXtrain and untainted by foreign source.

When you need something from another codebase:

  1. Load it at runtime via a documented env var / file path (e.g. the Codephreak persona via MINDXTRAIN_PERSONA_PATH), or
  2. Reimplement the behavior locally in mindXtrain style, citing the source as a reference, not pasting it.

This applies to the whole product, including the Coach: mindXtrain and Coach train models, and the training/dataset/persona machinery is mindXtrain-native β€” it consumes mindX artifacts (dream corpus, persona) through boundaries, it does not vendor mindX code.

Reuse boundaries

  • From /home/hacker/mindX/ (production codebase): Codephreak persona JSON loaded at runtime via MINDXTRAIN_PERSONA_PATH. Do not copy file bytes β€” load via env var (clean-room).
  • Not from /home/hacker/aglm/ β€” broken per its own README.

Adding things

  • New recipe β†’ drop YAML at mindxtrain/train/recipes/<name>.yaml; test_all_recipes_validate picks it up.
  • New training backend β†’ add mindxtrain/train/backend_<name>.py exposing run_<name>(cfg, plan, out_dir) -> Path; wire into train/dispatch.py; add to TrainingBackend literal in config/schema.py.
  • New operator backend β†’ subclass Backend in mindxtrain/operator/backends/<name>.py decorated @register_backend("<name>"); side-effect import from models/registry.py; add its base-URL branch to operator/app.py::backend_kwargs (shared by the operator route and the Coach chat stream) and, if it can be probed, to backend_reachable / backend_first_model. Registered backends: openai_compat, ollama, vllm, bankml (docs/bankml.md β€” the verified CPU engine, reached over HTTP/CLI only; receipts on ChatResponse.receipt, HTTP 400 β†’ BankmlRefusal, never retried).
  • New training method β†’ add _MethodBase subclass in config/schema.py with kind: Literal["<name>"]; extend TrainMethod discriminated union; add train/<name>.py runner; update dispatch; add a recipe.

Documentation hub

Doc What it covers
docs/NAV.md Docs index β€” start here; one line per doc, grouped.
docs/HANDOFF.md 11-step operator checklist (local β†’ MI300X droplet β†’ submission).
docs/architecture.md 5-layer architecture + MI300X invariants + data flow.
docs/development.md Toolchain, optional-deps, lazy-import pattern, debugging table.
docs/actualization_status.md Per-module map of what's real vs. requires extras.
docs/autotune.md The 60-second AOT probe β€” the differentiator.
docs/cli.md Every verb with synopsis, options, exit codes.
docs/yaml_schema.md Every field of the 10-section XTrainConfig.
docs/coach.md Interactive /coach/ web UI bundled in the operator.
docs/governance.md classroom / boardroom (any-N consensus) / dojo (prime-N dispute settlement).
docs/bankml.md bankml (github.com/cryptoAGI/bankml): serve --to bankml, the bankml backend, imprint-bankml (not comparable with the canonical gate).
docs/blueprints/ Frozen source design briefs (the spec the project was built against).