|
Download docs/HANDOFF.md from PYTHAI/mindXtrain: direct link, hf CLI and curl.
- Browser
- Download file 10.9 kB
-
https://huggingface.co/PYTHAI/mindXtrain/resolve/refs%2Fpr%2F1/docs/HANDOFF.md
- Command line
-
hf download hf://PYTHAI/mindXtrain@refs/pr/1/docs/HANDOFF.md
-
curl -L -o HANDOFF.md https://huggingface.co/PYTHAI/mindXtrain/resolve/refs%2Fpr%2F1/docs/HANDOFF.md
10.9 kB
| # HANDOFF β what you need to do next | |
| This is the ordered checklist for taking the mindxtrain repo from "code is | |
| done" to "demo is live." Each step is concrete; check it off when finished. | |
| The repo state at handoff: | |
| - Single canonical package at `mindxtrain/` (12 subpackages, ~100 modules). | |
| - All stub `NotImplementedError` paths replaced with real Python (lazy imports | |
| for heavyweight deps). | |
| - 112/112 tests pass on a CPU-only laptop (`uv sync` + `uv run pytest -q`). | |
| - Optional dep groups in `pyproject.toml`: `ml`, `eval`, `data`, `serve`, | |
| `chain`, `obs`. Install only what you need. | |
| - 12 YAML training recipes wired through the CLI. | |
| - Coach UI (`/coach/`) serves all 12 recipes without GPU. | |
| --- | |
| ## 1. Local setup (no GPU; 10 minutes) | |
| ```bash | |
| cd /home/hacker/Desktop/mindXtrain | |
| cp .env.example .env # then edit .env to fill in HF_TOKEN, etc. | |
| uv sync # base install | |
| uv run pytest -q # β 112 passed | |
| uv run mindxtrain --help # all 9 verbs listed | |
| ``` | |
| **What goes in `.env`** (rest of the file is sane defaults): | |
| | Var | Where to get it | | |
| |---|---| | |
| | `HF_TOKEN` | https://huggingface.co/settings/tokens (write scope) | | |
| | `HF_HUB_USERNAME` | your HF handle | | |
| | `LIGHTHOUSE_API_KEY` | https://files.lighthouse.storage/dashboard/apikey | | |
| | `MINDXTRAIN_OPENAI_API_KEY` | optional; only if you want to use openai_compat backend | | |
| > **Optional (on-chain anchors):** `MINDXTRAIN_REGISTRY_ADDR` (ERC-8004 contract), | |
| > `MINDXTRAIN_FACILITATOR_URL` (x402 facilitator). The publish path skips | |
| > these gracefully if unset. | |
| ## 2. Provision the MI300X droplet (sign-up + 30 min) | |
| > **Fast path (Coach UI):** if you've populated `GITHUB_TOKEN`, | |
| > `AMD_DEV_CLOUD_TOKEN`, and `AMD_DEV_CLOUD_SSH_KEY_ID` in `.env`, you can skip | |
| > the manual SSH dance entirely: | |
| > | |
| > 1. `uv run uvicorn mindxtrain.operator.app:app --port 8080` | |
| > 2. Open <http://localhost:8080/coach/>, scroll to step 6 ("Deploy"). | |
| > 3. Click β **Push to GitHub** β β‘ **Provision MI300X droplet**. The droplet | |
| > boots, cloud-init clones the repo from the SHA you just pushed, pulls the | |
| > container, and runs `mindxtrain bench` automatically. All output streams | |
| > live in the browser via SSE. | |
| > | |
| > Equivalent CLI: `mindxtrain github push && mindxtrain droplet provision`. | |
| > | |
| > The manual sequence below is preserved for scripted / CI use and as a | |
| > fallback when the Coach UI isn't available. | |
| ```bash | |
| # Sign up at https://devcloud.amd.com β request a single MI300X. | |
| # Wait for the droplet (typically same-day). | |
| # SSH in: | |
| ssh ubuntu@<droplet-ip> | |
| # Install podman if missing: | |
| sudo apt-get update && sudo apt-get install -y podman podman-compose | |
| # Pull the canonical training container: | |
| podman pull docker.io/rocm/primus:v26.2 | |
| # Snapshot the digest into the repo so others can reproduce: | |
| podman inspect --format '{{index .RepoDigests 0}}' rocm/primus:v26.2 \ | |
| | tee -a ops/containerfiles/digest.lock | |
| # Verify the GPU is visible: | |
| podman run --rm --device=/dev/kfd --device=/dev/dri rocm/primus:v26.2 \ | |
| rocminfo | head -50 | |
| # β should show gfx942, 192 GB HBM3 | |
| ``` | |
| > **Cost watch:** $1.99/hr Γ planned hours. Budget ~$30 for the full demo | |
| > pipeline (~15 GPU-hours). Leave the droplet **stopped** when not actively | |
| > training. | |
| ## 3. Install heavyweight deps inside the container | |
| ```bash | |
| # On the MI300X: | |
| git clone <your-repo-url> /workspace/mindxtrain | |
| cd /workspace/mindxtrain | |
| podman run -it --rm \ | |
| --device=/dev/kfd --device=/dev/dri \ | |
| -v /workspace/mindxtrain:/workspace/mindxtrain \ | |
| -w /workspace/mindxtrain \ | |
| rocm/primus:v26.2 bash | |
| # Inside the container: | |
| pip install -e ".[ml,eval,data,obs]" | |
| # (skip `serve` and `chain` until you need them β they pull large wheels) | |
| ``` | |
| ## 4. Run the autotune probe (real, ~60 s) | |
| ```bash | |
| mindxtrain bench --gpu 0 --out plan.json | |
| cat plan.json | jq '.attention_backend, .gemm_heuristic, .rccl_config' | |
| # β "ck", "hipblaslt_default", "1gpu_noop" | |
| ``` | |
| Snapshot `plan.json` into the repo so the run is reproducible: | |
| ```bash | |
| cp plan.json ops/k8s/plan-mi300x.json | |
| git add ops/k8s/plan-mi300x.json | |
| git commit -m "snapshot autotune plan from mi300x" | |
| ``` | |
| ## 5. Train + eval + quantize (~ 2 hours total for the demo recipe) | |
| ```bash | |
| # Pick a recipe: instella_3b_lora is the AMD-on-AMD demo path (~30 min). | |
| # Or qwen3_8b_sft_lora for the Qwen side prize (~75 min). | |
| mindxtrain init --template instella_3b_lora --out run.yaml | |
| # Optional: edit run.yaml for your project name, dataset, output path. | |
| $EDITOR run.yaml | |
| # Dataset prep (pulls + dedupes + tokenizes + packs): | |
| mindxtrain dataset prep run.yaml --out ./out/dataset | |
| # Training: | |
| mindxtrain train run.yaml --plan plan.json | |
| # β ./out/runs/<run_name>/checkpoint/ | |
| # Evaluation (MMLU subset): | |
| mindxtrain eval run.yaml | |
| # β ./out/runs/<run_name>/eval/lm_eval.json | |
| # Quantize to FP8: | |
| mindxtrain quantize run.yaml | |
| # β ./out/runs/<run_name>/quantized/ | |
| ``` | |
| If `mindxtrain train` fails with `accelerate not found`: you forgot | |
| `pip install -e ".[ml]"` inside the container (step 3). | |
| ## 6. Build the manifest + verify | |
| ```bash | |
| # Generate the provenance manifest by hashing every artifact: | |
| uv run python -c " | |
| from pathlib import Path | |
| from mindxtrain.config.loader import load_config | |
| from mindxtrain.provenance.manifest import emit_receipt, ProvenanceHashes | |
| cfg = load_config('run.yaml') | |
| run = Path('./out/runs') / cfg.meta.run_name | |
| m = emit_receipt( | |
| cfg, | |
| cfg.meta.run_name, | |
| config_yaml_path=Path('run.yaml'), | |
| dataset_manifest_path=run / 'dataset_manifest.json', | |
| checkpoint_dir=run / 'checkpoint', | |
| eval_json_path=run / 'eval/lm_eval.json', | |
| ) | |
| out = run / 'manifest.json' | |
| out.write_text(m.model_dump_json(indent=2)) | |
| print(out) | |
| " | |
| # Verify it round-trips: | |
| mindxtrain receipt ./out/runs/<run_name>/manifest.json --config run.yaml | |
| # β all BLAKE3 fields = true (config, checkpoint, autotune_plan; dataset/eval if present) | |
| ``` | |
| > **Auto-emitted receipts (operator + CPU lane).** Runs launched through the | |
| > operator β Coach UI or `POST /v1/training/jobs` β now write `manifest.json` | |
| > automatically at completion via `provenance.manifest.emit_receipt_for_run`, | |
| > alongside `config.snapshot.yaml` and `autotune_plan.json` in the run dir. The | |
| > receipt **binds the frozen AutotunePlan hash to the checkpoint hash** β this is | |
| > the AOT artifact that makes a run bitwise-verifiable (cf. Verde/RepOps). The | |
| > Coach "Verifiable receipt" card re-checks it live; `mindxtrain receipt` does the | |
| > same from a shell. The manual `emit_receipt` above remains the full GPU path | |
| > (dataset + eval JSON included). On MI300X, also snapshot the AOTriton / | |
| > hipBLASLt tuning caches next to `autotune_plan.json` so the compiled artifact β | |
| > not just the plan β is reproducible across machines. | |
| ## 7. Publish (HF Hub + Lighthouse + mindX register) | |
| ```bash | |
| # Push to HF (uses HF_TOKEN; private=False for the demo): | |
| mindxtrain publish run.yaml --manifest ./out/runs/<run_name>/manifest.json | |
| # β updates manifest.json in-place with hf_repo_id + lighthouse_cid | |
| ``` | |
| If `LIGHTHOUSE_API_KEY` is unset, the pin step skips gracefully and the | |
| manifest gets a `cid://stub-β¦` placeholder. | |
| ## 8. Deploy contracts (optional) | |
| The demo can ship without on-chain anchors. Do these once, when ready: | |
| ```bash | |
| cd contracts | |
| forge install | |
| forge test # local Foundry tests pass | |
| forge script script/Deploy.s.sol \ | |
| --rpc-url $MINDXTRAIN_BASE_RPC_URL \ | |
| --private-key $DEPLOYER_KEY \ | |
| --broadcast | |
| # β records contract address; paste into .env as MINDXTRAIN_REGISTRY_ADDR | |
| ``` | |
| Once `MINDXTRAIN_REGISTRY_ADDR` is set, `mindxtrain.provenance.erc8004.broadcast_attestation` | |
| can anchor the manifest BLAKE3 on-chain. | |
| ## 9. Serve the model + wire the production URL | |
| The production URL is `https://mindx.pythai.net` β the Coach UI is at `/coach/` | |
| and the public training-jobs API is at `/v1/training/jobs`. | |
| ```bash | |
| # Inside the rocm/vllm-dev container: | |
| podman-compose -f ops/compose/compose_dev.yaml up -d | |
| # β vLLM-ROCm at :8000, mindxtrain operator FastAPI at :8080 | |
| # Verify locally: | |
| curl http://localhost:8080/coach/api/health | |
| # β {"coach_version":"0.1.0", "recipes_available":>=14, ...} | |
| # Public training-jobs API smoke (bearer auth via MINDXTRAIN_API_KEY): | |
| curl -X POST http://localhost:8080/v1/training/jobs \ | |
| -H "Authorization: Bearer $MINDXTRAIN_API_KEY" \ | |
| -H "Content-Type: application/json" \ | |
| -d '{"recipe":"mindx_fallback_qwen3_1_5b_cpu_smoke"}' | |
| # β {"job_id":"...", "status":"running", "backend":"trl_cpu", ...} | |
| # Reverse-proxy mindx.pythai.net β MI300X:8080 (Caddy/Cloudflare). | |
| ``` | |
| Once the proxy is live, `curl https://mindx.pythai.net/coach/api/health` | |
| returns 200 from the public internet. | |
| ## 10. Publish & demo | |
| ```bash | |
| # Push code: | |
| git push origin main | |
| # End-to-end demo walk-through: | |
| # 1. mindxtrain init β show CLI verbs | |
| # 2. mindxtrain bench β 60-second autotune (the differentiator) | |
| # 3. mindxtrain train β timelapse of training | |
| # 4. mindxtrain quantize β FP8 weights | |
| # 5. curl /v1/chat/completions β live inference | |
| # 6. mindxtrain receipt β BLAKE3 reverify | |
| # 7. Open /coach/ β click through the UI | |
| # 8. Open /coach/dcoach β Imprint & Prove (CPU recall proof) | |
| ``` | |
| ## 11. Quality gates (run before every push) | |
| ```bash | |
| uv run ruff check . | |
| uv run mypy mindxtrain/config mindxtrain/provenance | |
| uv run pytest -q # β 112 passed | |
| ``` | |
| All three must pass before pushing to `main`. CI runs the same gates on the | |
| `main` branch. | |
| --- | |
| ## What's still TODO | |
| These paths are wired but require runtime/contracts/services to actually | |
| flow end-to-end: | |
| - **x402 metering** (`mindxtrain.provenance.x402`) β wired to httpx, needs | |
| a deployed facilitator URL. | |
| - **ERC-8004 broadcast** (`mindxtrain.provenance.erc8004.broadcast_attestation`) | |
| β needs deployed attestation registry + signer key. | |
| - **BANKON ENS** allocation (`mindxtrain.provenance.algorand.allocate_ens_subname`) | |
| β needs the BANKON allocation service deployed. | |
| - **AgenticPlace listing** (`mindxtrain.deploy.api_client.list_on_agenticplace`) | |
| β needs `agenticplace.pythai.net` live. | |
| - **mindX agent register** (`mindxtrain.deploy.api_client.register_with_mindx`) | |
| β needs `mindx.pythai.net/v1/agents` live. | |
| The framework itself ships as production-ready Apache-2.0; the integrations | |
| above are paid/external services you stand up at your own pace. | |
| --- | |
| ## Quick reference | |
| | What | Where | | |
| |---|---| | |
| | All CLI verbs | `mindxtrain --help` | | |
| | All recipes | `mindxtrain init --list` | | |
| | Coach UI | http://localhost:8080/coach/ | | |
| | Per-module status | `docs/actualization_status.md` | | |
| | Architecture | `docs/architecture.md` | | |
| | Autotune detail | `docs/autotune.md` | | |
| | Coach detail | `docs/coach.md` | | |
| | CLI reference | `docs/cli.md` | | |
| | YAML schema | `docs/yaml_schema.md` | | |
| | dcoach proof loop | `docs/dcoach.md` | | |
| | Frozen blueprints | `docs/blueprints/{mindXtrain,mindXtrain2}.md` | | |