File size: 7,048 Bytes
dfb775d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fe0e4a2
 
dfb775d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
---
license: apache-2.0
library_name: mindxtrain
pipeline_tag: text-generation
tags:
- mindx
- mindxtrain
- training-framework
- lora
- cpu-training
- proof-of-recall
datasets:
- PYTHAI/mindXascension
- PYTHAI/mindX-docs
---

> **mindXtrain for mindX — the Hugging Face fork.** This repository is the **mindX-specific** line of mindXtrain,
> forked on 2026-09-14 from the agnostic upstream
> [github.com/Professor-Codephreak/mindXtrain](https://github.com/Professor-Codephreak/mindXtrain) at commit
> [`661bd41`](https://github.com/Professor-Codephreak/mindXtrain/commit/661bd411738d11e633b25c681bbd5676556bab7d) (provenance in [`FORK.json`](FORK.json)).
> **mindXtrain continues here.** The GitHub repository is **archived for posterity** — the read-only record of the
> pioneering work of [Professor Codephreak](https://github.com/Professor-Codephreak). All new work lands on the Hub:
>
> ```bash
> git clone https://huggingface.co/PYTHAI/mindXtrain
> ```
>
> What this line trains for: the mindX lineage ([`PYTHAI/mindXascension`](https://huggingface.co/datasets/PYTHAI/mindXascension)),
> built from mindX's doctrine ([`PYTHAI/mindX-docs`](https://huggingface.co/datasets/PYTHAI/mindX-docs), with
> [the mapping](https://huggingface.co/datasets/PYTHAI/mindX-docs/blob/main/MAPPING.md)); the last accepted generation is
> [`PYTHAI/mindXtrain39`](https://huggingface.co/PYTHAI/mindXtrain39). The Hub footprint is mapped in
> [`examples/mindx/HUGGINGFACE_MAP.md`](examples/mindx/HUGGINGFACE_MAP.md). The upstream README follows unchanged.

# mindxtrain

Production training framework for fine-tuning open-weight LLMs on AMD MI300X
and serving them through an OpenAI-compatible API. Single ordered package,
canonical layout per [`docs/blueprints/mindXtrain2.md`](docs/blueprints/mindXtrain2.md)
§Part 4.

The single architectural feature that distinguishes mindxtrain from Axolotl,
LLaMA-Factory, Unsloth, torchtune and Primus is its **60-second AOT autotune
probe**: CK-vs-Triton attention, hipBLASLt heuristic, RCCL config — the plan
is fixed at training start, JIT autotune is forbidden in the production loop.

**Status**: production deployment in progress. The CPU-only base install passes
its full pytest suite (ruff + mypy clean); with the training extras installed the
suite is 672 green. Many modules ship as real Python on a CPU-only laptop;
heavyweight training, eval, and quantization paths gate on opt-in extra dep
groups. See [`docs/actualization_status.md`](docs/actualization_status.md) for the
per-module map and [`HANDOFF.md`](docs/HANDOFF.md) for the operator checklist.

## Where this runs

- **Operator + Coach UI:** [https://mindx.pythai.net/coach](https://mindx.pythai.net/coach)
- **Public training-jobs API:** `https://mindx.pythai.net/v1/training/jobs`
  (bearer auth via `MINDXTRAIN_API_KEY`)
- **mindX self-training loop:** mindX's dream cycle writes JSONL training
  data; this framework consumes it via the `mindx_dreams` data source and
  fine-tunes a small fallback model on a single MI300X.

## Prove it trains

mindXtrain doesn't just assert that training works — it proves recall. The
[**dcoach**](docs/dcoach.md) proof loop (`/coach/dcoach`) imprints a persona onto a
tiny model on CPU, then measures whether the model *recalls* it: the **classroom**
scores recall before vs after training, the **boardroom** rules success or failure,
and the verdict feeds an **autotune feedback loop** that tunes the next run. A clean
CPU run reports a positive imprint Δ (e.g. recall 0.07 → 0.28) and an approved
verdict. [`docs/NAV.md`](docs/NAV.md) is the full documentation hub.

## Quickstart

```bash
uv sync                                                    # base install
uv run pytest -q                                           # → 564 passed
uv run mindxtrain --help                                   # 9 verbs
uv run mindxtrain init --template qwen3_8b_sft_lora --out run.yaml
uv run mindxtrain bench --dry-run --out plan.json
uv run uvicorn mindxtrain.operator.app:app --host 0.0.0.0 --port 8080
# open http://localhost:8080/coach/  for the interactive UI
```

To unlock training / eval / quantize / publish, install the matching dep group:

```bash
uv sync --extra ml --extra eval --extra data         # train + eval + curate
# or
uv sync --all-extras                                  # everything except amd-quark
```

GPU steps (`bench` without `--dry-run`, `train`, `quantize`, `serve`) require
an AMD MI300X with ROCm 7.2.1; run inside `rocm/primus:v26.2`. The full
operator checklist lives in [`HANDOFF.md`](docs/HANDOFF.md).

## Layout

```
mindxtrain/{cli,config,data,models,train,eval,autotune,
            operator,storage,provenance,deploy,budget}/   # 99 modules
contracts/        Foundry workspace for ERC-8004 attestation registry
ops/              containerfiles, compose, k8s, vmm, gensyn
tests/            pytest suite — 566 tests, CPU-only smoke
examples/         demo YAML configs
docs/             user-facing documentation + frozen blueprints
scripts/          dev helpers
```

## Documentation

| Doc | What it covers |
|-----|----------------|
| [`HANDOFF.md`](docs/HANDOFF.md) | **Operator checklist** — ordered steps from local setup to live deployment. |
| [`docs/quickstart.md`](docs/quickstart.md) | Install + base-vs-extras command tour. |
| [`docs/architecture.md`](docs/architecture.md) | Canonical layout + 5-layer architecture + MI300X invariants. |
| [`docs/actualization_status.md`](docs/actualization_status.md) | Per-module map of what's real vs. requires extras. |
| [`docs/autotune.md`](docs/autotune.md) | The 60-second AOT probe — the architectural differentiator. |
| [`docs/coach.md`](docs/coach.md) | Interactive `/coach/` web UI bundled in the operator. |
| [`docs/dcoach.md`](docs/dcoach.md) | The dcoach proof loop — prove a CPU model recalls its training; decentralized-training fit. |
| [`docs/cli.md`](docs/cli.md) | Every `mindxtrain` verb with synopsis, options, exit codes. |
| [`docs/yaml_schema.md`](docs/yaml_schema.md) | Every field of the 10-section `XTrainConfig`. |
| [`docs/benchmarks.md`](docs/benchmarks.md) | Target metrics + the 7-cell framework comparison. |
| [`docs/development.md`](docs/development.md) | Toolchain, optional-deps, lazy-import pattern, invariants. |
| [`docs/blueprints/`](docs/blueprints/) | Source design briefs (frozen specification). |
| [`llm.txt`](llm.txt) | Orientation for another model — what is measured, what is not, the traps. |
| [`examples/mindx/`](examples/mindx/HUGGINGFACE_MAP.md) | Example consumer — mindX on the Hugging Face Hub: its lineage, [docs dataset + mapping](https://huggingface.co/datasets/PYTHAI/mindX-docs/blob/main/MAPPING.md), Spaces and licence-pinned base models. The framework stays agnostic. |

## License

Apache-2.0. See [LICENSE](LICENSE), [NOTICE](NOTICE), and the upstream-license
notices in [`LICENSE-MIT-upstream-glm51`](LICENSE-MIT-upstream-glm51) and
[`LICENSE-NOTICE.md`](docs/LICENSE-NOTICE.md). Version history in
[`CHANGELOG.md`](docs/CHANGELOG.md).