v1.0.2 — the Hub is reachable from the command line, and is not a chain
Browse filesThe hf/ extension existed and was complete, but nothing outside a Python import could reach it. New `mindxtrain hf` group: whoami / pull / warm / publish / lineage / dataset / space, each exiting nonzero when the result says ok:false so the ascent loop can warm a base and check $?.
huggingface-hub moves out of the `chain` extra into its own `hf` extra — reaching the Hub no longer means installing web3 and the Algorand SDK. `chain` pulls `mindxtrain[hf]` so existing installs keep working.
push_space dropped its commit receipt; it reports it now. First tests for the module.
765 tests pass (was 750).
- docs/CHANGELOG.md +32 -0
- docs/NAV.md +2 -1
- mindxtrain/__init__.py +1 -1
- mindxtrain/cli/main.py +120 -0
- mindxtrain/hf/__init__.py +2 -2
- mindxtrain/hf/extension.py +35 -34
- mindxtrain/storage/hf_hub.py +3 -3
- mindxtrain/ui/app.py +1 -1
- pyproject.toml +7 -3
- tests/test_hf_extension.py +159 -0
docs/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,38 @@ project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
| 6 |
|
| 7 |
## [Unreleased]
|
| 8 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
## [1.0.1] — 2026-09-15
|
| 10 |
|
| 11 |
|
|
|
|
| 6 |
|
| 7 |
## [Unreleased]
|
| 8 |
|
| 9 |
+
## [1.0.2] — 2026-09-15
|
| 10 |
+
|
| 11 |
+
### Added
|
| 12 |
+
|
| 13 |
+
- **mindXtrain can use Hugging Face from the command line.** The `mindxtrain/hf/` extension
|
| 14 |
+
existed and was complete — `account`, `pull_base`, `warm`, `publish_generation`,
|
| 15 |
+
`push_dataset`, `lineage`, `push_space` — but nothing outside a Python import could reach
|
| 16 |
+
it: there was no CLI verb, so on a node the Hub was effectively unavailable. New
|
| 17 |
+
`mindxtrain hf` group: `whoami · pull · warm · publish · lineage · dataset · space`.
|
| 18 |
+
Each prints the extension's own result dict and **exits nonzero when it reports
|
| 19 |
+
`ok: false`**, so the ascent loop can warm a base, check `$?`, and refuse to start a
|
| 20 |
+
three-hour run that would otherwise discover the missing base at the end of it.
|
| 21 |
+
- **Tests for the Hub extension**, which had none: token precedence across the three
|
| 22 |
+
accepted environment variables, the missing-dependency path (genuinely missing here, not
|
| 23 |
+
mocked), the `RepoFolder` trap that makes a scan report an empty repo, the guards that
|
| 24 |
+
must not need a network, and the CLI wiring including the nonzero exit.
|
| 25 |
+
|
| 26 |
+
### Changed
|
| 27 |
+
|
| 28 |
+
- **`huggingface-hub` moved out of the `chain` extra into its own `hf` extra.** Reaching the
|
| 29 |
+
Hub used to mean installing `web3` and the Algorand SDK, because the Hub was filed under
|
| 30 |
+
"chain". It is not a chain. `chain` now pulls `mindxtrain[hf]` so existing installs keep
|
| 31 |
+
working, `all` includes it, and every "not installed" message names `--extra hf` — the
|
| 32 |
+
extra that actually installs it — instead of sending the reader to the blockchain one.
|
| 33 |
+
- The `hf` module is ruff clean under the house config.
|
| 34 |
+
|
| 35 |
+
### Fixed
|
| 36 |
+
|
| 37 |
+
- **`push_space` captured the upload's commit receipt and discarded it**, while its two
|
| 38 |
+
sibling publishers both return it. A Space push now reports its commit like everything
|
| 39 |
+
else, so what was pushed is identifiable afterwards.
|
| 40 |
+
|
| 41 |
## [1.0.1] — 2026-09-15
|
| 42 |
|
| 43 |
|
docs/NAV.md
CHANGED
|
@@ -164,8 +164,9 @@ Target metrics + framework comparison.
|
|
| 164 |
- [What's not measured (yet)](benchmarks.md#whats-not-measured-yet)
|
| 165 |
|
| 166 |
### [CHANGELOG](CHANGELOG.md)
|
| 167 |
-
Version history (current: v1.0.
|
| 168 |
- [[Unreleased]](CHANGELOG.md#unreleased)
|
|
|
|
| 169 |
- [[1.0.1] — 2026-09-15](CHANGELOG.md#101--2026-09-15)
|
| 170 |
- [[1.0.0] — 2026-06-11](CHANGELOG.md#100--2026-06-11)
|
| 171 |
- [[0.1.0] — 2026-05-06](CHANGELOG.md#010--2026-05-06)
|
|
|
|
| 164 |
- [What's not measured (yet)](benchmarks.md#whats-not-measured-yet)
|
| 165 |
|
| 166 |
### [CHANGELOG](CHANGELOG.md)
|
| 167 |
+
Version history (current: v1.0.2).
|
| 168 |
- [[Unreleased]](CHANGELOG.md#unreleased)
|
| 169 |
+
- [[1.0.2] — 2026-09-15](CHANGELOG.md#102--2026-09-15)
|
| 170 |
- [[1.0.1] — 2026-09-15](CHANGELOG.md#101--2026-09-15)
|
| 171 |
- [[1.0.0] — 2026-06-11](CHANGELOG.md#100--2026-06-11)
|
| 172 |
- [[0.1.0] — 2026-05-06](CHANGELOG.md#010--2026-05-06)
|
mindxtrain/__init__.py
CHANGED
|
@@ -16,4 +16,4 @@ Layout per mindXtrain2.md §Part 4 "Repository layout":
|
|
| 16 |
budget/ psutil-derived ResourceBudget (carry-over from aGLM)
|
| 17 |
"""
|
| 18 |
|
| 19 |
-
__version__ = "1.0.
|
|
|
|
| 16 |
budget/ psutil-derived ResourceBudget (carry-over from aGLM)
|
| 17 |
"""
|
| 18 |
|
| 19 |
+
__version__ = "1.0.2"
|
mindxtrain/cli/main.py
CHANGED
|
@@ -25,10 +25,16 @@ mei_app = typer.Typer(
|
|
| 25 |
help="mindX Efficiency Index — score, history, promotion checks.",
|
| 26 |
no_args_is_help=True,
|
| 27 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
app.add_typer(dataset_app)
|
| 29 |
app.add_typer(github_app)
|
| 30 |
app.add_typer(droplet_app)
|
| 31 |
app.add_typer(mei_app)
|
|
|
|
| 32 |
console = Console()
|
| 33 |
|
| 34 |
|
|
@@ -710,6 +716,120 @@ def research(
|
|
| 710 |
console.print(f"[green]{result.summary()}[/green]")
|
| 711 |
|
| 712 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 713 |
# ---- github / droplet (source-tree publishing + remote provision) -------
|
| 714 |
|
| 715 |
|
|
|
|
| 25 |
help="mindX Efficiency Index — score, history, promotion checks.",
|
| 26 |
no_args_is_help=True,
|
| 27 |
)
|
| 28 |
+
hf_app = typer.Typer(
|
| 29 |
+
name="hf",
|
| 30 |
+
help="Hugging Face — who the token is, warm a base, publish a generation, read the lineage.",
|
| 31 |
+
no_args_is_help=True,
|
| 32 |
+
)
|
| 33 |
app.add_typer(dataset_app)
|
| 34 |
app.add_typer(github_app)
|
| 35 |
app.add_typer(droplet_app)
|
| 36 |
app.add_typer(mei_app)
|
| 37 |
+
app.add_typer(hf_app)
|
| 38 |
console = Console()
|
| 39 |
|
| 40 |
|
|
|
|
| 716 |
console.print(f"[green]{result.summary()}[/green]")
|
| 717 |
|
| 718 |
|
| 719 |
+
|
| 720 |
+
# ---- hugging face (the Hub as mindXtrain uses it) -----------------------------
|
| 721 |
+
|
| 722 |
+
|
| 723 |
+
def _hf_report(result: dict, *, quiet: bool = False) -> None:
|
| 724 |
+
"""Print a result dict and exit nonzero when it says `ok: false`.
|
| 725 |
+
|
| 726 |
+
Every `mindxtrain.hf` function returns `{"ok": bool, ...}` rather than raising, so the exit
|
| 727 |
+
code is what makes it scriptable: the ascent loop can warm a base, check `$?`, and refuse to
|
| 728 |
+
start a three-hour run that would only discover the missing base at the end.
|
| 729 |
+
"""
|
| 730 |
+
import json as _json
|
| 731 |
+
|
| 732 |
+
if not quiet:
|
| 733 |
+
console.print_json(_json.dumps(result, default=str))
|
| 734 |
+
if not result.get("ok"):
|
| 735 |
+
raise typer.Exit(code=1)
|
| 736 |
+
|
| 737 |
+
|
| 738 |
+
@hf_app.command("whoami")
|
| 739 |
+
def hf_whoami(
|
| 740 |
+
token: str = typer.Option("", "--token", help="Defaults to HF_TOKEN in the environment."),
|
| 741 |
+
) -> None:
|
| 742 |
+
"""Who the token is — and, separately, the namespaces it can actually WRITE.
|
| 743 |
+
|
| 744 |
+
Org membership is not write scope; the two are reported apart because assuming they are the
|
| 745 |
+
same is how a publish fails after the training finished.
|
| 746 |
+
"""
|
| 747 |
+
from mindxtrain.hf import account
|
| 748 |
+
|
| 749 |
+
_hf_report(account(token or None))
|
| 750 |
+
|
| 751 |
+
|
| 752 |
+
@hf_app.command("pull")
|
| 753 |
+
def hf_pull(
|
| 754 |
+
model_id: str = typer.Argument(..., help="Base model repo id, e.g. HuggingFaceTB/SmolLM2-135M."),
|
| 755 |
+
token: str = typer.Option("", "--token"),
|
| 756 |
+
allow: list[str] = typer.Option(None, "--allow", help="Glob to restrict the download."),
|
| 757 |
+
) -> None:
|
| 758 |
+
"""Fetch a base model into the local cache before training needs it."""
|
| 759 |
+
from mindxtrain.hf import pull_base
|
| 760 |
+
|
| 761 |
+
_hf_report(pull_base(model_id, token=token or None, allow_patterns=list(allow) if allow else None))
|
| 762 |
+
|
| 763 |
+
|
| 764 |
+
@hf_app.command("warm")
|
| 765 |
+
def hf_warm(
|
| 766 |
+
config: Path = typer.Argument(..., help="run.yaml — its model.name is what gets pulled."),
|
| 767 |
+
token: str = typer.Option("", "--token"),
|
| 768 |
+
) -> None:
|
| 769 |
+
"""Pull whatever a run config says it needs, so `train` starts cold-free."""
|
| 770 |
+
from mindxtrain.hf import warm
|
| 771 |
+
|
| 772 |
+
_hf_report(warm(config, token=token or None))
|
| 773 |
+
|
| 774 |
+
|
| 775 |
+
@hf_app.command("publish")
|
| 776 |
+
def hf_publish(
|
| 777 |
+
run_dir: Path = typer.Argument(..., help="A finished run directory."),
|
| 778 |
+
repo_id: str = typer.Argument(..., help="Target model repo, e.g. PYTHAI/mindXascension."),
|
| 779 |
+
token: str = typer.Option("", "--token"),
|
| 780 |
+
private: bool = typer.Option(False, "--private", help="Create the repo private."),
|
| 781 |
+
no_merged: bool = typer.Option(False, "--no-merged", help="Adapter only; skip merged weights."),
|
| 782 |
+
dry_run: bool = typer.Option(False, "--dry-run", help="Report what would upload; touch nothing."),
|
| 783 |
+
) -> None:
|
| 784 |
+
"""Publish a finished run as a model repo: weights, adapter, train.log, Modelfile, card."""
|
| 785 |
+
from mindxtrain.hf import publish_generation
|
| 786 |
+
|
| 787 |
+
_hf_report(publish_generation(run_dir, repo_id, token=token or None, private=private,
|
| 788 |
+
include_merged=not no_merged, dry_run=dry_run))
|
| 789 |
+
|
| 790 |
+
|
| 791 |
+
@hf_app.command("lineage")
|
| 792 |
+
def hf_lineage(
|
| 793 |
+
repo_id: str = typer.Argument(..., help="Repo to scan."),
|
| 794 |
+
repo_type: str = typer.Option("model", "--repo-type", help="model | dataset | space."),
|
| 795 |
+
local_runs: Path = typer.Option(None, "--local-runs", help="Run root, to list what is unpublished."),
|
| 796 |
+
token: str = typer.Option("", "--token"),
|
| 797 |
+
) -> None:
|
| 798 |
+
"""What is on the Hub for a project, reconciled against local runs."""
|
| 799 |
+
from mindxtrain.hf import lineage
|
| 800 |
+
|
| 801 |
+
_hf_report(lineage(repo_id, repo_type=repo_type, token=token or None, local_runs=local_runs))
|
| 802 |
+
|
| 803 |
+
|
| 804 |
+
@hf_app.command("dataset")
|
| 805 |
+
def hf_dataset(
|
| 806 |
+
path: Path = typer.Argument(..., help="A corpus folder or one JSONL."),
|
| 807 |
+
repo_id: str = typer.Argument(..., help="Target dataset repo."),
|
| 808 |
+
token: str = typer.Option("", "--token"),
|
| 809 |
+
private: bool = typer.Option(False, "--private"),
|
| 810 |
+
path_in_repo: str = typer.Option("", "--path-in-repo"),
|
| 811 |
+
) -> None:
|
| 812 |
+
"""Push a training corpus to a dataset repo."""
|
| 813 |
+
from mindxtrain.hf import push_dataset
|
| 814 |
+
|
| 815 |
+
_hf_report(push_dataset(path, repo_id, token=token or None, private=private,
|
| 816 |
+
path_in_repo=path_in_repo))
|
| 817 |
+
|
| 818 |
+
|
| 819 |
+
@hf_app.command("space")
|
| 820 |
+
def hf_space(
|
| 821 |
+
folder: Path = typer.Argument(..., help="A Gradio folder (must contain app.py)."),
|
| 822 |
+
space_id: str = typer.Argument(..., help="Target Space id."),
|
| 823 |
+
token: str = typer.Option("", "--token"),
|
| 824 |
+
public: bool = typer.Option(False, "--public", help="Public Space. Never park a write token on one."),
|
| 825 |
+
hardware: str = typer.Option("zero-a10g", "--hardware", help="Empty string for the free CPU tier."),
|
| 826 |
+
) -> None:
|
| 827 |
+
"""Push a Gradio folder as a Space."""
|
| 828 |
+
from mindxtrain.hf import push_space
|
| 829 |
+
|
| 830 |
+
_hf_report(push_space(folder, space_id, token=token or None, private=not public,
|
| 831 |
+
hardware=hardware or None))
|
| 832 |
+
|
| 833 |
# ---- github / droplet (source-tree publishing + remote provision) -------
|
| 834 |
|
| 835 |
|
mindxtrain/hf/__init__.py
CHANGED
|
@@ -12,9 +12,9 @@ the rest of the Hub as mindXtrain needs it, and nothing more:
|
|
| 12 |
- `datasets` — push/pull a training corpus with its manifest
|
| 13 |
|
| 14 |
Every function returns a plain dict: `{"ok": bool, …}`. Nothing here raises at import time, and
|
| 15 |
-
`huggingface_hub` is imported lazily so a CPU-only install without `--extra
|
| 16 |
"""
|
| 17 |
-
from .extension import (
|
| 18 |
account,
|
| 19 |
lineage,
|
| 20 |
publish_generation,
|
|
|
|
| 12 |
- `datasets` — push/pull a training corpus with its manifest
|
| 13 |
|
| 14 |
Every function returns a plain dict: `{"ok": bool, …}`. Nothing here raises at import time, and
|
| 15 |
+
`huggingface_hub` is imported lazily so a CPU-only install without `--extra hf` still loads.
|
| 16 |
"""
|
| 17 |
+
from .extension import (
|
| 18 |
account,
|
| 19 |
lineage,
|
| 20 |
publish_generation,
|
mindxtrain/hf/extension.py
CHANGED
|
@@ -20,12 +20,12 @@ import os
|
|
| 20 |
import re
|
| 21 |
import time
|
| 22 |
from pathlib import Path
|
| 23 |
-
from typing import Any
|
| 24 |
|
| 25 |
_HF_ENV = ("HF_TOKEN", "HUGGING_FACE_HUB_TOKEN", "HUGGINGFACEHUB_API_TOKEN")
|
| 26 |
|
| 27 |
|
| 28 |
-
def _token(explicit:
|
| 29 |
if explicit:
|
| 30 |
return explicit
|
| 31 |
for k in _HF_ENV:
|
|
@@ -35,26 +35,26 @@ def _token(explicit: Optional[str] = None) -> Optional[str]:
|
|
| 35 |
return None
|
| 36 |
|
| 37 |
|
| 38 |
-
def _api(token:
|
| 39 |
"""(HfApi, None) or (None, {"ok": False, …}) — never raises."""
|
| 40 |
try:
|
| 41 |
from huggingface_hub import HfApi
|
| 42 |
except ImportError:
|
| 43 |
-
return None, {"ok": False, "reason": "huggingface_hub not installed — `uv sync --extra
|
| 44 |
tok = _token(token)
|
| 45 |
if not tok:
|
| 46 |
return None, {"ok": False, "reason": f"no token — set one of {', '.join(_HF_ENV)}"}
|
| 47 |
return HfApi(token=tok), None
|
| 48 |
|
| 49 |
|
| 50 |
-
def tree_paths(api, repo_id: str, **kw) ->
|
| 51 |
"""File paths of a repo tree (folders skipped). See the module docstring: the tree's entries
|
| 52 |
carry `.path`, not `.rfilename`."""
|
| 53 |
return [f.path for f in api.list_repo_tree(repo_id, **kw) if type(f).__name__ != "RepoFolder"]
|
| 54 |
|
| 55 |
|
| 56 |
# ── who the token is ──────────────────────────────────────────────────────────
|
| 57 |
-
def account(token:
|
| 58 |
"""The identity behind the token, its role, the orgs it belongs to, and — separately — the
|
| 59 |
namespaces it can actually write."""
|
| 60 |
api, err = _api(token)
|
|
@@ -62,7 +62,7 @@ def account(token: Optional[str] = None) -> Dict[str, Any]:
|
|
| 62 |
return err
|
| 63 |
try:
|
| 64 |
me = api.whoami()
|
| 65 |
-
except Exception as e:
|
| 66 |
return {"ok": False, "reason": f"whoami failed: {type(e).__name__}: {str(e)[:200]}"}
|
| 67 |
auth = (me.get("auth") or {}).get("accessToken") or {}
|
| 68 |
perms = auth.get("fineGrained") or {}
|
|
@@ -82,29 +82,29 @@ def account(token: Optional[str] = None) -> Dict[str, Any]:
|
|
| 82 |
|
| 83 |
|
| 84 |
# ── bases, fetched before the run rather than during it ───────────────────────
|
| 85 |
-
def pull_base(model_id: str, *, token:
|
| 86 |
"""Download a base model to the local cache. Do this BEFORE training: a run that discovers a
|
| 87 |
missing base three hours in has wasted three hours."""
|
| 88 |
try:
|
| 89 |
from huggingface_hub import snapshot_download
|
| 90 |
except ImportError:
|
| 91 |
-
return {"ok": False, "reason": "huggingface_hub not installed — `uv sync --extra
|
| 92 |
t0 = time.time()
|
| 93 |
try:
|
| 94 |
path = snapshot_download(model_id, token=_token(token), allow_patterns=allow_patterns)
|
| 95 |
-
except Exception as e:
|
| 96 |
return {"ok": False, "model": model_id, "reason": f"{type(e).__name__}: {str(e)[:200]}"}
|
| 97 |
p = Path(path)
|
| 98 |
return {"ok": True, "model": model_id, "path": str(p), "seconds": round(time.time() - t0, 1),
|
| 99 |
"bytes": sum(f.stat().st_size for f in p.rglob("*") if f.is_file())}
|
| 100 |
|
| 101 |
|
| 102 |
-
def warm(config: Path | str, *, token:
|
| 103 |
"""Pull whatever a run.yaml says it needs (the base model) so `train` starts cold-free."""
|
| 104 |
try:
|
| 105 |
import yaml
|
| 106 |
cfg = yaml.safe_load(Path(config).read_text())
|
| 107 |
-
except Exception as e:
|
| 108 |
return {"ok": False, "reason": f"unreadable config: {type(e).__name__}: {str(e)[:160]}"}
|
| 109 |
base = ((cfg or {}).get("model") or {}).get("name")
|
| 110 |
if not base:
|
|
@@ -113,11 +113,11 @@ def warm(config: Path | str, *, token: Optional[str] = None) -> Dict[str, Any]:
|
|
| 113 |
|
| 114 |
|
| 115 |
# ── a finished run, published ─────────────────────────────────────────────────
|
| 116 |
-
def _card(repo_id: str, meta:
|
| 117 |
"""A model card written from the run's own numbers. No claim that is not in `meta`."""
|
| 118 |
m = meta
|
| 119 |
fm = {"license": m.get("license", "apache-2.0"), "library_name": "transformers",
|
| 120 |
-
"pipeline_tag": "text-generation", "tags": ["mindxtrain", "lora"
|
| 121 |
if m.get("base"):
|
| 122 |
fm["base_model"] = m["base"]
|
| 123 |
if m.get("dataset"):
|
|
@@ -137,9 +137,9 @@ def _card(repo_id: str, meta: Dict[str, Any]) -> str:
|
|
| 137 |
"the training log ships beside the weights.\n")
|
| 138 |
|
| 139 |
|
| 140 |
-
def publish_generation(run_dir: Path | str, repo_id: str, *, token:
|
| 141 |
-
meta:
|
| 142 |
-
include_merged: bool = True, dry_run: bool = False) ->
|
| 143 |
"""A finished run as a model repo: merged weights at the root (if present), the LoRA delta under
|
| 144 |
`adapter/`, `train.log`, a `Modelfile` for Ollama, and a card built from `meta`.
|
| 145 |
|
|
@@ -149,14 +149,14 @@ def publish_generation(run_dir: Path | str, repo_id: str, *, token: Optional[str
|
|
| 149 |
return {"ok": False, "reason": f"no run dir at {run}"}
|
| 150 |
ck = run / "checkpoint"
|
| 151 |
merged = run / "ollama_push" / "merged"
|
| 152 |
-
staged:
|
| 153 |
if include_merged and merged.is_dir():
|
| 154 |
for f in merged.iterdir():
|
| 155 |
if f.is_file():
|
| 156 |
staged[f.name] = f
|
| 157 |
if ck.is_dir():
|
| 158 |
for f in ck.iterdir():
|
| 159 |
-
if f.is_file() and f.name != "training_args.bin" or f.name == "training_args.bin":
|
| 160 |
staged[f"adapter/{f.name}"] = f
|
| 161 |
log = run / "train.log"
|
| 162 |
if log.is_file():
|
|
@@ -167,21 +167,21 @@ def publish_generation(run_dir: Path | str, repo_id: str, *, token: Optional[str
|
|
| 167 |
if not meta.get("base"):
|
| 168 |
try:
|
| 169 |
meta["base"] = json.loads((ck / "adapter_config.json").read_text()).get("base_model_name_or_path")
|
| 170 |
-
except Exception:
|
| 171 |
pass
|
| 172 |
card = _card(repo_id, meta)
|
| 173 |
modelfile = ("# ollama create <name> -f Modelfile (from this repo's directory)\nFROM .\n"
|
| 174 |
+ (f'SYSTEM """{persona_system}"""\n' if persona_system else "")
|
| 175 |
+ 'PARAMETER temperature 0.7\nPARAMETER repeat_penalty 1.3\nPARAMETER stop "<|im_end|>"\n')
|
| 176 |
-
plan = {"repo": repo_id, "private": private, "files": sorted(staged)
|
| 177 |
"bytes": sum(f.stat().st_size for f in staged.values())}
|
| 178 |
if dry_run:
|
| 179 |
return {"ok": True, "dry_run": True, "would_upload": plan, "card_preview": card[:400]}
|
| 180 |
api, err = _api(token)
|
| 181 |
if err:
|
| 182 |
return err
|
| 183 |
-
import tempfile
|
| 184 |
import shutil
|
|
|
|
| 185 |
with tempfile.TemporaryDirectory() as tmp:
|
| 186 |
stage = Path(tmp)
|
| 187 |
for rel, src in staged.items():
|
|
@@ -194,15 +194,15 @@ def publish_generation(run_dir: Path | str, repo_id: str, *, token: Optional[str
|
|
| 194 |
api.create_repo(repo_id, repo_type="model", private=private, exist_ok=True)
|
| 195 |
ci = api.upload_folder(folder_path=str(stage), repo_id=repo_id, repo_type="model",
|
| 196 |
commit_message=meta.get("commit_message") or "mindXtrain: a run, published with its evidence")
|
| 197 |
-
except Exception as e:
|
| 198 |
return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
|
| 199 |
return {"ok": True, "repo": repo_id, "url": f"https://huggingface.co/{repo_id}",
|
| 200 |
"commit": str(getattr(ci, "oid", ci))[:12], "uploaded": plan}
|
| 201 |
|
| 202 |
|
| 203 |
# ── the corpus ────────────────────────────────────────────────────────────────
|
| 204 |
-
def push_dataset(path: Path | str, repo_id: str, *, token:
|
| 205 |
-
path_in_repo: str = "", manifest:
|
| 206 |
"""Push a training corpus (a folder or one JSONL) with an optional manifest beside it."""
|
| 207 |
api, err = _api(token)
|
| 208 |
if err:
|
|
@@ -222,7 +222,7 @@ def push_dataset(path: Path | str, repo_id: str, *, token: Optional[str] = None,
|
|
| 222 |
api.upload_file(path_or_fileobj=json.dumps(manifest, indent=1).encode(),
|
| 223 |
path_in_repo=f"{path_in_repo}/MANIFEST.json".lstrip("/"),
|
| 224 |
repo_id=repo_id, repo_type="dataset", commit_message="mindXtrain: corpus manifest")
|
| 225 |
-
except Exception as e:
|
| 226 |
return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
|
| 227 |
return {"ok": True, "repo": repo_id, "url": f"https://huggingface.co/datasets/{repo_id}",
|
| 228 |
"commit": str(getattr(ci, "oid", ci))[:12]}
|
|
@@ -232,8 +232,8 @@ def push_dataset(path: Path | str, repo_id: str, *, token: Optional[str] = None,
|
|
| 232 |
_GEN = re.compile(r"(?:^|/)gen(\d+)(?:/|$)")
|
| 233 |
|
| 234 |
|
| 235 |
-
def lineage(repo_id: str, *, repo_type: str = "model", token:
|
| 236 |
-
local_runs:
|
| 237 |
"""Generations present in a repo, and — when `local_runs` is given — which local runs are not
|
| 238 |
published yet. Uses `tree_paths`, so folders never break the scan."""
|
| 239 |
api, err = _api(token)
|
|
@@ -241,9 +241,9 @@ def lineage(repo_id: str, *, repo_type: str = "model", token: Optional[str] = No
|
|
| 241 |
return err
|
| 242 |
try:
|
| 243 |
paths = tree_paths(api, repo_id, repo_type=repo_type, recursive=True)
|
| 244 |
-
except Exception as e:
|
| 245 |
return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:200]}"}
|
| 246 |
-
gens:
|
| 247 |
for p in paths:
|
| 248 |
m = _GEN.search(p)
|
| 249 |
if m:
|
|
@@ -262,9 +262,9 @@ def lineage(repo_id: str, *, repo_type: str = "model", token: Optional[str] = No
|
|
| 262 |
|
| 263 |
|
| 264 |
# ── Spaces ────────────────────────────────────────────────────────────────────
|
| 265 |
-
def push_space(folder: Path | str, space_id: str, *, token:
|
| 266 |
-
hardware:
|
| 267 |
-
secrets:
|
| 268 |
"""Push a Gradio folder as a Space. Existence is checked BEFORE creation (the Hub tests the
|
| 269 |
ZeroGPU quota first and answers 402 on a repo that already exists), and a README
|
| 270 |
`short_description` longer than 60 characters is refused by the Hub, so it is checked here."""
|
|
@@ -292,9 +292,10 @@ def push_space(folder: Path | str, space_id: str, *, token: Optional[str] = None
|
|
| 292 |
for k, v in (secrets or {}).items():
|
| 293 |
api.add_space_secret(space_id, k, v)
|
| 294 |
rt = api.get_space_runtime(space_id)
|
| 295 |
-
except Exception as e:
|
| 296 |
return {"ok": False, "space": space_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
|
| 297 |
return {"ok": True, "space": space_id, "existed": exists, "stage": str(rt.stage),
|
|
|
|
| 298 |
"url": f"https://huggingface.co/spaces/{space_id}",
|
| 299 |
"host": "https://" + space_id.replace("/", "-").replace("_", "-").lower() + ".hf.space",
|
| 300 |
"note": "free personal accounts host 2 ZeroGPU Spaces; a free org hosts none (402)"}
|
|
|
|
| 20 |
import re
|
| 21 |
import time
|
| 22 |
from pathlib import Path
|
| 23 |
+
from typing import Any
|
| 24 |
|
| 25 |
_HF_ENV = ("HF_TOKEN", "HUGGING_FACE_HUB_TOKEN", "HUGGINGFACEHUB_API_TOKEN")
|
| 26 |
|
| 27 |
|
| 28 |
+
def _token(explicit: str | None = None) -> str | None:
|
| 29 |
if explicit:
|
| 30 |
return explicit
|
| 31 |
for k in _HF_ENV:
|
|
|
|
| 35 |
return None
|
| 36 |
|
| 37 |
|
| 38 |
+
def _api(token: str | None = None):
|
| 39 |
"""(HfApi, None) or (None, {"ok": False, …}) — never raises."""
|
| 40 |
try:
|
| 41 |
from huggingface_hub import HfApi
|
| 42 |
except ImportError:
|
| 43 |
+
return None, {"ok": False, "reason": "huggingface_hub not installed — `uv sync --extra hf`"}
|
| 44 |
tok = _token(token)
|
| 45 |
if not tok:
|
| 46 |
return None, {"ok": False, "reason": f"no token — set one of {', '.join(_HF_ENV)}"}
|
| 47 |
return HfApi(token=tok), None
|
| 48 |
|
| 49 |
|
| 50 |
+
def tree_paths(api, repo_id: str, **kw) -> list[str]:
|
| 51 |
"""File paths of a repo tree (folders skipped). See the module docstring: the tree's entries
|
| 52 |
carry `.path`, not `.rfilename`."""
|
| 53 |
return [f.path for f in api.list_repo_tree(repo_id, **kw) if type(f).__name__ != "RepoFolder"]
|
| 54 |
|
| 55 |
|
| 56 |
# ── who the token is ──────────────────────────────────────────────────────────
|
| 57 |
+
def account(token: str | None = None) -> dict[str, Any]:
|
| 58 |
"""The identity behind the token, its role, the orgs it belongs to, and — separately — the
|
| 59 |
namespaces it can actually write."""
|
| 60 |
api, err = _api(token)
|
|
|
|
| 62 |
return err
|
| 63 |
try:
|
| 64 |
me = api.whoami()
|
| 65 |
+
except Exception as e:
|
| 66 |
return {"ok": False, "reason": f"whoami failed: {type(e).__name__}: {str(e)[:200]}"}
|
| 67 |
auth = (me.get("auth") or {}).get("accessToken") or {}
|
| 68 |
perms = auth.get("fineGrained") or {}
|
|
|
|
| 82 |
|
| 83 |
|
| 84 |
# ── bases, fetched before the run rather than during it ───────────────────────
|
| 85 |
+
def pull_base(model_id: str, *, token: str | None = None, allow_patterns: list[str] | None = None) -> dict[str, Any]:
|
| 86 |
"""Download a base model to the local cache. Do this BEFORE training: a run that discovers a
|
| 87 |
missing base three hours in has wasted three hours."""
|
| 88 |
try:
|
| 89 |
from huggingface_hub import snapshot_download
|
| 90 |
except ImportError:
|
| 91 |
+
return {"ok": False, "reason": "huggingface_hub not installed — `uv sync --extra hf`"}
|
| 92 |
t0 = time.time()
|
| 93 |
try:
|
| 94 |
path = snapshot_download(model_id, token=_token(token), allow_patterns=allow_patterns)
|
| 95 |
+
except Exception as e:
|
| 96 |
return {"ok": False, "model": model_id, "reason": f"{type(e).__name__}: {str(e)[:200]}"}
|
| 97 |
p = Path(path)
|
| 98 |
return {"ok": True, "model": model_id, "path": str(p), "seconds": round(time.time() - t0, 1),
|
| 99 |
"bytes": sum(f.stat().st_size for f in p.rglob("*") if f.is_file())}
|
| 100 |
|
| 101 |
|
| 102 |
+
def warm(config: Path | str, *, token: str | None = None) -> dict[str, Any]:
|
| 103 |
"""Pull whatever a run.yaml says it needs (the base model) so `train` starts cold-free."""
|
| 104 |
try:
|
| 105 |
import yaml
|
| 106 |
cfg = yaml.safe_load(Path(config).read_text())
|
| 107 |
+
except Exception as e:
|
| 108 |
return {"ok": False, "reason": f"unreadable config: {type(e).__name__}: {str(e)[:160]}"}
|
| 109 |
base = ((cfg or {}).get("model") or {}).get("name")
|
| 110 |
if not base:
|
|
|
|
| 113 |
|
| 114 |
|
| 115 |
# ── a finished run, published ─────────────────────────────────────────────────
|
| 116 |
+
def _card(repo_id: str, meta: dict[str, Any]) -> str:
|
| 117 |
"""A model card written from the run's own numbers. No claim that is not in `meta`."""
|
| 118 |
m = meta
|
| 119 |
fm = {"license": m.get("license", "apache-2.0"), "library_name": "transformers",
|
| 120 |
+
"pipeline_tag": "text-generation", "tags": ["mindxtrain", "lora", *list(m.get("tags") or [])]}
|
| 121 |
if m.get("base"):
|
| 122 |
fm["base_model"] = m["base"]
|
| 123 |
if m.get("dataset"):
|
|
|
|
| 137 |
"the training log ships beside the weights.\n")
|
| 138 |
|
| 139 |
|
| 140 |
+
def publish_generation(run_dir: Path | str, repo_id: str, *, token: str | None = None, private: bool = False,
|
| 141 |
+
meta: dict[str, Any] | None = None, persona_system: str | None = None,
|
| 142 |
+
include_merged: bool = True, dry_run: bool = False) -> dict[str, Any]:
|
| 143 |
"""A finished run as a model repo: merged weights at the root (if present), the LoRA delta under
|
| 144 |
`adapter/`, `train.log`, a `Modelfile` for Ollama, and a card built from `meta`.
|
| 145 |
|
|
|
|
| 149 |
return {"ok": False, "reason": f"no run dir at {run}"}
|
| 150 |
ck = run / "checkpoint"
|
| 151 |
merged = run / "ollama_push" / "merged"
|
| 152 |
+
staged: dict[str, Path] = {}
|
| 153 |
if include_merged and merged.is_dir():
|
| 154 |
for f in merged.iterdir():
|
| 155 |
if f.is_file():
|
| 156 |
staged[f.name] = f
|
| 157 |
if ck.is_dir():
|
| 158 |
for f in ck.iterdir():
|
| 159 |
+
if (f.is_file() and f.name != "training_args.bin") or f.name == "training_args.bin":
|
| 160 |
staged[f"adapter/{f.name}"] = f
|
| 161 |
log = run / "train.log"
|
| 162 |
if log.is_file():
|
|
|
|
| 167 |
if not meta.get("base"):
|
| 168 |
try:
|
| 169 |
meta["base"] = json.loads((ck / "adapter_config.json").read_text()).get("base_model_name_or_path")
|
| 170 |
+
except Exception:
|
| 171 |
pass
|
| 172 |
card = _card(repo_id, meta)
|
| 173 |
modelfile = ("# ollama create <name> -f Modelfile (from this repo's directory)\nFROM .\n"
|
| 174 |
+ (f'SYSTEM """{persona_system}"""\n' if persona_system else "")
|
| 175 |
+ 'PARAMETER temperature 0.7\nPARAMETER repeat_penalty 1.3\nPARAMETER stop "<|im_end|>"\n')
|
| 176 |
+
plan = {"repo": repo_id, "private": private, "files": [*sorted(staged), "README.md", "Modelfile"],
|
| 177 |
"bytes": sum(f.stat().st_size for f in staged.values())}
|
| 178 |
if dry_run:
|
| 179 |
return {"ok": True, "dry_run": True, "would_upload": plan, "card_preview": card[:400]}
|
| 180 |
api, err = _api(token)
|
| 181 |
if err:
|
| 182 |
return err
|
|
|
|
| 183 |
import shutil
|
| 184 |
+
import tempfile
|
| 185 |
with tempfile.TemporaryDirectory() as tmp:
|
| 186 |
stage = Path(tmp)
|
| 187 |
for rel, src in staged.items():
|
|
|
|
| 194 |
api.create_repo(repo_id, repo_type="model", private=private, exist_ok=True)
|
| 195 |
ci = api.upload_folder(folder_path=str(stage), repo_id=repo_id, repo_type="model",
|
| 196 |
commit_message=meta.get("commit_message") or "mindXtrain: a run, published with its evidence")
|
| 197 |
+
except Exception as e:
|
| 198 |
return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
|
| 199 |
return {"ok": True, "repo": repo_id, "url": f"https://huggingface.co/{repo_id}",
|
| 200 |
"commit": str(getattr(ci, "oid", ci))[:12], "uploaded": plan}
|
| 201 |
|
| 202 |
|
| 203 |
# ── the corpus ────────────────────────────────────────────────────────────────
|
| 204 |
+
def push_dataset(path: Path | str, repo_id: str, *, token: str | None = None, private: bool = False,
|
| 205 |
+
path_in_repo: str = "", manifest: dict[str, Any] | None = None) -> dict[str, Any]:
|
| 206 |
"""Push a training corpus (a folder or one JSONL) with an optional manifest beside it."""
|
| 207 |
api, err = _api(token)
|
| 208 |
if err:
|
|
|
|
| 222 |
api.upload_file(path_or_fileobj=json.dumps(manifest, indent=1).encode(),
|
| 223 |
path_in_repo=f"{path_in_repo}/MANIFEST.json".lstrip("/"),
|
| 224 |
repo_id=repo_id, repo_type="dataset", commit_message="mindXtrain: corpus manifest")
|
| 225 |
+
except Exception as e:
|
| 226 |
return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
|
| 227 |
return {"ok": True, "repo": repo_id, "url": f"https://huggingface.co/datasets/{repo_id}",
|
| 228 |
"commit": str(getattr(ci, "oid", ci))[:12]}
|
|
|
|
| 232 |
_GEN = re.compile(r"(?:^|/)gen(\d+)(?:/|$)")
|
| 233 |
|
| 234 |
|
| 235 |
+
def lineage(repo_id: str, *, repo_type: str = "model", token: str | None = None,
|
| 236 |
+
local_runs: Path | str | None = None) -> dict[str, Any]:
|
| 237 |
"""Generations present in a repo, and — when `local_runs` is given — which local runs are not
|
| 238 |
published yet. Uses `tree_paths`, so folders never break the scan."""
|
| 239 |
api, err = _api(token)
|
|
|
|
| 241 |
return err
|
| 242 |
try:
|
| 243 |
paths = tree_paths(api, repo_id, repo_type=repo_type, recursive=True)
|
| 244 |
+
except Exception as e:
|
| 245 |
return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:200]}"}
|
| 246 |
+
gens: dict[int, list[str]] = {}
|
| 247 |
for p in paths:
|
| 248 |
m = _GEN.search(p)
|
| 249 |
if m:
|
|
|
|
| 262 |
|
| 263 |
|
| 264 |
# ── Spaces ────────────────────────────────────────────────────────────────────
|
| 265 |
+
def push_space(folder: Path | str, space_id: str, *, token: str | None = None, private: bool = True,
|
| 266 |
+
hardware: str | None = "zero-a10g", variables: dict[str, str] | None = None,
|
| 267 |
+
secrets: dict[str, str] | None = None) -> dict[str, Any]:
|
| 268 |
"""Push a Gradio folder as a Space. Existence is checked BEFORE creation (the Hub tests the
|
| 269 |
ZeroGPU quota first and answers 402 on a repo that already exists), and a README
|
| 270 |
`short_description` longer than 60 characters is refused by the Hub, so it is checked here."""
|
|
|
|
| 292 |
for k, v in (secrets or {}).items():
|
| 293 |
api.add_space_secret(space_id, k, v)
|
| 294 |
rt = api.get_space_runtime(space_id)
|
| 295 |
+
except Exception as e:
|
| 296 |
return {"ok": False, "space": space_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
|
| 297 |
return {"ok": True, "space": space_id, "existed": exists, "stage": str(rt.stage),
|
| 298 |
+
"commit": str(getattr(ci, "oid", ci))[:12],
|
| 299 |
"url": f"https://huggingface.co/spaces/{space_id}",
|
| 300 |
"host": "https://" + space_id.replace("/", "-").replace("_", "-").lower() + ".hf.space",
|
| 301 |
"note": "free personal accounts host 2 ZeroGPU Spaces; a free org hosts none (402)"}
|
mindxtrain/storage/hf_hub.py
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
"""Push checkpoint + model card to Hugging Face Hub.
|
| 2 |
|
| 3 |
-
Lazy `import huggingface_hub` so users without `--extra
|
| 4 |
import this module. `HF_TOKEN` is read from env (or the token kwarg).
|
| 5 |
"""
|
| 6 |
|
|
@@ -24,7 +24,7 @@ def publish_to_hf(
|
|
| 24 |
try:
|
| 25 |
from huggingface_hub import HfApi
|
| 26 |
except ImportError as exc:
|
| 27 |
-
msg = "huggingface_hub not installed; run `uv sync --extra
|
| 28 |
raise RuntimeError(msg) from exc
|
| 29 |
|
| 30 |
api = HfApi(token=token or os.environ.get("HF_TOKEN"))
|
|
@@ -56,7 +56,7 @@ class HfHubProvider(StorageProvider):
|
|
| 56 |
try:
|
| 57 |
from huggingface_hub import snapshot_download
|
| 58 |
except ImportError as exc:
|
| 59 |
-
msg = "huggingface_hub not installed; run `uv sync --extra
|
| 60 |
raise RuntimeError(msg) from exc
|
| 61 |
# ref.uri is `https://huggingface.co/<repo_id>` — extract repo_id.
|
| 62 |
repo_id = ref.uri.removeprefix("https://huggingface.co/")
|
|
|
|
| 1 |
"""Push checkpoint + model card to Hugging Face Hub.
|
| 2 |
|
| 3 |
+
Lazy `import huggingface_hub` so users without `--extra hf` can still
|
| 4 |
import this module. `HF_TOKEN` is read from env (or the token kwarg).
|
| 5 |
"""
|
| 6 |
|
|
|
|
| 24 |
try:
|
| 25 |
from huggingface_hub import HfApi
|
| 26 |
except ImportError as exc:
|
| 27 |
+
msg = "huggingface_hub not installed; run `uv sync --extra hf`."
|
| 28 |
raise RuntimeError(msg) from exc
|
| 29 |
|
| 30 |
api = HfApi(token=token or os.environ.get("HF_TOKEN"))
|
|
|
|
| 56 |
try:
|
| 57 |
from huggingface_hub import snapshot_download
|
| 58 |
except ImportError as exc:
|
| 59 |
+
msg = "huggingface_hub not installed; run `uv sync --extra hf`."
|
| 60 |
raise RuntimeError(msg) from exc
|
| 61 |
# ref.uri is `https://huggingface.co/<repo_id>` — extract repo_id.
|
| 62 |
repo_id = ref.uri.removeprefix("https://huggingface.co/")
|
mindxtrain/ui/app.py
CHANGED
|
@@ -36,7 +36,7 @@ import gradio as gr
|
|
| 36 |
from .metrics import RunMetrics, parse_log # noqa: F401 (parse_log re-exported for tests)
|
| 37 |
from .theme import CSS, theme
|
| 38 |
|
| 39 |
-
VERSION = "1.0.
|
| 40 |
HOME = Path(os.environ.get("MINDXTRAIN_HOME") or Path(__file__).resolve().parents[2])
|
| 41 |
RECIPES = HOME / "mindxtrain" / "train" / "recipes"
|
| 42 |
TIERS = ["Basic", "Advanced", "Scientific"]
|
|
|
|
| 36 |
from .metrics import RunMetrics, parse_log # noqa: F401 (parse_log re-exported for tests)
|
| 37 |
from .theme import CSS, theme
|
| 38 |
|
| 39 |
+
VERSION = "1.0.2"
|
| 40 |
HOME = Path(os.environ.get("MINDXTRAIN_HOME") or Path(__file__).resolve().parents[2])
|
| 41 |
RECIPES = HOME / "mindxtrain" / "train" / "recipes"
|
| 42 |
TIERS = ["Basic", "Advanced", "Scientific"]
|
pyproject.toml
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
[project]
|
| 2 |
name = "mindxtrain"
|
| 3 |
-
version = "1.0.
|
| 4 |
description = "Production training framework for fine-tuning open-weight LLMs on AMD MI300X and serving them through an OpenAI-compatible API."
|
| 5 |
requires-python = ">=3.12,<3.13"
|
| 6 |
license = { text = "Apache-2.0" }
|
|
@@ -45,10 +45,14 @@ ui = [
|
|
| 45 |
"gradio[mcp]>=5.0",
|
| 46 |
"pyyaml>=6.0",
|
| 47 |
]
|
|
|
|
|
|
|
|
|
|
|
|
|
| 48 |
chain = [
|
| 49 |
"web3>=7.5",
|
| 50 |
"py-algorand-sdk>=2.7",
|
| 51 |
-
"
|
| 52 |
]
|
| 53 |
obs = [
|
| 54 |
"opentelemetry-sdk>=1.28",
|
|
@@ -56,7 +60,7 @@ obs = [
|
|
| 56 |
"psutil>=6.1",
|
| 57 |
]
|
| 58 |
all = [
|
| 59 |
-
"mindxtrain[ml,eval,data,serve,chain,obs,ui]",
|
| 60 |
]
|
| 61 |
|
| 62 |
[build-system]
|
|
|
|
| 1 |
[project]
|
| 2 |
name = "mindxtrain"
|
| 3 |
+
version = "1.0.2"
|
| 4 |
description = "Production training framework for fine-tuning open-weight LLMs on AMD MI300X and serving them through an OpenAI-compatible API."
|
| 5 |
requires-python = ">=3.12,<3.13"
|
| 6 |
license = { text = "Apache-2.0" }
|
|
|
|
| 45 |
"gradio[mcp]>=5.0",
|
| 46 |
"pyyaml>=6.0",
|
| 47 |
]
|
| 48 |
+
# The Hub is not a chain. It had been living inside `chain`, so reaching Hugging Face
|
| 49 |
+
# meant installing web3 and the Algorand SDK; `chain` still pulls it so existing
|
| 50 |
+
# installs keep working.
|
| 51 |
+
hf = ["huggingface-hub>=0.26"]
|
| 52 |
chain = [
|
| 53 |
"web3>=7.5",
|
| 54 |
"py-algorand-sdk>=2.7",
|
| 55 |
+
"mindxtrain[hf]",
|
| 56 |
]
|
| 57 |
obs = [
|
| 58 |
"opentelemetry-sdk>=1.28",
|
|
|
|
| 60 |
"psutil>=6.1",
|
| 61 |
]
|
| 62 |
all = [
|
| 63 |
+
"mindxtrain[ml,eval,data,serve,chain,obs,ui,hf]",
|
| 64 |
]
|
| 65 |
|
| 66 |
[build-system]
|
tests/test_hf_extension.py
ADDED
|
@@ -0,0 +1,159 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Units for the Hugging Face extension and its CLI surface.
|
| 2 |
+
|
| 3 |
+
`huggingface_hub` is an optional extra and is NOT installed in this environment, which makes
|
| 4 |
+
the missing-dependency path a real test rather than a mocked one. Everything else runs against
|
| 5 |
+
stubs: no token, no network, no Hub.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from __future__ import annotations
|
| 9 |
+
|
| 10 |
+
from typing import Any
|
| 11 |
+
|
| 12 |
+
import pytest
|
| 13 |
+
from typer.testing import CliRunner
|
| 14 |
+
|
| 15 |
+
from mindxtrain.cli.main import app
|
| 16 |
+
from mindxtrain.hf import extension as ext
|
| 17 |
+
|
| 18 |
+
runner = CliRunner()
|
| 19 |
+
|
| 20 |
+
|
| 21 |
+
@pytest.fixture(autouse=True)
|
| 22 |
+
def no_ambient_token(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 23 |
+
"""A developer's real HF_TOKEN must not change what these tests assert."""
|
| 24 |
+
for name in ext._HF_ENV:
|
| 25 |
+
monkeypatch.delenv(name, raising=False)
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
# --- token resolution --------------------------------------------------------------------
|
| 29 |
+
|
| 30 |
+
|
| 31 |
+
def test_explicit_token_beats_the_environment(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 32 |
+
monkeypatch.setenv("HF_TOKEN", "from-env")
|
| 33 |
+
assert ext._token("explicit") == "explicit"
|
| 34 |
+
|
| 35 |
+
|
| 36 |
+
def test_environment_names_are_tried_in_order(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 37 |
+
monkeypatch.setenv("HUGGINGFACEHUB_API_TOKEN", "third")
|
| 38 |
+
assert ext._token() == "third"
|
| 39 |
+
monkeypatch.setenv("HF_TOKEN", "first")
|
| 40 |
+
assert ext._token() == "first"
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
def test_no_token_anywhere_is_none() -> None:
|
| 44 |
+
assert ext._token() is None
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
# --- the optional dependency is reported, never raised -----------------------------------
|
| 48 |
+
|
| 49 |
+
|
| 50 |
+
def test_missing_dependency_is_a_result_not_an_exception() -> None:
|
| 51 |
+
"""`huggingface_hub` is genuinely absent here — this asserts the real path."""
|
| 52 |
+
api, err = ext._api("a-token")
|
| 53 |
+
assert api is None
|
| 54 |
+
assert err["ok"] is False
|
| 55 |
+
assert "--extra hf" in err["reason"] # the extra that actually installs it
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
def test_pull_base_reports_the_missing_dependency() -> None:
|
| 59 |
+
out = ext.pull_base("HuggingFaceTB/SmolLM2-135M")
|
| 60 |
+
assert out["ok"] is False and "huggingface_hub" in out["reason"]
|
| 61 |
+
|
| 62 |
+
|
| 63 |
+
def test_no_token_names_the_variables_to_set(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 64 |
+
"""With the library present but no token, the error must say which env vars are read."""
|
| 65 |
+
monkeypatch.setattr(ext, "_api", ext._api) # keep the real function
|
| 66 |
+
fake_hub = type("m", (), {"HfApi": lambda **kw: None})
|
| 67 |
+
monkeypatch.setitem(__import__("sys").modules, "huggingface_hub", fake_hub)
|
| 68 |
+
api, err = ext._api()
|
| 69 |
+
assert api is None
|
| 70 |
+
assert "HF_TOKEN" in err["reason"]
|
| 71 |
+
|
| 72 |
+
|
| 73 |
+
# --- tree_paths: the RepoFolder trap -----------------------------------------------------
|
| 74 |
+
|
| 75 |
+
|
| 76 |
+
# The filter matches on the CLASS NAME, so these stubs must carry the Hub's exact names —
|
| 77 |
+
# `tree_paths` deliberately avoids importing the optional type just to check it.
|
| 78 |
+
class RepoFile:
|
| 79 |
+
def __init__(self, path: str) -> None:
|
| 80 |
+
self.path = path
|
| 81 |
+
|
| 82 |
+
|
| 83 |
+
class RepoFolder:
|
| 84 |
+
def __init__(self, path: str) -> None:
|
| 85 |
+
self.path = path
|
| 86 |
+
|
| 87 |
+
|
| 88 |
+
def test_tree_paths_skips_folders() -> None:
|
| 89 |
+
"""Folders and files both carry `.path`; counting folders as files corrupts any scan."""
|
| 90 |
+
api = type("A", (), {"list_repo_tree": lambda self, r, **kw: [
|
| 91 |
+
RepoFile("gen1/adapter.safetensors"), RepoFolder("gen1"), RepoFile("README.md"),
|
| 92 |
+
]})()
|
| 93 |
+
assert ext.tree_paths(api, "PYTHAI/x") == ["gen1/adapter.safetensors", "README.md"]
|
| 94 |
+
|
| 95 |
+
|
| 96 |
+
def test_tree_paths_on_a_tree_of_only_folders_is_empty() -> None:
|
| 97 |
+
"""The bug this guards: a scan reporting "nothing on the Hub" while the repo is full."""
|
| 98 |
+
api = type("A", (), {"list_repo_tree": lambda self, r, **kw: [RepoFolder("gen1"), RepoFolder("gen2")]})()
|
| 99 |
+
assert ext.tree_paths(api, "PYTHAI/x") == []
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
# --- guards that must not need a network -------------------------------------------------
|
| 103 |
+
|
| 104 |
+
|
| 105 |
+
def test_warm_refuses_a_config_without_a_model(tmp_path) -> None: # type: ignore[no-untyped-def]
|
| 106 |
+
cfg = tmp_path / "run.yaml"
|
| 107 |
+
cfg.write_text("train:\n epochs: 1\n", encoding="utf-8")
|
| 108 |
+
out = ext.warm(cfg)
|
| 109 |
+
assert out["ok"] is False
|
| 110 |
+
assert "model.name" in out["reason"] or "huggingface_hub" in out["reason"]
|
| 111 |
+
|
| 112 |
+
|
| 113 |
+
def test_publish_refuses_a_missing_run_dir(tmp_path) -> None: # type: ignore[no-untyped-def]
|
| 114 |
+
out = ext.publish_generation(tmp_path / "nope", "PYTHAI/x")
|
| 115 |
+
assert out["ok"] is False and "no run dir" in out["reason"]
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
# --- the CLI surface ---------------------------------------------------------------------
|
| 119 |
+
|
| 120 |
+
|
| 121 |
+
def test_hf_is_registered_as_a_command_group() -> None:
|
| 122 |
+
result = runner.invoke(app, ["--help"])
|
| 123 |
+
assert "hf" in result.stdout
|
| 124 |
+
|
| 125 |
+
|
| 126 |
+
def test_every_hf_verb_is_reachable() -> None:
|
| 127 |
+
result = runner.invoke(app, ["hf", "--help"])
|
| 128 |
+
for verb in ("whoami", "pull", "warm", "publish", "lineage", "dataset", "space"):
|
| 129 |
+
assert verb in result.stdout
|
| 130 |
+
|
| 131 |
+
|
| 132 |
+
def test_failure_exits_nonzero_so_a_script_can_branch() -> None:
|
| 133 |
+
"""The ascent loop warms a base and checks $? — an `ok:false` that exits 0 would be a trap."""
|
| 134 |
+
result = runner.invoke(app, ["hf", "whoami"])
|
| 135 |
+
assert result.exit_code == 1
|
| 136 |
+
assert "huggingface_hub" in result.stdout
|
| 137 |
+
|
| 138 |
+
|
| 139 |
+
def test_success_exits_zero_and_prints_the_result(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 140 |
+
import mindxtrain.hf as hf_pkg
|
| 141 |
+
|
| 142 |
+
monkeypatch.setattr(hf_pkg, "account", lambda tok=None: {"ok": True, "name": "Gregory-L",
|
| 143 |
+
"can_write": ["PYTHAI"]})
|
| 144 |
+
result = runner.invoke(app, ["hf", "whoami"])
|
| 145 |
+
assert result.exit_code == 0
|
| 146 |
+
assert "Gregory-L" in result.stdout
|
| 147 |
+
|
| 148 |
+
|
| 149 |
+
def test_token_option_is_passed_through(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 150 |
+
seen: dict[str, Any] = {}
|
| 151 |
+
import mindxtrain.hf as hf_pkg
|
| 152 |
+
|
| 153 |
+
def fake_account(tok=None): # type: ignore[no-untyped-def]
|
| 154 |
+
seen["token"] = tok
|
| 155 |
+
return {"ok": True}
|
| 156 |
+
|
| 157 |
+
monkeypatch.setattr(hf_pkg, "account", fake_account)
|
| 158 |
+
assert runner.invoke(app, ["hf", "whoami", "--token", "abc"]).exit_code == 0
|
| 159 |
+
assert seen["token"] == "abc"
|