Gregory-L commited on
Commit
c4e0c8c
·
verified ·
1 Parent(s): f575b80

v1.0.2 — the Hub is reachable from the command line, and is not a chain

Browse files

The hf/ extension existed and was complete, but nothing outside a Python import could reach it. New `mindxtrain hf` group: whoami / pull / warm / publish / lineage / dataset / space, each exiting nonzero when the result says ok:false so the ascent loop can warm a base and check $?.

huggingface-hub moves out of the `chain` extra into its own `hf` extra — reaching the Hub no longer means installing web3 and the Algorand SDK. `chain` pulls `mindxtrain[hf]` so existing installs keep working.

push_space dropped its commit receipt; it reports it now. First tests for the module.

765 tests pass (was 750).

docs/CHANGELOG.md CHANGED
@@ -6,6 +6,38 @@ project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
6
 
7
  ## [Unreleased]
8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9
  ## [1.0.1] — 2026-09-15
10
 
11
 
 
6
 
7
  ## [Unreleased]
8
 
9
+ ## [1.0.2] — 2026-09-15
10
+
11
+ ### Added
12
+
13
+ - **mindXtrain can use Hugging Face from the command line.** The `mindxtrain/hf/` extension
14
+ existed and was complete — `account`, `pull_base`, `warm`, `publish_generation`,
15
+ `push_dataset`, `lineage`, `push_space` — but nothing outside a Python import could reach
16
+ it: there was no CLI verb, so on a node the Hub was effectively unavailable. New
17
+ `mindxtrain hf` group: `whoami · pull · warm · publish · lineage · dataset · space`.
18
+ Each prints the extension's own result dict and **exits nonzero when it reports
19
+ `ok: false`**, so the ascent loop can warm a base, check `$?`, and refuse to start a
20
+ three-hour run that would otherwise discover the missing base at the end of it.
21
+ - **Tests for the Hub extension**, which had none: token precedence across the three
22
+ accepted environment variables, the missing-dependency path (genuinely missing here, not
23
+ mocked), the `RepoFolder` trap that makes a scan report an empty repo, the guards that
24
+ must not need a network, and the CLI wiring including the nonzero exit.
25
+
26
+ ### Changed
27
+
28
+ - **`huggingface-hub` moved out of the `chain` extra into its own `hf` extra.** Reaching the
29
+ Hub used to mean installing `web3` and the Algorand SDK, because the Hub was filed under
30
+ "chain". It is not a chain. `chain` now pulls `mindxtrain[hf]` so existing installs keep
31
+ working, `all` includes it, and every "not installed" message names `--extra hf` — the
32
+ extra that actually installs it — instead of sending the reader to the blockchain one.
33
+ - The `hf` module is ruff clean under the house config.
34
+
35
+ ### Fixed
36
+
37
+ - **`push_space` captured the upload's commit receipt and discarded it**, while its two
38
+ sibling publishers both return it. A Space push now reports its commit like everything
39
+ else, so what was pushed is identifiable afterwards.
40
+
41
  ## [1.0.1] — 2026-09-15
42
 
43
 
docs/NAV.md CHANGED
@@ -164,8 +164,9 @@ Target metrics + framework comparison.
164
  - [What's not measured (yet)](benchmarks.md#whats-not-measured-yet)
165
 
166
  ### [CHANGELOG](CHANGELOG.md)
167
- Version history (current: v1.0.1).
168
  - [[Unreleased]](CHANGELOG.md#unreleased)
 
169
  - [[1.0.1] — 2026-09-15](CHANGELOG.md#101--2026-09-15)
170
  - [[1.0.0] — 2026-06-11](CHANGELOG.md#100--2026-06-11)
171
  - [[0.1.0] — 2026-05-06](CHANGELOG.md#010--2026-05-06)
 
164
  - [What's not measured (yet)](benchmarks.md#whats-not-measured-yet)
165
 
166
  ### [CHANGELOG](CHANGELOG.md)
167
+ Version history (current: v1.0.2).
168
  - [[Unreleased]](CHANGELOG.md#unreleased)
169
+ - [[1.0.2] — 2026-09-15](CHANGELOG.md#102--2026-09-15)
170
  - [[1.0.1] — 2026-09-15](CHANGELOG.md#101--2026-09-15)
171
  - [[1.0.0] — 2026-06-11](CHANGELOG.md#100--2026-06-11)
172
  - [[0.1.0] — 2026-05-06](CHANGELOG.md#010--2026-05-06)
mindxtrain/__init__.py CHANGED
@@ -16,4 +16,4 @@ Layout per mindXtrain2.md §Part 4 "Repository layout":
16
  budget/ psutil-derived ResourceBudget (carry-over from aGLM)
17
  """
18
 
19
- __version__ = "1.0.1"
 
16
  budget/ psutil-derived ResourceBudget (carry-over from aGLM)
17
  """
18
 
19
+ __version__ = "1.0.2"
mindxtrain/cli/main.py CHANGED
@@ -25,10 +25,16 @@ mei_app = typer.Typer(
25
  help="mindX Efficiency Index — score, history, promotion checks.",
26
  no_args_is_help=True,
27
  )
 
 
 
 
 
28
  app.add_typer(dataset_app)
29
  app.add_typer(github_app)
30
  app.add_typer(droplet_app)
31
  app.add_typer(mei_app)
 
32
  console = Console()
33
 
34
 
@@ -710,6 +716,120 @@ def research(
710
  console.print(f"[green]{result.summary()}[/green]")
711
 
712
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
713
  # ---- github / droplet (source-tree publishing + remote provision) -------
714
 
715
 
 
25
  help="mindX Efficiency Index — score, history, promotion checks.",
26
  no_args_is_help=True,
27
  )
28
+ hf_app = typer.Typer(
29
+ name="hf",
30
+ help="Hugging Face — who the token is, warm a base, publish a generation, read the lineage.",
31
+ no_args_is_help=True,
32
+ )
33
  app.add_typer(dataset_app)
34
  app.add_typer(github_app)
35
  app.add_typer(droplet_app)
36
  app.add_typer(mei_app)
37
+ app.add_typer(hf_app)
38
  console = Console()
39
 
40
 
 
716
  console.print(f"[green]{result.summary()}[/green]")
717
 
718
 
719
+
720
+ # ---- hugging face (the Hub as mindXtrain uses it) -----------------------------
721
+
722
+
723
+ def _hf_report(result: dict, *, quiet: bool = False) -> None:
724
+ """Print a result dict and exit nonzero when it says `ok: false`.
725
+
726
+ Every `mindxtrain.hf` function returns `{"ok": bool, ...}` rather than raising, so the exit
727
+ code is what makes it scriptable: the ascent loop can warm a base, check `$?`, and refuse to
728
+ start a three-hour run that would only discover the missing base at the end.
729
+ """
730
+ import json as _json
731
+
732
+ if not quiet:
733
+ console.print_json(_json.dumps(result, default=str))
734
+ if not result.get("ok"):
735
+ raise typer.Exit(code=1)
736
+
737
+
738
+ @hf_app.command("whoami")
739
+ def hf_whoami(
740
+ token: str = typer.Option("", "--token", help="Defaults to HF_TOKEN in the environment."),
741
+ ) -> None:
742
+ """Who the token is — and, separately, the namespaces it can actually WRITE.
743
+
744
+ Org membership is not write scope; the two are reported apart because assuming they are the
745
+ same is how a publish fails after the training finished.
746
+ """
747
+ from mindxtrain.hf import account
748
+
749
+ _hf_report(account(token or None))
750
+
751
+
752
+ @hf_app.command("pull")
753
+ def hf_pull(
754
+ model_id: str = typer.Argument(..., help="Base model repo id, e.g. HuggingFaceTB/SmolLM2-135M."),
755
+ token: str = typer.Option("", "--token"),
756
+ allow: list[str] = typer.Option(None, "--allow", help="Glob to restrict the download."),
757
+ ) -> None:
758
+ """Fetch a base model into the local cache before training needs it."""
759
+ from mindxtrain.hf import pull_base
760
+
761
+ _hf_report(pull_base(model_id, token=token or None, allow_patterns=list(allow) if allow else None))
762
+
763
+
764
+ @hf_app.command("warm")
765
+ def hf_warm(
766
+ config: Path = typer.Argument(..., help="run.yaml — its model.name is what gets pulled."),
767
+ token: str = typer.Option("", "--token"),
768
+ ) -> None:
769
+ """Pull whatever a run config says it needs, so `train` starts cold-free."""
770
+ from mindxtrain.hf import warm
771
+
772
+ _hf_report(warm(config, token=token or None))
773
+
774
+
775
+ @hf_app.command("publish")
776
+ def hf_publish(
777
+ run_dir: Path = typer.Argument(..., help="A finished run directory."),
778
+ repo_id: str = typer.Argument(..., help="Target model repo, e.g. PYTHAI/mindXascension."),
779
+ token: str = typer.Option("", "--token"),
780
+ private: bool = typer.Option(False, "--private", help="Create the repo private."),
781
+ no_merged: bool = typer.Option(False, "--no-merged", help="Adapter only; skip merged weights."),
782
+ dry_run: bool = typer.Option(False, "--dry-run", help="Report what would upload; touch nothing."),
783
+ ) -> None:
784
+ """Publish a finished run as a model repo: weights, adapter, train.log, Modelfile, card."""
785
+ from mindxtrain.hf import publish_generation
786
+
787
+ _hf_report(publish_generation(run_dir, repo_id, token=token or None, private=private,
788
+ include_merged=not no_merged, dry_run=dry_run))
789
+
790
+
791
+ @hf_app.command("lineage")
792
+ def hf_lineage(
793
+ repo_id: str = typer.Argument(..., help="Repo to scan."),
794
+ repo_type: str = typer.Option("model", "--repo-type", help="model | dataset | space."),
795
+ local_runs: Path = typer.Option(None, "--local-runs", help="Run root, to list what is unpublished."),
796
+ token: str = typer.Option("", "--token"),
797
+ ) -> None:
798
+ """What is on the Hub for a project, reconciled against local runs."""
799
+ from mindxtrain.hf import lineage
800
+
801
+ _hf_report(lineage(repo_id, repo_type=repo_type, token=token or None, local_runs=local_runs))
802
+
803
+
804
+ @hf_app.command("dataset")
805
+ def hf_dataset(
806
+ path: Path = typer.Argument(..., help="A corpus folder or one JSONL."),
807
+ repo_id: str = typer.Argument(..., help="Target dataset repo."),
808
+ token: str = typer.Option("", "--token"),
809
+ private: bool = typer.Option(False, "--private"),
810
+ path_in_repo: str = typer.Option("", "--path-in-repo"),
811
+ ) -> None:
812
+ """Push a training corpus to a dataset repo."""
813
+ from mindxtrain.hf import push_dataset
814
+
815
+ _hf_report(push_dataset(path, repo_id, token=token or None, private=private,
816
+ path_in_repo=path_in_repo))
817
+
818
+
819
+ @hf_app.command("space")
820
+ def hf_space(
821
+ folder: Path = typer.Argument(..., help="A Gradio folder (must contain app.py)."),
822
+ space_id: str = typer.Argument(..., help="Target Space id."),
823
+ token: str = typer.Option("", "--token"),
824
+ public: bool = typer.Option(False, "--public", help="Public Space. Never park a write token on one."),
825
+ hardware: str = typer.Option("zero-a10g", "--hardware", help="Empty string for the free CPU tier."),
826
+ ) -> None:
827
+ """Push a Gradio folder as a Space."""
828
+ from mindxtrain.hf import push_space
829
+
830
+ _hf_report(push_space(folder, space_id, token=token or None, private=not public,
831
+ hardware=hardware or None))
832
+
833
  # ---- github / droplet (source-tree publishing + remote provision) -------
834
 
835
 
mindxtrain/hf/__init__.py CHANGED
@@ -12,9 +12,9 @@ the rest of the Hub as mindXtrain needs it, and nothing more:
12
  - `datasets` — push/pull a training corpus with its manifest
13
 
14
  Every function returns a plain dict: `{"ok": bool, …}`. Nothing here raises at import time, and
15
- `huggingface_hub` is imported lazily so a CPU-only install without `--extra chain` still loads.
16
  """
17
- from .extension import ( # noqa: F401
18
  account,
19
  lineage,
20
  publish_generation,
 
12
  - `datasets` — push/pull a training corpus with its manifest
13
 
14
  Every function returns a plain dict: `{"ok": bool, …}`. Nothing here raises at import time, and
15
+ `huggingface_hub` is imported lazily so a CPU-only install without `--extra hf` still loads.
16
  """
17
+ from .extension import (
18
  account,
19
  lineage,
20
  publish_generation,
mindxtrain/hf/extension.py CHANGED
@@ -20,12 +20,12 @@ import os
20
  import re
21
  import time
22
  from pathlib import Path
23
- from typing import Any, Dict, List, Optional
24
 
25
  _HF_ENV = ("HF_TOKEN", "HUGGING_FACE_HUB_TOKEN", "HUGGINGFACEHUB_API_TOKEN")
26
 
27
 
28
- def _token(explicit: Optional[str] = None) -> Optional[str]:
29
  if explicit:
30
  return explicit
31
  for k in _HF_ENV:
@@ -35,26 +35,26 @@ def _token(explicit: Optional[str] = None) -> Optional[str]:
35
  return None
36
 
37
 
38
- def _api(token: Optional[str] = None):
39
  """(HfApi, None) or (None, {"ok": False, …}) — never raises."""
40
  try:
41
  from huggingface_hub import HfApi
42
  except ImportError:
43
- return None, {"ok": False, "reason": "huggingface_hub not installed — `uv sync --extra chain`"}
44
  tok = _token(token)
45
  if not tok:
46
  return None, {"ok": False, "reason": f"no token — set one of {', '.join(_HF_ENV)}"}
47
  return HfApi(token=tok), None
48
 
49
 
50
- def tree_paths(api, repo_id: str, **kw) -> List[str]:
51
  """File paths of a repo tree (folders skipped). See the module docstring: the tree's entries
52
  carry `.path`, not `.rfilename`."""
53
  return [f.path for f in api.list_repo_tree(repo_id, **kw) if type(f).__name__ != "RepoFolder"]
54
 
55
 
56
  # ── who the token is ──────────────────────────────────────────────────────────
57
- def account(token: Optional[str] = None) -> Dict[str, Any]:
58
  """The identity behind the token, its role, the orgs it belongs to, and — separately — the
59
  namespaces it can actually write."""
60
  api, err = _api(token)
@@ -62,7 +62,7 @@ def account(token: Optional[str] = None) -> Dict[str, Any]:
62
  return err
63
  try:
64
  me = api.whoami()
65
- except Exception as e: # noqa: BLE001
66
  return {"ok": False, "reason": f"whoami failed: {type(e).__name__}: {str(e)[:200]}"}
67
  auth = (me.get("auth") or {}).get("accessToken") or {}
68
  perms = auth.get("fineGrained") or {}
@@ -82,29 +82,29 @@ def account(token: Optional[str] = None) -> Dict[str, Any]:
82
 
83
 
84
  # ── bases, fetched before the run rather than during it ───────────────────────
85
- def pull_base(model_id: str, *, token: Optional[str] = None, allow_patterns: Optional[List[str]] = None) -> Dict[str, Any]:
86
  """Download a base model to the local cache. Do this BEFORE training: a run that discovers a
87
  missing base three hours in has wasted three hours."""
88
  try:
89
  from huggingface_hub import snapshot_download
90
  except ImportError:
91
- return {"ok": False, "reason": "huggingface_hub not installed — `uv sync --extra chain`"}
92
  t0 = time.time()
93
  try:
94
  path = snapshot_download(model_id, token=_token(token), allow_patterns=allow_patterns)
95
- except Exception as e: # noqa: BLE001
96
  return {"ok": False, "model": model_id, "reason": f"{type(e).__name__}: {str(e)[:200]}"}
97
  p = Path(path)
98
  return {"ok": True, "model": model_id, "path": str(p), "seconds": round(time.time() - t0, 1),
99
  "bytes": sum(f.stat().st_size for f in p.rglob("*") if f.is_file())}
100
 
101
 
102
- def warm(config: Path | str, *, token: Optional[str] = None) -> Dict[str, Any]:
103
  """Pull whatever a run.yaml says it needs (the base model) so `train` starts cold-free."""
104
  try:
105
  import yaml
106
  cfg = yaml.safe_load(Path(config).read_text())
107
- except Exception as e: # noqa: BLE001
108
  return {"ok": False, "reason": f"unreadable config: {type(e).__name__}: {str(e)[:160]}"}
109
  base = ((cfg or {}).get("model") or {}).get("name")
110
  if not base:
@@ -113,11 +113,11 @@ def warm(config: Path | str, *, token: Optional[str] = None) -> Dict[str, Any]:
113
 
114
 
115
  # ── a finished run, published ─────────────────────────────────────────────────
116
- def _card(repo_id: str, meta: Dict[str, Any]) -> str:
117
  """A model card written from the run's own numbers. No claim that is not in `meta`."""
118
  m = meta
119
  fm = {"license": m.get("license", "apache-2.0"), "library_name": "transformers",
120
- "pipeline_tag": "text-generation", "tags": ["mindxtrain", "lora"] + list(m.get("tags") or [])}
121
  if m.get("base"):
122
  fm["base_model"] = m["base"]
123
  if m.get("dataset"):
@@ -137,9 +137,9 @@ def _card(repo_id: str, meta: Dict[str, Any]) -> str:
137
  "the training log ships beside the weights.\n")
138
 
139
 
140
- def publish_generation(run_dir: Path | str, repo_id: str, *, token: Optional[str] = None, private: bool = False,
141
- meta: Optional[Dict[str, Any]] = None, persona_system: Optional[str] = None,
142
- include_merged: bool = True, dry_run: bool = False) -> Dict[str, Any]:
143
  """A finished run as a model repo: merged weights at the root (if present), the LoRA delta under
144
  `adapter/`, `train.log`, a `Modelfile` for Ollama, and a card built from `meta`.
145
 
@@ -149,14 +149,14 @@ def publish_generation(run_dir: Path | str, repo_id: str, *, token: Optional[str
149
  return {"ok": False, "reason": f"no run dir at {run}"}
150
  ck = run / "checkpoint"
151
  merged = run / "ollama_push" / "merged"
152
- staged: Dict[str, Path] = {}
153
  if include_merged and merged.is_dir():
154
  for f in merged.iterdir():
155
  if f.is_file():
156
  staged[f.name] = f
157
  if ck.is_dir():
158
  for f in ck.iterdir():
159
- if f.is_file() and f.name != "training_args.bin" or f.name == "training_args.bin":
160
  staged[f"adapter/{f.name}"] = f
161
  log = run / "train.log"
162
  if log.is_file():
@@ -167,21 +167,21 @@ def publish_generation(run_dir: Path | str, repo_id: str, *, token: Optional[str
167
  if not meta.get("base"):
168
  try:
169
  meta["base"] = json.loads((ck / "adapter_config.json").read_text()).get("base_model_name_or_path")
170
- except Exception: # noqa: BLE001
171
  pass
172
  card = _card(repo_id, meta)
173
  modelfile = ("# ollama create <name> -f Modelfile (from this repo's directory)\nFROM .\n"
174
  + (f'SYSTEM """{persona_system}"""\n' if persona_system else "")
175
  + 'PARAMETER temperature 0.7\nPARAMETER repeat_penalty 1.3\nPARAMETER stop "<|im_end|>"\n')
176
- plan = {"repo": repo_id, "private": private, "files": sorted(staged) + ["README.md", "Modelfile"],
177
  "bytes": sum(f.stat().st_size for f in staged.values())}
178
  if dry_run:
179
  return {"ok": True, "dry_run": True, "would_upload": plan, "card_preview": card[:400]}
180
  api, err = _api(token)
181
  if err:
182
  return err
183
- import tempfile
184
  import shutil
 
185
  with tempfile.TemporaryDirectory() as tmp:
186
  stage = Path(tmp)
187
  for rel, src in staged.items():
@@ -194,15 +194,15 @@ def publish_generation(run_dir: Path | str, repo_id: str, *, token: Optional[str
194
  api.create_repo(repo_id, repo_type="model", private=private, exist_ok=True)
195
  ci = api.upload_folder(folder_path=str(stage), repo_id=repo_id, repo_type="model",
196
  commit_message=meta.get("commit_message") or "mindXtrain: a run, published with its evidence")
197
- except Exception as e: # noqa: BLE001
198
  return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
199
  return {"ok": True, "repo": repo_id, "url": f"https://huggingface.co/{repo_id}",
200
  "commit": str(getattr(ci, "oid", ci))[:12], "uploaded": plan}
201
 
202
 
203
  # ── the corpus ────────────────────────────────────────────────────────────────
204
- def push_dataset(path: Path | str, repo_id: str, *, token: Optional[str] = None, private: bool = False,
205
- path_in_repo: str = "", manifest: Optional[Dict[str, Any]] = None) -> Dict[str, Any]:
206
  """Push a training corpus (a folder or one JSONL) with an optional manifest beside it."""
207
  api, err = _api(token)
208
  if err:
@@ -222,7 +222,7 @@ def push_dataset(path: Path | str, repo_id: str, *, token: Optional[str] = None,
222
  api.upload_file(path_or_fileobj=json.dumps(manifest, indent=1).encode(),
223
  path_in_repo=f"{path_in_repo}/MANIFEST.json".lstrip("/"),
224
  repo_id=repo_id, repo_type="dataset", commit_message="mindXtrain: corpus manifest")
225
- except Exception as e: # noqa: BLE001
226
  return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
227
  return {"ok": True, "repo": repo_id, "url": f"https://huggingface.co/datasets/{repo_id}",
228
  "commit": str(getattr(ci, "oid", ci))[:12]}
@@ -232,8 +232,8 @@ def push_dataset(path: Path | str, repo_id: str, *, token: Optional[str] = None,
232
  _GEN = re.compile(r"(?:^|/)gen(\d+)(?:/|$)")
233
 
234
 
235
- def lineage(repo_id: str, *, repo_type: str = "model", token: Optional[str] = None,
236
- local_runs: Optional[Path | str] = None) -> Dict[str, Any]:
237
  """Generations present in a repo, and — when `local_runs` is given — which local runs are not
238
  published yet. Uses `tree_paths`, so folders never break the scan."""
239
  api, err = _api(token)
@@ -241,9 +241,9 @@ def lineage(repo_id: str, *, repo_type: str = "model", token: Optional[str] = No
241
  return err
242
  try:
243
  paths = tree_paths(api, repo_id, repo_type=repo_type, recursive=True)
244
- except Exception as e: # noqa: BLE001
245
  return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:200]}"}
246
- gens: Dict[int, List[str]] = {}
247
  for p in paths:
248
  m = _GEN.search(p)
249
  if m:
@@ -262,9 +262,9 @@ def lineage(repo_id: str, *, repo_type: str = "model", token: Optional[str] = No
262
 
263
 
264
  # ── Spaces ────────────────────────────────────────────────────────────────────
265
- def push_space(folder: Path | str, space_id: str, *, token: Optional[str] = None, private: bool = True,
266
- hardware: Optional[str] = "zero-a10g", variables: Optional[Dict[str, str]] = None,
267
- secrets: Optional[Dict[str, str]] = None) -> Dict[str, Any]:
268
  """Push a Gradio folder as a Space. Existence is checked BEFORE creation (the Hub tests the
269
  ZeroGPU quota first and answers 402 on a repo that already exists), and a README
270
  `short_description` longer than 60 characters is refused by the Hub, so it is checked here."""
@@ -292,9 +292,10 @@ def push_space(folder: Path | str, space_id: str, *, token: Optional[str] = None
292
  for k, v in (secrets or {}).items():
293
  api.add_space_secret(space_id, k, v)
294
  rt = api.get_space_runtime(space_id)
295
- except Exception as e: # noqa: BLE001
296
  return {"ok": False, "space": space_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
297
  return {"ok": True, "space": space_id, "existed": exists, "stage": str(rt.stage),
 
298
  "url": f"https://huggingface.co/spaces/{space_id}",
299
  "host": "https://" + space_id.replace("/", "-").replace("_", "-").lower() + ".hf.space",
300
  "note": "free personal accounts host 2 ZeroGPU Spaces; a free org hosts none (402)"}
 
20
  import re
21
  import time
22
  from pathlib import Path
23
+ from typing import Any
24
 
25
  _HF_ENV = ("HF_TOKEN", "HUGGING_FACE_HUB_TOKEN", "HUGGINGFACEHUB_API_TOKEN")
26
 
27
 
28
+ def _token(explicit: str | None = None) -> str | None:
29
  if explicit:
30
  return explicit
31
  for k in _HF_ENV:
 
35
  return None
36
 
37
 
38
+ def _api(token: str | None = None):
39
  """(HfApi, None) or (None, {"ok": False, …}) — never raises."""
40
  try:
41
  from huggingface_hub import HfApi
42
  except ImportError:
43
+ return None, {"ok": False, "reason": "huggingface_hub not installed — `uv sync --extra hf`"}
44
  tok = _token(token)
45
  if not tok:
46
  return None, {"ok": False, "reason": f"no token — set one of {', '.join(_HF_ENV)}"}
47
  return HfApi(token=tok), None
48
 
49
 
50
+ def tree_paths(api, repo_id: str, **kw) -> list[str]:
51
  """File paths of a repo tree (folders skipped). See the module docstring: the tree's entries
52
  carry `.path`, not `.rfilename`."""
53
  return [f.path for f in api.list_repo_tree(repo_id, **kw) if type(f).__name__ != "RepoFolder"]
54
 
55
 
56
  # ── who the token is ──────────────────────────────────────────────────────────
57
+ def account(token: str | None = None) -> dict[str, Any]:
58
  """The identity behind the token, its role, the orgs it belongs to, and — separately — the
59
  namespaces it can actually write."""
60
  api, err = _api(token)
 
62
  return err
63
  try:
64
  me = api.whoami()
65
+ except Exception as e:
66
  return {"ok": False, "reason": f"whoami failed: {type(e).__name__}: {str(e)[:200]}"}
67
  auth = (me.get("auth") or {}).get("accessToken") or {}
68
  perms = auth.get("fineGrained") or {}
 
82
 
83
 
84
  # ── bases, fetched before the run rather than during it ───────────────────────
85
+ def pull_base(model_id: str, *, token: str | None = None, allow_patterns: list[str] | None = None) -> dict[str, Any]:
86
  """Download a base model to the local cache. Do this BEFORE training: a run that discovers a
87
  missing base three hours in has wasted three hours."""
88
  try:
89
  from huggingface_hub import snapshot_download
90
  except ImportError:
91
+ return {"ok": False, "reason": "huggingface_hub not installed — `uv sync --extra hf`"}
92
  t0 = time.time()
93
  try:
94
  path = snapshot_download(model_id, token=_token(token), allow_patterns=allow_patterns)
95
+ except Exception as e:
96
  return {"ok": False, "model": model_id, "reason": f"{type(e).__name__}: {str(e)[:200]}"}
97
  p = Path(path)
98
  return {"ok": True, "model": model_id, "path": str(p), "seconds": round(time.time() - t0, 1),
99
  "bytes": sum(f.stat().st_size for f in p.rglob("*") if f.is_file())}
100
 
101
 
102
+ def warm(config: Path | str, *, token: str | None = None) -> dict[str, Any]:
103
  """Pull whatever a run.yaml says it needs (the base model) so `train` starts cold-free."""
104
  try:
105
  import yaml
106
  cfg = yaml.safe_load(Path(config).read_text())
107
+ except Exception as e:
108
  return {"ok": False, "reason": f"unreadable config: {type(e).__name__}: {str(e)[:160]}"}
109
  base = ((cfg or {}).get("model") or {}).get("name")
110
  if not base:
 
113
 
114
 
115
  # ── a finished run, published ─────────────────────────────────────────────────
116
+ def _card(repo_id: str, meta: dict[str, Any]) -> str:
117
  """A model card written from the run's own numbers. No claim that is not in `meta`."""
118
  m = meta
119
  fm = {"license": m.get("license", "apache-2.0"), "library_name": "transformers",
120
+ "pipeline_tag": "text-generation", "tags": ["mindxtrain", "lora", *list(m.get("tags") or [])]}
121
  if m.get("base"):
122
  fm["base_model"] = m["base"]
123
  if m.get("dataset"):
 
137
  "the training log ships beside the weights.\n")
138
 
139
 
140
+ def publish_generation(run_dir: Path | str, repo_id: str, *, token: str | None = None, private: bool = False,
141
+ meta: dict[str, Any] | None = None, persona_system: str | None = None,
142
+ include_merged: bool = True, dry_run: bool = False) -> dict[str, Any]:
143
  """A finished run as a model repo: merged weights at the root (if present), the LoRA delta under
144
  `adapter/`, `train.log`, a `Modelfile` for Ollama, and a card built from `meta`.
145
 
 
149
  return {"ok": False, "reason": f"no run dir at {run}"}
150
  ck = run / "checkpoint"
151
  merged = run / "ollama_push" / "merged"
152
+ staged: dict[str, Path] = {}
153
  if include_merged and merged.is_dir():
154
  for f in merged.iterdir():
155
  if f.is_file():
156
  staged[f.name] = f
157
  if ck.is_dir():
158
  for f in ck.iterdir():
159
+ if (f.is_file() and f.name != "training_args.bin") or f.name == "training_args.bin":
160
  staged[f"adapter/{f.name}"] = f
161
  log = run / "train.log"
162
  if log.is_file():
 
167
  if not meta.get("base"):
168
  try:
169
  meta["base"] = json.loads((ck / "adapter_config.json").read_text()).get("base_model_name_or_path")
170
+ except Exception:
171
  pass
172
  card = _card(repo_id, meta)
173
  modelfile = ("# ollama create <name> -f Modelfile (from this repo's directory)\nFROM .\n"
174
  + (f'SYSTEM """{persona_system}"""\n' if persona_system else "")
175
  + 'PARAMETER temperature 0.7\nPARAMETER repeat_penalty 1.3\nPARAMETER stop "<|im_end|>"\n')
176
+ plan = {"repo": repo_id, "private": private, "files": [*sorted(staged), "README.md", "Modelfile"],
177
  "bytes": sum(f.stat().st_size for f in staged.values())}
178
  if dry_run:
179
  return {"ok": True, "dry_run": True, "would_upload": plan, "card_preview": card[:400]}
180
  api, err = _api(token)
181
  if err:
182
  return err
 
183
  import shutil
184
+ import tempfile
185
  with tempfile.TemporaryDirectory() as tmp:
186
  stage = Path(tmp)
187
  for rel, src in staged.items():
 
194
  api.create_repo(repo_id, repo_type="model", private=private, exist_ok=True)
195
  ci = api.upload_folder(folder_path=str(stage), repo_id=repo_id, repo_type="model",
196
  commit_message=meta.get("commit_message") or "mindXtrain: a run, published with its evidence")
197
+ except Exception as e:
198
  return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
199
  return {"ok": True, "repo": repo_id, "url": f"https://huggingface.co/{repo_id}",
200
  "commit": str(getattr(ci, "oid", ci))[:12], "uploaded": plan}
201
 
202
 
203
  # ── the corpus ────────────────────────────────────────────────────────────────
204
+ def push_dataset(path: Path | str, repo_id: str, *, token: str | None = None, private: bool = False,
205
+ path_in_repo: str = "", manifest: dict[str, Any] | None = None) -> dict[str, Any]:
206
  """Push a training corpus (a folder or one JSONL) with an optional manifest beside it."""
207
  api, err = _api(token)
208
  if err:
 
222
  api.upload_file(path_or_fileobj=json.dumps(manifest, indent=1).encode(),
223
  path_in_repo=f"{path_in_repo}/MANIFEST.json".lstrip("/"),
224
  repo_id=repo_id, repo_type="dataset", commit_message="mindXtrain: corpus manifest")
225
+ except Exception as e:
226
  return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
227
  return {"ok": True, "repo": repo_id, "url": f"https://huggingface.co/datasets/{repo_id}",
228
  "commit": str(getattr(ci, "oid", ci))[:12]}
 
232
  _GEN = re.compile(r"(?:^|/)gen(\d+)(?:/|$)")
233
 
234
 
235
+ def lineage(repo_id: str, *, repo_type: str = "model", token: str | None = None,
236
+ local_runs: Path | str | None = None) -> dict[str, Any]:
237
  """Generations present in a repo, and — when `local_runs` is given — which local runs are not
238
  published yet. Uses `tree_paths`, so folders never break the scan."""
239
  api, err = _api(token)
 
241
  return err
242
  try:
243
  paths = tree_paths(api, repo_id, repo_type=repo_type, recursive=True)
244
+ except Exception as e:
245
  return {"ok": False, "repo": repo_id, "reason": f"{type(e).__name__}: {str(e)[:200]}"}
246
+ gens: dict[int, list[str]] = {}
247
  for p in paths:
248
  m = _GEN.search(p)
249
  if m:
 
262
 
263
 
264
  # ── Spaces ────────────────────────────────────────────────────────────────────
265
+ def push_space(folder: Path | str, space_id: str, *, token: str | None = None, private: bool = True,
266
+ hardware: str | None = "zero-a10g", variables: dict[str, str] | None = None,
267
+ secrets: dict[str, str] | None = None) -> dict[str, Any]:
268
  """Push a Gradio folder as a Space. Existence is checked BEFORE creation (the Hub tests the
269
  ZeroGPU quota first and answers 402 on a repo that already exists), and a README
270
  `short_description` longer than 60 characters is refused by the Hub, so it is checked here."""
 
292
  for k, v in (secrets or {}).items():
293
  api.add_space_secret(space_id, k, v)
294
  rt = api.get_space_runtime(space_id)
295
+ except Exception as e:
296
  return {"ok": False, "space": space_id, "reason": f"{type(e).__name__}: {str(e)[:300]}"}
297
  return {"ok": True, "space": space_id, "existed": exists, "stage": str(rt.stage),
298
+ "commit": str(getattr(ci, "oid", ci))[:12],
299
  "url": f"https://huggingface.co/spaces/{space_id}",
300
  "host": "https://" + space_id.replace("/", "-").replace("_", "-").lower() + ".hf.space",
301
  "note": "free personal accounts host 2 ZeroGPU Spaces; a free org hosts none (402)"}
mindxtrain/storage/hf_hub.py CHANGED
@@ -1,6 +1,6 @@
1
  """Push checkpoint + model card to Hugging Face Hub.
2
 
3
- Lazy `import huggingface_hub` so users without `--extra chain` can still
4
  import this module. `HF_TOKEN` is read from env (or the token kwarg).
5
  """
6
 
@@ -24,7 +24,7 @@ def publish_to_hf(
24
  try:
25
  from huggingface_hub import HfApi
26
  except ImportError as exc:
27
- msg = "huggingface_hub not installed; run `uv sync --extra chain`."
28
  raise RuntimeError(msg) from exc
29
 
30
  api = HfApi(token=token or os.environ.get("HF_TOKEN"))
@@ -56,7 +56,7 @@ class HfHubProvider(StorageProvider):
56
  try:
57
  from huggingface_hub import snapshot_download
58
  except ImportError as exc:
59
- msg = "huggingface_hub not installed; run `uv sync --extra chain`."
60
  raise RuntimeError(msg) from exc
61
  # ref.uri is `https://huggingface.co/<repo_id>` — extract repo_id.
62
  repo_id = ref.uri.removeprefix("https://huggingface.co/")
 
1
  """Push checkpoint + model card to Hugging Face Hub.
2
 
3
+ Lazy `import huggingface_hub` so users without `--extra hf` can still
4
  import this module. `HF_TOKEN` is read from env (or the token kwarg).
5
  """
6
 
 
24
  try:
25
  from huggingface_hub import HfApi
26
  except ImportError as exc:
27
+ msg = "huggingface_hub not installed; run `uv sync --extra hf`."
28
  raise RuntimeError(msg) from exc
29
 
30
  api = HfApi(token=token or os.environ.get("HF_TOKEN"))
 
56
  try:
57
  from huggingface_hub import snapshot_download
58
  except ImportError as exc:
59
+ msg = "huggingface_hub not installed; run `uv sync --extra hf`."
60
  raise RuntimeError(msg) from exc
61
  # ref.uri is `https://huggingface.co/<repo_id>` — extract repo_id.
62
  repo_id = ref.uri.removeprefix("https://huggingface.co/")
mindxtrain/ui/app.py CHANGED
@@ -36,7 +36,7 @@ import gradio as gr
36
  from .metrics import RunMetrics, parse_log # noqa: F401 (parse_log re-exported for tests)
37
  from .theme import CSS, theme
38
 
39
- VERSION = "1.0.1"
40
  HOME = Path(os.environ.get("MINDXTRAIN_HOME") or Path(__file__).resolve().parents[2])
41
  RECIPES = HOME / "mindxtrain" / "train" / "recipes"
42
  TIERS = ["Basic", "Advanced", "Scientific"]
 
36
  from .metrics import RunMetrics, parse_log # noqa: F401 (parse_log re-exported for tests)
37
  from .theme import CSS, theme
38
 
39
+ VERSION = "1.0.2"
40
  HOME = Path(os.environ.get("MINDXTRAIN_HOME") or Path(__file__).resolve().parents[2])
41
  RECIPES = HOME / "mindxtrain" / "train" / "recipes"
42
  TIERS = ["Basic", "Advanced", "Scientific"]
pyproject.toml CHANGED
@@ -1,6 +1,6 @@
1
  [project]
2
  name = "mindxtrain"
3
- version = "1.0.1"
4
  description = "Production training framework for fine-tuning open-weight LLMs on AMD MI300X and serving them through an OpenAI-compatible API."
5
  requires-python = ">=3.12,<3.13"
6
  license = { text = "Apache-2.0" }
@@ -45,10 +45,14 @@ ui = [
45
  "gradio[mcp]>=5.0",
46
  "pyyaml>=6.0",
47
  ]
 
 
 
 
48
  chain = [
49
  "web3>=7.5",
50
  "py-algorand-sdk>=2.7",
51
- "huggingface-hub>=0.26",
52
  ]
53
  obs = [
54
  "opentelemetry-sdk>=1.28",
@@ -56,7 +60,7 @@ obs = [
56
  "psutil>=6.1",
57
  ]
58
  all = [
59
- "mindxtrain[ml,eval,data,serve,chain,obs,ui]",
60
  ]
61
 
62
  [build-system]
 
1
  [project]
2
  name = "mindxtrain"
3
+ version = "1.0.2"
4
  description = "Production training framework for fine-tuning open-weight LLMs on AMD MI300X and serving them through an OpenAI-compatible API."
5
  requires-python = ">=3.12,<3.13"
6
  license = { text = "Apache-2.0" }
 
45
  "gradio[mcp]>=5.0",
46
  "pyyaml>=6.0",
47
  ]
48
+ # The Hub is not a chain. It had been living inside `chain`, so reaching Hugging Face
49
+ # meant installing web3 and the Algorand SDK; `chain` still pulls it so existing
50
+ # installs keep working.
51
+ hf = ["huggingface-hub>=0.26"]
52
  chain = [
53
  "web3>=7.5",
54
  "py-algorand-sdk>=2.7",
55
+ "mindxtrain[hf]",
56
  ]
57
  obs = [
58
  "opentelemetry-sdk>=1.28",
 
60
  "psutil>=6.1",
61
  ]
62
  all = [
63
+ "mindxtrain[ml,eval,data,serve,chain,obs,ui,hf]",
64
  ]
65
 
66
  [build-system]
tests/test_hf_extension.py ADDED
@@ -0,0 +1,159 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Units for the Hugging Face extension and its CLI surface.
2
+
3
+ `huggingface_hub` is an optional extra and is NOT installed in this environment, which makes
4
+ the missing-dependency path a real test rather than a mocked one. Everything else runs against
5
+ stubs: no token, no network, no Hub.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ from typing import Any
11
+
12
+ import pytest
13
+ from typer.testing import CliRunner
14
+
15
+ from mindxtrain.cli.main import app
16
+ from mindxtrain.hf import extension as ext
17
+
18
+ runner = CliRunner()
19
+
20
+
21
+ @pytest.fixture(autouse=True)
22
+ def no_ambient_token(monkeypatch: pytest.MonkeyPatch) -> None:
23
+ """A developer's real HF_TOKEN must not change what these tests assert."""
24
+ for name in ext._HF_ENV:
25
+ monkeypatch.delenv(name, raising=False)
26
+
27
+
28
+ # --- token resolution --------------------------------------------------------------------
29
+
30
+
31
+ def test_explicit_token_beats_the_environment(monkeypatch: pytest.MonkeyPatch) -> None:
32
+ monkeypatch.setenv("HF_TOKEN", "from-env")
33
+ assert ext._token("explicit") == "explicit"
34
+
35
+
36
+ def test_environment_names_are_tried_in_order(monkeypatch: pytest.MonkeyPatch) -> None:
37
+ monkeypatch.setenv("HUGGINGFACEHUB_API_TOKEN", "third")
38
+ assert ext._token() == "third"
39
+ monkeypatch.setenv("HF_TOKEN", "first")
40
+ assert ext._token() == "first"
41
+
42
+
43
+ def test_no_token_anywhere_is_none() -> None:
44
+ assert ext._token() is None
45
+
46
+
47
+ # --- the optional dependency is reported, never raised -----------------------------------
48
+
49
+
50
+ def test_missing_dependency_is_a_result_not_an_exception() -> None:
51
+ """`huggingface_hub` is genuinely absent here — this asserts the real path."""
52
+ api, err = ext._api("a-token")
53
+ assert api is None
54
+ assert err["ok"] is False
55
+ assert "--extra hf" in err["reason"] # the extra that actually installs it
56
+
57
+
58
+ def test_pull_base_reports_the_missing_dependency() -> None:
59
+ out = ext.pull_base("HuggingFaceTB/SmolLM2-135M")
60
+ assert out["ok"] is False and "huggingface_hub" in out["reason"]
61
+
62
+
63
+ def test_no_token_names_the_variables_to_set(monkeypatch: pytest.MonkeyPatch) -> None:
64
+ """With the library present but no token, the error must say which env vars are read."""
65
+ monkeypatch.setattr(ext, "_api", ext._api) # keep the real function
66
+ fake_hub = type("m", (), {"HfApi": lambda **kw: None})
67
+ monkeypatch.setitem(__import__("sys").modules, "huggingface_hub", fake_hub)
68
+ api, err = ext._api()
69
+ assert api is None
70
+ assert "HF_TOKEN" in err["reason"]
71
+
72
+
73
+ # --- tree_paths: the RepoFolder trap -----------------------------------------------------
74
+
75
+
76
+ # The filter matches on the CLASS NAME, so these stubs must carry the Hub's exact names —
77
+ # `tree_paths` deliberately avoids importing the optional type just to check it.
78
+ class RepoFile:
79
+ def __init__(self, path: str) -> None:
80
+ self.path = path
81
+
82
+
83
+ class RepoFolder:
84
+ def __init__(self, path: str) -> None:
85
+ self.path = path
86
+
87
+
88
+ def test_tree_paths_skips_folders() -> None:
89
+ """Folders and files both carry `.path`; counting folders as files corrupts any scan."""
90
+ api = type("A", (), {"list_repo_tree": lambda self, r, **kw: [
91
+ RepoFile("gen1/adapter.safetensors"), RepoFolder("gen1"), RepoFile("README.md"),
92
+ ]})()
93
+ assert ext.tree_paths(api, "PYTHAI/x") == ["gen1/adapter.safetensors", "README.md"]
94
+
95
+
96
+ def test_tree_paths_on_a_tree_of_only_folders_is_empty() -> None:
97
+ """The bug this guards: a scan reporting "nothing on the Hub" while the repo is full."""
98
+ api = type("A", (), {"list_repo_tree": lambda self, r, **kw: [RepoFolder("gen1"), RepoFolder("gen2")]})()
99
+ assert ext.tree_paths(api, "PYTHAI/x") == []
100
+
101
+
102
+ # --- guards that must not need a network -------------------------------------------------
103
+
104
+
105
+ def test_warm_refuses_a_config_without_a_model(tmp_path) -> None: # type: ignore[no-untyped-def]
106
+ cfg = tmp_path / "run.yaml"
107
+ cfg.write_text("train:\n epochs: 1\n", encoding="utf-8")
108
+ out = ext.warm(cfg)
109
+ assert out["ok"] is False
110
+ assert "model.name" in out["reason"] or "huggingface_hub" in out["reason"]
111
+
112
+
113
+ def test_publish_refuses_a_missing_run_dir(tmp_path) -> None: # type: ignore[no-untyped-def]
114
+ out = ext.publish_generation(tmp_path / "nope", "PYTHAI/x")
115
+ assert out["ok"] is False and "no run dir" in out["reason"]
116
+
117
+
118
+ # --- the CLI surface ---------------------------------------------------------------------
119
+
120
+
121
+ def test_hf_is_registered_as_a_command_group() -> None:
122
+ result = runner.invoke(app, ["--help"])
123
+ assert "hf" in result.stdout
124
+
125
+
126
+ def test_every_hf_verb_is_reachable() -> None:
127
+ result = runner.invoke(app, ["hf", "--help"])
128
+ for verb in ("whoami", "pull", "warm", "publish", "lineage", "dataset", "space"):
129
+ assert verb in result.stdout
130
+
131
+
132
+ def test_failure_exits_nonzero_so_a_script_can_branch() -> None:
133
+ """The ascent loop warms a base and checks $? — an `ok:false` that exits 0 would be a trap."""
134
+ result = runner.invoke(app, ["hf", "whoami"])
135
+ assert result.exit_code == 1
136
+ assert "huggingface_hub" in result.stdout
137
+
138
+
139
+ def test_success_exits_zero_and_prints_the_result(monkeypatch: pytest.MonkeyPatch) -> None:
140
+ import mindxtrain.hf as hf_pkg
141
+
142
+ monkeypatch.setattr(hf_pkg, "account", lambda tok=None: {"ok": True, "name": "Gregory-L",
143
+ "can_write": ["PYTHAI"]})
144
+ result = runner.invoke(app, ["hf", "whoami"])
145
+ assert result.exit_code == 0
146
+ assert "Gregory-L" in result.stdout
147
+
148
+
149
+ def test_token_option_is_passed_through(monkeypatch: pytest.MonkeyPatch) -> None:
150
+ seen: dict[str, Any] = {}
151
+ import mindxtrain.hf as hf_pkg
152
+
153
+ def fake_account(tok=None): # type: ignore[no-untyped-def]
154
+ seen["token"] = tok
155
+ return {"ok": True}
156
+
157
+ monkeypatch.setattr(hf_pkg, "account", fake_account)
158
+ assert runner.invoke(app, ["hf", "whoami", "--token", "abc"]).exit_code == 0
159
+ assert seen["token"] == "abc"