|
Download docs/modules/train.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 9.26 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/train.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/modules/train.md
-
curl -L -o train.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/train.md
9.26 kB
| # `bankML/train/` β mindXtrain in Rust, one verified stage at a time | |
| ## Summary | |
| `bankML/train/` is bankML's training module. It ports mindXtrain (github.com/Professor-Codephreak/mindXtrain, | |
| continued at huggingface.co/PYTHAI/mindXtrain; Apache-2.0), which imprints a persona onto a model and proves it with a | |
| recall gate: | |
| ``` | |
| author (persona + exchanges β a chat-JSONL script) β imprint (LoRA SFT) β probe (the same inquiries before and | |
| after) β score (did the voice move toward the persona, and did the utterances change) β classroom β boardroom | |
| ``` | |
| Each stage reproduces mindXtrain's Python exactly, checked by an oracle that runs mindXtrain's own functions | |
| (`testing/train_oracle.py`), before anything new is built on it. Two stages are ported: **author** (`script.rs`, from | |
| `mindxtrain/data/scripts.py`) and **score** (`imprint.rs`, the lexical path of `mindxtrain/eval/imprint.py`). The | |
| module is a library API (`bankml::train`); it has no CLI command yet. It is exercised by its unit tests and its two | |
| oracles in the release gate. | |
| ## Technical usage | |
| ### `mod.rs` β the stage registry | |
| ```rust | |
| pub const STAGES: &[(&str, &str, &str)] | |
| pub fn py_json_str(s: &str, out: &mut String) | |
| ``` | |
| - `STAGES` lists each stage of the proof loop as (stage, module, status). Today: `author` β `script` (ported), | |
| `imprint` (not yet: LoRA SFT needs a backward pass), `probe` (not yet: needs adapter loading in the forward pass; | |
| the Llama graph runs since 0.3.4), `score` β `imprint` (ported), `classroom` and `boardroom` (not yet). A new stage is a | |
| new module and one line there. The probe's status string predates 0.3.4: the forward pass has played the Llama | |
| architecture since then (SmolLM2, `mindx-genN`; [forward.md](forward.md)), so what the probe still lacks is | |
| adapter loading and transformers' decoding rules (TODO 0.7.0). | |
| - `py_json_str` writes a JSON string as Python's `json.dumps(β¦, ensure_ascii=False)` does: `\"`, `\\`, `\n`, `\r`, | |
| `\t`, `\b`, `\f` escaped, other control characters as `\u00XX`, everything else as is. | |
| ### `script.rs` β the author stage | |
| ```rust | |
| pub struct Persona { pub name: String, pub system_prompt: String, pub voice_examples: Vec<String> } | |
| pub struct Exchange { pub user: String, pub assistant: String } | |
| pub fn persona_from_json(raw: &Json) -> Result<Persona, String> | |
| pub fn system_prompt(p: &Persona) -> String | |
| pub fn script_jsonl(p: &Persona, exchanges: &[Exchange], seed_voice: bool) -> String | |
| pub fn training_params(rows: usize) -> (u32, u32, u32) | |
| ``` | |
| - `persona_from_json` is `persona_from_dict`. It reads the same keys, clean-room, from any persona JSON object: | |
| - the name is the first truthy value of `name`, `persona`, `id`, `title` (default `actor`); | |
| - the system prompt is the first truthy value of `system_prompt`, `system`, `description`, `bio`, `summary`, | |
| `prompt`; | |
| - the voice examples are collected from `voice_examples`, `examples`, `utterances`, `samples`, `voice`: a list adds | |
| its strings, numbers and booleans as Python's `str()` writes them, and a string adds itself. | |
| - `system_prompt` is `persona_system_prompt`: the persona's own prompt stripped as Python's `str.strip()` strips, or | |
| `You are <name>. Stay in character and answer in your own voice.` | |
| - `script_jsonl` is `build_script_rows` + `write_script_jsonl`: one row per exchange, then, with `seed_voice`, one | |
| row per voice example answering `Say something as <name>.`. Each line is byte-identical to Python's | |
| `json.dumps(row, ensure_ascii=False)`. | |
| - `training_params` is `derive_training_params`: (epochs, grad_accum, per_device) by row count β 24 epochs up to 8 | |
| rows, 16 up to 32, 8 up to 128, 4 with grad_accum 2 up to 512, then 2 with grad_accum 4. Small scripts must overfit | |
| to imprint. | |
| ```rust | |
| use bankml::serve::Json; | |
| use bankml::train::script::{persona_from_json, script_jsonl, training_params, Exchange}; | |
| let p = persona_from_json(&Json::parse(r#"{"name": "Savante", "voice": "Verification beats permission."}"#).unwrap())?; | |
| let jsonl = script_jsonl(&p, &[Exchange { user: "Who are you?".into(), assistant: "Savante.".into() }], true); | |
| let (epochs, grad_accum, per_device) = training_params(jsonl.lines().count()); | |
| ``` | |
| ### `imprint.rs` β the score stage | |
| ```rust | |
| pub fn default_inquiries(name: &str) -> Vec<String> | |
| pub fn jaccard(a: &str, b: &str) -> f64 | |
| pub fn round4(x: f64) -> f64 | |
| pub struct ImprintReport { pub before_voice: f64, pub after_voice: f64, pub imprint_delta: f64, pub shift: f64, | |
| pub method: &'static str, pub imprinted: bool } | |
| pub fn score(before: &[String], after: &[String], baseline: &[String]) -> ImprintReport | |
| ``` | |
| - `default_inquiries` returns mindXtrain's five persona-agnostic recall probes. | |
| - Tokens are `[a-z0-9']+` of the Unicode-lowercased text; similarity is token Jaccard (0 when either side has no | |
| tokens). | |
| - The voice score is the mean, over utterances, of the best similarity to any voice example. The shift is the mean | |
| `1 β similarity(before, after)`. Means are summed left to right, then divided once, as Python's | |
| `sum(xs) / len(xs)`. | |
| - Every figure is rounded to 4 decimals as Python's `round` rounds (`round4`: the exact binary value, half to even). | |
| - The verdict is `imprinted = delta > 0 and shift > 0`, on the unrounded values. `method` is `lexical`, or `none` | |
| when there are no utterances or no voice examples. | |
| ## How it is verified | |
| `testing/train_oracle.py` runs mindXtrain's own Python with mindXtrain's interpreter and writes the cases to | |
| `.models/oracle-train/`: | |
| ```sh | |
| ~/mindxtrain/.venv/bin/python testing/train_oracle.py [~/mindxtrain] | |
| ``` | |
| - **`oracle_train_script`** (`script.rs`, in the release gate) reads `script.jsonl`: every persona in mindX and | |
| cryptoAGI plus edge cases (nameless, whitespace prompt, control characters and quotes, falsy name, a number as | |
| name), each with and without `seed_voice`. It requires the name, the system prompt, the JSONL text byte for byte, | |
| and the epochs. **84 of 84 scripts byte-identical** to `scripts.py` ([oracles.md](../oracles.md), 0.2.13). | |
| - **`oracle_train_imprint`** (`imprint.rs`, in the release gate) reads `imprint.jsonl`: 3,000 randomized | |
| before/after/baseline sets, including empty utterances and non-ASCII words. It requires every field of the report. | |
| **3,000 of 3,000 reports identical** to `score_imprint`. | |
| - Unit tests: `a_nameless_persona_gets_the_synthesised_prompt` and `jaccard_and_rounding_as_python` (including | |
| `round4(0.03125) == 0.0312`, a tie at the fourth decimal). | |
| ## Advantages and efficiency | |
| - **Identical, not similar.** Each stage is proved against mindXtrain's own functions on real personas and thousands | |
| of random cases before the next is built. A Rust stage can replace the Python one without changing a verdict. | |
| - **No Python, no packages.** The author and score stages run inside bankML with no interpreter and no external | |
| crate. The JSON is bankML's own (`serve::Json`), and Python's string escaping, truthiness, `str()`, `strip()` and | |
| `round` are reproduced explicitly where they matter to the bytes. | |
| - **Lexical path only, on purpose.** mindXtrain prefers sentence-transformer cosine when that package is installed; | |
| the dependency-free lexical path is the one ported, so the score needs no model. | |
| - **Rust practice visible in the code.** No `unsafe` in the module. Malformed input returns `Err` with a reason | |
| (`a persona is a JSON object`, `persona field "name" is not text`). Sets are `BTreeSet`, so token sets are | |
| deterministic. The oracles fail closed: every case must match. The toolchain is pinned to Rust 1.99.0 in | |
| `rust-toolchain.toml`. | |
| - **Where it goes next** ([TODO.md](../TODO.md), 0.7.0): the probe stage (PEFT adapters loaded from safetensors and | |
| merged, then recall probing with transformers' decoding rules, against mindXtrain's `probe_recall`); the classroom | |
| and boardroom verdicts, the dojo tie-break, the feedback ledger and the receipts; **LoRA training on the CPU** (a | |
| backward pass and AdamW for the 135Mβ0.6B imprint recipe, gradients checked against PyTorch f32); then one | |
| generation end to end in Rust: author β imprint β probe β score β verdict β GGUF. | |
| ## Limitations | |
| - Only the author and score stages are ported. Imprint (LoRA SFT), probe, classroom and boardroom are not. | |
| - The score stage ports only the lexical path; the sentence-transformer cosine path is not ported. | |
| - A persona's identity fields are expected to be strings. A number in a voice list is written as Python writes an | |
| int, so a float literal such as `3.0` would differ, because bankML's JSON keeps numbers as f64. | |
| - No CLI command; the module is reached through the library. | |
| - The probe waits for adapter loading and transformers' decoding rules (TODO 0.7.0). | |
| ## See also | |
| - [oracles.md](../oracles.md) β 0.2.13, the mindXtrain oracles | |
| - [TODO.md](../TODO.md) β 0.7.0, mindXtrain in Rust end to end | |
| - [usage.md](../usage.md) β testing and the release gate | |
| - [forward.md](forward.md) β the forward pass the probe stage will run on | |
| - [convert.md](convert.md), [create.md](create.md) β `bankml convert` and `bankml create` (0.3.5), which replace | |
| mindXtrain's `serve --to ollama` step: merged safetensors to a pinned GGUF and a persona layer over it | |