|
Download docs/modules/create.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 10.7 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/create.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/modules/create.md
-
curl -L -o create.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/create.md
10.7 kB
| # `bankML/create.rs` — `bankml create`: derived models over pinned bases | |
| ## Summary | |
| `create.rs` is Ollama's `ollama create` over bankML's gate (O5, 0.3.5). A **derived model** is a manifest that | |
| layers configuration on a **pinned base**: a system prompt, parameters, stop strings, example messages, a licence. | |
| It never copies weights. Loading one verifies the base exactly as a pinned model is verified (the guard, then the | |
| sha256 pin of its FORK.json), then applies the layer. | |
| It exists for mindX's flow: mindXtrain merges each generation into safetensors, and mindX's `promote.py` layers a | |
| persona with `ollama create`. `bankml create` takes the same Modelfile. `FROM` a safetensors directory converts it | |
| with `convert.rs` and pins the result, so the persona becomes a verified layer over a pinned GGUF. | |
| The module holds Ollama's Modelfile parser, the manifest format, the derived-model store, the layer's application to | |
| Ollama and OpenAI requests, the HTTP handlers for `/api/create`, `/api/delete` and `/api/copy`, and the CLI. | |
| ## Technical usage | |
| ### The manifest | |
| ```text | |
| <registry>/<name>.MODEL.json {kind, bankml, name, created_at, base: {name, file, sha256, fork, path}, | |
| layer: {base_sha256, system, parameters, stop, template_sha256, license, messages, requires}, | |
| digest: "sha256:…"} | |
| ``` | |
| `digest` is the sha256 of the layer's content (the base's sha256 and the layer), not of the name or the time. The | |
| same Modelfile over the same base gives the same digest, and a copy keeps it. A manifest whose digest is not its | |
| content's is refused when read: "edited by hand, refused — create it again". A manifest whose base pin has changed | |
| is refused too. It is written atomically (a `.json.part` file, then a rename). | |
| ### The Modelfile (`parse_modelfile`, `Spec::from_modelfile`) | |
| The parser is Ollama's `parser/parser.go`, state for state: case-insensitive instructions; `#` opens a comment only at | |
| a line's start; a value runs to the end of the line; `"…"` and `"""…"""` may span lines, with no escapes. The | |
| vendored reference for the format is mindX's `docs/ollama/setup/modelfile.md`. `quote()` | |
| writes values back as Ollama's `ollama show --modelfile` does, and they round-trip. | |
| | instruction | handling | | |
| |---|---| | |
| | `FROM` | exactly one: a registry name (pinned or derived), its suffix-less alias, a pinned GGUF path, or a safetensors directory (converted to `<name>-F16.gguf` and pinned with every input's sha256) | | |
| | `SYSTEM`, `TEMPLATE` | the last wins; `TEMPLATE` only when it is the base's own chat template (bankML renders that template byte-identically to llama.cpp; a different one would change every prompt) | | |
| | `PARAMETER` | one of `PARAMS` (below) or `stop`; the last value wins, `stop` accumulates | | |
| | `MESSAGE role content` | role `system`, `user` or `assistant`; recorded and applied, as Ollama applies them | | |
| | `LICENSE`, `REQUIRES` | recorded | | |
| | `ADAPTER` | refused | | |
| `FROM` a derived model inherits its layer: new values win, `stop` is replaced as a list, licences add up. A relative | |
| `FROM ./x` on the CLI is relative to the Modelfile. | |
| ### `PARAMS` | |
| ```rust | |
| pub const PARAMS: [(&str, bool); 13] = [("temperature", false), ("top_k", true), ("top_p", false), ("min_p", false), | |
| ("seed", true), ("num_ctx", true), ("num_predict", true), ("repeat_penalty", false), ("repeat_last_n", true), | |
| ("presence_penalty", false), ("frequency_penalty", false), ("typical_p", false), ("min_keep", true)]; | |
| ``` | |
| The `bool` says whether the value is an integer; the order is the order a manifest keeps. The four penalties joined | |
| in 0.3.6 (O2); `typical_p` and its `min_keep` in 0.3.7, when the engine came to reproduce typical-p. Floats are parsed at 32 bits, as Ollama parses them. Range checks at create: | |
| `num_ctx` ≥ 1; `num_predict` ≥ −2; `repeat_penalty` > 0; `frequency_penalty` and `presence_penalty` any finite value | |
| (negative too, as llama.cpp takes them); `top_k` and `seed` unchecked; the rest ≥ 0. A `PARAMETER` outside `PARAMS` | |
| is refused with one of three reasons: a sampler not reproduced (`tfs_z`, `mirostat`, `mirostat_eta`, | |
| `mirostat_tau`, `penalize_newline`); a resource option that belongs to `bankml serve`; or not a parameter bankML | |
| knows. | |
| ### Applying the layer (Ollama's `server/routes.go`) | |
| - `apply_ollama(&self, req, chat)`: the layer's parameters become defaults under the request's `options`, overridden | |
| key by key (`stop` as a whole list). For chat, the model's `MESSAGE`s go before the request's, and its `SYSTEM` goes | |
| first unless the request's first message is a system message. For a non-raw generate, the conversation goes in | |
| `bankml_messages`: the request's `system` if any, else the model's, then the `MESSAGE`s, then the prompt. A `raw` | |
| prompt gets neither. | |
| - `apply_openai(&self, req)`: the same message rule (Ollama's OpenAI endpoint goes through the same chat handler); the parameters fill absent top-level fields (`num_predict` as | |
| `max_tokens`); `num_ctx` is not a field but fits the conversation (`num_ctx()`). | |
| ### Resolution | |
| `resolve(rs, model)` tries a pin's exact name first, then a derived model's name, then the registry's suffix-less | |
| alias (0.3.5). So promote.py's re-create in place (`FROM mindx-gen39` as `mindx-gen39`, where `mindx-gen39` is the | |
| alias of the one pin `mindx-gen39-f16`) answers as itself with the layer, and `mindx-gen39-f16` gives the pin. `FROM` | |
| resolves the suffix-less alias the same way, and only when exactly one pin has it. A derived model loads through its base's entry, so it and its base share one residency. | |
| ### HTTP (needs `serve --native --registry`) | |
| ```sh | |
| curl -s $B/api/create -H "$J" -d '{"model": "mindx-persona", "from": "mindx-gen39", "system": "You are mindX.", "parameters": {"temperature": 0.7}}' | |
| curl -s $B/api/create -H "$J" -d '{"model": "p", "modelfile": "FROM mindx-gen39\nPARAMETER repeat_penalty 1.3"}' | |
| curl -s $B/api/copy -H "$J" -d '{"source": "mindx-persona", "destination": "mindx-persona-b"}' | |
| curl -s -X DELETE $B/api/delete -H "$J" -d '{"model": "mindx-persona-b"}' | |
| ``` | |
| - `/api/create`: `{model, modelfile}` or Ollama's structured `{model, from, system, template, license, parameters, | |
| messages}`; status lines streamed as NDJSON unless `stream: false`. Without `--registry`: 400, nowhere to write. | |
| - `/api/delete`: derived models only; a pin is refused ("bankml does not delete pins over HTTP"). | |
| - `/api/copy`: the same content and digest under a new name; a pin's copy is an empty layer over it. | |
| - `/api/show` of a derived model (`show_json`): the reconstructed Modelfile, `parameters`, `system`, `license`, | |
| `messages`, `details.parent_model`, and `bankml.layer` with the digest and the base's sha256 and pin. | |
| ### CLI | |
| `bankml create NAME -f Modelfile [--registry DIR] [--models DIR]` prints Ollama's status lines and exits 0, or | |
| `bankml create: refuse: …` and 2. The registry defaults to `$BANKML_FORKS`, else `~/.local/share/bankml/forks`. | |
| ## How it is verified | |
| - Unit tests: `modelfile_parses_as_ollama` (mindX's own Modelfiles, comments, CRLF, BOM, quoting, errors), | |
| `quote_round_trips_through_the_parser`, `spec_keeps_what_is_reproduced_and_refuses_the_rest`, | |
| `manifest_digest_is_the_content`, `the_layer_applies_as_ollama_applies_it`, `create_layers_over_a_pinned_base`, | |
| `http_create_show_copy_delete_round_trip`. | |
| - `testing/cli.rs`: `create_writes_a_layer_over_a_pin`. | |
| - `oracle_persona_layer` (`#[ignore]`, gate; `testing/persona_oracle.py`): promote.py's persona Modelfile created two | |
| ways, `FROM` the merged safetensors directory and `FROM mindx-gen39` in place. Asked the user's turns alone, each | |
| must give llama-server b11192's tokens for the same GGUF given the persona as the system message. CHANGELOG 0.3.5: | |
| 27 / 27 each. Live: `persona_oracle_live` through `/api/chat`, `/api/generate`, `/v1` and `/api/ps`. | |
| ## Advantages and efficiency | |
| - **No weights copied.** A layer is a small JSON file over a pinned GGUF; its digest is content-addressed, as | |
| Ollama's content addressing does. Every load of a derived model verifies its base (the guard, then the pin). | |
| - **mindX's Modelfiles unchanged.** The parser is Ollama's state machine, so promote.py's files create the same layer, | |
| and `/api/show` gives back a Modelfile that recreates it. | |
| - **Efficient.** Tamper detection is one sha256 over a short text. A derived model and its base share one load. A | |
| conversion never overwrites an existing pin. | |
| - **Rust practice.** No external crates; atomic manifest writes; `RwLock` poisoning recovered; errors carry the | |
| reason as `Result<_, String>`. | |
| - **Next** (docs/OLLAMA.md, docs/TODO.md): `ADAPTER` (LoRA merging) is a later O-phase; `promote.py --to bankml` in | |
| mindX and a bankml backend for mindXtrain (O5, proposed). | |
| ## Limitations | |
| - `ADAPTER`: "bankML does not merge LoRA adapters yet (a later O-phase); merge it into the base weights … then FROM | |
| the merged safetensors directory". | |
| - `TEMPLATE` other than the base's own: Ollama's are Go templates; bankML renders the GGUF's Jinja byte-identically. | |
| - `/api/create` `files`, `adapters`, `quantize`: refused, each with its reason. | |
| - More than one `FROM`; a name a pinned file already has; a `FROM` file not pinned in the registry directory. | |
| - Model names: letters, digits, `.`, `_`, `-`, at most 128, no tag but `latest`, not ending in `.gguf`. | |
| - A layer `num_ctx` above the server's `--ctx` is accepted at create (with a warning) and refused per request. | |
| - The 0.3.7 samplers that are not Ollama options (DRY, XTC, top-n-σ, dynamic temperature) are not Modelfile | |
| parameters; a request gives them through `/v1`. | |
| ## Design notes | |
| - `bankml create` is phase O5 of [../OLLAMA.md](../OLLAMA.md), first built in 0.3.5; since 0.3.5 a conversion under | |
| `create` never overwrites an existing pin, nor its FORK.json wherever the file is. | |
| - `is_print` approximates Go's `strconv.IsPrint` (no control characters, no space other than U+0020, no format | |
| characters), which the parser uses as Ollama's does. | |
| - A create changes what requests resolve to, so `POST /api/create` takes the engine's run lock, as a model load does. | |
| - Parameter floats are stored at 32 bits, as Ollama parses them, and written in the shortest form that reads back. | |
| ## See also | |
| - [../usage.md, derived models and conversion](../usage.md#derived-models-and-conversion-bankml-create-bankml-convert-o5-035) · | |
| [../OLLAMA.md, what O5's first cut built](../OLLAMA.md#what-o5s-first-cut-built-035) · [../oracles.md](../oracles.md) · | |
| [../TODO.md](../TODO.md) | |
| - [convert.md](convert.md) · [ollama.md](ollama.md) · [native.md](native.md) · [serve.md](serve.md) · [main.md](main.md) | |