File size: 10,293 Bytes
28c70af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
# `bankML/convert.rs` — `bankml convert`: safetensors to GGUF F16, byte-identical to llama.cpp

## Summary

`convert.rs` turns a Hugging Face safetensors directory into a GGUF F16 file for the **Llama architecture as
SmolLM2-135M and mindX's `mindx-genN` use it**. It is written to be **byte-identical** to llama.cpp b11192's
`convert_hf_to_gguf.py DIR --outtype f16`, read from the tag's `conversion/{base,llama}.py` and `gguf-py`, in its
order of operations.

It exists because mindXtrain (github.com/Professor-Codephreak/mindXtrain, continued at
huggingface.co/PYTHAI/mindXtrain; Apache-2.0) merges each generation's LoRA into safetensors (`ollama_push/merged`)
and serves it through Ollama with `serve --to ollama`; mindX's `promote.py` then layers the persona. This converter and `bankml create` take that merged directory instead: one pinned GGUF, with the
persona as a verified layer ([create.md](create.md)). `bankml create … FROM <directory>` calls `convert()` and pins
the result; `bankml convert` runs it alone.

What would make llama.cpp write something this does not reproduce is refused, with the reason. Nothing is
approximated.

## Technical usage

```sh
bankml convert DIR -o mindx-gen39-F16.gguf --fork mindx-gen39-F16.gguf.FORK.json \
    --source "PYTHAI/mindXascension@4bd31b9d weights/gen39"
```

| flag | meaning |
|---|---|
| `-o OUT.gguf` / `--outfile` | required |
| `--outtype f16` | the only type accepted; another is refused (exit 2) |
| `--model-name NAME` | `general.name` instead of the heuristics' (llama.cpp's flag of the same meaning) |
| `--fork FORK.json` | also write the pin: the output's sha256 and size, and every input file's |
| `--source SRC` | where the directory came from, recorded in the pin (default: the directory's canonical path) |
| `--ignore-model-card` | convert although a `README.md` is present; the result then differs from llama.cpp's |

On success it prints `<sha256>  <out>` on stdout and a summary (name, tensors, metadata keys, weights, bytes) on
stderr. A refusal prints `bankml convert: refuse: …` and exits 2.

### API

```rust
pub fn convert(dir: &Path, out: &Path, opt: &Options) -> Result<Report, String>
pub struct Options { pub model_name: Option<String>, pub ignore_model_card: bool }
pub fn fork_json(r: &Report, source: &str) -> String
```

`Report` carries `out`, `sha256`, `bytes`, `tensors`, `kv`, `params`, `name`, and `inputs` (each input's name, size
and sha256).

The pieces the oracles and `create` reach directly:

```rust
pub const KNOWN_TOKENIZERS: [(&str, &str, &str, &str); 1]          // (pre, vocab_merges_sha256, chkhsh, where measured)
pub fn vocab_merges_sha256(tokens: &[String], merges: &[String]) -> String
pub fn id_to_title(s: &str) -> String                             // gguf-py's metadata.py heuristics
pub fn model_id_components(model_id: &str, total_params: i64) -> IdParts   // (full name, org, basename, finetune, version, size label)
pub fn size_label(n: u64) -> String
pub enum Kv { U32(u32), I32(i32), F32(f32), Bool(bool), Str(String), Strs(Vec<String>), I32s(Vec<i32>) }
pub struct KvList(pub Vec<(String, Kv)>)                           // llama.cpp's kv_data: a key set twice keeps its first position
```

### What it writes

- **Metadata, in llama.cpp's order:** `general.architecture`, `general.type`, the sampling defaults of
  `generation_config.json` (`general.sampling.*`), the name heuristics of `gguf-py/metadata.py` over `_name_or_path`
  and then the directory's name (`general.name`, `organization`, `finetune`, `basename`, `version`, `size_label`, the
  size label otherwise from the weights counted), `TextModel.set_gguf_parameters` with the
  defaults `AutoConfig(LlamaConfig).to_dict()` supplies (`head_dim`, `rope_theta` 10000, `rms_norm_eps` 1e-6, the kv
  heads), `file_type` 1, Llama's `vocab_size` and `rope.dimension_count`, `quantization_version` 2, then the tokenizer.
- **Tensors:** every safetensors part in name order (gguf-py's reader, `SafetensorsLocal`, sorts by name); HF names
  mapped to GGUF names (`TensorNameMap` for `MODEL_ARCH.LLAMA`); the Q and K rows permuted from HF's rotate-half
  layout back to GGML's pairs (`LlamaModel.permute`); 1-D tensors and `*_norm.weight` in F32, every other matrix in
  F16 by round-to-nearest-even (numpy's `astype(float16)`); BF16 read exactly as F32 first; tied embeddings write no
  `output.weight`.
- **Tokenizer:** the `gpt2` byte-level BPE path (`_set_vocab_gpt2` with `SpecialVocab(load_merges=True)`, inside
  Llama's `set_vocab`): tokens in id order and their types, merges, the special token ids (`tokenizer_config.json`,
  then `config.json`), `add_*_token`, the
  chat template (`tokenizer_config.json` or `chat_template.jinja`), and Llama's `add_bos_token = false` for
  49,152-token vocabularies.
- **Layout:** GGUF v3, little-endian, alignment 32, tensor data in the order the tensors were added. The file is written to `OUT.gguf.part`, hashed as it is written,
  then renamed.

### `general.name` comes from the directory

llama.cpp names the model from the directory's name, and so does bankML. The bytes, and so the pin, depend on it.
CHANGELOG 0.3.5: gen39 converted from a directory named `merged` is `bb41f62d…`; from `mindx-gen39` it is the pin
`6b64c748…`. A directory called `merged` gives `general.name` `Merged`, as llama.cpp's would. `--model-name` overrides.

### Tokenizers

llama.cpp names a BPE pre-tokenizer by hashing the token ids its Python tokenizer gives one fixed text (`chkhsh`).
bankML has no Python, so `KNOWN_TOKENIZERS` keys on the sha256 of the vocabulary and merges
(`vocab_merges_sha256`: every token in id order, each followed by `\n`, a blank line, then every merge as `a b\n`)
plus the pre-tokenizer's shape, and records the `chkhsh` each entry gave, measured with
`testing/convert_oracle.py --chkhsh`. A tokenizer whose pre-tokenizer llama.cpp would not name `smollm` is refused. Today it has one
entry, `smollm` (SmolLM2-135M-Instruct and mindXascension gen39). gen39's tokenizer lost SmolLM2's `Digits`
pre-tokenizer when transformers 5.8 re-saved it; llama.cpp still names it `smollm`, and both shapes are accepted.

## How it is verified

- Unit tests: `name_heuristics_as_gguf_py`, `permute_matches_numpy_reshape_swapaxes`,
  `kv_set_keeps_the_first_position`, `regexes_as_gguf_py`.
- `oracle_convert_b11192` (`#[ignore]`, gate): `BANKML_CONVERT_DIR` against llama.cpp's GGUF in
  `BANKML_CONVERT_ORACLE`, sha256 for sha256. The gate runs it for each directory under `.models/convert/`.
- `oracle_name_heuristics` (gate): gguf-py's heuristics, written by `testing/convert_oracle.py --names`; 168 of 168
  Hub-style ids agree (CHANGELOG 0.3.5).
- `testing/convert_oracle.py`: `--record` runs llama.cpp's script, `--bankml` compares, `--compare` diffs two GGUFs
  key by key and tensor by tensor, `--chkhsh` measures a tokenizer for `KNOWN_TOKENIZERS`.
- `testing/cli.rs`: `convert_refuses_with_the_reason`.
- Recorded results (docs/OLLAMA.md): SmolLM2-135M-Instruct @ `12fd25f7` → `e9aba089…`, mindXascension gen39 @
  `4bd31b9d` → `6b64c748…`, the same as llama.cpp b11192. Run directly in 0.3.5 against `convert_hf_to_gguf.py`
  (torch 2.14.1+cpu, transformers 5.18.0) on fresh downloads: gen39 `6b64c748…` (the pin), SmolLM2 `ec30a679…` from
  both converters. SmolLM2's differs from its pin only because its directory had another name (see above).

## Advantages and efficiency

- **No Python, no torch.** llama.cpp's converter is a Python script that needs torch (usage.md §5); `bankml convert`
  is part of the zero-dependency binary and gives the same bytes.
- **Pinned at birth.** `--fork` (and `create`) write a FORK.json with the output's sha256 and every input's, so the
  result is verified like any imported model.
- **Efficient.** The safetensors are memory-mapped read-only; the output goes through a 1 MiB buffered writer that
  hashes as it writes, so no second pass is needed for the pin. docs/OLLAMA.md records a conversion at 2.6 s.
- **Rust practice.** Every unsupported case is a `Result` error with the reason; the output is written to a `.part`
  file and renamed, so a failed run leaves no partial GGUF under the final name.
- **Next:** docs/TODO.md (0.6.0) lists Llama 3.x (rope factors, the `llama3` pre-tokenizer, its template) as still
  open and refused today; the converter likewise refuses rope scaling and unknown pre-tokenizers.

## Limitations

Refused, with the reason in the source:

- Any architecture other than `LlamaForCausalLM`: "use llama.cpp's convert_hf_to_gguf.py for the rest".
- `config.json` with `text_config`, `quantization_config`, `num_local_experts` or `draft_vocab_size`;
  `is_causal: false`.
- Rope scaling (a `rope_type` other than `default`); attention or MLP biases.
- A `README.md` model card, unless `--ignore-model-card`: llama.cpp would copy its licence, tags, languages, base
  models and datasets in.
- SentencePiece (`tokenizer.model`) and Llama-HF BPE vocabularies; a non-BPE `tokenizer.json`; a normalizer; most
  post-processors; non-special added tokens; named or listed chat templates; any tokenizer not in `KNOWN_TOKENIZERS`.
- Tensor dtypes other than BF16, F16 and F32; tensor names it cannot map.
- Output types other than F16.
- A conversion under `create` never overwrites an existing pin.

## Design notes

- `bankml convert` and `bankml create` are phase O5 of [../OLLAMA.md](../OLLAMA.md).
- The name heuristics (`id_to_title`, `model_id_components`, `size_label` and the regular expressions behind them)
  are ported line for line from `gguf-py/metadata.py`; Python's `str.islower()`, `str.title()` and half-to-even
  `round()` are reproduced, because any difference changes `general.name` and so the bytes.
- `KvList` mirrors llama.cpp's `kv_data` dict: a key set twice keeps its first position and takes the new value;
  `add_string` skips empty strings.

## See also

- [../usage.md, derived models and conversion](../usage.md#derived-models-and-conversion-bankml-create-bankml-convert-o5-035) ·
  [../OLLAMA.md, what O5's first cut built](../OLLAMA.md#what-o5s-first-cut-built-035) · [../oracles.md](../oracles.md) ·
  [../TODO.md](../TODO.md)
- [create.md](create.md) · [gguf.md](gguf.md) · [tokenizer.md](tokenizer.md) · [main.md](main.md)