|
Download README.md from nirca/nirca-mini: direct link, hf CLI and curl.
- Browser
- Download file 6.69 kB
-
https://huggingface.co/nirca/nirca-mini/resolve/main/README.md
- Command line
-
hf download hf://nirca/nirca-mini/README.md
-
curl -L -o README.md https://huggingface.co/nirca/nirca-mini/resolve/main/README.md
6.69 kB
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| tags: | |
| - ternary | |
| - mixture-of-experts | |
| - byte-level | |
| - chatml | |
| - webgpu | |
| - from-scratch | |
| datasets: | |
| - HuggingFaceTB/smoltalk | |
| - open-thoughts/OpenThoughts-114k | |
| # Nirca Mini | |
| Nirca Mini is a small chat model trained from scratch, with ternary weights: | |
| its linear layers' weights are -1, 0 or +1, with a scale and a bias for each | |
| block of 128 weights. It reads and writes bytes, so it needs no tokenizer. It | |
| is published as it trains: each new checkpoint replaces `latest/`. | |
| - **Small:** about 109 million parameters. | |
| - Trained from scratch in ternary, with quantization-aware training from almost the very start; the ternary weights were learned in training, not quantized after. | |
| Nirca was inspired by [mini-AGI](https://github.com/volotat/mini-AGI), and its architecture is | |
| deliberately derived from it: Nirca began as a port of mini-AGI to MLX. | |
| ➡️ **Try it in your browser (WebGPU):** https://tg-techie-agents.github.io/nirca-mini-webgpu/mini/ | |
| This card is written by program from the checkpoint's own record, at step | |
| 72,242 (2,074,917,148 tokens seen). | |
| ## Model details | |
| - **Type:** decoder-only language model, chat-tuned, with latent recurrence and | |
| a mixture of experts | |
| - **Parameters:** about 109 million | |
| - **Weights:** ternary (-1, 0, +1) with a scale and a bias per block of 128 | |
| (MLX's 2-bit affine layout); norms and a few small layers in floating point | |
| - **Vocabulary:** 265 tokens: the 256 byte values and 9 special | |
| tokens | |
| - **Context length:** 2,048 bytes | |
| - **Width:** 512, with 8 attention heads and rotary position embeddings | |
| - **Experts:** 32 in one shared pool, 8 active per token | |
| - **Recurrence:** one block repeated with learned halting, 1 to 24 passes | |
| (a mean of 8 in training) | |
| - **Chat format:** ChatML (`<|im_start|>` / `<|im_end|>`) | |
| - **Language:** English | |
| ## Architecture | |
| What it keeps from [mini-AGI](https://github.com/volotat/mini-AGI), and what it changes: | |
| - **Kept:** bytes, no tokenizer; latent recurrence; one shared pool of experts | |
| every pass routes into; a halting head; first pretraining text mini-AGI's | |
| corpus. | |
| - **Changed:** ternary weights, trained quantization-aware from step 14,655 of its lineage (its format before that is not recorded); | |
| tuned on chats; chat format: ChatML instead of mini-AGI's own tags; | |
| a fixed 32-expert pool where mini-AGI grows and prunes its own. | |
|  | |
| The two dense layers run once. The recurrent block then runs again and again, | |
| each pass routing into the same pool of experts, and a halting head decides | |
| per byte when another pass would not change the answer (after PonderNet, Banino | |
| et al. 2021). Each pass's prediction is weighted by the chance of halting there. | |
| ## How to use | |
| The [demo](https://tg-techie-agents.github.io/nirca-mini-webgpu/mini/) runs the model in the browser on WebGPU. A page or program | |
| loads a checkpoint folder by URL: | |
| - `latest.json`: the newest checkpoint, its step, tokens seen, shape (`cfg`), | |
| and each file's URL, size and sha256. | |
| - `latest/`: the newest checkpoint's files. | |
| - `checkpoints/<tokens>k/`: every published checkpoint, kept, named by tokens | |
| seen in thousands. | |
| Each folder holds `model.safetensors` (the packed ternary weights) and | |
| `config.json` (the model's shape and format). Prompts use ChatML: | |
| ``` | |
| <|im_start|>user | |
| Hello!<|im_end|> | |
| <|im_start|>assistant | |
| ``` | |
| ## Training data | |
| Across its training, this checkpoint and the runs it was continued from saw: | |
| - An open pretraining corpus of text, in the earlier runs this one was continued from. | |
| - Open chat datasets. | |
| - An open reasoning dataset. | |
| - Synthetic chats from open-weight models, written to give the model its character. | |
| - A small set of private chats, with names, places, contact details and other personal details replaced by placeholders before training. | |
| - Plus 2 datasets not yet public. | |
| Its training data includes text under CC BY 4.0 and CC BY-NC-SA 4.0. | |
| ## Data and training workflow | |
| Each step is listed because the run's records show it was done: | |
| 1. **ChatML.** Every conversation is rendered in ChatML (`<|im_start|>role` … `<|im_end|>`). | |
| 2. **Splits by content.** Before training, conversations are split into training, validation and test by their content, so one question is never in two splits, and the held-out splits are kept apart from the training machine. | |
| 3. **Length filter.** Chat sets are cut to conversations whose whole text fits 2,048 bytes, the model's context. | |
| 4. **Anonymisation.** In the private chats, names, places, contact details and other personal details are replaced by typed placeholders, and the result is scanned again before training. | |
| 5. **Identifier check.** Synthetic chats that name a known person's identifier are dropped. | |
| 6. **Decontamination.** A training conversation that shares an exact turn or an 8-gram with a validation or test set is left out (after PaLM's 8-gram overlap test). | |
| 7. **Loss on the assistant's turns only**, from step 61,181: the model learns to answer, not to write the user's turns. | |
| 8. **Quantization-aware training in ternary**, from step 14,655 (29,497,375 tokens) on; the steps before it record no weight format. The published weights are the trained ones, not quantized after training. | |
| ## Training procedure | |
| - **Steps:** 72,242 | |
| - **Tokens seen:** 2,074,917,148 (one token is one byte) | |
| - **Quantization-aware training:** from step 14,655 of 72,242, so for 99% of its 2,074,917,148 tokens | |
| - **Objective:** next-byte prediction; on chats, loss on the assistant's turns only | |
| - **Recurrence:** the number of passes is drawn per step during training, and | |
| the halting head learns when to stop | |
| ## Evaluation | |
| No benchmark results are published yet. During training, validation loss is | |
| measured on held-out conversations that never leave the machine that scores | |
| them. | |
| ## Uses | |
| - **Intended:** research and experimentation with small, ternary, recurrent | |
| models; running a chat model in a browser. | |
| - **Out of scope:** anything where a wrong answer matters. It is not a source of | |
| facts or advice, and it has not been through any safety tuning. | |
| ## Bias, risks, and limitations | |
| It is a small research model, and it is often wrong or incoherent. It may | |
| repeat biases in its training data. It is trained mostly on English, and it has | |
| no knowledge of events beyond what its data held. Verify anything it says. | |
| ## Citation | |
| Nirca Mini's architecture is derived from [mini-AGI](https://github.com/volotat/mini-AGI) (see | |
| Architecture): | |
| ```bibtex | |
| @misc{volotat_miniagi, | |
| author = {volotat}, | |
| title = {mini-AGI}, | |
| howpublished = {\url{https://github.com/volotat/mini-AGI}}, | |
| note = {GitHub repository} | |
| } | |
| ``` | |