SLM-Archive
/

Novi-Micro-Base / README.md
GGUFGuy's picture NoviAIBot's picture
Duplicate from Novi-AI/Novi-Micro-Base
d39a477
|
Raw History Blame Contribute Delete
7.76 kB
---
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- novi
- novi-micro
- causal-lm
- from-scratch
- bananaall
---
# Novi-Micro-Base
![Novi-Micro Banner](banner.jpg)
**Novi-Micro-Base** is a tiny causal language model trained from scratch by **Novi-AI**.
With approximately **5.04 million parameters**, Novi-Micro explores language modeling at a small scale while using a substantially larger context window and training corpus than earlier Novi models.
โšก **5.04M parameters ยท 1B training tokens ยท 2,048-token context**
## Model Details
### Architecture
Novi-Micro-Base uses a custom **BananaMind 2-style decoder architecture** with RMSNorm, Rotary Position Embeddings (RoPE), grouped-query attention, QK normalization, SwiGLU feed-forward layers, and tied input/output embeddings.
| Property | Value |
| ------------------ | --------------------------------: |
| Model type | Causal Language Model |
| Architecture | **BananaMind 2-style** |
| Parameters | **5.04M** |
| Vocabulary size | **16,384** |
| Context length | **2,048 tokens** |
| Attention | **Grouped-Query Attention (GQA)** |
| Position encoding | **RoPE** |
| Normalization | **RMSNorm** |
| Feed-forward | **SwiGLU** |
| QK normalization | **Enabled** |
| Embeddings | **Tied** |
| Training precision | **FP32** |
## Training
Novi-Micro-Base was trained from scratch for approximately **1 billion tokens**.
The training configuration used a **5M parameter target**, a sequence length of **2,048 tokens**, and a single dataset source.
### Training Configuration
| Setting | Value |
| --------------------- | ----------------: |
| Training mode | **Pretraining** |
| Target size | **5M parameters** |
| Actual parameters | **5.04M** |
| Training tokens | **1,000,000,000** |
| Sequence length | **2,048** |
| Batch size | **2** |
| Gradient accumulation | **8** |
| Effective batch size | **16 sequences** |
| Learning rate | **2 ร— 10โปโด** |
| Scheduler | **Cosine** |
| Warmup ratio | **3%** |
| Weight decay | **0.01** |
| Maximum gradient norm | **1.0** |
| Optimizer steps | **30,518** |
| Precision | **FP32** |
| Torch compile | **Enabled** |
| Random seed | **1337** |
| Checkpoint interval | **100 steps** |
| Logging interval | **10 steps** |
The model was trained until the selected **1,000,000,000-token target** was consumed.
## Dataset
Novi-Micro-Base was trained using:
* **FineWeb-HQ**
The training configuration allocated a token target of exactly:
**1,000,000,000 tokens**
Documents were streamed from the dataset and packed into contiguous **2,048-token sequences** before being passed to the model.
## Tokenizer
Novi-Micro uses a custom **byte-level BPE tokenizer** with a vocabulary size of **16,384 tokens**.
The tokenizer was trained from samples drawn from the training dataset before model pretraining.
Special tokens include:
* `[PAD]`
* `[BOS]`
* `[EOS]`
* `[UNK]`
The tokenizer uses ByteLevel pre-tokenization and decoding.
## Architecture Details
Novi-Micro uses a decoder-only Transformer architecture.
### Attention
The model uses **grouped-query attention (GQA)**, where multiple query heads share key and value heads.
Rotary Position Embeddings (**RoPE**) are applied to the query and key representations.
The BananaMind 2-style architecture also applies **RMSNorm to the query and key head dimensions**.
### Feed-Forward Network
Each Transformer block uses a **SwiGLU** feed-forward network:
```text
SwiGLU(x) = SiLU(gate(x)) ร— up(x)
```
The result is projected back to the model's hidden dimension.
### Normalization
The model uses **RMSNorm** before both the attention and feed-forward sublayers.
### Embeddings
The input token embeddings and language-model output embeddings are **tied**, reducing the number of independent parameters.
## Intended Use
Novi-Micro-Base is primarily intended for:
* ๐Ÿ”ฌ Research and experimentation
* ๐Ÿงช Small-model language-model experiments
* ๐ŸŽ“ Educational purposes
* ๐Ÿ› ๏ธ Fine-tuning experiments
* ๐Ÿ’ป Lightweight local inference
* ๐Ÿค– Exploring language modeling at the million-parameter scale
As a **base model**, Novi-Micro-Base is not instruction-tuned and is not specifically trained to follow user commands or behave as a conversational assistant.
## Limitations
Novi-Micro-Base is a very small experimental language model.
With approximately **5 million parameters**, it is dramatically smaller than modern general-purpose language models and should not be expected to match their capabilities.
It may:
* Generate incoherent text
* Repeat phrases
* Produce factual errors
* Struggle with complex instructions
* Have limited world knowledge
* Perform poorly on difficult reasoning tasks
* Produce grammatically unusual text
* Fail to maintain coherent long-form generations
A 2,048-token context window does not eliminate the limitations caused by the model's small parameter count.
This model should be considered a **research and experimentation model**, rather than a production-ready general-purpose LLM.
## Usage
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "Novi-AI/Novi-Micro-Base"
tokenizer = AutoTokenizer.from_pretrained(
model_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
)
prompt = "Hello, my name is"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
**inputs,
max_new_tokens=50,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
Because Novi-Micro uses a custom architecture, `trust_remote_code=True` is required when loading the model through Transformers.
## Project History
Novi-Micro-Base is part of **Project Kairo**, the development codename for the Novi model project.
The Novi series follows the earlier **AppleMind** experiments and represents the primary model-development line of Novi-AI.
**AppleMind โ†’ Novi-Nano โ†’ Novi-Micro โ†’ future Novi models** ๐Ÿš€
Novi-Micro substantially expands upon Novi-Nano by increasing the model to approximately **5.04M parameters**, expanding the context window from 256 to **2,048 tokens**, and training on approximately **1 billion tokens**.
## Training Infrastructure
The model was trained using GPU compute through the **BananaAll** training environment.
The training worker supports GPU acceleration through CUDA and Intel XPU, with Torch compilation enabled for supported environments.
This model was trained from scratch rather than fine-tuned from an existing language model checkpoint.
## Acknowledgements
Novi-Micro was built using the open-source machine-learning ecosystem and datasets made available by the community.
Special thanks to:
* Hugging Face ๐Ÿค—
* FineWeb
* FineWeb-HQ
* The open-source Transformers ecosystem
## License
This model is released under the **Apache 2.0** license.
---
## ๐Ÿง  Novi AI
**Small models. Big experiments.**
Novi-Micro is intentionally tiny โ€” exploring how far a language model can go with only a few million parameters and approximately one billion training tokens.
*Novi AI 2026 โ€” Project Kairo*