--- language: - en library_name: transformers pipeline_tag: text-generation tags: - novi - novi-micro - causal-lm - from-scratch - bananaall --- # Novi-Micro-Base ![Novi-Micro Banner](banner.jpg) **Novi-Micro-Base** is a tiny causal language model trained from scratch by **Novi-AI**. With approximately **5.04 million parameters**, Novi-Micro explores language modeling at a small scale while using a substantially larger context window and training corpus than earlier Novi models. ⚡ **5.04M parameters · 1B training tokens · 2,048-token context** ## Model Details ### Architecture Novi-Micro-Base uses a custom **BananaMind 2-style decoder architecture** with RMSNorm, Rotary Position Embeddings (RoPE), grouped-query attention, QK normalization, SwiGLU feed-forward layers, and tied input/output embeddings. | Property | Value | | ------------------ | --------------------------------: | | Model type | Causal Language Model | | Architecture | **BananaMind 2-style** | | Parameters | **5.04M** | | Vocabulary size | **16,384** | | Context length | **2,048 tokens** | | Attention | **Grouped-Query Attention (GQA)** | | Position encoding | **RoPE** | | Normalization | **RMSNorm** | | Feed-forward | **SwiGLU** | | QK normalization | **Enabled** | | Embeddings | **Tied** | | Training precision | **FP32** | ## Training Novi-Micro-Base was trained from scratch for approximately **1 billion tokens**. The training configuration used a **5M parameter target**, a sequence length of **2,048 tokens**, and a single dataset source. ### Training Configuration | Setting | Value | | --------------------- | ----------------: | | Training mode | **Pretraining** | | Target size | **5M parameters** | | Actual parameters | **5.04M** | | Training tokens | **1,000,000,000** | | Sequence length | **2,048** | | Batch size | **2** | | Gradient accumulation | **8** | | Effective batch size | **16 sequences** | | Learning rate | **2 × 10⁻⁴** | | Scheduler | **Cosine** | | Warmup ratio | **3%** | | Weight decay | **0.01** | | Maximum gradient norm | **1.0** | | Optimizer steps | **30,518** | | Precision | **FP32** | | Torch compile | **Enabled** | | Random seed | **1337** | | Checkpoint interval | **100 steps** | | Logging interval | **10 steps** | The model was trained until the selected **1,000,000,000-token target** was consumed. ## Dataset Novi-Micro-Base was trained using: * **FineWeb-HQ** The training configuration allocated a token target of exactly: **1,000,000,000 tokens** Documents were streamed from the dataset and packed into contiguous **2,048-token sequences** before being passed to the model. ## Tokenizer Novi-Micro uses a custom **byte-level BPE tokenizer** with a vocabulary size of **16,384 tokens**. The tokenizer was trained from samples drawn from the training dataset before model pretraining. Special tokens include: * `[PAD]` * `[BOS]` * `[EOS]` * `[UNK]` The tokenizer uses ByteLevel pre-tokenization and decoding. ## Architecture Details Novi-Micro uses a decoder-only Transformer architecture. ### Attention The model uses **grouped-query attention (GQA)**, where multiple query heads share key and value heads. Rotary Position Embeddings (**RoPE**) are applied to the query and key representations. The BananaMind 2-style architecture also applies **RMSNorm to the query and key head dimensions**. ### Feed-Forward Network Each Transformer block uses a **SwiGLU** feed-forward network: ```text SwiGLU(x) = SiLU(gate(x)) × up(x) ``` The result is projected back to the model's hidden dimension. ### Normalization The model uses **RMSNorm** before both the attention and feed-forward sublayers. ### Embeddings The input token embeddings and language-model output embeddings are **tied**, reducing the number of independent parameters. ## Intended Use Novi-Micro-Base is primarily intended for: * 🔬 Research and experimentation * 🧪 Small-model language-model experiments * 🎓 Educational purposes * 🛠️ Fine-tuning experiments * 💻 Lightweight local inference * 🤖 Exploring language modeling at the million-parameter scale As a **base model**, Novi-Micro-Base is not instruction-tuned and is not specifically trained to follow user commands or behave as a conversational assistant. ## Limitations Novi-Micro-Base is a very small experimental language model. With approximately **5 million parameters**, it is dramatically smaller than modern general-purpose language models and should not be expected to match their capabilities. It may: * Generate incoherent text * Repeat phrases * Produce factual errors * Struggle with complex instructions * Have limited world knowledge * Perform poorly on difficult reasoning tasks * Produce grammatically unusual text * Fail to maintain coherent long-form generations A 2,048-token context window does not eliminate the limitations caused by the model's small parameter count. This model should be considered a **research and experimentation model**, rather than a production-ready general-purpose LLM. ## Usage ```python from transformers import AutoTokenizer, AutoModelForCausalLM model_id = "Novi-AI/Novi-Micro-Base" tokenizer = AutoTokenizer.from_pretrained( model_id, trust_remote_code=True, ) model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, ) prompt = "Hello, my name is" inputs = tokenizer(prompt, return_tensors="pt") outputs = model.generate( **inputs, max_new_tokens=50, ) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` Because Novi-Micro uses a custom architecture, `trust_remote_code=True` is required when loading the model through Transformers. ## Project History Novi-Micro-Base is part of **Project Kairo**, the development codename for the Novi model project. The Novi series follows the earlier **AppleMind** experiments and represents the primary model-development line of Novi-AI. **AppleMind → Novi-Nano → Novi-Micro → future Novi models** 🚀 Novi-Micro substantially expands upon Novi-Nano by increasing the model to approximately **5.04M parameters**, expanding the context window from 256 to **2,048 tokens**, and training on approximately **1 billion tokens**. ## Training Infrastructure The model was trained using GPU compute through the **BananaAll** training environment. The training worker supports GPU acceleration through CUDA and Intel XPU, with Torch compilation enabled for supported environments. This model was trained from scratch rather than fine-tuned from an existing language model checkpoint. ## Acknowledgements Novi-Micro was built using the open-source machine-learning ecosystem and datasets made available by the community. Special thanks to: * Hugging Face 🤗 * FineWeb * FineWeb-HQ * The open-source Transformers ecosystem ## License This model is released under the **Apache 2.0** license. --- ## 🧠 Novi AI **Small models. Big experiments.** Novi-Micro is intentionally tiny — exploring how far a language model can go with only a few million parameters and approximately one billion training tokens. *Novi AI 2026 — Project Kairo*