MiniTransformer-91M / README.md
Vivid86's picture
Upload README.md with huggingface_hub
7030274 verified
|
Raw History Blame Contribute Delete
9.21 kB
---
language:
- en
license: apache-2.0
tags:
- code
- developer-agent
- llama
- tensorrt
- blackwell
- rtx-5070
- pytorch
- slm
- text-generation
pipeline_tag: text-generation
inference:
parameters:
temperature: 0.2
top_p: 0.9
max_new_tokens: 160
repetition_penalty: 1.1
widget:
- text: "Human: Who are you?\n\nAssistant:"
example_title: "Who are you?"
- text: "Human: What are your core operating rules?\n\nAssistant:"
example_title: "Core Operating Rules"
- text: "Human: Write a Python function to compute the factorial of a number.\n\nAssistant:"
example_title: "Factorial in Python"
- text: "Human: How do you undo the last Git commit without losing your staged changes?\n\nAssistant:"
example_title: "Undo Git Commit"
- text: "Human: Explain what a binary search algorithm does and why it is fast.\n\nAssistant:"
example_title: "Binary Search"
- text: "Human: Write a SQL query to count customers by country.\n\nAssistant:"
example_title: "SQL Query"
---
# ⚑ MiniTransformer-91M (Vivid86 Local SLM)
[![PyTorch](https://img.shields.io/badge/PyTorch-2.11.0%2Bcu128-EE4C2C?logo=pytorch&logoColor=white)](https://pytorch.org/)
[![CUDA](https://img.shields.io/badge/CUDA-12.8%20(Blackwell%20sm__120)-76B900?logo=nvidia&logoColor=white)](https://developer.nvidia.com/cuda-toolkit)
[![TensorRT](https://img.shields.io/badge/TensorRT-11.3%20Enterprise-76B900?logo=nvidia&logoColor=white)](https://developer.nvidia.com/tensorrt)
[![Hardware](https://img.shields.io/badge/Hardware-NVIDIA%20GeForce%20RTX%205070-76B900)](https://www.nvidia.com/)
[![Parameters](https://img.shields.io/badge/Parameters-91.2M-blue)](#-technical-architecture)
[![License](https://img.shields.io/badge/License-Apache%202.0-green.svg)](LICENSE)
[![Workbench](https://img.shields.io/badge/Workbench-Vivid's%20Tech%20Bench-cyan)](https://vividstechhaven.pages.dev)
**MiniTransformer-91M** is an ultra-fast, local-first Small Language Model (SLM) trained completely from scratch on consumer hardware without third-party API dependencies or cloud GPU clusters.
Engineered specifically as a lightning-fast local co-processor for developer agent routing, intent classification, and structured code synthesis, it delivers **sub-2ms latency** and **up to 671 QPS** when compiled to NVIDIA TensorRT FP16.
> *"Sharp, confident, and relentless about writing clean code. Read before write, verify after edit, and never guess."*
---
## πŸ—οΈ Technical Architecture
MiniTransformer-91M follows a modern LLaMA-style autoregressive decoder architecture optimized for low-latency inference:
| Component | Specification | Architectural Purpose |
| :--- | :--- | :--- |
| **Total Parameters** | **91,245,312 (~91.2M)** | Right-sized for microsecond response on consumer GPUs |
| **Layers (`n_layers`)** | **12 Transformer Blocks** | Balanced depth for coherent multi-step agent reasoning |
| **Hidden Dim (`d_model`)** | **768** | Standard projection dimension matching GPT-2 scale |
| **Attention Heads** | **12 Query / 12 Key-Value** | FlashAttention Scaled Dot-Product Attention (SDPA) |
| **Feed-Forward (`d_ff`)** | **2,048 (SwiGLU)** | Swish-Gated Linear Units (`SiLU(W1(x)) * W3(x) -> W2`) |
| **Context Length** | **1,024 tokens** | Rotary Position Embeddings (RoPE) with strict KV-cache alignment |
| **Vocabulary Size** | **8,192 tokens** | Custom Byte-Level BPE trained on technical & code corpora |
| **Weight Tying** | **Enabled** | Embedding table shared with output language modeling head ($V \times D$) |
| **Normalization** | **RMSNorm ($\epsilon = 10^{-5}$)** | Scaled root-mean-square normalization without mean-centering overhead |
| **Precision** | **FP16 / BF16** | Native mixed-precision training and FP16 inference |
---
## ⚑ Workstation Inference Benchmarks
All benchmarks were measured on a single consumer workstation equipped with an **NVIDIA GeForce RTX 5070 12GB (Blackwell `sm_120`)** running CUDA 12.8 on Windows 11 / WSL2.
| Runtime / Engine | Execution Backend | Single-Token Latency | Throughput (QPS) | Inference VRAM |
| :--- | :--- | :---: | :---: | :---: |
| **PyTorch 2.11 (Eager)** | FP16 Autocast | 4.81 ms | 208 QPS | 420 MB |
| **PyTorch `torch.compile`** | Inductor + Triton | 3.10 ms | 322 QPS | 410 MB |
| **ONNX Runtime** | CUDA Execution Provider | 2.40 ms | 416 QPS | 360 MB |
| **NVIDIA TensorRT 11.3** | FP16 Engine (`sm_120`) | **1.24 ms** | **671 QPS** | **290 MB** |
*Note: TensorRT throughput was measured using concurrent asynchronous execution streams over batch size 1 with continuous generation.*
---
## πŸ”¬ Training Innovation: CPU AdamW Offload (< 1.8GB Peak VRAM)
Pre-training and fine-tuning models on consumer GPUs typically fails due to optimizer memory overhead:
* Standard 32-bit AdamW stores two state vectors ($m_t$ and $v_t$) per parameter. For larger models, optimizer states alone consume **10+ GB of VRAM**, crashing 12GB consumer cards.
* **The Solution**: AdamW optimizer states were offloaded directly to 32GB system DDR5 RAM over PCIe. The RTX 5070 GPU exclusively handles the forward and backward passes.
* **The Result**: Peak training VRAM stayed **under 1.78 GB**, leaving massive headroom for sequence lengths and batching on a standard desktop GPU.
Read the complete engineering deep-dive on the workbench:
πŸ‘‰ [I Trained a 458M Model on an RTX 5070 with Under 2GB VRAM: Crazy or Frugal?](https://vividstechhaven.pages.dev/blog/training-458m-under-2gb-vram)
---
## πŸš€ Quickstart & Usage
### 1. Standard Hugging Face `transformers` (Universal 3 Lines)
Because this model exports directly to `LlamaForCausalLM` format, it can be loaded with any standard Hugging Face pipeline:
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Vivid86/MiniTransformer-91M"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto"
)
# Format prompts using the standard Human / Assistant conversation format
prompt = "Human: Write a Python function to check if a word is a palindrome.\n\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=128,
temperature=0.2,
top_p=0.9,
repetition_penalty=1.1
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
### 2. High-Throughput Serving with `vLLM`
You can deploy MiniTransformer-91M as an OpenAI-compatible HTTP API server using `vLLM`:
```bash
pip install vllm
vllm serve Vivid86/MiniTransformer-91M \
--port 8000 \
--max-model-len 1024 \
--dtype float16
```
Once running, query it with any OpenAI-compatible client (Cursor, Continue.dev, LangChain, or curl):
```bash
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Vivid86/MiniTransformer-91M",
"messages": [{"role": "user", "content": "Explain binary search simply."}],
"temperature": 0.2
}'
```
---
## 🎯 Intended Use & Honest Engineering Boundaries
### βœ… Where MiniTransformer-91M Excels:
1. **Local Agent Routing**: Classifying user intent, selecting tools, and dispatching tasks in sub-2ms without round-trip network lag.
2. **Structured JSON Extraction**: Enforcing rigid schemas and parsing noisy inputs on-device.
3. **Local Co-Processor Loops**: Running continuous validation or verification passes in background agent swarms without burning cloud API budgets.
4. **Edge & Embedded Devices**: Low memory footprint (< 300MB VRAM) allows it to run on entry-level GPUs, laptops, and mini-PCs.
### ⚠️ Honest Limitations (The 1,024 Token Ceiling):
* **Context Ceiling**: The model operates with a hard 1,024 token rotary context window. It is not designed to absorb 50 pages of documentation in a single prompt.
* **Frontier Reasoning**: It will not match frontier models (Claude 3.5 Sonnet, GPT-4o) on abstract multi-hop mathematical proofs or complex full-stack architectural design.
* **Best Practice**: Pair MiniTransformer as a local high-speed routing and triage layer in front of larger models or dedicated RAG pipelines.
---
## 🌐 The Vivid Ecosystem
MiniTransformer-91M is part of the **Vivid Local SLM Family** developed for the **Vivid Developer Agent OS**:
* πŸ› οΈ **Workbench & Technical Debates**: [Vivid's Tech Bench](https://vividstechhaven.pages.dev)
* πŸ€– **Multi-Agent Runtime**: [Vivid Developer Agent GitHub](https://github.com/Ricky/Vivid_Developer_Agent)
* πŸ“¦ **Model Family**:
* `Vivid86/MiniTransformer-91M` (This model β€” 1.2ms latency, 671 QPS)
* `Vivid86/MiniTransformer-220M` (Balanced developer model β€” 369 QPS)
* `Vivid86/MiniTransformer-458M` (Deeper reasoning model trained with DDR5 offload)
---
## πŸ“œ License & Citation
Released under the **Apache 2.0 License**. Free for personal, research, and commercial use.
```bibtex
@misc{vivid86_minitransformer_2026,
author = {Vivid86},
title = {MiniTransformer-91M: High-Throughput Local SLM for Consumer Silicon},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Vivid86/MiniTransformer-91M}}
}
```