--- language: - en license: apache-2.0 tags: - code - developer-agent - llama - tensorrt - blackwell - rtx-5070 - pytorch - slm - text-generation pipeline_tag: text-generation inference: parameters: temperature: 0.2 top_p: 0.9 max_new_tokens: 160 repetition_penalty: 1.1 widget: - text: "Human: Who are you?\n\nAssistant:" example_title: "Who are you?" - text: "Human: What are your core operating rules?\n\nAssistant:" example_title: "Core Operating Rules" - text: "Human: Write a Python function to compute the factorial of a number.\n\nAssistant:" example_title: "Factorial in Python" - text: "Human: How do you undo the last Git commit without losing your staged changes?\n\nAssistant:" example_title: "Undo Git Commit" - text: "Human: Explain what a binary search algorithm does and why it is fast.\n\nAssistant:" example_title: "Binary Search" - text: "Human: Write a SQL query to count customers by country.\n\nAssistant:" example_title: "SQL Query" --- # ⚡ MiniTransformer-91M (Vivid86 Local SLM) [![PyTorch](https://img.shields.io/badge/PyTorch-2.11.0%2Bcu128-EE4C2C?logo=pytorch&logoColor=white)](https://pytorch.org/) [![CUDA](https://img.shields.io/badge/CUDA-12.8%20(Blackwell%20sm__120)-76B900?logo=nvidia&logoColor=white)](https://developer.nvidia.com/cuda-toolkit) [![TensorRT](https://img.shields.io/badge/TensorRT-11.3%20Enterprise-76B900?logo=nvidia&logoColor=white)](https://developer.nvidia.com/tensorrt) [![Hardware](https://img.shields.io/badge/Hardware-NVIDIA%20GeForce%20RTX%205070-76B900)](https://www.nvidia.com/) [![Parameters](https://img.shields.io/badge/Parameters-91.2M-blue)](#-technical-architecture) [![License](https://img.shields.io/badge/License-Apache%202.0-green.svg)](LICENSE) [![Workbench](https://img.shields.io/badge/Workbench-Vivid's%20Tech%20Bench-cyan)](https://vividstechhaven.pages.dev) **MiniTransformer-91M** is an ultra-fast, local-first Small Language Model (SLM) trained completely from scratch on consumer hardware without third-party API dependencies or cloud GPU clusters. Engineered specifically as a lightning-fast local co-processor for developer agent routing, intent classification, and structured code synthesis, it delivers **sub-2ms latency** and **up to 671 QPS** when compiled to NVIDIA TensorRT FP16. > *"Sharp, confident, and relentless about writing clean code. Read before write, verify after edit, and never guess."* --- ## 🏗️ Technical Architecture MiniTransformer-91M follows a modern LLaMA-style autoregressive decoder architecture optimized for low-latency inference: | Component | Specification | Architectural Purpose | | :--- | :--- | :--- | | **Total Parameters** | **91,245,312 (~91.2M)** | Right-sized for microsecond response on consumer GPUs | | **Layers (`n_layers`)** | **12 Transformer Blocks** | Balanced depth for coherent multi-step agent reasoning | | **Hidden Dim (`d_model`)** | **768** | Standard projection dimension matching GPT-2 scale | | **Attention Heads** | **12 Query / 12 Key-Value** | FlashAttention Scaled Dot-Product Attention (SDPA) | | **Feed-Forward (`d_ff`)** | **2,048 (SwiGLU)** | Swish-Gated Linear Units (`SiLU(W1(x)) * W3(x) -> W2`) | | **Context Length** | **1,024 tokens** | Rotary Position Embeddings (RoPE) with strict KV-cache alignment | | **Vocabulary Size** | **8,192 tokens** | Custom Byte-Level BPE trained on technical & code corpora | | **Weight Tying** | **Enabled** | Embedding table shared with output language modeling head ($V \times D$) | | **Normalization** | **RMSNorm ($\epsilon = 10^{-5}$)** | Scaled root-mean-square normalization without mean-centering overhead | | **Precision** | **FP16 / BF16** | Native mixed-precision training and FP16 inference | --- ## ⚡ Workstation Inference Benchmarks All benchmarks were measured on a single consumer workstation equipped with an **NVIDIA GeForce RTX 5070 12GB (Blackwell `sm_120`)** running CUDA 12.8 on Windows 11 / WSL2. | Runtime / Engine | Execution Backend | Single-Token Latency | Throughput (QPS) | Inference VRAM | | :--- | :--- | :---: | :---: | :---: | | **PyTorch 2.11 (Eager)** | FP16 Autocast | 4.81 ms | 208 QPS | 420 MB | | **PyTorch `torch.compile`** | Inductor + Triton | 3.10 ms | 322 QPS | 410 MB | | **ONNX Runtime** | CUDA Execution Provider | 2.40 ms | 416 QPS | 360 MB | | **NVIDIA TensorRT 11.3** | FP16 Engine (`sm_120`) | **1.24 ms** | **671 QPS** | **290 MB** | *Note: TensorRT throughput was measured using concurrent asynchronous execution streams over batch size 1 with continuous generation.* --- ## 🔬 Training Innovation: CPU AdamW Offload (< 1.8GB Peak VRAM) Pre-training and fine-tuning models on consumer GPUs typically fails due to optimizer memory overhead: * Standard 32-bit AdamW stores two state vectors ($m_t$ and $v_t$) per parameter. For larger models, optimizer states alone consume **10+ GB of VRAM**, crashing 12GB consumer cards. * **The Solution**: AdamW optimizer states were offloaded directly to 32GB system DDR5 RAM over PCIe. The RTX 5070 GPU exclusively handles the forward and backward passes. * **The Result**: Peak training VRAM stayed **under 1.78 GB**, leaving massive headroom for sequence lengths and batching on a standard desktop GPU. Read the complete engineering deep-dive on the workbench: 👉 [I Trained a 458M Model on an RTX 5070 with Under 2GB VRAM: Crazy or Frugal?](https://vividstechhaven.pages.dev/blog/training-458m-under-2gb-vram) --- ## 🚀 Quickstart & Usage ### 1. Standard Hugging Face `transformers` (Universal 3 Lines) Because this model exports directly to `LlamaForCausalLM` format, it can be loaded with any standard Hugging Face pipeline: ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "Vivid86/MiniTransformer-91M" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.float16, device_map="auto" ) # Format prompts using the standard Human / Assistant conversation format prompt = "Human: Write a Python function to check if a word is a palindrome.\n\nAssistant:" inputs = tokenizer(prompt, return_tensors="pt").to(model.device) outputs = model.generate( **inputs, max_new_tokens=128, temperature=0.2, top_p=0.9, repetition_penalty=1.1 ) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` ### 2. High-Throughput Serving with `vLLM` You can deploy MiniTransformer-91M as an OpenAI-compatible HTTP API server using `vLLM`: ```bash pip install vllm vllm serve Vivid86/MiniTransformer-91M \ --port 8000 \ --max-model-len 1024 \ --dtype float16 ``` Once running, query it with any OpenAI-compatible client (Cursor, Continue.dev, LangChain, or curl): ```bash curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "Vivid86/MiniTransformer-91M", "messages": [{"role": "user", "content": "Explain binary search simply."}], "temperature": 0.2 }' ``` --- ## 🎯 Intended Use & Honest Engineering Boundaries ### ✅ Where MiniTransformer-91M Excels: 1. **Local Agent Routing**: Classifying user intent, selecting tools, and dispatching tasks in sub-2ms without round-trip network lag. 2. **Structured JSON Extraction**: Enforcing rigid schemas and parsing noisy inputs on-device. 3. **Local Co-Processor Loops**: Running continuous validation or verification passes in background agent swarms without burning cloud API budgets. 4. **Edge & Embedded Devices**: Low memory footprint (< 300MB VRAM) allows it to run on entry-level GPUs, laptops, and mini-PCs. ### ⚠️ Honest Limitations (The 1,024 Token Ceiling): * **Context Ceiling**: The model operates with a hard 1,024 token rotary context window. It is not designed to absorb 50 pages of documentation in a single prompt. * **Frontier Reasoning**: It will not match frontier models (Claude 3.5 Sonnet, GPT-4o) on abstract multi-hop mathematical proofs or complex full-stack architectural design. * **Best Practice**: Pair MiniTransformer as a local high-speed routing and triage layer in front of larger models or dedicated RAG pipelines. --- ## 🌐 The Vivid Ecosystem MiniTransformer-91M is part of the **Vivid Local SLM Family** developed for the **Vivid Developer Agent OS**: * 🛠️ **Workbench & Technical Debates**: [Vivid's Tech Bench](https://vividstechhaven.pages.dev) * 🤖 **Multi-Agent Runtime**: [Vivid Developer Agent GitHub](https://github.com/Ricky/Vivid_Developer_Agent) * 📦 **Model Family**: * `Vivid86/MiniTransformer-91M` (This model — 1.2ms latency, 671 QPS) * `Vivid86/MiniTransformer-220M` (Balanced developer model — 369 QPS) * `Vivid86/MiniTransformer-458M` (Deeper reasoning model trained with DDR5 offload) --- ## 📜 License & Citation Released under the **Apache 2.0 License**. Free for personal, research, and commercial use. ```bibtex @misc{vivid86_minitransformer_2026, author = {Vivid86}, title = {MiniTransformer-91M: High-Throughput Local SLM for Consumer Silicon}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/Vivid86/MiniTransformer-91M}} } ```