Text Generation
TensorRT
Safetensors
PyTorch
English
llama
code
developer-agent
blackwell
rtx-5070
slm
conversational
Instructions to use Vivid86/MiniTransformer-91M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- TensorRT
How to use Vivid86/MiniTransformer-91M with TensorRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
|
Download README.md from Vivid86/MiniTransformer-91M: direct link, hf CLI and curl.
- Browser
- Download file 9.21 kB
-
https://huggingface.co/Vivid86/MiniTransformer-91M/resolve/main/README.md
- Command line
-
hf download hf://Vivid86/MiniTransformer-91M/README.md
-
curl -L -o README.md https://huggingface.co/Vivid86/MiniTransformer-91M/resolve/main/README.md
9.21 kB
| language: | |
| - en | |
| license: apache-2.0 | |
| tags: | |
| - code | |
| - developer-agent | |
| - llama | |
| - tensorrt | |
| - blackwell | |
| - rtx-5070 | |
| - pytorch | |
| - slm | |
| - text-generation | |
| pipeline_tag: text-generation | |
| inference: | |
| parameters: | |
| temperature: 0.2 | |
| top_p: 0.9 | |
| max_new_tokens: 160 | |
| repetition_penalty: 1.1 | |
| widget: | |
| - text: "Human: Who are you?\n\nAssistant:" | |
| example_title: "Who are you?" | |
| - text: "Human: What are your core operating rules?\n\nAssistant:" | |
| example_title: "Core Operating Rules" | |
| - text: "Human: Write a Python function to compute the factorial of a number.\n\nAssistant:" | |
| example_title: "Factorial in Python" | |
| - text: "Human: How do you undo the last Git commit without losing your staged changes?\n\nAssistant:" | |
| example_title: "Undo Git Commit" | |
| - text: "Human: Explain what a binary search algorithm does and why it is fast.\n\nAssistant:" | |
| example_title: "Binary Search" | |
| - text: "Human: Write a SQL query to count customers by country.\n\nAssistant:" | |
| example_title: "SQL Query" | |
| # β‘ MiniTransformer-91M (Vivid86 Local SLM) | |
| [](https://pytorch.org/) | |
| [-76B900?logo=nvidia&logoColor=white)](https://developer.nvidia.com/cuda-toolkit) | |
| [](https://developer.nvidia.com/tensorrt) | |
| [](https://www.nvidia.com/) | |
| [](#-technical-architecture) | |
| [](LICENSE) | |
| [](https://vividstechhaven.pages.dev) | |
| **MiniTransformer-91M** is an ultra-fast, local-first Small Language Model (SLM) trained completely from scratch on consumer hardware without third-party API dependencies or cloud GPU clusters. | |
| Engineered specifically as a lightning-fast local co-processor for developer agent routing, intent classification, and structured code synthesis, it delivers **sub-2ms latency** and **up to 671 QPS** when compiled to NVIDIA TensorRT FP16. | |
| > *"Sharp, confident, and relentless about writing clean code. Read before write, verify after edit, and never guess."* | |
| --- | |
| ## ποΈ Technical Architecture | |
| MiniTransformer-91M follows a modern LLaMA-style autoregressive decoder architecture optimized for low-latency inference: | |
| | Component | Specification | Architectural Purpose | | |
| | :--- | :--- | :--- | | |
| | **Total Parameters** | **91,245,312 (~91.2M)** | Right-sized for microsecond response on consumer GPUs | | |
| | **Layers (`n_layers`)** | **12 Transformer Blocks** | Balanced depth for coherent multi-step agent reasoning | | |
| | **Hidden Dim (`d_model`)** | **768** | Standard projection dimension matching GPT-2 scale | | |
| | **Attention Heads** | **12 Query / 12 Key-Value** | FlashAttention Scaled Dot-Product Attention (SDPA) | | |
| | **Feed-Forward (`d_ff`)** | **2,048 (SwiGLU)** | Swish-Gated Linear Units (`SiLU(W1(x)) * W3(x) -> W2`) | | |
| | **Context Length** | **1,024 tokens** | Rotary Position Embeddings (RoPE) with strict KV-cache alignment | | |
| | **Vocabulary Size** | **8,192 tokens** | Custom Byte-Level BPE trained on technical & code corpora | | |
| | **Weight Tying** | **Enabled** | Embedding table shared with output language modeling head ($V \times D$) | | |
| | **Normalization** | **RMSNorm ($\epsilon = 10^{-5}$)** | Scaled root-mean-square normalization without mean-centering overhead | | |
| | **Precision** | **FP16 / BF16** | Native mixed-precision training and FP16 inference | | |
| --- | |
| ## β‘ Workstation Inference Benchmarks | |
| All benchmarks were measured on a single consumer workstation equipped with an **NVIDIA GeForce RTX 5070 12GB (Blackwell `sm_120`)** running CUDA 12.8 on Windows 11 / WSL2. | |
| | Runtime / Engine | Execution Backend | Single-Token Latency | Throughput (QPS) | Inference VRAM | | |
| | :--- | :--- | :---: | :---: | :---: | | |
| | **PyTorch 2.11 (Eager)** | FP16 Autocast | 4.81 ms | 208 QPS | 420 MB | | |
| | **PyTorch `torch.compile`** | Inductor + Triton | 3.10 ms | 322 QPS | 410 MB | | |
| | **ONNX Runtime** | CUDA Execution Provider | 2.40 ms | 416 QPS | 360 MB | | |
| | **NVIDIA TensorRT 11.3** | FP16 Engine (`sm_120`) | **1.24 ms** | **671 QPS** | **290 MB** | | |
| *Note: TensorRT throughput was measured using concurrent asynchronous execution streams over batch size 1 with continuous generation.* | |
| --- | |
| ## π¬ Training Innovation: CPU AdamW Offload (< 1.8GB Peak VRAM) | |
| Pre-training and fine-tuning models on consumer GPUs typically fails due to optimizer memory overhead: | |
| * Standard 32-bit AdamW stores two state vectors ($m_t$ and $v_t$) per parameter. For larger models, optimizer states alone consume **10+ GB of VRAM**, crashing 12GB consumer cards. | |
| * **The Solution**: AdamW optimizer states were offloaded directly to 32GB system DDR5 RAM over PCIe. The RTX 5070 GPU exclusively handles the forward and backward passes. | |
| * **The Result**: Peak training VRAM stayed **under 1.78 GB**, leaving massive headroom for sequence lengths and batching on a standard desktop GPU. | |
| Read the complete engineering deep-dive on the workbench: | |
| π [I Trained a 458M Model on an RTX 5070 with Under 2GB VRAM: Crazy or Frugal?](https://vividstechhaven.pages.dev/blog/training-458m-under-2gb-vram) | |
| --- | |
| ## π Quickstart & Usage | |
| ### 1. Standard Hugging Face `transformers` (Universal 3 Lines) | |
| Because this model exports directly to `LlamaForCausalLM` format, it can be loaded with any standard Hugging Face pipeline: | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "Vivid86/MiniTransformer-91M" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| torch_dtype=torch.float16, | |
| device_map="auto" | |
| ) | |
| # Format prompts using the standard Human / Assistant conversation format | |
| prompt = "Human: Write a Python function to check if a word is a palindrome.\n\nAssistant:" | |
| inputs = tokenizer(prompt, return_tensors="pt").to(model.device) | |
| outputs = model.generate( | |
| **inputs, | |
| max_new_tokens=128, | |
| temperature=0.2, | |
| top_p=0.9, | |
| repetition_penalty=1.1 | |
| ) | |
| print(tokenizer.decode(outputs[0], skip_special_tokens=True)) | |
| ``` | |
| ### 2. High-Throughput Serving with `vLLM` | |
| You can deploy MiniTransformer-91M as an OpenAI-compatible HTTP API server using `vLLM`: | |
| ```bash | |
| pip install vllm | |
| vllm serve Vivid86/MiniTransformer-91M \ | |
| --port 8000 \ | |
| --max-model-len 1024 \ | |
| --dtype float16 | |
| ``` | |
| Once running, query it with any OpenAI-compatible client (Cursor, Continue.dev, LangChain, or curl): | |
| ```bash | |
| curl http://localhost:8000/v1/chat/completions \ | |
| -H "Content-Type: application/json" \ | |
| -d '{ | |
| "model": "Vivid86/MiniTransformer-91M", | |
| "messages": [{"role": "user", "content": "Explain binary search simply."}], | |
| "temperature": 0.2 | |
| }' | |
| ``` | |
| --- | |
| ## π― Intended Use & Honest Engineering Boundaries | |
| ### β Where MiniTransformer-91M Excels: | |
| 1. **Local Agent Routing**: Classifying user intent, selecting tools, and dispatching tasks in sub-2ms without round-trip network lag. | |
| 2. **Structured JSON Extraction**: Enforcing rigid schemas and parsing noisy inputs on-device. | |
| 3. **Local Co-Processor Loops**: Running continuous validation or verification passes in background agent swarms without burning cloud API budgets. | |
| 4. **Edge & Embedded Devices**: Low memory footprint (< 300MB VRAM) allows it to run on entry-level GPUs, laptops, and mini-PCs. | |
| ### β οΈ Honest Limitations (The 1,024 Token Ceiling): | |
| * **Context Ceiling**: The model operates with a hard 1,024 token rotary context window. It is not designed to absorb 50 pages of documentation in a single prompt. | |
| * **Frontier Reasoning**: It will not match frontier models (Claude 3.5 Sonnet, GPT-4o) on abstract multi-hop mathematical proofs or complex full-stack architectural design. | |
| * **Best Practice**: Pair MiniTransformer as a local high-speed routing and triage layer in front of larger models or dedicated RAG pipelines. | |
| --- | |
| ## π The Vivid Ecosystem | |
| MiniTransformer-91M is part of the **Vivid Local SLM Family** developed for the **Vivid Developer Agent OS**: | |
| * π οΈ **Workbench & Technical Debates**: [Vivid's Tech Bench](https://vividstechhaven.pages.dev) | |
| * π€ **Multi-Agent Runtime**: [Vivid Developer Agent GitHub](https://github.com/Ricky/Vivid_Developer_Agent) | |
| * π¦ **Model Family**: | |
| * `Vivid86/MiniTransformer-91M` (This model β 1.2ms latency, 671 QPS) | |
| * `Vivid86/MiniTransformer-220M` (Balanced developer model β 369 QPS) | |
| * `Vivid86/MiniTransformer-458M` (Deeper reasoning model trained with DDR5 offload) | |
| --- | |
| ## π License & Citation | |
| Released under the **Apache 2.0 License**. Free for personal, research, and commercial use. | |
| ```bibtex | |
| @misc{vivid86_minitransformer_2026, | |
| author = {Vivid86}, | |
| title = {MiniTransformer-91M: High-Throughput Local SLM for Consumer Silicon}, | |
| year = {2026}, | |
| publisher = {Hugging Face}, | |
| howpublished = {\url{https://huggingface.co/Vivid86/MiniTransformer-91M}} | |
| } | |
| ``` | |