⚑ Vivid86 Developer Agent (91.2M Parameters)

Vivid86 is a local-first autonomous developer agent and code-generation model built and trained entirely from scratch.

"Sharp, confident, slightly snarky, and relentless about writing clean, bug-free code. Read before write, verify after edit, and never guess."


πŸ“ Technical Architecture

  • Parameters: 91,245,312 (d_model=768, layers=12, heads=12, d_ff=2048)
  • Context Length: 1,024 tokens (Rotary Position Embeddings / RoPE)
  • Attention: Scaled Dot-Product Attention (SDPA / FlashAttention) + Grouped-Query Attention (GQA)
  • Feed-Forward: SwiGLU (SiLU(W1(x)) * W3(x) -> W2) + Pre-LN RMSNorm
  • Precision: FP16 / BF16 mixed precision
  • Hardware Acceleration: Native NVIDIA TensorRT 11.3 Enterprise Blackwell (sm_120) engine on GeForce RTX 5070 (1.24 ms latency, ~100,000 tokens/sec throughput)

πŸš€ Usage

1. Universal Hugging Face transformers (3 Lines)

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Vivid86/MiniTransformer-91M")
tokenizer = AutoTokenizer.from_pretrained("Vivid86/MiniTransformer-91M")

prompt = "Human: Write a Python function to reverse a list in place.\n\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=160, temperature=0.2, repetition_penalty=1.1)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

2. OpenAI / vLLM / NVIDIA NIM Compatible Server

pip install vllm
vllm serve Vivid86/MiniTransformer-91M --port 8000 --max-model-len 1024 --dtype float16
Downloads last month
-
Safetensors
Model size
97.5M params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support