Mimans M1 (514M) - 5.01B Tokens Milestone
Mimans M1 is a 514M parameter decoder-only transformer pretrained from scratch on code, math, and technical text using TileLang custom Blackwell GPU kernels and Muon + AdamW hybrid optimization.
Model Summary
- Parameters: 514,345,526 (~514M)
- Architecture: GQA (10:2), SwiGLU FFN, AttnRes Skip Gating, RMSNorm
- Context Length: 4,096 tokens (max 8,192, RoPE $\theta = 500,000$)
- Vocabulary: 49,152 (Byte-level BPE with FIM support)
- Training Tokens: 5.01B Tokens (Global Step 9,551)
- Latest Loss: 1.9756
- Hardware: NVIDIA GeForce RTX 5090 (Blackwell
sm_120)
Usage & Generation
Loading with Hugging Face Transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "ankushthakurr09/MimansM1_v1"
model = AutoModelForCausalLM.from_pretrained(repo_id, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
prompt = "def fibonacci(n):"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Running Standalone Generation Script:
python generate.py --prompt "def hello_world():" --max-tokens 50
- Downloads last month
- 1,283