Viu-1.5B-MoE (Frontier Sovereign 48-Layer Transformer)
Viu-1.5B-MoE is a sovereign Indian Large Language Model featuring a 48-Layer Ultra-Deep Frontier Mixture of Experts (MoE) architecture designed for deep hierarchical reasoning, Indic multilingual fluency (Hindi, Hinglish, English), and high-speed on-device inference.
- Total Parameters: 1,855,544,064 (~1.856 Billion with MTP / ~1.819 Billion Base)
- Active Parameters: 300,473,088 (~300.5 Million, 16.19% active compute per token)
- Attention Engine: Compressed Sparse Attention 2 (DeepSeek-V4.1 CSA2) with low-rank KV compression ($c_Q=256, c_{KV}=256$) + decoupled RoPE + Gemma 4 Dual QK-Norm.
- Residual Engine: DeepSeek-V4.1 Manifold-Constrained Hyper-Connections (mHC) eliminating signal explosion across 48 layers.
- MoE Topology: 1 Permanent Shared Expert + 28 Routed Micro-Experts across 48 Layers (1,392 Total Micro-Experts, Top-2 routing with Dynamic Router Bias).
- Frontier Stability: Gemma 4 (April 2026) Pure Dual-Head QK-Norm with zero gradient saturation.
- Tokenizer: Custom 48,000-vocabulary BPE optimized for Devanagari scripts, Hinglish colloquialisms, and code.
- Pretraining Dataset:
ViuAI/viu-mini-raw-pretrain(125B+ authentic tokens across 2,109 verified Parquet files).
ποΈ Architecture Highlights
- 48-Layer Ultra-Deep Reasoning: Deep hierarchical abstraction across 4 cognitive tiers (12 layers each: Token Mechanics $\rightarrow$ Multilingual Alignment $\rightarrow$ Domain Knowledge $\rightarrow$ Deep Deductive Logic).
- Compressed Sparse Attention 2 (CSA2): DeepSeek-V4.1 style low-rank KV compression saving 80% KV-cache memory for long-context inference on consumer GPUs.
- Manifold-Constrained Hyper-Connections (mHC): DeepSeek-V4.1 doubly stochastic manifold projection guaranteeing bounded signal propagation across all 48 layers.
- 1 Shared + 28 Routed Experts (1,392 Total): Dedicated shared expert captures universal grammar, while dynamic router bias enables auxiliary-loss-free expert specialization.
- Gemma 4 Pure QK-Norm: Eliminates legacy tanh soft-capping for maximum attention FLOPs and zero gradient saturation.
- Multi-Token Prediction (MTP): Lookahead training forcing the model to anticipate multiple tokens ahead.
π Quick Start
1. Verify Architecture Locally
python model/scripts/verify_arch.py
2. Pretrain on Single NVIDIA GeForce RTX 5090 (32GB)
cd model/scripts
python train.py --config ../configs/train_rtx5090.yaml
3. Pretrain on Single NVIDIA GeForce RTX 4090 (24GB)
cd model/scripts
python train.py --config ../configs/train_rtx4090.yaml