Viu-1.5B-MoE (Frontier Sovereign 48-Layer Transformer)

Viu-1.5B-MoE is a sovereign Indian Large Language Model featuring a 48-Layer Ultra-Deep Frontier Mixture of Experts (MoE) architecture designed for deep hierarchical reasoning, Indic multilingual fluency (Hindi, Hinglish, English), and high-speed on-device inference.

  • Total Parameters: 1,855,544,064 (~1.856 Billion with MTP / ~1.819 Billion Base)
  • Active Parameters: 300,473,088 (~300.5 Million, 16.19% active compute per token)
  • Attention Engine: Compressed Sparse Attention 2 (DeepSeek-V4.1 CSA2) with low-rank KV compression ($c_Q=256, c_{KV}=256$) + decoupled RoPE + Gemma 4 Dual QK-Norm.
  • Residual Engine: DeepSeek-V4.1 Manifold-Constrained Hyper-Connections (mHC) eliminating signal explosion across 48 layers.
  • MoE Topology: 1 Permanent Shared Expert + 28 Routed Micro-Experts across 48 Layers (1,392 Total Micro-Experts, Top-2 routing with Dynamic Router Bias).
  • Frontier Stability: Gemma 4 (April 2026) Pure Dual-Head QK-Norm with zero gradient saturation.
  • Tokenizer: Custom 48,000-vocabulary BPE optimized for Devanagari scripts, Hinglish colloquialisms, and code.
  • Pretraining Dataset: ViuAI/viu-mini-raw-pretrain (125B+ authentic tokens across 2,109 verified Parquet files).

πŸ›οΈ Architecture Highlights

  1. 48-Layer Ultra-Deep Reasoning: Deep hierarchical abstraction across 4 cognitive tiers (12 layers each: Token Mechanics $\rightarrow$ Multilingual Alignment $\rightarrow$ Domain Knowledge $\rightarrow$ Deep Deductive Logic).
  2. Compressed Sparse Attention 2 (CSA2): DeepSeek-V4.1 style low-rank KV compression saving 80% KV-cache memory for long-context inference on consumer GPUs.
  3. Manifold-Constrained Hyper-Connections (mHC): DeepSeek-V4.1 doubly stochastic manifold projection guaranteeing bounded signal propagation across all 48 layers.
  4. 1 Shared + 28 Routed Experts (1,392 Total): Dedicated shared expert captures universal grammar, while dynamic router bias enables auxiliary-loss-free expert specialization.
  5. Gemma 4 Pure QK-Norm: Eliminates legacy tanh soft-capping for maximum attention FLOPs and zero gradient saturation.
  6. Multi-Token Prediction (MTP): Lookahead training forcing the model to anticipate multiple tokens ahead.

πŸš€ Quick Start

1. Verify Architecture Locally

python model/scripts/verify_arch.py

2. Pretrain on Single NVIDIA GeForce RTX 5090 (32GB)

cd model/scripts
python train.py --config ../configs/train_rtx5090.yaml

3. Pretrain on Single NVIDIA GeForce RTX 4090 (24GB)

cd model/scripts
python train.py --config ../configs/train_rtx4090.yaml
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support