Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
ViuAI
/
ViuMini-MoE-242M
Like
0
Text Generation
PyTorch
Hindi
English
Mixture of Experts
mixture-of-experts
indic
hindi
hinglish
mla
deepseek-v3
gemma-2
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
ViuMini-MoE-242M
/
model
/
scripts
156 kB
Ctrl+K
Ctrl+K
2 contributors
History:
47 commits
ViuAI
Fix CUDA OOM on T4: cast model to amp_dtype directly, batch_size=2, safe RMSNorm
61b414e
verified
16 minutes ago
model.py
Safe
8.18 kB
init: 28L MoE 241M audit-fixed, smoke pass
16 days ago
push_to_hf.py
Safe
2.67 kB
init: 28L MoE 241M audit-fixed, smoke pass
16 days ago
toy_router_test.py
Safe
5.62 kB
step3: add F(0.10/50) + G(0.05/100)
15 days ago
train.py
98.6 kB
Fix CUDA OOM on T4: cast model to amp_dtype directly, batch_size=2, safe RMSNorm
16 minutes ago
verify_1_5b_arch.py
Safe
2.97 kB
Add 42-layer Viu-1.5B-MoE verification test suite
7 days ago
viu1_moe.py
Safe
15.6 kB
perf: free logits immediately in forward pass to release 3.1GB VRAM for mt_loss
7 days ago
viu_moe.py
22.6 kB
Fix CUDA OOM on T4: cast model to amp_dtype directly, batch_size=2, safe RMSNorm
41 minutes ago