Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
ViuAI
/
ViuMini-MoE-242M
Like
0
Text Generation
PyTorch
Hindi
English
Mixture of Experts
mixture-of-experts
indic
hindi
hinglish
mla
deepseek-v3
gemma-2
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
ViuMini-MoE-242M
/
model
2.9 GB
Ctrl+K
Ctrl+K
2 contributors
History:
76 commits
ViuAI
Fix CUDA OOM on T4: cast model to amp_dtype directly, batch_size=2, safe RMSNorm
a30bdca
verified
21 minutes ago
checkpoints
full ckpt step_2000 (resume)
7 days ago
configs
Fix CUDA OOM on T4: cast model to amp_dtype directly, batch_size=2, safe RMSNorm
33 minutes ago
scripts
Fix CUDA OOM on T4: cast model to amp_dtype directly, batch_size=2, safe RMSNorm
21 minutes ago
NOTES.md
Safe
2.07 kB
feat: SDPA FlashAttention, language tags conditioning, probe safety, slim checkpoints, and streaming clean
14 days ago