Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
ViuAI
/
ViuMini-Dense-360M
Like
1
Text Generation
PyTorch
Hindi
English
dense
transformer
indic
hindi
hinglish
english
gqa
gemma-4
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
ViuMini-Dense-360M
/
model
/
scripts
168 kB
Ctrl+K
Ctrl+K
2 contributors
History:
70 commits
ViuAI
fix: restore grad_ckpt and safe micro 8 accum 8 for zero-OOM guarantee on T4
96be75f
verified
about 7 hours ago
model.py
217 Bytes
refactor: consolidate pure 32L Dense architecture, purge MoE files, audit hardening
about 8 hours ago
push_to_hf.py
2.68 kB
refactor: consolidate pure 32L Dense architecture, purge MoE files, audit hardening
about 8 hours ago
toy_router_test.py
Safe
5.62 kB
step3: add F(0.10/50) + G(0.05/100)
17 days ago
train.py
101 kB
fix: restore grad_ckpt and safe micro 8 accum 8 for zero-OOM guarantee on T4
about 7 hours ago
verify_1_5b_arch.py
Safe
2.97 kB
Add 42-layer Viu-1.5B-MoE verification test suite
9 days ago
viu1_dense.py
12.7 kB
fix: restore grad_ckpt and safe micro 8 accum 8 for zero-OOM guarantee on T4
about 7 hours ago
viu1_moe.py
17.8 kB
P0 fix: gradient checkpointing, chunked cross-entropy, and vectorized MoE dispatch
1 day ago
viu_moe.py
24.5 kB
Vectorize MoE dispatch (argsort+bincount): eliminate 2,688 CUDA kernel launches per pass
2 days ago