Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
ViuAI
/
Viu-1.5B-MoE
Like
0
Text Generation
PyTorch
Hindi
English
Mixture of Experts
mixture-of-experts
indic
hindi
hinglish
mla
deepseek-v3
gemma-2
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
Viu-1.5B-MoE
6.17 MB
Ctrl+K
Ctrl+K
1 contributor
History:
163 commits
ViuAI
Optimize train_kaggle_t4.yaml: micro-batch 8 / accum 16 (8 passes/GPU on Dual T4), log_every 1
0914e9d
verified
6 days ago
data
Fix CUDA OOM on T4: cast model to amp_dtype directly, batch_size=2, safe RMSNorm
7 days ago
docs
Sync docs/PROGRESS.md after full codebase audit, retry hardening, and token security fixes
7 days ago
model
Optimize train_kaggle_t4.yaml: micro-batch 8 / accum 16 (8 passes/GPU on Dual T4), log_every 1
6 days ago
scratch
Fix CUDA OOM on T4: cast model to amp_dtype directly, batch_size=2, safe RMSNorm
7 days ago
tokenizer
Audit fix (grad_ckpt, 1392 comments, smoke experts, tokenizer mask) tokenizer/configs/tokenizer_config.json
12 days ago
.gitattributes
Safe
1.52 kB
initial commit
13 days ago
.gitignore
Safe
240 Bytes
Audit fix (grad_ckpt, 1392 comments, smoke experts, tokenizer mask) .gitignore
12 days ago
KAGGLE_INGEST_28GB_CORPUS.ipynb
Safe
21.3 kB
Sync KAGGLE_INGEST_28GB_CORPUS.ipynb - Kaggle 28GB ingestion notebook and docs
11 days ago
NOTES.md
Safe
40.9 kB
Auto-inject HF write token into RUN_ON_CLOUD.ipynb & record Section 19 directive in NOTES.md
7 days ago
README.md
Safe
2.9 kB
Fix CUDA OOM on T4: cast model to amp_dtype directly, batch_size=2, safe RMSNorm
7 days ago
RUN_ON_CLOUD.ipynb
Safe
11.9 kB
Auto-inject HF write token into RUN_ON_CLOUD.ipynb & record Section 19 directive in NOTES.md
7 days ago