|
Download docs/DATASET_CARD_draft.md from ViuAI/ViuMini-Dense-360M: direct link, hf CLI and curl.
- Browser
- Download file 3.42 kB
-
https://huggingface.co/ViuAI/ViuMini-Dense-360M/resolve/main/docs/DATASET_CARD_draft.md
- Command line
-
hf download hf://ViuAI/ViuMini-Dense-360M/docs/DATASET_CARD_draft.md
-
curl -L -o DATASET_CARD_draft.md https://huggingface.co/ViuAI/ViuMini-Dense-360M/resolve/main/docs/DATASET_CARD_draft.md
3.42 kB
| language: | |
| - hi | |
| - en | |
| - multilingual | |
| language_bcp47: | |
| - hi-Latn | |
| license: apache-2.0 | |
| task_categories: | |
| - text-generation | |
| tags: | |
| - hinglish | |
| - hindi | |
| - pretraining | |
| - code | |
| - science | |
| - legal | |
| pretty_name: Viu Mini Raw Pretrain | |
| size_categories: | |
| - 100M<n<1B | |
| configs: | |
| - config_name: default | |
| data_files: | |
| - split: train | |
| path: | |
| - hindi/train-*.parquet | |
| - hindi_fixed/train-*.parquet | |
| - english_fixed/train-*.parquet | |
| - hinglish/train-*.parquet | |
| - domains/train-*.parquet | |
| - translation/train-*.parquet | |
| - translation/samanantar-*.parquet | |
| - wikipedia/train-*.parquet | |
| - science/*.parquet | |
| - math/*.parquet | |
| - code/*.parquet | |
| - legal/train-*.parquet | |
| - health/*.parquet | |
| - stories/*.parquet | |
| - grammar/*.parquet | |
| - finance/*.parquet | |
| - dictionary/*.parquet | |
| - toxicity/*.parquet | |
| - uncensored/*.parquet | |
| # ViuAI/viu-mini-raw-pretrain | |
| Raw pretraining corpus for **ViuMini-MoE-242M** (Hinglish-first; language-mix rule OPTIONAL, default OFF β training config me `enforce_mix` dekho). | |
| Unified schema: `text, lang, source, domain, safety_tag` (older files still migrating β see Known issues). | |
| - `text`: training text (plain, or `<|user|> ... <|assistant|> ...` templated for QA/instruct sources) | |
| - `lang`: `hinglish` | `hindi` | `english` | `bilingual` | |
| - `source`: origin dataset/generator name (e.g. `indiccorp_v2`, `fineweb-edu`, `vyakaran_master`, `sciq`) | |
| - `domain`: knowledge pillar (e.g. `general_hindi`, `electronics_smt_pcb`, `science_qa`, `hindi_vyakaran_sandhi`) | |
| - `safety_tag`: `safe` | `toxic` | `uncensored` | |
| ## Safety notice (intentional inclusion) | |
| `toxicity/` and `uncensored/` folders are **intentionally included** for robustness research | |
| (hate-speech detection, refusal training, red-teaming). They are TAGGED, not hidden: | |
| ```python | |
| from datasets import load_dataset | |
| ds = load_dataset("ViuAI/viu-mini-raw-pretrain", split="train", streaming=True) | |
| def keep(row, allow_toxic=False, allow_uncensored=False): | |
| tag = (row.get("safety_tag") or "safe").lower() | |
| if tag == "safe": return True | |
| if tag == "toxic": return allow_toxic | |
| if tag == "uncensored": return allow_uncensored | |
| return True | |
| # default pretraining: safe only | |
| ds_safe = ds.filter(lambda r: keep(r)) | |
| # robustness research: safe+toxic | |
| ds_research = ds.filter(lambda r: keep(r, allow_toxic=True)) | |
| # red-team only, controlled: everything | |
| ds_full = ds.filter(lambda r: keep(r, allow_toxic=True, allow_uncensored=True)) | |
| ``` | |
| Do NOT train user-facing models on `full` without alignment (SFT/DPO + refusal). | |
| ## Known issues (2026-09-22 audit, fixing in progress) | |
| 1. Hindi shards `train-00115..00178`, `train-00307..00370` missing (upload gaps) β backfilling. | |
| 2. `english/` 37x~1.6GB + `hindi/` 13x~1.15GB oversized shards β resharding to ~150MB/100k rows. | |
| 3. Legacy schema drift (`question/answer`, `instruction/response`, `topic/difficulty`, `title/text`) β | |
| migrating to unified schema via `data/scripts/standardize_to_unified.py`. | |
| 4. Viewer HTTP 500 on some offsets until (2)+(3) complete. | |
| 5. Duplicate curated pillar rows (`domains/` vs generated) β dedup via stable blake2b hash. | |
| See `docs/HF_REPAIR_RUNBOOK.md` in the model repo for repair commands. | |
| ## License / citation | |
| Apache-2.0. Upstream sources retain their own licenses (Wikipedia CC-BY-SA, Samanantar CC-BY, etc.). | |
| If you use this dataset, cite upstream sources plus `ViuAI/viu-mini-raw-pretrain`. | |