--- language: - hi - en - multilingual language_bcp47: - hi-Latn license: apache-2.0 task_categories: - text-generation tags: - hinglish - hindi - pretraining - code - science - legal pretty_name: Viu Mini Raw Pretrain size_categories: - 100M ... <|assistant|> ...` templated for QA/instruct sources) - `lang`: `hinglish` | `hindi` | `english` | `bilingual` - `source`: origin dataset/generator name (e.g. `indiccorp_v2`, `fineweb-edu`, `vyakaran_master`, `sciq`) - `domain`: knowledge pillar (e.g. `general_hindi`, `electronics_smt_pcb`, `science_qa`, `hindi_vyakaran_sandhi`) - `safety_tag`: `safe` | `toxic` | `uncensored` ## Safety notice (intentional inclusion) `toxicity/` and `uncensored/` folders are **intentionally included** for robustness research (hate-speech detection, refusal training, red-teaming). They are TAGGED, not hidden: ```python from datasets import load_dataset ds = load_dataset("ViuAI/viu-mini-raw-pretrain", split="train", streaming=True) def keep(row, allow_toxic=False, allow_uncensored=False): tag = (row.get("safety_tag") or "safe").lower() if tag == "safe": return True if tag == "toxic": return allow_toxic if tag == "uncensored": return allow_uncensored return True # default pretraining: safe only ds_safe = ds.filter(lambda r: keep(r)) # robustness research: safe+toxic ds_research = ds.filter(lambda r: keep(r, allow_toxic=True)) # red-team only, controlled: everything ds_full = ds.filter(lambda r: keep(r, allow_toxic=True, allow_uncensored=True)) ``` Do NOT train user-facing models on `full` without alignment (SFT/DPO + refusal). ## Known issues (2026-09-22 audit, fixing in progress) 1. Hindi shards `train-00115..00178`, `train-00307..00370` missing (upload gaps) — backfilling. 2. `english/` 37x~1.6GB + `hindi/` 13x~1.15GB oversized shards — resharding to ~150MB/100k rows. 3. Legacy schema drift (`question/answer`, `instruction/response`, `topic/difficulty`, `title/text`) — migrating to unified schema via `data/scripts/standardize_to_unified.py`. 4. Viewer HTTP 500 on some offsets until (2)+(3) complete. 5. Duplicate curated pillar rows (`domains/` vs generated) — dedup via stable blake2b hash. See `docs/HF_REPAIR_RUNBOOK.md` in the model repo for repair commands. ## License / citation Apache-2.0. Upstream sources retain their own licenses (Wikipedia CC-BY-SA, Samanantar CC-BY, etc.). If you use this dataset, cite upstream sources plus `ViuAI/viu-mini-raw-pretrain`.