--- language: - hi - en - multilingual language_bcp47: - hi-Latn license: apache-2.0 task_categories: - text-generation tags: - hinglish - hindi - pretraining - code - science - legal - distillation - reasoning pretty_name: Viu Mini Raw Pretrain size_categories: - 100M ... <|assistant|> ...` templated for QA/instruct/distillation sources) - `lang`: `hinglish` | `hindi` | `english` | `bilingual` - `source`: origin dataset name (e.g. `indiccorp_v2`, `fineweb-edu`, `bespoke-stratos-r1`, `smoltalk`, `numina_math_cot`) - `domain`: knowledge pillar (e.g. `general_hindi`, `distilled_reasoning`, `distilled_math_cot`, `sports_cricket`, `automobile_rto`) - `safety_tag`: `safe` | `toxic` | `uncensored` ## Distillation & Reasoning Corpus (`distilled/`) Includes 33 high-quality reasoning shards (3.73 GB, ~1.7B tokens) distilled from frontier models (DeepSeek-R1, SmolTalk, NuminaMath, Magpie Llama 3.1, FineTome 100k, Bespoke Stratos R1) with native ` ... ` chain-of-thought isolation. ## Safety notice (intentional inclusion) `toxicity/` and `uncensored/` folders are **intentionally included** for robustness research (hate-speech detection, refusal training, red-teaming). They are TAGGED, not hidden: ```python from datasets import load_dataset ds = load_dataset("ViuAI/viu-mini-raw-pretrain", split="train", streaming=True) def keep(row, allow_toxic=False, allow_uncensored=False): tag = (row.get("safety_tag") or "safe").lower() if tag == "safe": return True if tag == "toxic": return allow_toxic if tag == "uncensored": return allow_uncensored return True # default pretraining: safe only ds_safe = ds.filter(lambda r: keep(r)) # robustness research: safe+toxic ds_research = ds.filter(lambda r: keep(r, allow_toxic=True)) # red-team only, controlled: everything ds_full = ds.filter(lambda r: keep(r, allow_toxic=True, allow_uncensored=True)) ``` Do NOT train user-facing models on `full` without alignment (SFT/DPO + refusal). ## Repair & Clean Status (2026-09-23) 1. Hindi missing shards: 100% Backfilled (447 shards total). 2. All oversized files resharded to 150MB standard shards (`english_fixed/` & `hindi_fixed/`). 3. 50 legacy oversized files deleted — ZERO storage duplication. 4. Schema unified across all domains (`grammar`, `math`, `science`, `toxicity`, `uncensored`, `domains`, `distilled`). 5. Frontier distillation corpus: 33 shards (3.73 GB) integrated into `distilled/`. 6. Net clean dataset size: ~229.5 GB (>105.1B tokens across 1,926 clean files). ## License / citation Apache-2.0. Upstream sources retain their own licenses (Wikipedia CC-BY-SA, Samanantar CC-BY, etc.). If you use this dataset, cite upstream sources plus `ViuAI/viu-mini-raw-pretrain`.