--- title: Data Quality emoji: ✅ colorFrom: blue colorTo: indigo sdk: static pinned: false --- # Data Quality ### Evaluate whether a dataset is ready for training, post-training, retrieval or evaluation **Data Quality** is an interactive Hugging Face Space for assessing the readiness of AI datasets. It helps structure the questions teams should ask before using data for: - pretraining - post-training - fine-tuning - retrieval-augmented generation - evaluation - agent training - multimodal AI - robotics and Physical AI > **A dataset can be large, clean and still be wrong for the task.** --- # What Is Data Quality in AI? Data quality is not a single score. A useful AI dataset should be evaluated across several dimensions: - correctness - relevance - completeness - diversity - duplication - provenance - licensing - privacy - contamination - freshness - consistency - label quality - representativeness - downstream utility The correct weighting depends on the intended AI system. --- # Dataset Readiness A practical readiness model can be expressed as: ```text Raw Dataset │ ├── Structure ├── Quality ├── Deduplication ├── Provenance ├── Privacy ├── Contamination ├── Coverage └── Task Fit │ ▼ Readiness Assessment │ ├── Ready ├── Needs Curation └── High Risk ``` --- # Core Quality Dimensions ## 1. Structural Quality Questions: - Is the schema consistent? - Are required fields present? - Are files parseable? - Are encodings valid? - Are media files readable? - Are timestamps usable? - Are IDs stable? Structural problems should usually be addressed before deeper quality analysis. --- ## 2. Content Quality Questions: - Is the content coherent? - Is it informative? - Is it relevant? - Is it factually plausible? - Is it malformed or spam-like? - Does it contain low-information repetition? Quality can be measured with: - heuristics - classifiers - LLM-based scoring - human review - source-level signals --- ## 3. Duplication Datasets should be checked for: - exact duplicates - near duplicates - semantic redundancy - repeated templates - mirrored sources Duplicate data can distort training distributions and waste compute. --- ## 4. Provenance A useful dataset should ideally preserve: - source - collection date - license - transformation history - filtering decisions - version - quality score Without provenance, governance and reproducibility become harder. --- ## 5. Licensing and Usage Rights Before training or redistribution, teams should understand: - source license - commercial-use conditions - attribution requirements - redistribution rights - derivative-work rules - jurisdictional restrictions --- ## 6. Privacy Potential concerns include: - names - emails - phone numbers - addresses - credentials - private records - personal identifiers Privacy handling may require: - detection - redaction - masking - removal - restricted access --- ## 7. Contamination Evaluation contamination can produce misleading benchmark results. Useful checks include: - exact overlap - n-gram overlap - fuzzy matching - code similarity - semantic similarity --- ## 8. Diversity A dataset should be checked for distributional concentration. Possible dimensions include: - language - topic - source - geography - domain - writing style - difficulty - modality --- ## 9. Freshness Some datasets degrade over time. Freshness matters especially for: - RAG - product data - software documentation - regulations - current events - enterprise knowledge - technical support --- ## 10. Label Quality For supervised datasets, quality depends on labels. Questions include: - Are labels correct? - Are instructions clear? - Do annotators agree? - Is ambiguity documented? - Are labels consistent across the dataset? --- ## 11. Task Fit A dataset may be high-quality but poorly matched to the target task. Task fit asks: - Does the dataset represent actual usage? - Are difficult examples included? - Are edge cases present? - Does the language match deployment? - Does the domain match deployment? - Are the expected capabilities covered? --- # Data Quality by AI Stage ## Pretraining Priorities often include: - scale - deduplication - quality filtering - language balance - provenance - privacy - mixture design - contamination control ## Post-Training Priorities often include: - instruction clarity - response correctness - preference reliability - reasoning quality - tool-use correctness - task diversity - difficulty balance ## Evaluation Priorities often include: - clean ground truth - decontamination - representative difficulty - judge reliability - coverage - freshness ## RAG Priorities often include: - relevance - freshness - source trust - metadata quality - chunk quality - access control - duplication ## Agent Data Priorities often include: - task success - tool correctness - action validity - recovery behavior - efficiency - safety --- # Readiness Scoring The interactive Space provides a structured readiness score. It evaluates multiple dimensions rather than collapsing everything into one simplistic metric. Example: ```text Structure 92 Content Quality 78 Provenance 55 Privacy 84 Deduplication 61 Task Fit 88 -------------------- Overall Readiness ``` The final score should always be interpreted with the individual dimensions. --- # Warning Signals A dataset should receive additional review if it has: - unknown source history - unclear licensing - high duplicate rates - benchmark overlap - extensive PII - missing metadata - stale content - severe class imbalance - uncertain labels - insufficient domain coverage --- # Data Quality and Curation Quality assessment and curation should form a loop: ```text Assess ↓ Find Weakness ↓ Curate ↓ Re-Assess ↓ Train / Evaluate ``` This creates a measurable data-centric workflow. --- # Data Quality and Synthetic Data Synthetic data should be assessed for: - correctness - diversity - duplication - difficulty - factuality - style collapse - model artifacts - verification success Generation alone does not guarantee useful training data. --- # Data Quality and Enterprise AI Enterprise datasets often introduce additional concerns: - access control - confidentiality - retention - provenance - source ownership - compliance - stale internal knowledge - duplicated documents - permission boundaries Data quality therefore overlaps with governance. --- # Data Quality and Multimodal AI Multimodal datasets may require additional checks: - image-text alignment - audio-text alignment - frame quality - corrupted files - timestamp synchronization - sensor calibration - missing modalities - duplicate media --- # Data Quality and Robotics Robotics datasets may require: - trajectory success labels - sensor completeness - action-state consistency - timestamp alignment - calibration - environment metadata - safety review - task coverage --- # Interactive Readiness Check The included `index.html` lets users score a dataset across key dimensions and receive: - an overall readiness score - a quality profile - highlighted risk areas - recommended curation actions - a stage-specific interpretation The tool is educational and vendor-neutral. It does not certify legal, regulatory or production readiness. --- # SEO & GEO Topic Map This Space is structured around explicit concepts relevant to search and generative retrieval: - AI data quality - dataset quality - dataset readiness - training data quality - post-training data - dataset curation - data provenance - data lineage - deduplication - data contamination - benchmark contamination - PII detection - dataset licensing - task fit - label quality - synthetic data quality - RAG data quality - agent data quality - multimodal data quality - data-centric AI --- # Planned Expansion Future versions may include: - automated dataset checks - dataset card parsing - richer scoring profiles - modality-specific checks - licensing metadata checks - benchmark contamination workflows - downloadable readiness reports - open-source integrations - community-contributed quality criteria --- # Collaboration & Partnerships **Data Quality is open to collaboration with companies, research teams, universities and open-source projects working on AI datasets and data-centric machine learning.** Relevant areas include: - dataset quality - dataset curation - filtering - deduplication - decontamination - provenance - privacy - synthetic data - annotation - post-training data - evaluation data - RAG - agent trajectories - multimodal data - enterprise AI data - governance - data infrastructure Possible collaboration formats include: - dataset quality case studies - joint Hugging Face Spaces - benchmark projects - open-source integrations - methodology comparisons - technical showcases - research collaborations - clearly disclosed partnerships and sponsorships ## Collaboration Contact **agenten@magenta.de** --- # Independence **Data Quality is an independent Hugging Face Space.** It is not an official project of Hugging Face or any dataset provider, annotation company, model provider or technology company referenced in future resources. --- # Long-Term Vision The long-term goal of **Data Quality** is to provide a practical reference for one of the most important questions in AI development: > **Is this dataset actually ready to improve the system we are building?** ### Assess. Curate. Validate. Improve.