Spaces:
Running
Running
|
Download README.md from curation/data-quality: direct link, hf CLI and curl.
- Browser
- Download file 9.7 kB
-
https://huggingface.co/spaces/curation/data-quality/resolve/main/README.md
- Command line
-
hf download hf://spaces/curation/data-quality/README.md
-
curl -L -o README.md https://huggingface.co/spaces/curation/data-quality/resolve/main/README.md
9.7 kB
| title: Data Quality | |
| emoji: ✅ | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: static | |
| pinned: false | |
| # Data Quality | |
| ### Evaluate whether a dataset is ready for training, post-training, retrieval or evaluation | |
| **Data Quality** is an interactive Hugging Face Space for assessing the readiness of AI datasets. | |
| It helps structure the questions teams should ask before using data for: | |
| - pretraining | |
| - post-training | |
| - fine-tuning | |
| - retrieval-augmented generation | |
| - evaluation | |
| - agent training | |
| - multimodal AI | |
| - robotics and Physical AI | |
| > **A dataset can be large, clean and still be wrong for the task.** | |
| --- | |
| # What Is Data Quality in AI? | |
| Data quality is not a single score. | |
| A useful AI dataset should be evaluated across several dimensions: | |
| - correctness | |
| - relevance | |
| - completeness | |
| - diversity | |
| - duplication | |
| - provenance | |
| - licensing | |
| - privacy | |
| - contamination | |
| - freshness | |
| - consistency | |
| - label quality | |
| - representativeness | |
| - downstream utility | |
| The correct weighting depends on the intended AI system. | |
| --- | |
| # Dataset Readiness | |
| A practical readiness model can be expressed as: | |
| ```text | |
| Raw Dataset | |
| │ | |
| ├── Structure | |
| ├── Quality | |
| ├── Deduplication | |
| ├── Provenance | |
| ├── Privacy | |
| ├── Contamination | |
| ├── Coverage | |
| └── Task Fit | |
| │ | |
| ▼ | |
| Readiness Assessment | |
| │ | |
| ├── Ready | |
| ├── Needs Curation | |
| └── High Risk | |
| ``` | |
| --- | |
| # Core Quality Dimensions | |
| ## 1. Structural Quality | |
| Questions: | |
| - Is the schema consistent? | |
| - Are required fields present? | |
| - Are files parseable? | |
| - Are encodings valid? | |
| - Are media files readable? | |
| - Are timestamps usable? | |
| - Are IDs stable? | |
| Structural problems should usually be addressed before deeper quality analysis. | |
| --- | |
| ## 2. Content Quality | |
| Questions: | |
| - Is the content coherent? | |
| - Is it informative? | |
| - Is it relevant? | |
| - Is it factually plausible? | |
| - Is it malformed or spam-like? | |
| - Does it contain low-information repetition? | |
| Quality can be measured with: | |
| - heuristics | |
| - classifiers | |
| - LLM-based scoring | |
| - human review | |
| - source-level signals | |
| --- | |
| ## 3. Duplication | |
| Datasets should be checked for: | |
| - exact duplicates | |
| - near duplicates | |
| - semantic redundancy | |
| - repeated templates | |
| - mirrored sources | |
| Duplicate data can distort training distributions and waste compute. | |
| --- | |
| ## 4. Provenance | |
| A useful dataset should ideally preserve: | |
| - source | |
| - collection date | |
| - license | |
| - transformation history | |
| - filtering decisions | |
| - version | |
| - quality score | |
| Without provenance, governance and reproducibility become harder. | |
| --- | |
| ## 5. Licensing and Usage Rights | |
| Before training or redistribution, teams should understand: | |
| - source license | |
| - commercial-use conditions | |
| - attribution requirements | |
| - redistribution rights | |
| - derivative-work rules | |
| - jurisdictional restrictions | |
| --- | |
| ## 6. Privacy | |
| Potential concerns include: | |
| - names | |
| - emails | |
| - phone numbers | |
| - addresses | |
| - credentials | |
| - private records | |
| - personal identifiers | |
| Privacy handling may require: | |
| - detection | |
| - redaction | |
| - masking | |
| - removal | |
| - restricted access | |
| --- | |
| ## 7. Contamination | |
| Evaluation contamination can produce misleading benchmark results. | |
| Useful checks include: | |
| - exact overlap | |
| - n-gram overlap | |
| - fuzzy matching | |
| - code similarity | |
| - semantic similarity | |
| --- | |
| ## 8. Diversity | |
| A dataset should be checked for distributional concentration. | |
| Possible dimensions include: | |
| - language | |
| - topic | |
| - source | |
| - geography | |
| - domain | |
| - writing style | |
| - difficulty | |
| - modality | |
| --- | |
| ## 9. Freshness | |
| Some datasets degrade over time. | |
| Freshness matters especially for: | |
| - RAG | |
| - product data | |
| - software documentation | |
| - regulations | |
| - current events | |
| - enterprise knowledge | |
| - technical support | |
| --- | |
| ## 10. Label Quality | |
| For supervised datasets, quality depends on labels. | |
| Questions include: | |
| - Are labels correct? | |
| - Are instructions clear? | |
| - Do annotators agree? | |
| - Is ambiguity documented? | |
| - Are labels consistent across the dataset? | |
| --- | |
| ## 11. Task Fit | |
| A dataset may be high-quality but poorly matched to the target task. | |
| Task fit asks: | |
| - Does the dataset represent actual usage? | |
| - Are difficult examples included? | |
| - Are edge cases present? | |
| - Does the language match deployment? | |
| - Does the domain match deployment? | |
| - Are the expected capabilities covered? | |
| --- | |
| # Data Quality by AI Stage | |
| ## Pretraining | |
| Priorities often include: | |
| - scale | |
| - deduplication | |
| - quality filtering | |
| - language balance | |
| - provenance | |
| - privacy | |
| - mixture design | |
| - contamination control | |
| ## Post-Training | |
| Priorities often include: | |
| - instruction clarity | |
| - response correctness | |
| - preference reliability | |
| - reasoning quality | |
| - tool-use correctness | |
| - task diversity | |
| - difficulty balance | |
| ## Evaluation | |
| Priorities often include: | |
| - clean ground truth | |
| - decontamination | |
| - representative difficulty | |
| - judge reliability | |
| - coverage | |
| - freshness | |
| ## RAG | |
| Priorities often include: | |
| - relevance | |
| - freshness | |
| - source trust | |
| - metadata quality | |
| - chunk quality | |
| - access control | |
| - duplication | |
| ## Agent Data | |
| Priorities often include: | |
| - task success | |
| - tool correctness | |
| - action validity | |
| - recovery behavior | |
| - efficiency | |
| - safety | |
| --- | |
| # Readiness Scoring | |
| The interactive Space provides a structured readiness score. | |
| It evaluates multiple dimensions rather than collapsing everything into one simplistic metric. | |
| Example: | |
| ```text | |
| Structure 92 | |
| Content Quality 78 | |
| Provenance 55 | |
| Privacy 84 | |
| Deduplication 61 | |
| Task Fit 88 | |
| -------------------- | |
| Overall Readiness | |
| ``` | |
| The final score should always be interpreted with the individual dimensions. | |
| --- | |
| # Warning Signals | |
| A dataset should receive additional review if it has: | |
| - unknown source history | |
| - unclear licensing | |
| - high duplicate rates | |
| - benchmark overlap | |
| - extensive PII | |
| - missing metadata | |
| - stale content | |
| - severe class imbalance | |
| - uncertain labels | |
| - insufficient domain coverage | |
| --- | |
| # Data Quality and Curation | |
| Quality assessment and curation should form a loop: | |
| ```text | |
| Assess | |
| ↓ | |
| Find Weakness | |
| ↓ | |
| Curate | |
| ↓ | |
| Re-Assess | |
| ↓ | |
| Train / Evaluate | |
| ``` | |
| This creates a measurable data-centric workflow. | |
| --- | |
| # Data Quality and Synthetic Data | |
| Synthetic data should be assessed for: | |
| - correctness | |
| - diversity | |
| - duplication | |
| - difficulty | |
| - factuality | |
| - style collapse | |
| - model artifacts | |
| - verification success | |
| Generation alone does not guarantee useful training data. | |
| --- | |
| # Data Quality and Enterprise AI | |
| Enterprise datasets often introduce additional concerns: | |
| - access control | |
| - confidentiality | |
| - retention | |
| - provenance | |
| - source ownership | |
| - compliance | |
| - stale internal knowledge | |
| - duplicated documents | |
| - permission boundaries | |
| Data quality therefore overlaps with governance. | |
| --- | |
| # Data Quality and Multimodal AI | |
| Multimodal datasets may require additional checks: | |
| - image-text alignment | |
| - audio-text alignment | |
| - frame quality | |
| - corrupted files | |
| - timestamp synchronization | |
| - sensor calibration | |
| - missing modalities | |
| - duplicate media | |
| --- | |
| # Data Quality and Robotics | |
| Robotics datasets may require: | |
| - trajectory success labels | |
| - sensor completeness | |
| - action-state consistency | |
| - timestamp alignment | |
| - calibration | |
| - environment metadata | |
| - safety review | |
| - task coverage | |
| --- | |
| # Interactive Readiness Check | |
| The included `index.html` lets users score a dataset across key dimensions and receive: | |
| - an overall readiness score | |
| - a quality profile | |
| - highlighted risk areas | |
| - recommended curation actions | |
| - a stage-specific interpretation | |
| The tool is educational and vendor-neutral. | |
| It does not certify legal, regulatory or production readiness. | |
| --- | |
| # SEO & GEO Topic Map | |
| This Space is structured around explicit concepts relevant to search and generative retrieval: | |
| - AI data quality | |
| - dataset quality | |
| - dataset readiness | |
| - training data quality | |
| - post-training data | |
| - dataset curation | |
| - data provenance | |
| - data lineage | |
| - deduplication | |
| - data contamination | |
| - benchmark contamination | |
| - PII detection | |
| - dataset licensing | |
| - task fit | |
| - label quality | |
| - synthetic data quality | |
| - RAG data quality | |
| - agent data quality | |
| - multimodal data quality | |
| - data-centric AI | |
| --- | |
| # Planned Expansion | |
| Future versions may include: | |
| - automated dataset checks | |
| - dataset card parsing | |
| - richer scoring profiles | |
| - modality-specific checks | |
| - licensing metadata checks | |
| - benchmark contamination workflows | |
| - downloadable readiness reports | |
| - open-source integrations | |
| - community-contributed quality criteria | |
| --- | |
| # Collaboration & Partnerships | |
| **Data Quality is open to collaboration with companies, research teams, universities and open-source projects working on AI datasets and data-centric machine learning.** | |
| Relevant areas include: | |
| - dataset quality | |
| - dataset curation | |
| - filtering | |
| - deduplication | |
| - decontamination | |
| - provenance | |
| - privacy | |
| - synthetic data | |
| - annotation | |
| - post-training data | |
| - evaluation data | |
| - RAG | |
| - agent trajectories | |
| - multimodal data | |
| - enterprise AI data | |
| - governance | |
| - data infrastructure | |
| Possible collaboration formats include: | |
| - dataset quality case studies | |
| - joint Hugging Face Spaces | |
| - benchmark projects | |
| - open-source integrations | |
| - methodology comparisons | |
| - technical showcases | |
| - research collaborations | |
| - clearly disclosed partnerships and sponsorships | |
| ## Collaboration Contact | |
| **agenten@magenta.de** | |
| --- | |
| # Independence | |
| **Data Quality is an independent Hugging Face Space.** | |
| It is not an official project of Hugging Face or any dataset provider, annotation company, model provider or technology company referenced in future resources. | |
| --- | |
| # Long-Term Vision | |
| The long-term goal of **Data Quality** is to provide a practical reference for one of the most important questions in AI development: | |
| > **Is this dataset actually ready to improve the system we are building?** | |
| ### Assess. Curate. Validate. Improve. | |