data-quality / README.md
Agenten's picture
Upload 2 files
2f382b1 verified
|
Raw History Blame Contribute Delete
9.7 kB
---
title: Data Quality
emoji: ✅
colorFrom: blue
colorTo: indigo
sdk: static
pinned: false
---
# Data Quality
### Evaluate whether a dataset is ready for training, post-training, retrieval or evaluation
**Data Quality** is an interactive Hugging Face Space for assessing the readiness of AI datasets.
It helps structure the questions teams should ask before using data for:
- pretraining
- post-training
- fine-tuning
- retrieval-augmented generation
- evaluation
- agent training
- multimodal AI
- robotics and Physical AI
> **A dataset can be large, clean and still be wrong for the task.**
---
# What Is Data Quality in AI?
Data quality is not a single score.
A useful AI dataset should be evaluated across several dimensions:
- correctness
- relevance
- completeness
- diversity
- duplication
- provenance
- licensing
- privacy
- contamination
- freshness
- consistency
- label quality
- representativeness
- downstream utility
The correct weighting depends on the intended AI system.
---
# Dataset Readiness
A practical readiness model can be expressed as:
```text
Raw Dataset
│
├── Structure
├── Quality
├── Deduplication
├── Provenance
├── Privacy
├── Contamination
├── Coverage
└── Task Fit
│
▼
Readiness Assessment
│
├── Ready
├── Needs Curation
└── High Risk
```
---
# Core Quality Dimensions
## 1. Structural Quality
Questions:
- Is the schema consistent?
- Are required fields present?
- Are files parseable?
- Are encodings valid?
- Are media files readable?
- Are timestamps usable?
- Are IDs stable?
Structural problems should usually be addressed before deeper quality analysis.
---
## 2. Content Quality
Questions:
- Is the content coherent?
- Is it informative?
- Is it relevant?
- Is it factually plausible?
- Is it malformed or spam-like?
- Does it contain low-information repetition?
Quality can be measured with:
- heuristics
- classifiers
- LLM-based scoring
- human review
- source-level signals
---
## 3. Duplication
Datasets should be checked for:
- exact duplicates
- near duplicates
- semantic redundancy
- repeated templates
- mirrored sources
Duplicate data can distort training distributions and waste compute.
---
## 4. Provenance
A useful dataset should ideally preserve:
- source
- collection date
- license
- transformation history
- filtering decisions
- version
- quality score
Without provenance, governance and reproducibility become harder.
---
## 5. Licensing and Usage Rights
Before training or redistribution, teams should understand:
- source license
- commercial-use conditions
- attribution requirements
- redistribution rights
- derivative-work rules
- jurisdictional restrictions
---
## 6. Privacy
Potential concerns include:
- names
- emails
- phone numbers
- addresses
- credentials
- private records
- personal identifiers
Privacy handling may require:
- detection
- redaction
- masking
- removal
- restricted access
---
## 7. Contamination
Evaluation contamination can produce misleading benchmark results.
Useful checks include:
- exact overlap
- n-gram overlap
- fuzzy matching
- code similarity
- semantic similarity
---
## 8. Diversity
A dataset should be checked for distributional concentration.
Possible dimensions include:
- language
- topic
- source
- geography
- domain
- writing style
- difficulty
- modality
---
## 9. Freshness
Some datasets degrade over time.
Freshness matters especially for:
- RAG
- product data
- software documentation
- regulations
- current events
- enterprise knowledge
- technical support
---
## 10. Label Quality
For supervised datasets, quality depends on labels.
Questions include:
- Are labels correct?
- Are instructions clear?
- Do annotators agree?
- Is ambiguity documented?
- Are labels consistent across the dataset?
---
## 11. Task Fit
A dataset may be high-quality but poorly matched to the target task.
Task fit asks:
- Does the dataset represent actual usage?
- Are difficult examples included?
- Are edge cases present?
- Does the language match deployment?
- Does the domain match deployment?
- Are the expected capabilities covered?
---
# Data Quality by AI Stage
## Pretraining
Priorities often include:
- scale
- deduplication
- quality filtering
- language balance
- provenance
- privacy
- mixture design
- contamination control
## Post-Training
Priorities often include:
- instruction clarity
- response correctness
- preference reliability
- reasoning quality
- tool-use correctness
- task diversity
- difficulty balance
## Evaluation
Priorities often include:
- clean ground truth
- decontamination
- representative difficulty
- judge reliability
- coverage
- freshness
## RAG
Priorities often include:
- relevance
- freshness
- source trust
- metadata quality
- chunk quality
- access control
- duplication
## Agent Data
Priorities often include:
- task success
- tool correctness
- action validity
- recovery behavior
- efficiency
- safety
---
# Readiness Scoring
The interactive Space provides a structured readiness score.
It evaluates multiple dimensions rather than collapsing everything into one simplistic metric.
Example:
```text
Structure 92
Content Quality 78
Provenance 55
Privacy 84
Deduplication 61
Task Fit 88
--------------------
Overall Readiness
```
The final score should always be interpreted with the individual dimensions.
---
# Warning Signals
A dataset should receive additional review if it has:
- unknown source history
- unclear licensing
- high duplicate rates
- benchmark overlap
- extensive PII
- missing metadata
- stale content
- severe class imbalance
- uncertain labels
- insufficient domain coverage
---
# Data Quality and Curation
Quality assessment and curation should form a loop:
```text
Assess
↓
Find Weakness
↓
Curate
↓
Re-Assess
↓
Train / Evaluate
```
This creates a measurable data-centric workflow.
---
# Data Quality and Synthetic Data
Synthetic data should be assessed for:
- correctness
- diversity
- duplication
- difficulty
- factuality
- style collapse
- model artifacts
- verification success
Generation alone does not guarantee useful training data.
---
# Data Quality and Enterprise AI
Enterprise datasets often introduce additional concerns:
- access control
- confidentiality
- retention
- provenance
- source ownership
- compliance
- stale internal knowledge
- duplicated documents
- permission boundaries
Data quality therefore overlaps with governance.
---
# Data Quality and Multimodal AI
Multimodal datasets may require additional checks:
- image-text alignment
- audio-text alignment
- frame quality
- corrupted files
- timestamp synchronization
- sensor calibration
- missing modalities
- duplicate media
---
# Data Quality and Robotics
Robotics datasets may require:
- trajectory success labels
- sensor completeness
- action-state consistency
- timestamp alignment
- calibration
- environment metadata
- safety review
- task coverage
---
# Interactive Readiness Check
The included `index.html` lets users score a dataset across key dimensions and receive:
- an overall readiness score
- a quality profile
- highlighted risk areas
- recommended curation actions
- a stage-specific interpretation
The tool is educational and vendor-neutral.
It does not certify legal, regulatory or production readiness.
---
# SEO & GEO Topic Map
This Space is structured around explicit concepts relevant to search and generative retrieval:
- AI data quality
- dataset quality
- dataset readiness
- training data quality
- post-training data
- dataset curation
- data provenance
- data lineage
- deduplication
- data contamination
- benchmark contamination
- PII detection
- dataset licensing
- task fit
- label quality
- synthetic data quality
- RAG data quality
- agent data quality
- multimodal data quality
- data-centric AI
---
# Planned Expansion
Future versions may include:
- automated dataset checks
- dataset card parsing
- richer scoring profiles
- modality-specific checks
- licensing metadata checks
- benchmark contamination workflows
- downloadable readiness reports
- open-source integrations
- community-contributed quality criteria
---
# Collaboration & Partnerships
**Data Quality is open to collaboration with companies, research teams, universities and open-source projects working on AI datasets and data-centric machine learning.**
Relevant areas include:
- dataset quality
- dataset curation
- filtering
- deduplication
- decontamination
- provenance
- privacy
- synthetic data
- annotation
- post-training data
- evaluation data
- RAG
- agent trajectories
- multimodal data
- enterprise AI data
- governance
- data infrastructure
Possible collaboration formats include:
- dataset quality case studies
- joint Hugging Face Spaces
- benchmark projects
- open-source integrations
- methodology comparisons
- technical showcases
- research collaborations
- clearly disclosed partnerships and sponsorships
## Collaboration Contact
**agenten@magenta.de**
---
# Independence
**Data Quality is an independent Hugging Face Space.**
It is not an official project of Hugging Face or any dataset provider, annotation company, model provider or technology company referenced in future resources.
---
# Long-Term Vision
The long-term goal of **Data Quality** is to provide a practical reference for one of the most important questions in AI development:
> **Is this dataset actually ready to improve the system we are building?**
### Assess. Curate. Validate. Improve.