gemproject / README.md
Muhammed Rishdin T
Update README.md
4c9ff28 verified
|
Raw History Blame Contribute Delete
4.93 kB
---
license: openrail
---
----------------------------------------------------------------------
license: cc-by-4.0
task_categories:
- text-generation
- text-classification
language:
- en
- code
tags:
- llm
- historical
- legacy-models
- nlp
- dataset
size_categories:
- 10K<n<100K
pretty_name: "Legacy AI: Historical Training Data & Model Outputs"
---
# πŸ€— Legacy AI Dataset & Model Archive
This repository contains a curated collection of training data and outputs generated by older/generational Large Language Models (LLMs). It is designed for researchers interested in the evolution of AI syntax, training data drift analysis, and reproducing historical AI benchmarks.
![License](https://img.shields.io/badge/license-MIT-blue.svg)
![Python](https://img.shields.io/badge/python-3.8+-green.svg)
## πŸ“š Dataset Description
This dataset aggregates text data used to train early-era LLMs (e.g., GPT-2 era, early BERT, or specific domain legacy models). It serves as a snapshot of the data landscape and synthetic text generation capabilities from [Insert Time Period, e.g., 2015-2020].
### Key Features
- **Historical Context:** Data reflects the internet corpus and synthetic text styles from [Year/Period].
- **Comparison Ready:** Structured to allow direct comparison between legacy model outputs and modern SOTA models.
- **Cleaned & Pre-processed:** HTML tags removed, deduplicated, and tokenized (optional).
## πŸ“– Uses
### Direct Use
You can use this dataset to:
- Train small, efficient language models for retro-style text generation.
- Benchmark how modern models align with or deviate from old training distributions.
- Study data contamination and concept drift in NLP over time.
### Out-of-Scope Use
- Using this data to generate misinformation, spam, or malicious content.
- Assuming the data represents current-world knowledge (it is dated).
## πŸš€ How to Use
### Loading the Data
To load this dataset using the Hugging Face `datasets` library:
```python
from datasets import load_dataset
dataset = load_dataset("[YOUR_USERNAME]/[REPO_NAME]")
# Print the first example
print(dataset['train'][0])
```
### Using with Transformers (If applicable)
If this repo includes a model trained on this data:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "[YOUR_USERNAME]/[REPO_NAME]"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
input_text = "Once upon a time in early AI research..."
inputs = tokenizer(input_text, return_tensors="pt")
outputs = model.generate(**inputs, max_length=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
## πŸ“Š Dataset Structure
The dataset is provided in JSON/Parquet format with the following fields:
| Column Name | Type | Description |
| :--- | :--- | :--- |
| `text` | string | The raw training text or model output. |
| `source` | string | The origin of the data (e.g., 'CommonCrawl_2018', 'GPT-2_Output'). |
| `model_version` | string | If applicable, the specific legacy model version used. |
| `timestamp` | string | Approximate date of data creation. |
## πŸ› οΈ Dataset Creation
### Curation Rationale
We selected these specific data points because they represent the "state of the art" inputs/outputs from the previous decade of AI development. This helps in analyzing the trajectory of language model bias and complexity.
### Source Data
The data is aggregated from:
1. **[Source 1]:** e.g., OpenWebText (2019 snapshot).
2. **[Source 2]:** e.g., BooksCorpus.
3. **Synthetic:** Outputs sampled from `gpt-2` and `bert-base-uncased`.
### Processing Steps
1. **Filtering:** Removed low-quality text based on perplexity scores.
2. **Deduplication:** Applied MinHash LSH to remove near-duplicates.
3. **Privacy:** PII (Personally Identifiable Information) was scrubbed using regex patterns.
## βš–οΈ Bias, Risks, and Limitations
**Important Note:** Since this dataset is based on "Old AI" and older internet snapshots, it contains significant biases that were prevalent in those eras.
- **Stereotypes:** Gender and racial biases common in pre-2020 internet text are present.
- **Inaccuracy:** Factual information may be outdated.
- **Toxicity:** Older models had fewer safety guardrails; some generated text may be toxic or offensive.
## πŸ“ Citation
If you use this dataset in your research, please cite:
```bibtex
@dataset{your_name_2023,
author = {Your Name},
title = {Legacy AI: Historical Training Data & Model Outputs},
year = {2023},
publisher = {Hugging Face},
version = {1.0.0}
}
```
## 🀝 Acknowledgments
- Hugging Face for the `datasets` library.
- The original creators of the legacy models referenced in this dataset.
- The open-source community for data cleaning tools.
---
*For questions or feedback, please open an issue in the repo or contact [@muhammedrishdin].*