jaydeepb's picture
Update README.md
6916d7e verified
|
Raw History Blame Contribute Delete
3.2 kB
---
title: Predicting Memorization Before Fine-Tuning
authors: Jaydeep Borkar, Niloofar Mireshghallah, and David A. Smith
colorFrom: indigo
colorTo: green
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
pinned: false
license: apache-2.0
---
# Predicting Memorization Before Fine-Tuning β€” demo
Interactive demo to the paper. A classifier trained only on pre-fine-tuning base-model features predicts which sequences a fine-tuned model will memorize.
- **Explore the held-out examples** β€” pick a random sequence from a held-out set and estimate its memorization risk.
- Datasets: FineWeb, PG-19, The Stack, OpenWebMath (all public). We do not include Enron and WildChat in the demo as they may contain sensitive text.
- **Score your own text** β€” paste a passage and get the classifier's predicted memorization
risk from base-model features. This is a forecast; there is no ground truth for arbitrary text.
## Files
- `app.py` β€” Gradio app.
- `data/lookup.parquet` β€” precomputed table (prefix, suffix, memorized flag, classifier score,
base-model features) sampled from the held-out 1M run (all memorized examples up to a cap plus a
random sample of non-memorized).
- `models/clf_*.joblib` β€” the trained GradientBoostingClassifier + scaler per dataset.
- `byo_features.py` β€” computes the six base-model features for arbitrary text (loads
`EleutherAI/pythia-1.4b`; on ZeroGPU it is placed on `cuda` at startup per HF's guidance).
- `assemble_lookup_data.py`, `train_classifiers.py` β€” offline scripts that produced the artifacts
above (not needed at runtime; kept for reproducibility).
## Run locally
```bash
pip install -r requirements.txt
python app.py # opens a local Gradio URL
```
Locally the app loads Pythia-1.4B (~3 GB) at startup on a GPU if present, otherwise CPU; the
`@spaces.GPU` decorator is a no-op off ZeroGPU.
## Notes
- The classifier uses only base-model (pre-fine-tuning) features and was trained on a separate
Run-1 fine-tuning run, then evaluated here on a disjoint Run-2 run.
## Data and model attribution
The demo code in this repository is released under Apache-2.0. The text excerpts shown in the
Explore tab are short passages drawn from public datasets and are displayed only to illustrate
research findings; each dataset remains under its own upstream license, held by its original
authors.
- FineWeb, [HuggingFaceFW/fineweb](https://huggingface.co/datasets/HuggingFaceFW/fineweb), under ODC-By 1.0.
- PG-19, [deepmind/pg19](https://huggingface.co/datasets/deepmind/pg19), public-domain books from Project Gutenberg.
- The Stack, [bigcode/the-stack](https://huggingface.co/datasets/bigcode/the-stack), permissively licensed source code collected by the BigCode project.
- OpenWebMath, [open-web-math/open-web-math](https://huggingface.co/datasets/open-web-math/open-web-math), openly released mathematical web text.
- Base model Pythia-1.4B, [EleutherAI/pythia-1.4b](https://huggingface.co/EleutherAI/pythia-1.4b), under Apache-2.0.
We thank the authors and maintainers of these datasets and of Pythia. If you are a rights holder
and want an excerpt removed, please open a discussion on the Space and we will take it down.