--- title: Predicting Memorization Before Fine-Tuning authors: Jaydeep Borkar, Niloofar Mireshghallah, and David A. Smith colorFrom: indigo colorTo: green sdk: gradio sdk_version: 5.49.1 app_file: app.py pinned: false license: apache-2.0 --- # Predicting Memorization Before Fine-Tuning — demo Interactive demo to the paper. A classifier trained only on pre-fine-tuning base-model features predicts which sequences a fine-tuned model will memorize. - **Explore the held-out examples** — pick a random sequence from a held-out set and estimate its memorization risk. - Datasets: FineWeb, PG-19, The Stack, OpenWebMath (all public). We do not include Enron and WildChat in the demo as they may contain sensitive text. - **Score your own text** — paste a passage and get the classifier's predicted memorization risk from base-model features. This is a forecast; there is no ground truth for arbitrary text. ## Files - `app.py` — Gradio app. - `data/lookup.parquet` — precomputed table (prefix, suffix, memorized flag, classifier score, base-model features) sampled from the held-out 1M run (all memorized examples up to a cap plus a random sample of non-memorized). - `models/clf_*.joblib` — the trained GradientBoostingClassifier + scaler per dataset. - `byo_features.py` — computes the six base-model features for arbitrary text (loads `EleutherAI/pythia-1.4b`; on ZeroGPU it is placed on `cuda` at startup per HF's guidance). - `assemble_lookup_data.py`, `train_classifiers.py` — offline scripts that produced the artifacts above (not needed at runtime; kept for reproducibility). ## Run locally ```bash pip install -r requirements.txt python app.py # opens a local Gradio URL ``` Locally the app loads Pythia-1.4B (~3 GB) at startup on a GPU if present, otherwise CPU; the `@spaces.GPU` decorator is a no-op off ZeroGPU. ## Notes - The classifier uses only base-model (pre-fine-tuning) features and was trained on a separate Run-1 fine-tuning run, then evaluated here on a disjoint Run-2 run. ## Data and model attribution The demo code in this repository is released under Apache-2.0. The text excerpts shown in the Explore tab are short passages drawn from public datasets and are displayed only to illustrate research findings; each dataset remains under its own upstream license, held by its original authors. - FineWeb, [HuggingFaceFW/fineweb](https://huggingface.co/datasets/HuggingFaceFW/fineweb), under ODC-By 1.0. - PG-19, [deepmind/pg19](https://huggingface.co/datasets/deepmind/pg19), public-domain books from Project Gutenberg. - The Stack, [bigcode/the-stack](https://huggingface.co/datasets/bigcode/the-stack), permissively licensed source code collected by the BigCode project. - OpenWebMath, [open-web-math/open-web-math](https://huggingface.co/datasets/open-web-math/open-web-math), openly released mathematical web text. - Base model Pythia-1.4B, [EleutherAI/pythia-1.4b](https://huggingface.co/EleutherAI/pythia-1.4b), under Apache-2.0. We thank the authors and maintainers of these datasets and of Pythia. If you are a rights holder and want an excerpt removed, please open a discussion on the Space and we will take it down.