jaydeepb's picture
Update README.md
6916d7e verified
|
Raw History Blame Contribute Delete
3.2 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: Predicting Memorization Before Fine-Tuning
authors: Jaydeep Borkar, Niloofar Mireshghallah, and David A. Smith
colorFrom: indigo
colorTo: green
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
pinned: false
license: apache-2.0

Predicting Memorization Before Fine-Tuning β€” demo

Interactive demo to the paper. A classifier trained only on pre-fine-tuning base-model features predicts which sequences a fine-tuned model will memorize.

  • Explore the held-out examples β€” pick a random sequence from a held-out set and estimate its memorization risk.
  • Datasets: FineWeb, PG-19, The Stack, OpenWebMath (all public). We do not include Enron and WildChat in the demo as they may contain sensitive text.
  • Score your own text β€” paste a passage and get the classifier's predicted memorization risk from base-model features. This is a forecast; there is no ground truth for arbitrary text.

Files

  • app.py β€” Gradio app.
  • data/lookup.parquet β€” precomputed table (prefix, suffix, memorized flag, classifier score, base-model features) sampled from the held-out 1M run (all memorized examples up to a cap plus a random sample of non-memorized).
  • models/clf_*.joblib β€” the trained GradientBoostingClassifier + scaler per dataset.
  • byo_features.py β€” computes the six base-model features for arbitrary text (loads EleutherAI/pythia-1.4b; on ZeroGPU it is placed on cuda at startup per HF's guidance).
  • assemble_lookup_data.py, train_classifiers.py β€” offline scripts that produced the artifacts above (not needed at runtime; kept for reproducibility).

Run locally

pip install -r requirements.txt
python app.py            # opens a local Gradio URL

Locally the app loads Pythia-1.4B (~3 GB) at startup on a GPU if present, otherwise CPU; the @spaces.GPU decorator is a no-op off ZeroGPU.

Notes

  • The classifier uses only base-model (pre-fine-tuning) features and was trained on a separate Run-1 fine-tuning run, then evaluated here on a disjoint Run-2 run.

Data and model attribution

The demo code in this repository is released under Apache-2.0. The text excerpts shown in the Explore tab are short passages drawn from public datasets and are displayed only to illustrate research findings; each dataset remains under its own upstream license, held by its original authors.

We thank the authors and maintainers of these datasets and of Pythia. If you are a rights holder and want an excerpt removed, please open a discussion on the Space and we will take it down.