Papers
arxiv:2610.00417

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

Published on Sep 30
· Submitted by
Joss Armstrong
on Oct 7
Authors:

Abstract

Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original passages but falls to 53.1% after paraphrasing and 29.0% after style rewriting. Generated-versus-human detection remains close to perfect against the tested human comparison set. We then compare two ways of selecting generated examples over three rounds of generation and retraining. One uses source information. The other uses a score from a separate reference model. The two rules select different examples, but the planned comparison does not detect a stable difference in the degradation of the resulting models. The results show that identifying where data came from and identifying which data are useful for training are separate problems. The experiment therefore separates source identity, criterion-facing selection, and recursive training outcome: neither the provenance score nor the tested criterion-facing proxy is established as sufficient for future recursive behaviour.

Community

Paper author Paper submitter

Repeated training on model-generated data can degrade later models, and one response is to use provenance when choosing which generated examples to reuse. This paper tests both halves of that idea on financial-risk text. Generator attribution is 98.7% accurate on original passages, 53.1% after paraphrasing and 29.0% after style rewriting, while generated-versus-human detection stays near perfect. Over three rounds of generation and retraining, a provenance-based selection rule and a reference-model score select different examples, but the planned comparison does not detect a stable difference in degradation. Identifying where data came from and identifying which data are useful for training are separate problems. A separate theoretical treatment frames the same issue in terms of task-relative information contracts: https://doi.org/10.5281/zenodo.22819849

I want provenance filtering to work, and I don't think it survives contact with a real pipeline. Attribution holds on raw generations, then mostly evaporates after a paraphrase pass — which is step one of nearly every prep script I've written. If the signal dies at the first transform, lineage tracking is a lot of bookkeeping for a label that's already gone.

The title is the honest part: knowing a passage came from a model tells you nothing about whether it's worth training on. Before I wire lineage into anything, I'd want the ablation — provenance filtering against a plain quality filter plus dedup, matched on data volume and compute. If it doesn't win there, it's metadata, not a filter.

Paper author Paper submitter

That comparison is in the paper. Section VII runs the provenance rule against a reference-model perplexity rule, which is the plain quality filter you describe, with a random draw and a fixed draw as controls. Every rule keeps 80 of 160 candidates on the same fine-tuning budget, over three rounds of generation and retraining.
It did not win. The two rules chose substantially different examples, fewer than half shared at each round, but the planned comparison found no stable difference in held-out degradation (p = 0.746). There was no dedup arm. The result does not show the rules are equivalent either, only that neither score was shown to carry the information that predicts training effect. That is the point of the title.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.00417
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.00417 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.00417 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.00417 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.