Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution
Abstract
Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original passages but falls to 53.1% after paraphrasing and 29.0% after style rewriting. Generated-versus-human detection remains close to perfect against the tested human comparison set. We then compare two ways of selecting generated examples over three rounds of generation and retraining. One uses source information. The other uses a score from a separate reference model. The two rules select different examples, but the planned comparison does not detect a stable difference in the degradation of the resulting models. The results show that identifying where data came from and identifying which data are useful for training are separate problems. The experiment therefore separates source identity, criterion-facing selection, and recursive training outcome: neither the provenance score nor the tested criterion-facing proxy is established as sufficient for future recursive behaviour.
Community
Repeated training on model-generated data can degrade later models, and one response is to use provenance when choosing which generated examples to reuse. This paper tests both halves of that idea on financial-risk text. Generator attribution is 98.7% accurate on original passages, 53.1% after paraphrasing and 29.0% after style rewriting, while generated-versus-human detection stays near perfect. Over three rounds of generation and retraining, a provenance-based selection rule and a reference-model score select different examples, but the planned comparison does not detect a stable difference in degradation. Identifying where data came from and identifying which data are useful for training are separate problems. A separate theoretical treatment frames the same issue in terms of task-relative information contracts: https://doi.org/10.5281/zenodo.22819849
I want provenance filtering to work, and I don't think it survives contact with a real pipeline. Attribution holds on raw generations, then mostly evaporates after a paraphrase pass — which is step one of nearly every prep script I've written. If the signal dies at the first transform, lineage tracking is a lot of bookkeeping for a label that's already gone.
The title is the honest part: knowing a passage came from a model tells you nothing about whether it's worth training on. Before I wire lineage into anything, I'd want the ablation — provenance filtering against a plain quality filter plus dedup, matched on data volume and compute. If it doesn't win there, it's metadata, not a filter.
That comparison is in the paper. Section VII runs the provenance rule against a reference-model perplexity rule, which is the plain quality filter you describe, with a random draw and a fixed draw as controls. Every rule keeps 80 of 160 candidates on the same fine-tuning budget, over three rounds of generation and retraining.
It did not win. The two rules chose substantially different examples, fewer than half shared at each round, but the planned comparison found no stable difference in held-out degradation (p = 0.746). There was no dedup arm. The result does not show the rules are equivalent either, only that neither score was shown to carry the information that predicts training effect. That is the point of the title.
Get this paper in your agent:
hf papers read 2610.00417 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper