Papers
arxiv:2605.21318

TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization

Published on Oct 5
ยท Submitted by
Lucheng Fu
on Oct 6
Authors:
,
,
,
,
,
,

Abstract

Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimization methods iteratively rewrite prompts using LLM-generated feedback, but the resulting prompts often become longer, accumulate narrow sample-specific rules, and generalize poorly beyond the training distribution. We study this failure mode as prompt distributional overfitting and argue that it reflects a lack of representation control in discrete text-space optimization. We formalize this view through representational inefficiency, a dual-factor measure that decomposes prompt inefficiency into capacity cost and scope narrowness, attributing distributional prompt overfitting to their coupled growth during optimization. We propose TextReg, a regularization framework that realizes a soft-penalty objective through regularized textual gradients, combining Dual-Evidence Gradient Purification, Semantic Edit Regularization, and Regularization-Guided Prompt Update. Across multiple reasoning benchmarks, TextReg substantially improves out-of-distribution (OOD) generalization, with accuracy gains of up to +11.8% over TextGrad and +16.5% over REVOLVE.

Community

Paper submitter

Can prompts overfit just like machine learning models do?

Prompt optimization has emerged as a powerful paradigm for improving LLM performance. However, we observed a recurring phenomenon: as optimization progresses, prompts often become longer, accumulate increasingly specific instructions, and generalize worse beyond the training distribution.

In this work, we study this failure mode as prompt distributional overfitting.

Rather than viewing prompts as arbitrary text, we treat them as structured representations of task knowledge and ask a fundamental question:

What makes an optimized prompt generalize?

To answer this question, we introduce representational inefficiency, a perspective that attributes prompt overfitting to the coupled growth of:

๐Ÿ”น Capacity Cost โ€” prompts becoming unnecessarily long

๐Ÿ”น Scope Narrowness โ€” rules becoming increasingly specific and less reusable

Building on this insight, we propose TextReg, a regularization framework for text-space prompt optimization that combines:

โœ… Dual-Evidence Gradient Purification
โœ… Semantic Edit Regularization
โœ… Regularization-Guided Prompt Update

Across multiple reasoning benchmarks and model families, TextReg consistently improves out-of-distribution generalization, achieving gains of up to +11.8% over TextGrad and +16.5% over REVOLVE.

More broadly, we hope this work contributes to a better understanding of how prompts evolve during optimization and how regularization principles can be extended from parameter space to text space.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2605.21318
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2605.21318 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2605.21318 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2605.21318 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.