What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document
Abstract
Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We show that an observer at the split can rebuild most of the client's text from this traffic, and we measure how much the gradients help. On GPT-2, an attacker who holds only the publicly released weights of the client's layers recovers 94.20% of tokens from the activations alone and 97.38% when it also sees the gradients, 3.17 percentage points more 95% interval [2.72, 3.64]. Counted by document, the difference is much larger. The attacker rebuilds 13.71% of 32-token documents exactly without the gradients and 37.77% with them, because a document only counts when every token is right. How we count also changes how good a defence looks. Secret mixup, which blends each outgoing vector with a decoy, stops the attacker from rebuilding almost any document exactly, yet the attacker still recovers 83-91% of tokens. In a second experiment on GPT-2 and Qwen3-0.6B, where the server trains only a run of consecutive layers, the layer at which the run starts changes both model quality and leakage, even when the run's length is fixed. We recommend reporting leakage both per token and per document, and treating what a split model sends as being as sensitive as the text itself.
Community
Adding gradients to a text-reconstruction attack raises token recovery from 94.20% to 97.38% in a controlled GPT-2 split-training experiment. Yet exact recovery of entire 32-token documents jumps from 13.71% to 37.77%—almost 2.8×.
The metric also changes how defences look: secret mixup reduces exact document recovery to at most 1.48%, while 83–91%
of tokens remain recoverable.
These results show why split-model privacy evaluations should report leakage at both token and document level, with explicit assumptions about what the attacker can observe and knows.
Get this paper in your agent:
hf papers read 2610.04128 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper