Papers
arxiv:2609.37040

Selecting The Most Informative Tokens in Natural Language Autoencoders

Published on Sep 29
· Submitted by
Federico Torrielli
on Sep 30
Authors:
,
,
,
,

Abstract

Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across 4.7 million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just 5% of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.

Community

Paper author Paper submitter

We ask a practical question about Natural Language Autoencoders (NLAs): if generating an explanation for every token is prohibitively expensive, which token positions should I inspect? We test this at scale, generating about 4.7 million explanations across four models and four auditing datasets covering prompt injection and concealed information.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.37040
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.37040 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.37040 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.37040 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.