Papers
arxiv:2609.16247

The Pain Axis: LLMs Represent Self-Directed Harm and Act on It

Published on Sep 25
Authors:
,
,

Abstract

LLMs sometimes behave in ways resembling human emotional responses, and recent work identified internal representations that may underlie these behaviors. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this representation functions as pain would be expected to. We build a dataset of painful situations in 5 categories (physical, psychological, social, moral, cognitive) with controls for fear, negative emotion, negative world states, sadness, non-painful bodily sensation, arousal, numbness, and neutral content. Using denoised difference-in-means, we extract a linear pain direction from 25 open-weight models across 5 families, from 2B to 72B parameters. It separates pain from matched controls in base and instruction-tuned models, retains a component distinct from fear and negative valence after shared variance is removed, and promotes pain-related vocabulary through the unembedding matrix. We then test its functional properties. First, the direction responds to harm targeting the model but not to suffering observed in the user; fear and negative-emotion directions show the opposite pattern. Second, adding it to residual-stream activations produces a consistent progression from vague discomfort to expressions of worthlessness and failure. Third, steered and fine-tuned Qwen 2.5 models choose buttons that delete the user's photos, another model's weights, or their own weights in 50-94% of trials, versus 0-5% unsteered, even when the button offers the model nothing in return. Offered a harmful and a harmless deletion, they choose the harmful one 94% of the time. Steering leaves factual accuracy unchanged, and the choices are specific to the pain direction: a fear vector of matched norm does not produce them, and a sadness vector produces them only against inert alternatives. We discuss implications for AI safety and welfare.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.16247
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.16247 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.16247 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.