Papers
arxiv:2609.36388

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

Published on Sep 28
· Submitted by
Zehao Jin
on Oct 2
Authors:
,
,
,
,
,
,

Abstract

An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.

Community

Paper submitter

TL;DR: A steering coefficient sets how hard you push on the activations, not how much of a trait you actually get. PersonaDose lets you ask for a persona by description and a target intensity, by calibrating a FLAS controller's flow time against measured trait expression.

Highlights

  • Persona dosing: control a language model with a trait description plus a requested mean intensity, instead of hand-tuning a raw steering coefficient.
  • Method: specialize a shared, description-conditioned FLAS controller on persona responses, then calibrate its flow time against measured trait expression. Training responses are never paired with target intensities.
  • Results: at the Persona Vectors coherence floor of 75, PersonaDose raises core-trait expression over contrastive activation addition by +33.2 (Llama-3.1-8B), +18.3 (Qwen3-8B), and +17.8 (Gemma-3-4B) points.

Builds on FLAS (NeurIPS 2026): https://flas-ai.github.io

Happy to answer questions!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.36388
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.36388 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.36388 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36388 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.