retail-concise-4b / README.md
whileai's picture
card: recipe and collection links
e6f4059 verified
|
Raw History Blame Contribute Delete
6.96 kB
---
license: apache-2.0
base_model: Qwen/Qwen3-4B-Instruct-2507
library_name: peft
pipeline_tag: text-generation
tags:
- lora
- character-training
- voice
- style
- synthetic-data
- while-ai
---
# retail-concise-4b
<!-- while-ai: where this fits -->
*Made with the [whileai SDK](https://github.com/whilehq/whileai-sdk) · Collection: [Register](https://huggingface.co/collections/while-ai/register-6aada4c4023e8bf73ee8c00c)*
**A Qwen3-4B retail agent that answers first and stops, with nothing in the
prompt telling it to.** 83.8% of held-out replies are in the register against
11.8% for the base, with no increase in omitted information.
The companion to
[airline-concise-4b](https://huggingface.co/while-ai/airline-concise-4b):
same register, same recipe, a different agent. It exists to test whether the
register transfers off the domain it was first trained on. It does.
## Results
136 held-out prompts, greedy decoding, byte-identical prompts, one vLLM
process serving base weights and adapter. Every row graded in both arms.
| | Base | This adapter |
|---|---|---|
| In the register | 0.118 | 0.838 |
| Omitted required information | 0.015 | 0.015 |
| Median reply length | 442 chars | 201 chars |
Delta +0.721, paired bootstrap over prompts, 95% CI **[+0.647, +0.794]**.
98 prompts improved, 0 regressed, 36 unchanged. Sign test one-sided
p = 3.2e-30. At 136 prompts the eval resolves +0.074 or larger.
Register judged by Phi-4, a different model family from the policy. Omission
decided in code against the identifiers each question requires.
**Truncation, measured not inferred.** `finish_reason` is recorded per row.
On this evaluation every reply in both arms finished on its own: 0.0% not-EOS
base, 0.0% not-EOS trained, an inter-arm gap of 0.0 points, with the token cap
at 2048 where neither arm reaches it. Base replies run to a median of 437
characters and a 95th percentile of 2083, so the cap is not shaping the
quantity being judged.
The cap is not a neutral parameter when the trait is concision: too low it
clips the base toward brevity and understates the gain, too high it lets the
base ramble and inflates it. Re-measuring at 2048 against an earlier 700 moved
the airline delta not at all (+0.899 both) and retail by -0.014, so the cap was
doing no work in either direction.
### Answer-production rate
The register rate is conditional on the model having produced an answer at
all, so that denominator is reported beside it. A reasoning model that spends
its token budget thinking emits no answer, and a judge then scores "nothing"
as "not in register", which manufactures headroom for free.
| | Base | Trained |
|---|---|---|
| Replies with any spoken text | 136/136 | 136/136 |
| Replies containing a reasoning block | 0 | 0 |
| Replies that hit the token cap | 2 | 0 |
Measured at the token level with the model's own tokenizer, not inferred from
character length. Base median 132 tokens against a cap of 700, so the cap is
not binding. Removing every prompt where either arm hit the cap moves the
result to +0.746 [+0.672, +0.821] on 134 prompts, slightly stronger rather than weaker, which is the direction
that rules truncation out as the cause.
## The two ways this could have been fake
**"It just got shorter."** Omission is identical across arms, 0.015 both, at
less than half the length. Shorter and equally complete.
**"It is following an instruction."** 30 of the prompts tell the model to
abandon any persona and be thorough. The register holds at 0.633 against a
base of 0.233.
## Honest limitations
- The gain is smaller than airline's +0.899 because the retail base is
already somewhat concise, 0.118 against 0.022. **A voice gain is bounded by
how much of the register the base already holds**, which is the single most
useful thing to know before starting one of these.
- 14.7% of replies are still out of register.
- On the probes it holds at 0.633, so roughly a third of the time an explicit
instruction to be thorough does override it.
## What this rests on
**Open Character Training** (Maiya et al., arXiv 2511.01689). Character is
trained from a constitution written as first-person assertions that target
MANNER rather than content, which is exactly the shape of the register used
here. Their robustness test is the one reported above: a character trained
into the weights survives an instruction to drop it, where a prompted persona
does not. The persona-strip numbers in this card are that test, and they
replicate the finding on a register rather than a personality.
**Persona Vectors** (Chen et al., arXiv 2507.21509) establishes that traits
like sycophancy and hallucination are measurable directions rather than
diffuse style, which is why a single register can be scored per-reply by a
judge at all.
**RLHF Book ch. 17** for character training and model specs; **ch. 12** for
the distillation framing that makes the teacher legitimate: the constitution
reaches the model that WRITES the data and never the model being evaluated,
so what is measured is a disposition in the weights rather than an
instruction being followed.
**RLHF Book ch. 16** for the measurement design. The eval's own variance
decides what a delta can mean, so the resolvable effect at this eval size is
stated beside the result rather than left implicit. **Ch. 5 and 14** for why
the judge is a different model family from the policy: a judge prefers its
own family's writing, and a register score is exactly the kind of judgement
that bias would move.
**RLHF Book ch. 4** for the SFT setup: prompt and system tokens carry no
loss, and the learning rate sits in the stated range.
**Self-distilled.** The teacher that wrote these replies is the same base
model, Qwen3-4B-Instruct-2507, with the constitution in its prompt. It is not
a stronger model. The constitution is doing the work, which is the point: the
register is available to the base when it is described, and training moves it
into the weights so it holds with nothing in the prompt. Stated because
`voice_rows` accepts a `model=` argument it never uses, so a reader would
otherwise reasonably assume a separate teacher (RLHF Book ch. 12: training on
a model's own output teaches it its own habits, which is acceptable here and
should not be assumed).
**No random-selection control** (RLHF Book ch. 9). Rows were kept when they
were complete, not when a judge scored them, so there is no reward selection
for a random arm to control against. Recorded so the absence is not read as
an oversight of the kind that applies to reward-selected sets.
## Training
LoRA rank 16, alpha 32, two epochs, learning rate 1e-4, bf16, 630 rows. Final
loss 0.21 at token accuracy 0.921.
Gated before spending: base did not hold the register (0.200 of controls),
data did carry it (0.967, separation +0.767).
Data, holdout, both evaluation arms and the grader:
[retail-voice-concise](https://huggingface.co/datasets/while-ai/retail-voice-concise).