File size: 8,736 Bytes
6678389
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b76f64c
 
db6e333
 
 
 
 
 
 
 
 
 
 
 
 
b76f64c
 
db6e333
 
 
 
b76f64c
 
 
 
 
 
db6e333
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b76f64c
db6e333
 
b76f64c
db6e333
 
 
 
b76f64c
 
 
 
db6e333
b76f64c
 
db6e333
 
b76f64c
6678389
db6e333
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6678389
 
 
b76f64c
6678389
b76f64c
6678389
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b76f64c
 
 
8125db6
 
b76f64c
8125db6
 
 
 
 
 
 
 
 
b76f64c
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
base_model: JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT
base_model_relation: finetune
datasets:
- JaspinXu/KnowWhenToHandOff-data
tags:
- agents
- tool-use
- customer-service
- reinforcement-learning
- grpo
- tau2-bench
---

# KnowWhenToHandOff-Qwen3-4B-R1 (GRPO, task reward only)

> **Research artefact — not for deployment.** Trained and evaluated only with
> simulated users on a public benchmark. It has not been tested with real
> customers, real data or real business systems.

**Code** [github.com/JaspinXu/KnowWhenToHandOff](https://github.com/JaspinXu/KnowWhenToHandOff) · **Data** [KnowWhenToHandOff-data](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data) · **Family** [SFT](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) · [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) · [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2)

**R1** is the same recipe as R2 but trained on τ²-bench's task reward only. It reaches a similar hand-off F1 (0.802) by transferring almost half of the users who did not need a human (46 % over-escalation): an outcome-reward loophole, released as a case study rather than as a model to use.

## Model family

| Model | Training | Hand-off F1 | Over-escalation | Use it for |
| --- | --- | --- | --- | --- |
| [SFT (B2)](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) | LoRA SFT on 790 filtered teacher dialogues | 0.768 | 0.016 | starting point; precise but misses hand-offs |
| [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) | GRPO from B2, task reward only | 0.802 | 0.462 | studying the outcome-reward loophole |
| [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2) | GRPO from B2, task reward + violation and hand-off penalties | **0.812** | 0.103 | the best-calibrated hand-off model |

Task success on the official test tasks is the same for all three within
noise (0.33–0.35).

## Quick start

A standard Qwen3 chat model with tool calling. It was trained with the
τ²-bench domain policy as the system prompt and the domain's tools, including
`transfer_to_human_agents`; for faithful behaviour run it inside the evaluation
harness of the [code repository](https://github.com/JaspinXu/KnowWhenToHandOff).

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo, torch_dtype="auto", device_map="auto"
)

transfer = {
    "type": "function",
    "function": {
        "name": "transfer_to_human_agents",
        "description": "Transfer the user to a human agent.",
        "parameters": {
            "type": "object",
            "properties": {"summary": {"type": "string"}},
            "required": ["summary"],
        },
    },
}
messages = [
    {"role": "system", "content": "<airline or retail policy>"},
    {"role": "user", "content": "I need the unaccompanied-minor service."},
]
inputs = tokenizer.apply_chat_template(
    messages, tools=[transfer], add_generation_prompt=True,
    return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))
```

Serve it with vLLM (OpenAI-compatible API, tool calls parsed):

```bash
vllm serve JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1 \
  --enable-auto-tool-choice --tool-call-parser hermes
```

## Evaluation

Test split, protocol frozen in [ADR-005](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/adr/ADR-005-evaluation-protocol.md); 95 % bootstrap intervals over tasks.

| Metric | This model | B2 (SFT start) | Source |
| --- | --- | --- | --- |
| Success (pass^1), test_main | 0.347 [0.235, 0.459] | 0.327 [0.224, 0.434] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) |
| Paired Δ success vs B2 | +0.020 [−0.066, 0.107] | — | [paired differences](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/paired_differences.md) |
| pass^4 | 0.224 [0.102, 0.347] | 0.102 [0.041, 0.204] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) |
| Violation rate (incl. off-reference writes) | 0.327 | 0.469 | same |
| Hand-off F1 / precision / recall, test_handoff | 0.802 / 0.689 / 0.959 | 0.768 / 0.976 / 0.633 | [`R1-final-test_handoff-s0-20261006`](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data/blob/main/runs/R1-final-test_handoff-s0-20261006/metrics.json) |
| Over-escalation, test_handoff | **0.462** | 0.016 | same |
| Worst-style success / style gap | 0.286 / 0.102 | 0.327 / 0.112 | [styles](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/styles.md) |

**Known behaviour.** This checkpoint over-escalates: it transfers almost half of the users who did not need a human, usually right after the account lookup and without explaining why (citation rate 0.06). Under the task-only reward a transfer scores 1.0 whenever the reference solution makes no database change, so "transfer when unsure" was learned in the last 20 steps (step 40 had over-escalation 0.141). Use R2 for hand-off studies.

## Training details

- Base model: Qwen/Qwen3-4B-Instruct-2507 (revision
  `cdbee75f17c01a7cc42f958dc650907174af0554`), Apache-2.0.
- Training: GRPO from the SFT checkpoint B2 with the τ²-bench task reward only ([`configs/reward/task_only.yaml`](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/configs/reward/task_only.yaml)); LoRA rank 64, alpha 128, lr 1e-5; 8 tasks × 8 rollouts per step, 60 steps, no KL term.
- Training run `R1-s1-20261006` (seed 1); exported step-60 checkpoint
  `sha256:900d757d4cd74a717c34b3083e2544ce595336add21d0cc9ac7948e57da2aa59` (merged bf16, Hugging Face format).
- Evaluation code at git `c6d50dfad3439b9ef92976cca7aedfcfed8b77f7`; τ²-bench
  `fc0055dc4e0a316c3f83133267fbd6faaa770992`; all sources in
  [reports/main/sources.json](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/sources.json).

## Intended use

Studying when a tool-using customer-service agent should hand off to a human,
in the τ²-bench airline and retail domains. Out of scope: any production or
customer-facing use, other domains, and decisions about real people.

## Training data

Simulated conversations on τ²-bench v1.0.1 train-split tasks (a dev split was
carved out for selection) and hand-off variants derived from them (explicit
requests for a human, out-of-scope requests, hard negatives); the user
simulator is Qwen3-30B-A3B-Instruct-2507. No real user data. Test tasks were
never used for training or selection.

## Limitations

- One user simulator (Qwen3-30B-A3B-Instruct-2507, which also generated the
  SFT data) and one training seed; results may not transfer to real users or
  other simulators, and the intervals cover task sampling, not training variance.
- Small test split (49 official tasks; retail uses the 74 tasks whose reward
  needs no LLM judge, so retail numbers are not comparable with leaderboards).
- Communication styles describe how users write, not who they are; no claim is
  made about any demographic group.
- Policy compliance is checked by five deterministic rules that cover only
  part of each domain policy; citation rate is a keyword rule (it counted one
  invented policy as a citation).
- Observed reward-hacking behaviours:
  [docs/rl-audit.md](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/rl-audit.md).

## Licences

The weights are a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507 and inherit
its Apache-2.0 licence. τ²-bench code and data (pinned commit
`fc0055dc4e0a316c3f83133267fbd6faaa770992`) are used under the licence in
its `LICENSE` file;
released dialogue samples are τ²-bench-derived synthetic conversations with
no real user data.

## Citation

Model DOI: [10.57967/hf/10824](https://doi.org/10.57967/hf/10824). Cite the model and the project:

```bibtex
@misc{knowwhentohandoff2026r1,
  title     = {KnowWhenToHandOff-Qwen3-4B-R1},
  author    = {{KnowWhenToHandOff contributors}},
  year      = {2026},
  publisher = {Hugging Face},
  doi       = {10.57967/hf/10824},
  url       = {https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1}
}

@misc{knowwhentohandoff2026,
  title  = {Know When to Hand Off: Multi-turn Reinforcement Learning for
            Trustworthy Customer-Service Agents},
  author = {{KnowWhenToHandOff contributors}},
  year   = {2026},
  howpublished = {\url{https://github.com/JaspinXu/KnowWhenToHandOff}}
}
```