File size: 2,000 Bytes
8c78df4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
---
license: other
license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen2.5-Omni-3B/blob/main/LICENSE
base_model: Qwen/Qwen2.5-Omni-3B
library_name: peft
tags:
- speech-emotion-recognition
- preference
- audio
- lora
---

# Comparative Reasoning — SFT (Qwen2.5-Omni-3B LoRA)

The **SFT** model from [Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions](https://arxiv.org/abs/2606.24082) (Interspeech 2026).
Given two utterances, it answers which one has higher **arousal**, **valence**, or **dominance**.

This repository contains a LoRA adapter (rank 64, alpha 128, all linear layers) for [Qwen/Qwen2.5-Omni-3B](https://huggingface.co/Qwen/Qwen2.5-Omni-3B).

| | |
|---|---|
| Training | SFT, lr 1e-4, total batch size 32 |
| Training data | 10k MSP-Podcast v2.0 preference pairs per attribute (30k total), target: answer only |
| MSP-Podcast test (A / V / D / Avg) | 88.1 / 87.8 / 86.7 / 87.5 |
| BIIC-Podcast / WHiSER (Avg) | 76.0 / 89.8 |

## Usage

Merge the adapter and evaluate with the code repository:

```bash
swift export --adapters Lab-MSP/comparative-reasoning-sft --merge_lora true \
  --model Qwen/Qwen2.5-Omni-3B --output_dir outputs/experiments/sft_merged
bash src/eval.sh msp_test outputs/experiments/sft_merged
```

Prompt format (two audios followed by the question):

```
system: You are a helpful assistant for emotion comparative reasoning.
user:   <audio><audio>You will hear two audio clips. Clip 1 is the first audio clip. Clip 2 is the second audio clip. <attribute definition> Which clip has higher <attribute>?
```

The model answers `<answer> Clip 1 </answer>` or `... Clip 2 ...`.

## Citation

```bibtex
@inproceedings{naini2026comparative,
  title     = {Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions},
  author    = {Naini, Abinay Reddy and Kim, Jaeyeon and Yang, Chao-Han Huck and Watanabe, Shinji and Busso, Carlos},
  booktitle = {Interspeech},
  year      = {2026}
}
```