Instructions to use Lab-MSP/comparative-reasoning-sft-cot with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Lab-MSP/comparative-reasoning-sft-cot with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Omni-3B") model = PeftModel.from_pretrained(base_model, "Lab-MSP/comparative-reasoning-sft-cot") - Notebooks
- Google Colab
- Kaggle
File size: 2,071 Bytes
06d878c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 | ---
license: other
license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen2.5-Omni-3B/blob/main/LICENSE
base_model: Qwen/Qwen2.5-Omni-3B
library_name: peft
tags:
- speech-emotion-recognition
- preference
- audio
- lora
---
# Comparative Reasoning — SFT-CoT (Qwen2.5-Omni-3B LoRA)
The **SFT-CoT** model from [Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions](https://arxiv.org/abs/2606.24082) (Interspeech 2026).
Given two utterances, it answers which one has higher **arousal**, **valence**, or **dominance** and explains why.
This repository contains a LoRA adapter (rank 64, alpha 128, all linear layers) for [Qwen/Qwen2.5-Omni-3B](https://huggingface.co/Qwen/Qwen2.5-Omni-3B).
| | |
|---|---|
| Training | SFT, lr 1e-4, total batch size 32 |
| Training data | 10k MSP-Podcast v2.0 preference pairs per attribute (30k total), target: reasoning trace + answer |
| MSP-Podcast test (A / V / D / Avg) | 85.5 / 86.5 / 84.6 / 85.5 |
| BIIC-Podcast / WHiSER (Avg) | 74.1 / 85.4 |
## Usage
Merge the adapter and evaluate with the code repository:
```bash
swift export --adapters Lab-MSP/comparative-reasoning-sft-cot --merge_lora true \
--model Qwen/Qwen2.5-Omni-3B --output_dir outputs/experiments/sft-cot_merged
bash src/eval.sh msp_test outputs/experiments/sft-cot_merged
```
Prompt format (two audios followed by the question):
```
system: You are a helpful assistant for emotion comparative reasoning.
user: <audio><audio>You will hear two audio clips. Clip 1 is the first audio clip. Clip 2 is the second audio clip. <attribute definition> Which clip has higher <attribute>?
```
The model answers `<think> ... </think> <answer> Clip 1 </answer>` or `... Clip 2 ...`.
## Citation
```bibtex
@inproceedings{naini2026comparative,
title = {Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions},
author = {Naini, Abinay Reddy and Kim, Jaeyeon and Yang, Chao-Han Huck and Watanabe, Shinji and Busso, Carlos},
booktitle = {Interspeech},
year = {2026}
}
```
|