--- license: other license_name: qwen-research license_link: https://huggingface.co/Qwen/Qwen2.5-Omni-3B/blob/main/LICENSE base_model: Qwen/Qwen2.5-Omni-3B library_name: peft tags: - speech-emotion-recognition - preference - audio - lora --- # Comparative Reasoning — SFT-CoT (Qwen2.5-Omni-3B LoRA) The **SFT-CoT** model from [Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions](https://arxiv.org/abs/2606.24082) (Interspeech 2026). Given two utterances, it answers which one has higher **arousal**, **valence**, or **dominance** and explains why. This repository contains a LoRA adapter (rank 64, alpha 128, all linear layers) for [Qwen/Qwen2.5-Omni-3B](https://huggingface.co/Qwen/Qwen2.5-Omni-3B). | | | |---|---| | Training | SFT, lr 1e-4, total batch size 32 | | Training data | 10k MSP-Podcast v2.0 preference pairs per attribute (30k total), target: reasoning trace + answer | | MSP-Podcast test (A / V / D / Avg) | 85.5 / 86.5 / 84.6 / 85.5 | | BIIC-Podcast / WHiSER (Avg) | 74.1 / 85.4 | ## Usage Merge the adapter and evaluate with the code repository: ```bash swift export --adapters Lab-MSP/comparative-reasoning-sft-cot --merge_lora true \ --model Qwen/Qwen2.5-Omni-3B --output_dir outputs/experiments/sft-cot_merged bash src/eval.sh msp_test outputs/experiments/sft-cot_merged ``` Prompt format (two audios followed by the question): ``` system: You are a helpful assistant for emotion comparative reasoning. user: