File size: 8,105 Bytes
6d07b7d
 
7b23a3a
 
 
 
 
 
 
 
 
 
 
 
 
6d07b7d
7b23a3a
 
 
8c39c63
7b23a3a
 
 
 
 
 
 
 
 
 
 
 
 
 
45337a3
 
7b23a3a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2340699
 
 
 
 
 
 
 
7b23a3a
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
---
license: apache-2.0
tags:
- music-editing
- audio
- multimodal
- conversational-ai
- lora
- musicgen
- qwen2.5-omni
language:
- en
pipeline_tag: audio-to-audio
datasets:
- OpenRB-Lab/AURA-Chat-Edit
---

# AURA: Unified Multimodal Framework for Conversational Music Editing

[![arXiv](https://img.shields.io/badge/arXiv-2609.14344-b31b1b.svg)](https://arxiv.org/abs/2609.14344)
[![GitHub](https://img.shields.io/badge/GitHub-AURA-blue.svg)](https://github.com/OpenRB-Lab/AURA)
[![Dataset](https://img.shields.io/badge/πŸ€—_Dataset-AURA--Chat--Edit-yellow.svg)](https://huggingface.co/datasets/OpenRB-Lab/AURA-Chat-Edit)


## Overview

**AURA** is a conversational music-editing agent that listens to a song and a natural-language instruction, replies conversationally, and renders the edited audio. The system consists of three components:

- **Thinker** β€” A [Qwen2.5-Omni-7B](https://huggingface.co/Qwen/Qwen2.5-Omni-7B) model fine-tuned with LoRA (r=16, alpha=32). It processes audio and text, generates conversational replies, and emits typed edit-token blocks `[EDIT_<KIND>][EDIT_0..7]` (7 kinds: ADD / REMOVE / REPLACE / EXTRACT / REBALANCE / EFFECT / MOOD).
- **Bridge** β€” A dual-stream fusion MusicGen decoder based on [facebook/musicgen-medium](https://huggingface.co/facebook/musicgen-medium). The 9 hidden states at the edit tokens condition the bridge via **BiFAM** (Bi-FiLM Attention Module: shared-query dual attention + FiLM modulation) and cross-attention K/V with LoRA (r=64, alpha=128). Includes learned projectors (258 MB) mapping from the thinker's hidden dimension to MusicGen's space.
- **Classifier** β€” An `EditSemanticClassifier` (two-head: edit kind + instrument) used by the programmatic planner for stem routing.

Localized edits are code-anchored outside the requested segment and seam-crossfaded via a stem-hybrid executor (HTDemucs-6s separation).

For the dataset, this open-weight is trained on larger scale data compared to our private one in the paper to ensure the better music quality. The private weight is used for the publications, so we need to ensure we have the fair comparison.

## Model Checkpoints

| File | Description | Size |
|------|-------------|------|
| `config.yaml` | Training configuration (paths, hyperparameters) | 1 KB |
| `thinker/adapter_config.json` | Thinker LoRA configuration | 1 KB |
| `thinker/adapter_model.safetensors` | Thinker LoRA weights (Qwen2.5-Omni-7B, r=16) | 2.0 GB |
| `bridge/projectors.pt` | Learned projectors (d_llm=3584 β†’ d_musicgen=2048) + FiLM MLPs/alphas/gates | 258 MB |
| `bridge/lora/adapter_config.json` | Bridge LoRA configuration | 1 KB |
| `bridge/lora/adapter_model.safetensors` | Bridge LoRA weights (MusicGen encoder_attn k/v, r=64) | 37 MB |
| `classifier/classifier.pt` | EditSemanticClassifier (kind + instrument heads) | 14 MB |

**Total checkpoint size: ~2.3 GB** (adapters only β€” base models downloaded separately)

### Thinker Details

- **Base model**: [Qwen/Qwen2.5-Omni-7B](https://huggingface.co/Qwen/Qwen2.5-Omni-7B)
- **LoRA config**: r=16, alpha=32, dropout=0.05
- **Target modules**: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`
- **Modules to save**: `embed_tokens`, `lm_head` (for custom edit tokens)
- **Training**: 2-epoch SFT on dialogue data, then joint training with bridge (4000 steps, Ξ»=0.5 convex loss)

### Bridge Details

- **Base model**: [facebook/musicgen-medium](https://huggingface.co/facebook/musicgen-medium) (1.5B params, frozen decoder)
- **Fusion mechanism**: BiFAM β€” shared-query dual cross-attention over edit-token hidden states + gated FiLM modulation at each decoder layer
- **LoRA config**: r=64, alpha=128, dropout=0.05, targeting `encoder_attn.{k_proj, v_proj}`
- **Projectors**: Linear projections from thinker hidden dim (3584) to MusicGen dim (2048), plus FiLM MLP layers
- **Cross-attention layers**: [0, 2, 4, 6, 8, 10, 12, 14]
- **Training**: 40k steps bridge-only, then 4000 steps joint with thinker

### Classifier Details

- **Architecture**: Two-head classifier (edit kind: 7 classes, instrument: multi-label)
- **Input**: 9 edit-token hidden states (pooled)
- **Used by**: Programmatic planner for stem routing decisions

## Quick Start

### 1. Download base models

The base models are downloaded automatically on first use, or you can pre-download them:

```python
from huggingface_hub import snapshot_download

# Thinker base model (~15 GB)
snapshot_download("Qwen/Qwen2.5-Omni-7B", cache_dir="weights")

# Bridge base model (~3.3 GB)
snapshot_download("facebook/musicgen-medium", cache_dir="weights")
```

### 2. Download AURA checkpoints

```python
from huggingface_hub import snapshot_download

# Download all AURA adapters (~2.3 GB)
repo_dir = snapshot_download("OpenRB-Lab/AURA")
```

Or download individual components:

```python
from huggingface_hub import hf_hub_download

# Thinker LoRA adapter
thinker_config = hf_hub_download("OpenRB-Lab/AURA", "thinker/adapter_config.json")
thinker_weights = hf_hub_download("OpenRB-Lab/AURA", "thinker/adapter_model.safetensors")

# Bridge projectors + LoRA
projectors = hf_hub_download("OpenRB-Lab/AURA", "bridge/projectors.pt")
bridge_config = hf_hub_download("OpenRB-Lab/AURA", "bridge/lora/adapter_config.json")
bridge_weights = hf_hub_download("OpenRB-Lab/AURA", "bridge/lora/adapter_model.safetensors")

# Classifier
classifier = hf_hub_download("OpenRB-Lab/AURA", "classifier/classifier.pt")
```

### 3. Usage

```python
# Point environment to your checkpoint directory
import os
os.environ["AURA_QWEN"] = "path/to/aura1/thinker"
os.environ["AURA_MG"] = "path/to/aura1/bridge"
os.environ["AURA_CLASSIFIER"] = "path/to/aura1/classifier/classifier.pt"

# Load the engine
from serving.engine import AuraEngine

engine = AuraEngine(device="cuda")
result = engine.edit(
    audio_path="path/to/song.wav",
    instruction="Add a jazzy saxophone melody to the chorus",
    guidance=2.0,
    seed=1234,
    max_seconds=5.0
)

# result["reply"]  -> conversational text response
# result["wav"]    -> edited audio (float32 numpy, 32 kHz)
# result["sr"]     -> 32000
```

### 4. Serving

```bash
# HTTP API
API_GPU=0 API_PORT=9004 bash src/scripts/serve_musicgen_api.sh

# Gradio web UI
WORKER_URL=http://127.0.0.1:9004 \
SFT_ADAPTER=path/to/aura1/thinker \
WEBAPP_PORT=7862 CUDA_VISIBLE_DEVICES=1 \
python src/edit_agent/webapp.py
```

## Training Configuration

Training uses a 3-stage pipeline:

1. **Stage 1 β€” Thinker SFT**: LoRA fine-tuning on conversational music-edit dialogues (2 epochs, lr=1e-4)
2. **Stage 2 β€” Bridge**: Fusion adapter training on cached thinker hidden states (40k steps, lr=5e-5)
3. **Stage 3 β€” Joint**: End-to-end training with live thinker + bridge (4k steps, loss = λ·CE_musicgen + (1βˆ’Ξ»)Β·CE_LM)

See `config.yaml` for the full training configuration.

## Results

Production checkpoints (`joint_fusion_r64/final`):

| Benchmark | FAD ↓ | CLAP ↑ | SSIM ↑ |
|-----------|-------|--------|--------|
| IMPG Add | 1.49 | β€” | β€” |
| IMPG Remove | 1.36 | β€” | β€” |
| IMPG Extract | 6.13 | β€” | β€” |
| Mixed 60-clip | 2.15 | 0.661 (MuLan cos) | β€” |

### Fusion Ablation (IMPG Benchmark)

| Method | FAD ↓ | CLAP ↑ | SSIM ↑ |
|--------|-------|--------|--------|
| **BiFAM (Ours)** | **0.41** | 0.218 | **0.776** |
| Cross-Attention Only | 2.48 | **0.263** | 0.099 |
| Concatenation Only | 13.64 | 0.181 | 0.115 |

## Citation

```bibtex
@misc{trinh2026auraunifiedmultimodalframework,
      title={AURA: Unified Multimodal Framework for Conversational Music Editing}, 
      author={Quoc-Huy Trinh and Minh-Van Nguyen and Debesh Jha},
      year={2026},
      eprint={2609.14344},
      archivePrefix={arXiv},
      primaryClass={cs.SD},
      url={https://arxiv.org/abs/2609.14344}, 
}
```

## License

This project is licensed under the Apache License 2.0. See [LICENSE](LICENSE) for details.

**Note**: The base models have their own licenses:
- Qwen2.5-Omni-7B: [Apache 2.0](https://huggingface.co/Qwen/Qwen2.5-Omni-7B)
- MusicGen-medium: [CC-BY-NC 4.0](https://huggingface.co/facebook/musicgen-medium)