File size: 8,105 Bytes
6d07b7d 7b23a3a 6d07b7d 7b23a3a 8c39c63 7b23a3a 45337a3 7b23a3a 2340699 7b23a3a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 | ---
license: apache-2.0
tags:
- music-editing
- audio
- multimodal
- conversational-ai
- lora
- musicgen
- qwen2.5-omni
language:
- en
pipeline_tag: audio-to-audio
datasets:
- OpenRB-Lab/AURA-Chat-Edit
---
# AURA: Unified Multimodal Framework for Conversational Music Editing
[](https://arxiv.org/abs/2609.14344)
[](https://github.com/OpenRB-Lab/AURA)
[](https://huggingface.co/datasets/OpenRB-Lab/AURA-Chat-Edit)
## Overview
**AURA** is a conversational music-editing agent that listens to a song and a natural-language instruction, replies conversationally, and renders the edited audio. The system consists of three components:
- **Thinker** β A [Qwen2.5-Omni-7B](https://huggingface.co/Qwen/Qwen2.5-Omni-7B) model fine-tuned with LoRA (r=16, alpha=32). It processes audio and text, generates conversational replies, and emits typed edit-token blocks `[EDIT_<KIND>][EDIT_0..7]` (7 kinds: ADD / REMOVE / REPLACE / EXTRACT / REBALANCE / EFFECT / MOOD).
- **Bridge** β A dual-stream fusion MusicGen decoder based on [facebook/musicgen-medium](https://huggingface.co/facebook/musicgen-medium). The 9 hidden states at the edit tokens condition the bridge via **BiFAM** (Bi-FiLM Attention Module: shared-query dual attention + FiLM modulation) and cross-attention K/V with LoRA (r=64, alpha=128). Includes learned projectors (258 MB) mapping from the thinker's hidden dimension to MusicGen's space.
- **Classifier** β An `EditSemanticClassifier` (two-head: edit kind + instrument) used by the programmatic planner for stem routing.
Localized edits are code-anchored outside the requested segment and seam-crossfaded via a stem-hybrid executor (HTDemucs-6s separation).
For the dataset, this open-weight is trained on larger scale data compared to our private one in the paper to ensure the better music quality. The private weight is used for the publications, so we need to ensure we have the fair comparison.
## Model Checkpoints
| File | Description | Size |
|------|-------------|------|
| `config.yaml` | Training configuration (paths, hyperparameters) | 1 KB |
| `thinker/adapter_config.json` | Thinker LoRA configuration | 1 KB |
| `thinker/adapter_model.safetensors` | Thinker LoRA weights (Qwen2.5-Omni-7B, r=16) | 2.0 GB |
| `bridge/projectors.pt` | Learned projectors (d_llm=3584 β d_musicgen=2048) + FiLM MLPs/alphas/gates | 258 MB |
| `bridge/lora/adapter_config.json` | Bridge LoRA configuration | 1 KB |
| `bridge/lora/adapter_model.safetensors` | Bridge LoRA weights (MusicGen encoder_attn k/v, r=64) | 37 MB |
| `classifier/classifier.pt` | EditSemanticClassifier (kind + instrument heads) | 14 MB |
**Total checkpoint size: ~2.3 GB** (adapters only β base models downloaded separately)
### Thinker Details
- **Base model**: [Qwen/Qwen2.5-Omni-7B](https://huggingface.co/Qwen/Qwen2.5-Omni-7B)
- **LoRA config**: r=16, alpha=32, dropout=0.05
- **Target modules**: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`
- **Modules to save**: `embed_tokens`, `lm_head` (for custom edit tokens)
- **Training**: 2-epoch SFT on dialogue data, then joint training with bridge (4000 steps, Ξ»=0.5 convex loss)
### Bridge Details
- **Base model**: [facebook/musicgen-medium](https://huggingface.co/facebook/musicgen-medium) (1.5B params, frozen decoder)
- **Fusion mechanism**: BiFAM β shared-query dual cross-attention over edit-token hidden states + gated FiLM modulation at each decoder layer
- **LoRA config**: r=64, alpha=128, dropout=0.05, targeting `encoder_attn.{k_proj, v_proj}`
- **Projectors**: Linear projections from thinker hidden dim (3584) to MusicGen dim (2048), plus FiLM MLP layers
- **Cross-attention layers**: [0, 2, 4, 6, 8, 10, 12, 14]
- **Training**: 40k steps bridge-only, then 4000 steps joint with thinker
### Classifier Details
- **Architecture**: Two-head classifier (edit kind: 7 classes, instrument: multi-label)
- **Input**: 9 edit-token hidden states (pooled)
- **Used by**: Programmatic planner for stem routing decisions
## Quick Start
### 1. Download base models
The base models are downloaded automatically on first use, or you can pre-download them:
```python
from huggingface_hub import snapshot_download
# Thinker base model (~15 GB)
snapshot_download("Qwen/Qwen2.5-Omni-7B", cache_dir="weights")
# Bridge base model (~3.3 GB)
snapshot_download("facebook/musicgen-medium", cache_dir="weights")
```
### 2. Download AURA checkpoints
```python
from huggingface_hub import snapshot_download
# Download all AURA adapters (~2.3 GB)
repo_dir = snapshot_download("OpenRB-Lab/AURA")
```
Or download individual components:
```python
from huggingface_hub import hf_hub_download
# Thinker LoRA adapter
thinker_config = hf_hub_download("OpenRB-Lab/AURA", "thinker/adapter_config.json")
thinker_weights = hf_hub_download("OpenRB-Lab/AURA", "thinker/adapter_model.safetensors")
# Bridge projectors + LoRA
projectors = hf_hub_download("OpenRB-Lab/AURA", "bridge/projectors.pt")
bridge_config = hf_hub_download("OpenRB-Lab/AURA", "bridge/lora/adapter_config.json")
bridge_weights = hf_hub_download("OpenRB-Lab/AURA", "bridge/lora/adapter_model.safetensors")
# Classifier
classifier = hf_hub_download("OpenRB-Lab/AURA", "classifier/classifier.pt")
```
### 3. Usage
```python
# Point environment to your checkpoint directory
import os
os.environ["AURA_QWEN"] = "path/to/aura1/thinker"
os.environ["AURA_MG"] = "path/to/aura1/bridge"
os.environ["AURA_CLASSIFIER"] = "path/to/aura1/classifier/classifier.pt"
# Load the engine
from serving.engine import AuraEngine
engine = AuraEngine(device="cuda")
result = engine.edit(
audio_path="path/to/song.wav",
instruction="Add a jazzy saxophone melody to the chorus",
guidance=2.0,
seed=1234,
max_seconds=5.0
)
# result["reply"] -> conversational text response
# result["wav"] -> edited audio (float32 numpy, 32 kHz)
# result["sr"] -> 32000
```
### 4. Serving
```bash
# HTTP API
API_GPU=0 API_PORT=9004 bash src/scripts/serve_musicgen_api.sh
# Gradio web UI
WORKER_URL=http://127.0.0.1:9004 \
SFT_ADAPTER=path/to/aura1/thinker \
WEBAPP_PORT=7862 CUDA_VISIBLE_DEVICES=1 \
python src/edit_agent/webapp.py
```
## Training Configuration
Training uses a 3-stage pipeline:
1. **Stage 1 β Thinker SFT**: LoRA fine-tuning on conversational music-edit dialogues (2 epochs, lr=1e-4)
2. **Stage 2 β Bridge**: Fusion adapter training on cached thinker hidden states (40k steps, lr=5e-5)
3. **Stage 3 β Joint**: End-to-end training with live thinker + bridge (4k steps, loss = λ·CE_musicgen + (1βΞ»)Β·CE_LM)
See `config.yaml` for the full training configuration.
## Results
Production checkpoints (`joint_fusion_r64/final`):
| Benchmark | FAD β | CLAP β | SSIM β |
|-----------|-------|--------|--------|
| IMPG Add | 1.49 | β | β |
| IMPG Remove | 1.36 | β | β |
| IMPG Extract | 6.13 | β | β |
| Mixed 60-clip | 2.15 | 0.661 (MuLan cos) | β |
### Fusion Ablation (IMPG Benchmark)
| Method | FAD β | CLAP β | SSIM β |
|--------|-------|--------|--------|
| **BiFAM (Ours)** | **0.41** | 0.218 | **0.776** |
| Cross-Attention Only | 2.48 | **0.263** | 0.099 |
| Concatenation Only | 13.64 | 0.181 | 0.115 |
## Citation
```bibtex
@misc{trinh2026auraunifiedmultimodalframework,
title={AURA: Unified Multimodal Framework for Conversational Music Editing},
author={Quoc-Huy Trinh and Minh-Van Nguyen and Debesh Jha},
year={2026},
eprint={2609.14344},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2609.14344},
}
```
## License
This project is licensed under the Apache License 2.0. See [LICENSE](LICENSE) for details.
**Note**: The base models have their own licenses:
- Qwen2.5-Omni-7B: [Apache 2.0](https://huggingface.co/Qwen/Qwen2.5-Omni-7B)
- MusicGen-medium: [CC-BY-NC 4.0](https://huggingface.co/facebook/musicgen-medium)
|