File size: 2,267 Bytes
26feeec a0ff11d 0dfc3ae a0ff11d 0dfc3ae a0ff11d 0dfc3ae a0ff11d 0dfc3ae a0ff11d 0dfc3ae a0ff11d 0dfc3ae 8c78290 0dfc3ae a0ff11d 0dfc3ae | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 | ---
tags:
- audio
- music-generation
- midi-to-audio
- audio-super-resolution
---
# CPR: COMBINING GLOBAL COMPOSING, LOCAL PERFORMING AND FULL-SEQUENCE REFINING IN PIANO RENDERING WITH CONTINUOUS AUTOREGRESSIVE MODELLING
1. **Composer–Performer (CP)** takes prompt audio, prompt MIDI, and target MIDI,
and renders **24 kHz mono audio**. The Composer is an autoregressive Qwen3
Transformer; the Performer renders local Mel spectrograms with flow matching.
A Vocos vocoder converts the generated Mel spectrograms to audio.
2. **Refiner (R)** takes a **24 kHz audio file** and writes **48 kHz mono audio**
using the LavaSR-based bandwidth-extension model and low-frequency fusion.
## Installation
Create a Conda environment with Python 3.10:
```bash
conda create -n cpr python=3.10
conda activate cpr
pip install -r requirements.txt
```
On the first CP run, CLAP automatically downloads and caches its RoBERTa initialization resources.
## Checkpoints
Download the inference assets from [Huggingface](https://huggingface.co/FEAfeatherTHER/CPR) into `checkpoints/`.
```bash
hf download FEAfeatherTHER/CPR \
composer_performer.safetensors \
composer_performer_config.json \
vocos.safetensors \
refiner.bin \
--local-dir checkpoints
```
| Component | File under `checkpoints/` |
|---|---|
| Composer–Performer | `composer_performer.safetensors` |
| Model architecture and Qwen configuration | `composer_performer_config.json` |
| 24 kHz Vocos vocoder | `vocos.safetensors` |
| Refiner | `refiner.bin` |
### Download CLAP
CLAP is required for Composer–Performer inference. Download the official music checkpoint
[music_audioset_epoch_15_esc_90.14.pt](https://huggingface.co/lukewys/laion_clap/blob/main/music_audioset_epoch_15_esc_90.14.pt) and save it as `checkpoints/clap.pt`:
```bash
curl --fail --location \
--output checkpoints/clap.pt \
https://huggingface.co/lukewys/laion_clap/resolve/main/music_audioset_epoch_15_esc_90.14.pt
```
## Inference
### 1. Composer–Performer → 24 kHz
```bash
bash scripts/infer_cp.sh
```
### 2. Refiner → 48 kHz
use any existing 24 kHz WAV or the output of Composer-Performer:
```bash
python infer_refiner.py --input /path/to/audio_24k.wav --output outputs/refined.wav
```
|