File size: 2,267 Bytes
26feeec
 
 
 
 
 
 
 
a0ff11d
0dfc3ae
 
a0ff11d
 
 
0dfc3ae
 
 
 
 
a0ff11d
0dfc3ae
 
a0ff11d
0dfc3ae
a0ff11d
0dfc3ae
 
 
 
 
 
a0ff11d
0dfc3ae
 
8c78290
0dfc3ae
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a0ff11d
0dfc3ae
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
---
tags:
  - audio
  - music-generation
  - midi-to-audio
  - audio-super-resolution
---

# CPR: COMBINING GLOBAL COMPOSING, LOCAL PERFORMING AND FULL-SEQUENCE REFINING IN PIANO RENDERING WITH CONTINUOUS AUTOREGRESSIVE MODELLING

1. **Composer–Performer (CP)** takes prompt audio, prompt MIDI, and target MIDI,
   and renders **24 kHz mono audio**. The Composer is an autoregressive Qwen3
   Transformer; the Performer renders local Mel spectrograms with flow matching.
   A Vocos vocoder converts the generated Mel spectrograms to audio.
2. **Refiner (R)** takes a **24 kHz audio file** and writes **48 kHz mono audio**
   using the LavaSR-based bandwidth-extension model and low-frequency fusion.

## Installation

Create a Conda environment with Python 3.10:

```bash
conda create -n cpr python=3.10
conda activate cpr
pip install -r requirements.txt
```

On the first CP run, CLAP automatically downloads and caches its RoBERTa initialization resources.

## Checkpoints

Download the inference assets from [Huggingface](https://huggingface.co/FEAfeatherTHER/CPR) into `checkpoints/`.

```bash
hf download FEAfeatherTHER/CPR \
  composer_performer.safetensors \
  composer_performer_config.json \
  vocos.safetensors \
  refiner.bin \
  --local-dir checkpoints
```

| Component | File under `checkpoints/` |
|---|---|
| Composer–Performer | `composer_performer.safetensors` |
| Model architecture and Qwen configuration | `composer_performer_config.json` |
| 24 kHz Vocos vocoder | `vocos.safetensors` |
| Refiner | `refiner.bin` |


### Download CLAP

CLAP is required for Composer–Performer inference. Download the official music checkpoint
[music_audioset_epoch_15_esc_90.14.pt](https://huggingface.co/lukewys/laion_clap/blob/main/music_audioset_epoch_15_esc_90.14.pt) and save it as `checkpoints/clap.pt`:

```bash
curl --fail --location \
  --output checkpoints/clap.pt \
  https://huggingface.co/lukewys/laion_clap/resolve/main/music_audioset_epoch_15_esc_90.14.pt
```

## Inference

### 1. Composer–Performer → 24 kHz

```bash
bash scripts/infer_cp.sh
```

### 2. Refiner → 48 kHz

use any existing 24 kHz WAV or the output of Composer-Performer:

```bash
python infer_refiner.py --input /path/to/audio_24k.wav --output outputs/refined.wav
```