File size: 2,222 Bytes
2577656
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
---
language:
- en
- zh
license: apache-2.0
base_model:
  - m-a-p/YuE2-3B
  - m-a-p/YuE2-Vae
tags:
- music-generation
- yue2
- modular-diffusers
pipeline_tag: text-to-audio
library_name: diffusers
---
# YuE2 Modular Diffusers

<audio controls src="https://huggingface.co/datasets/OzzyGT/diffusers-examples/resolve/main/yue2/city_lights.mp3"></audio>

*A 69-second song generated with the code below, seed `831001`.*

Modular Diffusers blocks for [m-a-p/YuE2-3B](https://huggingface.co/m-a-p/YuE2-3B): full songs from a style description and lyrics, planned as an ABC score first, as 48 kHz stereo audio. The model code and weights are in [OzzyGT/YuE2-3B-Diffusers](https://huggingface.co/OzzyGT/YuE2-3B-Diffusers).

Note: This model requires `tiktoken` and the example uses `soundfile`, install them with `pip install tiktoken soundfile`.

## Licensing

The code is Apache-2.0, adapted from [multimodal-art-projection/YuE](https://github.com/multimodal-art-projection/YuE). The weights are CC BY-NC 4.0 (non-commercial); see the [model repository](https://huggingface.co/OzzyGT/YuE2-3B-Diffusers).

## Sample song

The song above was generated with the following code:

```python
import soundfile as sf
import torch
from diffusers import ModularPipelineBlocks


blocks = ModularPipelineBlocks.from_pretrained(  # load the blocks first to avoid warnings
    "OzzyGT/YuE2-Modular",
    trust_remote_code=True,
    components_repo="OzzyGT/YuE2-3B-Diffusers",
    trust_components_code=True,
)
pipe = blocks.init_pipeline()
pipe.load_components(dtype={"transformer": torch.bfloat16, "vae": torch.float32})
pipe.to("cuda")

style = "English, warm piano pop, expressive female voice, acoustic piano, rounded bass and light drums, lyrical memorable melody, unhurried phrasing, 88 BPM"
lyrics = """[Verse]
Neon fades along the lane
Footsteps keep the time of rain
Fold the night and leave it here
Morning has a sky to clear

[Chorus]
Let the day come into view
Every road begins with you
Hold a little room for light
We will sing beyond the night"""

state = pipe(style=style, lyrics=lyrics, seed=831001, use_cuda_graph=True)
sf.write("song.flac", state.get("audios")[0].T.numpy(), state.get("sample_rate"), subtype="PCM_24")
```