Ethio-Music-YuE2
Ethio-Music-YuE2 is a LoRA adapter for YuE2-3B that steers it towards Ethiopian music:
- the six kiñit modes
- Ethiopian 6/8 and 4/4 grooves and Ethio-jazz
- Amharic lyrics
The repository also contains the code that turns tags and Amharic (Fidel) lyrics into a YuE2 request with a kiñit-constrained score.
Use it with the kiñit score, which is the default. With the score, the requested kiñit came out right in 72 of 72 test renders. From a caption alone, the adapter does much better than stock YuE2 but is still right only about half the time (see Limitations).
| Base model | m-a-p/YuE2-3B @ 8f24312, decoder m-a-p/YuE2-Vae @ 152733a |
| Adapter | LoRA, rank 64, on all 28 layers of YuE2's autoregressive (AR) branch. 69.7 M parameters, fp32, 279 MB. The NAR branch and the decoder are unchanged. |
| Training data | 756 songs (30.5 h) that YuE2-3B rendered from kiñit scores and that passed automatic quality gates, plus a regularizer of YuE2-generated songs. No real recordings. |
| Licence | CC BY-NC 4.0, non-commercial, inherited from YuE2-3B |
Kiñits and grooves
| Kiñit | Scale degrees |
|---|---|
| tizita major | 1 2 3 5 6 |
| tizita minor | 1 2 ♭3 5 ♭6 |
| bati major | 1 3 ♯4 5 7 |
| bati minor | 1 ♭3 4 5 ♭7 |
| ambassel | 1 ♭2 4 5 ♭6 |
| anchihoye | 1 ♭2 4 ♭5 ♭7 |
Sources disagree on bati major (♮4 or ♯4) and anchihoye (6 or ♭7). This release uses ♯4 and ♭7.
Tizita major and bati minor are modes of one pentatonic collection, and the other four kiñits are modes of a second one. Within a collection, only the tonic tells the kiñits apart, so the scores anchor it.
Grooves:
- chikchika (6/8)
- eskista (6/8)
- slow 6/8
- shegaye (4/4)
- Ethio-jazz (12/8). The score writes it in 6/8, because YuE2 plays a 12/8 score as straight 4/4.
How to use
You need the code in code/. Tested on Linux (WSL2) with an RTX 5090:
hf download Behailut/Ethio-Music-YuE2 --local-dir Ethio-Music-YuE2
cd Ethio-Music-YuE2/code
bash envs/setup_envs.sh ethioyue # ~/envs/ethioyue: Python 3.12, torch 2.10 (CUDA 12.8), yue2-infer; needs uv
# the exact base-model commits the adapter was trained on
~/envs/ethioyue/bin/python scripts/fetch_models.py --only yue2 yue2_vae \
--revision yue2=8f24312f187763b9854adeec87e437e2b03bbff1 \
--revision yue2_vae=152733a19ad43aa67e367f9b5503ef8075bb5126
~/envs/ethioyue/bin/python scripts/30_generate.py \
--lora ../ethio-music-yue2-ar-lora-r64.safetensors \
--tags "<|tizita_minor, krar, 6/8 chikchika|>" \
--lyrics my_song.txt --candidates 4 --out outputs/tizita_demo
my_song.txt contains Amharic lyrics in Fidel. Section lines are optional:
[Verse]
ትዝታ ትዝታ ልቤን ያስታውሰኛል
ፍቅርሽ ልቤን ወሰደው
Input:
- The lyrics are romanized automatically.
- Tags can name a kiñit, groove, genre, instrument, mood, vocal or production style (see
code/configs/taxonomy.yaml). - Leave out
--lyricsfor an instrumental.
Output: best.flac is the best of the candidates.
Speed and memory: a 40-second song took about 15 s per candidate and used about 8 GiB of VRAM.
Settings:
- Keep the defaults:
--composer kinit(our score) and captions that name the scale degrees. --composer none(caption only) and--composer planner(YuE2 writes the score and we snap it into the kiñit) are weaker; see the results below.- Don't use
--no-caption-degreeswith this adapter.
Without our code: the adapter keys are layers.{i}.{self_attn|mlp}.{q,k,v,o,gate,up,down}_proj.lora_{A|B}. They index YuE2's AR decoder layers (model.model.layers[i]).
- Merge with
W += lora_B @ lora_Aand scale 1.0, because alpha/rank is already folded intolora_B. Seecode/ethioyue/train/lora.py(merge_lora) andadapter_config.json. - Name the scale degrees in your caption, e.g.
Ethiopian tizita minor pentatonic mode (scale degrees 1 2 b3 5 b6). - Without our scores you get the caption-only or self-planned results below.
Results
Test set:
- 18 prompts: 6 kiñits × 3 grooves (chikchika 6/8, Ethio-jazz, shegaye 4/4), all with the same Amharic lyrics.
- 2 seeds per prompt, and 4 seeds for the adapter on the score path.
Metrics. Every render was transcribed with SheetSage2.
- Kiñit right: the requested kiñit ranks first of the six at the requested tonic. Songs YuE2 plans itself pick their own tonic, so they are judged on the pitch collection instead.
- Kiñit conformity: the duration-weighted share of melody notes inside the kiñit.
- Bar length held: the transcribed bar is within 5% of the requested length, or of half or double it.
- Tempo Acc2: the tempo is within ±4%, allowing octave errors.
- AudioBox PQ: AudioBox-Aesthetics production quality, on a 1–10 scale.
- Copy rate: the share of generated 16-token windows that also appear in the training songs.
| How the song is requested | Metric | Stock YuE2-3B | With this adapter |
|---|---|---|---|
| Tags → our kiñit score (recommended) | Kiñit right | 1.00 | 1.00 |
| Kiñit conformity | 0.994 | 0.998 | |
| Bar length held | 1.00 | 0.96 | |
| Tempo Acc2 | 0.69 | 0.79 | |
| AudioBox PQ | 8.23 | 8.29 | |
| Caption; YuE2 writes its own score | Right pitch collection | 0.33 | 0.61 |
| Kiñit conformity | 0.83 | 0.90 | |
| Bar length held | 0.56 | 0.83 | |
| Ran to the length cap | 11% | 0% | |
| AudioBox PQ | 8.29 | 8.24 | |
| Caption only (no score) | Right pitch collection | 0.33 | 0.53 |
| Kiñit conformity | 0.80 | 0.85 | |
| Bar length held | 0.28 | 0.56 | |
| AudioBox PQ | 8.29 | 8.26 | |
| All | Copy rate vs training songs | – | 0 |
Each score-path cell covers 36 renders for stock YuE2 and 72 for the adapter. Every other cell covers 36.
YuE2's own score snapped into the kiñit: the kiñit was right in only 25 of 36 renders. Bati major and anchihoye fail most, so this mode isn't recommended.
Right pitch collection by kiñit. Six prompts each, with stock YuE2 given the same captions:
| Kiñit | Caption only: stock → adapter | YuE2 writes its score: stock → adapter |
|---|---|---|
| tizita major | 1.00 → 0.67 | 1.00 → 0.67 |
| tizita minor | 0.00 → 0.50 | 0.00 → 0.67 |
| bati major | 0.00 → 0.17 | 0.00 → 0.50 |
| bati minor | 1.00 → 1.00 | 0.83 → 0.83 |
| ambassel | 0.00 → 0.50 | 0.00 → 0.50 |
| anchihoye | 0.00 → 0.33 | 0.00 → 0.50 |
Grooves from a caption alone (bar length held, 12 prompts each):
- chikchika: 0.25 → 0.75
- shegaye: 0.50 → 0.83
- Ethio-jazz: 0.08 → 0.08
With the score, Ethio-jazz holds its bar in 24 of 24 renders.
Limitations
- Kiñit choice from a caption alone is not reliable. Without a score, the right pitch collection comes out about half the time, with bati major at 17% and anchihoye at 33%. Use the kiñit score when the kiñit matters.
- Tizita major got slightly worse. The adapter traded some tizita major accuracy for the other kiñits: without a score it was right in 4 of 6 prompts instead of 6 of 6.
- Ethio-jazz from a caption alone rarely holds its meter. Use the score.
- About 1 chikchika render in 8 loses its 6/8 grouping and comes out grouped as 4/4 or 2/4, even with the score. Best-of-N selection checks the kiñit, not the meter, so listen and regenerate when this happens.
- Amharic intelligibility has not been measured reliably. Automatic speech recognition on sung vocals was too noisy to use: the character error rate was 0.64–0.86 on stock YuE2 and on an earlier adapter. No native-speaker listening test has been done yet.
- No new timbres. The training songs came from YuE2 itself, so the adapter re-weights what YuE2 can already do.
- Krar, masenqo, washint, begena and kebero are requested through descriptive proxies such as "krar-like bright plucked bowl lyre", and they won't sound authentic.
- Amharic ejectives are only approximated by the romanization.
- The evaluation is automatic and small. It relies on SheetSage2 transcription, beat tracking and AudioBox, over 18 prompts with 2–4 seeds. Treat small differences as noise.
- Numbers in lyrics aren't converted. Number-to-word normalization is off in this release, so write numbers and abbreviations as Amharic words.
Training
Data
- Prompts: 1,000 prompts over 6 kiñits × 5 grooves. 60% are sung in Amharic and 40% are instrumental.
- Scores: each prompt got a kiñit score. 87% came from the bundled composer, and 13% were YuE2's own score snapped into the kiñit.
- Rendering and gates: YuE2-3B rendered every score, and SheetSage2 transcribed every render. A render was kept only if it met all of these:
- kiñit conformity ≥ 0.90
- bar length held
- compound accents in 6/8 and 12/8 (salience ≥ 0.5)
- length within 30% of the score
- no truncation
- Result: 756 renders kept (76%, 30.5 h), split 719 for training and 37 for validation.
- Captions name the kiñit and its scale degrees, e.g. "Ethiopian tizita minor pentatonic mode (scale degrees 1 2 b3 5 b6), rolling 6/8 chikchika groove, 96 BPM".
- Regularizer: half of the training samples came from the regularizer pack of the YuE2 minted corpus, 4,732 songs generated by YuE2-3B. It preserves YuE2's token grammar and general ability.
- Lyrics: lines from the leyu-amharic transcripts (CC BY 4.0), plus 21 original lines, romanized with the bundled romanizer.
Procedure (full settings in training_config.yaml)
- Adapter: LoRA with rank 64, alpha 64 and no dropout, on the q/k/v/o and gate/up/down projections of all 28 AR layers.
- The base model was frozen in bf16, and the adapter was trained in fp32.
- The NAR branch and the decoder were not trained.
- How songs were shown: 50% caption only, 35% with the full score, 15% with the melody only. This way the adapter learns from text as well as from scores.
- Balance: sampling was split 50/50 between the two pitch collections. Otherwise the four-kiñit collection dominates.
- Optimizer: AdamW with learning rate 1e-4, betas 0.9/0.95 and no weight decay.
- 50 warm-up steps, then a cosine schedule (3,000-step horizon, 20% floor), stopped at step 1,500.
- Each step used 4 songs of up to 16,384 tokens, with gradient clipping at 1.0.
- Compute: 1 h 41 min on one RTX 5090, peaking at about 10 GiB.
- Checkpoint choice: the final checkpoint (step 1,500) was chosen over step 1,000 by caption-only generation results. Validation loss proved a misleading signal.
Files
| File | Contents |
|---|---|
ethio-music-yue2-ar-lora-r64.safetensors |
The adapter (format ethioyue-lora-v1; sha256 039d5cf3281af1d8…) |
adapter_config.json |
Base revisions, target modules, rank, merge rule, recommended settings |
training_config.yaml |
The training configuration |
code/ |
EthioYuE: kiñit score engine, Amharic romanizer, generator, LoRA and training code |
LICENSE |
Licence and attribution |
Licence and attribution
Released under CC BY-NC 4.0, for non-commercial use only. See LICENSE. It builds on:
- YuE2-3B by M-A-P (CC BY-NC 4.0). This repository contains no YuE2 weights.
- The YuE2 minted corpus (CC BY-NC 4.0).
- The leyu-amharic transcripts (CC BY 4.0).
The vendored YuE2 ABC validator in code/third_party/ is Apache-2.0.
Citation
If you use this adapter, please cite YuE2:
@article{yuan2026yue2,
title = {{YuE2}: Unifying Symbolic and Audio Music Generation at Frontier Quality},
author = {Yuan, Ruibin and Pan, Jiahao and Jiang, Junyan and Wu, Zhiyue and Zhou, Ziya and Sun, Jiankai and Li, Yizhi and Zhang, Ge and Gu, Yicheng and Tian, Zeyue and Dai, Junyu and Lin, Hanfeng and Li, Kai and Wu, Shangda and Liu, Xuanjie and Wang, Jiaming and Liu, Zihan and Wang, Yue and Ma, Yinghao and Yin, Hanzhi and Chen, Kangrui and Zhang, Xinyue and Ma, Ziyang and Liao, Mengqi and Zhao, Hejia and Huang, Guowei and Yan, Chao and Ke, Lei and Yu, Jianwei and Liu, Bei and Guo, Joe and Xue, Liumeng and Xia, Gus and Xue, Wei and Guo, Yike},
journal = {arXiv preprint arXiv:2609.33757},
year = {2026},
eprint = {2609.33757},
archivePrefix = {arXiv},
primaryClass = {eess.AS},
url = {https://arxiv.org/abs/2609.33757}
}
Model tree for Behailut/Ethio-Music-YuE2
Base model
m-a-p/YuE2-3B