lora-library / README.md
atomtanstudio's picture
Add Industrial v1 step-500 LoRA
af2250e verified
|
Raw History Blame Contribute Delete
16.7 kB
---
license: cc-by-nc-4.0
base_model: m-a-p/YuE2-3B
base_model_relation: adapter
pipeline_tag: text-to-audio
tags:
- lora
- yue2
- music-generation
- audio
- dream-pop
- old-school-hip-hop
- industrial
- industrial-metal
- industrial-dance
- sound-and-vision
---
# Atomtan Studio LoRA Library
A growing library of LoRAs by **Atomtan Studio**, with downloads, trigger words, settings, and compatibility notes together on one page.
The library currently includes **DreamPop v2**, **Old School Hip-Hop** and **Industrial v1** adapters for YuE2-3B. More adapters can be added to this same repository, organized by base model and style.
## Available LoRAs
| LoRA | Base model | Checkpoint | Trigger | Download |
| --- | --- | --- | --- | --- |
| **DreamPop v2** | YuE2-3B | 1,000 steps | `sv_dreampop` | [dreampop_sv_dreampop.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/dreampop/dreampop_sv_dreampop.safetensors?download=true) |
| **Old School Hip-Hop** | YuE2-3B | 800-step checkpoint | `sv_oldschoolhiphop` | [sv_oldschoolhiphop.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/oldschoolhiphop/sv_oldschoolhiphop.safetensors?download=true) |
| **Industrial v1** | YuE2-3B | 500-step checkpoint (experimental) | `sv_industrial` | [sv_industrial_step500.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/industrial/sv_industrial_step500.safetensors?download=true) |
## DreamPop v2
DreamPop v2 was trained from fresh adapters on 32 reviewed recordings across 12 artists, using the FL YuE2 `native_joint_v1` training recipe. It trains both the autoregressive (AR) and acoustic (NAR) components of YuE2. This release contains the final **1,000-step checkpoint** in a single combined export for Sound & Vision.
This is an adapter, not a standalone music model. Download and install the base model separately: [m-a-p/YuE2-3B](https://huggingface.co/m-a-p/YuE2-3B).
### Trigger and starting settings
Use **`sv_dreampop`** in the style prompt. For example:
~~~text
sv_dreampop, dream pop, shimmering guitars, soft vocals, spacious reverb, steady mid-tempo drums
~~~
| Setting | Starting value |
| --- | --- |
| LoRA strength | `1.0` |
| Symbolic planning / CoT | `off` |
| Guidance / CFG | `1.0` |
| Acoustic synthesis steps | `32` |
| Inference engine | Native YuE2 BF16 |
Use your own lyrics and adjust the musical description for the song. These settings are a starting point, not a claim that one setting or checkpoint is best for every prompt.
### Install in Sound & Vision
Use [Sound & Vision](https://github.com/atomtanstudio/sound-and-vision) with support for native YuE2 LoRA exports, available in commit [`440e4ee`](https://github.com/atomtanstudio/sound-and-vision/commit/440e4eed0391d94700f84e843669289cf7ec5cf2).
1. Download the checkpoint above and the accompanying [lora.json](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/dreampop/lora.json?download=true).
2. Place both files on the computer running the backend, under `<app>/loras/styles/DreamPop-v2/`.
3. Refresh the app or rescan the LoRA folder, then select **DreamPop v2** and **Step 1,000** under **Trained styles & artists**.
4. Start with strength `1.0`. Sound & Vision reads the trigger and generation defaults from the sidecar and checkpoint metadata.
If `SOUND_VISION_LORAS` points to a custom directory, use that directory instead of `<app>/loras`.
### Format and compatibility
The uploaded file uses **`sound-vision-yue2-native-export-v1`**. It combines the original AR and NAR adapter pair into fused projections while preserving their tensor values and scaling. Both learned components are included; Sound & Vision does not need a second adapter file for this release.
This is a custom YuE2 adapter layout. The `.safetensors` extension alone does not establish compatibility with other loaders. Direct loading in generic ComfyUI, Diffusers, or PEFT loaders has not been verified. FL YuE2 workflows that expect separate AR and NAR files need the original pair or an appropriate conversion; do not supply this combined file as either member of that pair.
### Training and verification
| Property | Value |
| --- | --- |
| Base-model revision | `1a96eca688d6ae5d7f0feb88573fec89920fcd19` |
| Training recipe | FL YuE2 `native_joint_v1` |
| Training steps | 1,000 |
| LoRA rank | 32 |
| Learning rate | `0.0001`, constant |
| Dataset | 32 recordings across 12 artists |
| Components | AR and NAR |
| Export precision | FP32 |
| Tensor count | 672 |
| File size | 234,969,024 bytes, approximately 235 MB |
The completed run recorded finite losses and active gradients for both components. Export verification checked every source tensor and preserved unit scaling. The downloadable file was checked against the export's SHA-256 digest after renaming.
These are training and file-integrity checks, not a held-out quality benchmark. Training previews used a 120-second diagnostic cap; that cap is not a recommended generation length or evidence that a full song will end naturally. Results vary with lyrics, prompt, seed, and runtime. The 1,000-step checkpoint is the final checkpoint from this run, not an independently established best checkpoint.
### Checksum
~~~text
617852d379c66b0407879421a353233d29cda8188cca82973a94c61b512ca03d dreampop_sv_dreampop.safetensors
~~~
Also available in [SHA256SUMS](https://huggingface.co/atomtanstudio/lora-library/blob/main/SHA256SUMS).
## Old School Hip-Hop
Old School Hip-Hop was trained on **90 recordings across 18 artists** from the Sound & Vision library, with the dataset selected to cover late-1980s and early-1990s boom-bap, sample-driven production, varied regional styles, and rhythmic rap delivery. It uses the FL YuE2 `native_joint_v1` training recipe and trains both the autoregressive (AR) and acoustic (NAR) components of YuE2. This release contains the **800-step checkpoint** for testing; it is not an independently established best checkpoint.
This is an adapter, not a standalone music model. Download and install the base model separately: [m-a-p/YuE2-3B](https://huggingface.co/m-a-p/YuE2-3B).
### Trigger and starting settings
Use **`sv_oldschoolhiphop`** in the style prompt. For example:
~~~text
sv_oldschoolhiphop, old-school hip-hop, boom-bap drums, dusty samples, swung hi-hats, warm bass, chopped soul and jazz textures, muted Rhodes, horn stabs, vinyl crackle, confident rhythmic rap delivery
~~~
| Setting | Starting value |
| --- | --- |
| LoRA strength | `1.0` |
| Symbolic planning / CoT | `off` |
| Guidance / CFG | `1.0` |
| Acoustic synthesis steps | `32` |
| Inference engine | Native YuE2 BF16 |
Use your own lyrics and adjust the musical description for the song. These settings are a starting point, not a claim that one setting or checkpoint is best for every prompt.
### Install in Sound & Vision
Use [Sound & Vision](https://github.com/atomtanstudio/sound-and-vision) with support for native YuE2 LoRA exports.
1. Download the checkpoint above and the accompanying [lora.json](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/oldschoolhiphop/lora.json?download=true).
2. Place both files on the computer running the backend, under `<app>/loras/styles/OldSchoolHipHop/`.
3. Refresh the app or rescan the LoRA folder, then select **Old School Hip-Hop** and **Step 800** under **Trained styles & artists**.
4. Start with strength `1.0`. Sound & Vision reads the trigger and generation defaults from the sidecar and checkpoint metadata.
If `SOUND_VISION_LORAS` points to a custom directory, use that directory instead of `<app>/loras`.
### Format and compatibility
The uploaded file uses **`sound-vision-yue2-native-export-v1`**. It combines the original AR and NAR adapter pair into fused projections while preserving their tensor values and scaling. Both learned components are included; Sound & Vision does not need a second adapter file for this release.
This is a custom YuE2 adapter layout. The `.safetensors` extension alone does not establish compatibility with other loaders. Direct loading in generic ComfyUI, Diffusers, or PEFT loaders has not been verified. FL YuE2 workflows that expect separate AR and NAR files need the original pair or an appropriate conversion; do not supply this combined file as either member of that pair.
### Training and verification
| Property | Value |
| --- | --- |
| Base-model revision | `1a96eca688d6ae5d7f0feb88573fec89920fcd19` |
| Training recipe | FL YuE2 `native_joint_v1` |
| Training steps | 1,000 |
| Released checkpoint | 800 |
| LoRA rank | 32 |
| Learning rate | `0.0001`, constant |
| Dataset | 90 recordings across 18 artists |
| Components | AR and NAR |
| File size | 234,969,032 bytes, approximately 235 MB |
The local checkpoint was verified against the SHA-256 digest shown below after the user-selected rename. These are training and file-integrity checks, not a held-out quality benchmark. Results vary with lyrics, prompt, seed, and runtime.
### Checksum
~~~text
ca176909d55998d4f78426439885ff79cc96ba4daaaaea22a9936590d6e72694 sv_oldschoolhiphop.safetensors
~~~
Also available in [SHA256SUMS](https://huggingface.co/atomtanstudio/lora-library/blob/main/SHA256SUMS).
## Industrial v1
An experimental combined industrial, industrial rock, industrial metal and industrial dance LoRA for **YuE2-3B**, with the shared trigger **`sv_industrial`**. This release contains the user-selected **500-step checkpoint** from a completed 1,000-step run.
**Quality status:** early listening found inconsistent fidelity, including thin or low-detail audio. Step 500 is released for testing; it is not an established best checkpoint. There was no held-out dataset or in-training listening evaluation.
### Download and starting settings
Download [sv_industrial_step500.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/industrial/sv_industrial_step500.safetensors?download=true) and [lora.json](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/industrial/lora.json?download=true).
| Setting | Starting value |
| --- | --- |
| Trigger | `sv_industrial` |
| Style LoRA strength | `0.70` |
| Symbolic planning / CoT | `full` (melody and chords) |
| Guidance / CFG | `1.0` |
| Acoustic synthesis steps | `32` |
| Inference engine | Native YuE2 BF16 with v9-bundle loader support |
| Embedded v9 decoder companion strength | Fixed at `1.0` |
Example style prompt:
```text
sv_industrial, industrial rock, industrial dance, dark, atmospheric, hypnotic, male vocals, clear melodic baritone, pulsing synthesizer bass, programmed drums, restrained distorted electric guitars, sparse verses, melodic chorus
```
Supply your own lyrics with section labels such as `[Verse]`, `[Chorus]` and `[Bridge]`. Artist and vocalist identities were not trained as separate selector labels. Describe the audible traits you want.
These are test settings, not a proven optimum. Compare checkpoints using the same prompt, new lyrics, seed and runtime. This file includes one checkpoint, not the dataset or optimizer state.
### Install in Sound & Vision
1. Use a Sound & Vision backend that explicitly supports **`sound-vision-yue2-v9-bundle-v1`**, including the embedded companion loader. Support for the older DreamPop/Old School Hip-Hop export format alone is insufficient. This bundle was verified in the local deployment; this release does not assert that every public Sound & Vision revision contains that loader.
2. Place the checkpoint and `lora.json` together under `<app>/loras/styles/Industrial-v1/`, or the equivalent folder under your `SOUND_VISION_LORAS` directory.
3. Refresh or rescan the library, select **Industrial v1 / Step 500**, and set style strength to **0.70** manually. The sidecar supplies the trigger and generation defaults; it does not set the strength slider.
4. Use the YuE2-3B base model and the default [YuE2-Vae listening decoder](https://huggingface.co/m-a-p/YuE2-Vae).
The sidecar's `preferred_step: 500` selects this published checkpoint. It is not a quality ranking.
### Format and compatibility
The file is a **`sound-vision-yue2-v9-bundle-v1`** bundle containing unchanged BF16 AI Toolkit style-adapter tensors for both AR and NAR, plus the matched FP32 v9 NAR companion. It contains **844 tensors**: 448 style tensors and 396 companion tensors.
The embedded companion uses [Mothersuperior's v9 joint pair](https://huggingface.co/Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4). It must be applied at fixed strength 1.0 before the trained style adapter. Its four `vae2llm` / `llm2vae` weight and bias tensors are **full replacements**, not scaled deltas. The style-strength slider must not scale these companion replacements.
The matching v9 semantic head was used to tokenize the training audio; it is not included in this inference bundle. Do not load a second copy of the companion on top of this bundle.
This is a custom native-loader format. Generic ComfyUI, Diffusers, PEFT, AI Toolkit training-file ingestion and older Sound & Vision loaders are **not verified compatible** with this packaged file. The `.safetensors` extension does not establish loader compatibility.
### Training and verification
| Property | Value |
| --- | --- |
| Base model | `m-a-p/YuE2-3B` |
| Base revision | `1a96eca688d6ae5d7f0feb88573fec89920fcd19` |
| Trainer | AI Toolkit 0.13.23, YuE2 joint AR + NAR |
| Completed run / released checkpoint | 1,000 steps / **500** |
| Rank / alpha | 32 / 32 |
| Main / AR learning rate | `5e-5` / `2e-5` |
| AR KL weight | `0.2` |
| Optimizer | AdamW 8-bit |
| Batch / accumulation | 1 / 1 |
| Dataset | 78 recordings, 48 credited acts, approximately 6 h 15 min |
| Prepared audio | 48 kHz stereo 16-bit FLAC |
| Acoustic training window | 60 seconds |
| Score conditioning | Full, 50% ABC dropout |
| Caption dropout / stem separation | 0 / disabled |
| File size | 258,066,232 bytes (about 258 MB) |
Prepared FLAC does not restore the fidelity of lossy source recordings. Captions used concise sound descriptions and reference lyrics; they were not all manually verified by listening.
The bundle's checksum, structure and loader compatibility were checked. Numerical merge checks on the bundle format verified sampled AR/NAR weight updates and all four decoder replacements. These are technical checks, not a held-out audio-quality benchmark.
### Checksums
```text
18459e0de71f7e1686c8d1b0a1e7d0348a26549c770032ca65d99efcb36bda0a sv_industrial_step500.safetensors
```
The embedded companion source SHA-256 is `585f303da1d5252d228d1e8ac6d4c4d11d970df9297406935cc8bdafa49cfa7e`.
See the repository [SHA256SUMS](https://huggingface.co/atomtanstudio/lora-library/blob/main/SHA256SUMS).
Detailed [Industrial model card](https://huggingface.co/atomtanstudio/lora-library/blob/main/yue2/industrial/README.md).
## Library layout
~~~text
README.md
SHA256SUMS
yue2/
dreampop/
dreampop_sv_dreampop.safetensors
lora.json
oldschoolhiphop/
sv_oldschoolhiphop.safetensors
lora.json
industrial/
sv_industrial_step500.safetensors
lora.json
README.md
~~~
New entries will be listed in the table above and stored in their own model/style folders within this repository. Check each entry's base model, loader requirements, and license before use.
## License and credits
The DreamPop, Old School Hip-Hop and Industrial adapters are shared under **CC BY-NC 4.0**, with the underlying YuE2 model terms applying. See the [CC BY-NC 4.0 license](https://creativecommons.org/licenses/by-nc/4.0/) and the [YuE2 model-weight license](https://github.com/multimodal-art-projection/YuE/blob/main/MODEL_LICENSE).
YuE2's September 16, 2026 additional permission allows individuals to generate and monetize outputs subject to its stated conditions. That permission does not extend to commercial redistribution or sale of the model weights. It also does not grant rights in third-party material used as inputs or outputs.
Base model and inference research: **YuE2 authors / Multimodal Art Projection**. Adapter training and Sound & Vision packaging: **Atomtan Studio**, using FL YuE2 tooling for DreamPop and Old School Hip-Hop, and [AI Toolkit](https://github.com/ostris/ai-toolkit) plus [Mothersuperior’s v9 tokenizer and decoder companion](https://huggingface.co/Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4) for Industrial. This is a community adapter, not an official YuE2 release or an endorsement by the base-model authors or recording artists. Training recordings are not included in this repository.