atomtanstudio commited on
Commit
af2250e
·
verified ·
1 Parent(s): b572762

Add Industrial v1 step-500 LoRA

Browse files

Publish the verified step-500 YuE2 bundle with its matched v9 decoder companion, sidecar, checksum, and documented experimental quality status.

README.md CHANGED
@@ -10,6 +10,9 @@ tags:
10
  - audio
11
  - dream-pop
12
  - old-school-hip-hop
 
 
 
13
  - sound-and-vision
14
  ---
15
 
@@ -17,7 +20,7 @@ tags:
17
 
18
  A growing library of LoRAs by **Atomtan Studio**, with downloads, trigger words, settings, and compatibility notes together on one page.
19
 
20
- The library currently includes **DreamPop v2** and **Old School Hip-Hop** adapters for YuE2-3B. More adapters can be added to this same repository, organized by base model and style.
21
 
22
  ## Available LoRAs
23
 
@@ -25,6 +28,7 @@ The library currently includes **DreamPop v2** and **Old School Hip-Hop** adapte
25
  | --- | --- | --- | --- | --- |
26
  | **DreamPop v2** | YuE2-3B | 1,000 steps | `sv_dreampop` | [dreampop_sv_dreampop.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/dreampop/dreampop_sv_dreampop.safetensors?download=true) |
27
  | **Old School Hip-Hop** | YuE2-3B | 800-step checkpoint | `sv_oldschoolhiphop` | [sv_oldschoolhiphop.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/oldschoolhiphop/sv_oldschoolhiphop.safetensors?download=true) |
 
28
 
29
  ## DreamPop v2
30
 
@@ -159,6 +163,91 @@ ca176909d55998d4f78426439885ff79cc96ba4daaaaea22a9936590d6e72694 sv_oldschoolhi
159
 
160
  Also available in [SHA256SUMS](https://huggingface.co/atomtanstudio/lora-library/blob/main/SHA256SUMS).
161
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
162
  ## Library layout
163
 
164
  ~~~text
@@ -171,14 +260,18 @@ yue2/
171
  oldschoolhiphop/
172
  sv_oldschoolhiphop.safetensors
173
  lora.json
 
 
 
 
174
  ~~~
175
 
176
  New entries will be listed in the table above and stored in their own model/style folders within this repository. Check each entry's base model, loader requirements, and license before use.
177
 
178
  ## License and credits
179
 
180
- The DreamPop and Old School Hip-Hop adapters are shared under **CC BY-NC 4.0**, with the underlying YuE2 model terms applying. See the [CC BY-NC 4.0 license](https://creativecommons.org/licenses/by-nc/4.0/) and the [YuE2 model-weight license](https://github.com/multimodal-art-projection/YuE/blob/main/MODEL_LICENSE).
181
 
182
  YuE2's September 16, 2026 additional permission allows individuals to generate and monetize outputs subject to its stated conditions. That permission does not extend to commercial redistribution or sale of the model weights. It also does not grant rights in third-party material used as inputs or outputs.
183
 
184
- Base model and inference research: **YuE2 authors / Multimodal Art Projection**. Adapter training and Sound & Vision packaging: **Atomtan Studio**, using FL YuE2 tooling. This is a community adapter, not an official YuE2 release or an endorsement by the base-model authors or recording artists. Training recordings are not included in this repository.
 
10
  - audio
11
  - dream-pop
12
  - old-school-hip-hop
13
+ - industrial
14
+ - industrial-metal
15
+ - industrial-dance
16
  - sound-and-vision
17
  ---
18
 
 
20
 
21
  A growing library of LoRAs by **Atomtan Studio**, with downloads, trigger words, settings, and compatibility notes together on one page.
22
 
23
+ The library currently includes **DreamPop v2**, **Old School Hip-Hop** and **Industrial v1** adapters for YuE2-3B. More adapters can be added to this same repository, organized by base model and style.
24
 
25
  ## Available LoRAs
26
 
 
28
  | --- | --- | --- | --- | --- |
29
  | **DreamPop v2** | YuE2-3B | 1,000 steps | `sv_dreampop` | [dreampop_sv_dreampop.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/dreampop/dreampop_sv_dreampop.safetensors?download=true) |
30
  | **Old School Hip-Hop** | YuE2-3B | 800-step checkpoint | `sv_oldschoolhiphop` | [sv_oldschoolhiphop.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/oldschoolhiphop/sv_oldschoolhiphop.safetensors?download=true) |
31
+ | **Industrial v1** | YuE2-3B | 500-step checkpoint (experimental) | `sv_industrial` | [sv_industrial_step500.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/industrial/sv_industrial_step500.safetensors?download=true) |
32
 
33
  ## DreamPop v2
34
 
 
163
 
164
  Also available in [SHA256SUMS](https://huggingface.co/atomtanstudio/lora-library/blob/main/SHA256SUMS).
165
 
166
+ ## Industrial v1
167
+
168
+ An experimental combined industrial, industrial rock, industrial metal and industrial dance LoRA for **YuE2-3B**, with the shared trigger **`sv_industrial`**. This release contains the user-selected **500-step checkpoint** from a completed 1,000-step run.
169
+
170
+ **Quality status:** early listening found inconsistent fidelity, including thin or low-detail audio. Step 500 is released for testing; it is not an established best checkpoint. There was no held-out dataset or in-training listening evaluation.
171
+
172
+ ### Download and starting settings
173
+
174
+ Download [sv_industrial_step500.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/industrial/sv_industrial_step500.safetensors?download=true) and [lora.json](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/industrial/lora.json?download=true).
175
+
176
+ | Setting | Starting value |
177
+ | --- | --- |
178
+ | Trigger | `sv_industrial` |
179
+ | Style LoRA strength | `0.70` |
180
+ | Symbolic planning / CoT | `full` (melody and chords) |
181
+ | Guidance / CFG | `1.0` |
182
+ | Acoustic synthesis steps | `32` |
183
+ | Inference engine | Native YuE2 BF16 with v9-bundle loader support |
184
+ | Embedded v9 decoder companion strength | Fixed at `1.0` |
185
+
186
+ Example style prompt:
187
+
188
+ ```text
189
+ sv_industrial, industrial rock, industrial dance, dark, atmospheric, hypnotic, male vocals, clear melodic baritone, pulsing synthesizer bass, programmed drums, restrained distorted electric guitars, sparse verses, melodic chorus
190
+ ```
191
+
192
+ Supply your own lyrics with section labels such as `[Verse]`, `[Chorus]` and `[Bridge]`. Artist and vocalist identities were not trained as separate selector labels. Describe the audible traits you want.
193
+
194
+ These are test settings, not a proven optimum. Compare checkpoints using the same prompt, new lyrics, seed and runtime. This file includes one checkpoint, not the dataset or optimizer state.
195
+
196
+ ### Install in Sound & Vision
197
+
198
+ 1. Use a Sound & Vision backend that explicitly supports **`sound-vision-yue2-v9-bundle-v1`**, including the embedded companion loader. Support for the older DreamPop/Old School Hip-Hop export format alone is insufficient. This bundle was verified in the local deployment; this release does not assert that every public Sound & Vision revision contains that loader.
199
+ 2. Place the checkpoint and `lora.json` together under `<app>/loras/styles/Industrial-v1/`, or the equivalent folder under your `SOUND_VISION_LORAS` directory.
200
+ 3. Refresh or rescan the library, select **Industrial v1 / Step 500**, and set style strength to **0.70** manually. The sidecar supplies the trigger and generation defaults; it does not set the strength slider.
201
+ 4. Use the YuE2-3B base model and the default [YuE2-Vae listening decoder](https://huggingface.co/m-a-p/YuE2-Vae).
202
+
203
+ The sidecar's `preferred_step: 500` selects this published checkpoint. It is not a quality ranking.
204
+
205
+ ### Format and compatibility
206
+
207
+ The file is a **`sound-vision-yue2-v9-bundle-v1`** bundle containing unchanged BF16 AI Toolkit style-adapter tensors for both AR and NAR, plus the matched FP32 v9 NAR companion. It contains **844 tensors**: 448 style tensors and 396 companion tensors.
208
+
209
+ The embedded companion uses [Mothersuperior's v9 joint pair](https://huggingface.co/Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4). It must be applied at fixed strength 1.0 before the trained style adapter. Its four `vae2llm` / `llm2vae` weight and bias tensors are **full replacements**, not scaled deltas. The style-strength slider must not scale these companion replacements.
210
+
211
+ The matching v9 semantic head was used to tokenize the training audio; it is not included in this inference bundle. Do not load a second copy of the companion on top of this bundle.
212
+
213
+ This is a custom native-loader format. Generic ComfyUI, Diffusers, PEFT, AI Toolkit training-file ingestion and older Sound & Vision loaders are **not verified compatible** with this packaged file. The `.safetensors` extension does not establish loader compatibility.
214
+
215
+ ### Training and verification
216
+
217
+ | Property | Value |
218
+ | --- | --- |
219
+ | Base model | `m-a-p/YuE2-3B` |
220
+ | Base revision | `1a96eca688d6ae5d7f0feb88573fec89920fcd19` |
221
+ | Trainer | AI Toolkit 0.13.23, YuE2 joint AR + NAR |
222
+ | Completed run / released checkpoint | 1,000 steps / **500** |
223
+ | Rank / alpha | 32 / 32 |
224
+ | Main / AR learning rate | `5e-5` / `2e-5` |
225
+ | AR KL weight | `0.2` |
226
+ | Optimizer | AdamW 8-bit |
227
+ | Batch / accumulation | 1 / 1 |
228
+ | Dataset | 78 recordings, 48 credited acts, approximately 6 h 15 min |
229
+ | Prepared audio | 48 kHz stereo 16-bit FLAC |
230
+ | Acoustic training window | 60 seconds |
231
+ | Score conditioning | Full, 50% ABC dropout |
232
+ | Caption dropout / stem separation | 0 / disabled |
233
+ | File size | 258,066,232 bytes (about 258 MB) |
234
+
235
+ Prepared FLAC does not restore the fidelity of lossy source recordings. Captions used concise sound descriptions and reference lyrics; they were not all manually verified by listening.
236
+
237
+ The bundle's checksum, structure and loader compatibility were checked. Numerical merge checks on the bundle format verified sampled AR/NAR weight updates and all four decoder replacements. These are technical checks, not a held-out audio-quality benchmark.
238
+
239
+ ### Checksums
240
+
241
+ ```text
242
+ 18459e0de71f7e1686c8d1b0a1e7d0348a26549c770032ca65d99efcb36bda0a sv_industrial_step500.safetensors
243
+ ```
244
+
245
+ The embedded companion source SHA-256 is `585f303da1d5252d228d1e8ac6d4c4d11d970df9297406935cc8bdafa49cfa7e`.
246
+
247
+ See the repository [SHA256SUMS](https://huggingface.co/atomtanstudio/lora-library/blob/main/SHA256SUMS).
248
+
249
+ Detailed [Industrial model card](https://huggingface.co/atomtanstudio/lora-library/blob/main/yue2/industrial/README.md).
250
+
251
  ## Library layout
252
 
253
  ~~~text
 
260
  oldschoolhiphop/
261
  sv_oldschoolhiphop.safetensors
262
  lora.json
263
+ industrial/
264
+ sv_industrial_step500.safetensors
265
+ lora.json
266
+ README.md
267
  ~~~
268
 
269
  New entries will be listed in the table above and stored in their own model/style folders within this repository. Check each entry's base model, loader requirements, and license before use.
270
 
271
  ## License and credits
272
 
273
+ The DreamPop, Old School Hip-Hop and Industrial adapters are shared under **CC BY-NC 4.0**, with the underlying YuE2 model terms applying. See the [CC BY-NC 4.0 license](https://creativecommons.org/licenses/by-nc/4.0/) and the [YuE2 model-weight license](https://github.com/multimodal-art-projection/YuE/blob/main/MODEL_LICENSE).
274
 
275
  YuE2's September 16, 2026 additional permission allows individuals to generate and monetize outputs subject to its stated conditions. That permission does not extend to commercial redistribution or sale of the model weights. It also does not grant rights in third-party material used as inputs or outputs.
276
 
277
+ Base model and inference research: **YuE2 authors / Multimodal Art Projection**. Adapter training and Sound & Vision packaging: **Atomtan Studio**, using FL YuE2 tooling for DreamPop and Old School Hip-Hop, and [AI Toolkit](https://github.com/ostris/ai-toolkit) plus [Mothersuperior’s v9 tokenizer and decoder companion](https://huggingface.co/Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4) for Industrial. This is a community adapter, not an official YuE2 release or an endorsement by the base-model authors or recording artists. Training recordings are not included in this repository.
SHA256SUMS CHANGED
@@ -1,2 +1,3 @@
1
  617852d379c66b0407879421a353233d29cda8188cca82973a94c61b512ca03d yue2/dreampop/dreampop_sv_dreampop.safetensors
2
- ca176909d55998d4f78426439885ff79cc96ba4daaaaea22a9936590d6e72694 yue2/oldschoolhiphop/sv_oldschoolhiphop.safetensors
 
 
1
  617852d379c66b0407879421a353233d29cda8188cca82973a94c61b512ca03d yue2/dreampop/dreampop_sv_dreampop.safetensors
2
+ ca176909d55998d4f78426439885ff79cc96ba4daaaaea22a9936590d6e72694 yue2/oldschoolhiphop/sv_oldschoolhiphop.safetensors
3
+ 18459e0de71f7e1686c8d1b0a1e7d0348a26549c770032ca65d99efcb36bda0a yue2/industrial/sv_industrial_step500.safetensors
yue2/industrial/README.md ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Industrial v1 — step 500
2
+
3
+ An experimental combined industrial, industrial rock, industrial metal and industrial dance LoRA for **YuE2-3B**, with the shared trigger **`sv_industrial`**. This release contains the user-selected **500-step checkpoint** from a completed 1,000-step run.
4
+
5
+ **Quality status:** early listening found inconsistent fidelity, including thin or low-detail audio. Step 500 is released for testing; it is not an established best checkpoint. There was no held-out dataset or in-training listening evaluation.
6
+
7
+ ## Download and starting settings
8
+
9
+ Download [sv_industrial_step500.safetensors](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/industrial/sv_industrial_step500.safetensors?download=true) and [lora.json](https://huggingface.co/atomtanstudio/lora-library/resolve/main/yue2/industrial/lora.json?download=true).
10
+
11
+ | Setting | Starting value |
12
+ | --- | --- |
13
+ | Trigger | `sv_industrial` |
14
+ | Style LoRA strength | `0.70` |
15
+ | Symbolic planning / CoT | `full` (melody and chords) |
16
+ | Guidance / CFG | `1.0` |
17
+ | Acoustic synthesis steps | `32` |
18
+ | Inference engine | Native YuE2 BF16 with v9-bundle loader support |
19
+ | Embedded v9 decoder companion strength | Fixed at `1.0` |
20
+
21
+ Example style prompt:
22
+
23
+ ```text
24
+ sv_industrial, industrial rock, industrial dance, dark, atmospheric, hypnotic, male vocals, clear melodic baritone, pulsing synthesizer bass, programmed drums, restrained distorted electric guitars, sparse verses, melodic chorus
25
+ ```
26
+
27
+ Supply your own lyrics with section labels such as `[Verse]`, `[Chorus]` and `[Bridge]`. Artist and vocalist identities were not trained as separate selector labels. Describe the audible traits you want.
28
+
29
+ These are test settings, not a proven optimum. Compare checkpoints using the same prompt, new lyrics, seed and runtime. This file includes one checkpoint, not the dataset or optimizer state.
30
+
31
+ ## Install in Sound & Vision
32
+
33
+ 1. Use a Sound & Vision backend that explicitly supports **`sound-vision-yue2-v9-bundle-v1`**, including the embedded companion loader. Support for the older DreamPop/Old School Hip-Hop export format alone is insufficient. This bundle was verified in the local deployment; this release does not assert that every public Sound & Vision revision contains that loader.
34
+ 2. Place the checkpoint and `lora.json` together under `<app>/loras/styles/Industrial-v1/`, or the equivalent folder under your `SOUND_VISION_LORAS` directory.
35
+ 3. Refresh or rescan the library, select **Industrial v1 / Step 500**, and set style strength to **0.70** manually. The sidecar supplies the trigger and generation defaults; it does not set the strength slider.
36
+ 4. Use the YuE2-3B base model and the default [YuE2-Vae listening decoder](https://huggingface.co/m-a-p/YuE2-Vae).
37
+
38
+ The sidecar's `preferred_step: 500` selects this published checkpoint. It is not a quality ranking.
39
+
40
+ ## Format and compatibility
41
+
42
+ The file is a **`sound-vision-yue2-v9-bundle-v1`** bundle containing unchanged BF16 AI Toolkit style-adapter tensors for both AR and NAR, plus the matched FP32 v9 NAR companion. It contains **844 tensors**: 448 style tensors and 396 companion tensors.
43
+
44
+ The embedded companion uses [Mothersuperior's v9 joint pair](https://huggingface.co/Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4). It must be applied at fixed strength 1.0 before the trained style adapter. Its four `vae2llm` / `llm2vae` weight and bias tensors are **full replacements**, not scaled deltas. The style-strength slider must not scale these companion replacements.
45
+
46
+ The matching v9 semantic head was used to tokenize the training audio; it is not included in this inference bundle. Do not load a second copy of the companion on top of this bundle.
47
+
48
+ This is a custom native-loader format. Generic ComfyUI, Diffusers, PEFT, AI Toolkit training-file ingestion and older Sound & Vision loaders are **not verified compatible** with this packaged file. The `.safetensors` extension does not establish loader compatibility.
49
+
50
+ ## Training and verification
51
+
52
+ | Property | Value |
53
+ | --- | --- |
54
+ | Base model | `m-a-p/YuE2-3B` |
55
+ | Base revision | `1a96eca688d6ae5d7f0feb88573fec89920fcd19` |
56
+ | Trainer | AI Toolkit 0.13.23, YuE2 joint AR + NAR |
57
+ | Completed run / released checkpoint | 1,000 steps / **500** |
58
+ | Rank / alpha | 32 / 32 |
59
+ | Main / AR learning rate | `5e-5` / `2e-5` |
60
+ | AR KL weight | `0.2` |
61
+ | Optimizer | AdamW 8-bit |
62
+ | Batch / accumulation | 1 / 1 |
63
+ | Dataset | 78 recordings, 48 credited acts, approximately 6 h 15 min |
64
+ | Prepared audio | 48 kHz stereo 16-bit FLAC |
65
+ | Acoustic training window | 60 seconds |
66
+ | Score conditioning | Full, 50% ABC dropout |
67
+ | Caption dropout / stem separation | 0 / disabled |
68
+ | File size | 258,066,232 bytes (about 258 MB) |
69
+
70
+ Prepared FLAC does not restore the fidelity of lossy source recordings. Captions used concise sound descriptions and reference lyrics; they were not all manually verified by listening.
71
+
72
+ The bundle's checksum, structure and loader compatibility were checked. Numerical merge checks on the bundle format verified sampled AR/NAR weight updates and all four decoder replacements. These are technical checks, not a held-out audio-quality benchmark.
73
+
74
+ ## Checksums
75
+
76
+ ```text
77
+ 18459e0de71f7e1686c8d1b0a1e7d0348a26549c770032ca65d99efcb36bda0a sv_industrial_step500.safetensors
78
+ ```
79
+
80
+ The embedded companion source SHA-256 is `585f303da1d5252d228d1e8ac6d4c4d11d970df9297406935cc8bdafa49cfa7e`.
81
+
82
+ See the repository [SHA256SUMS](https://huggingface.co/atomtanstudio/lora-library/blob/main/SHA256SUMS).
83
+
84
+ ## License and credits
85
+
86
+ Shared under **CC BY-NC 4.0**, subject to the underlying model and companion terms. Base model: [YuE2 authors / Multimodal Art Projection](https://huggingface.co/m-a-p/YuE2-3B). Real-audio tokenizer and v9 companion: [Mothersuperior](https://huggingface.co/Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4). Training toolkit: [ostris/ai-toolkit](https://github.com/ostris/ai-toolkit). Training and Sound & Vision packaging: Atomtan Studio.
87
+
88
+ This community release is not endorsed by the base-model authors or recording artists. Training recordings and lyric captions are not distributed in this repository.
89
+
yue2/industrial/lora.json ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "name": "Industrial v1",
3
+ "kind": "style",
4
+ "trigger": "sv_industrial",
5
+ "preferred_step": 500,
6
+ "generation": {
7
+ "cot": "full",
8
+ "cfg_scale": 1,
9
+ "ode_steps": 32
10
+ }
11
+ }
yue2/industrial/sv_industrial_step500.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:18459e0de71f7e1686c8d1b0a1e7d0348a26549c770032ca65d99efcb36bda0a
3
+ size 258066232