lora-library / README.md
atomtanstudio's picture
Add Industrial v1 step-500 LoRA
af2250e verified
|
Raw History Blame Contribute Delete
16.7 kB
metadata
license: cc-by-nc-4.0
base_model: m-a-p/YuE2-3B
base_model_relation: adapter
pipeline_tag: text-to-audio
tags:
  - lora
  - yue2
  - music-generation
  - audio
  - dream-pop
  - old-school-hip-hop
  - industrial
  - industrial-metal
  - industrial-dance
  - sound-and-vision

Atomtan Studio LoRA Library

A growing library of LoRAs by Atomtan Studio, with downloads, trigger words, settings, and compatibility notes together on one page.

The library currently includes DreamPop v2, Old School Hip-Hop and Industrial v1 adapters for YuE2-3B. More adapters can be added to this same repository, organized by base model and style.

Available LoRAs

LoRA Base model Checkpoint Trigger Download
DreamPop v2 YuE2-3B 1,000 steps sv_dreampop dreampop_sv_dreampop.safetensors
Old School Hip-Hop YuE2-3B 800-step checkpoint sv_oldschoolhiphop sv_oldschoolhiphop.safetensors
Industrial v1 YuE2-3B 500-step checkpoint (experimental) sv_industrial sv_industrial_step500.safetensors

DreamPop v2

DreamPop v2 was trained from fresh adapters on 32 reviewed recordings across 12 artists, using the FL YuE2 native_joint_v1 training recipe. It trains both the autoregressive (AR) and acoustic (NAR) components of YuE2. This release contains the final 1,000-step checkpoint in a single combined export for Sound & Vision.

This is an adapter, not a standalone music model. Download and install the base model separately: m-a-p/YuE2-3B.

Trigger and starting settings

Use sv_dreampop in the style prompt. For example:

sv_dreampop, dream pop, shimmering guitars, soft vocals, spacious reverb, steady mid-tempo drums
Setting Starting value
LoRA strength 1.0
Symbolic planning / CoT off
Guidance / CFG 1.0
Acoustic synthesis steps 32
Inference engine Native YuE2 BF16

Use your own lyrics and adjust the musical description for the song. These settings are a starting point, not a claim that one setting or checkpoint is best for every prompt.

Install in Sound & Vision

Use Sound & Vision with support for native YuE2 LoRA exports, available in commit 440e4ee.

  1. Download the checkpoint above and the accompanying lora.json.
  2. Place both files on the computer running the backend, under <app>/loras/styles/DreamPop-v2/.
  3. Refresh the app or rescan the LoRA folder, then select DreamPop v2 and Step 1,000 under Trained styles & artists.
  4. Start with strength 1.0. Sound & Vision reads the trigger and generation defaults from the sidecar and checkpoint metadata.

If SOUND_VISION_LORAS points to a custom directory, use that directory instead of <app>/loras.

Format and compatibility

The uploaded file uses sound-vision-yue2-native-export-v1. It combines the original AR and NAR adapter pair into fused projections while preserving their tensor values and scaling. Both learned components are included; Sound & Vision does not need a second adapter file for this release.

This is a custom YuE2 adapter layout. The .safetensors extension alone does not establish compatibility with other loaders. Direct loading in generic ComfyUI, Diffusers, or PEFT loaders has not been verified. FL YuE2 workflows that expect separate AR and NAR files need the original pair or an appropriate conversion; do not supply this combined file as either member of that pair.

Training and verification

Property Value
Base-model revision 1a96eca688d6ae5d7f0feb88573fec89920fcd19
Training recipe FL YuE2 native_joint_v1
Training steps 1,000
LoRA rank 32
Learning rate 0.0001, constant
Dataset 32 recordings across 12 artists
Components AR and NAR
Export precision FP32
Tensor count 672
File size 234,969,024 bytes, approximately 235 MB

The completed run recorded finite losses and active gradients for both components. Export verification checked every source tensor and preserved unit scaling. The downloadable file was checked against the export's SHA-256 digest after renaming.

These are training and file-integrity checks, not a held-out quality benchmark. Training previews used a 120-second diagnostic cap; that cap is not a recommended generation length or evidence that a full song will end naturally. Results vary with lyrics, prompt, seed, and runtime. The 1,000-step checkpoint is the final checkpoint from this run, not an independently established best checkpoint.

Checksum

617852d379c66b0407879421a353233d29cda8188cca82973a94c61b512ca03d  dreampop_sv_dreampop.safetensors

Also available in SHA256SUMS.

Old School Hip-Hop

Old School Hip-Hop was trained on 90 recordings across 18 artists from the Sound & Vision library, with the dataset selected to cover late-1980s and early-1990s boom-bap, sample-driven production, varied regional styles, and rhythmic rap delivery. It uses the FL YuE2 native_joint_v1 training recipe and trains both the autoregressive (AR) and acoustic (NAR) components of YuE2. This release contains the 800-step checkpoint for testing; it is not an independently established best checkpoint.

This is an adapter, not a standalone music model. Download and install the base model separately: m-a-p/YuE2-3B.

Trigger and starting settings

Use sv_oldschoolhiphop in the style prompt. For example:

sv_oldschoolhiphop, old-school hip-hop, boom-bap drums, dusty samples, swung hi-hats, warm bass, chopped soul and jazz textures, muted Rhodes, horn stabs, vinyl crackle, confident rhythmic rap delivery
Setting Starting value
LoRA strength 1.0
Symbolic planning / CoT off
Guidance / CFG 1.0
Acoustic synthesis steps 32
Inference engine Native YuE2 BF16

Use your own lyrics and adjust the musical description for the song. These settings are a starting point, not a claim that one setting or checkpoint is best for every prompt.

Install in Sound & Vision

Use Sound & Vision with support for native YuE2 LoRA exports.

  1. Download the checkpoint above and the accompanying lora.json.
  2. Place both files on the computer running the backend, under <app>/loras/styles/OldSchoolHipHop/.
  3. Refresh the app or rescan the LoRA folder, then select Old School Hip-Hop and Step 800 under Trained styles & artists.
  4. Start with strength 1.0. Sound & Vision reads the trigger and generation defaults from the sidecar and checkpoint metadata.

If SOUND_VISION_LORAS points to a custom directory, use that directory instead of <app>/loras.

Format and compatibility

The uploaded file uses sound-vision-yue2-native-export-v1. It combines the original AR and NAR adapter pair into fused projections while preserving their tensor values and scaling. Both learned components are included; Sound & Vision does not need a second adapter file for this release.

This is a custom YuE2 adapter layout. The .safetensors extension alone does not establish compatibility with other loaders. Direct loading in generic ComfyUI, Diffusers, or PEFT loaders has not been verified. FL YuE2 workflows that expect separate AR and NAR files need the original pair or an appropriate conversion; do not supply this combined file as either member of that pair.

Training and verification

Property Value
Base-model revision 1a96eca688d6ae5d7f0feb88573fec89920fcd19
Training recipe FL YuE2 native_joint_v1
Training steps 1,000
Released checkpoint 800
LoRA rank 32
Learning rate 0.0001, constant
Dataset 90 recordings across 18 artists
Components AR and NAR
File size 234,969,032 bytes, approximately 235 MB

The local checkpoint was verified against the SHA-256 digest shown below after the user-selected rename. These are training and file-integrity checks, not a held-out quality benchmark. Results vary with lyrics, prompt, seed, and runtime.

Checksum

ca176909d55998d4f78426439885ff79cc96ba4daaaaea22a9936590d6e72694  sv_oldschoolhiphop.safetensors

Also available in SHA256SUMS.

Industrial v1

An experimental combined industrial, industrial rock, industrial metal and industrial dance LoRA for YuE2-3B, with the shared trigger sv_industrial. This release contains the user-selected 500-step checkpoint from a completed 1,000-step run.

Quality status: early listening found inconsistent fidelity, including thin or low-detail audio. Step 500 is released for testing; it is not an established best checkpoint. There was no held-out dataset or in-training listening evaluation.

Download and starting settings

Download sv_industrial_step500.safetensors and lora.json.

Setting Starting value
Trigger sv_industrial
Style LoRA strength 0.70
Symbolic planning / CoT full (melody and chords)
Guidance / CFG 1.0
Acoustic synthesis steps 32
Inference engine Native YuE2 BF16 with v9-bundle loader support
Embedded v9 decoder companion strength Fixed at 1.0

Example style prompt:

sv_industrial, industrial rock, industrial dance, dark, atmospheric, hypnotic, male vocals, clear melodic baritone, pulsing synthesizer bass, programmed drums, restrained distorted electric guitars, sparse verses, melodic chorus

Supply your own lyrics with section labels such as [Verse], [Chorus] and [Bridge]. Artist and vocalist identities were not trained as separate selector labels. Describe the audible traits you want.

These are test settings, not a proven optimum. Compare checkpoints using the same prompt, new lyrics, seed and runtime. This file includes one checkpoint, not the dataset or optimizer state.

Install in Sound & Vision

  1. Use a Sound & Vision backend that explicitly supports sound-vision-yue2-v9-bundle-v1, including the embedded companion loader. Support for the older DreamPop/Old School Hip-Hop export format alone is insufficient. This bundle was verified in the local deployment; this release does not assert that every public Sound & Vision revision contains that loader.
  2. Place the checkpoint and lora.json together under <app>/loras/styles/Industrial-v1/, or the equivalent folder under your SOUND_VISION_LORAS directory.
  3. Refresh or rescan the library, select Industrial v1 / Step 500, and set style strength to 0.70 manually. The sidecar supplies the trigger and generation defaults; it does not set the strength slider.
  4. Use the YuE2-3B base model and the default YuE2-Vae listening decoder.

The sidecar's preferred_step: 500 selects this published checkpoint. It is not a quality ranking.

Format and compatibility

The file is a sound-vision-yue2-v9-bundle-v1 bundle containing unchanged BF16 AI Toolkit style-adapter tensors for both AR and NAR, plus the matched FP32 v9 NAR companion. It contains 844 tensors: 448 style tensors and 396 companion tensors.

The embedded companion uses Mothersuperior's v9 joint pair. It must be applied at fixed strength 1.0 before the trained style adapter. Its four vae2llm / llm2vae weight and bias tensors are full replacements, not scaled deltas. The style-strength slider must not scale these companion replacements.

The matching v9 semantic head was used to tokenize the training audio; it is not included in this inference bundle. Do not load a second copy of the companion on top of this bundle.

This is a custom native-loader format. Generic ComfyUI, Diffusers, PEFT, AI Toolkit training-file ingestion and older Sound & Vision loaders are not verified compatible with this packaged file. The .safetensors extension does not establish loader compatibility.

Training and verification

Property Value
Base model m-a-p/YuE2-3B
Base revision 1a96eca688d6ae5d7f0feb88573fec89920fcd19
Trainer AI Toolkit 0.13.23, YuE2 joint AR + NAR
Completed run / released checkpoint 1,000 steps / 500
Rank / alpha 32 / 32
Main / AR learning rate 5e-5 / 2e-5
AR KL weight 0.2
Optimizer AdamW 8-bit
Batch / accumulation 1 / 1
Dataset 78 recordings, 48 credited acts, approximately 6 h 15 min
Prepared audio 48 kHz stereo 16-bit FLAC
Acoustic training window 60 seconds
Score conditioning Full, 50% ABC dropout
Caption dropout / stem separation 0 / disabled
File size 258,066,232 bytes (about 258 MB)

Prepared FLAC does not restore the fidelity of lossy source recordings. Captions used concise sound descriptions and reference lyrics; they were not all manually verified by listening.

The bundle's checksum, structure and loader compatibility were checked. Numerical merge checks on the bundle format verified sampled AR/NAR weight updates and all four decoder replacements. These are technical checks, not a held-out audio-quality benchmark.

Checksums

18459e0de71f7e1686c8d1b0a1e7d0348a26549c770032ca65d99efcb36bda0a  sv_industrial_step500.safetensors

The embedded companion source SHA-256 is 585f303da1d5252d228d1e8ac6d4c4d11d970df9297406935cc8bdafa49cfa7e.

See the repository SHA256SUMS.

Detailed Industrial model card.

Library layout

README.md
SHA256SUMS
yue2/
  dreampop/
    dreampop_sv_dreampop.safetensors
    lora.json
  oldschoolhiphop/
    sv_oldschoolhiphop.safetensors
    lora.json
  industrial/
    sv_industrial_step500.safetensors
    lora.json
    README.md

New entries will be listed in the table above and stored in their own model/style folders within this repository. Check each entry's base model, loader requirements, and license before use.

License and credits

The DreamPop, Old School Hip-Hop and Industrial adapters are shared under CC BY-NC 4.0, with the underlying YuE2 model terms applying. See the CC BY-NC 4.0 license and the YuE2 model-weight license.

YuE2's September 16, 2026 additional permission allows individuals to generate and monetize outputs subject to its stated conditions. That permission does not extend to commercial redistribution or sale of the model weights. It also does not grant rights in third-party material used as inputs or outputs.

Base model and inference research: YuE2 authors / Multimodal Art Projection. Adapter training and Sound & Vision packaging: Atomtan Studio, using FL YuE2 tooling for DreamPop and Old School Hip-Hop, and AI Toolkit plus Mothersuperior’s v9 tokenizer and decoder companion for Industrial. This is a community adapter, not an official YuE2 release or an endorsement by the base-model authors or recording artists. Training recordings are not included in this repository.