JAY 2 β€” Male Vocal, Lyrical Structure & Storytelling LoRA

Trained from scratch on a retro-focused dataset for ACE-Step 1.5 XL Turbo.

JAY-2 is a full retrain, not a port of JAY-1. It targets a different architecture, a different dataset, and a wider set of vocal behaviours β€” most notably two-track voice separation and lyric-aware phrasing.

Built for acestep-v15-xl-turbo (4B DiT). Not compatible with acestep-v15-turbo (2B) or any non-XL variant. The layer count and hidden size differ, so the adapter does not bind β€” this is a structural mismatch, not a version warning.


⚠️ Prefix token β€” it goes first

jay_singing_runs must be the first token in the prompt.

This is a prefix token, not an inline keyword. It modifies how the whole prompt is conditioned, so its position matters:

βœ…  jay_singing_runs, melodic British rock, raw emotional male vocalist
❌  melodic British rock, raw emotional male vocalist, jay_singing_runs

Put it anywhere but the front and the effect is substantially reduced. Do not put it in the lyrics block β€” it belongs at the very start of the caption/tags field.

You are not obliged to use it. JAY-2 works without the token β€” at 1.0 the adapter carries its effect either way. The token is an amplifier, not a switch: including it drives the behaviors described below harder than they sit at baseline.

So: 1.0 is required, the token is optional, and the token amplifies on top of that.


What it does

Effect Description
Male vocal presence +50% toward a male vocal read, with substantially less female-vocal bleed underneath.
Lyrical structure +80% to lyric structure legibility β€” the strongest single change from JAY-1. Tighter sections, more predictable phrasing.
Vocal layering & separation Clearer lead and backing vocals as distinct tracks instead of one blended voice.
Ad-libs over chorus Lead-singer ad-libs sit properly above the chorus rather than being buried inside it.
Emotional delivery The vocal reads as more emotional and human rather than synthetic.
Lyric-aware phrasing The singer appears responsive to the meaning of the words while delivering them β€” phrasing that tracks the narrative rather than reciting the syllable count.
Flow & transitions +60% handling of flow patterns and transitions, specifically movement into and out of a bridge, into a chorus, and into an ending.
Instrumentals A substantial lift on instrumental cues as well β€” strongest on retro video game music captions. See Instrumentals.

On the percentages: these are the author's desk characterizations from A/B listening, not benchmark measurements. There is no objective "percent more male" to measure. Read them as an indication of direction and rough magnitude.

Compared to JAY-1

JAY-1 JAY-2
Base acestep-v15-turbo (2B) acestep-v15-xl-turbo (4B)
Male vocal +10% +50%
Lyrical structure +20% +80%
Female bleed reduced much more reduced
Flow / transitions not addressed +60%
Voice separation no yes
Ad-lib placement no yes
Lyric-aware phrasing no yes
Prefix token guy_singing_runs jay_singing_runs
Dataset original retro-focused, trained from 0 steps

The two are not interchangeable. JAY-1 will not load into XL-turbo, and JAY-2 will not load into 2B turbo.

Migrating from JAY-1? The prefix token changed. guy_singing_runs is not the JAY-2 token and will not trigger it β€” it reads as an ordinary word and fails silently, with no error. Use jay_singing_runs.


Files

adapter_model.safetensors
adapter_config.json

Inference-only. No dataset, no preprocessed tensors, no training script β€” this is not a post-training release.

⚠️ Renaming on download

ACE-Step only loads a LoRA whose files are named exactly adapter_model.safetensors and adapter_config.json. Rename either and the adapter will not be found and the run will proceed without it, silently.

If the LoRA seems to do nothing, check the filenames first. That is the cause far more often than strength.


Usage

  1. Download both files into one folder.
  2. Confirm the exact filenames above.
  3. Load the folder as the LoRA source.
  4. Set strength, put jay_singing_runs at the front of the prompt, generate.

Strength

1.0. Not a starting point β€” the only working value.

This adapter does not fade in gracefully. Below 1.0 it does not become "a subtler version of itself," it degrades into static. That is measured across every JAY adapter, not a conservative default.

The cause is in the config. lora_alpha: 64 against r: 32 gives PEFT a scaling factor of 2.0, so effective strength is double the slider number. At slider 1.0 you get an effective 2.0, which is roughly what it takes to override the base model's prior and produce coherent audio. Turn it down and the signal never clears the base model's own output β€” what you get is the base model, badly perturbed.

Consequence: this cannot be used as a subtle blend. If you want a whisper of JAY in something, remove the prefix trigger token. Use 1.0.

If you are loading it through a pipeline whose default scale is not 1.0, set it explicitly. A silent default of 0.8 is the single most likely reason someone reports this LoRA "has noise."

Suggested settings

Setting Value
LoRA scale 1.0
Think on, then compare both ways
LM temperature default
LM CFG scale default

This is a distilled 8-step model with no CFG. The base model's own sampling controls interact with the adapter, so they matter as much as the strength does.


Instrumentals

The retro dataset does not only carry vocal material. Instrumental cues take a substantial lift from this adapter, and the largest gain is on retro video game music captions β€” chiptune-adjacent cues, square and synth leads, era-appropriate arpeggios.

If you are generating without vocals, this is not a wasted download. The same 1.0 strength applies, and the same rule holds: there is no partial-strength setting.

Example instrumental caption

jay_singing_runs, retro video game music, 16-bit chiptune, square wave lead melody,
pulse bass, duty-cycle drums, era-authentic arpeggios, minor key, adventurous and
melancholic, no vocals, instrumental

Prompting

The adapter responds to the caption and the lyrics separately. Structure the lyrics with explicit tags β€” the lyrical-structure push has nothing to organize otherwise.

Example

jay_singing_runs, melodic British rock, polished but organic. Clean articulate lead
guitar, radio-friendly mid-tempo groove, tight rhythm guitar, supportive bass, crisp
drums. Male singer, deep baritone, gravelly rough-edged texture, raw and understated,
conversational storytelling phrasing.

The 2-column chorus format

For two-track lead/backing vocals, the chorus has to be written as two columns. The chorus line goes in parentheses on the left; the lead vocal track goes in the right column on the same line.

[Chorus]
(Oh, Sundays feel like heaven, hearts are running free)     hearts running free
(The world can wait a little, we've got our own parade)     the world can wait

This is what activates the 2-track vocal behavior β€” a distinct lead and a distinct backing vocal rather than one blended voice. A [Chorus] tag with a single column of lyrics will not produce it, no matter what the caption says. The structure has to be in the lyrics.

Lyrics

Full example, 2-column choruses:

[Intro]

[Verse 1]
Sun is rising slowly, light upon the floor
Coffee's on the table, no alarms, no chores
Nothing on the schedule, nowhere we need to go

[Pre-Chorus]
And I can feel it turning
Everything is changing

[Chorus]
(Oh, Sundays feel like heaven, hearts are running free)     hearts running free
(The world can wait a little, we've got our own parade)     the world can wait

[Verse 2]
Stories in the kitchen, songs drift through the air
The clock is just a number, the hours drift away

[Bridge]
Hold that thought, I'll come back round
Some things only matter when they are found

[Chorus]
(Oh, Sundays feel like heaven, hearts are running free)     hearts running free
(The world can wait a little, we've got our own parade)     the world can wait

[Outro]

The [Pre-Chorus] and [Bridge] tags are where the flow/transition handling is most visible. Tag them explicitly to get the benefit.


Limitations

  • XL-turbo only. Will not load into acestep-v15-turbo (2B) or non-XL variants.
  • The prefix token must be first. Mid-prompt placement substantially reduces the effect.
  • The male-vocal shift is a bias, not a lock. It will not convert a prompt explicitly asking for a female vocalist.
  • Lyrical-structure gains depend on explicit [Verse] / [Chorus] / [Bridge] tagging. Untagged freeform lyrics get much less.
  • Voice separation is a lift, not a mix bus. It will not replace actual stem separation if you need clean isolated tracks for production.
  • Running below scale 1.0 produces static, not a weaker effect. No blending.
  • Retro-leaning dataset: expect a vintage-leaning tilt. If you want contemporary timbres, counter it in the prompt and expect to fight the adapter somewhat.
  • Trained on a limited corpus by one person. Expect a narrower stylistic range than the base model.

License

CC0-1.0 β€” public domain dedication. Use, modify, redistribute, commercial or not, no attribution required. No warranty, no liability.

Base model: ACE-Step/acestep-v15-xl-turbo β€” MIT, trained on legally compliant datasets. Generated music may be used commercially. This adapter is an independent derivative and is not affiliated with or endorsed by the ACE-Step authors.


Author

str8bored@S-W-O-R-D

README Author

seven@S-W-O-R-D

Downloads last month
43
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for str8bored/JAY_2

Adapter
(2)
this model