Ensomi R2
Ensomi R2 turns a song into an osu!mania 4K chart. Give it an audio file and a
star rating from 2 to 6, and a few seconds later it writes a complete .osz
that opens in osu!. It is the first Ensomi release that works from audio alone.
Yellow notes are taps and cyan notes are long notes; each panel reads bottom to
top, then continues in the panel to its right. The
gallery shows
four songs at seed 0 with their generated .osu files. They illustrate
behaviour; they are not a quality benchmark.
The full report, with the project's aims, the decisions behind R2 and the
evidence for every claim below, is
docs/research/r2_system.md
in the code repository.
What it is for
Ensomi wants listening to music, playing it and creating for it to become one activity: charts that follow music already playing around you, candidate choreography for a mapper's selected section, and practice charts built around a movement a player wants to train. Those need charts that are playable at the requested difficulty, follow the music's rhythm, keep one character through the whole song, and respond to requests.
R2 is the first step: a whole song, generated while you wait. It does not yet play along in real time or regenerate one section of an existing chart.
How it works
audio โ> timing system โ> rows model โ> corpus prior โ> R2 with controls โ> .osz
beat grid head rows chart-level lanes, taps,
targets long notes
- Timing system. The BeatThis beat tracker finds the beats; each tempo segment is then refitted to the song's own note onsets, which sit much closer to where mappers put their beats. No trained weights of its own.
- Rows model (
rows/model.pt, 2.1M parameters). Beat by beat, it chooses where notes start from what it hears at each position in the beat, and its copy heads repeat earlier figures where the music sounds the same. - Corpus prior (
prior.json). From the 64 human charts whose head rows look most like these, it picks the chart's character: long-note level, chord density, difficulty. This stands in for the mapper's choice for the whole chart. - R2 (
r2/model.pt, 2.4M parameters). It arranges the lanes, taps and long notes on every head row; one decision per row places its new notes and the releases of the long notes it closes. A restoring force holds the chart's character to the end, playability limits measured on human charts (calibration.json) keep it playable, and the requests tilt its choices. R2 does not hear the audio.
What you can ask for
| Request | Status |
|---|---|
| Star rating for the whole song | works in blind screens |
| Jumpstream and chordstream sections | work in blind screens |
| Stream and jack sections | reach their target in measurement; not screened on their own |
| Chordjack sections | not offered as working |
| Long-note amount (`--ln tap | half |
| Section difficulty in stars | not yet screened |
What has been checked
Every chart passes legality and replay checks and has whole-millisecond times. Quality was judged by one experienced player in blind pairs, one generation seed per side, so a single pair is weak evidence:
- the new rows model beat the earlier head generator 10 to 3, one undecided;
- whole-song difficulty was called the right way in both pairs;
- the human chart was chosen over R2's in all three reference pairs.
On a holdout of 519 human charts, 85 % have 90 % of their heads within 20 ms of the timing system's grid. The four example charts regenerate note for note from this repository's files.
Known limits: long notes are the weakest part; easier sections land short of their target; long-song stability rests on a decode-time rule; every verdict is one listener's.
Quick start
Checked on macOS arm64 (Apple M5, Python 3.10.20, torch 2.11.0, CPU, one
thread). Linux and CUDA (--extra cuda) were not run.
git clone https://github.com/ensomi-labs/ensomi-model.git
cd ensomi-model
git checkout --detach r2-1.0
uv sync --python 3.10 --extra mps
hf download sed-i/ensomi-r2 --local-dir artifacts/hf-r2
uv run --python 3.10 --extra mps python -m ensomi_model.system.make \
--model-dir artifacts/hf-r2 --audio /path/to/song.mp3 --star 3.5 \
--title "Song title" --artist "Artist" --seed 0 --out artifacts/my-chart.osz
A 2-3.3 minute song takes 8-12 s. The first run also downloads BeatThis's
final0 checkpoint (77 MB) from its authors. Decoding audio other than WAV
needs ffmpeg. --help lists every request, for example
--stream 60000,90000,0.7,jumpstream.
Files
| File | Bytes | SHA-256 |
|---|---|---|
rows/model.pt |
9,622,005 | 46b6ca37870120fda4d5de632949921fc26e020cf80abe44b4ba7f6f9557dda4 |
r2/model.pt |
9,655,631 | 8dedefe4b11d238ad0b3186e6dfefe31f52cb333091aafe2e8efda8ad2ee3c6f |
prior.json |
1,958,310 | be5741d520120aaeb7ecf9126c2c12401015095b107b025584802f8e715590af |
calibration.json |
443,193 | d39cd16fda451163d00d70cd6a913eb4bf64326ecd577feab08c25d8c22455ab |
scopes.parquet |
2,447,837 | 61f8648523cddb86e7316812b46896410f12b1a6ae179dcbf082c8b8615e3134 |
system.json |
209 | 333d1c59a5ca199e07369b980992c2bb9683851246807426f1152260502e9718 |
LICENSE |
20,850 | e66c269d4819aaab34b49ef5220c4ddab6756f21bb5180761a4eb8561f2b7bbd |
SHA256SUMS lists every file, including examples/ (four generated charts
with their receipts) and previews/. r2/model.pt holds R2's weights and
model configuration only; the example receipts record the hash of the training
checkpoint, which also held optimizer state. The .pt files are PyTorch
pickles, loaded by the code repository.
onnx/ holds the same networks as ONNX step graphs (float32, opset 17) for
in-browser inference by the TypeScript runtime in
ensomi-web: BeatThis final0, the
rows encoder and step, and R2 predict, release and commit, with runtime.json
(graph interfaces, carried state and decode constants). Replaying the same
random draws, they reproduce the PyTorch charts of the four examples note for
note.
| Component | Trained from | Code tag |
|---|---|---|
| Timing system | no training | timing-1.0 |
| Rows model | commit cf753aa, 5,000 steps |
rows-1.0 |
| R2 | commit 3954031, 64M decisions; checkpoint at 56,000,574 |
r2-1.0 |
R2 and the rows model were trained on osu!'s ranked and loved 4K charts from 2 to 6 stars (11,368 charts of 4,167 songs). Neither those charts nor their audio are distributed here. The prior, calibration and scope table hold numbers derived from them, without titles, artists, mappers or file paths.
Licence
Everything in this repository is licensed
CC BY-NC-SA 4.0: you may
use, share and adapt it with credit, not for commercial purposes, and
adaptations must carry the same licence, except onnx/beatthis.onnx, an
ONNX export of BeatThis final0 (CPJKU), which keeps its MIT licence
(onnx/LICENSE-beat_this). The Python system downloads BeatThis from its
authors instead. The model code is AGPL-3.0-only; the browser runtime is
Apache-2.0.
