Papers
arxiv:2608.04378

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

Published on Aug 5
ยท Submitted by
Scott Hawley
on Aug 6
Authors:

Abstract

Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model'' for symbolic music: a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives (pitch- and time-shift equivariance, masked embedding prediction, and a distributional regularizer), using no labels and no music-theory vocabulary. Probing the frozen embeddings shows that the level at which a musical property becomes decodable tracks its musical time scale: phrase boundaries are read off the coarsest levels, note density and harmonic detail off the finest. Temporal and phrase structure emerge from the self-supervised objectives alone, while harmonic content must be asked for; a small chord-supervision head raises joint chord recovery from .18 to .54, and key detection, which is never supervised, from .16 to .70. Following the Representation AutoEncoder paradigm, a conditional flow-matching model stands in for a trained decoder, flowing in pixel space from PCA-reduced conditioning: it reproduces a target window at pixel F1 0.996, and the same per-level conditioning dropout that controls how far variations stray also enables graphical prompting for masked inpainting with no inpainting-specific sampler. The pipeline runs on CPU producing a suggestion in 2.8 s, or 0.6 s on Apple MPS, which we demonstrate in a live interactive demo. In concert with an LLM-based brain, these capabilities supply the core of a collaborative music creation agent in service of, rather than in place of, human agency.

Community

Paper author Paper submitter

Excited to share a new Preprint & LIVE DEMO! This is >2 years of my life building fast, lightweight CPU-friendly models to facilitate songwriters' iterative workflows, with Representations usable for Understanding and Generation, trained in a Self-Supervised way. First, the demo:
https://drscotthawley-midi-rae-jepa-son.hf.space

This turned out so fun that I had to pause paper-writing to revamp it into a user-friendly app to send to friends! You draw an inpainting mask that controls where spatial dropout occurs, and the "EQ"-looking sliders scale the abstraction level.

The preprint is here: https://arxiv.org/abs/2608.04378 The first 6 pages + refs are submitted to the NeurIPS Creative AI Track ("single-blind, preprints ok" ๐Ÿ‘) Preprint adds 10 pages of Supplemental Materials: ablation studies, hyperparameter surveys, etc.

You may have seen another preprint from me a couple weeks ago (https://arxiv.org/abs/2607.14537), which was the project state mid-April 2026 submitted to ISMIR (double-blind, gag order) but the new preprint is the more mature one (despite being the "-SON"!)

The story started with the "graphical prompts" of "Pictures of MIDI" ca. Jan. 2024 (https://picturesofmidi.github.io/PicturesOfMIDI/) but that was big & slow. The quest to streamline it led me to Flow Models (fast sampling), Representation AutoEncoders (reusable rep's), and LeJEPA (to resist collapse).

You could wade through my messy "midi-rae" work repo but maybe best to wait til I push a clean distro-repo in the coming weeks. Stay tuned via the links on the project website: https://drscotthawley.github.io/midi-rae-jepa-son/ Weights are closed for now; I'll probably open them closer to NeurIPS time.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.04378
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.04378 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.04378 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.04378 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.