Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation
Abstract
Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model'' for symbolic music: a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives (pitch- and time-shift equivariance, masked embedding prediction, and a distributional regularizer), using no labels and no music-theory vocabulary. Probing the frozen embeddings shows that the level at which a musical property becomes decodable tracks its musical time scale: phrase boundaries are read off the coarsest levels, note density and harmonic detail off the finest. Temporal and phrase structure emerge from the self-supervised objectives alone, while harmonic content must be asked for; a small chord-supervision head raises joint chord recovery from .18 to .54, and key detection, which is never supervised, from .16 to .70. Following the Representation AutoEncoder paradigm, a conditional flow-matching model stands in for a trained decoder, flowing in pixel space from PCA-reduced conditioning: it reproduces a target window at pixel F1 0.996, and the same per-level conditioning dropout that controls how far variations stray also enables graphical prompting for masked inpainting with no inpainting-specific sampler. The pipeline runs on CPU producing a suggestion in 2.8 s, or 0.6 s on Apple MPS, which we demonstrate in a live interactive demo. In concert with an LLM-based brain, these capabilities supply the core of a collaborative music creation agent in service of, rather than in place of, human agency.
Community
Excited to share a new Preprint & LIVE DEMO! This is >2 years of my life building fast, lightweight CPU-friendly models to facilitate songwriters' iterative workflows, with Representations usable for Understanding and Generation, trained in a Self-Supervised way. First, the demo:
https://drscotthawley-midi-rae-jepa-son.hf.space
This turned out so fun that I had to pause paper-writing to revamp it into a user-friendly app to send to friends! You draw an inpainting mask that controls where spatial dropout occurs, and the "EQ"-looking sliders scale the abstraction level.
The preprint is here: https://arxiv.org/abs/2608.04378 The first 6 pages + refs are submitted to the NeurIPS Creative AI Track ("single-blind, preprints ok" ๐) Preprint adds 10 pages of Supplemental Materials: ablation studies, hyperparameter surveys, etc.
You may have seen another preprint from me a couple weeks ago (https://arxiv.org/abs/2607.14537), which was the project state mid-April 2026 submitted to ISMIR (double-blind, gag order) but the new preprint is the more mature one (despite being the "-SON"!)
The story started with the "graphical prompts" of "Pictures of MIDI" ca. Jan. 2024 (https://picturesofmidi.github.io/PicturesOfMIDI/) but that was big & slow. The quest to streamline it led me to Flow Models (fast sampling), Representation AutoEncoders (reusable rep's), and LeJEPA (to resist collapse).
You could wade through my messy "midi-rae" work repo but maybe best to wait til I push a clean distro-repo in the coming weeks. Stay tuned via the links on the project website: https://drscotthawley.github.io/midi-rae-jepa-son/ Weights are closed for now; I'll probably open them closer to NeurIPS time.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music (2026)
- ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music (2026)
- Music-JEPA: Learning a World Model of Sound from Action (2026)
- Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping (2026)
- Frequency-Aware Self-Supervised Music Representation Learning (2026)
- BeatEdit: Symbolic Music Generation as Explicit Editing (2026)
- BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.04378 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper