Papers
arxiv:2609.30548

MuseTimbre: Zero-Shot Timbre Transfer by Controlling a Frozen Music Generator

Published on Sep 24
Authors:
,

Abstract

Instrument timbre transfer re-voices a performance using the timbre of another instrument. Extracting the target timbre from an audio reference capture more nuances than inferring it from a text prompt. Systems that read timbre from such a clip train a dedicated model for the task, which captures the timbre cleanly but stays a narrow, single-purpose system. More versatile approaches add control to a pretrained music generator, yet a reference clip entangles timbre with genre and melody, so these systems fall back on text to name the timbre. We present MuseTimbre, the first system, to our best knowledge, that transfers timbre from an audio reference to a polyphonic source through conditioning a pretrained music generator. This system employs a multi-pitch estimator to extract pitch information from the source and finetune a CLAP encoder to extract timbre information from the reference audio. Experiments show that across four datasets of real polyphonic recordings, MuseTimbre achieves pitch alignment on par with the baselines while matching the reference timbre far more closely. Results also show that the finetuned CLAP-based timbre extractor is robust to pitch variations, making it useful in timbre similarity measures.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.30548
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.30548 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.30548 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.