Exploring the Design Space of Representation Learning for Audio Transformations
Paper • 2608.28127 • Published
Paper (ISMIR 2026)
Inference weights of the main RLAT model from the paper, as config.json and model.safetensors:
audio_encoder: the trainable adaptor that runs on precomputed latents of the frozen Stable Audio Open VAE (revision f21265c1e2710b3bd2386596943f0007f55f802e, 64 channels at 21.5 Hz for 44.1 kHz audio) and returns the audio embedding z_y (1024-d).inverse_predictor: maps z_y to the transformation embedding z_t (1024-d). This is the blind variant (blind_inverse: true in config.json): the dry-audio input is zeroed, so z_t is computed from the wet audio alone.proj_LT: the projection head used by the contrastive objective (128-d).The model definitions needed to instantiate these weights are part of the training and evaluation code, which is not public yet. Until it is, this repository is weights only; there is no installable package and no loader here.
Encoding waveforms requires the Stable Audio Open VAE; accept its access terms separately.