Exploring the Design Space of Representation Learning for Audio Transformations

Paper (ISMIR 2026)

What this repository contains

Inference weights of the main RLAT model from the paper, as config.json and model.safetensors:

  • audio_encoder: the trainable adaptor that runs on precomputed latents of the frozen Stable Audio Open VAE (revision f21265c1e2710b3bd2386596943f0007f55f802e, 64 channels at 21.5 Hz for 44.1 kHz audio) and returns the audio embedding z_y (1024-d).
  • inverse_predictor: maps z_y to the transformation embedding z_t (1024-d). This is the blind variant (blind_inverse: true in config.json): the dry-audio input is zeroed, so z_t is computed from the wet audio alone.
  • proj_LT: the projection head used by the contrastive objective (128-d).

The model definitions needed to instantiate these weights are part of the training and evaluation code, which is not public yet. Until it is, this repository is weights only; there is no installable package and no loader here.

What this repository does not contain

  • Training checkpoints, the other ablation variants, or the processor encoder.
  • Evaluation data. None of the evaluation sets used in the paper are published here.

Encoding waveforms requires the Stable Audio Open VAE; accept its access terms separately.

Downloads last month
77
Safetensors
Model size
85.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for shlee-97/rlat