PEACE: effect code and audio embeddings

PEACE embeds Faust audio effect code and effected ("wet") audio in one space, so an audio query retrieves the effect chain that produced it, and a program retrieves audio processed by it. It accompanies PEACE: Joint Embeddings of DSP Effects Code and Audio, by David Braun and Adam Finkelstein (ISMIR 2026).

Inference code · Interactive map

Each subfolder holds a complete model: the AFx-Rep audio encoder (a mid/side CNN14 trained from scratch), a code encoder, and both modalities' SLAP projector and predictor heads. Retrieval uses cosine similarity of the predictions q.

Subfolder Code encoder Audio→Code R@1 R@10 Code→Audio R@1 R@10
boxgraph BoxGraph GNN over the Faust Box API graph 48.5% 77.7% 50.8% 77.5%
t5 T5 v1.1 small, fine-tuned, over Faust source text 50.6% 78.1% 48.4% 76.3%

Scores are the paper's, on the 4,096-pair test set from 6-second excerpts, with exactly one correct item per query. BoxGraph generalizes less well than T5 to chains longer than the one to three effects seen in training, but its embeddings do not depend on how the program is written.

Usage

pip install "peace @ git+https://github.com/DBraun/PEACE"
import numpy as np
from audiotree import AudioTree
from peace import PEACE

model = PEACE.from_pretrained("davidbraun/peace", subfolder="boxgraph")
programs = [
    model.library.program([("reverb_freeverb", [0.8, 0.5, 0.3, 0.4])]),
    model.library.program(
        [("distortion", [1, 3, 0.6, 0.5, 0.2, 0.7, 1.0]), ("delay", [0, 0.3, 0.5, 0.2, 0.5, 0.2, 0.5, 0.5, 0.4])]
    ),
]
gallery = np.asarray(model.encode_code(programs))  # [2, 768]
query = np.asarray(model.encode_audio(AudioTree.from_file("wet.wav", duration=6)))
ranking = np.argsort(-(query @ gallery.T), axis=-1)

encode_code(programs, mask_parameters=True) hides every parameter value, so embeddings describe only which effects are used and in what order. The package also ships the 21 effects as Faust source and renders programs with DawDreamer.

Training

  • Data: 200K pairs of Faust code and 10-second, 48 kHz wet audio. Dry audio comes from 23 public datasets across bass, drums, guitar, instruments, multitrack stems, music, piano, singing, and speech. Each program chains one to three of 21 Faust effects with randomly sampled parameters.
  • Audio: mid/side log-mel spectrograms (2048-point FFT, hop 1024, 128 bands, 20–20,000 Hz) of 6-second excerpts, after normalizing clips to −18 LUFS.
  • Objective: SLAP (Guinot et al., 2025), a BYOL-style objective without negative samples, for 50,000 steps at effective batch size 256.
  • Selection: the checkpoint with the best validation mean reciprocal rank.

PEACE's audio encoder shares AFx-Rep's architecture but not its exact input features: its training log-mel frames are zero-padded at the ends of each clip, where AFx-Rep pads by reflection. The ISMIR proceedings computed PEACE's features with reflection padding in the RIR and chain retrieval evaluations; the arXiv version recomputes them on the training features, so those numbers differ slightly between the two versions.

Limitations

  • The model knows only the 21 effects in its library, applied in series. Programs with other effects or with parallel or feedback routing are rejected.
  • Effects with near-identical outputs, such as commuting linear effects, cannot be told apart from audio alone, so some retrieval "errors" are equivalent.
  • Retrieval has not been evaluated in a user study.

Licensing

The inference code is MIT-licensed. The weights are released under CC BY-NC 4.0: non-commercial use with attribution. The t5 weights are fine-tuned from google/t5-v1_1-small, Apache License 2.0, and its tokenizer.json is derived from that model's SentencePiece vocabulary. The Faust effect sources in the package keep their own terms, listed in its NOTICE file; compressor.dsp is GPL-3.0-or-later.

Citation

@inproceedings{braun2026peace,
  title     = {{PEACE}: Joint Embeddings of {DSP} Effects Code and Audio},
  author    = {Braun, David and Finkelstein, Adam},
  booktitle = {Proc. of the 27th Int. Society for Music Information Retrieval Conference (ISMIR)},
  year      = {2026},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support