Toward Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning

CochCNN9 checkpoints from the CCN 2026 paper Toward Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning. This repository hosts the eight models used in the data-scaling experiments: supervised word / auditory-event / multi-task controls, invariant SSL (iSSL), contrastive-equivariant SSL (CE-SSL), and AudioSet-scaled variants of the event and SSL models.

Checkpoints

All encoders share the same CochCNN9 backbone and cochleagram frontend (50 ERB filters, 50–10 kHz). Input is a mono waveform at 20 kHz, 2 s clips ((batch, 1, 40000)). Scaled means training on full unbalanced AudioSet; the others use matched Word-Speaker-Noise speech-in-noise (Feather et al., NeurIPS 2019).

Plot name Folder Key Objective Data
Supervised word cochcnn9-supervised-word word Supervised word classification (794 classes) Matched Word-Speaker-Noise
Supervised auditory events cochcnn9-supervised-auditory-events aud_events Supervised auditory-event classification (517 classes) Matched Word-Speaker-Noise
Supervised multi-task cochcnn9-supervised-multitask multitask Supervised multi-task: word (794), speaker (433), and auditory events (517) Matched Word-Speaker-Noise
Scaled supervised auditory events cochcnn9-supervised-auditory-events-scaled scaled_aud_events Supervised auditory-event classification (527 classes) AudioSet (scaled)
iSSL cochcnn9-issl issl Invariant SSL (Barlow Twins, λ=0) Matched Word-Speaker-Noise
CE-SSL cochcnn9-ce-ssl ce_ssl Contrastive-equivariant SSL (Barlow Twins, λ=0.5) Matched Word-Speaker-Noise
Scaled iSSL cochcnn9-issl-scaled scaled_issl Invariant SSL (Barlow Twins, λ=0) AudioSet (scaled)
Scaled CE-SSL cochcnn9-ce-ssl-scaled scaled_ce_ssl Contrastive-equivariant SSL (Barlow Twins, λ=0.5) AudioSet (scaled)

Each folder contains config.yaml (training recipe) and model.safetensors (Lightning state_dict, no optimizer). Model checkpoints should support replication of linear probes, zero-shot evaluations, and brain-model comparisons.

How to use

Install the code from GitHub (pip install -e ".[hub]"), then load a checkpoint by registry key:

from lightning_scripts.zero_shot_utils import load_single_cochdnn_model

encoder, name, layer_names = load_single_cochdnn_model(
    "ce_ssl", from_hub=True, device="cpu",
)
# waveform: (batch, 1, time) at 20 kHz
activations = encoder(waveform)  # dict[str, Tensor], flattened (batch, dim)

Keys: word, aud_events, multitask, scaled_aud_events, issl, ce_ssl, scaled_issl, scaled_ce_ssl. Layers such as relu4 and relufc match the paper evaluations.

To download files directly:

from huggingface_hub import hf_hub_download

config_path = hf_hub_download("imgriff/ce-ssl-ccn2026", "cochcnn9-ce-ssl/config.yaml")
weights_path = hf_hub_download("imgriff/ce-ssl-ccn2026", "cochcnn9-ce-ssl/model.safetensors")

Training

Training code is in the GitHub repository.

Shared settings across these checkpoints: CochCNN9 backbone, LARS optimizer, base learning rate 0.2. Per-model YAML files (epochs, batch size, λ, dataset class) live in each folder as config.yaml and in the GitHub repo under model_configs/. Dataset paths use COCHDNN_* environment variables.

Non-scaled models are trained on matched speech-in-noise from the Word-Speaker-Noise dataset (Feather et al., NeurIPS 2019). Scaled models are trained on full unbalanced AudioSet (Gemmeke et al., ICASSP 2017).

Intended use

Research on auditory representation learning, including linear probes (ESC-50, Speech Commands, Word-Speaker-Noise word, NSynth), zero-shot triplet evaluations, and brain–model comparisons.

Limitations

  • Trained on English speech and AudioSet environmental audio
  • Expects 20 kHz waveforms and the in-repo cochleagram frontend
  • Task heads (classifiers, SSL projectors) are in the state_dict for exact reconstruction; downstream work typically uses intermediate encoder layers

Citation

@inproceedings{griffith2026humanaligned,
  title={Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning},
  author={Ian M. Griffith and Thomas Edward Yerxa and Josh McDermott and Jenelle Feather},
  booktitle={9th Annual Conference on Cognitive Computational Neuroscience},
  year={2026},
  doi={10.32470/uqprhu8},
  url={https://openreview.net/forum?id=qaNtSV4PGm}
}

Word-Speaker-Noise dataset (Feather et al., NeurIPS 2019):

@inproceedings{feather2019metamers,
  title={Metamers of neural networks reveal divergence from human perceptual systems},
  author={Feather, Jenelle and Durango, Alex and Gonzalez, Ray and McDermott, Josh},
  booktitle={Advances in Neural Information Processing Systems},
  year={2019}
}

AudioSet (Gemmeke et al., ICASSP 2017):

@inproceedings{gemmeke2017audioset,
  title={{Audio Set}: An ontology and human-labeled dataset for audio events},
  author={Gemmeke, Jort F. and Ellis, Daniel P. W. and Freedman, Dylan and Jansen, Aren and Lawrence, Wade and Moore, R. Channing and Plakal, Manoj and Ritter, Marvin},
  booktitle={Proc. IEEE ICASSP},
  year={2017}
}

License

MIT. See LICENSE in this repository and in the GitHub repo.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support