Toward Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning
CochCNN9 checkpoints from the CCN 2026 paper Toward Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning. This repository hosts the eight models used in the data-scaling experiments: supervised word / auditory-event / multi-task controls, invariant SSL (iSSL), contrastive-equivariant SSL (CE-SSL), and AudioSet-scaled variants of the event and SSL models.
Checkpoints
All encoders share the same CochCNN9 backbone and cochleagram frontend (50 ERB filters, 50–10 kHz). Input is a mono waveform at 20 kHz, 2 s clips ((batch, 1, 40000)). Scaled means training on full unbalanced AudioSet; the others use matched Word-Speaker-Noise speech-in-noise (Feather et al., NeurIPS 2019).
| Plot name | Folder | Key | Objective | Data |
|---|---|---|---|---|
| Supervised word | cochcnn9-supervised-word |
word |
Supervised word classification (794 classes) | Matched Word-Speaker-Noise |
| Supervised auditory events | cochcnn9-supervised-auditory-events |
aud_events |
Supervised auditory-event classification (517 classes) | Matched Word-Speaker-Noise |
| Supervised multi-task | cochcnn9-supervised-multitask |
multitask |
Supervised multi-task: word (794), speaker (433), and auditory events (517) | Matched Word-Speaker-Noise |
| Scaled supervised auditory events | cochcnn9-supervised-auditory-events-scaled |
scaled_aud_events |
Supervised auditory-event classification (527 classes) | AudioSet (scaled) |
| iSSL | cochcnn9-issl |
issl |
Invariant SSL (Barlow Twins, λ=0) | Matched Word-Speaker-Noise |
| CE-SSL | cochcnn9-ce-ssl |
ce_ssl |
Contrastive-equivariant SSL (Barlow Twins, λ=0.5) | Matched Word-Speaker-Noise |
| Scaled iSSL | cochcnn9-issl-scaled |
scaled_issl |
Invariant SSL (Barlow Twins, λ=0) | AudioSet (scaled) |
| Scaled CE-SSL | cochcnn9-ce-ssl-scaled |
scaled_ce_ssl |
Contrastive-equivariant SSL (Barlow Twins, λ=0.5) | AudioSet (scaled) |
Each folder contains config.yaml (training recipe) and model.safetensors (Lightning state_dict, no optimizer). Model checkpoints should support replication of linear probes, zero-shot evaluations, and brain-model comparisons.
How to use
Install the code from GitHub (pip install -e ".[hub]"), then load a checkpoint by registry key:
from lightning_scripts.zero_shot_utils import load_single_cochdnn_model
encoder, name, layer_names = load_single_cochdnn_model(
"ce_ssl", from_hub=True, device="cpu",
)
# waveform: (batch, 1, time) at 20 kHz
activations = encoder(waveform) # dict[str, Tensor], flattened (batch, dim)
Keys: word, aud_events, multitask, scaled_aud_events, issl, ce_ssl, scaled_issl, scaled_ce_ssl. Layers such as relu4 and relufc match the paper evaluations.
To download files directly:
from huggingface_hub import hf_hub_download
config_path = hf_hub_download("imgriff/ce-ssl-ccn2026", "cochcnn9-ce-ssl/config.yaml")
weights_path = hf_hub_download("imgriff/ce-ssl-ccn2026", "cochcnn9-ce-ssl/model.safetensors")
Training
Training code is in the GitHub repository.
Shared settings across these checkpoints: CochCNN9 backbone, LARS optimizer, base learning rate 0.2. Per-model YAML files (epochs, batch size, λ, dataset class) live in each folder as config.yaml and in the GitHub repo under model_configs/. Dataset paths use COCHDNN_* environment variables.
Non-scaled models are trained on matched speech-in-noise from the Word-Speaker-Noise dataset (Feather et al., NeurIPS 2019). Scaled models are trained on full unbalanced AudioSet (Gemmeke et al., ICASSP 2017).
Intended use
Research on auditory representation learning, including linear probes (ESC-50, Speech Commands, Word-Speaker-Noise word, NSynth), zero-shot triplet evaluations, and brain–model comparisons.
Limitations
- Trained on English speech and AudioSet environmental audio
- Expects 20 kHz waveforms and the in-repo cochleagram frontend
- Task heads (classifiers, SSL projectors) are in the
state_dictfor exact reconstruction; downstream work typically uses intermediate encoder layers
Citation
@inproceedings{griffith2026humanaligned,
title={Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning},
author={Ian M. Griffith and Thomas Edward Yerxa and Josh McDermott and Jenelle Feather},
booktitle={9th Annual Conference on Cognitive Computational Neuroscience},
year={2026},
doi={10.32470/uqprhu8},
url={https://openreview.net/forum?id=qaNtSV4PGm}
}
Word-Speaker-Noise dataset (Feather et al., NeurIPS 2019):
@inproceedings{feather2019metamers,
title={Metamers of neural networks reveal divergence from human perceptual systems},
author={Feather, Jenelle and Durango, Alex and Gonzalez, Ray and McDermott, Josh},
booktitle={Advances in Neural Information Processing Systems},
year={2019}
}
AudioSet (Gemmeke et al., ICASSP 2017):
@inproceedings{gemmeke2017audioset,
title={{Audio Set}: An ontology and human-labeled dataset for audio events},
author={Gemmeke, Jort F. and Ellis, Daniel P. W. and Freedman, Dylan and Jansen, Aren and Lawrence, Wade and Moore, R. Channing and Plakal, Manoj and Ritter, Marvin},
booktitle={Proc. IEEE ICASSP},
year={2017}
}
License
MIT. See LICENSE in this repository and in the GitHub repo.