Model Card for DualCounter
Class-agnostic audio repetition counting: counting repeated sounds in an audio waveform (a clock striking, hammer blows, a heartbeat, etc.) without training on those specific sound classes.
Code: https://github.com/Hamozwa/audio-counting
Paper: [link] · Demo: https://huggingface.co/spaces/Hamozwa/dualcounter
Datasets: https://huggingface.co/datasets/Hamozwa/RepeatSynth · https://huggingface.co/datasets/Hamozwa/RepeatReal
Model Description
DualCounter runs two independent counting architectures and combines their predictions:
- WavCounter regresses a repetition heatmap over the waveform and derives a count from it via a Schmitt trigger.
- TSSMCounter builds a temporal self-similarity matrix from Wav2Vec2 features and classifies the count directly with a DINO vision transformer.
DualCounter runs both and falls back to TSSMCounter's prediction when the WavCounter heatmap is judged low quality.
Repository Contents
wav_counter/ checkpoint for the WavCounter model
tssm_counter/ checkpoint for the TSSMCounter model
Usage
Load and run these checkpoints using the training/inference code in the linked GitHub repo (WavCounter_tester.py, TSSMCounter_tester.py, DualCounter.py). DINO weights are pulled automatically via torch.hub (facebookresearch/dino) when running TSSMCounter.
Training Data
Trained on the synthetic datasets in RepeatSynth. Evaluated zero-shot on the real-world recordings in RepeatReal, covering mechanical, medical, and ecological domains.
Citation
[Your paper's BibTeX entry here]