Model Card for DualCounter

Class-agnostic audio repetition counting: counting repeated sounds in an audio waveform (a clock striking, hammer blows, a heartbeat, etc.) without training on those specific sound classes.

Code: https://github.com/Hamozwa/audio-counting

Paper: [link] · Demo: https://huggingface.co/spaces/Hamozwa/dualcounter

Datasets: https://huggingface.co/datasets/Hamozwa/RepeatSynth · https://huggingface.co/datasets/Hamozwa/RepeatReal

Model Description

DualCounter runs two independent counting architectures and combines their predictions:

  • WavCounter regresses a repetition heatmap over the waveform and derives a count from it via a Schmitt trigger.
  • TSSMCounter builds a temporal self-similarity matrix from Wav2Vec2 features and classifies the count directly with a DINO vision transformer.

DualCounter runs both and falls back to TSSMCounter's prediction when the WavCounter heatmap is judged low quality.

Repository Contents

wav_counter/    checkpoint for the WavCounter model
tssm_counter/   checkpoint for the TSSMCounter model

Usage

Load and run these checkpoints using the training/inference code in the linked GitHub repo (WavCounter_tester.py, TSSMCounter_tester.py, DualCounter.py). DINO weights are pulled automatically via torch.hub (facebookresearch/dino) when running TSSMCounter.

Training Data

Trained on the synthetic datasets in RepeatSynth. Evaluated zero-shot on the real-world recordings in RepeatReal, covering mechanical, medical, and ecological domains.

Citation

[Your paper's BibTeX entry here]
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Hamozwa/DualCounter 1