--- license: mit pipeline_tag: audio-classification --- # Model Card for DualCounter Class-agnostic audio repetition counting: counting repeated sounds in an audio waveform (a clock striking, hammer blows, a heartbeat, etc.) without training on those specific sound classes. Code: https://github.com/Hamozwa/audio-counting Paper: [link] · Demo: https://huggingface.co/spaces/Hamozwa/dualcounter Datasets: https://huggingface.co/datasets/Hamozwa/RepeatSynth · https://huggingface.co/datasets/Hamozwa/RepeatReal ## Model Description DualCounter runs two independent counting architectures and combines their predictions: - **WavCounter** regresses a repetition heatmap over the waveform and derives a count from it via a Schmitt trigger. - **TSSMCounter** builds a temporal self-similarity matrix from Wav2Vec2 features and classifies the count directly with a DINO vision transformer. DualCounter runs both and falls back to TSSMCounter's prediction when the WavCounter heatmap is judged low quality. ## Repository Contents ``` wav_counter/ checkpoint for the WavCounter model tssm_counter/ checkpoint for the TSSMCounter model ``` ## Usage Load and run these checkpoints using the training/inference code in the linked GitHub repo (`WavCounter_tester.py`, `TSSMCounter_tester.py`, `DualCounter.py`). DINO weights are pulled automatically via `torch.hub` (`facebookresearch/dino`) when running TSSMCounter. ## Training Data Trained on the synthetic datasets in RepeatSynth. Evaluated zero-shot on the real-world recordings in RepeatReal, covering mechanical, medical, and ecological domains. ## Citation ```bibtex [Your paper's BibTeX entry here] ```