DualCounter / README.md
Hamozwa's picture
Create README.md
868df03 verified
|
Raw History Blame Contribute Delete
1.67 kB
metadata
license: mit
pipeline_tag: audio-classification

Model Card for DualCounter

Class-agnostic audio repetition counting: counting repeated sounds in an audio waveform (a clock striking, hammer blows, a heartbeat, etc.) without training on those specific sound classes.

Code: https://github.com/Hamozwa/audio-counting

Paper: [link] · Demo: https://huggingface.co/spaces/Hamozwa/dualcounter

Datasets: https://huggingface.co/datasets/Hamozwa/RepeatSynth · https://huggingface.co/datasets/Hamozwa/RepeatReal

Model Description

DualCounter runs two independent counting architectures and combines their predictions:

  • WavCounter regresses a repetition heatmap over the waveform and derives a count from it via a Schmitt trigger.
  • TSSMCounter builds a temporal self-similarity matrix from Wav2Vec2 features and classifies the count directly with a DINO vision transformer.

DualCounter runs both and falls back to TSSMCounter's prediction when the WavCounter heatmap is judged low quality.

Repository Contents

wav_counter/    checkpoint for the WavCounter model
tssm_counter/   checkpoint for the TSSMCounter model

Usage

Load and run these checkpoints using the training/inference code in the linked GitHub repo (WavCounter_tester.py, TSSMCounter_tester.py, DualCounter.py). DINO weights are pulled automatically via torch.hub (facebookresearch/dino) when running TSSMCounter.

Training Data

Trained on the synthetic datasets in RepeatSynth. Evaluated zero-shot on the real-world recordings in RepeatReal, covering mechanical, medical, and ecological domains.

Citation

[Your paper's BibTeX entry here]