DualCounter / README.md
Hamozwa's picture
Create README.md
868df03 verified
|
Raw History Blame Contribute Delete
1.67 kB
---
license: mit
pipeline_tag: audio-classification
---
# Model Card for DualCounter
Class-agnostic audio repetition counting: counting repeated sounds in an audio waveform (a clock striking, hammer blows, a heartbeat, etc.) without training on those specific sound classes.
Code: https://github.com/Hamozwa/audio-counting
Paper: [link] 路 Demo: https://huggingface.co/spaces/Hamozwa/dualcounter
Datasets: https://huggingface.co/datasets/Hamozwa/RepeatSynth 路 https://huggingface.co/datasets/Hamozwa/RepeatReal
## Model Description
DualCounter runs two independent counting architectures and combines their predictions:
- **WavCounter** regresses a repetition heatmap over the waveform and derives a count from it via a Schmitt trigger.
- **TSSMCounter** builds a temporal self-similarity matrix from Wav2Vec2 features and classifies the count directly with a DINO vision transformer.
DualCounter runs both and falls back to TSSMCounter's prediction when the WavCounter heatmap is judged low quality.
## Repository Contents
```
wav_counter/ checkpoint for the WavCounter model
tssm_counter/ checkpoint for the TSSMCounter model
```
## Usage
Load and run these checkpoints using the training/inference code in the linked GitHub repo (`WavCounter_tester.py`, `TSSMCounter_tester.py`, `DualCounter.py`). DINO weights are pulled automatically via `torch.hub` (`facebookresearch/dino`) when running TSSMCounter.
## Training Data
Trained on the synthetic datasets in RepeatSynth. Evaluated zero-shot on the real-world recordings in RepeatReal, covering mechanical, medical, and ecological domains.
## Citation
```bibtex
[Your paper's BibTeX entry here]
```