WebDoNhac models

Converted to ONNX for WebDoNhac (in-browser music detection; nothing uploaded).

These four audio-tagging models run inside the browser with ONNX Runtime Web. Your video and audio stay on your computer; the page only downloads the app code and these model files.

File Source model WebDoNhac preset Size (bytes) SHA-256
mn10_as.onnx EfficientAT mn10_as (AudioSet, mAP .471) fast 24,018,506 f11d0be24823294117798e3c385f53c8193e61e68637fcdd78df32313c4463bd
panns_mnv2.onnx PANNs MobileNetV2 (AudioSet, mAP .383) fast 20,714,472 1e6724b3242f247d16c8e08e3706d54ee1475220c6cc291d07132a4d284671bd
mn40_as_ext.onnx EfficientAT mn40_as_ext (AudioSet, mAP .487) accurate 278,074,159 64c934027ab91c8c4ec5f7f4450d6009c1d2fbf33343a374c63c7d4ec98ffd51
panns_cnn14.onnx PANNs Cnn14 (AudioSet, mAP .431) accurate 327,330,782 be2badf3f285866f6b389ac8a27bf705a0f31a11b6f5ffbace51d81b4d4127a9

Interface

  • Input wave: float32 [B, 160000], mono 32 kHz audio (10 s windows). Feature extraction (log-mel) is built into the graph.
  • Output probs: float32 [B, 527], AudioSet class probabilities. "Music" is index 137.

Changes from the originals

The released PyTorch checkpoints were exported to ONNX together with their feature-extraction front end so they run in the browser. The weights were not retrained or fine-tuned.

Licenses and attribution

This repository has mixed licenses, so the metadata says license: other. Each file is under the license of the model it was converted from.

EfficientAT: mn10_as.onnx, mn40_as_ext.onnx: MIT License

Florian Schmid, Khaled Koutini, Gerhard Widmer. "Efficient Large-Scale Audio Tagging via Transformer-to-CNN Knowledge Distillation." ICASSP 2023. Code and pretrained weights: https://github.com/fschmid56/EfficientAT

MIT License

Copyright (c) 2022 Florian Schmid

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

PANNs: panns_cnn14.onnx, panns_mnv2.onnx: CC BY 4.0

Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, Mark D. Plumbley. "PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition." IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28 (2020), 2880–2894. arXiv:1912.10211. Code: https://github.com/qiuqiangkong/audioset_tagging_cnn · Pretrained weights (Cnn14_mAP=0.431.pth, MobileNetV2_mAP=0.383.pth): https://zenodo.org/record/3987831

The weights are licensed under Creative Commons Attribution 4.0 International (CC BY 4.0), https://creativecommons.org/licenses/by/4.0/. Change: converted to ONNX with the feature-extraction front end included; not retrained.

AudioSet ontology and labels: CC BY-SA 4.0

The 527 output classes and their names come from Google's AudioSet ontology (https://research.google.com/audioset/), licensed under Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0), https://creativecommons.org/licenses/by-sa/4.0/.

Citation

@inproceedings{schmid2023efficient,
  title={Efficient Large-Scale Audio Tagging via Transformer-to-CNN Knowledge Distillation},
  author={Schmid, Florian and Koutini, Khaled and Widmer, Gerhard},
  booktitle={ICASSP 2023},
  year={2023}
}
@article{kong2020panns,
  title={PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition},
  author={Kong, Qiuqiang and Cao, Yin and Iqbal, Turab and Wang, Yuxuan and Wang, Wenwu and Plumbley, Mark D.},
  journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing},
  volume={28}, pages={2880--2894}, year={2020}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for minhducjp/webdonhac-models