WebDoNhac models
Converted to ONNX for WebDoNhac (in-browser music detection; nothing uploaded).
These four audio-tagging models run inside the browser with ONNX Runtime Web. Your video and audio stay on your computer; the page only downloads the app code and these model files.
| File | Source model | WebDoNhac preset | Size (bytes) | SHA-256 |
|---|---|---|---|---|
mn10_as.onnx |
EfficientAT mn10_as (AudioSet, mAP .471) | fast | 24,018,506 | f11d0be24823294117798e3c385f53c8193e61e68637fcdd78df32313c4463bd |
panns_mnv2.onnx |
PANNs MobileNetV2 (AudioSet, mAP .383) | fast | 20,714,472 | 1e6724b3242f247d16c8e08e3706d54ee1475220c6cc291d07132a4d284671bd |
mn40_as_ext.onnx |
EfficientAT mn40_as_ext (AudioSet, mAP .487) | accurate | 278,074,159 | 64c934027ab91c8c4ec5f7f4450d6009c1d2fbf33343a374c63c7d4ec98ffd51 |
panns_cnn14.onnx |
PANNs Cnn14 (AudioSet, mAP .431) | accurate | 327,330,782 | be2badf3f285866f6b389ac8a27bf705a0f31a11b6f5ffbace51d81b4d4127a9 |
Interface
- Input
wave: float32[B, 160000], mono 32 kHz audio (10 s windows). Feature extraction (log-mel) is built into the graph. - Output
probs: float32[B, 527], AudioSet class probabilities. "Music" is index 137.
Changes from the originals
The released PyTorch checkpoints were exported to ONNX together with their feature-extraction front end so they run in the browser. The weights were not retrained or fine-tuned.
Licenses and attribution
This repository has mixed licenses, so the metadata says license: other. Each file is under the license of the model it was converted from.
EfficientAT: mn10_as.onnx, mn40_as_ext.onnx: MIT License
Florian Schmid, Khaled Koutini, Gerhard Widmer. "Efficient Large-Scale Audio Tagging via Transformer-to-CNN Knowledge Distillation." ICASSP 2023. Code and pretrained weights: https://github.com/fschmid56/EfficientAT
MIT License
Copyright (c) 2022 Florian Schmid
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
PANNs: panns_cnn14.onnx, panns_mnv2.onnx: CC BY 4.0
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, Mark D. Plumbley. "PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition." IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28 (2020), 2880–2894. arXiv:1912.10211.
Code: https://github.com/qiuqiangkong/audioset_tagging_cnn · Pretrained weights (Cnn14_mAP=0.431.pth, MobileNetV2_mAP=0.383.pth): https://zenodo.org/record/3987831
The weights are licensed under Creative Commons Attribution 4.0 International (CC BY 4.0), https://creativecommons.org/licenses/by/4.0/. Change: converted to ONNX with the feature-extraction front end included; not retrained.
AudioSet ontology and labels: CC BY-SA 4.0
The 527 output classes and their names come from Google's AudioSet ontology (https://research.google.com/audioset/), licensed under Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0), https://creativecommons.org/licenses/by-sa/4.0/.
Citation
@inproceedings{schmid2023efficient,
title={Efficient Large-Scale Audio Tagging via Transformer-to-CNN Knowledge Distillation},
author={Schmid, Florian and Koutini, Khaled and Widmer, Gerhard},
booktitle={ICASSP 2023},
year={2023}
}
@article{kong2020panns,
title={PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition},
author={Kong, Qiuqiang and Cao, Yin and Iqbal, Turab and Wang, Yuxuan and Wang, Wenwu and Plumbley, Mark D.},
journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing},
volume={28}, pages={2880--2894}, year={2020}
}