SONAR is an evaluation toolkit for multilingual ASR that goes beyond WER/CER. It combines semantic similarity, the Poseidon Score, and analysis across dialect, demographic, and metadata-based failure modes. ๐
Our goal is to make it easier for everyone to understand why an ASR model fails, not just how often. ๐ You can plug in your own models + audio, extend it to new languages and datasets, or contribute directly. ๐ ๏ธ
MIT licensed. Would love feedback from the HF community! ๐ค