Parseh aligner de v1: ONNX int8 from jonatasgrosman/wav2vec2-large-xlsr-53-german@4b8a02957378d0f2da2ef74091156b032c485a89
b7d6adf verified |
Download README.md from parseh/aligner-de: direct link, hf CLI and curl.
- Browser
- Download file 2.38 kB
-
https://huggingface.co/parseh/aligner-de/resolve/main/README.md
- Command line
-
hf download hf://parseh/aligner-de/README.md
-
curl -L -o README.md https://huggingface.co/parseh/aligner-de/resolve/main/README.md
2.38 kB
| license: apache-2.0 | |
| language: de | |
| library_name: onnx | |
| tags: | |
| - onnx | |
| - wav2vec2 | |
| - ctc | |
| - forced-alignment | |
| - quantized | |
| - parseh | |
| base_model: jonatasgrosman/wav2vec2-large-xlsr-53-german | |
| base_model_relation: quantized | |
| datasets: | |
| - common_voice | |
| # Parseh German forced-alignment network | |
| This CTC network aligns known German caption text to 16 kHz mono audio in Parseh. It is not a general-purpose speech-recognition model. Parseh runs it locally with NumPy and ONNX Runtime, then uses CTC Viterbi alignment for word spans or character spans. | |
| ## Provenance and conversion | |
| Derived from [jonatasgrosman/wav2vec2-large-xlsr-53-german](https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53-german) at `4b8a02957378d0f2da2ef74091156b032c485a89`, based on `facebook/wav2vec2-large-xlsr-53`. It was converted to ONNX opset 17 and dynamically quantized to int8 (`QUInt8`, per-channel `MatMul` weights); no retraining was performed. Exact source and output hashes are in `meta.json` and `SHA256SUMS`. | |
| ## Verification | |
| - ONNX Runtime CPU versions: 1.23.2, 1.30.0; dynamic 0.5 s and 60 s inputs passed. | |
| - fp32 ONNX/PyTorch log-softmax max absolute difference: `0.000301361083984375`; frame-argmax agreement: `None`. | |
| - int8/fp32 frame-argmax agreement: `1.0`; generated-speech span-start delta median/p95/max: `0.0` / `0.0` / `0.0` ms. | |
| - Generated 30-second CPU alignment cost: `5.648164` s; peak RSS: `2879397888` bytes. | |
| `espeak-ng` generated audio was used only for pipeline monotonicity, not a real-speech accuracy claim. Unknown characters use a wildcard score to preserve a monotonic path. | |
| ## Limitations | |
| Input must be 16 kHz mono audio and known caption text. Noise, accents, text errors, vocabulary gaps, and quantization can affect timestamps. No accuracy guarantee is made. | |
| ## Licence and attribution | |
| The source tag is `apache-2.0`. `LICENSE` and `NOTICE` retain source, base-model, dataset attribution, and changes. No endorsement by original authors, Meta, Mozilla, or Parseh is implied. | |
| ## Citation from the source card | |
| ```bibtex | |
| @misc{grosman2021xlsr53-large-german, | |
| title={Fine-tuned {XLSR}-53 large model for speech recognition in {G}erman}, | |
| author={Grosman, Jonatas}, | |
| howpublished={\url{https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53-german}}, | |
| year={2021} | |
| } | |
| ``` | |
| ## Contact | |
| Contact: [development@parseh.io](mailto:development@parseh.io). | |