--- license: apache-2.0 language: it library_name: onnx tags: - onnx - wav2vec2 - ctc - forced-alignment - quantized - parseh base_model: jonatasgrosman/wav2vec2-large-xlsr-53-italian base_model_relation: quantized datasets: - common_voice --- # Parseh Italian forced-alignment network This CTC network aligns known Italian caption text to 16 kHz mono audio in Parseh. It is not a general-purpose speech-recognition model. Parseh runs it locally with NumPy and ONNX Runtime, then uses CTC Viterbi alignment for word spans or character spans. ## Provenance and conversion Derived from [jonatasgrosman/wav2vec2-large-xlsr-53-italian](https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53-italian) at `dab04a3e00d8326052f3fb22a6ff276b822f6131`, based on `facebook/wav2vec2-large-xlsr-53`. It was converted to ONNX opset 17 and dynamically quantized to int8 (`QUInt8`, per-channel `MatMul` weights); no retraining was performed. Exact source and output hashes are in `meta.json` and `SHA256SUMS`. ## Verification - ONNX Runtime CPU versions: 1.23.2, 1.30.0; dynamic 0.5 s and 60 s inputs passed. - fp32 ONNX/PyTorch log-softmax max absolute difference: `0.0001678466796875`; frame-argmax agreement: `None`. - int8/fp32 frame-argmax agreement: `1.0`; generated-speech span-start delta median/p95/max: `0.0` / `0.0` / `0.0` ms. - Generated 30-second CPU alignment cost: `4.674252` s; peak RSS: `2874867712` bytes. `espeak-ng` generated audio was used only for pipeline monotonicity, not a real-speech accuracy claim. Unknown characters use a wildcard score to preserve a monotonic path. ## Limitations Input must be 16 kHz mono audio and known caption text. Noise, accents, text errors, vocabulary gaps, and quantization can affect timestamps. No accuracy guarantee is made. ## Licence and attribution The source tag is `apache-2.0`. `LICENSE` and `NOTICE` retain source, base-model, dataset attribution, and changes. No endorsement by original authors, Meta, Mozilla, or Parseh is implied. ## Citation from the source card ```bibtex @misc{grosman2021xlsr53-large-italian, title={Fine-tuned {XLSR}-53 large model for speech recognition in {I}talian}, author={Grosman, Jonatas}, howpublished={\url{https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53-italian}}, year={2021} } ``` ## Contact Contact: [development@parseh.io](mailto:development@parseh.io).