File size: 2,390 Bytes
c0f8960
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
---
license: apache-2.0
language: en
library_name: onnx
tags:
- onnx
- wav2vec2
- ctc
- forced-alignment
- quantized
- parseh
base_model: jonatasgrosman/wav2vec2-large-xlsr-53-english
base_model_relation: quantized
datasets:
- common_voice
---

# Parseh English forced-alignment network

This CTC network aligns known English caption text to 16 kHz mono audio in Parseh. It is not a general-purpose speech-recognition model. Parseh runs it locally with NumPy and ONNX Runtime, then uses CTC Viterbi alignment for word spans or character spans.

## Provenance and conversion

Derived from [jonatasgrosman/wav2vec2-large-xlsr-53-english](https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53-english) at `569a6236e92bd5f7652a0420bfe9bb94c5664080`, based on `facebook/wav2vec2-large-xlsr-53`. It was converted to ONNX opset 17 and dynamically quantized to int8 (`QUInt8`, per-channel `MatMul` weights); no retraining was performed. Exact source and output hashes are in `meta.json` and `SHA256SUMS`.

## Verification

- ONNX Runtime CPU versions: 1.23.2, 1.30.0; dynamic 0.5 s and 60 s inputs passed.
- fp32 ONNX/PyTorch log-softmax max absolute difference: `7.295608520507812e-05`; frame-argmax agreement: `None`.
- int8/fp32 frame-argmax agreement: `1.0`; generated-speech span-start delta median/p95/max: `0.0` / `0.0` / `0.0` ms.
- Generated 30-second CPU alignment cost: `5.709291` s; peak RSS: `2924044288` bytes.

`espeak-ng` generated audio was used only for pipeline monotonicity, not a real-speech accuracy claim. Unknown characters use a wildcard score to preserve a monotonic path.

## Limitations

Input must be 16 kHz mono audio and known caption text. Noise, accents, text errors, vocabulary gaps, and quantization can affect timestamps. No accuracy guarantee is made. 

## Licence and attribution

The source tag is `apache-2.0`. `LICENSE` and `NOTICE` retain source, base-model, dataset attribution, and changes. No endorsement by original authors, Meta, Mozilla, or Parseh is implied.

## Citation from the source card

```bibtex
@misc{grosman2021xlsr53-large-english,
  title={Fine-tuned {XLSR}-53 large model for speech recognition in {E}nglish},
  author={Grosman, Jonatas},
  howpublished={\url{https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53-english}},
  year={2021}
}
```

## Contact

Contact: [development@parseh.io](mailto:development@parseh.io).