aligner-zh / README.md
br1-ursino's picture
Parseh aligner zh v1: ONNX int8 from ydshieh/wav2vec2-large-xlsr-53-chinese-zh-cn-gpt@cf511f7e00de089e67e80fbedd5fbfb8e76ea067
920ed68 verified
|
Raw History Blame Contribute Delete
2.27 kB
---
license: apache-2.0
language: zh
library_name: onnx
tags:
- onnx
- wav2vec2
- ctc
- forced-alignment
- quantized
- parseh
base_model: ydshieh/wav2vec2-large-xlsr-53-chinese-zh-cn-gpt
base_model_relation: quantized
datasets:
- common_voice
---
# Parseh Chinese forced-alignment network
This CTC network aligns known Chinese caption text to 16 kHz mono audio in Parseh. It is not a general-purpose speech-recognition model. Parseh runs it locally with NumPy and ONNX Runtime, then uses CTC Viterbi alignment for word spans or character spans.
## Provenance and conversion
Derived from [ydshieh/wav2vec2-large-xlsr-53-chinese-zh-cn-gpt](https://huggingface.co/ydshieh/wav2vec2-large-xlsr-53-chinese-zh-cn-gpt) at `cf511f7e00de089e67e80fbedd5fbfb8e76ea067`, based on `facebook/wav2vec2-large-xlsr-53`. It was converted to ONNX opset 17 and dynamically quantized to int8 (`QUInt8`, per-channel `MatMul` weights); no retraining was performed. Exact source and output hashes are in `meta.json` and `SHA256SUMS`.
## Verification
- ONNX Runtime CPU versions: 1.23.2, 1.30.0; dynamic 0.5 s and 60 s inputs passed.
- fp32 ONNX/PyTorch log-softmax max absolute difference: `0.000202178955078125`; frame-argmax agreement: `1.0`.
- int8/fp32 frame-argmax agreement: `1.0`; generated-speech span-start delta median/p95/max: `0.0` / `0.0` / `0.0` ms.
- Generated 30-second CPU alignment cost: `5.006784` s; peak RSS: `3979509760` bytes.
`espeak-ng` generated audio was used only for pipeline monotonicity, not a real-speech accuracy claim. Unknown characters use a wildcard score to preserve a monotonic path.
## Limitations
Input must be 16 kHz mono audio and known caption text. Noise, accents, text errors, vocabulary gaps, and quantization can affect timestamps. No accuracy guarantee is made. This language returns character spans because captions are not space-delimited.
## Licence and attribution
The source tag is `apache-2.0`. `LICENSE` and `NOTICE` retain source, base-model, dataset attribution, and changes. No endorsement by original authors, Meta, Mozilla, or Parseh is implied.
## Citation from the source card
No separate citation block was provided by the source card.
## Contact
Contact: [development@parseh.io](mailto:development@parseh.io).