exp38 character 6-gram LMs
Character 6-gram Kneser-Ney language models (own implementation, charlm.py of the TrOCR_Hebrew project; arrays in .npz, no pickle) used for
CTC prefix beam search fusion with the TrOCR-CTC models.
exp38_char6_hhdfull.npz: trained ONLY on the HHD train transcripts (4,740 lines; no Ben-Yehuda or other text, no HHD test text). 198-character alphabet. Decoding weights tuned on the HHD dev writers for the HHD-only TrOCR-CTC model: alpha 2.0 (LM weight), beta 6.0 (per-character bonus), beam 12.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support