punct_cap_seg_47_language, Core ML, fp32

A Core ML conversion of 1-800-BAD-CODE/punct_cap_seg_47_language: punctuation, true-casing and sentence boundary detection for 47 languages, in one pass over lower-cased, unpunctuated text.

Converted from the original ONNX model at commit 1b9d51fc7989ebc61e844d407d9dadd08ff4ba28 (ONNX โ†’ PyTorch with onnx2torch โ†’ Core ML with coremltools 8.3) by conversion/convert.py, with the versions in conversion/requirements.txt. The weights stay fp32: half precision changes the predictions. The script checks that the punctuation and sentence ends it predicts match the ONNX model's.

Files

  • PunctCapSeg47.mlpackage: the model. Input input_ids, int32 [1, 128]: <s> (1), the text's token ids, </s> (2), padded with <pad> (3). Outputs, one per token:
    • pre_preds, post_preds: indices into pre_labels and post_labels of the original config.yaml โ€” the mark before and after the token;
    • cap_preds ([1, 128, 16]): 1 where that character of the token is upper case;
    • seg_preds: 1 where a sentence ends after the token.
  • vocab.json: the SentencePiece unigram vocabulary (piece, score, type) of spe_unigram_64k_lowercase_47lang.model, and the special ids.
  • ORIGINAL_README.md: the original model card.
  • conversion/convert.py, conversion/requirements.txt: the conversion.

License

Apache 2.0, as the original model. All credit for the model goes to its author.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for aidsoid/punct_cap_seg_47_language-CoreML-fp32

Quantized
(1)
this model