|
Download README.md from Ransaka/sinlib: direct link, hf CLI and curl.
- Browser
- Download file 3.13 kB
-
https://huggingface.co/Ransaka/sinlib/resolve/main/README.md
- Command line
-
hf download hf://Ransaka/sinlib/README.md
-
curl -L -o README.md https://huggingface.co/Ransaka/sinlib/resolve/main/README.md
3.13 kB
| language: | |
| - si | |
| license: mit | |
| tags: | |
| - sinhala | |
| - nlp | |
| - spellcheck | |
| - typo-detection | |
| - seq2seq | |
| - sequence-labeling | |
| - pytorch | |
| metrics: | |
| - accuracy | |
|  | |
| # sinlib: Pretrained Sinhala NLP and Spellchecking Models | |
| This repository hosts the official pretrained model checkpoints, vocabularies, and statistical datasets for the **`sinlib`** library—a comprehensive Sinhala NLP toolkit. | |
| ## Models & Artifacts Hosted | |
| This repository contains the following files loaded dynamically by `sinlib.spellcheck.TypoDetector`: | |
| * **`bigru_detector.pt`**: A bidirectional GRU sequence labeling model (`BiGRUSequenceLabeler`) trained to detect spelling errors and character substitutions at the akshara level. | |
| * **`bigru_corrector.pt`**: A sequence-to-sequence bidirectional GRU encoder-decoder model with Attention (`BiGRUSeq2Seq`) that performs generative character/akshara corrections. | |
| * **`akshara_vocab.json`**: Vocabulary mappings mapping Sinhala phonological units (aksharas) and basic punctuation to token IDs for the neural models. | |
| * **`akshara_ngram.json`**: Stored counts and vocabularies used by the statistical Trigram model (`AksharaNGram`) for likelihood evaluation. | |
| * **`news_unigrams.json` & `news_bigrams.json`**: Pre-calculated unigram and bigram word-level frequencies extracted from Sinhala news corpora, used for context-aware candidate re-ranking (Stupid Backoff). | |
| * **`dictionary.npy`**: The canonical baseline dictionary containing ~61K valid Sinhala words. | |
| * **`ngram_probs.npy`**: N-gram probabilities used by the default `PreTrainedTokenizer` fallback checks. | |
| Additional models, configuration files, and vocabulary resources required by the `sinlib` package are also included in this repository. | |
| --- | |
| ## Intended Use | |
| These weights are designed to be loaded directly through the `sinlib` Python package. | |
| ### Installation | |
| ```bash | |
| pip install sinlib | |
| ``` | |
| ### Example Inference (Spellcheck) | |
| ```python | |
| from sinlib.spellcheck import TypoDetector | |
| # Automatically downloads and caches the model files from this HF repository | |
| detector = TypoDetector.from_pretrained("Ransaka/sinlib") | |
| # Run spellcheck (with punctuation preservation & unicode normalization) | |
| sentence = "කොළඹ වරායේ සිට බස්නහිර දෙසින් නාවික සැතපුම් 17ක් පමණ දුරන්." | |
| corrected = detector(sentence) | |
| print(corrected) | |
| # Output: "කොළඹ වරායේ සිට බස්නාහිර දෙසින් නාවික සැතපුම් 17ක් පමණ දුරින්." | |
| ``` | |
| ### Typo Detection Fallback | |
| If PyTorch is not installed in the target environment, `sinlib` falls back automatically to the statistical N-Gram model (`akshara_ngram.json`) to perform spelling checks. | |
| ```python | |
| # Check word suspicion level | |
| is_typo = detector.is_word_suspicious("පසලට") # True | |
| ``` | |
| For more details, usage examples, and API references, please refer to the official documentation: | |
| https://sinlib.readthedocs.io/en/latest/ |