Fill-Mask
Transformers
Safetensors
pakmosaic
pakistan
multilingual
encoder
research-preview
sota-target
custom_code
Instructions to use ProximaAI/PakMosaic-Small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ProximaAI/PakMosaic-Small with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="ProximaAI/PakMosaic-Small", trust_remote_code=True)# Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("ProximaAI/PakMosaic-Small", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from ProximaAI/PakMosaic-Small: direct link, hf CLI and curl.
- Browser
- Download file 1.45 kB
-
https://huggingface.co/ProximaAI/PakMosaic-Small/resolve/main/README.md
- Command line
-
hf download hf://ProximaAI/PakMosaic-Small/README.md
-
curl -L -o README.md https://huggingface.co/ProximaAI/PakMosaic-Small/resolve/main/README.md
1.45 kB
| library_name: transformers | |
| license: apache-2.0 | |
| pipeline_tag: fill-mask | |
| language: | |
| - ur | |
| - ps | |
| - pa | |
| - sd | |
| - skr | |
| - ks | |
| - gu | |
| tags: | |
| - pakmosaic | |
| - pakistan | |
| - multilingual | |
| - encoder | |
| - research-preview | |
| - sota-target | |
| - fill-mask | |
| # PakMosaic-Small | |
| This is the finished ~67M PakMosaic encoder after a **2B-token** serious train on the clean scale mix. | |
| ## Architecture | |
| Encoder-only MLM: pre-norm, RoPE, GeGLU, full attention, tied embeddings. | |
| | | | | |
| | --- | ---: | | |
| | hidden size | 512 | | |
| | layers | 12 | | |
| | heads | 8 | | |
| | intermediate | 2048 | | |
| | vocab | 32,000 | | |
| ## How to load | |
| ```python | |
| from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline | |
| repo = "ProximaAI/PakMosaic-Small" | |
| tok = AutoTokenizer.from_pretrained(repo) | |
| model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True) | |
| fill = pipeline("fill-mask", model=model, tokenizer=tok) | |
| print(fill("یہ ایک <mask> ہے۔")) | |
| ``` | |
| `AutoModel` also works (`trust_remote_code=True`) and returns encoder hidden states. | |
| ## Data | |
| Training mix is Wikimedia plus third-party web crawls used under their original terms. **Only Wikimedia / CC BY-SA is redistributed on the Hub:** [PakMosaic-Wikimedia-v0.2](https://huggingface.co/datasets/ProximaAI/PakMosaic-Wikimedia-v0.2). | |
| ## Citation | |
| ## Acknowledgements | |
| Wikipedia volunteer editors, native reviewers, and [Proxima AI](https://huggingface.co/ProximaAI). | |