--- library_name: transformers license: apache-2.0 pipeline_tag: fill-mask language: - ur - ps - pa - sd - skr - ks - gu tags: - pakmosaic - pakistan - multilingual - encoder - research-preview - sota-target - fill-mask --- # PakMosaic-Small This is the finished ~67M PakMosaic encoder after a **2B-token** serious train on the clean scale mix. ## Architecture Encoder-only MLM: pre-norm, RoPE, GeGLU, full attention, tied embeddings. | | | | --- | ---: | | hidden size | 512 | | layers | 12 | | heads | 8 | | intermediate | 2048 | | vocab | 32,000 | ## How to load ```python from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline repo = "ProximaAI/PakMosaic-Small" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True) fill = pipeline("fill-mask", model=model, tokenizer=tok) print(fill("یہ ایک ہے۔")) ``` `AutoModel` also works (`trust_remote_code=True`) and returns encoder hidden states. ## Data Training mix is Wikimedia plus third-party web crawls used under their original terms. **Only Wikimedia / CC BY-SA is redistributed on the Hub:** [PakMosaic-Wikimedia-v0.2](https://huggingface.co/datasets/ProximaAI/PakMosaic-Wikimedia-v0.2). ## Citation ## Acknowledgements Wikipedia volunteer editors, native reviewers, and [Proxima AI](https://huggingface.co/ProximaAI).