Pakistan NLP Suite / PakMosaic
Collection
Research preview for Pakistan multilingual encoder work. SOTA-target, not a SOTA claim. Wikimedia-only public text. • 3 items • Updated
How to use ProximaAI/PakMosaic-Small with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("fill-mask", model="ProximaAI/PakMosaic-Small", trust_remote_code=True) # Load model directly
from transformers import AutoModelForMaskedLM
model = AutoModelForMaskedLM.from_pretrained("ProximaAI/PakMosaic-Small", trust_remote_code=True, device_map="auto")# Load model directly
from transformers import AutoModelForMaskedLM
model = AutoModelForMaskedLM.from_pretrained("ProximaAI/PakMosaic-Small", trust_remote_code=True, device_map="auto")This is the finished ~67M PakMosaic encoder after a 2B-token serious train on the clean scale mix.
Encoder-only MLM: pre-norm, RoPE, GeGLU, full attention, tied embeddings.
| hidden size | 512 |
| layers | 12 |
| heads | 8 |
| intermediate | 2048 |
| vocab | 32,000 |
from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline
repo = "ProximaAI/PakMosaic-Small"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)
fill = pipeline("fill-mask", model=model, tokenizer=tok)
print(fill("یہ ایک <mask> ہے۔"))
AutoModel also works (trust_remote_code=True) and returns encoder hidden states.
Training mix is Wikimedia plus third-party web crawls used under their original terms. Only Wikimedia / CC BY-SA is redistributed on the Hub: PakMosaic-Wikimedia-v0.2.
Wikipedia volunteer editors, native reviewers, and Proxima AI.
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="ProximaAI/PakMosaic-Small", trust_remote_code=True)