How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("fill-mask", model="ProximaAI/PakMosaic-Small", trust_remote_code=True)
# Load model directly
from transformers import AutoModelForMaskedLM
model = AutoModelForMaskedLM.from_pretrained("ProximaAI/PakMosaic-Small", trust_remote_code=True, device_map="auto")
Quick Links

PakMosaic-Small

This is the finished ~67M PakMosaic encoder after a 2B-token serious train on the clean scale mix.

Architecture

Encoder-only MLM: pre-norm, RoPE, GeGLU, full attention, tied embeddings.

hidden size 512
layers 12
heads 8
intermediate 2048
vocab 32,000

How to load

from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline

repo = "ProximaAI/PakMosaic-Small"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)

fill = pipeline("fill-mask", model=model, tokenizer=tok)
print(fill("یہ ایک <mask> ہے۔"))

AutoModel also works (trust_remote_code=True) and returns encoder hidden states.

Data

Training mix is Wikimedia plus third-party web crawls used under their original terms. Only Wikimedia / CC BY-SA is redistributed on the Hub: PakMosaic-Wikimedia-v0.2.

Citation

Acknowledgements

Wikipedia volunteer editors, native reviewers, and Proxima AI.

Downloads last month
-
Safetensors
Model size
83.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including ProximaAI/PakMosaic-Small