PakMosaic-Small / README.md
mahwizzzz's picture
Update README.md
d159f5f verified
|
Raw History Blame Contribute Delete
1.45 kB
metadata
library_name: transformers
license: apache-2.0
pipeline_tag: fill-mask
language:
  - ur
  - ps
  - pa
  - sd
  - skr
  - ks
  - gu
tags:
  - pakmosaic
  - pakistan
  - multilingual
  - encoder
  - research-preview
  - sota-target
  - fill-mask

PakMosaic-Small

This is the finished ~67M PakMosaic encoder after a 2B-token serious train on the clean scale mix.

Architecture

Encoder-only MLM: pre-norm, RoPE, GeGLU, full attention, tied embeddings.

hidden size 512
layers 12
heads 8
intermediate 2048
vocab 32,000

How to load

from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline

repo = "ProximaAI/PakMosaic-Small"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)

fill = pipeline("fill-mask", model=model, tokenizer=tok)
print(fill("یہ ایک <mask> ہے۔"))

AutoModel also works (trust_remote_code=True) and returns encoder hidden states.

Data

Training mix is Wikimedia plus third-party web crawls used under their original terms. Only Wikimedia / CC BY-SA is redistributed on the Hub: PakMosaic-Wikimedia-v0.2.

Citation

Acknowledgements

Wikipedia volunteer editors, native reviewers, and Proxima AI.