Fill-Mask
Transformers
Safetensors
pakmosaic
pakistan
multilingual
encoder
research-preview
sota-target
custom_code
Instructions to use ProximaAI/PakMosaic-Small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ProximaAI/PakMosaic-Small with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="ProximaAI/PakMosaic-Small", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("ProximaAI/PakMosaic-Small", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -22,43 +22,7 @@ tags:
|
|
| 22 |
|
| 23 |
# PakMosaic-Small
|
| 24 |
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
**Status:** research preview 路 **SOTA-target** 路 not a factual SOTA claim
|
| 28 |
-
`SOTA_VERIFIED = false`
|
| 29 |
-
|
| 30 |
-
This is the finished **~67M** PakMosaic encoder after a **2B-token** serious train on the clean scale mix. It is not PakMosaic-Base-132M (that run is still in progress). It is not a published SOTA model.
|
| 31 |
-
|
| 32 |
-
## Training result
|
| 33 |
-
|
| 34 |
-
| item | value |
|
| 35 |
-
| --- | --- |
|
| 36 |
-
| run | `pakmosaic-small_20260925T090924Z` |
|
| 37 |
-
| last step | 2,904,769 |
|
| 38 |
-
| tokens seen | 2,000,000,068 / 2,000,000,000 |
|
| 39 |
-
| last-500 mean MLM loss | 1.958 |
|
| 40 |
-
| last-500 mean masked accuracy | 0.640 |
|
| 41 |
-
| hardware | NVIDIA GeForce RTX 4060 8GB (solo) |
|
| 42 |
-
| objective | masked language modeling, mask prob 0.15 |
|
| 43 |
-
| seq length | 512 |
|
| 44 |
-
| tokenizer | PakMosaic-Tokenizer-32k (`hf_byte_bpe_32k_identity`) |
|
| 45 |
-
|
| 46 |
-
Last-batch loss jumps by language. The 500-step mean is the number to quote.
|
| 47 |
-
|
| 48 |
-
## Diagnostic evals (not PakBench)
|
| 49 |
-
|
| 50 |
-
These are engineering sanity checks from `python -m paknlp_registry encoder evaluate`. They are **not** frozen benchmark scores.
|
| 51 |
-
|
| 52 |
-
| check | result |
|
| 53 |
-
| --- | ---: |
|
| 54 |
-
| Same Urdu sentence vs spacing/punctuation noise (cosine) | 0.760 |
|
| 55 |
-
| Near paraphrase pair (cosine) | 0.900 |
|
| 56 |
-
| Unrelated pair (cosine) | 0.689 |
|
| 57 |
-
| Urdu vs Pashto mean-pool centroid cosine | 0.890 |
|
| 58 |
-
|
| 59 |
-
Near > far is the expected direction. Centroid cosine is **not** a language-ID score.
|
| 60 |
-
|
| 61 |
-
SOTA gate (required PakBench evidence): **failed closed**. Missing frozen evaluation, contamination certificate, multi-language task results, macros, and contemporary baselines. Do not write SOTA on this card until that gate returns true.
|
| 62 |
|
| 63 |
## Architecture
|
| 64 |
|
|
@@ -72,7 +36,6 @@ Encoder-only MLM: pre-norm, RoPE, GeGLU, full attention, tied embeddings.
|
|
| 72 |
| intermediate | 2048 |
|
| 73 |
| vocab | 32,000 |
|
| 74 |
|
| 75 |
-
This is a **custom PakMosaic encoder**, not BERT. Hugging Face Transformers loads it with `trust_remote_code=True` (RoPE + GeGLU + pre-norm live in this repo).
|
| 76 |
|
| 77 |
## How to load
|
| 78 |
|
|
@@ -89,37 +52,14 @@ print(fill("蹖蹃 丕蹖讴 <mask> 蹃蹝蹟"))
|
|
| 89 |
|
| 90 |
`AutoModel` also works (`trust_remote_code=True`) and returns encoder hidden states.
|
| 91 |
|
| 92 |
-
Tokenizer: PakMosaic-Tokenizer-32k / `hf_byte_bpe_32k_identity` (hash `0f55e47920211588e6985172cd71338c8aff6226fe912c9668055b44bacc7ce7`). Frozen Urdu fertility 1.250 vs XLM-R 1.382 is a **tokenizer** metric, not model SOTA.
|
| 93 |
|
| 94 |
## Data
|
| 95 |
|
| 96 |
Training mix is Wikimedia plus third-party web crawls used under their original terms. **Only Wikimedia / CC BY-SA is redistributed on the Hub:** [PakMosaic-Wikimedia-v0.2](https://huggingface.co/datasets/ProximaAI/PakMosaic-Wikimedia-v0.2).
|
| 97 |
|
| 98 |
-
FineWeb2, HPLT 2.0, and Common Voice are **not** re-uploaded here. Get those from the original steward.
|
| 99 |
-
|
| 100 |
-
Languages actually sampled in this run: Pashto, Urdu, Punjabi Shahmukhi, Saraiki, Sindhi, Kashmiri, Gujarati. 66 registered identities are not 66 trained languages.
|
| 101 |
-
|
| 102 |
-
## Intended use
|
| 103 |
-
|
| 104 |
-
Research on Pakistan-language representation learning and reproduction of the public Wikimedia subset.
|
| 105 |
-
|
| 106 |
-
## Out of scope
|
| 107 |
-
|
| 108 |
-
Production deployment, surveillance, claiming 66-language training, or calling this model SOTA.
|
| 109 |
-
|
| 110 |
-
## Family
|
| 111 |
-
|
| 112 |
-
| model | params | this release |
|
| 113 |
-
| --- | ---: | --- |
|
| 114 |
-
| PakMosaic-Micro | ~26.6M | earlier engineering run |
|
| 115 |
-
| **PakMosaic-Small** | **~67M** | **this repo, 2B tokens** |
|
| 116 |
-
| PakMosaic-Base-132M | ~132.5M | training, not published |
|
| 117 |
-
|
| 118 |
-
Project landing: [ProximaAI/PakMosaic](https://huggingface.co/ProximaAI/PakMosaic)
|
| 119 |
|
| 120 |
## Citation
|
| 121 |
|
| 122 |
-
Cite the Pakistan NLP Suite / PakMosaic repository version until a paper exists.
|
| 123 |
|
| 124 |
## Acknowledgements
|
| 125 |
|
|
|
|
| 22 |
|
| 23 |
# PakMosaic-Small
|
| 24 |
|
| 25 |
+
This is the finished ~67M PakMosaic encoder after a **2B-token** serious train on the clean scale mix.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 26 |
|
| 27 |
## Architecture
|
| 28 |
|
|
|
|
| 36 |
| intermediate | 2048 |
|
| 37 |
| vocab | 32,000 |
|
| 38 |
|
|
|
|
| 39 |
|
| 40 |
## How to load
|
| 41 |
|
|
|
|
| 52 |
|
| 53 |
`AutoModel` also works (`trust_remote_code=True`) and returns encoder hidden states.
|
| 54 |
|
|
|
|
| 55 |
|
| 56 |
## Data
|
| 57 |
|
| 58 |
Training mix is Wikimedia plus third-party web crawls used under their original terms. **Only Wikimedia / CC BY-SA is redistributed on the Hub:** [PakMosaic-Wikimedia-v0.2](https://huggingface.co/datasets/ProximaAI/PakMosaic-Wikimedia-v0.2).
|
| 59 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 60 |
|
| 61 |
## Citation
|
| 62 |
|
|
|
|
| 63 |
|
| 64 |
## Acknowledgements
|
| 65 |
|