PakMosaic-Small / README.md
mahwizzzz's picture
Update README.md
d159f5f verified
|
Raw History Blame Contribute Delete
1.45 kB
---
library_name: transformers
license: apache-2.0
pipeline_tag: fill-mask
language:
- ur
- ps
- pa
- sd
- skr
- ks
- gu
tags:
- pakmosaic
- pakistan
- multilingual
- encoder
- research-preview
- sota-target
- fill-mask
---
# PakMosaic-Small
This is the finished ~67M PakMosaic encoder after a **2B-token** serious train on the clean scale mix.
## Architecture
Encoder-only MLM: pre-norm, RoPE, GeGLU, full attention, tied embeddings.
| | |
| --- | ---: |
| hidden size | 512 |
| layers | 12 |
| heads | 8 |
| intermediate | 2048 |
| vocab | 32,000 |
## How to load
```python
from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline
repo = "ProximaAI/PakMosaic-Small"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)
fill = pipeline("fill-mask", model=model, tokenizer=tok)
print(fill("یہ ایک <mask> ہے۔"))
```
`AutoModel` also works (`trust_remote_code=True`) and returns encoder hidden states.
## Data
Training mix is Wikimedia plus third-party web crawls used under their original terms. **Only Wikimedia / CC BY-SA is redistributed on the Hub:** [PakMosaic-Wikimedia-v0.2](https://huggingface.co/datasets/ProximaAI/PakMosaic-Wikimedia-v0.2).
## Citation
## Acknowledgements
Wikipedia volunteer editors, native reviewers, and [Proxima AI](https://huggingface.co/ProximaAI).