File size: 1,453 Bytes
9ae456b
24c100d
9ae456b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d159f5f
9ae456b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24c100d
9ae456b
 
24c100d
 
 
 
 
9ae456b
 
24c100d
 
9ae456b
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
---

library_name: transformers
license: apache-2.0
pipeline_tag: fill-mask
language:
- ur
- ps
- pa
- sd
- skr
- ks
- gu
tags:
- pakmosaic
- pakistan
- multilingual
- encoder
- research-preview
- sota-target
- fill-mask
---


# PakMosaic-Small

This is the finished ~67M PakMosaic encoder after a **2B-token** serious train on the clean scale mix. 

## Architecture

Encoder-only MLM: pre-norm, RoPE, GeGLU, full attention, tied embeddings.

| | |
| --- | ---: |
| hidden size | 512 |
| layers | 12 |
| heads | 8 |
| intermediate | 2048 |
| vocab | 32,000 |


## How to load

```python

from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline



repo = "ProximaAI/PakMosaic-Small"

tok = AutoTokenizer.from_pretrained(repo)

model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)



fill = pipeline("fill-mask", model=model, tokenizer=tok)

print(fill("یہ ایک <mask> ہے۔"))

```

`AutoModel` also works (`trust_remote_code=True`) and returns encoder hidden states.


## Data

Training mix is Wikimedia plus third-party web crawls used under their original terms. **Only Wikimedia / CC BY-SA is redistributed on the Hub:** [PakMosaic-Wikimedia-v0.2](https://huggingface.co/datasets/ProximaAI/PakMosaic-Wikimedia-v0.2).


## Citation


## Acknowledgements

Wikipedia volunteer editors, native reviewers, and [Proxima AI](https://huggingface.co/ProximaAI).