mahwizzzz commited on
Commit
d159f5f
路
verified 路
1 Parent(s): 2dcd089

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +1 -61
README.md CHANGED
@@ -22,43 +22,7 @@ tags:
22
 
23
  # PakMosaic-Small
24
 
25
- **Pakistan NLP Suite** 路 published by [Proxima AI](https://huggingface.co/ProximaAI)
26
-
27
- **Status:** research preview 路 **SOTA-target** 路 not a factual SOTA claim
28
- `SOTA_VERIFIED = false`
29
-
30
- This is the finished **~67M** PakMosaic encoder after a **2B-token** serious train on the clean scale mix. It is not PakMosaic-Base-132M (that run is still in progress). It is not a published SOTA model.
31
-
32
- ## Training result
33
-
34
- | item | value |
35
- | --- | --- |
36
- | run | `pakmosaic-small_20260925T090924Z` |
37
- | last step | 2,904,769 |
38
- | tokens seen | 2,000,000,068 / 2,000,000,000 |
39
- | last-500 mean MLM loss | 1.958 |
40
- | last-500 mean masked accuracy | 0.640 |
41
- | hardware | NVIDIA GeForce RTX 4060 8GB (solo) |
42
- | objective | masked language modeling, mask prob 0.15 |
43
- | seq length | 512 |
44
- | tokenizer | PakMosaic-Tokenizer-32k (`hf_byte_bpe_32k_identity`) |
45
-
46
- Last-batch loss jumps by language. The 500-step mean is the number to quote.
47
-
48
- ## Diagnostic evals (not PakBench)
49
-
50
- These are engineering sanity checks from `python -m paknlp_registry encoder evaluate`. They are **not** frozen benchmark scores.
51
-
52
- | check | result |
53
- | --- | ---: |
54
- | Same Urdu sentence vs spacing/punctuation noise (cosine) | 0.760 |
55
- | Near paraphrase pair (cosine) | 0.900 |
56
- | Unrelated pair (cosine) | 0.689 |
57
- | Urdu vs Pashto mean-pool centroid cosine | 0.890 |
58
-
59
- Near > far is the expected direction. Centroid cosine is **not** a language-ID score.
60
-
61
- SOTA gate (required PakBench evidence): **failed closed**. Missing frozen evaluation, contamination certificate, multi-language task results, macros, and contemporary baselines. Do not write SOTA on this card until that gate returns true.
62
 
63
  ## Architecture
64
 
@@ -72,7 +36,6 @@ Encoder-only MLM: pre-norm, RoPE, GeGLU, full attention, tied embeddings.
72
  | intermediate | 2048 |
73
  | vocab | 32,000 |
74
 
75
- This is a **custom PakMosaic encoder**, not BERT. Hugging Face Transformers loads it with `trust_remote_code=True` (RoPE + GeGLU + pre-norm live in this repo).
76
 
77
  ## How to load
78
 
@@ -89,37 +52,14 @@ print(fill("蹖蹃 丕蹖讴 <mask> 蹃蹝蹟"))
89
 
90
  `AutoModel` also works (`trust_remote_code=True`) and returns encoder hidden states.
91
 
92
- Tokenizer: PakMosaic-Tokenizer-32k / `hf_byte_bpe_32k_identity` (hash `0f55e47920211588e6985172cd71338c8aff6226fe912c9668055b44bacc7ce7`). Frozen Urdu fertility 1.250 vs XLM-R 1.382 is a **tokenizer** metric, not model SOTA.
93
 
94
  ## Data
95
 
96
  Training mix is Wikimedia plus third-party web crawls used under their original terms. **Only Wikimedia / CC BY-SA is redistributed on the Hub:** [PakMosaic-Wikimedia-v0.2](https://huggingface.co/datasets/ProximaAI/PakMosaic-Wikimedia-v0.2).
97
 
98
- FineWeb2, HPLT 2.0, and Common Voice are **not** re-uploaded here. Get those from the original steward.
99
-
100
- Languages actually sampled in this run: Pashto, Urdu, Punjabi Shahmukhi, Saraiki, Sindhi, Kashmiri, Gujarati. 66 registered identities are not 66 trained languages.
101
-
102
- ## Intended use
103
-
104
- Research on Pakistan-language representation learning and reproduction of the public Wikimedia subset.
105
-
106
- ## Out of scope
107
-
108
- Production deployment, surveillance, claiming 66-language training, or calling this model SOTA.
109
-
110
- ## Family
111
-
112
- | model | params | this release |
113
- | --- | ---: | --- |
114
- | PakMosaic-Micro | ~26.6M | earlier engineering run |
115
- | **PakMosaic-Small** | **~67M** | **this repo, 2B tokens** |
116
- | PakMosaic-Base-132M | ~132.5M | training, not published |
117
-
118
- Project landing: [ProximaAI/PakMosaic](https://huggingface.co/ProximaAI/PakMosaic)
119
 
120
  ## Citation
121
 
122
- Cite the Pakistan NLP Suite / PakMosaic repository version until a paper exists.
123
 
124
  ## Acknowledgements
125
 
 
22
 
23
  # PakMosaic-Small
24
 
25
+ This is the finished ~67M PakMosaic encoder after a **2B-token** serious train on the clean scale mix.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
26
 
27
  ## Architecture
28
 
 
36
  | intermediate | 2048 |
37
  | vocab | 32,000 |
38
 
 
39
 
40
  ## How to load
41
 
 
52
 
53
  `AutoModel` also works (`trust_remote_code=True`) and returns encoder hidden states.
54
 
 
55
 
56
  ## Data
57
 
58
  Training mix is Wikimedia plus third-party web crawls used under their original terms. **Only Wikimedia / CC BY-SA is redistributed on the Hub:** [PakMosaic-Wikimedia-v0.2](https://huggingface.co/datasets/ProximaAI/PakMosaic-Wikimedia-v0.2).
59
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60
 
61
  ## Citation
62
 
 
63
 
64
  ## Acknowledgements
65