UncleanCode commited on
Commit
424aef9
·
verified ·
1 Parent(s): b935288

Upload via Autoresearch export_data

Browse files
.gitattributes CHANGED
@@ -34,3 +34,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  models/labse-ig-ha-yo/tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  models/labse-ig-ha-yo/tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
1_Pooling/config.json ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {
2
+ "embedding_dimension": 768,
3
+ "pooling_mode": "cls",
4
+ "include_prompt": true
5
+ }
2_Dense/config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "in_features": 768,
3
+ "out_features": 768,
4
+ "bias": true,
5
+ "activation_function": "torch.nn.modules.activation.Tanh",
6
+ "module_input_name": "sentence_embedding",
7
+ "module_output_name": "sentence_embedding"
8
+ }
2_Dense/model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c7d0e218e49743fefa25884a33cdcdb52d2f119c693db1bd7e0d149fc5932c85
3
+ size 2362528
3_Normalize/config.json ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ {
2
+ "module_input_name": "sentence_embedding",
3
+ "module_output_name": "sentence_embedding"
4
+ }
MODEL_CARD.md ADDED
@@ -0,0 +1,232 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - ig
4
+ - ha
5
+ - yo
6
+ - en
7
+ library_name: sentence-transformers
8
+ pipeline_tag: sentence-similarity
9
+ tags:
10
+ - embeddings
11
+ - sentence-transformers
12
+ - cross-lingual
13
+ - multilingual
14
+ - igbo
15
+ - hausa
16
+ - yoruba
17
+ - information-retrieval
18
+ - semantic-search
19
+ base_model: sentence-transformers/LaBSE
20
+ ---
21
+
22
+ # Native-Bird
23
+
24
+ **Native-Bird** is a 768-dimensional multilingual sentence-embedding model fine-tuned from **LaBSE** for cross-lingual retrieval involving **Igbo, Hausa, Yoruba, and English**.
25
+
26
+ Its primary use case is **native-language query → English document retrieval**. A user can describe what they are looking for in Igbo while the indexed document is written in English, and Native-Bird maps both into the same embedding space for semantic search.
27
+
28
+ ## What it is designed to do
29
+
30
+ Example:
31
+
32
+ > **Igbo query:** “Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn.”
33
+ >
34
+ > **English document:** “The nucleus consists of two particles - neutrons and protons.”
35
+
36
+ The model should assign a high cosine similarity to the matching English text even though the languages differ.
37
+
38
+ ## Model details
39
+
40
+ - **Model name:** Native-Bird
41
+ - **Architecture:** SentenceTransformer based on LaBSE
42
+ - **Embedding dimension:** 768
43
+ - **Maximum sequence length used during training:** 128 tokens
44
+ - **Similarity:** cosine similarity
45
+ - **Base model:** sentence-transformers/LaBSE
46
+ - **Training objective:** Cached Multiple Negatives Ranking Loss (CachedMNRL)
47
+ - **Languages:** Igbo, Hausa, Yoruba, English
48
+ - **Training examples:** 76,150 paired examples
49
+ - **Epochs:** 4
50
+ - **Per-device batch size:** 256
51
+ - **Learning rate:** 2e-5
52
+ - **Warmup:** 10%
53
+ - **Scheduler:** cosine
54
+ - **Training precision:** BF16
55
+ - **Negative sampling:** in-batch negatives with NO_DUPLICATES; CachedMNRL mini-batch size 128
56
+
57
+ ## Evaluation
58
+
59
+ The retrieval evaluation uses an English corpus of **1,004 sentences** and **204 queries per language**. For each query, the correct English sentence is known. We report Recall@1, Recall@10, and Mean Reciprocal Rank (MRR).
60
+
61
+ ### Fine-tuned Native-Bird
62
+
63
+ | Language | Query condition | R@1 | R@10 | MRR |
64
+ |---|---|---:|---:|---:|
65
+ | Igbo | Clean | 1.0000 | 1.0000 | 1.0000 |
66
+ | Igbo | ASR-style noise | 1.0000 | 1.0000 | 1.0000 |
67
+ | Hausa | Clean | 1.0000 | 1.0000 | 1.0000 |
68
+ | Hausa | ASR-style noise | 1.0000 | 1.0000 | 1.0000 |
69
+ | Yoruba | Clean | 0.9804 | 1.0000 | 0.9902 |
70
+ | Yoruba | ASR-style noise | 0.9804 | 1.0000 | 0.9867 |
71
+
72
+ ### Before fine-tuning: LaBSE baseline
73
+
74
+ | Language | Query condition | R@1 | R@10 | MRR |
75
+ |---|---|---:|---:|---:|
76
+ | Igbo | Clean | 1.0000 | 1.0000 | 1.0000 |
77
+ | Igbo | ASR-style noise | 0.9853 | 0.9951 | 0.9888 |
78
+ | Hausa | Clean | 1.0000 | 1.0000 | 1.0000 |
79
+ | Hausa | ASR-style noise | 0.9951 | 1.0000 | 0.9975 |
80
+ | Yoruba | Clean | 0.9804 | 1.0000 | 0.9871 |
81
+ | Yoruba | ASR-style noise | 0.9559 | 0.9951 | 0.9683 |
82
+
83
+ ### Interpretation
84
+
85
+ Native-Bird is **strong on the benchmark it was trained and evaluated against**, particularly for Igbo and Hausa. Fine-tuning improved noisy-query retrieval for Igbo and Hausa and slightly improved Yoruba.
86
+
87
+ The scores do **not** prove that Native-Bird will retrieve arbitrary long English documents from arbitrary Igbo descriptions with the same accuracy. The current evaluation is sentence-level. Real document search should therefore split documents into passages/chunks before embedding.
88
+
89
+ ## Direct Igbo → English retrieval checks
90
+
91
+ The trained model was directly run against the English evaluation corpus. Examples:
92
+
93
+ **Igbo:** “Mgbanwe n'ọdịdị na-agbakwunye ọdịdị kejenetiki ọhụrụ, ma nhọpụta na-ewepụ ya n'ogwu ọdịdị nke a kọwapụtara.”
94
+
95
+ → **Top English result:** “Mutation adds new genetic variation, and selection removes it from the pool of expressed variation.”
96
+
97
+ Cosine similarity: **0.7079**
98
+
99
+ **Igbo:** “Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn.”
100
+
101
+ → **Top English result:** “The nucleus consists of two particles - neutrons and protons.”
102
+
103
+ Cosine similarity: **0.7049**
104
+
105
+ **Igbo:** “Enwere ọtụtụ ihe mere ha ji kara web prọgzi mma: ha na-agbanwe ụzọ trafik ịntanetị niile, ọ bụghị naanị http.”
106
+
107
+ → **Top English result:** “They are superior to web proxies for several reasons: They re-route all Internet traffic, not only http.”
108
+
109
+ Cosine similarity: **0.6774**
110
+
111
+ These are direct demonstrations of the intended behavior: the query is Igbo while the retrieved target is English.
112
+
113
+ ## Installation
114
+
115
+ ~~~bash
116
+ pip install -U sentence-transformers torch
117
+ ~~~
118
+
119
+ ## Basic usage
120
+
121
+ ~~~python
122
+ from sentence_transformers import SentenceTransformer
123
+
124
+ model = SentenceTransformer("Modularcomputing/Native-Bird")
125
+
126
+ query = "Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn."
127
+
128
+ documents = [
129
+ "The nucleus consists of two particles - neutrons and protons.",
130
+ "The liver is an organ responsible for many metabolic functions.",
131
+ "Photosynthesis converts light energy into chemical energy.",
132
+ ]
133
+
134
+ q = model.encode(query, normalize_embeddings=True)
135
+ d = model.encode(documents, normalize_embeddings=True)
136
+
137
+ scores = d @ q
138
+
139
+ for i in scores.argsort()[::-1]:
140
+ print(float(scores[i]), documents[i])
141
+ ~~~
142
+
143
+ The matching English sentence should rank first.
144
+
145
+ ## Testing the model
146
+
147
+ A minimal test is:
148
+
149
+ 1. Load Native-Bird.
150
+ 2. Encode an Igbo query.
151
+ 3. Encode several English candidate documents.
152
+ 4. Compute cosine similarity.
153
+ 5. Sort candidates by similarity.
154
+ 6. Check whether the semantically matching English document is ranked first.
155
+
156
+ For a real search system, replace the small list with a vector database or FAISS index.
157
+
158
+ ## Searching a real English document collection
159
+
160
+ Do **not** embed a multi-page document as one vector. Split each document into passages, normally around 100–300 words with overlap, and store each passage embedding together with its parent document ID.
161
+
162
+ At query time:
163
+
164
+ 1. Encode the user's Igbo query with Native-Bird.
165
+ 2. Search the English passage vectors using cosine similarity or an ANN index such as FAISS.
166
+ 3. Return the highest-scoring passages.
167
+ 4. Group or rerank passages by their parent document.
168
+
169
+ Conceptually:
170
+
171
+ ~~~text
172
+ Igbo query
173
+ ↓
174
+ Native-Bird encoder
175
+ ↓
176
+ 768-dimensional query vector
177
+ ↓ cosine similarity
178
+ English passage index
179
+ ↓
180
+ Top-k English passages
181
+ ↓
182
+ Parent documents
183
+ ~~~
184
+
185
+ This is the architecture needed for the intended “describe it in Igbo, find the English document” product.
186
+
187
+ ## Important limitation
188
+
189
+ Native-Bird is an **embedding/retrieval model**, not a translator and not a generative language model. It does not generate an English answer. It represents text as vectors so that semantically related text can be found across languages.
190
+
191
+ The current model uses a 128-token training sequence length and was evaluated on sentence-level retrieval. For long documents, chunking is therefore important.
192
+
193
+ A production deployment should be evaluated on the actual target domain. Nigerian government documents, university material, medical documents, legal documents, and technical documentation can have vocabulary and writing styles that differ substantially from the current benchmark.
194
+
195
+ ## Evaluation methodology
196
+
197
+ The evaluation script embeds the 1,004-item English corpus, embeds each language's 204 queries, computes cosine similarities, and measures the rank of the known matching English sentence.
198
+
199
+ Two query variants are evaluated:
200
+
201
+ - **clean:** original native-language query
202
+ - **ASR-style noise:** Unicode diacritics removed, lowercased, punctuation stripped, and whitespace normalized
203
+
204
+ The benchmark is useful for measuring cross-lingual retrieval and robustness to transcription-style normalization, but should not be interpreted as a universal real-world retrieval score.
205
+
206
+ ## Intended applications
207
+
208
+ - Igbo → English semantic document search
209
+ - Hausa → English semantic document search
210
+ - Yoruba → English semantic document search
211
+ - Multilingual educational search
212
+ - Native-language interfaces for English knowledge bases
213
+ - Cross-lingual RAG retrieval
214
+ - Search over Nigerian institutional documents
215
+
216
+ ## Training provenance
217
+
218
+ Native-Bird was fine-tuned from LaBSE using paired native-language/English retrieval data prepared for this project. Training used 76,150 pairs and a contrastive retrieval objective.
219
+
220
+ The final training run completed successfully for four epochs, saved the model, and produced the evaluation results reported above.
221
+
222
+ ## Model status
223
+
224
+ **Status: research / early release.**
225
+
226
+ The core cross-lingual retrieval capability is demonstrated. The next meaningful validation step is a human-annotated benchmark of real Nigerian documents and real Igbo/Hausa/Yoruba information needs, especially for multi-paragraph documents.
227
+
228
+ ## Citation
229
+
230
+ If you use Native-Bird in research or a project, please reference the model as:
231
+
232
+ **Native-Bird: Cross-lingual Native-Language Embeddings for English Document Retrieval.**
README.md CHANGED
@@ -1,3 +1,232 @@
1
  ---
2
- license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ language:
3
+ - ig
4
+ - ha
5
+ - yo
6
+ - en
7
+ library_name: sentence-transformers
8
+ pipeline_tag: sentence-similarity
9
+ tags:
10
+ - embeddings
11
+ - sentence-transformers
12
+ - cross-lingual
13
+ - multilingual
14
+ - igbo
15
+ - hausa
16
+ - yoruba
17
+ - information-retrieval
18
+ - semantic-search
19
+ base_model: sentence-transformers/LaBSE
20
  ---
21
+
22
+ # Native-Bird
23
+
24
+ **Native-Bird** is a 768-dimensional multilingual sentence-embedding model fine-tuned from **LaBSE** for cross-lingual retrieval involving **Igbo, Hausa, Yoruba, and English**.
25
+
26
+ Its primary use case is **native-language query → English document retrieval**. A user can describe what they are looking for in Igbo while the indexed document is written in English, and Native-Bird maps both into the same embedding space for semantic search.
27
+
28
+ ## What it is designed to do
29
+
30
+ Example:
31
+
32
+ > **Igbo query:** “Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn.”
33
+ >
34
+ > **English document:** “The nucleus consists of two particles - neutrons and protons.”
35
+
36
+ The model should assign a high cosine similarity to the matching English text even though the languages differ.
37
+
38
+ ## Model details
39
+
40
+ - **Model name:** Native-Bird
41
+ - **Architecture:** SentenceTransformer based on LaBSE
42
+ - **Embedding dimension:** 768
43
+ - **Maximum sequence length used during training:** 128 tokens
44
+ - **Similarity:** cosine similarity
45
+ - **Base model:** sentence-transformers/LaBSE
46
+ - **Training objective:** Cached Multiple Negatives Ranking Loss (CachedMNRL)
47
+ - **Languages:** Igbo, Hausa, Yoruba, English
48
+ - **Training examples:** 76,150 paired examples
49
+ - **Epochs:** 4
50
+ - **Per-device batch size:** 256
51
+ - **Learning rate:** 2e-5
52
+ - **Warmup:** 10%
53
+ - **Scheduler:** cosine
54
+ - **Training precision:** BF16
55
+ - **Negative sampling:** in-batch negatives with NO_DUPLICATES; CachedMNRL mini-batch size 128
56
+
57
+ ## Evaluation
58
+
59
+ The retrieval evaluation uses an English corpus of **1,004 sentences** and **204 queries per language**. For each query, the correct English sentence is known. We report Recall@1, Recall@10, and Mean Reciprocal Rank (MRR).
60
+
61
+ ### Fine-tuned Native-Bird
62
+
63
+ | Language | Query condition | R@1 | R@10 | MRR |
64
+ |---|---|---:|---:|---:|
65
+ | Igbo | Clean | 1.0000 | 1.0000 | 1.0000 |
66
+ | Igbo | ASR-style noise | 1.0000 | 1.0000 | 1.0000 |
67
+ | Hausa | Clean | 1.0000 | 1.0000 | 1.0000 |
68
+ | Hausa | ASR-style noise | 1.0000 | 1.0000 | 1.0000 |
69
+ | Yoruba | Clean | 0.9804 | 1.0000 | 0.9902 |
70
+ | Yoruba | ASR-style noise | 0.9804 | 1.0000 | 0.9867 |
71
+
72
+ ### Before fine-tuning: LaBSE baseline
73
+
74
+ | Language | Query condition | R@1 | R@10 | MRR |
75
+ |---|---|---:|---:|---:|
76
+ | Igbo | Clean | 1.0000 | 1.0000 | 1.0000 |
77
+ | Igbo | ASR-style noise | 0.9853 | 0.9951 | 0.9888 |
78
+ | Hausa | Clean | 1.0000 | 1.0000 | 1.0000 |
79
+ | Hausa | ASR-style noise | 0.9951 | 1.0000 | 0.9975 |
80
+ | Yoruba | Clean | 0.9804 | 1.0000 | 0.9871 |
81
+ | Yoruba | ASR-style noise | 0.9559 | 0.9951 | 0.9683 |
82
+
83
+ ### Interpretation
84
+
85
+ Native-Bird is **strong on the benchmark it was trained and evaluated against**, particularly for Igbo and Hausa. Fine-tuning improved noisy-query retrieval for Igbo and Hausa and slightly improved Yoruba.
86
+
87
+ The scores do **not** prove that Native-Bird will retrieve arbitrary long English documents from arbitrary Igbo descriptions with the same accuracy. The current evaluation is sentence-level. Real document search should therefore split documents into passages/chunks before embedding.
88
+
89
+ ## Direct Igbo → English retrieval checks
90
+
91
+ The trained model was directly run against the English evaluation corpus. Examples:
92
+
93
+ **Igbo:** “Mgbanwe n'ọdịdị na-agbakwunye ọdịdị kejenetiki ọhụrụ, ma nhọpụta na-ewepụ ya n'ogwu ọdịdị nke a kọwapụtara.”
94
+
95
+ → **Top English result:** “Mutation adds new genetic variation, and selection removes it from the pool of expressed variation.”
96
+
97
+ Cosine similarity: **0.7079**
98
+
99
+ **Igbo:** “Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn.”
100
+
101
+ → **Top English result:** “The nucleus consists of two particles - neutrons and protons.”
102
+
103
+ Cosine similarity: **0.7049**
104
+
105
+ **Igbo:** “Enwere ọtụtụ ihe mere ha ji kara web prọgzi mma: ha na-agbanwe ụzọ trafik ịntanetị niile, ọ bụghị naanị http.”
106
+
107
+ → **Top English result:** “They are superior to web proxies for several reasons: They re-route all Internet traffic, not only http.”
108
+
109
+ Cosine similarity: **0.6774**
110
+
111
+ These are direct demonstrations of the intended behavior: the query is Igbo while the retrieved target is English.
112
+
113
+ ## Installation
114
+
115
+ ~~~bash
116
+ pip install -U sentence-transformers torch
117
+ ~~~
118
+
119
+ ## Basic usage
120
+
121
+ ~~~python
122
+ from sentence_transformers import SentenceTransformer
123
+
124
+ model = SentenceTransformer("Modularcomputing/Native-Bird")
125
+
126
+ query = "Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn."
127
+
128
+ documents = [
129
+ "The nucleus consists of two particles - neutrons and protons.",
130
+ "The liver is an organ responsible for many metabolic functions.",
131
+ "Photosynthesis converts light energy into chemical energy.",
132
+ ]
133
+
134
+ q = model.encode(query, normalize_embeddings=True)
135
+ d = model.encode(documents, normalize_embeddings=True)
136
+
137
+ scores = d @ q
138
+
139
+ for i in scores.argsort()[::-1]:
140
+ print(float(scores[i]), documents[i])
141
+ ~~~
142
+
143
+ The matching English sentence should rank first.
144
+
145
+ ## Testing the model
146
+
147
+ A minimal test is:
148
+
149
+ 1. Load Native-Bird.
150
+ 2. Encode an Igbo query.
151
+ 3. Encode several English candidate documents.
152
+ 4. Compute cosine similarity.
153
+ 5. Sort candidates by similarity.
154
+ 6. Check whether the semantically matching English document is ranked first.
155
+
156
+ For a real search system, replace the small list with a vector database or FAISS index.
157
+
158
+ ## Searching a real English document collection
159
+
160
+ Do **not** embed a multi-page document as one vector. Split each document into passages, normally around 100–300 words with overlap, and store each passage embedding together with its parent document ID.
161
+
162
+ At query time:
163
+
164
+ 1. Encode the user's Igbo query with Native-Bird.
165
+ 2. Search the English passage vectors using cosine similarity or an ANN index such as FAISS.
166
+ 3. Return the highest-scoring passages.
167
+ 4. Group or rerank passages by their parent document.
168
+
169
+ Conceptually:
170
+
171
+ ~~~text
172
+ Igbo query
173
+ ↓
174
+ Native-Bird encoder
175
+ ↓
176
+ 768-dimensional query vector
177
+ ↓ cosine similarity
178
+ English passage index
179
+ ↓
180
+ Top-k English passages
181
+ ↓
182
+ Parent documents
183
+ ~~~
184
+
185
+ This is the architecture needed for the intended “describe it in Igbo, find the English document” product.
186
+
187
+ ## Important limitation
188
+
189
+ Native-Bird is an **embedding/retrieval model**, not a translator and not a generative language model. It does not generate an English answer. It represents text as vectors so that semantically related text can be found across languages.
190
+
191
+ The current model uses a 128-token training sequence length and was evaluated on sentence-level retrieval. For long documents, chunking is therefore important.
192
+
193
+ A production deployment should be evaluated on the actual target domain. Nigerian government documents, university material, medical documents, legal documents, and technical documentation can have vocabulary and writing styles that differ substantially from the current benchmark.
194
+
195
+ ## Evaluation methodology
196
+
197
+ The evaluation script embeds the 1,004-item English corpus, embeds each language's 204 queries, computes cosine similarities, and measures the rank of the known matching English sentence.
198
+
199
+ Two query variants are evaluated:
200
+
201
+ - **clean:** original native-language query
202
+ - **ASR-style noise:** Unicode diacritics removed, lowercased, punctuation stripped, and whitespace normalized
203
+
204
+ The benchmark is useful for measuring cross-lingual retrieval and robustness to transcription-style normalization, but should not be interpreted as a universal real-world retrieval score.
205
+
206
+ ## Intended applications
207
+
208
+ - Igbo → English semantic document search
209
+ - Hausa → English semantic document search
210
+ - Yoruba → English semantic document search
211
+ - Multilingual educational search
212
+ - Native-language interfaces for English knowledge bases
213
+ - Cross-lingual RAG retrieval
214
+ - Search over Nigerian institutional documents
215
+
216
+ ## Training provenance
217
+
218
+ Native-Bird was fine-tuned from LaBSE using paired native-language/English retrieval data prepared for this project. Training used 76,150 pairs and a contrastive retrieval objective.
219
+
220
+ The final training run completed successfully for four epochs, saved the model, and produced the evaluation results reported above.
221
+
222
+ ## Model status
223
+
224
+ **Status: research / early release.**
225
+
226
+ The core cross-lingual retrieval capability is demonstrated. The next meaningful validation step is a human-annotated benchmark of real Nigerian documents and real Igbo/Hausa/Yoruba information needs, especially for multi-paragraph documents.
227
+
228
+ ## Citation
229
+
230
+ If you use Native-Bird in research or a project, please reference the model as:
231
+
232
+ **Native-Bird: Cross-lingual Native-Language Embeddings for English Document Retrieval.**
config.json ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_cross_attention": false,
3
+ "architectures": [
4
+ "BertModel"
5
+ ],
6
+ "attention_probs_dropout_prob": 0.1,
7
+ "bos_token_id": null,
8
+ "classifier_dropout": null,
9
+ "directionality": "bidi",
10
+ "dtype": "float32",
11
+ "eos_token_id": null,
12
+ "gradient_checkpointing": false,
13
+ "hidden_act": "gelu",
14
+ "hidden_dropout_prob": 0.1,
15
+ "hidden_size": 768,
16
+ "initializer_range": 0.02,
17
+ "intermediate_size": 3072,
18
+ "is_decoder": false,
19
+ "layer_norm_eps": 1e-12,
20
+ "max_position_embeddings": 512,
21
+ "model_type": "bert",
22
+ "num_attention_heads": 12,
23
+ "num_hidden_layers": 12,
24
+ "pad_token_id": 0,
25
+ "pooler_fc_size": 768,
26
+ "pooler_num_attention_heads": 12,
27
+ "pooler_num_fc_layers": 3,
28
+ "pooler_size_per_head": 128,
29
+ "pooler_type": "first_token_transform",
30
+ "position_embedding_type": "absolute",
31
+ "tie_word_embeddings": true,
32
+ "transformers_version": "5.19.0",
33
+ "type_vocab_size": 2,
34
+ "use_cache": false,
35
+ "vocab_size": 501153
36
+ }
config_sentence_transformers.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "__version__": {
3
+ "pytorch": "2.13.0+cu129",
4
+ "sentence_transformers": "6.1.0",
5
+ "transformers": "5.19.0"
6
+ },
7
+ "default_prompt_name": null,
8
+ "model_type": "SentenceTransformer",
9
+ "prompts": {
10
+ "document": "",
11
+ "query": ""
12
+ },
13
+ "similarity_fn_name": "cosine"
14
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cb68b8f24b63156c6faa4dde3c178f31e0ab942a09a361c73bc6cb9ade8eb935
3
+ size 1883730160
modules.json ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "idx": 0,
4
+ "name": "0",
5
+ "path": "",
6
+ "type": "sentence_transformers.base.modules.transformer.Transformer"
7
+ },
8
+ {
9
+ "idx": 1,
10
+ "name": "1",
11
+ "path": "1_Pooling",
12
+ "type": "sentence_transformers.sentence_transformer.modules.pooling.Pooling"
13
+ },
14
+ {
15
+ "idx": 2,
16
+ "name": "2",
17
+ "path": "2_Dense",
18
+ "type": "sentence_transformers.base.modules.dense.Dense"
19
+ },
20
+ {
21
+ "idx": 3,
22
+ "name": "3",
23
+ "path": "3_Normalize",
24
+ "type": "sentence_transformers.base.modules.normalize.Normalize"
25
+ }
26
+ ]
results.json ADDED
@@ -0,0 +1,66 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "baseline": {
3
+ "ig/clean": {
4
+ "R@1": 1.0,
5
+ "R@10": 1.0,
6
+ "MRR": 1.0
7
+ },
8
+ "ig/asr_noise": {
9
+ "R@1": 0.9853,
10
+ "R@10": 0.9951,
11
+ "MRR": 0.9888
12
+ },
13
+ "ha/clean": {
14
+ "R@1": 1.0,
15
+ "R@10": 1.0,
16
+ "MRR": 1.0
17
+ },
18
+ "ha/asr_noise": {
19
+ "R@1": 0.9951,
20
+ "R@10": 1.0,
21
+ "MRR": 0.9975
22
+ },
23
+ "yo/clean": {
24
+ "R@1": 0.9804,
25
+ "R@10": 1.0,
26
+ "MRR": 0.9871
27
+ },
28
+ "yo/asr_noise": {
29
+ "R@1": 0.9559,
30
+ "R@10": 0.9951,
31
+ "MRR": 0.9683
32
+ }
33
+ },
34
+ "finetuned": {
35
+ "ig/clean": {
36
+ "R@1": 1.0,
37
+ "R@10": 1.0,
38
+ "MRR": 1.0
39
+ },
40
+ "ig/asr_noise": {
41
+ "R@1": 1.0,
42
+ "R@10": 1.0,
43
+ "MRR": 1.0
44
+ },
45
+ "ha/clean": {
46
+ "R@1": 1.0,
47
+ "R@10": 1.0,
48
+ "MRR": 1.0
49
+ },
50
+ "ha/asr_noise": {
51
+ "R@1": 1.0,
52
+ "R@10": 1.0,
53
+ "MRR": 1.0
54
+ },
55
+ "yo/clean": {
56
+ "R@1": 0.9804,
57
+ "R@10": 1.0,
58
+ "MRR": 0.9902
59
+ },
60
+ "yo/asr_noise": {
61
+ "R@1": 0.9804,
62
+ "R@10": 1.0,
63
+ "MRR": 0.9867
64
+ }
65
+ }
66
+ }
sentence_bert_config.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "transformer_task": "feature-extraction",
3
+ "modality_config": {
4
+ "text": {
5
+ "method": "forward",
6
+ "method_output_name": "last_hidden_state"
7
+ }
8
+ },
9
+ "module_output_name": "token_embeddings"
10
+ }
test_native_bird.py ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from sentence_transformers import SentenceTransformer
2
+
3
+ MODEL = "Modularcomputing/Native-Bird"
4
+
5
+ query = "Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn."
6
+ documents = [
7
+ "The nucleus consists of two particles - neutrons and protons.",
8
+ "The liver is an organ responsible for many metabolic functions.",
9
+ "Photosynthesis converts light energy into chemical energy.",
10
+ ]
11
+
12
+ model = SentenceTransformer(MODEL)
13
+ q = model.encode(query, normalize_embeddings=True)
14
+ d = model.encode(documents, normalize_embeddings=True)
15
+ scores = d @ q
16
+
17
+ print("Native-Bird cross-lingual retrieval test")
18
+ for rank, i in enumerate(scores.argsort()[::-1], 1):
19
+ print(f"{rank}. score={scores[i]:.4f} | {documents[i]}")
20
+
21
+ assert scores.argmax() == 0, "The expected English match was not ranked first."
22
+ print("PASS: Igbo query retrieved the matching English document first.")
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:edba4e57ec22a2a74bbdb601d3f908e4699c34f8386d52ed055e6fe6bd2b51ac
3
+ size 13632172
tokenizer_config.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "cls_token": "[CLS]",
4
+ "do_basic_tokenize": true,
5
+ "do_lower_case": false,
6
+ "full_tokenizer_file": null,
7
+ "is_local": false,
8
+ "local_files_only": false,
9
+ "mask_token": "[MASK]",
10
+ "model_max_length": 128,
11
+ "never_split": null,
12
+ "pad_token": "[PAD]",
13
+ "sep_token": "[SEP]",
14
+ "strip_accents": null,
15
+ "tokenize_chinese_chars": true,
16
+ "tokenizer_class": "BertTokenizer",
17
+ "unk_token": "[UNK]"
18
+ }