UncleanCode commited on
Commit
b935288
·
verified ·
1 Parent(s): 9f414f6

Upload via Autoresearch export_data

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ models/labse-ig-ha-yo/tokenizer.json filter=lfs diff=lfs merge=lfs -text
models/labse-ig-ha-yo/1_Pooling/config.json ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {
2
+ "embedding_dimension": 768,
3
+ "pooling_mode": "cls",
4
+ "include_prompt": true
5
+ }
models/labse-ig-ha-yo/2_Dense/config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "in_features": 768,
3
+ "out_features": 768,
4
+ "bias": true,
5
+ "activation_function": "torch.nn.modules.activation.Tanh",
6
+ "module_input_name": "sentence_embedding",
7
+ "module_output_name": "sentence_embedding"
8
+ }
models/labse-ig-ha-yo/2_Dense/model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c7d0e218e49743fefa25884a33cdcdb52d2f119c693db1bd7e0d149fc5932c85
3
+ size 2362528
models/labse-ig-ha-yo/3_Normalize/config.json ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ {
2
+ "module_input_name": "sentence_embedding",
3
+ "module_output_name": "sentence_embedding"
4
+ }
models/labse-ig-ha-yo/README.md ADDED
@@ -0,0 +1,442 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ tags:
3
+ - sentence-transformers
4
+ - sentence-similarity
5
+ - feature-extraction
6
+ - dense
7
+ - generated_from_trainer
8
+ - dataset_size:76150
9
+ - loss:CachedMultipleNegativesRankingLoss
10
+ - loss:MultipleNegativesRankingLoss
11
+ base_model: sentence-transformers/LaBSE
12
+ widget:
13
+ - source_sentence: po pelu ni akoko pelu awon
14
+ sentences:
15
+ - About time with them too.
16
+ - 'Us: We''re just like stars!'
17
+ - It can get you to the next day.
18
+ - source_sentence: Ki ló n ṣẹlẹ / Ki lo n shele?
19
+ sentences:
20
+ - Sure this time it's fine.
21
+ - What's going on/happened?
22
+ - (I've got something in my eye!
23
+ - source_sentence: ban ga laihi gare su int mm
24
+ sentences:
25
+ - '"Cities have been paralyzed"'
26
+ - I wouldn't blame them. (NM)
27
+ - Inside, there are no paths.
28
+ - source_sentence: '"A cikin gõnaki da marẽmari."'
29
+ sentences:
30
+ - How Many Days Are In A 2020?
31
+ - 'And they would say: "Our Lord!'
32
+ - —amid gardens and springs,
33
+ - source_sentence: Mo ti ri pe ninu ara mi ."
34
+ sentences:
35
+ - Bring my Soul out of Prison.
36
+ - I've found it within myself'."
37
+ - I looked and couldn't believe it!
38
+ pipeline_tag: sentence-similarity
39
+ library_name: sentence-transformers
40
+ ---
41
+
42
+ # SentenceTransformer based on sentence-transformers/LaBSE
43
+
44
+ This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [sentence-transformers/LaBSE](https://huggingface.co/sentence-transformers/LaBSE). It maps inputs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.
45
+
46
+ ## Model Details
47
+
48
+ ### Model Description
49
+ - **Model Type:** Sentence Transformer
50
+ - **Base model:** [sentence-transformers/LaBSE](https://huggingface.co/sentence-transformers/LaBSE) <!-- at revision 836121a0533e5664b21c7aacc5d22951f2b8b25b -->
51
+ - **Maximum Sequence Length:** 128 tokens
52
+ - **Output Dimensionality:** 768 dimensions
53
+ - **Similarity Function:** Cosine Similarity
54
+ - **Supported Modality:** Text
55
+ <!-- - **Training Dataset:** Unknown -->
56
+ <!-- - **Language:** Unknown -->
57
+ <!-- - **License:** Unknown -->
58
+
59
+ ### Model Sources
60
+
61
+ - **Documentation:** [Sentence Transformers Documentation](https://sbert.net)
62
+ - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers)
63
+ - **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers)
64
+
65
+ ### Full Model Architecture
66
+
67
+ ```
68
+ SentenceTransformer(
69
+ (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
70
+ (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'cls', 'include_prompt': True})
71
+ (2): Dense({'in_features': 768, 'out_features': 768, 'bias': True, 'activation_function': 'torch.nn.modules.activation.Tanh', 'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
72
+ (3): Normalize({'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
73
+ )
74
+ ```
75
+
76
+ ## Usage
77
+
78
+ ### Direct Usage (Sentence Transformers)
79
+
80
+ First install the Sentence Transformers library:
81
+
82
+ ```bash
83
+ pip install -U sentence-transformers
84
+ ```
85
+ Then you can load this model and run inference.
86
+ ```python
87
+ from sentence_transformers import SentenceTransformer
88
+
89
+ # Download from the 🤗 Hub
90
+ model = SentenceTransformer("sentence_transformers_model_id")
91
+ # Run inference
92
+ sentences = [
93
+ 'Mo ti ri pe ninu ara mi ."',
94
+ 'I\'ve found it within myself\'."',
95
+ 'Bring my Soul out of Prison.',
96
+ ]
97
+ embeddings = model.encode(sentences)
98
+ print(embeddings.shape)
99
+ # [3, 768]
100
+
101
+ # Get the similarity scores for the embeddings
102
+ similarities = model.similarity(embeddings, embeddings)
103
+ print(similarities)
104
+ # tensor([[1.0000, 0.8338, 0.0731],
105
+ # [0.8338, 1.0000, 0.1770],
106
+ # [0.0731, 0.1770, 1.0000]])
107
+ ```
108
+ <!--
109
+ ### Direct Usage (Transformers)
110
+
111
+ <details><summary>Click to see the direct usage in Transformers</summary>
112
+
113
+ </details>
114
+ -->
115
+
116
+ <!--
117
+ ### Downstream Usage (Sentence Transformers)
118
+
119
+ You can finetune this model on your own dataset.
120
+
121
+ <details><summary>Click to expand</summary>
122
+
123
+ </details>
124
+ -->
125
+
126
+ <!--
127
+ ### Out-of-Scope Use
128
+
129
+ *List how the model may foreseeably be misused and address what users ought not to do with the model.*
130
+ -->
131
+
132
+ <!--
133
+ ## Bias, Risks and Limitations
134
+
135
+ *What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*
136
+ -->
137
+
138
+ <!--
139
+ ### Recommendations
140
+
141
+ *What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*
142
+ -->
143
+
144
+ ## Training Details
145
+
146
+ ### Training Dataset
147
+
148
+ #### Unnamed Dataset
149
+
150
+ * Size: 76,150 training samples
151
+ * Columns: <code>anchor</code> and <code>positive</code>
152
+ * Approximate statistics based on the first 100 samples:
153
+ | | anchor | positive |
154
+ |:---------|:-----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------|
155
+ | type | string | string |
156
+ | modality | text | text |
157
+ | details | <ul><li>min: 9 tokens</li><li>mean: 36.91 tokens</li><li>max: 128 tokens</li></ul> | <ul><li>min: 9 tokens</li><li>mean: 35.99 tokens</li><li>max: 128 tokens</li></ul> |
158
+ * Samples:
159
+ | anchor | positive |
160
+ |:-------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
161
+ | <code>Ilé Ẹjọ́ Gíga Jù Lọ Nílẹ̀ Korea á lè lo ìdájọ́ tí Ilé Ẹjọ́ yìí ṣe nínú ọ̀rọ̀ ọ̀kọ̀ọ̀kan àwọn tí ẹ̀rí ọkàn wọn ò jẹ́ kí wọ́n ṣiṣẹ́ ológun.</code> | <code>The Constitutional Court’s decision now opens the door for the Supreme Court of Korea to apply this ruling to specific cases involving conscientious objectors. Hundreds of thousands of people were evacuated, a process that proved to be especially complicated because of government-mandated physical distancing.</code> |
162
+ | <code>"Wanda Ya sanya muku ƙasa shimfiɗa, kuma Ya shigar muku da hanyõyi a cikinta, kuma Ya saukar da ruwa daga sama."</code> | <code>Who has made earth for you like a bed (spread out); and has opened roads (ways and paths etc.) for you therein; and has sent down water (rain) from the sky.</code> |
163
+ | <code>Ìwọ ni Èlíjà bí?"+ Ó sì wí pé: "Èmi kọ́."</code> | <code>Are you Elijah?" and he says, "I am not."</code> |
164
+ * Loss: [<code>CachedMultipleNegativesRankingLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#cachedmultiplenegativesrankingloss) with these parameters:
165
+ ```json
166
+ {
167
+ "scale": 30.0,
168
+ "similarity_fct": "cos_sim",
169
+ "mini_batch_size": 128,
170
+ "mini_batch_num_tokens": null,
171
+ "gather_across_devices": false,
172
+ "directions": [
173
+ "query_to_doc"
174
+ ],
175
+ "partition_mode": "joint",
176
+ "hardness_mode": null,
177
+ "hardness_strength": 0.0
178
+ }
179
+ ```
180
+
181
+ ### Training Hyperparameters
182
+ #### Non-Default Hyperparameters
183
+
184
+ - `per_device_train_batch_size`: 256
185
+ - `num_train_epochs`: 4.0
186
+ - `learning_rate`: 2e-05
187
+ - `lr_scheduler_type`: cosine
188
+ - `warmup_steps`: 0.1
189
+ - `bf16`: True
190
+ - `dataloader_num_workers`: 4
191
+ - `batch_sampler`: no_duplicates
192
+
193
+ #### All Hyperparameters
194
+ <details><summary>Click to expand</summary>
195
+
196
+ - `per_device_train_batch_size`: 256
197
+ - `num_train_epochs`: 4.0
198
+ - `max_steps`: -1
199
+ - `learning_rate`: 2e-05
200
+ - `lr_scheduler_type`: cosine
201
+ - `lr_scheduler_kwargs`: None
202
+ - `warmup_steps`: 0.1
203
+ - `optim`: adamw_torch_fused
204
+ - `optim_args`: None
205
+ - `weight_decay`: 0.0
206
+ - `adam_beta1`: 0.9
207
+ - `adam_beta2`: 0.999
208
+ - `adam_epsilon`: 1e-08
209
+ - `optim_target_modules`: None
210
+ - `gradient_accumulation_steps`: 1
211
+ - `average_tokens_across_devices`: True
212
+ - `max_grad_norm`: 1.0
213
+ - `label_smoothing_factor`: 0.0
214
+ - `bf16`: True
215
+ - `fp16`: False
216
+ - `bf16_full_eval`: False
217
+ - `fp16_full_eval`: False
218
+ - `tf32`: None
219
+ - `gradient_checkpointing`: False
220
+ - `gradient_checkpointing_kwargs`: None
221
+ - `torch_compile`: False
222
+ - `torch_compile_backend`: None
223
+ - `torch_compile_mode`: None
224
+ - `use_liger_kernel`: False
225
+ - `liger_kernel_config`: None
226
+ - `use_cache`: False
227
+ - `neftune_noise_alpha`: None
228
+ - `torch_empty_cache_steps`: None
229
+ - `auto_find_batch_size`: False
230
+ - `log_on_each_node`: True
231
+ - `logging_nan_inf_filter`: True
232
+ - `include_num_input_tokens_seen`: no
233
+ - `log_level`: passive
234
+ - `log_level_replica`: warning
235
+ - `disable_tqdm`: False
236
+ - `project`: huggingface
237
+ - `trackio_space_id`: None
238
+ - `trackio_bucket_id`: None
239
+ - `trackio_static_space_id`: None
240
+ - `per_device_eval_batch_size`: 8
241
+ - `prediction_loss_only`: True
242
+ - `eval_on_start`: False
243
+ - `eval_do_concat_batches`: True
244
+ - `eval_use_gather_object`: False
245
+ - `eval_accumulation_steps`: None
246
+ - `include_for_metrics`: []
247
+ - `batch_eval_metrics`: False
248
+ - `save_only_model`: False
249
+ - `save_on_each_node`: False
250
+ - `enable_jit_checkpoint`: False
251
+ - `push_to_hub`: False
252
+ - `hub_private_repo`: None
253
+ - `hub_model_id`: None
254
+ - `hub_strategy`: every_save
255
+ - `hub_always_push`: False
256
+ - `hub_revision`: None
257
+ - `load_best_model_at_end`: False
258
+ - `ignore_data_skip`: False
259
+ - `restore_callback_states_from_checkpoint`: False
260
+ - `full_determinism`: False
261
+ - `seed`: 42
262
+ - `data_seed`: None
263
+ - `use_cpu`: False
264
+ - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
265
+ - `parallelism_config`: None
266
+ - `dataloader_drop_last`: False
267
+ - `dataloader_num_workers`: 4
268
+ - `dataloader_pin_memory`: True
269
+ - `dataloader_persistent_workers`: False
270
+ - `dataloader_prefetch_factor`: None
271
+ - `dataloader_multiprocessing_context`: None
272
+ - `dataloader_in_order`: True
273
+ - `remove_unused_columns`: True
274
+ - `label_names`: None
275
+ - `train_sampling_strategy`: random
276
+ - `length_column_name`: length
277
+ - `ddp_find_unused_parameters`: None
278
+ - `ddp_bucket_cap_mb`: None
279
+ - `ddp_broadcast_buffers`: False
280
+ - `ddp_static_graph`: None
281
+ - `ddp_backend`: None
282
+ - `ddp_timeout`: 1800
283
+ - `fsdp`: None
284
+ - `fsdp_config`: None
285
+ - `deepspeed`: None
286
+ - `debug`: []
287
+ - `skip_memory_metrics`: True
288
+ - `do_predict`: False
289
+ - `resume_from_checkpoint`: None
290
+ - `local_rank`: -1
291
+ - `prompts`: None
292
+ - `batch_sampler`: no_duplicates
293
+ - `multi_dataset_batch_sampler`: proportional
294
+ - `router_mapping`: {}
295
+ - `learning_rate_mapping`: {}
296
+ - `warmup_ratio`: None
297
+
298
+ </details>
299
+
300
+ ### Training Logs
301
+ | Epoch | Step | Training Loss |
302
+ |:------:|:----:|:-------------:|
303
+ | 0.0671 | 20 | 0.4747 |
304
+ | 0.1342 | 40 | 0.3165 |
305
+ | 0.2013 | 60 | 0.2757 |
306
+ | 0.2685 | 80 | 0.2269 |
307
+ | 0.3356 | 100 | 0.2092 |
308
+ | 0.4027 | 120 | 0.1850 |
309
+ | 0.4698 | 140 | 0.1616 |
310
+ | 0.5369 | 160 | 0.1554 |
311
+ | 0.6040 | 180 | 0.1499 |
312
+ | 0.6711 | 200 | 0.1600 |
313
+ | 0.7383 | 220 | 0.1278 |
314
+ | 0.8054 | 240 | 0.1107 |
315
+ | 0.8725 | 260 | 0.1230 |
316
+ | 0.9396 | 280 | 0.1177 |
317
+ | 1.0067 | 300 | 0.0924 |
318
+ | 1.0738 | 320 | 0.0679 |
319
+ | 1.1409 | 340 | 0.0665 |
320
+ | 1.2081 | 360 | 0.0771 |
321
+ | 1.2752 | 380 | 0.0646 |
322
+ | 1.3423 | 400 | 0.0757 |
323
+ | 1.4094 | 420 | 0.0728 |
324
+ | 1.4765 | 440 | 0.0767 |
325
+ | 1.5436 | 460 | 0.0732 |
326
+ | 1.6107 | 480 | 0.0615 |
327
+ | 1.6779 | 500 | 0.0639 |
328
+ | 1.7450 | 520 | 0.0576 |
329
+ | 1.8121 | 540 | 0.0686 |
330
+ | 1.8792 | 560 | 0.0585 |
331
+ | 1.9463 | 580 | 0.0655 |
332
+ | 2.0134 | 600 | 0.0615 |
333
+ | 2.0805 | 620 | 0.0430 |
334
+ | 2.1477 | 640 | 0.0376 |
335
+ | 2.2148 | 660 | 0.0377 |
336
+ | 2.2819 | 680 | 0.0384 |
337
+ | 2.3490 | 700 | 0.0393 |
338
+ | 2.4161 | 720 | 0.0365 |
339
+ | 2.4832 | 740 | 0.0421 |
340
+ | 2.5503 | 760 | 0.0367 |
341
+ | 2.6174 | 780 | 0.0433 |
342
+ | 2.6846 | 800 | 0.0360 |
343
+ | 2.7517 | 820 | 0.0363 |
344
+ | 2.8188 | 840 | 0.0370 |
345
+ | 2.8859 | 860 | 0.0295 |
346
+ | 2.9530 | 880 | 0.0327 |
347
+ | 3.0201 | 900 | 0.0321 |
348
+ | 3.0872 | 920 | 0.0318 |
349
+ | 3.1544 | 940 | 0.0249 |
350
+ | 3.2215 | 960 | 0.0248 |
351
+ | 3.2886 | 980 | 0.0236 |
352
+ | 3.3557 | 1000 | 0.0249 |
353
+ | 3.4228 | 1020 | 0.0332 |
354
+ | 3.4899 | 1040 | 0.0298 |
355
+ | 3.5570 | 1060 | 0.0283 |
356
+ | 3.6242 | 1080 | 0.0261 |
357
+ | 3.6913 | 1100 | 0.0346 |
358
+ | 3.7584 | 1120 | 0.0270 |
359
+ | 3.8255 | 1140 | 0.0284 |
360
+ | 3.8926 | 1160 | 0.0321 |
361
+ | 3.9597 | 1180 | 0.0267 |
362
+
363
+
364
+ ### Training Time
365
+ - **Training**: 5.0 minutes
366
+
367
+ ### Framework Versions
368
+ - Python: 3.12.3
369
+ - Sentence Transformers: 6.1.0
370
+ - Transformers: 5.19.0
371
+ - PyTorch: 2.13.0+cu129
372
+ - Accelerate: 1.15.0
373
+ - Datasets: 5.1.0
374
+ - Tokenizers: 0.23.2
375
+
376
+ ## Additional Resources
377
+
378
+ - [Training and Finetuning Embedding Models with Sentence Transformers](https://huggingface.co/blog/train-sentence-transformers): the end-to-end guide for training or finetuning Sentence Transformer models.
379
+ - [Introduction to Matryoshka Embedding Models](https://huggingface.co/blog/matryoshka): variable-size embeddings that can be truncated with minimal quality loss.
380
+ - [Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval](https://huggingface.co/blog/embedding-quantization): post-training compression of embedding vectors.
381
+ - [Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/multimodal-sentence-transformers): use text, image, audio, and video models through the same API.
382
+ - [Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/train-multimodal-sentence-transformers): train multimodal embedding models, with a Visual Document Retrieval walkthrough.
383
+
384
+ ## Citation
385
+
386
+ ### BibTeX
387
+
388
+ #### Sentence Transformers
389
+ ```bibtex
390
+ @inproceedings{reimers-2019-sentence-bert,
391
+ title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
392
+ author = "Reimers, Nils and Gurevych, Iryna",
393
+ booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
394
+ month = "11",
395
+ year = "2019",
396
+ publisher = "Association for Computational Linguistics",
397
+ url = "https://arxiv.org/abs/1908.10084",
398
+ }
399
+ ```
400
+
401
+ #### CachedMultipleNegativesRankingLoss
402
+ ```bibtex
403
+ @misc{gao2021scaling,
404
+ title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
405
+ author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
406
+ year={2021},
407
+ eprint={2101.06983},
408
+ archivePrefix={arXiv},
409
+ primaryClass={cs.LG}
410
+ }
411
+ ```
412
+
413
+ #### MultipleNegativesRankingLoss
414
+ ```bibtex
415
+ @misc{oord2019representationlearningcontrastivepredictive,
416
+ title={Representation Learning with Contrastive Predictive Coding},
417
+ author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
418
+ year={2019},
419
+ eprint={1807.03748},
420
+ archivePrefix={arXiv},
421
+ primaryClass={cs.LG},
422
+ url={https://arxiv.org/abs/1807.03748},
423
+ }
424
+ ```
425
+
426
+ <!--
427
+ ## Glossary
428
+
429
+ *Clearly define terms in order to be accessible across audiences.*
430
+ -->
431
+
432
+ <!--
433
+ ## Model Card Authors
434
+
435
+ *Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
436
+ -->
437
+
438
+ <!--
439
+ ## Model Card Contact
440
+
441
+ *Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
442
+ -->
models/labse-ig-ha-yo/config.json ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_cross_attention": false,
3
+ "architectures": [
4
+ "BertModel"
5
+ ],
6
+ "attention_probs_dropout_prob": 0.1,
7
+ "bos_token_id": null,
8
+ "classifier_dropout": null,
9
+ "directionality": "bidi",
10
+ "dtype": "float32",
11
+ "eos_token_id": null,
12
+ "gradient_checkpointing": false,
13
+ "hidden_act": "gelu",
14
+ "hidden_dropout_prob": 0.1,
15
+ "hidden_size": 768,
16
+ "initializer_range": 0.02,
17
+ "intermediate_size": 3072,
18
+ "is_decoder": false,
19
+ "layer_norm_eps": 1e-12,
20
+ "max_position_embeddings": 512,
21
+ "model_type": "bert",
22
+ "num_attention_heads": 12,
23
+ "num_hidden_layers": 12,
24
+ "pad_token_id": 0,
25
+ "pooler_fc_size": 768,
26
+ "pooler_num_attention_heads": 12,
27
+ "pooler_num_fc_layers": 3,
28
+ "pooler_size_per_head": 128,
29
+ "pooler_type": "first_token_transform",
30
+ "position_embedding_type": "absolute",
31
+ "tie_word_embeddings": true,
32
+ "transformers_version": "5.19.0",
33
+ "type_vocab_size": 2,
34
+ "use_cache": false,
35
+ "vocab_size": 501153
36
+ }
models/labse-ig-ha-yo/config_sentence_transformers.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "__version__": {
3
+ "pytorch": "2.13.0+cu129",
4
+ "sentence_transformers": "6.1.0",
5
+ "transformers": "5.19.0"
6
+ },
7
+ "default_prompt_name": null,
8
+ "model_type": "SentenceTransformer",
9
+ "prompts": {
10
+ "document": "",
11
+ "query": ""
12
+ },
13
+ "similarity_fn_name": "cosine"
14
+ }
models/labse-ig-ha-yo/model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cb68b8f24b63156c6faa4dde3c178f31e0ab942a09a361c73bc6cb9ade8eb935
3
+ size 1883730160
models/labse-ig-ha-yo/modules.json ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "idx": 0,
4
+ "name": "0",
5
+ "path": "",
6
+ "type": "sentence_transformers.base.modules.transformer.Transformer"
7
+ },
8
+ {
9
+ "idx": 1,
10
+ "name": "1",
11
+ "path": "1_Pooling",
12
+ "type": "sentence_transformers.sentence_transformer.modules.pooling.Pooling"
13
+ },
14
+ {
15
+ "idx": 2,
16
+ "name": "2",
17
+ "path": "2_Dense",
18
+ "type": "sentence_transformers.base.modules.dense.Dense"
19
+ },
20
+ {
21
+ "idx": 3,
22
+ "name": "3",
23
+ "path": "3_Normalize",
24
+ "type": "sentence_transformers.base.modules.normalize.Normalize"
25
+ }
26
+ ]
models/labse-ig-ha-yo/results.json ADDED
@@ -0,0 +1,66 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "baseline": {
3
+ "ig/clean": {
4
+ "R@1": 1.0,
5
+ "R@10": 1.0,
6
+ "MRR": 1.0
7
+ },
8
+ "ig/asr_noise": {
9
+ "R@1": 0.9853,
10
+ "R@10": 0.9951,
11
+ "MRR": 0.9888
12
+ },
13
+ "ha/clean": {
14
+ "R@1": 1.0,
15
+ "R@10": 1.0,
16
+ "MRR": 1.0
17
+ },
18
+ "ha/asr_noise": {
19
+ "R@1": 0.9951,
20
+ "R@10": 1.0,
21
+ "MRR": 0.9975
22
+ },
23
+ "yo/clean": {
24
+ "R@1": 0.9804,
25
+ "R@10": 1.0,
26
+ "MRR": 0.9871
27
+ },
28
+ "yo/asr_noise": {
29
+ "R@1": 0.9559,
30
+ "R@10": 0.9951,
31
+ "MRR": 0.9683
32
+ }
33
+ },
34
+ "finetuned": {
35
+ "ig/clean": {
36
+ "R@1": 1.0,
37
+ "R@10": 1.0,
38
+ "MRR": 1.0
39
+ },
40
+ "ig/asr_noise": {
41
+ "R@1": 1.0,
42
+ "R@10": 1.0,
43
+ "MRR": 1.0
44
+ },
45
+ "ha/clean": {
46
+ "R@1": 1.0,
47
+ "R@10": 1.0,
48
+ "MRR": 1.0
49
+ },
50
+ "ha/asr_noise": {
51
+ "R@1": 1.0,
52
+ "R@10": 1.0,
53
+ "MRR": 1.0
54
+ },
55
+ "yo/clean": {
56
+ "R@1": 0.9804,
57
+ "R@10": 1.0,
58
+ "MRR": 0.9902
59
+ },
60
+ "yo/asr_noise": {
61
+ "R@1": 0.9804,
62
+ "R@10": 1.0,
63
+ "MRR": 0.9867
64
+ }
65
+ }
66
+ }
models/labse-ig-ha-yo/sentence_bert_config.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "transformer_task": "feature-extraction",
3
+ "modality_config": {
4
+ "text": {
5
+ "method": "forward",
6
+ "method_output_name": "last_hidden_state"
7
+ }
8
+ },
9
+ "module_output_name": "token_embeddings"
10
+ }
models/labse-ig-ha-yo/tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:edba4e57ec22a2a74bbdb601d3f908e4699c34f8386d52ed055e6fe6bd2b51ac
3
+ size 13632172
models/labse-ig-ha-yo/tokenizer_config.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "cls_token": "[CLS]",
4
+ "do_basic_tokenize": true,
5
+ "do_lower_case": false,
6
+ "full_tokenizer_file": null,
7
+ "is_local": false,
8
+ "local_files_only": false,
9
+ "mask_token": "[MASK]",
10
+ "model_max_length": 128,
11
+ "never_split": null,
12
+ "pad_token": "[PAD]",
13
+ "sep_token": "[SEP]",
14
+ "strip_accents": null,
15
+ "tokenize_chinese_chars": true,
16
+ "tokenizer_class": "BertTokenizer",
17
+ "unk_token": "[UNK]"
18
+ }