--- tags: - sentence-transformers - sentence-similarity - feature-extraction - dense - generated_from_trainer - dataset_size:76150 - loss:CachedMultipleNegativesRankingLoss - loss:MultipleNegativesRankingLoss base_model: sentence-transformers/LaBSE widget: - source_sentence: po pelu ni akoko pelu awon sentences: - About time with them too. - 'Us: We''re just like stars!' - It can get you to the next day. - source_sentence: Ki ló n ṣẹlẹ / Ki lo n shele? sentences: - Sure this time it's fine. - What's going on/happened? - (I've got something in my eye! - source_sentence: ban ga laihi gare su int mm sentences: - '"Cities have been paralyzed"' - I wouldn't blame them. (NM) - Inside, there are no paths. - source_sentence: '"A cikin gõnaki da marẽmari."' sentences: - How Many Days Are In A 2020? - 'And they would say: "Our Lord!' - —amid gardens and springs, - source_sentence: Mo ti ri pe ninu ara mi ." sentences: - Bring my Soul out of Prison. - I've found it within myself'." - I looked and couldn't believe it! pipeline_tag: sentence-similarity library_name: sentence-transformers --- # SentenceTransformer based on sentence-transformers/LaBSE This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [sentence-transformers/LaBSE](https://huggingface.co/sentence-transformers/LaBSE). It maps inputs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more. ## Model Details ### Model Description - **Model Type:** Sentence Transformer - **Base model:** [sentence-transformers/LaBSE](https://huggingface.co/sentence-transformers/LaBSE) - **Maximum Sequence Length:** 128 tokens - **Output Dimensionality:** 768 dimensions - **Similarity Function:** Cosine Similarity - **Supported Modality:** Text ### Model Sources - **Documentation:** [Sentence Transformers Documentation](https://sbert.net) - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers) - **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers) ### Full Model Architecture ``` SentenceTransformer( (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'}) (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'cls', 'include_prompt': True}) (2): Dense({'in_features': 768, 'out_features': 768, 'bias': True, 'activation_function': 'torch.nn.modules.activation.Tanh', 'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'}) (3): Normalize({'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'}) ) ``` ## Usage ### Direct Usage (Sentence Transformers) First install the Sentence Transformers library: ```bash pip install -U sentence-transformers ``` Then you can load this model and run inference. ```python from sentence_transformers import SentenceTransformer # Download from the 🤗 Hub model = SentenceTransformer("sentence_transformers_model_id") # Run inference sentences = [ 'Mo ti ri pe ninu ara mi ."', 'I\'ve found it within myself\'."', 'Bring my Soul out of Prison.', ] embeddings = model.encode(sentences) print(embeddings.shape) # [3, 768] # Get the similarity scores for the embeddings similarities = model.similarity(embeddings, embeddings) print(similarities) # tensor([[1.0000, 0.8338, 0.0731], # [0.8338, 1.0000, 0.1770], # [0.0731, 0.1770, 1.0000]]) ``` ## Training Details ### Training Dataset #### Unnamed Dataset * Size: 76,150 training samples * Columns: anchor and positive * Approximate statistics based on the first 100 samples: | | anchor | positive | |:---------|:-----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | type | string | string | | modality | text | text | | details | | | * Samples: | anchor | positive | |:-------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | Ilé Ẹjọ́ Gíga Jù Lọ Nílẹ̀ Korea á lè lo ìdájọ́ tí Ilé Ẹjọ́ yìí ṣe nínú ọ̀rọ̀ ọ̀kọ̀ọ̀kan àwọn tí ẹ̀rí ọkàn wọn ò jẹ́ kí wọ́n ṣiṣẹ́ ológun. | The Constitutional Court’s decision now opens the door for the Supreme Court of Korea to apply this ruling to specific cases involving conscientious objectors. Hundreds of thousands of people were evacuated, a process that proved to be especially complicated because of government-mandated physical distancing. | | "Wanda Ya sanya muku ƙasa shimfiɗa, kuma Ya shigar muku da hanyõyi a cikinta, kuma Ya saukar da ruwa daga sama." | Who has made earth for you like a bed (spread out); and has opened roads (ways and paths etc.) for you therein; and has sent down water (rain) from the sky. | | Ìwọ ni Èlíjà bí?"+ Ó sì wí pé: "Èmi kọ́." | Are you Elijah?" and he says, "I am not." | * Loss: [CachedMultipleNegativesRankingLoss](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#cachedmultiplenegativesrankingloss) with these parameters: ```json { "scale": 30.0, "similarity_fct": "cos_sim", "mini_batch_size": 128, "mini_batch_num_tokens": null, "gather_across_devices": false, "directions": [ "query_to_doc" ], "partition_mode": "joint", "hardness_mode": null, "hardness_strength": 0.0 } ``` ### Training Hyperparameters #### Non-Default Hyperparameters - `per_device_train_batch_size`: 256 - `num_train_epochs`: 4.0 - `learning_rate`: 2e-05 - `lr_scheduler_type`: cosine - `warmup_steps`: 0.1 - `bf16`: True - `dataloader_num_workers`: 4 - `batch_sampler`: no_duplicates #### All Hyperparameters
Click to expand - `per_device_train_batch_size`: 256 - `num_train_epochs`: 4.0 - `max_steps`: -1 - `learning_rate`: 2e-05 - `lr_scheduler_type`: cosine - `lr_scheduler_kwargs`: None - `warmup_steps`: 0.1 - `optim`: adamw_torch_fused - `optim_args`: None - `weight_decay`: 0.0 - `adam_beta1`: 0.9 - `adam_beta2`: 0.999 - `adam_epsilon`: 1e-08 - `optim_target_modules`: None - `gradient_accumulation_steps`: 1 - `average_tokens_across_devices`: True - `max_grad_norm`: 1.0 - `label_smoothing_factor`: 0.0 - `bf16`: True - `fp16`: False - `bf16_full_eval`: False - `fp16_full_eval`: False - `tf32`: None - `gradient_checkpointing`: False - `gradient_checkpointing_kwargs`: None - `torch_compile`: False - `torch_compile_backend`: None - `torch_compile_mode`: None - `use_liger_kernel`: False - `liger_kernel_config`: None - `use_cache`: False - `neftune_noise_alpha`: None - `torch_empty_cache_steps`: None - `auto_find_batch_size`: False - `log_on_each_node`: True - `logging_nan_inf_filter`: True - `include_num_input_tokens_seen`: no - `log_level`: passive - `log_level_replica`: warning - `disable_tqdm`: False - `project`: huggingface - `trackio_space_id`: None - `trackio_bucket_id`: None - `trackio_static_space_id`: None - `per_device_eval_batch_size`: 8 - `prediction_loss_only`: True - `eval_on_start`: False - `eval_do_concat_batches`: True - `eval_use_gather_object`: False - `eval_accumulation_steps`: None - `include_for_metrics`: [] - `batch_eval_metrics`: False - `save_only_model`: False - `save_on_each_node`: False - `enable_jit_checkpoint`: False - `push_to_hub`: False - `hub_private_repo`: None - `hub_model_id`: None - `hub_strategy`: every_save - `hub_always_push`: False - `hub_revision`: None - `load_best_model_at_end`: False - `ignore_data_skip`: False - `restore_callback_states_from_checkpoint`: False - `full_determinism`: False - `seed`: 42 - `data_seed`: None - `use_cpu`: False - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None} - `parallelism_config`: None - `dataloader_drop_last`: False - `dataloader_num_workers`: 4 - `dataloader_pin_memory`: True - `dataloader_persistent_workers`: False - `dataloader_prefetch_factor`: None - `dataloader_multiprocessing_context`: None - `dataloader_in_order`: True - `remove_unused_columns`: True - `label_names`: None - `train_sampling_strategy`: random - `length_column_name`: length - `ddp_find_unused_parameters`: None - `ddp_bucket_cap_mb`: None - `ddp_broadcast_buffers`: False - `ddp_static_graph`: None - `ddp_backend`: None - `ddp_timeout`: 1800 - `fsdp`: None - `fsdp_config`: None - `deepspeed`: None - `debug`: [] - `skip_memory_metrics`: True - `do_predict`: False - `resume_from_checkpoint`: None - `local_rank`: -1 - `prompts`: None - `batch_sampler`: no_duplicates - `multi_dataset_batch_sampler`: proportional - `router_mapping`: {} - `learning_rate_mapping`: {} - `warmup_ratio`: None
### Training Logs | Epoch | Step | Training Loss | |:------:|:----:|:-------------:| | 0.0671 | 20 | 0.4747 | | 0.1342 | 40 | 0.3165 | | 0.2013 | 60 | 0.2757 | | 0.2685 | 80 | 0.2269 | | 0.3356 | 100 | 0.2092 | | 0.4027 | 120 | 0.1850 | | 0.4698 | 140 | 0.1616 | | 0.5369 | 160 | 0.1554 | | 0.6040 | 180 | 0.1499 | | 0.6711 | 200 | 0.1600 | | 0.7383 | 220 | 0.1278 | | 0.8054 | 240 | 0.1107 | | 0.8725 | 260 | 0.1230 | | 0.9396 | 280 | 0.1177 | | 1.0067 | 300 | 0.0924 | | 1.0738 | 320 | 0.0679 | | 1.1409 | 340 | 0.0665 | | 1.2081 | 360 | 0.0771 | | 1.2752 | 380 | 0.0646 | | 1.3423 | 400 | 0.0757 | | 1.4094 | 420 | 0.0728 | | 1.4765 | 440 | 0.0767 | | 1.5436 | 460 | 0.0732 | | 1.6107 | 480 | 0.0615 | | 1.6779 | 500 | 0.0639 | | 1.7450 | 520 | 0.0576 | | 1.8121 | 540 | 0.0686 | | 1.8792 | 560 | 0.0585 | | 1.9463 | 580 | 0.0655 | | 2.0134 | 600 | 0.0615 | | 2.0805 | 620 | 0.0430 | | 2.1477 | 640 | 0.0376 | | 2.2148 | 660 | 0.0377 | | 2.2819 | 680 | 0.0384 | | 2.3490 | 700 | 0.0393 | | 2.4161 | 720 | 0.0365 | | 2.4832 | 740 | 0.0421 | | 2.5503 | 760 | 0.0367 | | 2.6174 | 780 | 0.0433 | | 2.6846 | 800 | 0.0360 | | 2.7517 | 820 | 0.0363 | | 2.8188 | 840 | 0.0370 | | 2.8859 | 860 | 0.0295 | | 2.9530 | 880 | 0.0327 | | 3.0201 | 900 | 0.0321 | | 3.0872 | 920 | 0.0318 | | 3.1544 | 940 | 0.0249 | | 3.2215 | 960 | 0.0248 | | 3.2886 | 980 | 0.0236 | | 3.3557 | 1000 | 0.0249 | | 3.4228 | 1020 | 0.0332 | | 3.4899 | 1040 | 0.0298 | | 3.5570 | 1060 | 0.0283 | | 3.6242 | 1080 | 0.0261 | | 3.6913 | 1100 | 0.0346 | | 3.7584 | 1120 | 0.0270 | | 3.8255 | 1140 | 0.0284 | | 3.8926 | 1160 | 0.0321 | | 3.9597 | 1180 | 0.0267 | ### Training Time - **Training**: 5.0 minutes ### Framework Versions - Python: 3.12.3 - Sentence Transformers: 6.1.0 - Transformers: 5.19.0 - PyTorch: 2.13.0+cu129 - Accelerate: 1.15.0 - Datasets: 5.1.0 - Tokenizers: 0.23.2 ## Additional Resources - [Training and Finetuning Embedding Models with Sentence Transformers](https://huggingface.co/blog/train-sentence-transformers): the end-to-end guide for training or finetuning Sentence Transformer models. - [Introduction to Matryoshka Embedding Models](https://huggingface.co/blog/matryoshka): variable-size embeddings that can be truncated with minimal quality loss. - [Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval](https://huggingface.co/blog/embedding-quantization): post-training compression of embedding vectors. - [Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/multimodal-sentence-transformers): use text, image, audio, and video models through the same API. - [Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/train-multimodal-sentence-transformers): train multimodal embedding models, with a Visual Document Retrieval walkthrough. ## Citation ### BibTeX #### Sentence Transformers ```bibtex @inproceedings{reimers-2019-sentence-bert, title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", author = "Reimers, Nils and Gurevych, Iryna", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing", month = "11", year = "2019", publisher = "Association for Computational Linguistics", url = "https://arxiv.org/abs/1908.10084", } ``` #### CachedMultipleNegativesRankingLoss ```bibtex @misc{gao2021scaling, title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup}, author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan}, year={2021}, eprint={2101.06983}, archivePrefix={arXiv}, primaryClass={cs.LG} } ``` #### MultipleNegativesRankingLoss ```bibtex @misc{oord2019representationlearningcontrastivepredictive, title={Representation Learning with Contrastive Predictive Coding}, author={Aaron van den Oord and Yazhe Li and Oriol Vinyals}, year={2019}, eprint={1807.03748}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/1807.03748}, } ```