Sentence Similarity
sentence-transformers
Safetensors
bert
embeddings
cross-lingual
multilingual
igbo
hausa
yoruba
information-retrieval
semantic-search
text-embeddings-inference
Instructions to use Modularcomputing/Native-Bird with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use Modularcomputing/Native-Bird with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Modularcomputing/Native-Bird") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
|
Download models/labse-ig-ha-yo/README.md from Modularcomputing/Native-Bird: direct link, hf CLI and curl.
- Browser
- Download file 17.2 kB
-
https://huggingface.co/Modularcomputing/Native-Bird/resolve/main/models/labse-ig-ha-yo/README.md
- Command line
-
hf download hf://Modularcomputing/Native-Bird/models/labse-ig-ha-yo/README.md
-
curl -L -o README.md https://huggingface.co/Modularcomputing/Native-Bird/resolve/main/models/labse-ig-ha-yo/README.md
17.2 kB
| tags: | |
| - sentence-transformers | |
| - sentence-similarity | |
| - feature-extraction | |
| - dense | |
| - generated_from_trainer | |
| - dataset_size:76150 | |
| - loss:CachedMultipleNegativesRankingLoss | |
| - loss:MultipleNegativesRankingLoss | |
| base_model: sentence-transformers/LaBSE | |
| widget: | |
| - source_sentence: po pelu ni akoko pelu awon | |
| sentences: | |
| - About time with them too. | |
| - 'Us: We''re just like stars!' | |
| - It can get you to the next day. | |
| - source_sentence: Ki ló n ṣẹlẹ / Ki lo n shele? | |
| sentences: | |
| - Sure this time it's fine. | |
| - What's going on/happened? | |
| - (I've got something in my eye! | |
| - source_sentence: ban ga laihi gare su int mm | |
| sentences: | |
| - '"Cities have been paralyzed"' | |
| - I wouldn't blame them. (NM) | |
| - Inside, there are no paths. | |
| - source_sentence: '"A cikin gõnaki da marẽmari."' | |
| sentences: | |
| - How Many Days Are In A 2020? | |
| - 'And they would say: "Our Lord!' | |
| - —amid gardens and springs, | |
| - source_sentence: Mo ti ri pe ninu ara mi ." | |
| sentences: | |
| - Bring my Soul out of Prison. | |
| - I've found it within myself'." | |
| - I looked and couldn't believe it! | |
| pipeline_tag: sentence-similarity | |
| library_name: sentence-transformers | |
| # SentenceTransformer based on sentence-transformers/LaBSE | |
| This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [sentence-transformers/LaBSE](https://huggingface.co/sentence-transformers/LaBSE). It maps inputs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more. | |
| ## Model Details | |
| ### Model Description | |
| - **Model Type:** Sentence Transformer | |
| - **Base model:** [sentence-transformers/LaBSE](https://huggingface.co/sentence-transformers/LaBSE) <!-- at revision 836121a0533e5664b21c7aacc5d22951f2b8b25b --> | |
| - **Maximum Sequence Length:** 128 tokens | |
| - **Output Dimensionality:** 768 dimensions | |
| - **Similarity Function:** Cosine Similarity | |
| - **Supported Modality:** Text | |
| <!-- - **Training Dataset:** Unknown --> | |
| <!-- - **Language:** Unknown --> | |
| <!-- - **License:** Unknown --> | |
| ### Model Sources | |
| - **Documentation:** [Sentence Transformers Documentation](https://sbert.net) | |
| - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers) | |
| - **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers) | |
| ### Full Model Architecture | |
| ``` | |
| SentenceTransformer( | |
| (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'}) | |
| (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'cls', 'include_prompt': True}) | |
| (2): Dense({'in_features': 768, 'out_features': 768, 'bias': True, 'activation_function': 'torch.nn.modules.activation.Tanh', 'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'}) | |
| (3): Normalize({'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'}) | |
| ) | |
| ``` | |
| ## Usage | |
| ### Direct Usage (Sentence Transformers) | |
| First install the Sentence Transformers library: | |
| ```bash | |
| pip install -U sentence-transformers | |
| ``` | |
| Then you can load this model and run inference. | |
| ```python | |
| from sentence_transformers import SentenceTransformer | |
| # Download from the 🤗 Hub | |
| model = SentenceTransformer("sentence_transformers_model_id") | |
| # Run inference | |
| sentences = [ | |
| 'Mo ti ri pe ninu ara mi ."', | |
| 'I\'ve found it within myself\'."', | |
| 'Bring my Soul out of Prison.', | |
| ] | |
| embeddings = model.encode(sentences) | |
| print(embeddings.shape) | |
| # [3, 768] | |
| # Get the similarity scores for the embeddings | |
| similarities = model.similarity(embeddings, embeddings) | |
| print(similarities) | |
| # tensor([[1.0000, 0.8338, 0.0731], | |
| # [0.8338, 1.0000, 0.1770], | |
| # [0.0731, 0.1770, 1.0000]]) | |
| ``` | |
| <!-- | |
| ### Direct Usage (Transformers) | |
| <details><summary>Click to see the direct usage in Transformers</summary> | |
| </details> | |
| --> | |
| <!-- | |
| ### Downstream Usage (Sentence Transformers) | |
| You can finetune this model on your own dataset. | |
| <details><summary>Click to expand</summary> | |
| </details> | |
| --> | |
| <!-- | |
| ### Out-of-Scope Use | |
| *List how the model may foreseeably be misused and address what users ought not to do with the model.* | |
| --> | |
| <!-- | |
| ## Bias, Risks and Limitations | |
| *What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.* | |
| --> | |
| <!-- | |
| ### Recommendations | |
| *What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.* | |
| --> | |
| ## Training Details | |
| ### Training Dataset | |
| #### Unnamed Dataset | |
| * Size: 76,150 training samples | |
| * Columns: <code>anchor</code> and <code>positive</code> | |
| * Approximate statistics based on the first 100 samples: | |
| | | anchor | positive | | |
| |:---------|:-----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | |
| | type | string | string | | |
| | modality | text | text | | |
| | details | <ul><li>min: 9 tokens</li><li>mean: 36.91 tokens</li><li>max: 128 tokens</li></ul> | <ul><li>min: 9 tokens</li><li>mean: 35.99 tokens</li><li>max: 128 tokens</li></ul> | | |
| * Samples: | |
| | anchor | positive | | |
| |:-------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | |
| | <code>Ilé Ẹjọ́ Gíga Jù Lọ Nílẹ̀ Korea á lè lo ìdájọ́ tí Ilé Ẹjọ́ yìí ṣe nínú ọ̀rọ̀ ọ̀kọ̀ọ̀kan àwọn tí ẹ̀rí ọkàn wọn ò jẹ́ kí wọ́n ṣiṣẹ́ ológun.</code> | <code>The Constitutional Court’s decision now opens the door for the Supreme Court of Korea to apply this ruling to specific cases involving conscientious objectors. Hundreds of thousands of people were evacuated, a process that proved to be especially complicated because of government-mandated physical distancing.</code> | | |
| | <code>"Wanda Ya sanya muku ƙasa shimfiɗa, kuma Ya shigar muku da hanyõyi a cikinta, kuma Ya saukar da ruwa daga sama."</code> | <code>Who has made earth for you like a bed (spread out); and has opened roads (ways and paths etc.) for you therein; and has sent down water (rain) from the sky.</code> | | |
| | <code>Ìwọ ni Èlíjà bí?"+ Ó sì wí pé: "Èmi kọ́."</code> | <code>Are you Elijah?" and he says, "I am not."</code> | | |
| * Loss: [<code>CachedMultipleNegativesRankingLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#cachedmultiplenegativesrankingloss) with these parameters: | |
| ```json | |
| { | |
| "scale": 30.0, | |
| "similarity_fct": "cos_sim", | |
| "mini_batch_size": 128, | |
| "mini_batch_num_tokens": null, | |
| "gather_across_devices": false, | |
| "directions": [ | |
| "query_to_doc" | |
| ], | |
| "partition_mode": "joint", | |
| "hardness_mode": null, | |
| "hardness_strength": 0.0 | |
| } | |
| ``` | |
| ### Training Hyperparameters | |
| #### Non-Default Hyperparameters | |
| - `per_device_train_batch_size`: 256 | |
| - `num_train_epochs`: 4.0 | |
| - `learning_rate`: 2e-05 | |
| - `lr_scheduler_type`: cosine | |
| - `warmup_steps`: 0.1 | |
| - `bf16`: True | |
| - `dataloader_num_workers`: 4 | |
| - `batch_sampler`: no_duplicates | |
| #### All Hyperparameters | |
| <details><summary>Click to expand</summary> | |
| - `per_device_train_batch_size`: 256 | |
| - `num_train_epochs`: 4.0 | |
| - `max_steps`: -1 | |
| - `learning_rate`: 2e-05 | |
| - `lr_scheduler_type`: cosine | |
| - `lr_scheduler_kwargs`: None | |
| - `warmup_steps`: 0.1 | |
| - `optim`: adamw_torch_fused | |
| - `optim_args`: None | |
| - `weight_decay`: 0.0 | |
| - `adam_beta1`: 0.9 | |
| - `adam_beta2`: 0.999 | |
| - `adam_epsilon`: 1e-08 | |
| - `optim_target_modules`: None | |
| - `gradient_accumulation_steps`: 1 | |
| - `average_tokens_across_devices`: True | |
| - `max_grad_norm`: 1.0 | |
| - `label_smoothing_factor`: 0.0 | |
| - `bf16`: True | |
| - `fp16`: False | |
| - `bf16_full_eval`: False | |
| - `fp16_full_eval`: False | |
| - `tf32`: None | |
| - `gradient_checkpointing`: False | |
| - `gradient_checkpointing_kwargs`: None | |
| - `torch_compile`: False | |
| - `torch_compile_backend`: None | |
| - `torch_compile_mode`: None | |
| - `use_liger_kernel`: False | |
| - `liger_kernel_config`: None | |
| - `use_cache`: False | |
| - `neftune_noise_alpha`: None | |
| - `torch_empty_cache_steps`: None | |
| - `auto_find_batch_size`: False | |
| - `log_on_each_node`: True | |
| - `logging_nan_inf_filter`: True | |
| - `include_num_input_tokens_seen`: no | |
| - `log_level`: passive | |
| - `log_level_replica`: warning | |
| - `disable_tqdm`: False | |
| - `project`: huggingface | |
| - `trackio_space_id`: None | |
| - `trackio_bucket_id`: None | |
| - `trackio_static_space_id`: None | |
| - `per_device_eval_batch_size`: 8 | |
| - `prediction_loss_only`: True | |
| - `eval_on_start`: False | |
| - `eval_do_concat_batches`: True | |
| - `eval_use_gather_object`: False | |
| - `eval_accumulation_steps`: None | |
| - `include_for_metrics`: [] | |
| - `batch_eval_metrics`: False | |
| - `save_only_model`: False | |
| - `save_on_each_node`: False | |
| - `enable_jit_checkpoint`: False | |
| - `push_to_hub`: False | |
| - `hub_private_repo`: None | |
| - `hub_model_id`: None | |
| - `hub_strategy`: every_save | |
| - `hub_always_push`: False | |
| - `hub_revision`: None | |
| - `load_best_model_at_end`: False | |
| - `ignore_data_skip`: False | |
| - `restore_callback_states_from_checkpoint`: False | |
| - `full_determinism`: False | |
| - `seed`: 42 | |
| - `data_seed`: None | |
| - `use_cpu`: False | |
| - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None} | |
| - `parallelism_config`: None | |
| - `dataloader_drop_last`: False | |
| - `dataloader_num_workers`: 4 | |
| - `dataloader_pin_memory`: True | |
| - `dataloader_persistent_workers`: False | |
| - `dataloader_prefetch_factor`: None | |
| - `dataloader_multiprocessing_context`: None | |
| - `dataloader_in_order`: True | |
| - `remove_unused_columns`: True | |
| - `label_names`: None | |
| - `train_sampling_strategy`: random | |
| - `length_column_name`: length | |
| - `ddp_find_unused_parameters`: None | |
| - `ddp_bucket_cap_mb`: None | |
| - `ddp_broadcast_buffers`: False | |
| - `ddp_static_graph`: None | |
| - `ddp_backend`: None | |
| - `ddp_timeout`: 1800 | |
| - `fsdp`: None | |
| - `fsdp_config`: None | |
| - `deepspeed`: None | |
| - `debug`: [] | |
| - `skip_memory_metrics`: True | |
| - `do_predict`: False | |
| - `resume_from_checkpoint`: None | |
| - `local_rank`: -1 | |
| - `prompts`: None | |
| - `batch_sampler`: no_duplicates | |
| - `multi_dataset_batch_sampler`: proportional | |
| - `router_mapping`: {} | |
| - `learning_rate_mapping`: {} | |
| - `warmup_ratio`: None | |
| </details> | |
| ### Training Logs | |
| | Epoch | Step | Training Loss | | |
| |:------:|:----:|:-------------:| | |
| | 0.0671 | 20 | 0.4747 | | |
| | 0.1342 | 40 | 0.3165 | | |
| | 0.2013 | 60 | 0.2757 | | |
| | 0.2685 | 80 | 0.2269 | | |
| | 0.3356 | 100 | 0.2092 | | |
| | 0.4027 | 120 | 0.1850 | | |
| | 0.4698 | 140 | 0.1616 | | |
| | 0.5369 | 160 | 0.1554 | | |
| | 0.6040 | 180 | 0.1499 | | |
| | 0.6711 | 200 | 0.1600 | | |
| | 0.7383 | 220 | 0.1278 | | |
| | 0.8054 | 240 | 0.1107 | | |
| | 0.8725 | 260 | 0.1230 | | |
| | 0.9396 | 280 | 0.1177 | | |
| | 1.0067 | 300 | 0.0924 | | |
| | 1.0738 | 320 | 0.0679 | | |
| | 1.1409 | 340 | 0.0665 | | |
| | 1.2081 | 360 | 0.0771 | | |
| | 1.2752 | 380 | 0.0646 | | |
| | 1.3423 | 400 | 0.0757 | | |
| | 1.4094 | 420 | 0.0728 | | |
| | 1.4765 | 440 | 0.0767 | | |
| | 1.5436 | 460 | 0.0732 | | |
| | 1.6107 | 480 | 0.0615 | | |
| | 1.6779 | 500 | 0.0639 | | |
| | 1.7450 | 520 | 0.0576 | | |
| | 1.8121 | 540 | 0.0686 | | |
| | 1.8792 | 560 | 0.0585 | | |
| | 1.9463 | 580 | 0.0655 | | |
| | 2.0134 | 600 | 0.0615 | | |
| | 2.0805 | 620 | 0.0430 | | |
| | 2.1477 | 640 | 0.0376 | | |
| | 2.2148 | 660 | 0.0377 | | |
| | 2.2819 | 680 | 0.0384 | | |
| | 2.3490 | 700 | 0.0393 | | |
| | 2.4161 | 720 | 0.0365 | | |
| | 2.4832 | 740 | 0.0421 | | |
| | 2.5503 | 760 | 0.0367 | | |
| | 2.6174 | 780 | 0.0433 | | |
| | 2.6846 | 800 | 0.0360 | | |
| | 2.7517 | 820 | 0.0363 | | |
| | 2.8188 | 840 | 0.0370 | | |
| | 2.8859 | 860 | 0.0295 | | |
| | 2.9530 | 880 | 0.0327 | | |
| | 3.0201 | 900 | 0.0321 | | |
| | 3.0872 | 920 | 0.0318 | | |
| | 3.1544 | 940 | 0.0249 | | |
| | 3.2215 | 960 | 0.0248 | | |
| | 3.2886 | 980 | 0.0236 | | |
| | 3.3557 | 1000 | 0.0249 | | |
| | 3.4228 | 1020 | 0.0332 | | |
| | 3.4899 | 1040 | 0.0298 | | |
| | 3.5570 | 1060 | 0.0283 | | |
| | 3.6242 | 1080 | 0.0261 | | |
| | 3.6913 | 1100 | 0.0346 | | |
| | 3.7584 | 1120 | 0.0270 | | |
| | 3.8255 | 1140 | 0.0284 | | |
| | 3.8926 | 1160 | 0.0321 | | |
| | 3.9597 | 1180 | 0.0267 | | |
| ### Training Time | |
| - **Training**: 5.0 minutes | |
| ### Framework Versions | |
| - Python: 3.12.3 | |
| - Sentence Transformers: 6.1.0 | |
| - Transformers: 5.19.0 | |
| - PyTorch: 2.13.0+cu129 | |
| - Accelerate: 1.15.0 | |
| - Datasets: 5.1.0 | |
| - Tokenizers: 0.23.2 | |
| ## Additional Resources | |
| - [Training and Finetuning Embedding Models with Sentence Transformers](https://huggingface.co/blog/train-sentence-transformers): the end-to-end guide for training or finetuning Sentence Transformer models. | |
| - [Introduction to Matryoshka Embedding Models](https://huggingface.co/blog/matryoshka): variable-size embeddings that can be truncated with minimal quality loss. | |
| - [Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval](https://huggingface.co/blog/embedding-quantization): post-training compression of embedding vectors. | |
| - [Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/multimodal-sentence-transformers): use text, image, audio, and video models through the same API. | |
| - [Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers](https://huggingface.co/blog/train-multimodal-sentence-transformers): train multimodal embedding models, with a Visual Document Retrieval walkthrough. | |
| ## Citation | |
| ### BibTeX | |
| #### Sentence Transformers | |
| ```bibtex | |
| @inproceedings{reimers-2019-sentence-bert, | |
| title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", | |
| author = "Reimers, Nils and Gurevych, Iryna", | |
| booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing", | |
| month = "11", | |
| year = "2019", | |
| publisher = "Association for Computational Linguistics", | |
| url = "https://arxiv.org/abs/1908.10084", | |
| } | |
| ``` | |
| #### CachedMultipleNegativesRankingLoss | |
| ```bibtex | |
| @misc{gao2021scaling, | |
| title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup}, | |
| author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan}, | |
| year={2021}, | |
| eprint={2101.06983}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.LG} | |
| } | |
| ``` | |
| #### MultipleNegativesRankingLoss | |
| ```bibtex | |
| @misc{oord2019representationlearningcontrastivepredictive, | |
| title={Representation Learning with Contrastive Predictive Coding}, | |
| author={Aaron van den Oord and Yazhe Li and Oriol Vinyals}, | |
| year={2019}, | |
| eprint={1807.03748}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.LG}, | |
| url={https://arxiv.org/abs/1807.03748}, | |
| } | |
| ``` | |
| <!-- | |
| ## Glossary | |
| *Clearly define terms in order to be accessible across audiences.* | |
| --> | |
| <!-- | |
| ## Model Card Authors | |
| *Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.* | |
| --> | |
| <!-- | |
| ## Model Card Contact | |
| *Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.* | |
| --> |