UncleanCode's picture
Upload via Autoresearch export_data
b935288 verified
|
Raw History Blame Contribute Delete
17.2 kB
metadata
tags:
  - sentence-transformers
  - sentence-similarity
  - feature-extraction
  - dense
  - generated_from_trainer
  - dataset_size:76150
  - loss:CachedMultipleNegativesRankingLoss
  - loss:MultipleNegativesRankingLoss
base_model: sentence-transformers/LaBSE
widget:
  - source_sentence: po pelu ni akoko pelu awon
    sentences:
      - About time with them too.
      - 'Us: We''re just like stars!'
      - It can get you to the next day.
  - source_sentence: Ki ló n ṣẹlẹ / Ki lo n shele?
    sentences:
      - Sure this time it's fine.
      - What's going on/happened?
      - (I've got something in my eye!
  - source_sentence: ban ga laihi gare su int mm
    sentences:
      - '"Cities have been paralyzed"'
      - I wouldn't blame them. (NM)
      - Inside, there are no paths.
  - source_sentence: '"A cikin gõnaki da marẽmari."'
    sentences:
      - How Many Days Are In A 2020?
      - 'And they would say: "Our Lord!'
      - —amid gardens and springs,
  - source_sentence: Mo ti ri pe ninu ara mi ."
    sentences:
      - Bring my Soul out of Prison.
      - I've found it within myself'."
      - I looked and couldn't believe it!
pipeline_tag: sentence-similarity
library_name: sentence-transformers

SentenceTransformer based on sentence-transformers/LaBSE

This is a sentence-transformers model finetuned from sentence-transformers/LaBSE. It maps inputs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: sentence-transformers/LaBSE
  • Maximum Sequence Length: 128 tokens
  • Output Dimensionality: 768 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'cls', 'include_prompt': True})
  (2): Dense({'in_features': 768, 'out_features': 768, 'bias': True, 'activation_function': 'torch.nn.modules.activation.Tanh', 'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
  (3): Normalize({'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'Mo ti ri pe ninu ara mi ."',
    'I\'ve found it within myself\'."',
    'Bring my Soul out of Prison.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8338, 0.0731],
#         [0.8338, 1.0000, 0.1770],
#         [0.0731, 0.1770, 1.0000]])

Training Details

Training Dataset

Unnamed Dataset

  • Size: 76,150 training samples
  • Columns: anchor and positive
  • Approximate statistics based on the first 100 samples:
    anchor positive
    type string string
    modality text text
    details
    • min: 9 tokens
    • mean: 36.91 tokens
    • max: 128 tokens
    • min: 9 tokens
    • mean: 35.99 tokens
    • max: 128 tokens
  • Samples:
    anchor positive
    Ilé Ẹjọ́ Gíga Jù Lọ Nílẹ̀ Korea á lè lo ìdájọ́ tí Ilé Ẹjọ́ yìí ṣe nínú ọ̀rọ̀ ọ̀kọ̀ọ̀kan àwọn tí ẹ̀rí ọkàn wọn ò jẹ́ kí wọ́n ṣiṣẹ́ ológun. The Constitutional Court’s decision now opens the door for the Supreme Court of Korea to apply this ruling to specific cases involving conscientious objectors. Hundreds of thousands of people were evacuated, a process that proved to be especially complicated because of government-mandated physical distancing.
    "Wanda Ya sanya muku ƙasa shimfiɗa, kuma Ya shigar muku da hanyõyi a cikinta, kuma Ya saukar da ruwa daga sama." Who has made earth for you like a bed (spread out); and has opened roads (ways and paths etc.) for you therein; and has sent down water (rain) from the sky.
    Ìwọ ni Èlíjà bí?"+ Ó sì wí pé: "Èmi kọ́." Are you Elijah?" and he says, "I am not."
  • Loss: CachedMultipleNegativesRankingLoss with these parameters:
    {
        "scale": 30.0,
        "similarity_fct": "cos_sim",
        "mini_batch_size": 128,
        "mini_batch_num_tokens": null,
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 256
  • num_train_epochs: 4.0
  • learning_rate: 2e-05
  • lr_scheduler_type: cosine
  • warmup_steps: 0.1
  • bf16: True
  • dataloader_num_workers: 4
  • batch_sampler: no_duplicates

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 256
  • num_train_epochs: 4.0
  • max_steps: -1
  • learning_rate: 2e-05
  • lr_scheduler_type: cosine
  • lr_scheduler_kwargs: None
  • warmup_steps: 0.1
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: True
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 8
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 4
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • dataloader_multiprocessing_context: None
  • dataloader_in_order: True
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: no_duplicates
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}
  • warmup_ratio: None

Training Logs

Epoch Step Training Loss
0.0671 20 0.4747
0.1342 40 0.3165
0.2013 60 0.2757
0.2685 80 0.2269
0.3356 100 0.2092
0.4027 120 0.1850
0.4698 140 0.1616
0.5369 160 0.1554
0.6040 180 0.1499
0.6711 200 0.1600
0.7383 220 0.1278
0.8054 240 0.1107
0.8725 260 0.1230
0.9396 280 0.1177
1.0067 300 0.0924
1.0738 320 0.0679
1.1409 340 0.0665
1.2081 360 0.0771
1.2752 380 0.0646
1.3423 400 0.0757
1.4094 420 0.0728
1.4765 440 0.0767
1.5436 460 0.0732
1.6107 480 0.0615
1.6779 500 0.0639
1.7450 520 0.0576
1.8121 540 0.0686
1.8792 560 0.0585
1.9463 580 0.0655
2.0134 600 0.0615
2.0805 620 0.0430
2.1477 640 0.0376
2.2148 660 0.0377
2.2819 680 0.0384
2.3490 700 0.0393
2.4161 720 0.0365
2.4832 740 0.0421
2.5503 760 0.0367
2.6174 780 0.0433
2.6846 800 0.0360
2.7517 820 0.0363
2.8188 840 0.0370
2.8859 860 0.0295
2.9530 880 0.0327
3.0201 900 0.0321
3.0872 920 0.0318
3.1544 940 0.0249
3.2215 960 0.0248
3.2886 980 0.0236
3.3557 1000 0.0249
3.4228 1020 0.0332
3.4899 1040 0.0298
3.5570 1060 0.0283
3.6242 1080 0.0261
3.6913 1100 0.0346
3.7584 1120 0.0270
3.8255 1140 0.0284
3.8926 1160 0.0321
3.9597 1180 0.0267

Training Time

  • Training: 5.0 minutes

Framework Versions

  • Python: 3.12.3
  • Sentence Transformers: 6.1.0
  • Transformers: 5.19.0
  • PyTorch: 2.13.0+cu129
  • Accelerate: 1.15.0
  • Datasets: 5.1.0
  • Tokenizers: 0.23.2

Additional Resources

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

CachedMultipleNegativesRankingLoss

@misc{gao2021scaling,
    title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
    author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
    year={2021},
    eprint={2101.06983},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}