sentence-transformers/all-nli
Viewer • Updated • 2.86M • 3.82k • 54
How to use sparse-encoder/example-splade-cocondenser-ensembledistil-nli with sentence-transformers:
from sentence_transformers import SparseEncoder
model = SparseEncoder("sparse-encoder/example-splade-cocondenser-ensembledistil-nli")
queries = ["Which planet is known as the Red Planet?"]
documents = [
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)This is a SPLADE Sparse Encoder model finetuned from naver/splade-cocondenser-ensembledistil on the all-nli dataset using the sentence-transformers library. It maps sentences & paragraphs to a 30522-dimensional sparse vector space and can be used for semantic search and sparse retrieval.
SparseEncoder(
(0): MLMTransformer({'max_seq_length': 256, 'do_lower_case': False}) with MLMTransformer model: BertForMaskedLM
(1): SpladePooling({'pooling_strategy': 'max', 'activation_function': 'relu', 'word_embedding_dimension': 30522})
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SparseEncoder
# Download from the 🤗 Hub
model = SparseEncoder("arthurbresnu/example-splade-cocondenser-ensembledistil-nli")
# Run inference
sentences = [
'A man is sitting in on the side of the street with brass pots.',
'A man is playing with the brass pots',
'A group of adults are swimming at the beach.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# (3, 30522)
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
sts-dev and sts-testSparseEmbeddingSimilarityEvaluator| Metric | sts-dev | sts-test |
|---|---|---|
| pearson_cosine | 0.8554 | 0.8223 |
| spearman_cosine | 0.8486 | 0.8068 |
| active_dims | 99.1247 | 95.4228 |
| sparsity_ratio | 0.9968 | 0.9969 |
sentence1, sentence2, and score| sentence1 | sentence2 | score | |
|---|---|---|---|
| type | string | string | float |
| details |
|
|
|
| sentence1 | sentence2 | score |
|---|---|---|
A person on a horse jumps over a broken down airplane. |
A person is training his horse for a competition. |
0.5 |
A person on a horse jumps over a broken down airplane. |
A person is at a diner, ordering an omelette. |
0.0 |
A person on a horse jumps over a broken down airplane. |
A person is outdoors, on a horse. |
1.0 |
SpladeLoss with these parameters:{
"loss": "SparseMultipleNegativesRankingLoss(scale=1, similarity_fct='dot_score')",
"lambda_corpus": 0.003
}
sentence1, sentence2, and score| sentence1 | sentence2 | score | |
|---|---|---|---|
| type | string | string | float |
| details |
|
|
|
| sentence1 | sentence2 | score |
|---|---|---|
Two women are embracing while holding to go packages. |
The sisters are hugging goodbye while holding to go packages after just eating lunch. |
0.5 |
Two women are embracing while holding to go packages. |
Two woman are holding packages. |
1.0 |
Two women are embracing while holding to go packages. |
The men are fighting outside a deli. |
0.0 |
SpladeLoss with these parameters:{
"loss": "SparseMultipleNegativesRankingLoss(scale=1, similarity_fct='dot_score')",
"lambda_corpus": 0.003
}
eval_strategy: stepsper_device_train_batch_size: 16per_device_eval_batch_size: 16learning_rate: 4e-06num_train_epochs: 1bf16: Trueload_best_model_at_end: Truebatch_sampler: no_duplicatesoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 16per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 4e-06weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 1max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Truefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Trueignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}tp_size: 0fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters: auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional| Epoch | Step | Training Loss | Validation Loss | sts-dev_spearman_cosine | sts-test_spearman_cosine |
|---|---|---|---|---|---|
| -1 | -1 | - | - | 0.8366 | - |
| 0.032 | 20 | 1.0832 | - | - | - |
| 0.064 | 40 | 0.8212 | - | - | - |
| 0.096 | 60 | 0.796 | - | - | - |
| 0.128 | 80 | 0.7953 | - | - | - |
| 0.16 | 100 | 0.7574 | - | - | - |
| 0.192 | 120 | 0.6197 | 0.6750 | 0.8443 | - |
| 0.224 | 140 | 0.7125 | - | - | - |
| 0.256 | 160 | 0.817 | - | - | - |
| 0.288 | 180 | 0.7309 | - | - | - |
| 0.32 | 200 | 0.639 | - | - | - |
| 0.352 | 220 | 0.6873 | - | - | - |
| 0.384 | 240 | 0.6973 | 0.6253 | 0.8471 | - |
| 0.416 | 260 | 0.7197 | - | - | - |
| 0.448 | 280 | 0.5894 | - | - | - |
| 0.48 | 300 | 0.6682 | - | - | - |
| 0.512 | 320 | 0.6064 | - | - | - |
| 0.544 | 340 | 0.648 | - | - | - |
| 0.576 | 360 | 0.6344 | 0.6071 | 0.8483 | - |
| 0.608 | 380 | 0.5742 | - | - | - |
| 0.64 | 400 | 0.4962 | - | - | - |
| 0.672 | 420 | 0.4863 | - | - | - |
| 0.704 | 440 | 0.5547 | - | - | - |
| 0.736 | 460 | 0.6097 | - | - | - |
| 0.768 | 480 | 0.6307 | 0.6027 | 0.8471 | - |
| 0.8 | 500 | 0.6226 | - | - | - |
| 0.832 | 520 | 0.6607 | - | - | - |
| 0.864 | 540 | 0.526 | - | - | - |
| 0.896 | 560 | 0.6036 | - | - | - |
| 0.928 | 580 | 0.5897 | - | - | - |
| 0.96 | 600 | 0.6395 | 0.5892 | 0.8486 | - |
| 0.992 | 620 | 0.6069 | - | - | - |
| -1 | -1 | - | - | - | 0.8068 |
Carbon emissions were measured using CodeCarbon.
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{formal2022distillationhardnegativesampling,
title={From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective},
author={Thibault Formal and Carlos Lassance and Benjamin Piwowarski and Stéphane Clinchant},
year={2022},
eprint={2205.04733},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2205.04733},
}
@misc{henderson2017efficient,
title={Efficient Natural Language Response Suggestion for Smart Reply},
author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
year={2017},
eprint={1705.00652},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
@article{paria2020minimizing,
title={Minimizing flops to learn efficient sparse representations},
author={Paria, Biswajit and Yeh, Chih-Kuan and Yen, Ian EH and Xu, Ning and Ravikumar, Pradeep and P{'o}czos, Barnab{'a}s},
journal={arXiv preprint arXiv:2004.05665},
year={2020}
}
Base model
naver/splade-cocondenser-ensembledistil