abelcetina/simcse-bert-base-snli-sup

A supervised SimCSE sentence encoder, trained from bert-base-uncased on a 100k-record subset of SNLI. Built for U2T02 at Universidad PolitΓ©cnica de YucatΓ‘n (Data Engineering, 9th Quarter Group B).

Team: Christian CarreΓ±o Β· RaΓΊl Cetina Β· Christopher QuiΓ±ones Β· Daniel GΓ³mez

Usage

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("abelcetina/simcse-bert-base-snli-sup")
embeddings = model.encode(["A man is playing a guitar.",
                           "Someone is performing music."])

Similarity is the cosine between normalised embeddings. No regressor on top.

Training data

A provided 100,000-record subset of the SNLI train split, from which two sets were built:

  • Unsupervised: 165,529 unique sentences.
  • Supervised: 33,351 (premise, entailment) pairs, of which 28.4% also carry a contradiction hypothesis for the same premise β€” the paper's hard negatives.

This model was trained on the supervised objective.

Recipe

Base model bert-base-uncased
Batch size 128
Learning rate 5e-05
Epochs 3
Max sequence length 32
Temperature Ο„ 0.05
Dropout 0.1
Pooling (train β†’ eval) cls_mlp β†’ cls_mlp
Seed 42
Hardware Tesla T4

In a contrastive objective the batch size is the number of in-batch negatives, so it is a property of the method rather than a throughput setting.

Evaluation β€” STS-B (Spearman Γ—100)

Embed both sentences, L2-normalise, cosine similarity, Spearman against the human scores. No regressor. The test split was read exactly once for this model.

Model dev test alignment uniformity source
Ours: unsup-SimCSE measured 75.49 66.51 0.3123 -2.8188 unsup_unsup_base_s42_20260930-154313
Ours: sup-SimCSE measured 82.14 77.82 0.1930 -3.0463 sup_sup_base_s42_20260930-155328
bert-base-uncased (mean pooling) measured 59.31 47.29 0.1948 -1.6497 assignment brief, Notes
SBERT-2019 (bert-base-nli-mean-tokens) measured 80.77 76.98 0.1929 -3.0487 assignment brief, Notes
unsup-SimCSE-BERT-base (paper) quoted 82.50 76.85 β€” β€” Gao et al. 2021, Table 3 (dev) / Table 5 (test)
sup-SimCSE-BERT-base (paper) quoted 86.20 84.25 β€” β€” Gao et al. 2021, Tables 3 and 7 (dev) / Table 5 (test)

Alignment and uniformity are both lower is better.

Limitations

  • Trained on image captions. SNLI premises are short (median 9 words) and concrete. STS-B is mixed-genre, so this model faces a domain shift the original SimCSE β€” trained on Wikipedia β€” did not.
  • Below the paper's figures, by design. The provided data is 6Γ— smaller than the paper's unsupervised corpus and 9.4Γ— smaller than its supervised one, and only 28.4% of our pairs carry a hard negative against its full triplets.
  • English only, and inherits bert-base-uncased's biases.
  • Not tuned for retrieval at scale or for any domain-specific corpus.

Citation

SimCSE: Gao, Yao & Chen, SimCSE: Simple Contrastive Learning of Sentence Embeddings, EMNLP 2021. https://arxiv.org/abs/2104.08821

Downloads last month
26
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for abelcetina/simcse-bert-base-snli-sup

Finetuned
(7016)
this model

Dataset used to train abelcetina/simcse-bert-base-snli-sup

Paper for abelcetina/simcse-bert-base-snli-sup