TSDAE Insurance Embedding Model (bge-base)

A 768-dimensional English sentence-embedding model adapted to Indian insurance policy wordings, built by Turtlemint.

It starts from BAAI/bge-base-en-v1.5 and is further trained with TSDAE (Transformer-based Sequential Denoising Auto-Encoder), an unsupervised domain-adaptation method, on passages from insurance policy wording documents. The model also has an extended vocabulary: 195 insurance and medical abbreviations (for example AYUSH, ABHA, CABG, ABDM, APPD) are added as whole tokens, so they are no longer split into sub-word pieces.

The goal is embeddings that capture the terms, clause structure and phrasing of policy documents (definitions, waiting periods, exclusions, sub-limits, claim procedures and so on) better than a general-purpose model does.

Model Details

Developed by Turtlemint (InsuranceGPT team)
Model type Sentence Transformer (BERT encoder)
Base model BAAI/bge-base-en-v1.5
Training method TSDAE, unsupervised denoising auto-encoder
Language English (Indian insurance domain)
Parameters ~110M
Embedding dimension 768
Max sequence length 512 tokens
Pooling CLS token, followed by L2 normalization
Similarity function Cosine
Vocabulary 30,522 (base) + 195 domain abbreviations = 30,717
License MIT (same as the base model)

Architecture

SentenceTransformer(
  (0): Transformer(BertModel, max_seq_length=512)
  (1): Pooling(embedding_dimension=768, pooling_mode='cls')
  (2): Normalize()
)

Usage

Sentence Transformers

pip install -U sentence-transformers
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("Turtlemint-India/TSDA_Embedding_model")

query = "What is the waiting period for pre-existing diseases?"
passages = [
    "Pre-existing Diseases will be covered after a waiting period of 36 months "
    "of continuous coverage since inception of the first policy with Us.",
    "Room rent is payable up to 1% of the Sum Insured per day.",
    "AYUSH treatment expenses are covered up to the Sum Insured for in-patient hospitalisation.",
]

q_emb = model.encode([query])
p_emb = model.encode(passages)

scores = model.similarity(q_emb, p_emb)   # cosine similarity
print(scores)
# tensor([[0.9881, 0.9719, 0.9798]])  -> the pre-existing-disease clause ranks first

No query or passage prefix/instruction is needed.

Hugging Face Transformers (without sentence-transformers)

import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel

tok = AutoTokenizer.from_pretrained("Turtlemint-India/TSDA_Embedding_model")
model = AutoModel.from_pretrained("Turtlemint-India/TSDA_Embedding_model").eval()

texts = ["Maternity benefit has a waiting period of 9 months."]
batch = tok(texts, padding=True, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
    out = model(**batch)
emb = F.normalize(out.last_hidden_state[:, 0], p=2, dim=1)   # CLS pooling + L2 norm

Important: Interpreting Similarity Scores

TSDAE is a reconstruction objective, not a contrastive one. Like most TSDAE-only models, this model produces cosine similarities in a narrow, high band; in the example above all scores fall between about 0.97 and 0.99.

  • Relative ranking is what matters. Use the scores to rank or re-rank candidates (top-k retrieval, nearest neighbours).
  • Don't use fixed absolute thresholds such as "relevant if score > 0.8"; nearly everything will pass. If you need a threshold, calibrate it on your own labelled data.
  • For best retrieval quality, we recommend treating this model as a domain-adapted starting point and fine-tuning it further with a contrastive objective (for example MultipleNegativesRankingLoss on question–clause pairs), which spreads the score distribution out.

Intended Uses

Recommended

  • Semantic search and retrieval over health and general insurance policy wordings, brochures and prospectuses (RAG pipelines).
  • Clustering, de-duplication and near-duplicate detection of policy clauses across insurers and products.
  • A domain-adapted backbone for supervised fine-tuning on insurance retrieval, classification or matching tasks.

Out of scope

  • Giving insurance, legal, medical or financial advice. Embeddings only measure textual similarity. They do not decide coverage, claim eligibility or the legal meaning of a clause.
  • Non-English text, and domains far from insurance (general-purpose performance may be lower than the base model's).
  • Using absolute similarity values as calibrated relevance probabilities (see above).

Training Details

Training Data

A private corpus of Indian insurance policy wording documents (mainly health insurance, plus some general insurance). The PDFs were parsed to Markdown and split into passage-level chunks:

  • 83 source JSONL files, ~14.2K raw chunks
  • Filtered to chunks of 10–150 words
  • 10,781 training and 220 evaluation passages
  • Content includes definitions, coverage and benefit clauses, waiting periods, exclusions, sub-limits and co-payments, claim and grievance procedures, UIN references and jurisdiction clauses.

The training corpus is not released.

Noise Function

TSDAE learns by reconstructing an original passage from a corrupted version of it, using only the sentence embedding. We used a domain-specific noise function:

Corruption Setting
Random token deletion 50% of tokens deleted
Number masking 50% of numeric values (amounts, percentages, days/months) replaced with a [NUM] placeholder
Clause shuffling enabled with probability 0.4

Example:

Noisy: 19. OF SUBROGATION It is hereby and that the terms and conditions in the or shall waive all their of or which they may have …

Original: 19. WAIVER OF SUBROGATION CLAUSE It is hereby agreed and understood that otherwise subject to the terms exclusions, provisions and conditions contained in the Policy or endorsed thereon, the Insurers shall waive all their rights of subrogation …

Vocabulary Extension

Before training, 195 insurance and medical abbreviations taken from a curated abbreviation dictionary were added to the tokenizer, and the embedding matrix was resized to match. The new embeddings were learned during TSDAE training.

Training Procedure

  • Loss: DenoisingAutoEncoderLoss (decoder initialised from BAAI/bge-base-en-v1.5, encoder–decoder weights tied)
  • Epochs: 15 (10,110 steps)
  • Batch size: 16
  • Learning rate: 3e-5, linear schedule, 10% warm-up
  • Optimizer: AdamW (fused), weight decay 0.0, max grad norm 1.0
  • Precision: bf16
  • Seed: 42
  • Framework versions: sentence-transformers 5.6.0, transformers 5.12.1, PyTorch 2.12.1 (CUDA 13.0)

Training Curve

The model is from the final checkpoint (step 10,110, epoch 15). Reconstruction loss on the held-out set:

Epoch Step Eval loss
0.74 500 5.901
2.97 2,000 3.465
5.19 3,500 2.604
7.42 5,000 2.227
9.64 6,500 2.028
11.87 8,000 1.903
14.09 9,500 1.776
15.00 10,110 1.756

Final training loss was about 1.42.

Evaluation

Only the denoising reconstruction loss shown above has been measured so far. No downstream retrieval benchmark (for example nDCG or Recall@k on labelled insurance question–clause pairs) has been run yet. Evaluate on your own task before using the model in production.

Limitations and Bias

  • Narrow score range: see Interpreting Similarity Scores.
  • Domain and geography: trained mainly on Indian (IRDAI-regulated) health insurance wordings. Terms from other markets, other lines of business (motor, life, marine) or other regulators may be represented less well.
  • Parsing artifacts: the source text came from PDF parsing and contains some noise (for example LaTeX-like fragments such as $^{\mathrm{TM}}$, table fragments and broken sentences). The model has learned from this noise.
  • Lower-casing: the tokenizer is uncased, so AYUSH and ayush map to the same token.
  • Not a decision system: embeddings must not be the sole basis for coverage, underwriting or claim decisions.

Citation

If you use this model, please cite the base model and the TSDAE method:

@inproceedings{wang-2021-TSDAE,
  title     = "TSDAE: Using Transformer-based Sequential Denoising Auto-Encoder for Unsupervised Sentence Embedding Learning",
  author    = "Wang, Kexin and Reimers, Nils and Gurevych, Iryna",
  booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2021",
  year      = "2021",
  pages     = "671--688",
  url       = "https://arxiv.org/abs/2104.06979",
}

@misc{bge_embedding,
  title         = {C-Pack: Packaged Resources To Advance General Chinese Embedding},
  author        = {Shitao Xiao and Zheng Liu and Peitian Zhang and Niklas Muennighoff},
  year          = {2023},
  eprint        = {2309.07597},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL}
}

@inproceedings{reimers-2019-sentence-bert,
  title     = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
  author    = "Reimers, Nils and Gurevych, Iryna",
  booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
  year      = "2019",
  url       = "https://arxiv.org/abs/1908.10084",
}

Contact

Maintained by the Turtlemint InsuranceGPT team. For questions or issues, please open a discussion on this repository.

Downloads last month
6
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Turtlemint-India/TSDA_Embedding_model

Finetuned
(495)
this model

Papers for Turtlemint-India/TSDA_Embedding_model