Instructions to use bugBug04S/legal-embed-modernbert-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use bugBug04S/legal-embed-modernbert-v2 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("bugBug04S/legal-embed-modernbert-v2") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
legal-embed-modernbert-v2
Embedding model for legal retrieval: matches plain-English legal questions to the passages that answer them, and clause-type descriptions to the contract clauses that match. Fine-tuned from nomic-ai/modernbert-embed-base.
Use it for: legal Q&A search, contract clause lookup, RAG over legal corpora. Supports Matryoshka embeddings — truncate to 512/256/128/64 dims for cheaper search at modest quality cost.
v2 changes over v1: hard-negative mining, contract-clause training data, Australian legal QA, and a general-domain slice to limit domain drift.
Evaluation
All three models evaluated under identical settings (max_seq_length=256,
nomic prefixes applied). Best per column in bold.
Legal Q&A — 1,000 held-out Law StackExchange pairs, never seen in training, retrieval among 1,000 candidates:
| Model | Accuracy@1 | Recall@10 | MRR |
|---|---|---|---|
| legal-embed-modernbert-v2 | 0.904 | 0.994 | 0.938 |
| legal-embed-modernbert-v1 | 0.893 | 0.994 | 0.931 |
| nomic-ai/modernbert-embed-base | 0.800 | 0.948 | 0.856 |
Contractual Clause Retrieval (isaacus) — 45 clause types, 90 clauses:
| Model | Accuracy@1 | Recall@10 | MRR |
|---|---|---|---|
| legal-embed-modernbert-v2 | 0.844 | 0.978 | 0.900 |
| legal-embed-modernbert-v1 | 0.622 | 0.889 | 0.708 |
| nomic-ai/modernbert-embed-base | 0.644 | 0.933 | 0.733 |
LegalBench Consumer Contracts QA (mteb):
| Model | Accuracy@1 | Recall@10 | MRR |
|---|---|---|---|
| legal-embed-modernbert-v1 | 0.556 | 0.922 | 0.672 |
| legal-embed-modernbert-v2 | 0.528 | 0.902 | 0.643 |
| nomic-ai/modernbert-embed-base | 0.525 | 0.907 | 0.650 |
How to read these results
- Clause retrieval: the large gain reflects training on CUAD clause-type data. The evaluation texts are from a different source, but the task format (clause-type query → clause text) matches the training data, so this is not a pure zero-shot result. Judge it as "trained for this task and does it well," not "generalises to unseen tasks."
- Consumer contracts QA: v2 is marginally below v1 and level with the base. This benchmark asks natural questions about contract content rather than matching clause types, and v2's broader training mix traded a little of this for the clause-retrieval gain. If consumer-contract QA is your primary use case, evaluate v1 as well.
- Numbers here are not comparable to those on the v1 model card, which used a different sequence-length setting for the baseline.
Usage
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("bugBug04S/legal-embed-modernbert-v2")
query = "search_query: can the client terminate early without cause?"
docs = [
"search_document: Either party may terminate this Agreement upon thirty (30) days prior written notice.",
"search_document: The Licensee shall indemnify and hold harmless the Licensor against all claims.",
]
q = model.encode(query, normalize_embeddings=True)
d = model.encode(docs, normalize_embeddings=True)
print(q @ d.T) # higher = more relevant
Required: prepend search_query: to queries and search_document: to
passages. The model is trained with these prefixes and degrades noticeably
without them.
Training
- Base: nomic-ai/modernbert-embed-base (149M params)
- 24,413 triplets (anchor, positive, mined hard negative):
- ~16K Law StackExchange Q&A pairs (CC-BY-SA 4.0)
- ~2.1K Australian legal QA (isaacus)
- ~2.4K CUAD contract clauses, capped at 60 per clause type (CC-BY 4.0)
- 4K general-domain pairs (Natural Questions) to limit domain drift
- Hard negatives mined with v1; mean positive/negative similarity gap 0.15
- Loss: MatryoshkaLoss(CachedMultipleNegativesRankingLoss), dims 768/512/256/128/64
- 1 epoch, batch 64, lr 2e-5, fp16, max_seq_length 256, NO_DUPLICATES batch sampler
- Single T4 GPU, ~37 minutes
Limitations
- Trained and evaluated at 256 tokens; long contracts should be chunked.
- Predominantly US/UK/Australian sources. Other jurisdictions are underrepresented.
- Law StackExchange answers are community-written and not authoritative.
- Retrieval only — this model does not generate legal advice, and its output is not a substitute for a qualified lawyer.
- Downloads last month
- 137
Model tree for bugBug04S/legal-embed-modernbert-v2
Base model
answerdotai/ModernBERT-base