Sentence Similarity
sentence-transformers
Safetensors
Swedish
English
eurobert
diabase
embeddings
semantic-search
retrieval
rag
swedish
european-ai
sovereign-ai
custom_code
Instructions to use Diabase/embedding-1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use Diabase/embedding-1 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Diabase/embedding-1", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| language: | |
| - sv | |
| - en | |
| library_name: sentence-transformers | |
| pipeline_tag: sentence-similarity | |
| base_model: EuroBERT/EuroBERT-210m | |
| tags: | |
| - diabase | |
| - embeddings | |
| - sentence-transformers | |
| - semantic-search | |
| - retrieval | |
| - rag | |
| - swedish | |
| - european-ai | |
| - sovereign-ai | |
| # diabase-embedding-1 | |
| > **Research preview.** An initial document-hash split was invalidated after a | |
| > semantic-leakage audit. This model was retrained from the Swedish base on corrected | |
| > topic-grouped v2 data. Independent MTEB evaluation confirms a large improvement over | |
| > that base, but the model remains well behind a strong public Swedish baseline. | |
| **A Swedish-first European embedding model for local semantic search and RAG.** | |
| `diabase-embedding-1` maps queries and documents into the same dense vector space. It | |
| is designed to retrieve a broad candidate set quickly; `diabase-reranker-1` can then | |
| rerank those candidates when higher precision is required. | |
| The model is built on the open European | |
| [EuroBERT-210m](https://huggingface.co/EuroBERT/EuroBERT-210m) encoder. Diabase first | |
| adapts that foundation to Swedish with masked-language training, then performs | |
| retrieval-specific contrastive training on Swedish query/document pairs. | |
| | | | | |
| |---|---| | |
| | **Status** | Research preview; not recommended as a production default | | |
| | **Architecture** | EuroBERT encoder + mean pooling + L2 normalization | | |
| | **Parameters** | Approximately 210M in the released encoder | | |
| | **Embedding dimension** | 768 | | |
| | **Primary language** | Swedish | | |
| | **Secondary language** | English replay is retained during adaptation | | |
| | **Initial retrieval context** | Up to 512 tokens per query/document | | |
| | **Deployment** | Local CPU/GPU; Sentence Transformers; ONNX planned | | |
| | **Companion model** | `Diabase/reranker-1` | | |
| ## Quick start | |
| ```python | |
| from sentence_transformers import SentenceTransformer | |
| model = SentenceTransformer("Diabase/embedding-1") | |
| queries = model.encode_query( | |
| ["Vilka krav ställer EU:s AI-förordning på högrisksystem?"], | |
| normalize_embeddings=True, | |
| ) | |
| documents = model.encode_document( | |
| [ | |
| "Leverantörer av högrisksystem måste bland annat arbeta med " | |
| "riskhantering, dokumentation, datakvalitet och mänsklig tillsyn.", | |
| "Stockholm är Sveriges huvudstad och ligger vid Mälaren.", | |
| ], | |
| normalize_embeddings=True, | |
| ) | |
| scores = queries @ documents.T | |
| print(scores) | |
| ``` | |
| The final release will document whether explicit query/document prompts are required. | |
| Use the model's bundled Sentence Transformers configuration rather than manually | |
| constructing pooling layers. | |
| ## Intended use | |
| - Swedish semantic search over private document collections | |
| - First-stage retrieval for local and on-prem RAG systems | |
| - Search over Swedish government, legal, technical and enterprise documents | |
| - Query/document clustering, duplicate detection and related-text discovery | |
| - Candidate generation before `diabase-reranker-1` | |
| ## Out of scope | |
| - Treating cosine similarity as a calibrated probability | |
| - Using retrieval score as evidence that a document is true or trustworthy | |
| - Unsupervised high-stakes decisions about people | |
| - Generating text or answering questions directly | |
| - Replacing access control, document authorization or privacy filtering | |
| ## Model design | |
| The product is created in three stages: | |
| 1. **European foundation:** EuroBERT-210m, released under Apache-2.0. | |
| 2. **Swedish language adaptation:** masked-language training with a controlled mixture | |
| of cleaned Swedish web text, Swedish Wikipedia and English replay. | |
| 3. **Retrieval training:** contrastive query/document learning with in-batch negatives, | |
| followed by hard-negative training. | |
| The masked-language head used during adaptation is discarded. The released artifact is | |
| the approximately 210M-parameter encoder with pooling and normalization. | |
| ## Training data | |
| The Swedish v1 adaptation run used: | |
| ### Language adaptation | |
| - 70% cleaned HPLT Swedish | |
| - 15% Swedish Wikipedia | |
| - 15% cleaned HPLT English replay | |
| - Actual budget: 300,007,424 fixed-length training tokens | |
| Source weights are applied after tokenization to fixed-length chunks, so they represent | |
| token shares rather than document shares. | |
| ### Retrieval tuning | |
| - Project-generated Swedish extractive-QA query/document positives | |
| - Additional Swedish document/query pairs after validation and deduplication | |
| - In-batch negatives | |
| - Hard negatives retrieved by an earlier embedding checkpoint | |
| Documents, not individual questions, are assigned to train/dev/test. This prevents | |
| questions sharing the same passage from leaking that passage across splits. | |
| ### Data caveat | |
| HPLT licenses its dataset packaging under CC0 but does not claim ownership of the | |
| underlying web text. Wikipedia and every additional source retain their respective | |
| terms. Deployers are responsible for evaluating data, privacy and jurisdictional | |
| requirements for their use case. A final data-provenance appendix will accompany the | |
| release. | |
| ## Evaluation | |
| ### Swedish language adaptation | |
| Before retrieval-specific tuning, Diabase ran a 16.4M-token masked-language pilot and a | |
| 300M-token controlled Swedish adaptation. The comparison below uses mean-pooled | |
| EuroBERT embeddings on a bootstrap held-out set of 810 Swedish questions and 807 | |
| candidate passages. | |
| | Metric | EuroBERT-210m | 16.4M pilot | 300M adaptation | Change vs base | | |
| |--------|---------------:|-------------:|----------------:|---------------:| | |
| | Recall@1 | 3.09% | 5.06% | **5.80%** | +88% | | |
| | Recall@5 | 8.27% | 12.96% | **15.68%** | +90% | | |
| | Recall@10 | 12.72% | 19.26% | **22.96%** | +81% | | |
| | Recall@20 | 20.12% | 28.89% | **32.96%** | +64% | | |
| | MRR@10 | 5.37% | 8.52% | **10.29%** | +92% | | |
| | nDCG@10 | 7.07% | 11.01% | **13.26%** | +88% | | |
| | Median relevant-document rank | 107 | **55** | 57 | -47% | | |
| | Mean relevant-document rank | 174.4 | **129.9** | 144.3 | -17% | | |
| The 300M run clearly improves top-k retrieval over both unmodified EuroBERT and the | |
| pilot. Its mean and median rank are slightly worse than the pilot, indicating a weaker | |
| tail for queries that miss the top results. These are **adaptation results, not final | |
| product scores**. The bootstrap dataset derives from project QA data and is useful for | |
| iteration, but it is not sufficiently independent or representative for a release claim. | |
| ### Final release gates | |
| The final model card will report: | |
| - Recall@1/5/10/20 | |
| - MRR@10 | |
| - nDCG@10 | |
| - Swedish and cross-lingual retrieval | |
| - Latency, throughput and memory on CPU and GPU | |
| - Comparison against unmodified EuroBERT and relevant open baselines | |
| - Results on a separate human-reviewed product benchmark | |
| The model will not be released as v1 unless it improves held-out Recall@10 and nDCG@10 | |
| over the unmodified foundation without an unacceptable cross-lingual regression. | |
| ### Retrieval-tuned model | |
| `diabase-embedding-1` was then trained for three epochs on 6,705 query/document | |
| positives with Cached Multiple Negatives Ranking Loss and 127 in-batch negatives per | |
| query. | |
| | Metric | Swedish-adapted base | **diabase-embedding-1** | Absolute change | | |
| |--------|---------------------:|------------------------:|----------------:| | |
| | Recall@1 | 5.80% | **42.10%** | +36.30 points | | |
| | Recall@5 | 15.68% | **78.64%** | +62.96 points | | |
| | Recall@10 | 22.96% | **89.26%** | +66.30 points | | |
| | Recall@20 | 32.96% | **98.27%** | +65.31 points | | |
| | MRR@10 | 10.29% | **57.18%** | +46.89 points | | |
| | nDCG@10 | 13.26% | **64.91%** | +51.65 points | | |
| | Median relevant-document rank | 57 | **2** | -55 | | |
| | Mean relevant-document rank | 144.3 | **4.2** | -140.1 | | |
| The final dev and test Recall@10 values were 90.26% and 89.26%, respectively. Their | |
| close agreement rules out a simple dev-only anomaly, but does not rule out shared-topic | |
| leakage. | |
| After training, an audit found many paraphrased passages about the same facts and topics | |
| across splits. The split grouped exact normalized documents but did not group semantic | |
| duplicates. For example, the corpus contains multiple differently worded passages that | |
| answer the same Göta kanal question. Some were also incorrectly selected as hard | |
| negatives. | |
| The table is retained for experiment transparency, but **must not be presented as an | |
| independent generalization result**. It measures performance on unseen wording inside a | |
| heavily shared topic distribution and may substantially overstate real-world quality. | |
| ### Release status | |
| The v2 weights may be published as an explicitly labelled research preview. A stable | |
| release requires: | |
| - Semantic clustering or source/topic grouping before train/dev/test assignment | |
| - Multi-positive relevance labels for paraphrased passages | |
| - A separate human-reviewed benchmark built from real Swedish documents | |
| - Comparison with relevant public Swedish and multilingual embedding baselines | |
| - Swedish/English cross-lingual regression checks | |
| - CPU/GPU latency, throughput and memory measurements | |
| - Final data-license review and artifact checksums | |
| ### Corrected v2 evaluation | |
| After the audit, Diabase rebuilt the data into 371 semantic/factual groups before | |
| splitting. The retrieval model was retrained from the Swedish-adapted encoder, not from | |
| the contaminated v1 retrieval checkpoint. The held-out test contains 455 queries, | |
| 455 documents and 32 topic groups. | |
| | Metric | Swedish-adapted base | **v2 candidate** | Absolute change | | |
| |--------|---------------------:|-----------------:|----------------:| | |
| | Recall@1 | 22.86% | **49.01%** | +26.15 points | | |
| | Recall@5 | 43.08% | **77.58%** | +34.50 points | | |
| | Recall@10 | 48.79% | **86.81%** | +38.02 points | | |
| | Recall@20 | 60.22% | **95.16%** | +34.94 points | | |
| | MRR@10 | 30.70% | **61.27%** | +30.57 points | | |
| | nDCG@10 | 35.03% | **67.40%** | +32.37 points | | |
| | Median relevant-document rank | 11 | **2** | -9 | | |
| | Mean relevant-document rank | 47.6 | **4.8** | -42.8 | | |
| This corrected result is substantially more credible than the original bootstrap: | |
| topic families are assigned as groups, paraphrased passages can be labelled as multiple | |
| positives, and the candidate was trained only on the new training split. It remains an | |
| internal synthetic benchmark, not an independent claim about arbitrary Swedish corpora. | |
| ### Independent Swedish MTEB evaluation | |
| The primary public evaluation uses MTEB 2.18.1, a maximum sequence length of 512 and | |
| the fixed dataset revisions recorded in the result files under `runs/`. nDCG@10 is the | |
| main score for both tasks. | |
| | Model | SweFAQ nDCG@10 | SweDN nDCG@10 | | |
| |---|---:|---:| | |
| | Swedish-adapted EuroBERT base | 10.46% | 2.96% | | |
| | **diabase-embedding-1** | **47.08%** | **34.03%** | | |
| | KBLab sentence-bert-swedish-cased | 73.30% | 70.80% | | |
| The Diabase model improves the adapted base by 36.62 and 31.07 percentage points, so | |
| retrieval tuning clearly learned transferable Swedish retrieval behavior. It does not | |
| match the public KBLab baseline: the remaining deficits are 26.22 points on SweFAQ and | |
| 36.78 points on SweDN. This result blocks a stable or best-in-class claim. | |
| ### Local latency | |
| Measured with PyTorch 2.11.0, batch size 32 for document encoding and 20 timed | |
| single-query iterations: | |
| - CPU: 89.2 ms mean query latency; 25.1 documents/second over 1,000 documents. | |
| - NVIDIA RTX 5070: 26.0 ms mean query latency; 1,104.3 documents/second. | |
| - Combined embedding/reranker benchmark peak GPU allocation: approximately 1.67 GB. | |
| These are local measurements, not service-level guarantees. | |
| ## Limitations | |
| - Swedish is the primary optimization target; other languages may perform worse. | |
| - The first release is trained mainly on web, encyclopedic and QA-shaped text. | |
| - Long documents must initially be chunked; native architecture context does not imply | |
| that retrieval quality has been validated at 8,192 tokens. | |
| - Dense retrieval can miss exact identifiers, numbers and rare terms. Hybrid lexical | |
| retrieval may improve those cases. | |
| - Similarity models can reproduce social and topical biases in their training data. | |
| - Retrieval quality depends strongly on chunking, document quality and query style. | |
| ## Security and privacy | |
| The model runs locally and does not require a hosted inference API. Local execution does | |
| not itself provide compliance: callers must enforce document authorization before | |
| retrieval, protect vector indexes, remove secrets where appropriate and defend against | |
| retrieval-corpus poisoning and prompt injection in retrieved documents. | |
| ## Relationship to other Diabase models | |
| ```text | |
| query | |
| -> diabase-embedding-1 | |
| -> top candidate documents | |
| -> diabase-reranker-1 (optional) | |
| -> best grounded context | |
| -> Diabase Europa | |
| ``` | |
| `diabase-embedding-1` is optimized for fast candidate recall. | |
| `diabase-reranker-1` is optimized for slower, more precise pairwise relevance scoring. | |
| ## License and attribution | |
| The Diabase derivative is released under Apache-2.0, consistent with the EuroBERT | |
| foundation. Release weight SHA-256: | |
| `669e5fe8c969471ea9bdec83a262f587eb7a43faea4ed423193d36ae9d4d4789`. | |
| EuroBERT citation: | |
| ```bibtex | |
| @misc{boizard2025eurobertscalingmultilingualencoders, | |
| title={EuroBERT: Scaling Multilingual Encoders for European Languages}, | |
| author={Nicolas Boizard and Hippolyte Gisserot-Boukhlef and Duarte M. Alves | |
| and André Martins and Ayoub Hammal and Caio Corro and Céline Hudelot | |
| and Emmanuel Malherbe and Etienne Malaboeuf and Fanny Jourdan | |
| and Gabriel Hautreux and João Alves and Kevin El-Haddad and Manuel Faysse | |
| and Maxime Peyrard and Nuno M. Guerreiro and Patrick Fernandes | |
| and Ricardo Rei and Pierre Colombo}, | |
| year={2025}, | |
| eprint={2503.05500}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.CL}, | |
| url={https://arxiv.org/abs/2503.05500} | |
| } | |
| ``` | |
| ## Maintainer | |
| Built by [Diabase](https://diabase.ai). | |