SSE Retrieval MRL v2: Regularization of Representation Space and Performance Improvement via Hyperparameter Optimization
Rikka Botan
Independent Researcher, Japan
https://rikka-botan.github.io
Abstract
1 Introduction
In the field of Natural Language Processing (NLP), techniques for embedding sentence semantics into vector spaces are fundamental to information retrieval and semantic search. While Transformer-based models demonstrate high accuracy, their computational costs remain a significant challenge for real-time inference on edge devices. Conversely, approaches utilizing static embeddings offer speed but have been noted for limitations in expressiveness and difficulties in regularization.
This study reports on the development of SSE Retrieval MRL v2 to address these challenges. We emphasize improvements over version 1 (v1), specifically focusing on enhanced representation space regularization through hyperparameter optimization to achieve a balance between accuracy and efficiency.
SSE Retrieval MRL v2:
https://huggingface.co/RikkaBotan/stable-static-embedding-fast-retrieval-mrl-en-v2
version 1 technical report:
https://huggingface.co/blog/RikkaBotan/stable-static-embedding-technical-report
2 Methodology
2.1 Architecture
Figure 1 | SSE Architecture
2.2 Training Configuration
The model was trained with the following hyperparameters:
- Batch Size: 2048 (
per_device_train_batch_size) - Gradient Accumulation Steps: 4 (
gradient_accumulation_steps) - Learning Rate: 0.1
- Optimizer: AdamW (beta2: 0.9999, epsilon: 1e-10)
- Scheduler: Cosine with Warmup (ratio: 0.1)
- Epochs: 1
- Computer: A100 SXM4 (80GB) (vast ai)
2.3 Datasets and Loss Functions
Training was conducted using 15 datasets, including SQuAD, TriviaQA, and AllNLI. We employed a combination of MatryoshkaLoss—aimed at maintaining performance across diverse dimensions—and MultipleNegativesRankingLoss.
3 Results
3.1 Training results
Figure 2 | Training Loss Across Training Steps
Figure 3 | NanoBEIR mean nDCG@10 Across Training Steps
3.2 Evaluation results
Evaluation results using the NanoBEIR benchmark are presented below.
Table 1 | Model Comparison (NanoBEIR Mean NDCG@10)
| Model | NanoBEIR Mean nDCG@10 | Dimensions | Parameters | Notes |
|---|---|---|---|---|
| SSE Retrieval MRL v2 | 0.5159 | 512 | 15.6M | Released model |
| SSE Retrieval MRL v1 | 0.5124 | 512 | ~16M | Previous version |
| static-retrieval-mrl-en-v1 | 0.5030 | 1024 | 31.3M | Prior static reference |
3.3 Differences between v2 and v1
Version 2 demonstrates superiority over previous versions (SSE Retrieval MRL and static-retrieval-mrl-en-v1) in the following areas:
Retrieval Quality: SSE Retrieval MRL v2 achieves a NanoBEIR mean nDCG@10 of 0.5159 at 512 dimensions. This corresponds to an absolute improvement of 0.0035 over v1 (0.5124) and 0.0129 over the prior static reference at its native 1024-dimensional width (0.5030).
Dimensional Efficiency: The 256-dimensional prefix of SSE Retrieval MRL v2 reaches 0.5036, matching the prior static reference's native 1024-dimensional score of 0.5030 with one quarter of the vector storage. This result establishes a storage and dense-similarity-computation advantage; encoder throughput is evaluated separately in the paper.
3.4 Matryoshka Truncation and Spectral Analysis
The performance improvement of SSE Retrieval MRL v2 is attributed to the synergistic effect of gradient control via DyT layers and hyperparameter tuning. Notably, the ability to maintain a score of 0.5036 even when down-sampled to 256 dimensions due to Matryoshka properties suggests flexible applicability under resource constraints.
Figure 4 | NanoBEIR English mean nDCG@10 vs Matryoshka Embedding Truncation.
Table 2 | NanoBEIR English mean nDCG@10 vs Matryoshka Embedding Truncation.
| Model | 32 | 64 | 128 | 256 | 512 | 1024 |
|---|---|---|---|---|---|---|
| SSE (Static Embedding + Separable DyT) v2 | 0.3487 | 0.4236 | 0.4732 | 0.5036 | 0.5159 | - |
| SSE (Static Embedding + Separable DyT) v1 | 0.3448 | 0.4275 | 0.4659 | 0.4969 | 0.5124 | - |
| Static Embedding + DyT | 0.3338 | 0.4134 | 0.4622 | 0.4919 | 0.5025 | - |
| Static Embedding (no DyT) | 0.3367 | 0.4161 | 0.4625 | 0.4912 | 0.5068 | - |
| static-retrieval-mrl-en-v1 (For reference) | 0.3532 | 0.4177 | 0.4621 | 0.4818 | 0.4958 | 0.5030 |
Figure 5 | PCA Spectrum on the 13 NanoBEIR English Datasets: Normalized Eigenvalue Decay (Logarithmic Scale).
4 Discussion
Compared to previous models, SSE Retrieval MRL v2 shows even stronger low-rank regularization. These results provide further evidence supporting the hypothesis that there is a correlation between low-rank regularization in the representation space and model performance when training with matryoshka loss and target learning.
However, this is a trend observed with SSE models, and it is not confirmed whether it holds for general embedding models.
5 Conclusion
SSE Retrieval MRL v2 has been demonstrated as a lightweight and high-performance information retrieval model by achieving regularization of the representation space through hyperparameter optimization. Improvements from version 1 have yielded significant progress in both accuracy and speed.
Acknowledgements
Our interest in this topic originated from reading Tom Aarsen's seminal article, Train 400x faster Static Embedding Models with Sentence Transformers, which motivated us to investigate on static embedding.
I thank the developers of sentence-transformers, python and pytorch.
I thank all the researchers for their efforts to date.
I thank Japan's high standard of education.
And most of all, thank you for your interest in this blog.
About us
Japanese independent researcher having shy and pampered personality. Twin-tail hair is a charm point. Interested in nlp. Usually using python and C.
Please contact us if you have any requests for joint research, writing, speaking engagements, or employment.







