Mirror Memory Default Embedding Model

A fine-tuned English text embedding model developed for Mirror Memory, a Rust-based retrieval-augmented generation (RAG) and long-term AI memory library.

This model is based on BAAI/bge-small-en-v1.5 and was fine-tuned on a cleaned subset of the GooAQ question-answer dataset using MultipleNegativesRankingLoss.

The resulting model produces 384-dimensional dense embeddings and is intended for semantic retrieval workloads such as vector search, RAG, document retrieval, and long-term AI memory.


Model Details

Property Value
Model Mirror Memory Default Embedding Model
Base model BAAI/bge-small-en-v1.5
Architecture BERT
Language English
Embedding dimension 384
Maximum sequence length 512 tokens
Pooling CLS
Normalization L2 normalization
Training objective Multiple Negatives Ranking Loss
Training dataset GooAQ
Training subset ~30,000 cleaned samples
Weight format SafeTensors
Intended tasks Semantic search, retrieval, RAG, AI memory

Motivation

Mirror Memory uses vector embeddings to convert documents and text into representations that can be efficiently stored and retrieved from a vector database.

The default embedding model is designed to provide a compact embedding representation while remaining practical for local and resource-conscious inference.

It is intended to work particularly well for:

  • Semantic search
  • Retrieval-Augmented Generation (RAG)
  • Long-term AI memory
  • Document retrieval
  • Question-answer retrieval
  • Similarity search
  • Knowledge-base retrieval
  • Vector databases

Training

Base Model

The model starts from:

BAAI/bge-small-en-v1.5

The base model is an English BERT-based sentence embedding model with a 384-dimensional output space.

Pooling and Normalization

The exported model package contains 1_Pooling/config.json with:

{
  "embedding_dimension": 384,
  "pooling_mode": "cls",
  "include_prompt": true
}

Therefore, the shipped model uses CLS pooling, not mean pooling.

The package also contains the 2_Normalize Sentence Transformers module, so the final sentence embeddings are L2-normalized.

Dataset

Training used a cleaned subset of the GooAQ dataset.

  • Dataset: sentence-transformers/gooaq
  • Training subset: approximately 30,000 samples
  • Duplicate entries were removed during preprocessing.
  • The training data consists of question-answer pairs.

GooAQ is a large-scale question-answer dataset originally collected from Google search/autocomplete and answer boxes. See the original dataset card for details.

Objective

The model was fine-tuned using:

MultipleNegativesRankingLoss

This training objective encourages semantically related query/document pairs to have closer representations while using other examples in the batch as negative examples.


Evaluation

The model was evaluated using retrieval-oriented metrics on the project's evaluation setup. The structured model-index metadata above exposes these same results to the Hugging Face Hub.

Metric Result
Recall@1 82.37%
NDCG@10 90.75%
MRR@10 88.37%

These are project-specific evaluation results, not official MTEB benchmark results.

For reproducibility, benchmark details such as the exact evaluation split, preprocessing, and evaluation code should be considered alongside the Mirror Memory project source.


Inference

Sentence Transformers

Install:

pip install -U sentence-transformers

Load the model:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "KeZZ08/mirror-memory-default-embedding-model"
)

texts = [
    "Rust provides memory safety and high performance.",
    "Retrieval augmented generation combines search with language models.",
    "Vector databases store embeddings for semantic retrieval.",
]

embeddings = model.encode(
    texts,
    normalize_embeddings=True
)

print(embeddings.shape)
# (3, 384)

The returned embeddings have 384 dimensions.


Single Text Embedding

from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "KeZZ08/mirror-memory-default-embedding-model"
)

text = "Mirror Memory provides long-term retrieval for AI systems."

embedding = model.encode(
    text,
    normalize_embeddings=True
)

print(embedding.shape)
# (384,)

Batch Embedding

For retrieval systems, batch inference is recommended:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "KeZZ08/mirror-memory-default-embedding-model"
)

documents = [
    "Rust is a systems programming language.",
    "Qdrant is a vector database.",
    "RAG combines retrieval with generation.",
    "Embeddings represent text as numerical vectors.",
]

embeddings = model.encode(
    documents,
    batch_size=32,
    normalize_embeddings=True,
    show_progress_bar=True
)

print(embeddings.shape)
# (4, 384)

Similarity Search Example

from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer(
    "KeZZ08/mirror-memory-default-embedding-model"
)

query = "How does vector search work?"
documents = [
    "Vector databases store numerical representations of data.",
    "Python is commonly used for machine learning.",
    "Rust provides memory safety without garbage collection.",
]

query_embedding = model.encode(
    query,
    normalize_embeddings=True
)

document_embeddings = model.encode(
    documents,
    normalize_embeddings=True
)

scores = document_embeddings @ query_embedding

for document, score in sorted(
    zip(documents, scores),
    key=lambda x: x[1],
    reverse=True
):
    print(f"{score:.4f}  {document}")

Because the embeddings are normalized, the dot product can be used as cosine similarity.


Mirror Memory

This model serves as the default embedding model for the Mirror Memory project.

Mirror Memory is a Rust library designed for high-performance:

  • Document processing
  • Text extraction
  • Text cleaning
  • Text chunking
  • Embedding generation
  • Vector storage
  • Retrieval-Augmented Generation
  • Long-term AI memory

The project supports multiple embedding backends, including custom Hugging Face/Candle models and FastEmbed models.

The model in this repository is the default embedding model intended for Mirror Memory's retrieval pipeline.


Model Architecture

The model retains the BERT-based architecture of the BGE-small-en-v1.5 family.

Key characteristics:

Architecture:        BERT
Hidden size:         384
Transformer layers:  12
Attention heads:     12
Intermediate size:   1536
Maximum sequence:    512 tokens
Embedding size:      384
Pooling:             CLS
Normalization:       L2

The exported repository contains the Sentence Transformers configuration required to load the model, including:

1_Pooling/
2_Normalize/
config.json
config_sentence_transformers.json
model.safetensors
modules.json
sentence_bert_config.json
tokenizer.json
tokenizer_config.json

Intended Use

This model is intended for English-language embedding and retrieval applications, including:

  • Semantic search
  • Document similarity
  • Question-answer retrieval
  • RAG systems
  • AI memory systems
  • Knowledge-base search
  • Vector database indexing
  • Clustering and similarity analysis

Limitations

This model should not be treated as a general-purpose generative language model.

It generates embeddings rather than natural-language responses.

Important limitations include:

  • The model is optimized for English text.
  • Performance can vary across domains that differ substantially from the training data.
  • The reported evaluation results come from the project's own evaluation setup.
  • The model has not been presented here as an official MTEB benchmark submission.
  • As with other embedding models, retrieval quality depends on preprocessing, chunking, query formulation, vector similarity method, and downstream retrieval configuration.

Files

The repository contains the exported Sentence Transformers model:

.
├── 1_Pooling/
│   └── config.json
├── 2_Normalize/
│   └── config.json
├── config.json
├── config_sentence_transformers.json
├── model.safetensors
├── modules.json
├── README.md
├── sentence_bert_config.json
├── tokenizer.json
└── tokenizer_config.json

The model weights are provided in SafeTensors format.

No Python training checkpoint is required for inference.


Relationship to BGE

This model is a fine-tuned version of:

BAAI/bge-small-en-v1.5

The original BGE model and its documentation are maintained by the Beijing Academy of Artificial Intelligence (BAAI).

This repository contains the fine-tuned model used by Mirror Memory and should not be confused with the original BGE release.

For the original model, architecture details, benchmarks, and upstream documentation, refer to the official BGE model card.


License and Attribution

This repository is released under the MIT License.

The model is based on BAAI/bge-small-en-v1.5, which is also listed by its official Hugging Face repository under the MIT License.

The training data includes a cleaned subset of GooAQ. The original AllenAI GooAQ dataset is listed under the Apache License 2.0. Users should review the original dataset terms and attribution requirements when using this model or building derivative systems.

Upstream references


Citation

If you use this model in your project, please reference this repository:

KeZZ08/mirror-memory-default-embedding-model

For work involving the original BGE model, please also refer to the official BGE publication and model repository.


Disclaimer

This model is provided for research and software-development purposes. Users are responsible for evaluating the model for their own application, data, domain, and deployment requirements.

Downloads last month
40
Safetensors
Model size
33.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KeZZ08/mirror-memory-default-embedding-model

Finetuned
(412)
this model

Dataset used to train KeZZ08/mirror-memory-default-embedding-model

Evaluation results