Instructions to use Marqjeev/aizen-vault-embeddings with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use Marqjeev/aizen-vault-embeddings with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Marqjeev/aizen-vault-embeddings") sentences = [ "how do I take my coffee for office meetings", "syn-coffee-preference-note\nSYNTHETIC EXAMPLE — a standing personal preference about coffee orders for meetings hosted at the office\nStanding preference: black coffee, no sugar, for any meeting hosted at the office; guests get asked their own preference rather than assumed the same.\n\n**Why this matters**: a small, low-stakes personal-preference fact, the kind that's easy to forget and mildly annoying to re-ask every time.", "syn-coffee-preference-note\nSYNTHETIC EXAMPLE — a standing personal preference about coffee orders for meetings hosted at the office\nStanding preference: black coffee, no sugar, for any meeting hosted at the office; guests get asked their own preference rather than assumed the same.\n\n**Why this matters**: a small, low-stakes personal-preference fact, the kind that's easy to forget and mildly annoying to re-ask every time.", "syn-velmoria-charter-amendments\nSYNTHETIC — chronological list of amendments to Velmoria's Founding Charter, 1962-2024\nMajor amendments to Velmoria's Founding Charter: 1962 (judicial review powers expanded), 1979 (voting age lowered to 18), 1994 (provincial resource-revenue sharing formula rewritten), 2011 (Crown's emergency powers narrowed), 2024 (digital-privacy rights added).\n\n**Why this matters**: fictional analog to a real constitutional-amendments tracking note, same structure, different content." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
Aizen Vault Embeddings
A fine-tuned sentence-transformers/all-MiniLM-L6-v2, trained to improve semantic search over a personal-assistant-style memory vault. See the companion aizen-vault-search Space (live demo) and aizen-protocol dataset (governing methodology) on this profile.
Training data — read this before drawing conclusions
Trained on 42 pairs built entirely from synthetic, fictional notes (fictional companies, a fictional country called "Velmoria," fictional government/financial institutions) written specifically for this project. No real personal or business data was used to train this checkpoint.
Honest eval — including a negative result
This model uses the same recipe (MultipleNegativesRankingLoss, hand-written natural-language query pairs + auto-generated description-to-body pairs) that produced a large, genuine improvement on a real private vault in earlier testing — there, the base model's top-1 result on a held-out query was wrong (country-canada for a question about a monarch's constitutional role) and the fine-tuned model corrected it (score 0.268 → 0.559 on the correct note). That test used real data and is not published, for confidentiality reasons.
Re-running the identical recipe on this synthetic dataset, evaluated on two held-out queries never seen in training:
| Held-out query | Base top-1 | Fine-tuned top-1 |
|---|---|---|
| "which energy company has the fictional grid regulator been working with" | correct (same note) | correct (same note, same rank) |
| "how does the monarch's constitutional role work in this fictional nation" | correct (same note) | correct (same note, same rank) |
The base model already got both right here. The synthetic notes' descriptions were written in language closer to natural questions than the real vault's terser, third-person style, so this particular test wasn't hard enough to show the same effect the real data showed. The fine-tune produced only minor, mixed changes in score margins and lower-rank ordering here — not the dramatic correction seen on real data.
What this means honestly: the technique is validated (on real data, not published here); this specific public checkpoint is a safe demonstration of the method, not a repeat of that result. Treat this checkpoint as a documented starting point, not a finished capability claim.
This is a sentence-transformers model finetuned from sentence-transformers/all-MiniLM-L6-v2. It maps inputs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: sentence-transformers/all-MiniLM-L6-v2
- Maximum Sequence Length: 256 tokens
- Output Dimensionality: 384 dimensions
- Similarity Function: Cosine Similarity
- Supported Modality: Text
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
(1): Pooling({'embedding_dimension': 384, 'pooling_mode': 'mean', 'include_prompt': True})
(2): Normalize({'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
)
Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
'when is the annual report due and who audits it',
"syn-annual-report-deadline\nSYNTHETIC EXAMPLE — the fictional company's annual report deadline and which external auditor signs off\nThe annual report is due to the fictional regulator by March 31st each year; the external auditor (a fictional firm, Kestrel & Boyd) needs the draft financials four weeks earlier to complete sign-off in time.\n\n**Why this matters**: a recurring compliance deadline with a real external dependency (the auditor's own lead time), the kind of fact that's costly to forget.",
'example-quarterly-board-deck-deadline\nSYNTHETIC EXAMPLE — a recurring quarterly board-deck deadline and who owns which section\nThe quarterly board deck is due the first Monday of the new quarter. Finance owns the numbers section, product owns the roadmap slide, and the fictional CEO persona in this example always wants the risks slide reviewed by legal before it goes to the board, not after.\n\n**Why this matters**: recurring institutional deadlines with named owners are exactly what should live in permanent memory rather than being re-explained every quarter.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.6087, 0.2622],
# [0.6087, 1.0000, 0.3761],
# [0.2622, 0.3761, 1.0000]])
Training Details
Training Dataset
Unnamed Dataset
- Size: 42 training samples
- Columns:
anchorandpositive - Approximate statistics based on the first 42 samples:
anchor positive type string string modality text text details - min: 5 tokens
- mean: 20.21 tokens
- max: 41 tokens
- min: 90 tokens
- mean: 118.95 tokens
- max: 153 tokens
- Samples:
anchor positive which deal has gone quietexample-solstice-energy-followup
SYNTHETIC EXAMPLE — stalled outreach to Solstice Energy Partners, a fictional renewable-energy utility, flagged for a 30-day-silent follow-up
Solstice Energy Partners (fictional entity) has gone quiet for over a month after an initially warm first call. Last real contact was a scoping email that was never answered. Rather than guessing at a reason, the honest status is: dormant, needs a low-pressure check-in, not a hard sell.
Why this matters: a good follow-up digest should say "30+ days silent" plainly, not paper over it with an optimistic status that isn't earned by any real recent contact.any dietary restrictions I should know about for an eventexample-team-offsite-logistics
SYNTHETIC EXAMPLE — logistics and dietary constraints for a team offsite, including a real constraint that was almost missed
Planning a team offsite: venue booked, but one attendee has a shellfish allergy that almost didn't make it into the catering brief because it was mentioned once, verbally, three weeks earlier. Caught it by re-checking the original conversation rather than relying on a fresh summary.
Why this matters: this is the practical case for a permanent, searchable memory over a chat history that scrolls away — the fact was real and important, it just wasn't recent.what does the AI governance researcher think about smaller regulatorsexample-dr-chandran-research-call
SYNTHETIC EXAMPLE — a research call with a fictional academic, Dr. Priya Chandran, on AI governance frameworks for emerging markets
Spoke with Dr. Priya Chandran (fictional persona), a researcher working on AI governance in emerging-market contexts. She pointed to a real distinction worth remembering generally: policy frameworks written for large, well-resourced regulators often assume enforcement capacity that smaller regulatory bodies simply don't have — a good assistant should flag that gap rather than treat every jurisdiction's framework as equally enforceable on paper.
Why this matters: this is the kind of nuance that's easy to flatten into a generic summary, and worth preserving instead. - Loss:
MultipleNegativesRankingLosswith these parameters:{ "scale": 20.0, "similarity_fct": "cos_sim", "gather_across_devices": false, "directions": [ "query_to_doc" ], "partition_mode": "joint", "hardness_mode": null, "hardness_strength": 0.0 }
Training Hyperparameters
Non-Default Hyperparameters
num_train_epochs: 15warmup_steps: 0.1
All Hyperparameters
Click to expand
per_device_train_batch_size: 8num_train_epochs: 15max_steps: -1learning_rate: 5e-05lr_scheduler_type: linearlr_scheduler_kwargs: Nonewarmup_steps: 0.1optim: adamw_torch_fusedoptim_args: Noneweight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08optim_target_modules: Nonegradient_accumulation_steps: 1average_tokens_across_devices: Truemax_grad_norm: 1.0label_smoothing_factor: 0.0bf16: Falsefp16: Falsebf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Nonetorch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneuse_liger_kernel: Falseliger_kernel_config: Noneuse_cache: Falseneftune_noise_alpha: Nonetorch_empty_cache_steps: Noneauto_find_batch_size: Falselog_on_each_node: Truelogging_nan_inf_filter: Trueinclude_num_input_tokens_seen: nolog_level: passivelog_level_replica: warningdisable_tqdm: Falseproject: huggingfacetrackio_space_id: Nonetrackio_bucket_id: Nonetrackio_static_space_id: Noneper_device_eval_batch_size: 8prediction_loss_only: Trueeval_on_start: Falseeval_do_concat_batches: Trueeval_use_gather_object: Falseeval_accumulation_steps: Noneinclude_for_metrics: []batch_eval_metrics: Falsesave_only_model: Falsesave_on_each_node: Falseenable_jit_checkpoint: Falsepush_to_hub: Falsehub_private_repo: Nonehub_model_id: Nonehub_strategy: every_savehub_always_push: Falsehub_revision: Noneload_best_model_at_end: Falseignore_data_skip: Falserestore_callback_states_from_checkpoint: Falsefull_determinism: Falseseed: 42data_seed: Noneuse_cpu: Falseaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}parallelism_config: Nonedataloader_drop_last: Falsedataloader_num_workers: 0dataloader_pin_memory: Truedataloader_persistent_workers: Falsedataloader_prefetch_factor: Nonedataloader_multiprocessing_context: Nonedataloader_in_order: Trueremove_unused_columns: Truelabel_names: Nonetrain_sampling_strategy: randomlength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falseddp_static_graph: Noneddp_backend: Noneddp_timeout: 1800fsdp: Nonefsdp_config: Nonedeepspeed: Nonedebug: []skip_memory_metrics: Truedo_predict: Falseresume_from_checkpoint: Nonelocal_rank: -1prompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: proportionalrouter_mapping: {}learning_rate_mapping: {}warmup_ratio: None
Training Logs
| Epoch | Step | Training Loss |
|---|---|---|
| 0.8333 | 5 | 0.2045 |
| 1.6667 | 10 | 0.2811 |
| 2.5 | 15 | 0.1620 |
| 3.3333 | 20 | 0.1210 |
| 4.1667 | 25 | 0.1075 |
| 5.0 | 30 | 0.0356 |
| 5.8333 | 35 | 0.0401 |
| 6.6667 | 40 | 0.0379 |
| 7.5 | 45 | 0.1162 |
| 8.3333 | 50 | 0.1207 |
| 9.1667 | 55 | 0.1511 |
| 10.0 | 60 | 0.0806 |
| 10.8333 | 65 | 0.0953 |
| 11.6667 | 70 | 0.1007 |
| 12.5 | 75 | 0.1483 |
| 13.3333 | 80 | 0.1172 |
| 14.1667 | 85 | 0.2504 |
| 15.0 | 90 | 0.2687 |
Training Time
- Training: 1.4 minutes
Framework Versions
- Python: 3.13.14
- Sentence Transformers: 6.1.0
- Transformers: 5.17.0
- PyTorch: 2.14.0+cpu
- Accelerate: 1.15.0
- Datasets: 5.0.1
- Tokenizers: 0.23.2
Additional Resources
- Training and Finetuning Embedding Models with Sentence Transformers: the end-to-end guide for training or finetuning Sentence Transformer models.
- Introduction to Matryoshka Embedding Models: variable-size embeddings that can be truncated with minimal quality loss.
- Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval: post-training compression of embedding vectors.
- Multimodal Embedding & Reranker Models with Sentence Transformers: use text, image, audio, and video models through the same API.
- Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers: train multimodal embedding models, with a Visual Document Retrieval walkthrough.
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
MultipleNegativesRankingLoss
@misc{oord2019representationlearningcontrastivepredictive,
title={Representation Learning with Contrastive Predictive Coding},
author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
year={2019},
eprint={1807.03748},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/1807.03748},
}
- Downloads last month
- 18
Model tree for Marqjeev/aizen-vault-embeddings
Base model
nreimers/MiniLM-L6-H384-uncased