Aizen Vault Embeddings

A fine-tuned sentence-transformers/all-MiniLM-L6-v2, trained to improve semantic search over a personal-assistant-style memory vault. See the companion aizen-vault-search Space (live demo) and aizen-protocol dataset (governing methodology) on this profile.

Training data — read this before drawing conclusions

Trained on 42 pairs built entirely from synthetic, fictional notes (fictional companies, a fictional country called "Velmoria," fictional government/financial institutions) written specifically for this project. No real personal or business data was used to train this checkpoint.

Honest eval — including a negative result

This model uses the same recipe (MultipleNegativesRankingLoss, hand-written natural-language query pairs + auto-generated description-to-body pairs) that produced a large, genuine improvement on a real private vault in earlier testing — there, the base model's top-1 result on a held-out query was wrong (country-canada for a question about a monarch's constitutional role) and the fine-tuned model corrected it (score 0.268 → 0.559 on the correct note). That test used real data and is not published, for confidentiality reasons.

Re-running the identical recipe on this synthetic dataset, evaluated on two held-out queries never seen in training:

Held-out query Base top-1 Fine-tuned top-1
"which energy company has the fictional grid regulator been working with" correct (same note) correct (same note, same rank)
"how does the monarch's constitutional role work in this fictional nation" correct (same note) correct (same note, same rank)

The base model already got both right here. The synthetic notes' descriptions were written in language closer to natural questions than the real vault's terser, third-person style, so this particular test wasn't hard enough to show the same effect the real data showed. The fine-tune produced only minor, mixed changes in score margins and lower-rank ordering here — not the dramatic correction seen on real data.

What this means honestly: the technique is validated (on real data, not published here); this specific public checkpoint is a safe demonstration of the method, not a repeat of that result. Treat this checkpoint as a documented starting point, not a finished capability claim.


This is a sentence-transformers model finetuned from sentence-transformers/all-MiniLM-L6-v2. It maps inputs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: sentence-transformers/all-MiniLM-L6-v2
  • Maximum Sequence Length: 256 tokens
  • Output Dimensionality: 384 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 384, 'pooling_mode': 'mean', 'include_prompt': True})
  (2): Normalize({'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'when is the annual report due and who audits it',
    "syn-annual-report-deadline\nSYNTHETIC EXAMPLE — the fictional company's annual report deadline and which external auditor signs off\nThe annual report is due to the fictional regulator by March 31st each year; the external auditor (a fictional firm, Kestrel & Boyd) needs the draft financials four weeks earlier to complete sign-off in time.\n\n**Why this matters**: a recurring compliance deadline with a real external dependency (the auditor's own lead time), the kind of fact that's costly to forget.",
    'example-quarterly-board-deck-deadline\nSYNTHETIC EXAMPLE — a recurring quarterly board-deck deadline and who owns which section\nThe quarterly board deck is due the first Monday of the new quarter. Finance owns the numbers section, product owns the roadmap slide, and the fictional CEO persona in this example always wants the risks slide reviewed by legal before it goes to the board, not after.\n\n**Why this matters**: recurring institutional deadlines with named owners are exactly what should live in permanent memory rather than being re-explained every quarter.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.6087, 0.2622],
#         [0.6087, 1.0000, 0.3761],
#         [0.2622, 0.3761, 1.0000]])

Training Details

Training Dataset

Unnamed Dataset

  • Size: 42 training samples
  • Columns: anchor and positive
  • Approximate statistics based on the first 42 samples:
    anchor positive
    type string string
    modality text text
    details
    • min: 5 tokens
    • mean: 20.21 tokens
    • max: 41 tokens
    • min: 90 tokens
    • mean: 118.95 tokens
    • max: 153 tokens
  • Samples:
    anchor positive
    which deal has gone quiet example-solstice-energy-followup
    SYNTHETIC EXAMPLE — stalled outreach to Solstice Energy Partners, a fictional renewable-energy utility, flagged for a 30-day-silent follow-up
    Solstice Energy Partners (fictional entity) has gone quiet for over a month after an initially warm first call. Last real contact was a scoping email that was never answered. Rather than guessing at a reason, the honest status is: dormant, needs a low-pressure check-in, not a hard sell.

    Why this matters: a good follow-up digest should say "30+ days silent" plainly, not paper over it with an optimistic status that isn't earned by any real recent contact.
    any dietary restrictions I should know about for an event example-team-offsite-logistics
    SYNTHETIC EXAMPLE — logistics and dietary constraints for a team offsite, including a real constraint that was almost missed
    Planning a team offsite: venue booked, but one attendee has a shellfish allergy that almost didn't make it into the catering brief because it was mentioned once, verbally, three weeks earlier. Caught it by re-checking the original conversation rather than relying on a fresh summary.

    Why this matters: this is the practical case for a permanent, searchable memory over a chat history that scrolls away — the fact was real and important, it just wasn't recent.
    what does the AI governance researcher think about smaller regulators example-dr-chandran-research-call
    SYNTHETIC EXAMPLE — a research call with a fictional academic, Dr. Priya Chandran, on AI governance frameworks for emerging markets
    Spoke with Dr. Priya Chandran (fictional persona), a researcher working on AI governance in emerging-market contexts. She pointed to a real distinction worth remembering generally: policy frameworks written for large, well-resourced regulators often assume enforcement capacity that smaller regulatory bodies simply don't have — a good assistant should flag that gap rather than treat every jurisdiction's framework as equally enforceable on paper.

    Why this matters: this is the kind of nuance that's easy to flatten into a generic summary, and worth preserving instead.
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • num_train_epochs: 15
  • warmup_steps: 0.1

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 8
  • num_train_epochs: 15
  • max_steps: -1
  • learning_rate: 5e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0.1
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 8
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • dataloader_multiprocessing_context: None
  • dataloader_in_order: True
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}
  • warmup_ratio: None

Training Logs

Epoch Step Training Loss
0.8333 5 0.2045
1.6667 10 0.2811
2.5 15 0.1620
3.3333 20 0.1210
4.1667 25 0.1075
5.0 30 0.0356
5.8333 35 0.0401
6.6667 40 0.0379
7.5 45 0.1162
8.3333 50 0.1207
9.1667 55 0.1511
10.0 60 0.0806
10.8333 65 0.0953
11.6667 70 0.1007
12.5 75 0.1483
13.3333 80 0.1172
14.1667 85 0.2504
15.0 90 0.2687

Training Time

  • Training: 1.4 minutes

Framework Versions

  • Python: 3.13.14
  • Sentence Transformers: 6.1.0
  • Transformers: 5.17.0
  • PyTorch: 2.14.0+cpu
  • Accelerate: 1.15.0
  • Datasets: 5.0.1
  • Tokenizers: 0.23.2

Additional Resources

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Marqjeev/aizen-vault-embeddings

Papers for Marqjeev/aizen-vault-embeddings