Instructions to use nau-ai/Nau-Embed-1.2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use nau-ai/Nau-Embed-1.2 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("nau-ai/Nau-Embed-1.2") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Nau Embed is licensed under CC BY-NC-SA 4.0 and is also subject to Google's Gemma Terms of Use, because it is built on EmbeddingGemma. Non-commercial use is free: credit Nau and share anything you build from it under the same licence. Commercial use, including running it inside a paid product or service, needs a separate commercial licence from the Nau authors; ask by opening a discussion titled "Commercial licence" on this repository.
Log in or Sign Up to review the conditions and access this model content.
Nau Embed 1.2 (ނައޫ)
A Dhivehi text-embedding model for search and retrieval. It turns Dhivehi (Thaana), romanised Dhivehi and English text into 768-number vectors, so a question in any of the three finds the right Dhivehi passage.
It is a fine-tune of Google's EmbeddingGemma-300m: 300M parameters, small enough to embed an archive on an ordinary CPU.
New in 1.2: better search over long, real documents. On a document-search benchmark it beats gemini-embedding-2, and it reads passages up to 1,024 tokens (1.0: 512).
Results
Document search (417 search questions — 240 English, 86 Thaana, 89 romanised Dhivehi — over 19,886 passages of 900 characters cut from 1,977 Maldivian official documents, scanned and born-digital; relevance graded per page). R@k is the share of relevant pages found in the top k; S@20 the share of questions with at least one relevant page in the top 20. None of these documents was used for training.
| Model | Width | R@20 | R@40 | R@160 | S@20 | MRR | nDCG@10 |
|---|---|---|---|---|---|---|---|
| Nau Embed 1.2 | 768 | 0.463 | 0.541 | 0.666 | 0.655 | 0.387 | 0.326 |
| Nau Embed 1.2 | 256 | 0.451 | 0.533 | 0.649 | 0.643 | 0.377 | 0.314 |
| gemini-embedding-2 | 3072 | 0.432 | 0.505 | 0.648 | 0.619 | 0.351 | 0.298 |
| gemini-embedding-2 | 768 | 0.418 | 0.485 | 0.642 | 0.609 | 0.342 | 0.294 |
| Nau Embed 1.0 | 768 | 0.370 | 0.436 | 0.572 | 0.542 | 0.280 | 0.237 |
By question language (R@20, Nau Embed 1.2 vs gemini-embedding-2 at 3072): English 0.432 vs 0.431, Thaana 0.514 vs 0.480, romanised Dhivehi 0.498 vs 0.388. English R@160 is level (0.665 each).
NTREX English↔Dhivehi bitext mining (MMTEB task NTREXBitextMining, eng_Latn-div_Thaa, 1,997 sentence pairs, macro F1). NTREX was never used for training.
| Model | Size | en→dv | dv→en |
|---|---|---|---|
| Nau Embed 1.2 | 300M | 0.984 | 0.983 |
| Nau Embed 1.0 | 300M | 0.979 | 0.980 |
| gemini-embedding-001 | API | 0.952 | 0.971 |
| EmbeddingGemma-300m (base) | 300M | 0.386 | 0.394 |
| Qwen3-Embedding 0.6B–8B (MMTEB) | 0.6–8B | 0.09–0.16 | |
| bge-m3 (MMTEB) | 568M | 0.025 | 0.068 |
Dhivehi retrieval (held-out dev set: 6,751 queries against one pool of 12,378 Dhivehi passages; Recall@10).
| Query type | Nau Embed 1.2 | Nau Embed 1.0 | gemini-embedding-001 | Base |
|---|---|---|---|---|
| English question → Dhivehi passage | 0.998 | 0.997 | 0.997 | 0.584 |
| English keyword search → Dhivehi passage | 0.993 | 0.990 | 0.991 | 0.511 |
| Romanised Dhivehi headline → article | 0.957 | 0.951 | 0.762 | 0.394 |
| Thaana headline → article | 0.931 | 0.920 | 0.893 | 0.715 |
| Dhivehi question → answer | 0.682 | 0.750 | 0.601 | 0.365 |
| Word → dictionary definition | 0.655 | 0.632 | 0.549 | 0.357 |
Known weakness: short Dhivehi question → answer matching (148 questions) is below 1.0 (0.682 vs 0.750). If that is your main use, Nau Embed 1.0 is better at it.
English-only retrieval is close to 1.0 and below the base model (NanoBEIR mean nDCG@10: 0.522; 1.0: 0.545; base: 0.614). For English documents with English queries, the base model is better.
Usage
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("nau-ai/Nau-Embed-1.2") # gated: log in and accept the licence first
docs = ["2004 ވަނަ އަހަރު އޭޝިއާ ސަރަހައްދަށް އެރި ސުނާމީގައި މަރުވި ކަމަށް ކަނޑައެޅި އިންޑޮނޭޝިއާގެ ފުލުހަކު ދިރިހުއްޓާ ފެނިއްޖެ އެވެ. ...",
"ކާބޯތަކެއްޗާއި އާންމު ކޮށް ގޭތެރޭގައި ބޭނުން ކުރާ ތަކެތީގެ އަގުތައް އަނެއްކާ ވެސް މައްޗަށް ހިނގައްޖެ ..."]
doc_vecs = model.encode(docs, prompt_name="document") # "title: none | text: "
q_vecs = model.encode(["Was a police officer thought killed in the 2004 tsunami found alive?",
"2004 Ge Tsunami Gai Maruvi Kamah Ninmmi Fuluhaku Dhirihuttaa Fenijje"], prompt_name="query") # "task: search result | query: "
scores = model.similarity(q_vecs, doc_vecs)
- Always encode questions with
prompt_name="query"and passages withprompt_name="document". Leaving the prompts out lowers accuracy. - Query and document vectors must come from this same model. They are not comparable with vectors from any other model, including Nau Embed 1.0: re-embed your documents when you upgrade.
- Passages up to 1,024 tokens. It was trained on 900-character chunks with 150 characters of overlap; chunks of that size work well.
- To match English and Dhivehi sentences one-to-one (translation pairs), use
prompt_name="BitextMining"on both sides. - Vectors can be truncated to 512, 256 or 128 numbers (Matryoshka training). Re-normalise after truncating, and truncate queries and documents the same way. At 256 it still beats gemini-embedding-2 at full width on the benchmark above.
- Use float32 or bfloat16, not float16 (EmbeddingGemma does not support float16).
Training
Three rounds, each starting from the last.
- Nau Embed 1.0: contrastive fine-tuning with in-batch and mined hard negatives inside a Matryoshka loss (768/512/256/128), one epoch over 772k pairs, then merged with the base model (0.8 × fine-tuned + 0.2 × base). The pairs are listed in the 1.0 card; their Dhivehi side is all human-written.
- Nau Embed 1.1 and 1.2, contrastive: 162k search questions over 41,568 passages from Maldivian official documents (separate from the benchmark documents), in four kinds: English, Thaana, romanised Dhivehi and general. The passages are human-written; the questions were written by Gemini. A question was kept only if the passage contains the sentence it was written from, and Thaana questions only if at least 85% of their words are in the Radheef dictionary. A quarter of the 1.0 pairs were replayed alongside so earlier skills were kept. In 1.2 the hard negatives were chosen with gemini-embedding-2, skipping passages it scored almost as high as the right one (likely second right answers).
- Nau Embed 1.2, distillation: 60k questions, each with its right passage and 6 hard negatives. The model learned to rank the 7 passages the way gemini-embedding-2 (3072 dimensions) scores them (KL divergence between the two score distributions).
- Nau Embed 1.2, final weights: the average of the models after steps 2 and 3 (0.5 × each). Distillation alone improved document search but weakened short Dhivehi question → answer matching; the average keeps most of both, and scores higher on document search than either.
All benchmark documents, all dev-set documents and all of NTREX were removed before training.
Licence and attribution
CC BY-NC-SA 4.0, with a commercial licence on request, and subject to the Gemma Terms of Use (see LICENSE and NOTICE).
- Non-commercial use (research, teaching, personal and public-interest projects) is free. Credit Nau Embed, and share anything you build from it (fine-tunes, merges, quantisations, distilled models) under the same licence.
- Commercial use, including running it inside a paid product or service, needs a commercial licence from the Nau authors. Its terms include publishing the improvements you make. To ask for one, open a discussion titled "Commercial licence" on this repository.
- Because it is built on EmbeddingGemma, use is also governed by Google's Gemma Terms of Use and Prohibited Use Policy. These apply to everyone, including commercial licensees.
Access requires accepting these terms in the form at the top of this page.
Related: Nau Embed 1.0, Nau 1.0, the Dhivehi chat model.
Questions or problems: open a discussion on this repository.
- Downloads last month
- 2
Model tree for nau-ai/Nau-Embed-1.2
Base model
google/embeddinggemma-300m