Native-Bird

Native-Bird is a 768-dimensional multilingual sentence-embedding model fine-tuned from LaBSE for cross-lingual retrieval involving Igbo, Hausa, Yoruba, and English.

Its primary use case is native-language query → English document retrieval. A user can describe what they are looking for in Igbo while the indexed document is written in English, and Native-Bird maps both into the same embedding space for semantic search.

What it is designed to do

Example:

Igbo query: “Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn.”

English document: “The nucleus consists of two particles - neutrons and protons.”

The model should assign a high cosine similarity to the matching English text even though the languages differ.

Model details

  • Model name: Native-Bird
  • Architecture: SentenceTransformer based on LaBSE
  • Embedding dimension: 768
  • Maximum sequence length used during training: 128 tokens
  • Similarity: cosine similarity
  • Base model: sentence-transformers/LaBSE
  • Training objective: Cached Multiple Negatives Ranking Loss (CachedMNRL)
  • Languages: Igbo, Hausa, Yoruba, English
  • Training examples: 76,150 paired examples
  • Epochs: 4
  • Per-device batch size: 256
  • Learning rate: 2e-5
  • Warmup: 10%
  • Scheduler: cosine
  • Training precision: BF16
  • Negative sampling: in-batch negatives with NO_DUPLICATES; CachedMNRL mini-batch size 128

Evaluation

The retrieval evaluation uses an English corpus of 1,004 sentences and 204 queries per language. For each query, the correct English sentence is known. We report Recall@1, Recall@10, and Mean Reciprocal Rank (MRR).

Fine-tuned Native-Bird

Language Query condition R@1 R@10 MRR
Igbo Clean 1.0000 1.0000 1.0000
Igbo ASR-style noise 1.0000 1.0000 1.0000
Hausa Clean 1.0000 1.0000 1.0000
Hausa ASR-style noise 1.0000 1.0000 1.0000
Yoruba Clean 0.9804 1.0000 0.9902
Yoruba ASR-style noise 0.9804 1.0000 0.9867

Before fine-tuning: LaBSE baseline

Language Query condition R@1 R@10 MRR
Igbo Clean 1.0000 1.0000 1.0000
Igbo ASR-style noise 0.9853 0.9951 0.9888
Hausa Clean 1.0000 1.0000 1.0000
Hausa ASR-style noise 0.9951 1.0000 0.9975
Yoruba Clean 0.9804 1.0000 0.9871
Yoruba ASR-style noise 0.9559 0.9951 0.9683

Interpretation

Native-Bird is strong on the benchmark it was trained and evaluated against, particularly for Igbo and Hausa. Fine-tuning improved noisy-query retrieval for Igbo and Hausa and slightly improved Yoruba.

The scores do not prove that Native-Bird will retrieve arbitrary long English documents from arbitrary Igbo descriptions with the same accuracy. The current evaluation is sentence-level. Real document search should therefore split documents into passages/chunks before embedding.

Direct Igbo → English retrieval checks

The trained model was directly run against the English evaluation corpus. Examples:

Igbo: “Mgbanwe n'ọdịdị na-agbakwunye ọdịdị kejenetiki ọhụrụ, ma nhọpụta na-ewepụ ya n'ogwu ọdịdị nke a kọwapụtara.”

→ Top English result: “Mutation adds new genetic variation, and selection removes it from the pool of expressed variation.”

Cosine similarity: 0.7079

Igbo: “Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn.”

→ Top English result: “The nucleus consists of two particles - neutrons and protons.”

Cosine similarity: 0.7049

Igbo: “Enwere ọtụtụ ihe mere ha ji kara web prọgzi mma: ha na-agbanwe ụzọ trafik ịntanetị niile, ọ bụghị naanị http.”

→ Top English result: “They are superior to web proxies for several reasons: They re-route all Internet traffic, not only http.”

Cosine similarity: 0.6774

These are direct demonstrations of the intended behavior: the query is Igbo while the retrieved target is English.

Installation

pip install -U sentence-transformers torch

Basic usage

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("Modularcomputing/Native-Bird")

query = "Neuklọs ahụ bụ ahụ ihe abụọ mejupụtara ya - neutrọn na protọn."

documents = [
    "The nucleus consists of two particles - neutrons and protons.",
    "The liver is an organ responsible for many metabolic functions.",
    "Photosynthesis converts light energy into chemical energy.",
]

q = model.encode(query, normalize_embeddings=True)
d = model.encode(documents, normalize_embeddings=True)

scores = d @ q

for i in scores.argsort()[::-1]:
    print(float(scores[i]), documents[i])

The matching English sentence should rank first.

Testing the model

A minimal test is:

  1. Load Native-Bird.
  2. Encode an Igbo query.
  3. Encode several English candidate documents.
  4. Compute cosine similarity.
  5. Sort candidates by similarity.
  6. Check whether the semantically matching English document is ranked first.

For a real search system, replace the small list with a vector database or FAISS index.

Searching a real English document collection

Do not embed a multi-page document as one vector. Split each document into passages, normally around 100–300 words with overlap, and store each passage embedding together with its parent document ID.

At query time:

  1. Encode the user's Igbo query with Native-Bird.
  2. Search the English passage vectors using cosine similarity or an ANN index such as FAISS.
  3. Return the highest-scoring passages.
  4. Group or rerank passages by their parent document.

Conceptually:

Igbo query
    ↓
Native-Bird encoder
    ↓
768-dimensional query vector
    ↓ cosine similarity
English passage index
    ↓
Top-k English passages
    ↓
Parent documents

This is the architecture needed for the intended “describe it in Igbo, find the English document” product.

Important limitation

Native-Bird is an embedding/retrieval model, not a translator and not a generative language model. It does not generate an English answer. It represents text as vectors so that semantically related text can be found across languages.

The current model uses a 128-token training sequence length and was evaluated on sentence-level retrieval. For long documents, chunking is therefore important.

A production deployment should be evaluated on the actual target domain. Nigerian government documents, university material, medical documents, legal documents, and technical documentation can have vocabulary and writing styles that differ substantially from the current benchmark.

Evaluation methodology

The evaluation script embeds the 1,004-item English corpus, embeds each language's 204 queries, computes cosine similarities, and measures the rank of the known matching English sentence.

Two query variants are evaluated:

  • clean: original native-language query
  • ASR-style noise: Unicode diacritics removed, lowercased, punctuation stripped, and whitespace normalized

The benchmark is useful for measuring cross-lingual retrieval and robustness to transcription-style normalization, but should not be interpreted as a universal real-world retrieval score.

Intended applications

  • Igbo → English semantic document search
  • Hausa → English semantic document search
  • Yoruba → English semantic document search
  • Multilingual educational search
  • Native-language interfaces for English knowledge bases
  • Cross-lingual RAG retrieval
  • Search over Nigerian institutional documents

Training provenance

Native-Bird was fine-tuned from LaBSE using paired native-language/English retrieval data prepared for this project. Training used 76,150 pairs and a contrastive retrieval objective.

The final training run completed successfully for four epochs, saved the model, and produced the evaluation results reported above.

Model status

Status: research / early release.

The core cross-lingual retrieval capability is demonstrated. The next meaningful validation step is a human-annotated benchmark of real Nigerian documents and real Igbo/Hausa/Yoruba information needs, especially for multi-paragraph documents.

Citation

If you use Native-Bird in research or a project, please reference the model as:

Native-Bird: Cross-lingual Native-Language Embeddings for English Document Retrieval.

Downloads last month
-
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Modularcomputing/Native-Bird

Finetuned
(100)
this model