FADA: Arabic text embeddings

FADA: Arabic text embeddings, mapped like stars

Full-resolution poster (4K PNG)

FADA maps Arabic text to a 768-dimensional vector. It is a Matryoshka model: the first 512, 256, 128, 64 dimensions can be used on their own with little loss. It was built in two stages:

  1. Contrastive fine-tuning. UBC-NLP/ARBERTv2 fine-tuned on Arab3M-Triplets with an in-batch contrastive (InfoNCE) loss and Matryoshka representation learning.
  2. Distillation. That model was then trained to reproduce the embeddings of Qwen/Qwen3-Embedding-8B on the texts of the same dataset.

The stage-1 recipe is inspired by Swan-Small (Bhatia et al., 2024), which also fine-tunes ARBERTv2 with InfoNCE. This is an independent implementation on different data.

Results

Semantic textual similarity on Arabic, from MTEB (mteb==2.22.0): STS17 (ar-ar) and STS22.v2 (ar). Score = Spearman correlation x 100 between the model's cosine similarity and human ratings. Higher is better. Evaluated with max_seq_length=512.

Model STS17 (ar-ar) STS22.v2 (ar) average
raw ARBERTv2, 768 dims 49.24 47.80 48.52
FADA before distillation, 768 dims 74.14 52.89 63.51
FADA before distillation, 64 dims 73.11 49.87 61.49
this model, 768 dims 79.72 58.43 69.07
this model, 256 dims 79.61 58.22 68.91
this model, 64 dims 78.29 58.37 68.33
intfloat/multilingual-e5-small, 384 dims 74.62 61.44 68.03
intfloat/multilingual-e5-large, 1024 dims 78.71 63.73 71.22
Qwen/Qwen3-Embedding-8B (the teacher), 4096 dims 88.95 69.57 79.26

Training

Stage 1: contrastive fine-tuning

Base model UBC-NLP/ARBERTv2
Data Arab3M-Triplets (anchor / positive / negative), cleaned and de-duplicated
Steps 23,000 steps (the run was configured for 50,000 and stopped early), batch size 2,048 triplets
Loss MatryoshkaLoss (dims 768, 512, 256, 128, 64, equal weights) around CachedMultipleNegativesRankingLoss (scale 20, GradCache chunk size 512)
Optimizer Lion, learning rate 5e-6, betas (0.9, 0.99), weight decay 0.1 (not applied to biases and norm layers), cosine schedule with 10% warm-up
Time and hardware about 10 hours on one NVIDIA RTX PRO 6000

Stage 2: distillation from Qwen/Qwen3-Embedding-8B

Student the stage-1 model
Teacher Qwen/Qwen3-Embedding-8B: 4096-dim, last-token pooling, bf16, no instruction prompt, max length 256. Its vectors were computed once and cached (float16).
Data texts only: the unique anchors, positives and negatives of Arab3M-Triplets (1,068,221 texts). The teacher vectors are the only labels; the triplet structure is not used.
Steps 5,636 steps (configured: 10,000), batch size 2,048 = 11,542,528 text samples, about 10.8 passes over the 1,068,221 texts
Loss for each Matryoshka size d in (768, 512, 256, 128, 64): (a) cosine loss, 1 - cos(head_d(normalized student[:d]), teacher); (b) similarity loss, MSE between the student's and the teacher's in-batch cosine-similarity matrices (off-diagonal entries). Both are averaged over the sizes and weighted 10 and 200, the cosine : similarity ratio of Jasper & Stella. Their third (relative-similarity) loss is not used.
Heads one linear layer d -> 4096 per Matryoshka size, initialised by ridge regression on 50,000 texts. Used only during training; not part of this repository.
Optimizer AdamW, learning rate 2e-05 (encoder) and 0.0001 (heads), weight decay 0.01 (not applied to biases and norm layers), gradient clipping 1
Schedule cosine decay with 10% warm-up, configured for 10,000 steps; training was stopped at step 5,636, so the learning rate had not decayed to zero (last logged value 9.57e-06)
Precision / length bf16, gradient checkpointing, max sequence length 200 tokens during training
Hardware NVIDIA RTX PRO 6000 Blackwell Server Edition (about 2 hours)
Seed 42
Libraries sentence-transformers 6.1.0, transformers 5.18.0, torch 2.11.0+cu130

Distillation loss

distillation loss

kd_cos and kd_sim are the two loss terms after weighting (x10 and x200), each averaged over the 5 Matryoshka sizes; the values are averages over each 25-step logging window.

Step loss kd_cos kd_sim
25 20.623 3.816 16.807
500 3.994 3.021 0.973
1,000 3.633 2.779 0.854
2,500 3.020 2.427 0.592
5,000 2.697 2.225 0.472
5,625 (last logged) 2.657 2.198 0.459

Usage

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("masterofaudio2077/Fada_ar_embedding")                      # 768 dims
# model = SentenceTransformer("masterofaudio2077/Fada_ar_embedding", truncate_dim=256)  # Matryoshka: 512, 256, 128 or 64 also work

sentences = ["ما أجمل الليل الهادئ", "سماء الليل مليئة بالنجوم", "أحب القهوة في الصباح"]
embeddings = model.encode(sentences)
print(model.similarity(embeddings, embeddings))

Without Sentence-Transformers (plain transformers, mean pooling over the tokens; slice the first d numbers for a smaller Matryoshka size):

import torch
from transformers import AutoTokenizer, AutoModel

tok = AutoTokenizer.from_pretrained("masterofaudio2077/Fada_ar_embedding")
model = AutoModel.from_pretrained("masterofaudio2077/Fada_ar_embedding").eval()

def embed(texts, dim=768):
    batch = tok(texts, padding=True, truncation=True, max_length=512, return_tensors="pt")
    with torch.no_grad():
        hidden = model(**batch).last_hidden_state
    mask = batch["attention_mask"].unsqueeze(-1).float()
    pooled = (hidden * mask).sum(1) / mask.sum(1)
    return pooled[:, :dim]

vecs = embed(["ما أجمل الليل الهادئ", "سماء الليل مليئة بالنجوم"])
print(torch.nn.functional.cosine_similarity(vecs[0], vecs[1], dim=0))

Notes:

  • No query or passage prefix is needed (none was used in training).
  • The output vectors are not L2-normalised. Cosine similarity (the default of model.similarity) is unaffected; use normalize_embeddings=True if you need unit vectors.
  • The saved default max_seq_length is 200; set model.max_seq_length = 512 for longer texts, as in the benchmark above.
  • The weights are stored as model.safetensors. The repository was saved with sentence-transformers 6.1.0; if loading fails on an older version, run pip install -U sentence-transformers.

Architecture:

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
)

License and terms

Released under apache-2.0. Please note: the base model ARBERTv2 states no license on its model page, and the training dataset Arab3M-Triplets is labelled Apache-2.0 on the Hub but its access form asks users to agree to non-commercial use only. The teacher Qwen3-Embedding-8B is Apache-2.0. Check these terms before any commercial use.

Citations

If you use this model, please also cite the work it builds on. The training data (Arab3M-Triplets) provides no paper or BibTeX entry; cite it by name and link.

@inproceedings{abdul-mageed-etal-2021-arbert,
    title = "{ARBERT} {\&} {MARBERT}: Deep Bidirectional Transformers for {A}rabic",
    author = "Abdul-Mageed, Muhammad and Elmadany, AbdelRahim and Nagoudi, El Moatez Billah",
    booktitle = "Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)",
    month = aug, year = "2021", address = "Online", publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.acl-long.551", doi = "10.18653/v1/2021.acl-long.551", pages = "7088--7105"
}

@article{elmadany2022orca,
  title={ORCA: A Challenging Benchmark for Arabic Language Understanding},
  author={Elmadany, AbdelRahim and Nagoudi, El Moatez Billah and Abdul-Mageed, Muhammad},
  journal={arXiv preprint arXiv:2212.10758}, year={2022}
}

@article{qwen3embedding,
  title={Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models},
  author={Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren},
  journal={arXiv preprint arXiv:2506.05176},
  year={2025}
}

@misc{zhang2025jasperstelladistillationsota,
  title={Jasper and Stella: distillation of SOTA embedding models},
  author={Dun Zhang and Jiacheng Li and Ziyang Zeng and Fulong Wang},
  year={2025}, eprint={2412.19048}, archivePrefix={arXiv}, primaryClass={cs.IR},
  url={https://arxiv.org/abs/2412.19048}
}

@article{bhatia2024swan,
  title={Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks},
  author={Bhatia, Gagan and Nagoudi, El Moatez Billah and El Mekki, Abdellah and Alwajih, Fakhraddin and Abdul-Mageed, Muhammad},
  journal={arXiv preprint arXiv:2411.01192}, year={2024}
}

@article{kusupati2022matryoshka,
  title={Matryoshka Representation Learning},
  author={Kusupati, Aditya and Bhatt, Gantavya and Rege, Aniket and Wallingford, Matthew and Sinha, Aditya and Ramanujan, Vivek and Howard-Snyder, William and Chen, Kaifeng and Kakade, Sham and Jain, Prateek and Farhadi, Ali},
  journal={arXiv preprint arXiv:2205.13147}, year={2022}
}

@inproceedings{gao2021scaling,
  title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
  author={Gao, Luyu and Zhang, Yunyi and Han, Jiawei and Callan, Jamie},
  booktitle={RepL4NLP 2021}, year={2021}, note={arXiv:2101.06983}
}

@article{oord2018representation,
  title={Representation Learning with Contrastive Predictive Coding},
  author={van den Oord, Aaron and Li, Yazhe and Vinyals, Oriol},
  journal={arXiv preprint arXiv:1807.03748}, year={2018}
}

@article{chen2023lion,
  title={Symbolic Discovery of Optimization Algorithms},
  author={Chen, Xiangning and Liang, Chen and Huang, Da and Real, Esteban and Wang, Kaiyuan and Liu, Yao and Pham, Hieu and Dong, Xuanyi and Luong, Thang and Hsieh, Cho-Jui and Lu, Yifeng and Le, Quoc V.},
  journal={arXiv preprint arXiv:2302.06675}, year={2023}
}

@inproceedings{loshchilov2019decoupled,
  title={Decoupled Weight Decay Regularization},
  author={Loshchilov, Ilya and Hutter, Frank},
  booktitle={International Conference on Learning Representations}, year={2019}, note={arXiv:1711.05101}
}

@inproceedings{reimers2019sentence,
  title={Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks},
  author={Reimers, Nils and Gurevych, Iryna},
  booktitle={EMNLP 2019}, year={2019}, note={arXiv:1908.10084}
}

@article{muennighoff2022mteb,
  title={MTEB: Massive Text Embedding Benchmark},
  author={Muennighoff, Niklas and Tazi, Nouamane and Magne, Lo{\"i}c and Reimers, Nils},
  journal={arXiv preprint arXiv:2210.07316}, year={2022}
}

@article{wang2024multilingual,
  title={Multilingual E5 Text Embeddings: A Technical Report},
  author={Wang, Liang and Yang, Nan and Huang, Xiaolong and Yang, Linjun and Majumder, Rangan and Wei, Furu},
  journal={arXiv preprint arXiv:2402.05672}, year={2024}
}

Dataset: Omartificial-Intelligence-Space, Arab3M-Triplets, https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arab3M-Triplets

Card generated on 2026-10-02.

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masterofaudio2077/Fada_ar_embedding

Base model

UBC-NLP/ARBERTv2
Finetuned
(12)
this model

Dataset used to train masterofaudio2077/Fada_ar_embedding

Papers for masterofaudio2077/Fada_ar_embedding