Instructions to use masterofaudio2077/Fada_ar_embedding with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use masterofaudio2077/Fada_ar_embedding with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("masterofaudio2077/Fada_ar_embedding") sentences = [ "هذا شخص سعيد", "هذا كلب سعيد", "هذا شخص سعيد جدا", "اليوم هو يوم مشمس" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
FADA: Arabic text embeddings
Full-resolution poster (4K PNG)
FADA maps Arabic text to a 768-dimensional vector. It is a Matryoshka model: the first 512, 256, 128, 64 dimensions can be used on their own with little loss. It was built in two stages:
- Contrastive fine-tuning. UBC-NLP/ARBERTv2 fine-tuned on Arab3M-Triplets with an in-batch contrastive (InfoNCE) loss and Matryoshka representation learning.
- Distillation. That model was then trained to reproduce the embeddings of Qwen/Qwen3-Embedding-8B on the texts of the same dataset.
The stage-1 recipe is inspired by Swan-Small (Bhatia et al., 2024), which also fine-tunes ARBERTv2 with InfoNCE. This is an independent implementation on different data.
Results
Semantic textual similarity on Arabic, from MTEB (mteb==2.22.0): STS17 (ar-ar) and STS22.v2 (ar).
Score = Spearman correlation x 100 between the model's cosine similarity and human ratings. Higher is better. Evaluated with max_seq_length=512.
| Model | STS17 (ar-ar) | STS22.v2 (ar) | average |
|---|---|---|---|
| raw ARBERTv2, 768 dims | 49.24 | 47.80 | 48.52 |
| FADA before distillation, 768 dims | 74.14 | 52.89 | 63.51 |
| FADA before distillation, 64 dims | 73.11 | 49.87 | 61.49 |
| this model, 768 dims | 79.72 | 58.43 | 69.07 |
| this model, 256 dims | 79.61 | 58.22 | 68.91 |
| this model, 64 dims | 78.29 | 58.37 | 68.33 |
| intfloat/multilingual-e5-small, 384 dims | 74.62 | 61.44 | 68.03 |
| intfloat/multilingual-e5-large, 1024 dims | 78.71 | 63.73 | 71.22 |
| Qwen/Qwen3-Embedding-8B (the teacher), 4096 dims | 88.95 | 69.57 | 79.26 |
Training
Stage 1: contrastive fine-tuning
| Base model | UBC-NLP/ARBERTv2 |
| Data | Arab3M-Triplets (anchor / positive / negative), cleaned and de-duplicated |
| Steps | 23,000 steps (the run was configured for 50,000 and stopped early), batch size 2,048 triplets |
| Loss | MatryoshkaLoss (dims 768, 512, 256, 128, 64, equal weights) around CachedMultipleNegativesRankingLoss (scale 20, GradCache chunk size 512) |
| Optimizer | Lion, learning rate 5e-6, betas (0.9, 0.99), weight decay 0.1 (not applied to biases and norm layers), cosine schedule with 10% warm-up |
| Time and hardware | about 10 hours on one NVIDIA RTX PRO 6000 |
Stage 2: distillation from Qwen/Qwen3-Embedding-8B
| Student | the stage-1 model |
| Teacher | Qwen/Qwen3-Embedding-8B: 4096-dim, last-token pooling, bf16, no instruction prompt, max length 256. Its vectors were computed once and cached (float16). |
| Data | texts only: the unique anchors, positives and negatives of Arab3M-Triplets (1,068,221 texts). The teacher vectors are the only labels; the triplet structure is not used. |
| Steps | 5,636 steps (configured: 10,000), batch size 2,048 = 11,542,528 text samples, about 10.8 passes over the 1,068,221 texts |
| Loss | for each Matryoshka size d in (768, 512, 256, 128, 64): (a) cosine loss, 1 - cos(head_d(normalized student[:d]), teacher); (b) similarity loss, MSE between the student's and the teacher's in-batch cosine-similarity matrices (off-diagonal entries). Both are averaged over the sizes and weighted 10 and 200, the cosine : similarity ratio of Jasper & Stella. Their third (relative-similarity) loss is not used. |
| Heads | one linear layer d -> 4096 per Matryoshka size, initialised by ridge regression on 50,000 texts. Used only during training; not part of this repository. |
| Optimizer | AdamW, learning rate 2e-05 (encoder) and 0.0001 (heads), weight decay 0.01 (not applied to biases and norm layers), gradient clipping 1 |
| Schedule | cosine decay with 10% warm-up, configured for 10,000 steps; training was stopped at step 5,636, so the learning rate had not decayed to zero (last logged value 9.57e-06) |
| Precision / length | bf16, gradient checkpointing, max sequence length 200 tokens during training |
| Hardware | NVIDIA RTX PRO 6000 Blackwell Server Edition (about 2 hours) |
| Seed | 42 |
| Libraries | sentence-transformers 6.1.0, transformers 5.18.0, torch 2.11.0+cu130 |
Distillation loss
kd_cos and kd_sim are the two loss terms after weighting (x10 and x200), each averaged over the 5 Matryoshka sizes; the values are averages over each 25-step logging window.
| Step | loss | kd_cos | kd_sim |
|---|---|---|---|
| 25 | 20.623 | 3.816 | 16.807 |
| 500 | 3.994 | 3.021 | 0.973 |
| 1,000 | 3.633 | 2.779 | 0.854 |
| 2,500 | 3.020 | 2.427 | 0.592 |
| 5,000 | 2.697 | 2.225 | 0.472 |
| 5,625 (last logged) | 2.657 | 2.198 | 0.459 |
Usage
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("masterofaudio2077/Fada_ar_embedding") # 768 dims
# model = SentenceTransformer("masterofaudio2077/Fada_ar_embedding", truncate_dim=256) # Matryoshka: 512, 256, 128 or 64 also work
sentences = ["ما أجمل الليل الهادئ", "سماء الليل مليئة بالنجوم", "أحب القهوة في الصباح"]
embeddings = model.encode(sentences)
print(model.similarity(embeddings, embeddings))
Without Sentence-Transformers (plain transformers, mean pooling over the tokens; slice the first d numbers for a smaller Matryoshka size):
import torch
from transformers import AutoTokenizer, AutoModel
tok = AutoTokenizer.from_pretrained("masterofaudio2077/Fada_ar_embedding")
model = AutoModel.from_pretrained("masterofaudio2077/Fada_ar_embedding").eval()
def embed(texts, dim=768):
batch = tok(texts, padding=True, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
hidden = model(**batch).last_hidden_state
mask = batch["attention_mask"].unsqueeze(-1).float()
pooled = (hidden * mask).sum(1) / mask.sum(1)
return pooled[:, :dim]
vecs = embed(["ما أجمل الليل الهادئ", "سماء الليل مليئة بالنجوم"])
print(torch.nn.functional.cosine_similarity(vecs[0], vecs[1], dim=0))
Notes:
- No query or passage prefix is needed (none was used in training).
- The output vectors are not L2-normalised. Cosine similarity (the default of
model.similarity) is unaffected; usenormalize_embeddings=Trueif you need unit vectors. - The saved default
max_seq_lengthis 200; setmodel.max_seq_length = 512for longer texts, as in the benchmark above. - The weights are stored as
model.safetensors. The repository was saved with sentence-transformers 6.1.0; if loading fails on an older version, runpip install -U sentence-transformers.
Architecture:
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
(1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
)
License and terms
Released under apache-2.0. Please note: the base model ARBERTv2 states no license on its model page, and the training dataset
Arab3M-Triplets is labelled Apache-2.0 on the Hub but its access form asks users to agree to non-commercial use only.
The teacher Qwen3-Embedding-8B is Apache-2.0. Check these terms before any commercial use.
Citations
If you use this model, please also cite the work it builds on. The training data (Arab3M-Triplets) provides no paper or BibTeX entry; cite it by name and link.
@inproceedings{abdul-mageed-etal-2021-arbert,
title = "{ARBERT} {\&} {MARBERT}: Deep Bidirectional Transformers for {A}rabic",
author = "Abdul-Mageed, Muhammad and Elmadany, AbdelRahim and Nagoudi, El Moatez Billah",
booktitle = "Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)",
month = aug, year = "2021", address = "Online", publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2021.acl-long.551", doi = "10.18653/v1/2021.acl-long.551", pages = "7088--7105"
}
@article{elmadany2022orca,
title={ORCA: A Challenging Benchmark for Arabic Language Understanding},
author={Elmadany, AbdelRahim and Nagoudi, El Moatez Billah and Abdul-Mageed, Muhammad},
journal={arXiv preprint arXiv:2212.10758}, year={2022}
}
@article{qwen3embedding,
title={Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models},
author={Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren},
journal={arXiv preprint arXiv:2506.05176},
year={2025}
}
@misc{zhang2025jasperstelladistillationsota,
title={Jasper and Stella: distillation of SOTA embedding models},
author={Dun Zhang and Jiacheng Li and Ziyang Zeng and Fulong Wang},
year={2025}, eprint={2412.19048}, archivePrefix={arXiv}, primaryClass={cs.IR},
url={https://arxiv.org/abs/2412.19048}
}
@article{bhatia2024swan,
title={Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks},
author={Bhatia, Gagan and Nagoudi, El Moatez Billah and El Mekki, Abdellah and Alwajih, Fakhraddin and Abdul-Mageed, Muhammad},
journal={arXiv preprint arXiv:2411.01192}, year={2024}
}
@article{kusupati2022matryoshka,
title={Matryoshka Representation Learning},
author={Kusupati, Aditya and Bhatt, Gantavya and Rege, Aniket and Wallingford, Matthew and Sinha, Aditya and Ramanujan, Vivek and Howard-Snyder, William and Chen, Kaifeng and Kakade, Sham and Jain, Prateek and Farhadi, Ali},
journal={arXiv preprint arXiv:2205.13147}, year={2022}
}
@inproceedings{gao2021scaling,
title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
author={Gao, Luyu and Zhang, Yunyi and Han, Jiawei and Callan, Jamie},
booktitle={RepL4NLP 2021}, year={2021}, note={arXiv:2101.06983}
}
@article{oord2018representation,
title={Representation Learning with Contrastive Predictive Coding},
author={van den Oord, Aaron and Li, Yazhe and Vinyals, Oriol},
journal={arXiv preprint arXiv:1807.03748}, year={2018}
}
@article{chen2023lion,
title={Symbolic Discovery of Optimization Algorithms},
author={Chen, Xiangning and Liang, Chen and Huang, Da and Real, Esteban and Wang, Kaiyuan and Liu, Yao and Pham, Hieu and Dong, Xuanyi and Luong, Thang and Hsieh, Cho-Jui and Lu, Yifeng and Le, Quoc V.},
journal={arXiv preprint arXiv:2302.06675}, year={2023}
}
@inproceedings{loshchilov2019decoupled,
title={Decoupled Weight Decay Regularization},
author={Loshchilov, Ilya and Hutter, Frank},
booktitle={International Conference on Learning Representations}, year={2019}, note={arXiv:1711.05101}
}
@inproceedings{reimers2019sentence,
title={Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks},
author={Reimers, Nils and Gurevych, Iryna},
booktitle={EMNLP 2019}, year={2019}, note={arXiv:1908.10084}
}
@article{muennighoff2022mteb,
title={MTEB: Massive Text Embedding Benchmark},
author={Muennighoff, Niklas and Tazi, Nouamane and Magne, Lo{\"i}c and Reimers, Nils},
journal={arXiv preprint arXiv:2210.07316}, year={2022}
}
@article{wang2024multilingual,
title={Multilingual E5 Text Embeddings: A Technical Report},
author={Wang, Liang and Yang, Nan and Huang, Xiaolong and Yang, Linjun and Majumder, Rangan and Wei, Furu},
journal={arXiv preprint arXiv:2402.05672}, year={2024}
}
Dataset: Omartificial-Intelligence-Space, Arab3M-Triplets, https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arab3M-Triplets
Card generated on 2026-10-02.
- Downloads last month
- -
Model tree for masterofaudio2077/Fada_ar_embedding
Base model
UBC-NLP/ARBERTv2
