Arabic Triplet Matryoshka V2 Model [ATM2]
26
23 commits
2 linked in READMEs
updated Sep 7, 2025

Arabic-Triplet-Matryoshka-V2-Model is a state-of-the-art Arabic language embedding model based on the sentence-transformers framework. It is fine-tuned from aubmindlab/bert-base-arabertv02 and specifically designed to capture the rich semantic nuances of Arabic text.
It is described in detail in the paper GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Hybrid Loss Training.
This model maps sentences and paragraphs to a 768-dimensional dense vector space, enabling high-quality semantic text operations including:
The model was trained using a combination of two loss functions:
Training parameters:
The model demonstrates exceptional performance on standard Arabic semantic textual similarity benchmarks:
| Model | Dim | # Params. | STS17 | STS22-v2 | Average |
|---|---|---|---|---|---|
| Arabic-Triplet-Matryoshka-V2 | 768 | 135M | 85 | 64 | 75 |
| Arabert-all-nli-triplet-Matryoshka | 768 | 135M | 83 | 64 | 74 |
| AraGemma-Embedding-300m | 768 | 303M | 84 | 62 | 73 |
| GATE-AraBert-V1 | 767 | 135M | 83 | 63 | 73 |
| Marbert-all-nli-triplet-Matryoshka | 768 | 163M | 82 | 61 | 72 |
| Arabic-labse-Matryoshka | 768 | 471M | 82 | 61 | 72 |
| AraEuroBert-Small | 768 | 210M | 80 | 61 | 71 |
| E5-all-nli-triplet-Matryoshka | 384 | 278M | 80 | 60 | 70 |
| text-embedding-3-large | 3072 | - | 81 | 59 | 70 |
| Arabic-all-nli-triplet-Matryoshka | 768 | 135M | 82 | 54 | 68 |
| AraEuroBert-Mid | 1151 | 610M | 83 | 53 | 68 |
| paraphrase-multilingual-mpnet-base-v2 | 768 | 135M | 79 | 55 | 67 |
| AraEuroBert-Large | 2304 | 2.1B | 79 | 55 | 67 |
| text-embedding-ada-002 | 1536 | - | 71 | 62 | 66 |
| text-embedding-3-small | 1536 | - | 72 | 57 | 65 |
This represents the current state-of-the-art for Arabic embedding models, outperforming previous approaches by a significant margin.
This model is particularly well-suited for:
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("Omartificial-Intelligence-Space/Arabic-Triplet-Matryoshka-V2")
# Run inference
sentences = [
'SENTENCE 1',
'SENTENCE 2',
'SENTENCE 3',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
Despite its strong performance, users should be aware of the following limitations:
This model is intended for research and applications that benefit Arabic language processing. Users should be mindful of potential biases that may exist in the training data and the resulting embeddings. We encourage responsible use of this technology and welcome feedback on ways to improve fairness and representation.
If you use the Arabic Matryoshka Embeddings Model in your research or applications, please cite it as follows:
@article{nacar2025gate,
title={GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training},
author={Nacar, Omer and Koubaa, Anis and Sibaee, Serry and Al-Habashi, Yasser and Ammar, Adel and Boulila, Wadii},
journal={arXiv preprint arXiv:2505.24581},
year={2025}
}
We would like to acknowledge AraBERT for the base model and akhooli for the valuable dataset that made this work possible.
Arabic Triplet Matryoshka V2 Model [ATM2]
26
23 commits
2 linked in READMEs
updated Sep 7, 2025

Arabic-Triplet-Matryoshka-V2-Model is a state-of-the-art Arabic language embedding model based on the sentence-transformers framework. It is fine-tuned from aubmindlab/bert-base-arabertv02 and specifically designed to capture the rich semantic nuances of Arabic text.
It is described in detail in the paper GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Hybrid Loss Training.
This model maps sentences and paragraphs to a 768-dimensional dense vector space, enabling high-quality semantic text operations including:
The model was trained using a combination of two loss functions:
Training parameters:
The model demonstrates exceptional performance on standard Arabic semantic textual similarity benchmarks:
| Model | Dim | # Params. | STS17 | STS22-v2 | Average |
|---|---|---|---|---|---|
| Arabic-Triplet-Matryoshka-V2 | 768 | 135M | 85 | 64 | 75 |
| Arabert-all-nli-triplet-Matryoshka | 768 | 135M | 83 | 64 | 74 |
| AraGemma-Embedding-300m | 768 | 303M | 84 | 62 | 73 |
| GATE-AraBert-V1 | 767 | 135M | 83 | 63 | 73 |
| Marbert-all-nli-triplet-Matryoshka | 768 | 163M | 82 | 61 | 72 |
| Arabic-labse-Matryoshka | 768 | 471M | 82 | 61 | 72 |
| AraEuroBert-Small | 768 | 210M | 80 | 61 | 71 |
| E5-all-nli-triplet-Matryoshka | 384 | 278M | 80 | 60 | 70 |
| text-embedding-3-large | 3072 | - | 81 | 59 | 70 |
| Arabic-all-nli-triplet-Matryoshka | 768 | 135M | 82 | 54 | 68 |
| AraEuroBert-Mid | 1151 | 610M | 83 | 53 | 68 |
| paraphrase-multilingual-mpnet-base-v2 | 768 | 135M | 79 | 55 | 67 |
| AraEuroBert-Large | 2304 | 2.1B | 79 | 55 | 67 |
| text-embedding-ada-002 | 1536 | - | 71 | 62 | 66 |
| text-embedding-3-small | 1536 | - | 72 | 57 | 65 |
This represents the current state-of-the-art for Arabic embedding models, outperforming previous approaches by a significant margin.
This model is particularly well-suited for:
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("Omartificial-Intelligence-Space/Arabic-Triplet-Matryoshka-V2")
# Run inference
sentences = [
'SENTENCE 1',
'SENTENCE 2',
'SENTENCE 3',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
Despite its strong performance, users should be aware of the following limitations:
This model is intended for research and applications that benefit Arabic language processing. Users should be mindful of potential biases that may exist in the training data and the resulting embeddings. We encourage responsible use of this technology and welcome feedback on ways to improve fairness and representation.
If you use the Arabic Matryoshka Embeddings Model in your research or applications, please cite it as follows:
@article{nacar2025gate,
title={GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training},
author={Nacar, Omer and Koubaa, Anis and Sibaee, Serry and Al-Habashi, Yasser and Ammar, Adel and Boulila, Wadii},
journal={arXiv preprint arXiv:2505.24581},
year={2025}
}
We would like to acknowledge AraBERT for the base model and akhooli for the valuable dataset that made this work possible.