State-of-the-art Arabic Sentence Embeddings
This model is a Matryoshka representation learning version of AraBERT specifically fine-tuned for Arabic Natural Language Inference (NLI) tasks. It generates embeddings that can be truncated to different dimensions (768, 512, 256, 128, 64) while maintaining strong performance across all sizes.
The model is based on aubmindlab/bert-base-arabertv02 and trained using the Matryoshka Representation Learning approach, which allows for flexible embedding dimensions without retraining.
Our model shows significant improvements over the base AraBERT model across all embedding dimensions:
| Dimension | Matryoshka Accuracy | Base Accuracy | Matryoshka F1 | Base F1 | Improvement |
|---|---|---|---|---|---|
| 768 | 80.3% | 56.8% | 81.15% | 41.94% | +39.21% |
| 512 | 80.6% | 56.9% | 81.36% | 44.32% | +37.05% |
| 256 | 80.95% | 55.65% | 81.42% | 38.7% | +42.72% |
| 128 | 81.25% | 56.7% | 81.37% | 40.6% | +40.77% |
| 64 | 81.0% | 55.8% | 80.51% | 37.92% | +42.59% |
pip install sentence-transformers torch
from sentence_transformers import SentenceTransformer
# Load the model
model = SentenceTransformer('AhmedZaky1/arabic-bert-nli-matryoshka')
# Example sentences
sentences = [
"الطقس جميل اليوم",
"إنه يوم مشمس وجميل",
"أحب قراءة الكتب"
]
# Generate embeddings (default: full 768 dimensions)
embeddings = model.encode(sentences)
print(f"Full embeddings shape: {embeddings.shape}")
# Use different dimensions by truncating
embeddings_256 = embeddings[:, :256] # Use first 256 dimensions
embeddings_128 = embeddings[:, :128] # Use first 128 dimensions
embeddings_64 = embeddings[:, :64] # Use first 64 dimensions
print(f"256-dim embeddings shape: {embeddings_256.shape}")
from sentence_transformers import util
# Compute similarity between sentences
sentence1 = "القطة تجلس على السجادة"
sentence2 = "الكلب يلعب في الحديقة"
embeddings = model.encode([sentence1, sentence2])
similarity = util.cos_sim(embeddings[0], embeddings[1])
print(f"Similarity: {similarity.item():.4f}")
def classify_nli_pair(premise, hypothesis, threshold=0.6):
"""
Classify Natural Language Inference relationship
Args:
premise: The premise sentence
hypothesis: The hypothesis sentence
threshold: Similarity threshold for classification
Returns:
str: 'entailment' if similarity > threshold, else 'contradiction'
"""
embeddings = model.encode([premise, hypothesis])
similarity = util.cos_sim(embeddings[0], embeddings[1]).item()
return 'entailment' if similarity > threshold else 'contradiction'
# Example usage
premise = "الرجل يقرأ كتاباً في المكتبة"
hypothesis = "شخص يقرأ في مكان هادئ"
result = classify_nli_pair(premise, hypothesis)
print(f"Relationship: {result}")
If you use this model in your research, please cite:
@model{arabic-bert-nli-matryoshka,
title={Arabic BERT NLI Matryoshka Embeddings},
author={Ahmed Mouad},
year={2025},
url={https://huggingface.co/AhmedZaky1/arabic-bert-nli-matryoshka}
}
This model is released under the Apache 2.0 License.
Model Version: 1.0
Last Updated: May 2025
Framework: sentence-transformers
Language: Arabic (العربية)
State-of-the-art Arabic Sentence Embeddings
This model is a Matryoshka representation learning version of AraBERT specifically fine-tuned for Arabic Natural Language Inference (NLI) tasks. It generates embeddings that can be truncated to different dimensions (768, 512, 256, 128, 64) while maintaining strong performance across all sizes.
The model is based on aubmindlab/bert-base-arabertv02 and trained using the Matryoshka Representation Learning approach, which allows for flexible embedding dimensions without retraining.
Our model shows significant improvements over the base AraBERT model across all embedding dimensions:
| Dimension | Matryoshka Accuracy | Base Accuracy | Matryoshka F1 | Base F1 | Improvement |
|---|---|---|---|---|---|
| 768 | 80.3% | 56.8% | 81.15% | 41.94% | +39.21% |
| 512 | 80.6% | 56.9% | 81.36% | 44.32% | +37.05% |
| 256 | 80.95% | 55.65% | 81.42% | 38.7% | +42.72% |
| 128 | 81.25% | 56.7% | 81.37% | 40.6% | +40.77% |
| 64 | 81.0% | 55.8% | 80.51% | 37.92% | +42.59% |
pip install sentence-transformers torch
from sentence_transformers import SentenceTransformer
# Load the model
model = SentenceTransformer('AhmedZaky1/arabic-bert-nli-matryoshka')
# Example sentences
sentences = [
"الطقس جميل اليوم",
"إنه يوم مشمس وجميل",
"أحب قراءة الكتب"
]
# Generate embeddings (default: full 768 dimensions)
embeddings = model.encode(sentences)
print(f"Full embeddings shape: {embeddings.shape}")
# Use different dimensions by truncating
embeddings_256 = embeddings[:, :256] # Use first 256 dimensions
embeddings_128 = embeddings[:, :128] # Use first 128 dimensions
embeddings_64 = embeddings[:, :64] # Use first 64 dimensions
print(f"256-dim embeddings shape: {embeddings_256.shape}")
from sentence_transformers import util
# Compute similarity between sentences
sentence1 = "القطة تجلس على السجادة"
sentence2 = "الكلب يلعب في الحديقة"
embeddings = model.encode([sentence1, sentence2])
similarity = util.cos_sim(embeddings[0], embeddings[1])
print(f"Similarity: {similarity.item():.4f}")
def classify_nli_pair(premise, hypothesis, threshold=0.6):
"""
Classify Natural Language Inference relationship
Args:
premise: The premise sentence
hypothesis: The hypothesis sentence
threshold: Similarity threshold for classification
Returns:
str: 'entailment' if similarity > threshold, else 'contradiction'
"""
embeddings = model.encode([premise, hypothesis])
similarity = util.cos_sim(embeddings[0], embeddings[1]).item()
return 'entailment' if similarity > threshold else 'contradiction'
# Example usage
premise = "الرجل يقرأ كتاباً في المكتبة"
hypothesis = "شخص يقرأ في مكان هادئ"
result = classify_nli_pair(premise, hypothesis)
print(f"Relationship: {result}")
If you use this model in your research, please cite:
@model{arabic-bert-nli-matryoshka,
title={Arabic BERT NLI Matryoshka Embeddings},
author={Ahmed Mouad},
year={2025},
url={https://huggingface.co/AhmedZaky1/arabic-bert-nli-matryoshka}
}
This model is released under the Apache 2.0 License.
Model Version: 1.0
Last Updated: May 2025
Framework: sentence-transformers
Language: Arabic (العربية)