shyngys879/kazakh-e5-rag-embedding

Model

Kazakh-E5-RAG-Embedding: Kazakh Embedding Model for RAG

8

17 commits

1 linked in READMEs

updated Aug 14, 2026

See the code

README

Kazakh-E5-RAG-Embedding: Kazakh Embedding Model for RAG

Kazakh-E5-RAG-Embedding is an E5-style text embedding model for Kazakh RAG, semantic search, FAQ search, question-answer retrieval, and document search.

🏆 Best evaluated BASE-size embedding model on our Kazakh hard-negative retrieval benchmark.

The model is built on the multilingual-e5-base architecture and further optimized for Kazakh question-passage matching, hard-negative retrieval, and Kazakh Wikipedia-style document search.

Why use this model?

  • Kazakh-focused retrieval: optimized for Kazakh questions, passages, and document search
  • 🔍 Hard-negative ranking: designed to distinguish correct passages from very similar incorrect passages
  • ⚡ Efficient BASE-size model: 278M parameters
  • 🧠 E5-style format: uses query: and passage: prefixes
  • 🔧 RAG-ready: works with sentence-transformers, vector databases, and retrieval pipelines
  • 📊 Evaluated on 3 Kazakh retrieval benchmarks

Usage

Installation

pip install sentence-transformers numpy

Convert Text to Embeddings

This model converts Kazakh text into 768-dimensional embedding vectors that can be used for semantic search, retrieval, and RAG.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("shyngys879/kazakh-e5-rag-embedding")

text = "passage: Астана — Қазақстан Республикасының астанасы."

embedding = model.encode(text, normalize_embeddings=True)

print(embedding.shape)
# (768,)

print(embedding[:5])
# Example output: [0.021, -0.034, 0.056, ...]

Basic Usage

from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("shyngys879/kazakh-e5-rag-embedding")

query = "query: Қазақстанның астанасы қай қала?"
passage = "passage: Астана — Қазақстан Республикасының астанасы."

query_embedding = model.encode(query, normalize_embeddings=True)
passage_embedding = model.encode(passage, normalize_embeddings=True)

# Since embeddings are normalized, dot product = cosine similarity
similarity = np.dot(query_embedding, passage_embedding)

print(f"Similarity: {similarity:.4f}")
from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("shyngys879/kazakh-e5-rag-embedding")

documents = [
    "Астана — Қазақстан Республикасының астанасы.",
    "Алматы — Қазақстанның ең үлкен қаласы.",
    "Қазақстан — Орталық Азиядағы мемлекет.",
    "Абай Құнанбайұлы — қазақтың ұлы ақыны және ағартушысы.",
]

query = "Қазақстанның астанасы қай қала?"

doc_embeddings = model.encode(
    ["passage: " + doc for doc in documents],
    normalize_embeddings=True
)

query_embedding = model.encode(
    "query: " + query,
    normalize_embeddings=True
)

scores = np.dot(doc_embeddings, query_embedding)
best_idx = np.argmax(scores)

print("Question:", query)
print("Best match:", documents[best_idx])
print("Score:", float(scores[best_idx]))

This same pattern can be used for FAQ search, document retrieval, and RAG pipelines: encode your documents, retrieve the most relevant passages, then pass them to an LLM as context.

Important Prefix Format

For best results, use E5-style prefixes:

Input typePrefix
Query / questionquery: ...
Passage / documentpassage: ...

Benchmark Results

The model was evaluated on three Kazakh retrieval settings:

  1. OfficialKazQAD-HardTFIDF99 — hard-negative question-passage retrieval.
  2. WikiFullCorpus — Kazakh Wikipedia-style full-corpus retrieval.
  3. KazQAD-100 local — KazQAD-style retrieval with 100 candidates per query.

MRR means Mean Reciprocal Rank: higher is better, and it rewards models that rank the correct passage closer to the top.


OfficialKazQAD-HardTFIDF99

This benchmark uses 1,929 KazQAD test queries.
For each query, the model must rank the correct passage among 100 candidate passages.

The candidate set contains:

  • 1 correct passage
  • 99 TF-IDF hard negatives

A hard negative is an incorrect passage that is textually or topically similar to the query. This makes the task harder than retrieval with random negative passages, because the model must identify which similar passage actually answers the question.

ModelHits@1Hits@5MRRParams
Kazakh-E5-RAG-Embedding35.98%72.47%0.5189278M
KazEmbed-V5 original30.07%65.68%0.4619278M
multilingual-e5-large30.17%62.99%0.4490560M
multilingual-e5-base25.87%56.71%0.4048278M
paraphrase-multilingual-mpnet-base-v210.01%29.91%0.2082278M
LaBSE7.47%26.75%0.1821471M

Additional Benchmarks

BenchmarkModelHits@1Hits@5MRRParams
WikiFullCorpusmultilingual-e5-large69.13%85.79%0.7656560M
WikiFullCorpusmultilingual-e5-base65.85%80.87%0.7290278M
WikiFullCorpusKazakh-E5-RAG-Embedding60.38%74.32%0.6689278M
WikiFullCorpusKazEmbed-V5 original56.83%70.49%0.6276278M
KazQAD-100 localKazakh-E5-RAG-Embedding91.55%100.00%0.9554278M
KazQAD-100 localKazEmbed-V5 original87.32%100.00%0.9334278M
KazQAD-100 localmultilingual-e5-large86.27%98.24%0.9189560M
KazQAD-100 localmultilingual-e5-base85.56%97.54%0.9105278M

Benchmark notes:

  • WikiFullCorpus evaluates full-corpus retrieval over Kazakh Wikipedia-style passages. The model must find the correct passage from a larger document corpus, which makes this closer to practical semantic search and RAG.
  • KazQAD-100 local evaluates Kazakh question-passage retrieval with 100 candidates per query. It is a supporting benchmark for checking whether the model ranks the correct answer passage near the top.

Model Details

FieldValue
Model nameshyngys879/kazakh-e5-rag-embedding
Model typeText embedding / bi-encoder retrieval model
Architecture familymultilingual-e5-base / XLM-RoBERTa
Continued fine-tuning fromNurlykhan/kazembed-v5
Model lineageintfloat/multilingual-e5-base → Nurlykhan/kazembed-v5 → this model
Parameters278M
Embedding dimension768
Main languageKazakh
TaskRetrieval, semantic search, RAG, question-passage matching
Training objectiveKazakh retrieval / question-passage matching
Training dataKazQAD-style retrieval data, TF-IDF hard negatives, Kazakh Wikipedia-style retrieval examples

  • Kazakh RAG systems
  • Kazakh semantic search
  • FAQ search
  • document retrieval
  • question-answer retrieval
  • educational search
  • Kazakh Wikipedia / encyclopedic search
  • multilingual projects involving Kazakh

Limitations

  • The model is optimized mainly for Kazakh retrieval and RAG.
  • On WikiFullCorpus, multilingual-e5-base and multilingual-e5-large perform better.

Citation

@misc{kazakh-e5-rag-embedding-2026,
  title={Kazakh-E5-RAG-Embedding: A Kazakh Retrieval Embedding Model for RAG},
  author={Shyngys Sovetkhan},
  year={2026},
  howpublished={Hugging Face Model Hub},
  url={https://huggingface.co/shyngys879/kazakh-e5-rag-embedding}
}

Acknowledgements

This model builds on:

  • Nurlykhan/kazembed-v5
  • intfloat/multilingual-e5-base
  • KazQAD / KazQAD-style retrieval data
  • Kazakh Wikipedia-style retrieval examples

Contact, Custom Fine-Tuning, and Model Use

For commercial deployment, RAG or semantic-search integration, model evaluation, or custom fine-tuning on your organization's documents and domain-specific data, please contact:

Shyngys Sovetkhan
Email: shyngyssovetkhan0@gmail.com

If your company or organization uses this model, has evaluated it, or has found it useful, I would greatly appreciate a brief official confirmation letter describing the use case or evaluation results. Such documentation helps demonstrate the model's real-world impact in my university applications. The letter can be sent to the email address above.

e5
embeddings
endpoints_compatible
feature-extraction
kazakh
retrieval
safetensors
semantic-search
sentence-similarity
sentence-transformers
text-embeddings-inference
xlm-roberta

shyngys879/kazakh-e5-rag-embedding

Model

Kazakh-E5-RAG-Embedding: Kazakh Embedding Model for RAG

8

17 commits

1 linked in READMEs

updated Aug 14, 2026

See the code

README

Kazakh-E5-RAG-Embedding: Kazakh Embedding Model for RAG

Kazakh-E5-RAG-Embedding is an E5-style text embedding model for Kazakh RAG, semantic search, FAQ search, question-answer retrieval, and document search.

🏆 Best evaluated BASE-size embedding model on our Kazakh hard-negative retrieval benchmark.

The model is built on the multilingual-e5-base architecture and further optimized for Kazakh question-passage matching, hard-negative retrieval, and Kazakh Wikipedia-style document search.

Why use this model?

  • Kazakh-focused retrieval: optimized for Kazakh questions, passages, and document search
  • 🔍 Hard-negative ranking: designed to distinguish correct passages from very similar incorrect passages
  • ⚡ Efficient BASE-size model: 278M parameters
  • 🧠 E5-style format: uses query: and passage: prefixes
  • 🔧 RAG-ready: works with sentence-transformers, vector databases, and retrieval pipelines
  • 📊 Evaluated on 3 Kazakh retrieval benchmarks

Usage

Installation

pip install sentence-transformers numpy

Convert Text to Embeddings

This model converts Kazakh text into 768-dimensional embedding vectors that can be used for semantic search, retrieval, and RAG.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("shyngys879/kazakh-e5-rag-embedding")

text = "passage: Астана — Қазақстан Республикасының астанасы."

embedding = model.encode(text, normalize_embeddings=True)

print(embedding.shape)
# (768,)

print(embedding[:5])
# Example output: [0.021, -0.034, 0.056, ...]

Basic Usage

from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("shyngys879/kazakh-e5-rag-embedding")

query = "query: Қазақстанның астанасы қай қала?"
passage = "passage: Астана — Қазақстан Республикасының астанасы."

query_embedding = model.encode(query, normalize_embeddings=True)
passage_embedding = model.encode(passage, normalize_embeddings=True)

# Since embeddings are normalized, dot product = cosine similarity
similarity = np.dot(query_embedding, passage_embedding)

print(f"Similarity: {similarity:.4f}")
from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("shyngys879/kazakh-e5-rag-embedding")

documents = [
    "Астана — Қазақстан Республикасының астанасы.",
    "Алматы — Қазақстанның ең үлкен қаласы.",
    "Қазақстан — Орталық Азиядағы мемлекет.",
    "Абай Құнанбайұлы — қазақтың ұлы ақыны және ағартушысы.",
]

query = "Қазақстанның астанасы қай қала?"

doc_embeddings = model.encode(
    ["passage: " + doc for doc in documents],
    normalize_embeddings=True
)

query_embedding = model.encode(
    "query: " + query,
    normalize_embeddings=True
)

scores = np.dot(doc_embeddings, query_embedding)
best_idx = np.argmax(scores)

print("Question:", query)
print("Best match:", documents[best_idx])
print("Score:", float(scores[best_idx]))

This same pattern can be used for FAQ search, document retrieval, and RAG pipelines: encode your documents, retrieve the most relevant passages, then pass them to an LLM as context.

Important Prefix Format

For best results, use E5-style prefixes:

Input typePrefix
Query / questionquery: ...
Passage / documentpassage: ...

Benchmark Results

The model was evaluated on three Kazakh retrieval settings:

  1. OfficialKazQAD-HardTFIDF99 — hard-negative question-passage retrieval.
  2. WikiFullCorpus — Kazakh Wikipedia-style full-corpus retrieval.
  3. KazQAD-100 local — KazQAD-style retrieval with 100 candidates per query.

MRR means Mean Reciprocal Rank: higher is better, and it rewards models that rank the correct passage closer to the top.


OfficialKazQAD-HardTFIDF99

This benchmark uses 1,929 KazQAD test queries.
For each query, the model must rank the correct passage among 100 candidate passages.

The candidate set contains:

  • 1 correct passage
  • 99 TF-IDF hard negatives

A hard negative is an incorrect passage that is textually or topically similar to the query. This makes the task harder than retrieval with random negative passages, because the model must identify which similar passage actually answers the question.

ModelHits@1Hits@5MRRParams
Kazakh-E5-RAG-Embedding35.98%72.47%0.5189278M
KazEmbed-V5 original30.07%65.68%0.4619278M
multilingual-e5-large30.17%62.99%0.4490560M
multilingual-e5-base25.87%56.71%0.4048278M
paraphrase-multilingual-mpnet-base-v210.01%29.91%0.2082278M
LaBSE7.47%26.75%0.1821471M

Additional Benchmarks

BenchmarkModelHits@1Hits@5MRRParams
WikiFullCorpusmultilingual-e5-large69.13%85.79%0.7656560M
WikiFullCorpusmultilingual-e5-base65.85%80.87%0.7290278M
WikiFullCorpusKazakh-E5-RAG-Embedding60.38%74.32%0.6689278M
WikiFullCorpusKazEmbed-V5 original56.83%70.49%0.6276278M
KazQAD-100 localKazakh-E5-RAG-Embedding91.55%100.00%0.9554278M
KazQAD-100 localKazEmbed-V5 original87.32%100.00%0.9334278M
KazQAD-100 localmultilingual-e5-large86.27%98.24%0.9189560M
KazQAD-100 localmultilingual-e5-base85.56%97.54%0.9105278M

Benchmark notes:

  • WikiFullCorpus evaluates full-corpus retrieval over Kazakh Wikipedia-style passages. The model must find the correct passage from a larger document corpus, which makes this closer to practical semantic search and RAG.
  • KazQAD-100 local evaluates Kazakh question-passage retrieval with 100 candidates per query. It is a supporting benchmark for checking whether the model ranks the correct answer passage near the top.

Model Details

FieldValue
Model nameshyngys879/kazakh-e5-rag-embedding
Model typeText embedding / bi-encoder retrieval model
Architecture familymultilingual-e5-base / XLM-RoBERTa
Continued fine-tuning fromNurlykhan/kazembed-v5
Model lineageintfloat/multilingual-e5-base → Nurlykhan/kazembed-v5 → this model
Parameters278M
Embedding dimension768
Main languageKazakh
TaskRetrieval, semantic search, RAG, question-passage matching
Training objectiveKazakh retrieval / question-passage matching
Training dataKazQAD-style retrieval data, TF-IDF hard negatives, Kazakh Wikipedia-style retrieval examples

  • Kazakh RAG systems
  • Kazakh semantic search
  • FAQ search
  • document retrieval
  • question-answer retrieval
  • educational search
  • Kazakh Wikipedia / encyclopedic search
  • multilingual projects involving Kazakh

Limitations

  • The model is optimized mainly for Kazakh retrieval and RAG.
  • On WikiFullCorpus, multilingual-e5-base and multilingual-e5-large perform better.

Citation

@misc{kazakh-e5-rag-embedding-2026,
  title={Kazakh-E5-RAG-Embedding: A Kazakh Retrieval Embedding Model for RAG},
  author={Shyngys Sovetkhan},
  year={2026},
  howpublished={Hugging Face Model Hub},
  url={https://huggingface.co/shyngys879/kazakh-e5-rag-embedding}
}

Acknowledgements

This model builds on:

  • Nurlykhan/kazembed-v5
  • intfloat/multilingual-e5-base
  • KazQAD / KazQAD-style retrieval data
  • Kazakh Wikipedia-style retrieval examples

Contact, Custom Fine-Tuning, and Model Use

For commercial deployment, RAG or semantic-search integration, model evaluation, or custom fine-tuning on your organization's documents and domain-specific data, please contact:

Shyngys Sovetkhan
Email: shyngyssovetkhan0@gmail.com

If your company or organization uses this model, has evaluated it, or has found it useful, I would greatly appreciate a brief official confirmation letter describing the use case or evaluation results. Such documentation helps demonstrate the model's real-world impact in my university applications. The letter can be sent to the email address above.

e5
embeddings
endpoints_compatible
feature-extraction
kazakh
retrieval
safetensors
semantic-search
sentence-similarity
sentence-transformers
text-embeddings-inference
xlm-roberta