Evidence-based benchmark for Kazakh information retrieval β the independent proof base for the Kazakh Stemmer.
Corpus: 8,370 passages from Kazakh Wikipedia
Queries: 300 queries Γ 3 categories (natural / inflected / vocabulary-gap)
Format: BEIR-compatible β three subsets: corpus, queries, qrels
Browse the data: use the subset switcher at the top of the Data Studio viewer to move between
corpus(Kazakh passages),queries(the 300 questions), andqrels(relevance judgements). The default view is the corpus.
A Kazakh morphological stemmer significantly improves lexical search: +16% nDCG@10 on inflected queries (p=0.0017) and +9% overall (p=0.0001), on 300 queries with paired-bootstrap significance (10k resamples). BM25+stemmer is also the most balanced retriever overall and beats Dense LaBSE on every category (nDCG@10 0.754 vs 0.481).
| System | inflected | natural | vocab-gap | ALL |
|---|---|---|---|---|
| BM25 (no normalization) | 0.627 | 0.703 | 0.741 | 0.690 |
| BM25 + Kazakh stemmer | 0.727 | 0.772 | 0.764 | 0.754 |
| Dense β LaBSE | 0.477 | 0.546 | 0.419 | 0.481 |
| Dense β Granite-278m | 0.791 | 0.923 | 0.303 | 0.672 |
| Dense β multilingual-E5-base | 0.845 | 0.947 | 0.562 | 0.785 |
End-to-end RAG (Qwen2.5-7B, n=300) is reported honestly: the stemmer raises retrieval hit@3 (0.737 β 0.803), but the end-to-end accuracy gain is not statistically significant (McNemar p=0.63) β the bottleneck is the Kazakh-language generator, not the retriever. The stemmer's value is proven at the retrieval level.
See github.com/Tim2190/Kaz-RAG-search-benchmark for full results, code, and reproduction instructions.
This dataset is the in-domain (Wikipedia) half of a larger evaluation. The full study now covers 14 embedding systems across two domains β this Wikipedia set (in-domain) and an out-of-domain set of Akorda formal speeches β with a second preprint on hybrid retrieval and out-of-domain robustness.
| # | Model | Wiki nDCG@10 | Akorda nDCG@10 |
|---|---|---|---|
| 1 | BGE-M3 | 0.866 | 0.679 |
| 2 | Jina v3 | 0.821 | 0.613 |
| 3 | Cohere embed-v4 | 0.800 | 0.367 |
| 4 | multilingual-e5-base | 0.785 | 0.509 |
| 5 | BM25 + Kazakh stemmer | 0.754 | 0.517 |
(top 5 shown β full 14-system table with hybrids and significance in the repo)
Headline findings:
[UNK] collapse (Nomic) wreck Kazakh retrieval regardless of model size.β‘οΈ Full leaderboard (14 systems, hybrids, significance): github.com/Tim2190/Kaz-RAG-search-benchmark
π Second preprint (hybrid retrieval & OOD robustness): https://doi.org/10.5281/zenodo.20781386
The corpus is derived from Kazakh Wikipedia (the wikimedia/wikipedia dump on
Hugging Face), licensed under CC BY-SA 4.0. This dataset inherits the same license:
attribute the source and share derivatives alike.
from datasets import load_dataset
corpus = load_dataset("Tim2190/kaz-rag-search-benchmark", "corpus", split="corpus")
queries = load_dataset("Tim2190/kaz-rag-search-benchmark", "queries", split="queries")
qrels = load_dataset("Tim2190/kaz-rag-search-benchmark", "qrels", split="test")
This dataset accompanies a preprint archived on Zenodo with a permanent DOI. If you use the benchmark, please cite:
Seidalin, T. (2026). Morphology Beats Multilingual Embeddings for Kazakh Retrieval: A 300-Query Benchmark with Honest Negative Results. Zenodo. https://doi.org/10.5281/zenodo.20605663
@misc{seidalin2026kazakh,
author = {Seidalin, Timur},
title = {Morphology Beats Multilingual Embeddings for Kazakh Retrieval: A 300-Query Benchmark with Honest Negative Results},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.20605663},
url = {https://doi.org/10.5281/zenodo.20605663}
}
Evidence-based benchmark for Kazakh information retrieval β the independent proof base for the Kazakh Stemmer.
Corpus: 8,370 passages from Kazakh Wikipedia
Queries: 300 queries Γ 3 categories (natural / inflected / vocabulary-gap)
Format: BEIR-compatible β three subsets: corpus, queries, qrels
Browse the data: use the subset switcher at the top of the Data Studio viewer to move between
corpus(Kazakh passages),queries(the 300 questions), andqrels(relevance judgements). The default view is the corpus.
A Kazakh morphological stemmer significantly improves lexical search: +16% nDCG@10 on inflected queries (p=0.0017) and +9% overall (p=0.0001), on 300 queries with paired-bootstrap significance (10k resamples). BM25+stemmer is also the most balanced retriever overall and beats Dense LaBSE on every category (nDCG@10 0.754 vs 0.481).
| System | inflected | natural | vocab-gap | ALL |
|---|---|---|---|---|
| BM25 (no normalization) | 0.627 | 0.703 | 0.741 | 0.690 |
| BM25 + Kazakh stemmer | 0.727 | 0.772 | 0.764 | 0.754 |
| Dense β LaBSE | 0.477 | 0.546 | 0.419 | 0.481 |
| Dense β Granite-278m | 0.791 | 0.923 | 0.303 | 0.672 |
| Dense β multilingual-E5-base | 0.845 | 0.947 | 0.562 | 0.785 |
End-to-end RAG (Qwen2.5-7B, n=300) is reported honestly: the stemmer raises retrieval hit@3 (0.737 β 0.803), but the end-to-end accuracy gain is not statistically significant (McNemar p=0.63) β the bottleneck is the Kazakh-language generator, not the retriever. The stemmer's value is proven at the retrieval level.
See github.com/Tim2190/Kaz-RAG-search-benchmark for full results, code, and reproduction instructions.
This dataset is the in-domain (Wikipedia) half of a larger evaluation. The full study now covers 14 embedding systems across two domains β this Wikipedia set (in-domain) and an out-of-domain set of Akorda formal speeches β with a second preprint on hybrid retrieval and out-of-domain robustness.
| # | Model | Wiki nDCG@10 | Akorda nDCG@10 |
|---|---|---|---|
| 1 | BGE-M3 | 0.866 | 0.679 |
| 2 | Jina v3 | 0.821 | 0.613 |
| 3 | Cohere embed-v4 | 0.800 | 0.367 |
| 4 | multilingual-e5-base | 0.785 | 0.509 |
| 5 | BM25 + Kazakh stemmer | 0.754 | 0.517 |
(top 5 shown β full 14-system table with hybrids and significance in the repo)
Headline findings:
[UNK] collapse (Nomic) wreck Kazakh retrieval regardless of model size.β‘οΈ Full leaderboard (14 systems, hybrids, significance): github.com/Tim2190/Kaz-RAG-search-benchmark
π Second preprint (hybrid retrieval & OOD robustness): https://doi.org/10.5281/zenodo.20781386
The corpus is derived from Kazakh Wikipedia (the wikimedia/wikipedia dump on
Hugging Face), licensed under CC BY-SA 4.0. This dataset inherits the same license:
attribute the source and share derivatives alike.
from datasets import load_dataset
corpus = load_dataset("Tim2190/kaz-rag-search-benchmark", "corpus", split="corpus")
queries = load_dataset("Tim2190/kaz-rag-search-benchmark", "queries", split="queries")
qrels = load_dataset("Tim2190/kaz-rag-search-benchmark", "qrels", split="test")
This dataset accompanies a preprint archived on Zenodo with a permanent DOI. If you use the benchmark, please cite:
Seidalin, T. (2026). Morphology Beats Multilingual Embeddings for Kazakh Retrieval: A 300-Query Benchmark with Honest Negative Results. Zenodo. https://doi.org/10.5281/zenodo.20605663
@misc{seidalin2026kazakh,
author = {Seidalin, Timur},
title = {Morphology Beats Multilingual Embeddings for Kazakh Retrieval: A 300-Query Benchmark with Honest Negative Results},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.20605663},
url = {https://doi.org/10.5281/zenodo.20605663}
}