Tim2190/Kaz-RAG-search-benchmark

Evidence-based benchmark for Kazakh information retrieval: 300 queries × 3 categories, BM25/dense/hybrid, with honest negative results. The proof base for a Kazakh morphological stemmer.

Python

0

176 commits

updated Jun 24, 2026

See the code

README

Kazakh Stemmer — Effectiveness, Proven

DOI

This repository is the independent evidence base for my Kazakh Stemmer. A reproducible, statistically validated benchmark showing the stemmer measurably improves Kazakh search — on 300 queries over 8 370 Wikipedia passages.

What the stemmer does: give it a word in any grammatical form → it returns the root, plus the suffixes it stripped. балаларымызда → бала (суффиксы: да, ымыз, лар). Kazakh is agglutinative — one word appears in hundreds of forms, and ordinary search misses them. The stemmer lets search see through that.

5-system comparison

The proof: stemming improves search quality by +16% nDCG@10 on inflected queries (p=0.0017) and +9% overall (p=0.0001), on 300 queries with paired-bootstrap significance. It also outperforms zero-shot Google LaBSE embeddings (0.754 vs 0.481, n=300) — for Kazakh, morphological normalization matters more than naive multilingual vectors.

→ Try the Kazakh Stemmer · full methodology and numbers below


Headline Result

Corpus: 8 370 passages (Kazakh Wikipedia) · Queries: 300 (100 entities × 3 categories) Baseline: BM25 (Okapi). "Before" — no normalization; "After" — corpus and query tokens stemmed with the Kazakh stemmer.

Statistical Significance — nDCG@10 (paired bootstrap, 10 000 resamples, n=300)

categoryBeforeAfterΔgainp-value
inflected0.6270.727+0.101+16%p=0.0017 ✅
natural0.7030.772+0.068+10%p=0.0063 ✅
vocabulary-gap0.7410.764+0.023+3%p=0.21 ✗
ALL0.6900.754+0.064+9%p=0.0001 ✅

Vocabulary-gap is not significantly improved — expected: stemming fixes morphological mismatch, not semantic gaps between synonyms. This is the honest, theoretically correct result.

Full metrics with recall@{1,5,10}, MRR@10 — in results/RESULTS.md.


Why This Matters

Kazakh is agglutinative: 800 Wikipedia articles produce 102 408 unique surface forms. A single word like теңіз (sea) appears as теңіздерге, теңізде, теңіздің in text, but a search query may use a different form. Exact word matching fails completely on inflected queries (recall@1 = 0.10 without stemmer).

The stemmer reduces forms to their root (Бакуде → баку, теңіздерге → теңіз), letting BM25 see through morphological variation.

Dense Models: What They Get Right and Wrong

inflected (morphology)naturalvocabulary-gap (synonyms)
BM25+Stemmer✅ fixed✅✅
Dense Granite-278M✅✅ best inflected (0.791)✅✅ best natural (0.923)❌ collapses (0.303)
Dense E5-base✅ good (0.845)✅✅ best (0.947)✅ best (0.562)
Dense LaBSEweak (0.477)ok (0.546)weak (0.419)

BM25+Stemmer (0.754) outperforms zero-shot LaBSE (0.481) on n=300. LaBSE is a strong multilingual model — but it receives no Kazakh-specific fine-tuning here, and for a highly agglutinative language, morphological normalization turns out to matter more than raw multilingual embeddings. E5-base (0.785) does beat the stemmer overall at the cost of GPU inference and 15–30 min embedding time; Granite collapses on vocabulary-gap (0.303) despite leading on inflected queries.


RAG End-to-End: Does Better Retrieval Reach the Answer? (Qwen2.5-7B, n=300)

We ran the full RAG chain (BM25 → context → Qwen2.5-7B 4-bit) on the same 300 queries, changing only the retriever. The stemmer improves retrieval (hit@3 0.737 → 0.803) — but does that reach the final answer?

Queries and scoring. All 300 queries are short factual questions with a one-to-two word gold answer (e.g. "Қазақстанның астанасы қайда?" → "Астана"). Accuracy is measured by substring match: the answer is correct if the gold string appears in the LLM's response, abstain if the model wrote "Ақпарат жоқ" (no information), and hallucination otherwise. This approach requires no LLM judge and is fully reproducible, but it is conservative — a semantically correct answer phrased differently counts as hallucination.

RAG end-to-end: Qwen2.5-7B, n=300

StemmerRetrieval hit@3AccuracyHallucinationAbstain
No stemmer0.7370.4830.2170.300
Kazakh stemmer0.8030.5000.2400.260
Δ+0.066+0.017+0.023−0.040

✅ The stemmer's effectiveness is proven — for retrieval. It significantly improves search quality (nDCG@10 +9% overall, +16% on inflected, p≤0.0017, n=300) and raises the RAG retrieval hit-rate 0.737 → 0.803. That part is solid. The null result below is about the generator, not the stemmer: the bottleneck in Kazakh RAG today is the LLM's Kazakh comprehension, not the search step.

End-to-end accuracy gain is not statistically significant (McNemar exact p = 0.63; net +5 correct of 300). Better retrieval is necessary but not sufficient — the generator still has to extract the answer, and Qwen2.5-7B (4-bit) often fails to even with the right passage. The stemmer mostly makes Qwen more willing to answer (abstain 0.300 → 0.260), and those recovered answers split between correct and hallucinated, so net accuracy barely moves. A stronger Kazakh generator would likely convert the proven retrieval gain into real accuracy.

The directionally strongest effect is on inflected (morphology) queries — the biggest retrieval jump (+0.15 hit@3) and +0.07 accuracy — exactly where theory predicts, but even there p=0.25. A trend, not a proof.

⚠️ Replication note: an earlier version claimed a Qwen +0.100 accuracy gain that "scales with model competence," measured on n=60. That signal did not replicate at n=300 (+0.100 → +0.017, p=0.63) — it was sampling noise. The retrieval result survived the jump to 300 queries and got stronger; the end-to-end RAG claim did not. We keep the honest negative result. Full breakdown in results/RESULTS.md.


Methodology

  • Corpus. 800 random Kazakh Wikipedia articles from wikimedia/wikipedia (Hugging Face) → cleaning → chunking (~120 words) → language filter → 8 370 passages.
  • Queries. 100 entities (countries, cities, people, concepts) × 3 categories:
    • inflected — key word in an oblique grammatical case (morphology stress-test);
    • vocabulary-gap — synonyms/paraphrase (semantic stress-test); see caveat below;
    • natural — standard questions. Each query has a known ground-truth passage (qrels).

    ⚠️ Caveat on vocabulary-gap. Despite its name, this category was later found to have the highest query↔gold lexical overlap of the three (≈0.56), because the LLM generator reused key terms from the gold passage. Strong lexical (BM25/stemmer) scores here reflect that overlap, not semantic-gap closing. The genuine low-overlap test is the Akorda low_overlap category (≈0.32) — see results/akorda/AKORDA_RESULTS.md.

  • Systems. BM25 Okapi (k₁=1.5, b=0.75) ± Kazakh stemmer. Dense retrieval: three models evaluated zero-shot (no fine-tuning, no hard-negative training):
    • sentence-transformers/LaBSE — no query/passage prefixes (symmetric model);
    • intfloat/multilingual-e5-base — query: / passage: prefixes per model card;
    • ibm-granite/granite-embedding-278m-multilingual — no prefixes (symmetric). Similarity: cosine (L2-normalized embeddings → dot product). Brute-force exact search, no FAISS, no hybrid, no re-ranking. Same ~120-word chunks for all systems. Embeddings pre-computed and cached; BM25 re-indexing takes seconds.
  • Metrics. Recall@{1,5,10}, MRR@10, nDCG@10 + statistical significance (paired bootstrap, 10 000 resamples).
  • Why not lemmatization? The Kazakh stemmer performs full morphological analysis (dictionary lookup → base form + suffix list), which for an agglutinative language is functionally equivalent to lemmatization for retrieval purposes. No standalone Kazakh lemmatizer with a public API exists for direct comparison; this is an acknowledged limitation. The stemmer's base forms are the same forms that appear in corpus text, so the normalization is symmetric and consistent — which is what retrieval requires.
  • Reproducibility. Stem cache committed (102k tokens); BM25 results require no network. Dense results require GPU (~15–30 min in Colab).

Quickstart

git clone https://github.com/Tim2190/Kaz-RAG-search-benchmark.git
cd Kaz-RAG-search-benchmark
pip install -r requirements.txt

# BM25 before/after (no network — stem cache in repo)
python -m src.eval.run_benchmark --stemmer identity --out results/bm25_identity.json
python -m src.eval.run_benchmark --stemmer kazakh   --out results/bm25_kazakh.json

# Delta table + chart
python -m src.eval.compare --before results/bm25_identity.json \
    --after results/bm25_kazakh.json --chart results/before_after.png

# Statistical significance
python -m src.eval.significance

# 5-system comparison chart
python -m src.eval.chart_all --out results/systems_ndcg.png

Dense retrieval and RAG require GPU (Colab T4 is sufficient). See PIPELINE.md.

Tests

Core logic (metrics, tokenization, chunking, stem client, retrieval) covered by tests:

python -m unittest discover tests      # 79 tests, no network

Repository Structure

src/
  scraping/    # corpus collection (wiki dump via HF; news/gov fallback)
  corpus/      # cleaning, chunking, language filter
  queries/     # queries + qrels, dataset loader
  retrieval/   # bm25 (lexical) + dense (embeddings)
  preprocess/  # Kazakh stemmer (HTTP client + cache), tokenizer
  eval/        # metrics, benchmark runner, compare, significance, charts
  rag/         # LLM prompt, scorer, hallucination harness
data/
  corpus/      # corpus.jsonl — 8 370 passages
  queries/     # queries.jsonl — 300 queries with qrels
  resources/   # stem_cache.json — stemmer cache (102k words)
results/       # metrics JSON, charts, RESULTS.md
tests/         # 79 unit tests

What We Tried That Didn't Help: Synonym Query Expansion

We also tested a synonym expansion layer on top of the stemmer: for every query term, its stemmed synonyms (from a ~10k-word dictionary) are appended to the query before BM25 search (corpus untouched). The intuition was that synonyms might close the vocabulary-gap that morphology alone can't.

It made retrieval worse across the board — including on vocabulary-gap, where we expected it to help:

nDCG@10 (n=300)StemmerStemmer + SynonymsΔ
ALL0.7540.539−0.215
vocabulary-gap0.7640.627−0.137
inflected0.7270.537−0.190
natural0.7720.451−0.321

End-to-end RAG confirmed it: accuracy 0.500 → 0.393 (McNemar exact p = 0.0002 — significantly worse).

Why: the stemmer already gives a strong lexical signal (0.754 nDCG@10). Expansion bloats each query from ~6.5 to ~26.5 tokens, and most of those added synonyms are not in the relevant passage — they pull in spurious documents and bury the right one. When the lexical match is already good, unweighted synonym expansion is just noise. Synonyms are a tool for a weak lexical signal (short documents, no normalization, domain jargon); here the signal is strong, so the stemmer works better on its own.

Reproducible: python -m src.eval.run_synonyms and python -m src.eval.hit_at_k.

Diagnosing the failure — vocabulary-gap subanalysis. To distinguish between two hypotheses ("dictionary doesn't cover the right words" vs "expansion dilutes the signal"), we split the 100 vocabulary-gap queries by whether the synonym cache bridged the actual gap to the gold passage (python -m src.eval.vocab_gap_analysis):

subgroupnkazakh nDCG@10synonym nDCG@10Δ
uncovered (no synonyms found)21.0001.000≈0
covered_noise (synonyms found, none in gold passage)730.7510.570▼0.181
covered_bridge (synonyms found, ≥1 in gold passage)250.7830.766▼0.017

The mechanism is the problem, not the dictionary. The dictionary covers 98% of queries (only 2 uncovered). But in 73% of cases it returns synonyms for a different sense of the word — contextually wrong, pulling in spurious documents (▼0.181). Critically, even in the 25% where the correct synonym IS added (covered_bridge), retrieval still slightly drops (▼0.017) because the other query terms each add their own wrong-context synonyms. Unweighted expansion is the wrong tool here: what's needed is context-aware disambiguation, not a flat synonym lookup.


What's proven

The central claim — Kazakh morphology breaks lexical search, and a stemmer fixes it — is statistically proven and fully reproducible (nDCG@10 +9% overall, +16% on inflected, p≤0.0017, n=300). Five retrieval systems are benchmarked; the end-to-end RAG effect on Qwen2.5-7B is measured and honestly reported (retrieval improves, end-to-end accuracy gain not significant — the bottleneck is the generator). Full numbers in results/RESULTS.md.


Citation

A preprint of this benchmark is archived on Zenodo with a permanent DOI:

Seidalin, T. (2026). Morphology Beats Multilingual Embeddings for Kazakh Retrieval: A 300-Query Benchmark with Honest Negative Results. Zenodo. https://doi.org/10.5281/zenodo.20605663

@misc{seidalin2026kazakh,
  author       = {Seidalin, Timur},
  title        = {Morphology Beats Multilingual Embeddings for Kazakh Retrieval:
                  A 300-Query Benchmark with Honest Negative Results},
  year         = {2026},
  publisher    = {Zenodo},
  doi          = {10.5281/zenodo.20605663},
  url          = {https://doi.org/10.5281/zenodo.20605663}
}

Follow-up experiments

Work extending the original 300-query study (the conclusions above are unchanged):

  • New embedding models on the same n=300 benchmark — IBM Granite R2 (97M/311M), a Kazakh-fine-tuned E5, and lexical–dense RRF hybrids, plus a tokenization analysis: results/SPRINT2_NEW_MODELS.md.
  • IBM Granite R2 review (R2 vs R1 on Kazakh, with limitations) — English · Русский.
  • Validated semantic-gap benchmark — 127 native-speaker-validated low-overlap queries, passage- and article-level scoring with bootstrap CI: results/SPRINT3_SYNONYM.md.
  • Akorda OOD confirmatory study — the same 7 systems re-run on an entirely different corpus (official presidential speeches, akorda.kz, n=244 queries). Rankings largely hold (Spearman ρ=0.89); dense model ordering is OOD-stable, BM25 vs dense balance shifts with domain lexical overlap: results/akorda/AKORDA_RESULTS.md.
  • Input preprocessing without fine-tuning (attempt #1 to fix the input) — can stemming or Cyrillic→Latin transliteration repair byte-fallback tokenizers (Qwen3) without training? Self-contained study (4 lines × e5/qwen3 × Wiki/Akorda, paired bootstrap): no — every transform hurts; a direct fertility measurement shows the byte-fallback failure is representational, not script-level. English · Русский.
  • Preprint 2 — full write-up combining new models (Granite R2, kazakh-e5), hybrid RRF results, OOD validation, and tokenizer fertility analysis. DOI

Русская версия

benchmark
kazakh
nlp
rag
stemmer

Tim2190/Kaz-RAG-search-benchmark

Evidence-based benchmark for Kazakh information retrieval: 300 queries × 3 categories, BM25/dense/hybrid, with honest negative results. The proof base for a Kazakh morphological stemmer.

Python

0

176 commits

updated Jun 24, 2026

See the code

README

Kazakh Stemmer — Effectiveness, Proven

DOI

This repository is the independent evidence base for my Kazakh Stemmer. A reproducible, statistically validated benchmark showing the stemmer measurably improves Kazakh search — on 300 queries over 8 370 Wikipedia passages.

What the stemmer does: give it a word in any grammatical form → it returns the root, plus the suffixes it stripped. балаларымызда → бала (суффиксы: да, ымыз, лар). Kazakh is agglutinative — one word appears in hundreds of forms, and ordinary search misses them. The stemmer lets search see through that.

5-system comparison

The proof: stemming improves search quality by +16% nDCG@10 on inflected queries (p=0.0017) and +9% overall (p=0.0001), on 300 queries with paired-bootstrap significance. It also outperforms zero-shot Google LaBSE embeddings (0.754 vs 0.481, n=300) — for Kazakh, morphological normalization matters more than naive multilingual vectors.

→ Try the Kazakh Stemmer · full methodology and numbers below


Headline Result

Corpus: 8 370 passages (Kazakh Wikipedia) · Queries: 300 (100 entities × 3 categories) Baseline: BM25 (Okapi). "Before" — no normalization; "After" — corpus and query tokens stemmed with the Kazakh stemmer.

Statistical Significance — nDCG@10 (paired bootstrap, 10 000 resamples, n=300)

categoryBeforeAfterΔgainp-value
inflected0.6270.727+0.101+16%p=0.0017 ✅
natural0.7030.772+0.068+10%p=0.0063 ✅
vocabulary-gap0.7410.764+0.023+3%p=0.21 ✗
ALL0.6900.754+0.064+9%p=0.0001 ✅

Vocabulary-gap is not significantly improved — expected: stemming fixes morphological mismatch, not semantic gaps between synonyms. This is the honest, theoretically correct result.

Full metrics with recall@{1,5,10}, MRR@10 — in results/RESULTS.md.


Why This Matters

Kazakh is agglutinative: 800 Wikipedia articles produce 102 408 unique surface forms. A single word like теңіз (sea) appears as теңіздерге, теңізде, теңіздің in text, but a search query may use a different form. Exact word matching fails completely on inflected queries (recall@1 = 0.10 without stemmer).

The stemmer reduces forms to their root (Бакуде → баку, теңіздерге → теңіз), letting BM25 see through morphological variation.

Dense Models: What They Get Right and Wrong

inflected (morphology)naturalvocabulary-gap (synonyms)
BM25+Stemmer✅ fixed✅✅
Dense Granite-278M✅✅ best inflected (0.791)✅✅ best natural (0.923)❌ collapses (0.303)
Dense E5-base✅ good (0.845)✅✅ best (0.947)✅ best (0.562)
Dense LaBSEweak (0.477)ok (0.546)weak (0.419)

BM25+Stemmer (0.754) outperforms zero-shot LaBSE (0.481) on n=300. LaBSE is a strong multilingual model — but it receives no Kazakh-specific fine-tuning here, and for a highly agglutinative language, morphological normalization turns out to matter more than raw multilingual embeddings. E5-base (0.785) does beat the stemmer overall at the cost of GPU inference and 15–30 min embedding time; Granite collapses on vocabulary-gap (0.303) despite leading on inflected queries.


RAG End-to-End: Does Better Retrieval Reach the Answer? (Qwen2.5-7B, n=300)

We ran the full RAG chain (BM25 → context → Qwen2.5-7B 4-bit) on the same 300 queries, changing only the retriever. The stemmer improves retrieval (hit@3 0.737 → 0.803) — but does that reach the final answer?

Queries and scoring. All 300 queries are short factual questions with a one-to-two word gold answer (e.g. "Қазақстанның астанасы қайда?" → "Астана"). Accuracy is measured by substring match: the answer is correct if the gold string appears in the LLM's response, abstain if the model wrote "Ақпарат жоқ" (no information), and hallucination otherwise. This approach requires no LLM judge and is fully reproducible, but it is conservative — a semantically correct answer phrased differently counts as hallucination.

RAG end-to-end: Qwen2.5-7B, n=300

StemmerRetrieval hit@3AccuracyHallucinationAbstain
No stemmer0.7370.4830.2170.300
Kazakh stemmer0.8030.5000.2400.260
Δ+0.066+0.017+0.023−0.040

✅ The stemmer's effectiveness is proven — for retrieval. It significantly improves search quality (nDCG@10 +9% overall, +16% on inflected, p≤0.0017, n=300) and raises the RAG retrieval hit-rate 0.737 → 0.803. That part is solid. The null result below is about the generator, not the stemmer: the bottleneck in Kazakh RAG today is the LLM's Kazakh comprehension, not the search step.

End-to-end accuracy gain is not statistically significant (McNemar exact p = 0.63; net +5 correct of 300). Better retrieval is necessary but not sufficient — the generator still has to extract the answer, and Qwen2.5-7B (4-bit) often fails to even with the right passage. The stemmer mostly makes Qwen more willing to answer (abstain 0.300 → 0.260), and those recovered answers split between correct and hallucinated, so net accuracy barely moves. A stronger Kazakh generator would likely convert the proven retrieval gain into real accuracy.

The directionally strongest effect is on inflected (morphology) queries — the biggest retrieval jump (+0.15 hit@3) and +0.07 accuracy — exactly where theory predicts, but even there p=0.25. A trend, not a proof.

⚠️ Replication note: an earlier version claimed a Qwen +0.100 accuracy gain that "scales with model competence," measured on n=60. That signal did not replicate at n=300 (+0.100 → +0.017, p=0.63) — it was sampling noise. The retrieval result survived the jump to 300 queries and got stronger; the end-to-end RAG claim did not. We keep the honest negative result. Full breakdown in results/RESULTS.md.


Methodology

  • Corpus. 800 random Kazakh Wikipedia articles from wikimedia/wikipedia (Hugging Face) → cleaning → chunking (~120 words) → language filter → 8 370 passages.
  • Queries. 100 entities (countries, cities, people, concepts) × 3 categories:
    • inflected — key word in an oblique grammatical case (morphology stress-test);
    • vocabulary-gap — synonyms/paraphrase (semantic stress-test); see caveat below;
    • natural — standard questions. Each query has a known ground-truth passage (qrels).

    ⚠️ Caveat on vocabulary-gap. Despite its name, this category was later found to have the highest query↔gold lexical overlap of the three (≈0.56), because the LLM generator reused key terms from the gold passage. Strong lexical (BM25/stemmer) scores here reflect that overlap, not semantic-gap closing. The genuine low-overlap test is the Akorda low_overlap category (≈0.32) — see results/akorda/AKORDA_RESULTS.md.

  • Systems. BM25 Okapi (k₁=1.5, b=0.75) ± Kazakh stemmer. Dense retrieval: three models evaluated zero-shot (no fine-tuning, no hard-negative training):
    • sentence-transformers/LaBSE — no query/passage prefixes (symmetric model);
    • intfloat/multilingual-e5-base — query: / passage: prefixes per model card;
    • ibm-granite/granite-embedding-278m-multilingual — no prefixes (symmetric). Similarity: cosine (L2-normalized embeddings → dot product). Brute-force exact search, no FAISS, no hybrid, no re-ranking. Same ~120-word chunks for all systems. Embeddings pre-computed and cached; BM25 re-indexing takes seconds.
  • Metrics. Recall@{1,5,10}, MRR@10, nDCG@10 + statistical significance (paired bootstrap, 10 000 resamples).
  • Why not lemmatization? The Kazakh stemmer performs full morphological analysis (dictionary lookup → base form + suffix list), which for an agglutinative language is functionally equivalent to lemmatization for retrieval purposes. No standalone Kazakh lemmatizer with a public API exists for direct comparison; this is an acknowledged limitation. The stemmer's base forms are the same forms that appear in corpus text, so the normalization is symmetric and consistent — which is what retrieval requires.
  • Reproducibility. Stem cache committed (102k tokens); BM25 results require no network. Dense results require GPU (~15–30 min in Colab).

Quickstart

git clone https://github.com/Tim2190/Kaz-RAG-search-benchmark.git
cd Kaz-RAG-search-benchmark
pip install -r requirements.txt

# BM25 before/after (no network — stem cache in repo)
python -m src.eval.run_benchmark --stemmer identity --out results/bm25_identity.json
python -m src.eval.run_benchmark --stemmer kazakh   --out results/bm25_kazakh.json

# Delta table + chart
python -m src.eval.compare --before results/bm25_identity.json \
    --after results/bm25_kazakh.json --chart results/before_after.png

# Statistical significance
python -m src.eval.significance

# 5-system comparison chart
python -m src.eval.chart_all --out results/systems_ndcg.png

Dense retrieval and RAG require GPU (Colab T4 is sufficient). See PIPELINE.md.

Tests

Core logic (metrics, tokenization, chunking, stem client, retrieval) covered by tests:

python -m unittest discover tests      # 79 tests, no network

Repository Structure

src/
  scraping/    # corpus collection (wiki dump via HF; news/gov fallback)
  corpus/      # cleaning, chunking, language filter
  queries/     # queries + qrels, dataset loader
  retrieval/   # bm25 (lexical) + dense (embeddings)
  preprocess/  # Kazakh stemmer (HTTP client + cache), tokenizer
  eval/        # metrics, benchmark runner, compare, significance, charts
  rag/         # LLM prompt, scorer, hallucination harness
data/
  corpus/      # corpus.jsonl — 8 370 passages
  queries/     # queries.jsonl — 300 queries with qrels
  resources/   # stem_cache.json — stemmer cache (102k words)
results/       # metrics JSON, charts, RESULTS.md
tests/         # 79 unit tests

What We Tried That Didn't Help: Synonym Query Expansion

We also tested a synonym expansion layer on top of the stemmer: for every query term, its stemmed synonyms (from a ~10k-word dictionary) are appended to the query before BM25 search (corpus untouched). The intuition was that synonyms might close the vocabulary-gap that morphology alone can't.

It made retrieval worse across the board — including on vocabulary-gap, where we expected it to help:

nDCG@10 (n=300)StemmerStemmer + SynonymsΔ
ALL0.7540.539−0.215
vocabulary-gap0.7640.627−0.137
inflected0.7270.537−0.190
natural0.7720.451−0.321

End-to-end RAG confirmed it: accuracy 0.500 → 0.393 (McNemar exact p = 0.0002 — significantly worse).

Why: the stemmer already gives a strong lexical signal (0.754 nDCG@10). Expansion bloats each query from ~6.5 to ~26.5 tokens, and most of those added synonyms are not in the relevant passage — they pull in spurious documents and bury the right one. When the lexical match is already good, unweighted synonym expansion is just noise. Synonyms are a tool for a weak lexical signal (short documents, no normalization, domain jargon); here the signal is strong, so the stemmer works better on its own.

Reproducible: python -m src.eval.run_synonyms and python -m src.eval.hit_at_k.

Diagnosing the failure — vocabulary-gap subanalysis. To distinguish between two hypotheses ("dictionary doesn't cover the right words" vs "expansion dilutes the signal"), we split the 100 vocabulary-gap queries by whether the synonym cache bridged the actual gap to the gold passage (python -m src.eval.vocab_gap_analysis):

subgroupnkazakh nDCG@10synonym nDCG@10Δ
uncovered (no synonyms found)21.0001.000≈0
covered_noise (synonyms found, none in gold passage)730.7510.570▼0.181
covered_bridge (synonyms found, ≥1 in gold passage)250.7830.766▼0.017

The mechanism is the problem, not the dictionary. The dictionary covers 98% of queries (only 2 uncovered). But in 73% of cases it returns synonyms for a different sense of the word — contextually wrong, pulling in spurious documents (▼0.181). Critically, even in the 25% where the correct synonym IS added (covered_bridge), retrieval still slightly drops (▼0.017) because the other query terms each add their own wrong-context synonyms. Unweighted expansion is the wrong tool here: what's needed is context-aware disambiguation, not a flat synonym lookup.


What's proven

The central claim — Kazakh morphology breaks lexical search, and a stemmer fixes it — is statistically proven and fully reproducible (nDCG@10 +9% overall, +16% on inflected, p≤0.0017, n=300). Five retrieval systems are benchmarked; the end-to-end RAG effect on Qwen2.5-7B is measured and honestly reported (retrieval improves, end-to-end accuracy gain not significant — the bottleneck is the generator). Full numbers in results/RESULTS.md.


Citation

A preprint of this benchmark is archived on Zenodo with a permanent DOI:

Seidalin, T. (2026). Morphology Beats Multilingual Embeddings for Kazakh Retrieval: A 300-Query Benchmark with Honest Negative Results. Zenodo. https://doi.org/10.5281/zenodo.20605663

@misc{seidalin2026kazakh,
  author       = {Seidalin, Timur},
  title        = {Morphology Beats Multilingual Embeddings for Kazakh Retrieval:
                  A 300-Query Benchmark with Honest Negative Results},
  year         = {2026},
  publisher    = {Zenodo},
  doi          = {10.5281/zenodo.20605663},
  url          = {https://doi.org/10.5281/zenodo.20605663}
}

Follow-up experiments

Work extending the original 300-query study (the conclusions above are unchanged):

  • New embedding models on the same n=300 benchmark — IBM Granite R2 (97M/311M), a Kazakh-fine-tuned E5, and lexical–dense RRF hybrids, plus a tokenization analysis: results/SPRINT2_NEW_MODELS.md.
  • IBM Granite R2 review (R2 vs R1 on Kazakh, with limitations) — English · Русский.
  • Validated semantic-gap benchmark — 127 native-speaker-validated low-overlap queries, passage- and article-level scoring with bootstrap CI: results/SPRINT3_SYNONYM.md.
  • Akorda OOD confirmatory study — the same 7 systems re-run on an entirely different corpus (official presidential speeches, akorda.kz, n=244 queries). Rankings largely hold (Spearman ρ=0.89); dense model ordering is OOD-stable, BM25 vs dense balance shifts with domain lexical overlap: results/akorda/AKORDA_RESULTS.md.
  • Input preprocessing without fine-tuning (attempt #1 to fix the input) — can stemming or Cyrillic→Latin transliteration repair byte-fallback tokenizers (Qwen3) without training? Self-contained study (4 lines × e5/qwen3 × Wiki/Akorda, paired bootstrap): no — every transform hurts; a direct fertility measurement shows the byte-fallback failure is representational, not script-level. English · Русский.
  • Preprint 2 — full write-up combining new models (Granite R2, kazakh-e5), hybrid RRF results, OOD validation, and tokenizer fertility analysis. DOI

Русская версия

benchmark
kazakh
nlp
rag
stemmer

Languages

Python

99.3%