KazBERT for Named Entity Recognition (NER)
4
11 commits
1 linked in READMEs
updated Feb 20, 2026
This model is a fine-tuned version of KazBERT (a specialized BERT model for the Kazakh language) on the KazNERD dataset. It is designed to identify and categorize named entities in Kazakh text into 25 distinct classes.
While multilingual models like XLM-RoBERTa cover many languages, KazBERT-NER is specifically optimized for the nuances of the Kazakh language. Despite having significantly fewer parameters than "Large" multilingual models, it achieves competitive performance, demonstrating superior efficiency and domain-specific knowledge (especially in complex categories like Adages).
The model was trained with a focus on stability and fine-tuning the pre-existing semantic knowledge of KazBERT:
The model shows exceptional stability across both Validation and Test sets, proving its ability to generalize to unseen Kazakh text.
| Metric | Value |
|---|---|
| F1-Score | 95.22% |
| Precision | 95.16% |
| Recall | 95.28% |
The model excels in identifying core entities such as Persons, Dates, and Monetary values.
| Entity Class | Precision | Recall | F1-Score |
|---|---|---|---|
| PERSON | 98.46% | 98.25% | 98.35% |
| MONEY | 98.86% | 98.41% | 98.64% |
| GPE (Geopolitics) | 97.05% | 96.21% | 96.63% |
| CARDINAL | 97.44% | 98.30% | 97.87% |
| DATE | 96.79% | 96.90% | 96.84% |
| ADAGE | 50.00% | 36.84% | 42.42% |
While models like xlm-roberta-large-kaznerd (560M params) may show slightly higher overall F1, KazBERT-NER (~110M params) offers:
Efficiency: 5x fewer parameters, leading to much faster inference and lower deployment costs.
Native Understanding: Better performance on culture-specific entities like ADAGE compared to many multilingual alternatives.
Clean Embeddings: Contextual representations focused purely on Kazakh syntax and semantics.
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
model_name = "Eraly-ml/KazBERT-NERD"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)
nlp = pipeline("ner", model=model, tokenizer=tokenizer, aggregation_strategy="simple")
example = "Қазақстан Республикасы — Орталық Азияда орналасқан мемлекет."
print(nlp(example))
# [{'entity_group': 'GPE', 'score': 0.9998361, 'word': 'қазақстан республикасы', 'start': 0, 'end': 22},
# {'entity_group': 'LOCATION', 'score': 0.9993429, 'word': 'орталық азияда', 'start': 25, 'end': 39}]
KazBERT for Named Entity Recognition (NER)
4
11 commits
1 linked in READMEs
updated Feb 20, 2026
This model is a fine-tuned version of KazBERT (a specialized BERT model for the Kazakh language) on the KazNERD dataset. It is designed to identify and categorize named entities in Kazakh text into 25 distinct classes.
While multilingual models like XLM-RoBERTa cover many languages, KazBERT-NER is specifically optimized for the nuances of the Kazakh language. Despite having significantly fewer parameters than "Large" multilingual models, it achieves competitive performance, demonstrating superior efficiency and domain-specific knowledge (especially in complex categories like Adages).
The model was trained with a focus on stability and fine-tuning the pre-existing semantic knowledge of KazBERT:
The model shows exceptional stability across both Validation and Test sets, proving its ability to generalize to unseen Kazakh text.
| Metric | Value |
|---|---|
| F1-Score | 95.22% |
| Precision | 95.16% |
| Recall | 95.28% |
The model excels in identifying core entities such as Persons, Dates, and Monetary values.
| Entity Class | Precision | Recall | F1-Score |
|---|---|---|---|
| PERSON | 98.46% | 98.25% | 98.35% |
| MONEY | 98.86% | 98.41% | 98.64% |
| GPE (Geopolitics) | 97.05% | 96.21% | 96.63% |
| CARDINAL | 97.44% | 98.30% | 97.87% |
| DATE | 96.79% | 96.90% | 96.84% |
| ADAGE | 50.00% | 36.84% | 42.42% |
While models like xlm-roberta-large-kaznerd (560M params) may show slightly higher overall F1, KazBERT-NER (~110M params) offers:
Efficiency: 5x fewer parameters, leading to much faster inference and lower deployment costs.
Native Understanding: Better performance on culture-specific entities like ADAGE compared to many multilingual alternatives.
Clean Embeddings: Contextual representations focused purely on Kazakh syntax and semantics.
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
model_name = "Eraly-ml/KazBERT-NERD"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)
nlp = pipeline("ner", model=model, tokenizer=tokenizer, aggregation_strategy="simple")
example = "Қазақстан Республикасы — Орталық Азияда орналасқан мемлекет."
print(nlp(example))
# [{'entity_group': 'GPE', 'score': 0.9998361, 'word': 'қазақстан республикасы', 'start': 0, 'end': 22},
# {'entity_group': 'LOCATION', 'score': 0.9993429, 'word': 'орталық азияда', 'start': 25, 'end': 39}]