yeshpanovrustem/xlm-roberta-large-kaznerd

Model

A Named Entity Recognition Model for Kazakh

11

60 commits

2 linked in READMEs

updated Aug 7, 2025

See the code

README

A Named Entity Recognition Model for Kazakh

How to use

You can use this model with the Transformers pipeline for NER.

from transformers import AutoTokenizer, AutoModelForTokenClassification
from transformers import pipeline

tokenizer = AutoTokenizer.from_pretrained("yeshpanovrustem/xlm-roberta-large-kaznerd")
model = AutoModelForTokenClassification.from_pretrained("yeshpanovrustem/xlm-roberta-large-kaznerd")

# aggregation_strategy = "none"
nlp = pipeline("ner", model = model, tokenizer = tokenizer, aggregation_strategy = "none")
example = "Қазақстан Республикасы — Шығыс Еуропа мен Орталық Азияда орналасқан мемлекет."

ner_results = nlp(example)
for result in ner_results:
    print(result)

# output:
# {'entity': 'B-GPE', 'score': 0.9995646, 'index': 1, 'word': '▁Қазақстан', 'start': 0, 'end': 9}
# {'entity': 'I-GPE', 'score': 0.9994935, 'index': 2, 'word': '▁Республикасы', 'start': 10, 'end': 22}
# {'entity': 'B-LOCATION', 'score': 0.99906737, 'index': 4, 'word': '▁Шығыс', 'start': 25, 'end': 30}
# {'entity': 'I-LOCATION', 'score': 0.999153, 'index': 5, 'word': '▁Еуропа', 'start': 31, 'end': 37}
# {'entity': 'B-LOCATION', 'score': 0.9991597, 'index': 7, 'word': '▁Орталық', 'start': 42, 'end': 49}
# {'entity': 'I-LOCATION', 'score': 0.9991725, 'index': 8, 'word': '▁Азия', 'start': 50, 'end': 54}
# {'entity': 'I-LOCATION', 'score': 0.9992299, 'index': 9, 'word': 'да', 'start': 54, 'end': 56}

token = ""
label_list = []
token_list = []

for result in ner_results:
    if result["word"].startswith("▁"):
        if token:
            token_list.append(token.replace("▁", ""))
        token = result["word"]
        label_list.append(result["entity"])
    else:
        token += result["word"]

token_list.append(token.replace("▁", ""))

for token, label in zip(token_list, label_list):
    print(f"{token}\t{label}")

# output:
# Қазақстан	B-GPE
# Республикасы	I-GPE
# Шығыс	B-LOCATION
# Еуропа	I-LOCATION
# Орталық	B-LOCATION
# Азияда	I-LOCATION

# aggregation_strategy = "simple"
nlp = pipeline("ner", model = model, tokenizer = tokenizer, aggregation_strategy = "simple")
example = "Қазақстан Республикасы — Шығыс Еуропа мен Орталық Азияда орналасқан мемлекет."

ner_results = nlp(example)
for result in ner_results:
    print(result)

# output:
# {'entity_group': 'GPE', 'score': 0.999529, 'word': 'Қазақстан Республикасы', 'start': 0, 'end': 22}
# {'entity_group': 'LOCATION', 'score': 0.9991102, 'word': 'Шығыс Еуропа', 'start': 25, 'end': 37}
# {'entity_group': 'LOCATION', 'score': 0.9991874, 'word': 'Орталық Азияда', 'start': 42, 'end': 56}

Evaluation results on the validation and test sets

Validation setTest set
PrecisionRecallF1-scorePrecisionRecallF1-score
96.58%96.66%96.62%96.49%96.86%96.67%

Model performance for the NE classes of the validation set

NE ClassPrecisionRecallF1-scoreSupport
ADAGE90.00%47.37%62.07%19
ART91.36%95.48%93.38%155
CARDINAL98.44%98.37%98.40%2,878
CONTACT100.00%83.33%90.91%18
DATE97.38%97.27%97.33%2,603
DISEASE96.72%97.52%97.12%121
EVENT83.24%93.51%88.07%154
FACILITY68.95%84.83%76.07%178
GPE98.46%96.50%97.47%1,656
LANGUAGE95.45%89.36%92.31%47
LAW87.50%87.50%87.50%56
LOCATION92.49%93.81%93.14%210
MISCELLANEOUS100.00%76.92%86.96%26
MONEY99.56%100.00%99.78%455
NON_HUMAN0.00%0.00%0.00%1
NORP95.71%95.45%95.58%374
ORDINAL98.14%95.84%96.98%385
ORGANISATION92.19%90.97%91.58%753
PERCENTAGE99.08%99.08%99.08%437
PERSON98.47%98.72%98.60%1,175
POSITION96.15%97.79%96.96%587
PRODUCT89.06%78.08%83.21%73
PROJECT92.13%95.22%93.65%209
QUANTITY97.58%98.30%97.94%411
TIME94.81%96.63%95.71%208
micro avg96.58%96.66%96.62%13,189
macro avg90.12%87.51%88.39%13,189
weighted avg96.67%96.66%96.63%13,189

Model performance for the NE classes of the test set

NE ClassPrecisionRecallF1-scoreSupport
ADAGE71.43%29.41%41.67%17
ART95.71%96.89%96.30%161
CARDINAL98.43%98.60%98.51%2,789
CONTACT94.44%85.00%89.47%20
DATE96.59%97.60%97.09%2,584
DISEASE87.69%95.80%91.57%119
EVENT86.67%92.86%89.66%154
FACILITY74.88%81.73%78.16%197
GPE98.57%97.81%98.19%1,691
LANGUAGE90.70%95.12%92.86%41
LAW93.33%76.36%84.00%55
LOCATION92.08%89.42%90.73%208
MISCELLANEOUS86.21%96.15%90.91%26
MONEY100.00%100.00%100.00%427
NON_HUMAN0.00%0.00%0.00%1
NORP99.46%99.18%99.32%368
ORDINAL96.63%97.64%97.14%382
ORGANISATION90.97%91.23%91.10%718
PERCENTAGE98.05%98.05%98.05%462
PERSON98.70%99.13%98.92%1,151
POSITION96.36%97.65%97.00%597
PRODUCT89.23%77.33%82.86%75
PROJECT93.69%93.69%93.69%206
QUANTITY97.26%97.02%97.14%403
TIME94.95%94.09%94.52%220
micro avg96.54%96.85%96.69%13,072
macro avg88.88%87.11%87.55%13,072
weighted avg96.55%96.85%96.67%13,072
endpoints_compatible
Named Entity Recognition
pytorch
safetensors
token-classification
transformers
xlm-roberta

yeshpanovrustem/xlm-roberta-large-kaznerd

Model

A Named Entity Recognition Model for Kazakh

11

60 commits

2 linked in READMEs

updated Aug 7, 2025

See the code

README

A Named Entity Recognition Model for Kazakh

How to use

You can use this model with the Transformers pipeline for NER.

from transformers import AutoTokenizer, AutoModelForTokenClassification
from transformers import pipeline

tokenizer = AutoTokenizer.from_pretrained("yeshpanovrustem/xlm-roberta-large-kaznerd")
model = AutoModelForTokenClassification.from_pretrained("yeshpanovrustem/xlm-roberta-large-kaznerd")

# aggregation_strategy = "none"
nlp = pipeline("ner", model = model, tokenizer = tokenizer, aggregation_strategy = "none")
example = "Қазақстан Республикасы — Шығыс Еуропа мен Орталық Азияда орналасқан мемлекет."

ner_results = nlp(example)
for result in ner_results:
    print(result)

# output:
# {'entity': 'B-GPE', 'score': 0.9995646, 'index': 1, 'word': '▁Қазақстан', 'start': 0, 'end': 9}
# {'entity': 'I-GPE', 'score': 0.9994935, 'index': 2, 'word': '▁Республикасы', 'start': 10, 'end': 22}
# {'entity': 'B-LOCATION', 'score': 0.99906737, 'index': 4, 'word': '▁Шығыс', 'start': 25, 'end': 30}
# {'entity': 'I-LOCATION', 'score': 0.999153, 'index': 5, 'word': '▁Еуропа', 'start': 31, 'end': 37}
# {'entity': 'B-LOCATION', 'score': 0.9991597, 'index': 7, 'word': '▁Орталық', 'start': 42, 'end': 49}
# {'entity': 'I-LOCATION', 'score': 0.9991725, 'index': 8, 'word': '▁Азия', 'start': 50, 'end': 54}
# {'entity': 'I-LOCATION', 'score': 0.9992299, 'index': 9, 'word': 'да', 'start': 54, 'end': 56}

token = ""
label_list = []
token_list = []

for result in ner_results:
    if result["word"].startswith("▁"):
        if token:
            token_list.append(token.replace("▁", ""))
        token = result["word"]
        label_list.append(result["entity"])
    else:
        token += result["word"]

token_list.append(token.replace("▁", ""))

for token, label in zip(token_list, label_list):
    print(f"{token}\t{label}")

# output:
# Қазақстан	B-GPE
# Республикасы	I-GPE
# Шығыс	B-LOCATION
# Еуропа	I-LOCATION
# Орталық	B-LOCATION
# Азияда	I-LOCATION

# aggregation_strategy = "simple"
nlp = pipeline("ner", model = model, tokenizer = tokenizer, aggregation_strategy = "simple")
example = "Қазақстан Республикасы — Шығыс Еуропа мен Орталық Азияда орналасқан мемлекет."

ner_results = nlp(example)
for result in ner_results:
    print(result)

# output:
# {'entity_group': 'GPE', 'score': 0.999529, 'word': 'Қазақстан Республикасы', 'start': 0, 'end': 22}
# {'entity_group': 'LOCATION', 'score': 0.9991102, 'word': 'Шығыс Еуропа', 'start': 25, 'end': 37}
# {'entity_group': 'LOCATION', 'score': 0.9991874, 'word': 'Орталық Азияда', 'start': 42, 'end': 56}

Evaluation results on the validation and test sets

Validation setTest set
PrecisionRecallF1-scorePrecisionRecallF1-score
96.58%96.66%96.62%96.49%96.86%96.67%

Model performance for the NE classes of the validation set

NE ClassPrecisionRecallF1-scoreSupport
ADAGE90.00%47.37%62.07%19
ART91.36%95.48%93.38%155
CARDINAL98.44%98.37%98.40%2,878
CONTACT100.00%83.33%90.91%18
DATE97.38%97.27%97.33%2,603
DISEASE96.72%97.52%97.12%121
EVENT83.24%93.51%88.07%154
FACILITY68.95%84.83%76.07%178
GPE98.46%96.50%97.47%1,656
LANGUAGE95.45%89.36%92.31%47
LAW87.50%87.50%87.50%56
LOCATION92.49%93.81%93.14%210
MISCELLANEOUS100.00%76.92%86.96%26
MONEY99.56%100.00%99.78%455
NON_HUMAN0.00%0.00%0.00%1
NORP95.71%95.45%95.58%374
ORDINAL98.14%95.84%96.98%385
ORGANISATION92.19%90.97%91.58%753
PERCENTAGE99.08%99.08%99.08%437
PERSON98.47%98.72%98.60%1,175
POSITION96.15%97.79%96.96%587
PRODUCT89.06%78.08%83.21%73
PROJECT92.13%95.22%93.65%209
QUANTITY97.58%98.30%97.94%411
TIME94.81%96.63%95.71%208
micro avg96.58%96.66%96.62%13,189
macro avg90.12%87.51%88.39%13,189
weighted avg96.67%96.66%96.63%13,189

Model performance for the NE classes of the test set

NE ClassPrecisionRecallF1-scoreSupport
ADAGE71.43%29.41%41.67%17
ART95.71%96.89%96.30%161
CARDINAL98.43%98.60%98.51%2,789
CONTACT94.44%85.00%89.47%20
DATE96.59%97.60%97.09%2,584
DISEASE87.69%95.80%91.57%119
EVENT86.67%92.86%89.66%154
FACILITY74.88%81.73%78.16%197
GPE98.57%97.81%98.19%1,691
LANGUAGE90.70%95.12%92.86%41
LAW93.33%76.36%84.00%55
LOCATION92.08%89.42%90.73%208
MISCELLANEOUS86.21%96.15%90.91%26
MONEY100.00%100.00%100.00%427
NON_HUMAN0.00%0.00%0.00%1
NORP99.46%99.18%99.32%368
ORDINAL96.63%97.64%97.14%382
ORGANISATION90.97%91.23%91.10%718
PERCENTAGE98.05%98.05%98.05%462
PERSON98.70%99.13%98.92%1,151
POSITION96.36%97.65%97.00%597
PRODUCT89.23%77.33%82.86%75
PROJECT93.69%93.69%93.69%206
QUANTITY97.26%97.02%97.14%403
TIME94.95%94.09%94.52%220
micro avg96.54%96.85%96.69%13,072
macro avg88.88%87.11%87.55%13,072
weighted avg96.55%96.85%96.67%13,072
endpoints_compatible
Named Entity Recognition
pytorch
safetensors
token-classification
transformers
xlm-roberta