This model is a fine-tuned version of kazars24/trocr-base-handwritten-ru specifically optimized for recognizing Kazakh printed text. It leverages the TrOCR (Transformer-based Optical Character Recognition) architecture, utilizing a Vision Transformer (ViT) encoder and a RoBERTa-based decoder.
The model was adapted to handle Kazakh-specific Cyrillic characters (ә, ғ, қ, ң, ө, ұ, ү, һ, і) by resizing token embeddings and training on synthetic data.
The model was trained on thekamilya/kazakh-printed-dataset, which was synthetically generated using text from the ISSAI KazPARC corpus.
To overcome the scarcity of labeled Kazakh OCR data, I developed a robust synthetic generation engine:
import torch
from PIL import Image
from transformers import TrOCRProcessor, VisionEncoderDecoderModel
processor = TrOCRProcessor.from_pretrained("thekamilya/kazakh-trocr-fine-tuned")
model = VisionEncoderDecoderModel.from_pretrained("thekamilya/kazakh-trocr-fine-tuned")
# Move model to GPU
device = "cuda" if torch.cuda.is_available() else "cpu"
model = model.to(device)
# Load image
image = Image.open("zheke.jpg").convert("RGB")
# Prepare input and move to GPU
pixel_values = processor(
images=image,
return_tensors="pt"
).pixel_values.to(device)
# Inference on GPU
with torch.no_grad():
generated_ids = model.generate(pixel_values)
# Decode text
generated_text = processor.batch_decode(
generated_ids,
skip_special_tokens=True
)[0]
print(f"Recognized Text: {generated_text}")
This model is a fine-tuned version of kazars24/trocr-base-handwritten-ru specifically optimized for recognizing Kazakh printed text. It leverages the TrOCR (Transformer-based Optical Character Recognition) architecture, utilizing a Vision Transformer (ViT) encoder and a RoBERTa-based decoder.
The model was adapted to handle Kazakh-specific Cyrillic characters (ә, ғ, қ, ң, ө, ұ, ү, һ, і) by resizing token embeddings and training on synthetic data.
The model was trained on thekamilya/kazakh-printed-dataset, which was synthetically generated using text from the ISSAI KazPARC corpus.
To overcome the scarcity of labeled Kazakh OCR data, I developed a robust synthetic generation engine:
import torch
from PIL import Image
from transformers import TrOCRProcessor, VisionEncoderDecoderModel
processor = TrOCRProcessor.from_pretrained("thekamilya/kazakh-trocr-fine-tuned")
model = VisionEncoderDecoderModel.from_pretrained("thekamilya/kazakh-trocr-fine-tuned")
# Move model to GPU
device = "cuda" if torch.cuda.is_available() else "cpu"
model = model.to(device)
# Load image
image = Image.open("zheke.jpg").convert("RGB")
# Prepare input and move to GPU
pixel_values = processor(
images=image,
return_tensors="pt"
).pixel_values.to(device)
# Inference on GPU
with torch.no_grad():
generated_ids = model.generate(pixel_values)
# Decode text
generated_text = processor.batch_decode(
generated_ids,
skip_special_tokens=True
)[0]
print(f"Recognized Text: {generated_text}")