PolynomeAI/Llama-3.1-8B-kkru

Model

Llama-3.1-8B Fine-tuned for Russian-Kazakh Translation

0

11 commits

1 linked in READMEs

updated Feb 18, 2025

See the code

README

Llama-3.1-8B Fine-tuned for Russian-Kazakh Translation

This model is a fine-tuned version of Meta's Llama-3.1-8B, optimized for bidirectional translation between Russian and Kazakh languages. The model demonstrates strong performance in both translation directions, particularly excelling in Russian to Kazakh translation where it outperforms several baseline models.

Model Details

  • Base Model: unsloth/Meta-Llama-3.1-8B-bnb-4bit
  • Training Duration: 166 hours
  • Hardware: 1x NVIDIA A100 SXM GPU
  • Training Framework: Unsloth library for efficient fine-tuning
  • Model Size: 8B parameters

Training Data

Data Sources and Distribution

SourceSamplesPercentage
nu848,50955.38%
kazparc (kk-ru)290,78518.98%
kazparc (kk-en)290,87518.99%
kaznu78,9205.15%
News-Commentary9,0750.59%
TED20206,8870.45%
QED4,6640.30%
Tatoeba2,3010.15%
Total1,532,016100%

Evaluation Results – Russian to Kazakh (ru-kk)

MT Metrics (100 samples)

ModelTypeBLEUCOMET
PolynomeAIOpen-source13.310.85
issai/LLama-3.1-KazLLM-1.0-8BOpen-source4.510.75
meta/LLama-3.1-1.0-8BOpen-source4.660.65

LLM Judge Evaluation (GPT-4o mini)

Comparison with Yandex:

  • Yandex Better: 78.0%
  • PolynomeAI Better: 16.5%
  • Both Good: 2.5%
  • Both Bad: 3%

Evaluation Results – Kazakh to Russian (kk-ru)

MT Metrics (100 samples)

ModelTypeBLEUCOMET
PolynomeAIOpen-source28.720.91
issai/LLama-3.1-KazLLM-1.0-8BOpen-source28.060.91
meta/LLama-3.1-1.0-8BOpen-source16.640.87

We are not including deepvk_kazRush-ru-kk in this evaluation because it was specifically trained for Ru-Kk direction.

LLM Judge Evaluation (GPT-4o mini)

Comparison with Yandex:

  • Yandex Better: 52.0%
  • PolynomeAI Better: 41.0%
  • Both Good: 4.5%
  • Both Bad: 2.5%

Usage

To Do

Limitations and Bias

  • The model's performance has been primarily evaluated on a test set of 100 samples
  • Performance may vary depending on domain and complexity of the input text
  • The model inherits any biases present in the Llama-3.1-8B base model and training data
  • The training data is heavily skewed towards the 'nu' source (89.3%)
  • Most of the training data (97.1%) falls in the moderate quality range based on COMET scores (0.4 - 0.6)
endpoints_compatible
kazakh
pytorch
russian
safetensors
transformers
translation

PolynomeAI/Llama-3.1-8B-kkru

Model

Llama-3.1-8B Fine-tuned for Russian-Kazakh Translation

0

11 commits

1 linked in READMEs

updated Feb 18, 2025

See the code

README

Llama-3.1-8B Fine-tuned for Russian-Kazakh Translation

This model is a fine-tuned version of Meta's Llama-3.1-8B, optimized for bidirectional translation between Russian and Kazakh languages. The model demonstrates strong performance in both translation directions, particularly excelling in Russian to Kazakh translation where it outperforms several baseline models.

Model Details

  • Base Model: unsloth/Meta-Llama-3.1-8B-bnb-4bit
  • Training Duration: 166 hours
  • Hardware: 1x NVIDIA A100 SXM GPU
  • Training Framework: Unsloth library for efficient fine-tuning
  • Model Size: 8B parameters

Training Data

Data Sources and Distribution

SourceSamplesPercentage
nu848,50955.38%
kazparc (kk-ru)290,78518.98%
kazparc (kk-en)290,87518.99%
kaznu78,9205.15%
News-Commentary9,0750.59%
TED20206,8870.45%
QED4,6640.30%
Tatoeba2,3010.15%
Total1,532,016100%

Evaluation Results – Russian to Kazakh (ru-kk)

MT Metrics (100 samples)

ModelTypeBLEUCOMET
PolynomeAIOpen-source13.310.85
issai/LLama-3.1-KazLLM-1.0-8BOpen-source4.510.75
meta/LLama-3.1-1.0-8BOpen-source4.660.65

LLM Judge Evaluation (GPT-4o mini)

Comparison with Yandex:

  • Yandex Better: 78.0%
  • PolynomeAI Better: 16.5%
  • Both Good: 2.5%
  • Both Bad: 3%

Evaluation Results – Kazakh to Russian (kk-ru)

MT Metrics (100 samples)

ModelTypeBLEUCOMET
PolynomeAIOpen-source28.720.91
issai/LLama-3.1-KazLLM-1.0-8BOpen-source28.060.91
meta/LLama-3.1-1.0-8BOpen-source16.640.87

We are not including deepvk_kazRush-ru-kk in this evaluation because it was specifically trained for Ru-Kk direction.

LLM Judge Evaluation (GPT-4o mini)

Comparison with Yandex:

  • Yandex Better: 52.0%
  • PolynomeAI Better: 41.0%
  • Both Good: 4.5%
  • Both Bad: 2.5%

Usage

To Do

Limitations and Bias

  • The model's performance has been primarily evaluated on a test set of 100 samples
  • Performance may vary depending on domain and complexity of the input text
  • The model inherits any biases present in the Llama-3.1-8B base model and training data
  • The training data is heavily skewed towards the 'nu' source (89.3%)
  • Most of the training data (97.1%) falls in the moderate quality range based on COMET scores (0.4 - 0.6)
endpoints_compatible
kazakh
pytorch
russian
safetensors
transformers
translation