HiTZ/alpaca-lora-13b-en-pt-es-ca-eu-gl-at

Model

should probably proofread and complete it, then remove this comment. -->

0

32 commits

1 linked in READMEs

updated Mar 25, 2023

See the code

README

alpaca-lora-13b-en-pt-es-ca-eu-gl-at

This model is a fine-tuned version of decapoda-research/llama-13b-hf on the HiTZ/alpaca_mt ['en', 'pt', 'es', 'ca', 'eu', 'gl', 'at'] dataset. It achieves the following results on the evaluation set:

  • Loss: 0.9967

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0003
  • train_batch_size: 16
  • eval_batch_size: 16
  • seed: 42
  • gradient_accumulation_steps: 8
  • total_train_batch_size: 128
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_ratio: 0.03
  • num_epochs: 1
  • mixed_precision_training: Native AMP

Training results

Training LossEpochStepValidation Loss
1.3030.041001.2875
1.21530.072001.2016
1.15840.113001.1560
1.14260.154001.1277
1.11980.185001.1063
1.06310.226001.0911
1.07140.267001.0773
1.05050.298001.0667
1.04750.339001.0562
1.04110.3710001.0485
1.04180.411001.0413
1.04190.4412001.0339
1.03150.4813001.0290
1.02350.5114001.0238
1.03080.5515001.0189
1.00390.5916001.0157
1.00480.6217001.0110
0.99820.6618001.0080
1.01960.719001.0049
1.0190.7320001.0030
1.00370.7721001.0009
1.00030.8122000.9995
0.99420.8423000.9982
0.99860.8824000.9974
0.99870.9225000.9969
0.97630.9526000.9967
0.97330.9927000.9967

Framework versions

  • Transformers 4.28.0.dev0
  • Pytorch 2.0.0+cu117
  • Datasets 2.10.1
  • Tokenizers 0.13.2
generated_from_trainer

Contributors

juletxara

32 commits

HiTZ/alpaca-lora-13b-en-pt-es-ca-eu-gl-at

Model

should probably proofread and complete it, then remove this comment. -->

0

32 commits

1 linked in READMEs

updated Mar 25, 2023

See the code

README

alpaca-lora-13b-en-pt-es-ca-eu-gl-at

This model is a fine-tuned version of decapoda-research/llama-13b-hf on the HiTZ/alpaca_mt ['en', 'pt', 'es', 'ca', 'eu', 'gl', 'at'] dataset. It achieves the following results on the evaluation set:

  • Loss: 0.9967

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0003
  • train_batch_size: 16
  • eval_batch_size: 16
  • seed: 42
  • gradient_accumulation_steps: 8
  • total_train_batch_size: 128
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_ratio: 0.03
  • num_epochs: 1
  • mixed_precision_training: Native AMP

Training results

Training LossEpochStepValidation Loss
1.3030.041001.2875
1.21530.072001.2016
1.15840.113001.1560
1.14260.154001.1277
1.11980.185001.1063
1.06310.226001.0911
1.07140.267001.0773
1.05050.298001.0667
1.04750.339001.0562
1.04110.3710001.0485
1.04180.411001.0413
1.04190.4412001.0339
1.03150.4813001.0290
1.02350.5114001.0238
1.03080.5515001.0189
1.00390.5916001.0157
1.00480.6217001.0110
0.99820.6618001.0080
1.01960.719001.0049
1.0190.7320001.0030
1.00370.7721001.0009
1.00030.8122000.9995
0.99420.8423000.9982
0.99860.8824000.9974
0.99870.9225000.9969
0.97630.9526000.9967
0.97330.9927000.9967

Framework versions

  • Transformers 4.28.0.dev0
  • Pytorch 2.0.0+cu117
  • Datasets 2.10.1
  • Tokenizers 0.13.2
generated_from_trainer

Contributors

juletxara

32 commits