Tensoic/Tiny-Llama-openhermes-1.1B-step-715k-1.5T

Model

4

stars

6

commits

1

linked in READMEs

Nov 20, 2023

updated

endpoints_compatible
generated_from_trainer
llama
pytorch
safetensors
text-generation
text-generation-inference
transformers

README

This model is a fine-tuned version of PY007/TinyLlama-1.1B-intermediate-step-715k-1.5T on the openhermes dataset. It achieves the following results on the evaluation set:

  • Loss: 1.2355

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0002
  • train_batch_size: 1
  • eval_batch_size: 1
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 8
  • total_train_batch_size: 8
  • total_eval_batch_size: 8
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_steps: 100
  • num_epochs: 1

Training results

Training LossEpochStepValidation Loss
1.46540.013.5326
1.21620.0515031.9335
1.19180.130061.7391
1.41880.1545091.7574
1.82810.260121.6704
0.86390.2575151.7459
1.37640.390181.6832
2.11720.35105211.6398
1.18550.4120241.6007
1.56040.45135271.5256
1.02240.5150301.4891
1.55820.55165331.4903
0.94890.6180361.4179
1.670.65195391.4585
0.85420.7210421.3810
1.53010.75225451.3645
0.9510.8240481.3087
1.17910.85255511.3018
1.33420.9270541.2595
1.12210.95285571.2355

Framework versions

  • Transformers 4.34.0
  • Pytorch 2.0.1+cu117
  • Datasets 2.14.6
  • Tokenizers 0.14.1

Contributors

adarshxs

5 commits

SFconvertbot

1 commits

Tensoic/Tiny-Llama-openhermes-1.1B-step-715k-1.5T

Model

4

stars

6

commits

1

linked in READMEs

Nov 20, 2023

updated

endpoints_compatible
generated_from_trainer
llama
pytorch
safetensors
text-generation
text-generation-inference
transformers

README

This model is a fine-tuned version of PY007/TinyLlama-1.1B-intermediate-step-715k-1.5T on the openhermes dataset. It achieves the following results on the evaluation set:

  • Loss: 1.2355

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0002
  • train_batch_size: 1
  • eval_batch_size: 1
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 8
  • total_train_batch_size: 8
  • total_eval_batch_size: 8
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_steps: 100
  • num_epochs: 1

Training results

Training LossEpochStepValidation Loss
1.46540.013.5326
1.21620.0515031.9335
1.19180.130061.7391
1.41880.1545091.7574
1.82810.260121.6704
0.86390.2575151.7459
1.37640.390181.6832
2.11720.35105211.6398
1.18550.4120241.6007
1.56040.45135271.5256
1.02240.5150301.4891
1.55820.55165331.4903
0.94890.6180361.4179
1.670.65195391.4585
0.85420.7210421.3810
1.53010.75225451.3645
0.9510.8240481.3087
1.17910.85255511.3018
1.33420.9270541.2595
1.12210.95285571.2355

Framework versions

  • Transformers 4.34.0
  • Pytorch 2.0.1+cu117
  • Datasets 2.14.6
  • Tokenizers 0.14.1

Contributors

adarshxs

5 commits

SFconvertbot

1 commits