TurkuNLP/gpt3-finnish-large

Model

Generative Pretrained Transformer with 881M parameteres for Finnish.

9

7 commits

1 linked in READMEs

updated Jun 27, 2023

See the code

README

Generative Pretrained Transformer with 881M parameteres for Finnish.

TurkuNLP Finnish GPT-3-models are a model family of pretrained monolingual GPT-style language models that are based on BLOOM-architecture. Note that the models are pure language models, meaning that they are not instruction finetuned for dialogue or answering questions.

These models are intended to be used as foundational models that can be e.g. instruction finetuned to serve as modern chat-models.

All models are trained for 300B tokens.

Parameters

ModelLayersDimHeadsParams
Small1276812186M
Medium24102416437M
Large24153616881M
XL242064241.5B
”3B”322560322.8B
”8B”324096327.5B
"13B"4051204013.3B

Datasets

We used a combination of multiple Finnish resources.

Sampling ratios

DatasetCharsRatioWeightW.Ratio
Parsebank35.0B16.9%1.522.7%
mC4-Fi46.3B22.4%1.020.0%
CC-Fi79.6B38.5%1.034.4%
Fiwiki0.8B0.4%3.01.0%
Lönnrot0.8B0.4%3.01.0%
Yle1.6B0.8%2.01.4%
STT2.2B1.1%2.01.9%
ePub13.5B6.5%1.05.8%
Lehdet5.8B2.8%1.02.5%
Suomi2420.6B9.9%1.08.9%
Reddit-Fi0.7B0.4%1.00.3%
TOTAL207.0B100.0%N/A100.0%

More documentation and a paper coming soon.

bloom
endpoints_compatible
feature-extraction
pytorch
text-generation
text-generation-inference
transformers

Contributors

spyysalo

4 commits

rluukkon

3 commits

TurkuNLP/gpt3-finnish-large

Model

Generative Pretrained Transformer with 881M parameteres for Finnish.

9

7 commits

1 linked in READMEs

updated Jun 27, 2023

See the code

README

Generative Pretrained Transformer with 881M parameteres for Finnish.

TurkuNLP Finnish GPT-3-models are a model family of pretrained monolingual GPT-style language models that are based on BLOOM-architecture. Note that the models are pure language models, meaning that they are not instruction finetuned for dialogue or answering questions.

These models are intended to be used as foundational models that can be e.g. instruction finetuned to serve as modern chat-models.

All models are trained for 300B tokens.

Parameters

ModelLayersDimHeadsParams
Small1276812186M
Medium24102416437M
Large24153616881M
XL242064241.5B
”3B”322560322.8B
”8B”324096327.5B
"13B"4051204013.3B

Datasets

We used a combination of multiple Finnish resources.

Sampling ratios

DatasetCharsRatioWeightW.Ratio
Parsebank35.0B16.9%1.522.7%
mC4-Fi46.3B22.4%1.020.0%
CC-Fi79.6B38.5%1.034.4%
Fiwiki0.8B0.4%3.01.0%
Lönnrot0.8B0.4%3.01.0%
Yle1.6B0.8%2.01.4%
STT2.2B1.1%2.01.9%
ePub13.5B6.5%1.05.8%
Lehdet5.8B2.8%1.02.5%
Suomi2420.6B9.9%1.08.9%
Reddit-Fi0.7B0.4%1.00.3%
TOTAL207.0B100.0%N/A100.0%

More documentation and a paper coming soon.

bloom
endpoints_compatible
feature-extraction
pytorch
text-generation
text-generation-inference
transformers

Contributors

spyysalo

4 commits

rluukkon

3 commits