Based model but uses layernorm instead of QK.sum(-1) for the normalization, for better hardware efficiency.
2
stars
6
commits
1
linked in READMEs
Oct 18, 2024
updated
6 commits
ylwt/PaddlePaddle-gpt2
0
zhouzhen6246/LIRA
ylwt/PaddlePaddle-gpt-neo-125m
Tongjilibo/gpt2-ml_30g_corpus
jingqun/textharmony
Tongjilibo/gpt2-ml_15g_corpus
CognitiveKernel/ck-70b-v1
binxia/LLMGA-pretrained-mlp
HazyResearch/ThunderKittens
Tile primitives for speedy kernels
3,674