SlyEcho/open_llama_3b_ggml

Model

ggml versions of OpenLLaMa 3B

28

10 commits

3 linked in READMEs

updated Jun 22, 2023

See the code

README

ggml versions of OpenLLaMa 3B

Use with llama.cpp

Support is now merged to master branch.

Newer quantizations

There are now more quantization types in llama.cpp, some lower than 4 bits. Currently these are not supported, maybe because some weights have shapes that don't divide by 256.

Perplexity on wiki.test.raw

Qchunk600BT1000BT
F16[616]8.46567.7861
Q8_0[616]8.46677.7874
Q5_1[616]8.50727.8424
Q5_0[616]8.51567.8474
Q4_1[616]8.61028.0483
Q4_0[616]8.66748.0962

SlyEcho/open_llama_3b_ggml

Model

ggml versions of OpenLLaMa 3B

28

10 commits

3 linked in READMEs

updated Jun 22, 2023

See the code

README

ggml versions of OpenLLaMa 3B

Use with llama.cpp

Support is now merged to master branch.

Newer quantizations

There are now more quantization types in llama.cpp, some lower than 4 bits. Currently these are not supported, maybe because some weights have shapes that don't divide by 256.

Perplexity on wiki.test.raw

Qchunk600BT1000BT
F16[616]8.46567.7861
Q8_0[616]8.46677.7874
Q5_1[616]8.50727.8424
Q5_0[616]8.51567.8474
Q4_1[616]8.61028.0483
Q4_0[616]8.66748.0962