hfl/chinese-alpaca-2-1.3b-rlhf-gguf

Model

6

stars

25

commits

1

linked in READMEs

Jan 24, 2024

updated

endpoints_compatible
gguf
Browse cluster: Chinese LLM Model Distribution

README

Chinese-Alpaca-2-1.3B-RLHF-GGUF

This repository contains GGUF-v3 version (llama.cpp compatible) of Chinese-Alpaca-2-1.3B-RLHF, which is tuned on Chinese-Alpaca-2-1.3B with RLHF using DeepSpeed-Chat.

The optimal context length is 1K for this model. Specify -c 1024 when using with llama.cpp

Performance

Metric: PPL, lower is better

Quantoriginalimatrix (-im)
Q2_K20.1066 +/- 0.2923618.8209 +/- 0.27561
Q3_K16.9214 +/- 0.2613316.5729 +/- 0.25706
Q4_015.8056 +/- 0.23749-
Q4_K16.1579 +/- 0.2506415.7746 +/- 0.24476
Q5_015.4528 +/- 0.23911-
Q5_K15.3198 +/- 0.2362715.4791 +/- 0.23959
Q6_K15.3718 +/- 0.2376415.2572 +/- 0.23549
Q8_015.3302 +/- 0.23727-
F1615.3291 +/- 0.23728-

The model with -im suffix is generated with important matrix, which has generally better performance (not always though).

Others

For full model in HuggingFace format, please see: https://huggingface.co/hfl/chinese-alpaca-2-1.3b-rlhf

Please refer to https://github.com/ymcui/Chinese-LLaMA-Alpaca-2/ for more details.

Contributors

hfl-rc

25 commits

hfl/chinese-alpaca-2-1.3b-rlhf-gguf

Model

6

stars

25

commits

1

linked in READMEs

Jan 24, 2024

updated

endpoints_compatible
gguf
Browse cluster: Chinese LLM Model Distribution

README

Chinese-Alpaca-2-1.3B-RLHF-GGUF

This repository contains GGUF-v3 version (llama.cpp compatible) of Chinese-Alpaca-2-1.3B-RLHF, which is tuned on Chinese-Alpaca-2-1.3B with RLHF using DeepSpeed-Chat.

The optimal context length is 1K for this model. Specify -c 1024 when using with llama.cpp

Performance

Metric: PPL, lower is better

Quantoriginalimatrix (-im)
Q2_K20.1066 +/- 0.2923618.8209 +/- 0.27561
Q3_K16.9214 +/- 0.2613316.5729 +/- 0.25706
Q4_015.8056 +/- 0.23749-
Q4_K16.1579 +/- 0.2506415.7746 +/- 0.24476
Q5_015.4528 +/- 0.23911-
Q5_K15.3198 +/- 0.2362715.4791 +/- 0.23959
Q6_K15.3718 +/- 0.2376415.2572 +/- 0.23549
Q8_015.3302 +/- 0.23727-
F1615.3291 +/- 0.23728-

The model with -im suffix is generated with important matrix, which has generally better performance (not always though).

Others

For full model in HuggingFace format, please see: https://huggingface.co/hfl/chinese-alpaca-2-1.3b-rlhf

Please refer to https://github.com/ymcui/Chinese-LLaMA-Alpaca-2/ for more details.

Contributors

hfl-rc

25 commits