A simple script to test if GPT tokenizers from various popular models can encode+decode documents losslessly.
Documents tested are:
Does it matter if tokenizers can/can't reproduce the input exactly? I guess this is subjective, but I'd say it's at least a nice feature. A feature that some tokenizers out there don't seem to have.
NOTE: When decoding, it's important to pass clean_up_tokenization_spaces=False to tokenizer.decode. Otherwise you will get a lot of missmatches related to tokenizer eating whitespace.
| Tokenizer Name | Vocab Size | QQQ_0 | QQQ_1 | QQQ_2 | QQQ_3 | QQQ_4 | QQQ_5 | QQQ_6 | QQQ_7 |
|---|---|---|---|---|---|---|---|---|---|
| EleutherAI/polyglot-ko-5.8b | 30003 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| huggyllama/llama-7b | 32000 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| openlm-research/open_llama_7b_400bt_preview | 32000 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| bigcode/starcoder | 49152 | ✔️ | ✔️ | ✔️ | ❌ | ✔️ | ✔️ | ✔️ | ✔️ |
| facebook/galactica-6.7b | 50000 | ✔️ | ✔️ | ✔️ | ❌ | ❌ | ✔️ | ✔️ | ❌ |
| gpt2-large | 50257 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| EleutherAI/gpt-neox-20b | 50277 | ✔️ | ✔️ | ✔️ | ❌ | ✔️ | ✔️ | ✔️ | ✔️ |
| tiiuae/falcon-40b | 65024 | ✔️ | ✔️ | ✔️ | ❌ | ✔️ | ✔️ | ✔️ | ✔️ |
| bigscience/bloom-560m | 250680 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| Tokenizer Name | Vocab Size | QQQ_0 | QQQ_1 | QQQ_2 | QQQ_3 | QQQ_4 | QQQ_5 | QQQ_6 | QQQ_7 |
|---|---|---|---|---|---|---|---|---|---|
| EleutherAI/polyglot-ko-5.8b | 30003 | 36 | 436 | 428 | 18310 | 292044 | 390318 | 197049 | 144043 |
| huggyllama/llama-7b | 32000 | 12 | 231 | 223 | 5465 | 187793 | 196813 | 115284 | 88333 |
| openlm-research/open_llama_7b_400bt_preview | 32000 | 30 | 225 | 218 | 15452 | 189579 | 197971 | 148109 | 210463 |
| bigcode/starcoder | 49152 | 11 | 240 | 240 | 4678 | 179713 | 206842 | 104817 | 72613 |
| facebook/galactica-6.7b | 50000 | 11 | 248 | 236 | 5158 | 216860 | 196171 | 127599 | 145842 |
| gpt2-large | 50257 | 28 | 217 | 213 | 14781 | 175021 | 196107 | 139985 | 99792 |
| EleutherAI/gpt-neox-20b | 50277 | 11 | 216 | 210 | 4897 | 167877 | 183997 | 106364 | 76410 |
| tiiuae/falcon-40b | 65024 | 11 | 226 | 214 | 5486 | 171417 | 181700 | 112559 | 102247 |
| bigscience/bloom-560m | 250680 | 10 | 230 | 210 | 4082 | 127679 | 182074 | 92153 | 66894 |
Many models use the same tokenizers. Listing some of the relationships here for future reference to myself.
6 commits
Python
100.0%
A simple script to test if GPT tokenizers from various popular models can encode+decode documents losslessly.
Documents tested are:
Does it matter if tokenizers can/can't reproduce the input exactly? I guess this is subjective, but I'd say it's at least a nice feature. A feature that some tokenizers out there don't seem to have.
NOTE: When decoding, it's important to pass clean_up_tokenization_spaces=False to tokenizer.decode. Otherwise you will get a lot of missmatches related to tokenizer eating whitespace.
| Tokenizer Name | Vocab Size | QQQ_0 | QQQ_1 | QQQ_2 | QQQ_3 | QQQ_4 | QQQ_5 | QQQ_6 | QQQ_7 |
|---|---|---|---|---|---|---|---|---|---|
| EleutherAI/polyglot-ko-5.8b | 30003 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| huggyllama/llama-7b | 32000 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| openlm-research/open_llama_7b_400bt_preview | 32000 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| bigcode/starcoder | 49152 | ✔️ | ✔️ | ✔️ | ❌ | ✔️ | ✔️ | ✔️ | ✔️ |
| facebook/galactica-6.7b | 50000 | ✔️ | ✔️ | ✔️ | ❌ | ❌ | ✔️ | ✔️ | ❌ |
| gpt2-large | 50257 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| EleutherAI/gpt-neox-20b | 50277 | ✔️ | ✔️ | ✔️ | ❌ | ✔️ | ✔️ | ✔️ | ✔️ |
| tiiuae/falcon-40b | 65024 | ✔️ | ✔️ | ✔️ | ❌ | ✔️ | ✔️ | ✔️ | ✔️ |
| bigscience/bloom-560m | 250680 | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ | ✔️ |
| Tokenizer Name | Vocab Size | QQQ_0 | QQQ_1 | QQQ_2 | QQQ_3 | QQQ_4 | QQQ_5 | QQQ_6 | QQQ_7 |
|---|---|---|---|---|---|---|---|---|---|
| EleutherAI/polyglot-ko-5.8b | 30003 | 36 | 436 | 428 | 18310 | 292044 | 390318 | 197049 | 144043 |
| huggyllama/llama-7b | 32000 | 12 | 231 | 223 | 5465 | 187793 | 196813 | 115284 | 88333 |
| openlm-research/open_llama_7b_400bt_preview | 32000 | 30 | 225 | 218 | 15452 | 189579 | 197971 | 148109 | 210463 |
| bigcode/starcoder | 49152 | 11 | 240 | 240 | 4678 | 179713 | 206842 | 104817 | 72613 |
| facebook/galactica-6.7b | 50000 | 11 | 248 | 236 | 5158 | 216860 | 196171 | 127599 | 145842 |
| gpt2-large | 50257 | 28 | 217 | 213 | 14781 | 175021 | 196107 | 139985 | 99792 |
| EleutherAI/gpt-neox-20b | 50277 | 11 | 216 | 210 | 4897 | 167877 | 183997 | 106364 | 76410 |
| tiiuae/falcon-40b | 65024 | 11 | 226 | 214 | 5486 | 171417 | 181700 | 112559 | 102247 |
| bigscience/bloom-560m | 250680 | 10 | 230 | 210 | 4082 | 127679 | 182074 | 92153 | 66894 |
Many models use the same tokenizers. Listing some of the relationships here for future reference to myself.
6 commits
Python
100.0%