An unofficial implementation of the Infini-gram model proposed by Liu et al. (2024)
Go
33
73 commits
updated Jun 19, 2024
This repo contains two (unofficial) implementations of the infini-gram model described in Liu et al. (2024). This branch contains the Golang implementation. The main branch contains a Python implementation.
The tokenizers used here are the Go bindings to the official Rust library.
First, build the rust tokenizers binary:
cd tokenizers
make
Then, you can build the infinigram binary:
cd ../
go build -ldflags "-s"
./infinigram --train_file corpus.txt --out_dir output --tokenizer_config tokenizer.json
where corpus.txt contains one document per line. tokenizer.json corresponds to the HuggingFace pretrained Tokenizers file (e.g., for gpt2).
This implementation features:
--interactive_mode {0,1})mmap to access both the tokenized documents and the suffix array; memory usage during inference should be minimal.--max_mem): you should hypothetically be able to train (and infer) on any sized corpus regardless of how much memory you have--min_matches). e.g., you may set this at a value >= 2 to avoid sparse predictions where the $(n-1)$-gram corresponds to only a single document.Run ./infinigram --help for more information.
I use the text_64 function implemented in the Go suffixarray library---the files under suffixarray/ are from this library with minor modifications.
Go
99.6%
An unofficial implementation of the Infini-gram model proposed by Liu et al. (2024)
Go
33
73 commits
updated Jun 19, 2024
This repo contains two (unofficial) implementations of the infini-gram model described in Liu et al. (2024). This branch contains the Golang implementation. The main branch contains a Python implementation.
The tokenizers used here are the Go bindings to the official Rust library.
First, build the rust tokenizers binary:
cd tokenizers
make
Then, you can build the infinigram binary:
cd ../
go build -ldflags "-s"
./infinigram --train_file corpus.txt --out_dir output --tokenizer_config tokenizer.json
where corpus.txt contains one document per line. tokenizer.json corresponds to the HuggingFace pretrained Tokenizers file (e.g., for gpt2).
This implementation features:
--interactive_mode {0,1})mmap to access both the tokenized documents and the suffix array; memory usage during inference should be minimal.--max_mem): you should hypothetically be able to train (and infer) on any sized corpus regardless of how much memory you have--min_matches). e.g., you may set this at a value >= 2 to avoid sparse predictions where the $(n-1)$-gram corresponds to only a single document.Run ./infinigram --help for more information.
I use the text_64 function implemented in the Go suffixarray library---the files under suffixarray/ are from this library with minor modifications.
Go
99.6%