Haiyang-W/TokenFormer-900M

Model

The *TokenFormer* is a fully attention-based architecture

3

10 commits

3 linked in READMEs

updated Nov 9, 2024

See the code

README

The TokenFormer is a fully attention-based architecture that unifies the computations of token-token and token-parameter interactions by entirely employing the attention mechanism, maximizes the flexibility of neural network.(see paper). It contains four models of sizes 150M, 450M, 900M, 1.5B. For each size, it's trained based on gpt-neox code base and uses Pile with 300B tokens. All 4 model sizes are trained on the exact same data, in the exact same order.

TokenFormer-900M

Model Details

  • Developed by: Haiyang Wang
  • Model type: TokenFormer-based Language Model
  • Language: English
  • Learn more: TokenFormer's GitHub repository for training procedure, config files, and details on how to use. See paper for more evals and implementation details.
  • Library: GPT-NeoX
  • License: Apache 2.0
  • Contact: to ask questions about this model, please email Haiyang Wang.
TokenFormer modelLayers#QKV Param Tokens#Output Param Tokens#FFN Param TokensModel DimHeadsBatch SizeLearning RateTraining Iterations
150M127687683072768122M6.0 x 10-4143000
450M241024102440961024162M6.0 x 10-4143000
900M321280128051201280162M6.0 x 10-4143000
1.5B401536153661441536162M6.0 x 10-4143000
Engineering details for the TokenFormer.

Training

Training data

The Pile is a 825GiB general-purpose dataset in English. It was created by EleutherAI specifically for training large language models. It contains texts from 22 diverse sources, roughly broken down into five categories: academic writing (e.g. arXiv), internet (e.g. CommonCrawl), prose (e.g. Project Gutenberg), dialogue (e.g. YouTube subtitles), and miscellaneous (e.g. GitHub, Enron Emails). See the Pile paper for a breakdown of all data sources, methodology, and a discussion of ethical implications. Consult the datasheet for more detailed documentation about the Pile and its component datasets. The Pile can be downloaded from the official website, or from a community mirror.

Training procedure

We follow the default training strategy of Pythia in gpt-neox, including the dataset processing, hyper-parameter and code base. All models were trained on the exact same data, in the exact same order. Each model saw 299,892,736,000 tokens during training.

All TokenFormer models trained for 143000 steps at a batch size of 2M (2,097,152 tokens).
See GitHub for more details on training procedure.
TokenFormer uses the same tokenizer as GPT-NeoX- 20B.

Evaluations

All TokenFormer models were evaluated using the LM Evaluation Harness. You can run the evaluation with our instruction.
Expand the sections below to see plots of evaluation results for all TokenFormer compared with Opensource Transformer-based LLMs.

Model#ParamLAMBADAHellaSwagPIQAArc-EArc-CWinoGrandeAverage
Pythia150M35.430.362.343.623.651.340.1
TokenFormer150M45.035.564.947.324.950.444.7
Pythia410M51.440.666.952.124.653.848.2
TokenFormer450M57.347.569.556.226.754.652.0
Pythia1B56.147.270.757.027.153.551.9
TokenFormer900M64.055.372.459.930.656.456.4
GPT-Neo1.3B57.248.971.156.225.954.952.4
OPT1.3B58.053.772.456.729.659.555.0
Pythia1.3B61.752.171.060.528.557.255.2
GPT-Neo2.7B62.255.871.161.130.257.656.5
OPT2.7B63.660.674.860.831.361.058.7
Pythia2.8B64.759.374.064.132.959.759.1
TokenFormer1.5B64.760.074.864.832.059.759.3
Zero-shot evaluation of Language Modeling.
pytorch

Haiyang-W/TokenFormer-900M

Model

The *TokenFormer* is a fully attention-based architecture

3

10 commits

3 linked in READMEs

updated Nov 9, 2024

See the code

README

The TokenFormer is a fully attention-based architecture that unifies the computations of token-token and token-parameter interactions by entirely employing the attention mechanism, maximizes the flexibility of neural network.(see paper). It contains four models of sizes 150M, 450M, 900M, 1.5B. For each size, it's trained based on gpt-neox code base and uses Pile with 300B tokens. All 4 model sizes are trained on the exact same data, in the exact same order.

TokenFormer-900M

Model Details

  • Developed by: Haiyang Wang
  • Model type: TokenFormer-based Language Model
  • Language: English
  • Learn more: TokenFormer's GitHub repository for training procedure, config files, and details on how to use. See paper for more evals and implementation details.
  • Library: GPT-NeoX
  • License: Apache 2.0
  • Contact: to ask questions about this model, please email Haiyang Wang.
TokenFormer modelLayers#QKV Param Tokens#Output Param Tokens#FFN Param TokensModel DimHeadsBatch SizeLearning RateTraining Iterations
150M127687683072768122M6.0 x 10-4143000
450M241024102440961024162M6.0 x 10-4143000
900M321280128051201280162M6.0 x 10-4143000
1.5B401536153661441536162M6.0 x 10-4143000
Engineering details for the TokenFormer.

Training

Training data

The Pile is a 825GiB general-purpose dataset in English. It was created by EleutherAI specifically for training large language models. It contains texts from 22 diverse sources, roughly broken down into five categories: academic writing (e.g. arXiv), internet (e.g. CommonCrawl), prose (e.g. Project Gutenberg), dialogue (e.g. YouTube subtitles), and miscellaneous (e.g. GitHub, Enron Emails). See the Pile paper for a breakdown of all data sources, methodology, and a discussion of ethical implications. Consult the datasheet for more detailed documentation about the Pile and its component datasets. The Pile can be downloaded from the official website, or from a community mirror.

Training procedure

We follow the default training strategy of Pythia in gpt-neox, including the dataset processing, hyper-parameter and code base. All models were trained on the exact same data, in the exact same order. Each model saw 299,892,736,000 tokens during training.

All TokenFormer models trained for 143000 steps at a batch size of 2M (2,097,152 tokens).
See GitHub for more details on training procedure.
TokenFormer uses the same tokenizer as GPT-NeoX- 20B.

Evaluations

All TokenFormer models were evaluated using the LM Evaluation Harness. You can run the evaluation with our instruction.
Expand the sections below to see plots of evaluation results for all TokenFormer compared with Opensource Transformer-based LLMs.

Model#ParamLAMBADAHellaSwagPIQAArc-EArc-CWinoGrandeAverage
Pythia150M35.430.362.343.623.651.340.1
TokenFormer150M45.035.564.947.324.950.444.7
Pythia410M51.440.666.952.124.653.848.2
TokenFormer450M57.347.569.556.226.754.652.0
Pythia1B56.147.270.757.027.153.551.9
TokenFormer900M64.055.372.459.930.656.456.4
GPT-Neo1.3B57.248.971.156.225.954.952.4
OPT1.3B58.053.772.456.729.659.555.0
Pythia1.3B61.752.171.060.528.557.255.2
GPT-Neo2.7B62.255.871.161.130.257.656.5
OPT2.7B63.660.674.860.831.361.058.7
Pythia2.8B64.759.374.064.132.959.759.1
TokenFormer1.5B64.760.074.864.832.059.759.3
Zero-shot evaluation of Language Modeling.
pytorch