andreuka18/DeepSeek-R1-Distill-Llama-8B-lmsys-openthoughts-tokenized

Dataset

0

stars

4

commits

1

linked in READMEs

Mar 31, 2025

updated

README

This dataset is used for training Sparse Autoencoders (SAEs) to identify reasoning features in Large Language Models (LLMs), as described in the paper I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders.

Code for the paper is available at: https://github.com/AIRI-Institute/SAE-Reasoning

The dataset consists of tokenized text data used for training the SAEs.

dataset_info: features:

  • name: tokens sequence: int64 splits:
  • name: train num_bytes: 6403125000.0 num_examples: 781250 download_size: 1440049400 dataset_size: 6403125000.0 configs:
  • config_name: default data_files:
    • split: train path: data/train-*

Contributors

andreuka18

3 commits

nielsr

1 commits

andreuka18/DeepSeek-R1-Distill-Llama-8B-lmsys-openthoughts-tokenized

Dataset

0

stars

4

commits

1

linked in READMEs

Mar 31, 2025

updated

README

This dataset is used for training Sparse Autoencoders (SAEs) to identify reasoning features in Large Language Models (LLMs), as described in the paper I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders.

Code for the paper is available at: https://github.com/AIRI-Institute/SAE-Reasoning

The dataset consists of tokenized text data used for training the SAEs.

dataset_info: features:

  • name: tokens sequence: int64 splits:
  • name: train num_bytes: 6403125000.0 num_examples: 781250 download_size: 1440049400 dataset_size: 6403125000.0 configs:
  • config_name: default data_files:
    • split: train path: data/train-*

Contributors

andreuka18

3 commits

nielsr

1 commits