kyutai/ARC_finetuning

Dataset

ARC-Encoder finetuning dataset

1

15 commits

3 linked in READMEs

updated Jul 30, 2026

See the code

README

ARC-Encoder finetuning dataset

This dataset gathers the sub-datasets of supervised and synthetized samples necessary to fine-tune on context compression tasks an ARC-Encoder as described in the paper ARC-Encoder: learning compressed text representations for large language models available here.

Dataset Details

Dataset Description

It consists in 12 jsonl files separated in 4 task categories: Translation, Question-Answering, Reading Comprehension and Summarization. To fine-tune your ARC-Encoder from the HF collection ARC-Encoders follow the recipe described in the paper and use the following codebase ARC-Encoder. Proportion for sampling among these datasets are described in the Appendix.

Dataset Sources

We gathered already existing datasets which sources are listed below:

For the first 5 datasets (QA samples), we retrieved 5 passages of KILT (MIT license) Wikipedia passage chunks using NVEmbed v.2, CC BY-NC 4.0.

For the translations, we used passages from ATLAS, CC-BY-SA, and translate them using Gemma 3 27B, Gemma licence, in:

  • Spanish, French, German and Danish
  • Hindi, Russian, Swahili, Arabic, Turkish, Japanese, Finnish and Chinese (simplified)

Uses

Sub-datasets are kept separated as at training time we want to be able to gather in-context example from each dataset independantly to design the final fine-tuning samples.

Licensing

ARC-Encoder fine-tuning is licensed under the CC-BY 4.0 license.

Citations

If you use this dataset, please cite:

@article{
pilchen2026arcencoder,
title={{ARC}-Encoder: learning compressed text representations for large language models},
author={Hippolyte Pilchen and Edouard Grave and Patrick Perez},
journal={Transactions on Machine Learning Research},
issn={2835-8856},
year={2026},
url={https://openreview.net/forum?id=lU1P9dsqfn},
note={Featured Certification}
}

kyutai/ARC_finetuning

Dataset

ARC-Encoder finetuning dataset

1

15 commits

3 linked in READMEs

updated Jul 30, 2026

See the code

README

ARC-Encoder finetuning dataset

This dataset gathers the sub-datasets of supervised and synthetized samples necessary to fine-tune on context compression tasks an ARC-Encoder as described in the paper ARC-Encoder: learning compressed text representations for large language models available here.

Dataset Details

Dataset Description

It consists in 12 jsonl files separated in 4 task categories: Translation, Question-Answering, Reading Comprehension and Summarization. To fine-tune your ARC-Encoder from the HF collection ARC-Encoders follow the recipe described in the paper and use the following codebase ARC-Encoder. Proportion for sampling among these datasets are described in the Appendix.

Dataset Sources

We gathered already existing datasets which sources are listed below:

For the first 5 datasets (QA samples), we retrieved 5 passages of KILT (MIT license) Wikipedia passage chunks using NVEmbed v.2, CC BY-NC 4.0.

For the translations, we used passages from ATLAS, CC-BY-SA, and translate them using Gemma 3 27B, Gemma licence, in:

  • Spanish, French, German and Danish
  • Hindi, Russian, Swahili, Arabic, Turkish, Japanese, Finnish and Chinese (simplified)

Uses

Sub-datasets are kept separated as at training time we want to be able to gather in-context example from each dataset independantly to design the final fine-tuning samples.

Licensing

ARC-Encoder fine-tuning is licensed under the CC-BY 4.0 license.

Citations

If you use this dataset, please cite:

@article{
pilchen2026arcencoder,
title={{ARC}-Encoder: learning compressed text representations for large language models},
author={Hippolyte Pilchen and Edouard Grave and Patrick Perez},
journal={Transactions on Machine Learning Research},
issn={2835-8856},
year={2026},
url={https://openreview.net/forum?id=lU1P9dsqfn},
note={Featured Certification}
}