This dataset gathers the sub-datasets of supervised and synthetized samples necessary to fine-tune on context compression tasks an ARC-Encoder as described in the paper ARC-Encoder: learning compressed text representations for large language models available here.
It consists in 12 jsonl files separated in 4 task categories: Translation, Question-Answering, Reading Comprehension and Summarization. To fine-tune your ARC-Encoder from the HF collection ARC-Encoders follow the recipe described in the paper and use the following codebase ARC-Encoder. Proportion for sampling among these datasets are described in the Appendix.
We gathered already existing datasets which sources are listed below:
For the first 5 datasets (QA samples), we retrieved 5 passages of KILT (MIT license) Wikipedia passage chunks using NVEmbed v.2, CC BY-NC 4.0.
For the translations, we used passages from ATLAS, CC-BY-SA, and translate them using Gemma 3 27B, Gemma licence, in:
Sub-datasets are kept separated as at training time we want to be able to gather in-context example from each dataset independantly to design the final fine-tuning samples.
ARC-Encoder fine-tuning is licensed under the CC-BY 4.0 license.
If you use this dataset, please cite:
@article{
pilchen2026arcencoder,
title={{ARC}-Encoder: learning compressed text representations for large language models},
author={Hippolyte Pilchen and Edouard Grave and Patrick Perez},
journal={Transactions on Machine Learning Research},
issn={2835-8856},
year={2026},
url={https://openreview.net/forum?id=lU1P9dsqfn},
note={Featured Certification}
}
This dataset gathers the sub-datasets of supervised and synthetized samples necessary to fine-tune on context compression tasks an ARC-Encoder as described in the paper ARC-Encoder: learning compressed text representations for large language models available here.
It consists in 12 jsonl files separated in 4 task categories: Translation, Question-Answering, Reading Comprehension and Summarization. To fine-tune your ARC-Encoder from the HF collection ARC-Encoders follow the recipe described in the paper and use the following codebase ARC-Encoder. Proportion for sampling among these datasets are described in the Appendix.
We gathered already existing datasets which sources are listed below:
For the first 5 datasets (QA samples), we retrieved 5 passages of KILT (MIT license) Wikipedia passage chunks using NVEmbed v.2, CC BY-NC 4.0.
For the translations, we used passages from ATLAS, CC-BY-SA, and translate them using Gemma 3 27B, Gemma licence, in:
Sub-datasets are kept separated as at training time we want to be able to gather in-context example from each dataset independantly to design the final fine-tuning samples.
ARC-Encoder fine-tuning is licensed under the CC-BY 4.0 license.
If you use this dataset, please cite:
@article{
pilchen2026arcencoder,
title={{ARC}-Encoder: learning compressed text representations for large language models},
author={Hippolyte Pilchen and Edouard Grave and Patrick Perez},
journal={Transactions on Machine Learning Research},
issn={2835-8856},
year={2026},
url={https://openreview.net/forum?id=lU1P9dsqfn},
note={Featured Certification}
}