LE

Lemoncoke/Marathon

Dataset

4

stars

8

commits

3

linked in READMEs

May 16, 2024

updated

long context

README

Dataset Card for Marathon

Release

  • [2024/05/15] πŸ”₯ Marathon is accepted by ACL 2024 Main Conference.

Dataset Summary

Marathon benchmark is a new long-context multiple-choice benchmark, mainly based on LooGLE, with some original data from LongBench. The context length can reach up to 200K+. Marathon benchmark comprises six tasks: Comprehension and Reasoning, Multiple Information Retrieval, Timeline Reorder, Computation, Passage Retrieval, and Short Dependency Question Answering. Each test case includes a Long Context, a question, and multiple candidate options. Large Language Models (LLMs) need to select the correct answer from the given options based on the Long Context in the test.

Github

Marathon is also available at Github: Marathon.

Data Instances

An example of test looks as follows. This is a toy example.

{
	"id": "7",
  "type": "comprehension_and_reasoning",
  "context": " Early life. Picardo was born in Jerez de la Frontera, in the Province of CΓ‘diz in AndalucΓ­a, Spain on 18 June 1919. His father was Alvaro Picardo de Celis and his mother's family name was CastellΓ³n. He had four brothers, one of whom died in infancy. His father died in 1929 when Picardo was ten years old. With his mother and his brothers he moved to Madrid, Spain. [Truncated for display purpose] ",
  "question": "How many people were in Picardo's family when he was twelve?",
  "options": {
    "A": "five",
    "B": "eight",
    "C": "nine",
    "D": "ten"
  },
  "length": 268760
}
  • Methods (optimizing methods):
    • 🏐 Vanilla
    • 🎾 RAG (Retrieval Augmented Generation)
    • πŸ€ PC (LongLLMLingua Prompt Compression)
  • Embedding Models:
    • 🍿 OpenAI: text-embedding-ada-002
    • πŸ” Jina: Jina-Embedding-base
TagModelParametersContext WindowMethodEmbeddingAvg. Accuracy ⬆️
🏐GPT-4-128K🏐 Vanilla-78.59
πŸŽΎπŸ”Yi-chat34B200K🎾 RAGπŸ” Jina63.81
🎾🍿Yi-chat34B200K🎾 RAG🍿 OpenAI63.56
🎾🍿Tutu2-DPO70B8K🎾 RAG🍿 OpenAI61.97
πŸŽΎπŸ”Tutu2-DPO70B8K🎾 RAGπŸ” Jina61.52
πŸŽΎπŸ”Qwen14B8K🎾 RAGπŸ” Jina58.12
🏐ChatGPT-16K🏐 Vanilla-57.37
🏐Yi-chat34B200K🏐 Vanilla-55.91
πŸŽΎπŸ”Beluga270B4K🎾 RAGπŸ” Jina55.72
🏐ChatGLM36B32K🏐 Vanilla-55.05
πŸŽΎπŸ”Zephyr7B32K🎾 RAGπŸ” Jina53.79
🎾🍿Qwen14B8K🎾 RAG🍿 OpenAI53.46
πŸ€Beluga270B4KπŸ€ PC-52.29
πŸŽΎπŸ”Mistral7B32K🎾 RAGπŸ” Jina52.04
🎾🍿Alfred40B8K🎾 RAG🍿 OpenAI51.35
πŸŽΎπŸ”Alfred40B8K🎾 RAGπŸ” Jina51.24
🎾🍿ChatGLM36B32K🎾 RAG🍿 OpenAI50.99
πŸŽΎπŸ”ChatGLM36B32K🎾 RAGπŸ” Jina50.60
🎾🍿Mistral7B32K🎾 RAG🍿 OpenAI50.18
🎾🍿Zephyr7B32K🎾 RAG🍿 OpenAI49.63
🏐Beluga270B4K🏐 Vanilla-49.51
πŸ€Yi34B200KπŸ€ PC-48.66
🎾🍿Beluga270B4K🎾 RAG🍿 OpenAI48.24
πŸ€ChatGLM36B32KπŸ€ PC-47.91
πŸ€Tulu2-DPO70B8KπŸ€ PC-46.56
πŸ€Qwen14B8KπŸ€ PC-44.12
🏐Mistral7B32K🏐 Vanilla-39.81
🏐Qwen14B8K🏐 Vanilla-39.27
πŸ€Alfred40B8KπŸ€ PC-38.82
🏐Zephyr7B32K🏐 Vanilla-37.97
🏐Tulu2-DPO7B8K🏐 Vanilla-37.92
πŸŽΎπŸ”Longchat13B16K🎾 RAGπŸ” Jina37.78
🏐Alfred40B8K🏐 Vanilla-37.31
πŸ€Mistral7B32KπŸ€ PC-37.01
🏐Longchat13B16K🏐 Vanilla-35.87
πŸ€Longchat13B16KπŸ€ PC-35.61
πŸ€Zephyr7B32KπŸ€ PC-30.23
🎾🍿Longchat13B16K🎾 RAG🍿 OpenAI29.95

Online Evaluation

Welcome to Marathon Race, online evaluation is now available at https://openbenchmark.online/marathon.

Answer File Format

The file should be a JSON file containing a list of dictionaries with a length of 1530. Each dictionary must include at least two fields: 'id' and 'answer'. Here is a sample answer file:

[
  {
    "id": "0",
    "answer": "C"
  },
  {
    "id": "1",
    "answer": "B"
  },
  {
    "id": "2",
    "answer": "B"
  },
  ...
   {
    "id": "1529",
    "answer": "C"
  }
]

Results File Format

The Results file is a JSON file that includes the accuracy of the LLM (Language Learning Model) in 6 tasks within the Marathon, as well as the average accuracy across all tasks. Here is a sample results file:

{
    "comprehension_and_reasoning": {
        "accuracy": 0.46218487394957986,
        "correct": 165,
        "total": 357
    },
    "multiple_information_retrieval": {
        "accuracy": 0.41935483870967744,
        "correct": 143,
        "total": 341
    },
    "timeline_reorder": {
        "accuracy": 0.2894736842105263,
        "correct": 44,
        "total": 152
    },
    "computation": {
        "accuracy": 0.23711340206185566,
        "correct": 23,
        "total": 97
    },
    "passage_retrieval": {
        "accuracy": 0.49666666666666665,
        "correct": 149,
        "total": 300
    },
    "shortdep_qa": {
        "accuracy": 0.4840989399293286,
        "correct": 137,
        "total": 283
    },
    "average": 0.39814873425460573
}

Citations

If you find our work useful, please cite us.

@article{zhang2023marathon,
  title={Marathon: A Race Through the Realm of Long Context with Large Language Models},
  author={Zhang, Lei and Li, Yunshui and Liu, Ziqiang and Liu, Junhao and Yang, Jiaxi and Yang, Min},
  url={https://huggingface.co/datasets/Lemoncoke/Marathon},
  year={2023}
}

When citing our work, please kindly consider citing the original dataset papers.

@misc{li2023loogle,
  title={Can Long-Context Language Models Understand Long Contexts?},
  author={ Li, Jiaqi and Wang, Mengmeng and Zheng, Zilong and Zhang, Muhan },
  url={https://github.com/bigai-nlco/LooGLE},
  year={2023}
}
@article{bai2023longbench,
  title={LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding},
  author={Bai, Yushi and Lv, Xin and Zhang, Jiajie and Lyu, Hongchang and Tang, Jiankai and Huang, Zhidian and Du, Zhengxiao and Liu, Xiao and Zeng, Aohan and Hou, Lei and Dong, Yuxiao and Tang, Jie and Li, Juanzi},
  journal={arXiv preprint arXiv:2308.14508},
  year={2023}
}

Contributors

LZ
Lei Zhang

7 commits

Hambaobao

1 commits

LE

Lemoncoke/Marathon

Dataset

4

stars

8

commits

3

linked in READMEs

May 16, 2024

updated

long context

README

Dataset Card for Marathon

Release

  • [2024/05/15] πŸ”₯ Marathon is accepted by ACL 2024 Main Conference.

Dataset Summary

Marathon benchmark is a new long-context multiple-choice benchmark, mainly based on LooGLE, with some original data from LongBench. The context length can reach up to 200K+. Marathon benchmark comprises six tasks: Comprehension and Reasoning, Multiple Information Retrieval, Timeline Reorder, Computation, Passage Retrieval, and Short Dependency Question Answering. Each test case includes a Long Context, a question, and multiple candidate options. Large Language Models (LLMs) need to select the correct answer from the given options based on the Long Context in the test.

Github

Marathon is also available at Github: Marathon.

Data Instances

An example of test looks as follows. This is a toy example.

{
	"id": "7",
  "type": "comprehension_and_reasoning",
  "context": " Early life. Picardo was born in Jerez de la Frontera, in the Province of CΓ‘diz in AndalucΓ­a, Spain on 18 June 1919. His father was Alvaro Picardo de Celis and his mother's family name was CastellΓ³n. He had four brothers, one of whom died in infancy. His father died in 1929 when Picardo was ten years old. With his mother and his brothers he moved to Madrid, Spain. [Truncated for display purpose] ",
  "question": "How many people were in Picardo's family when he was twelve?",
  "options": {
    "A": "five",
    "B": "eight",
    "C": "nine",
    "D": "ten"
  },
  "length": 268760
}
  • Methods (optimizing methods):
    • 🏐 Vanilla
    • 🎾 RAG (Retrieval Augmented Generation)
    • πŸ€ PC (LongLLMLingua Prompt Compression)
  • Embedding Models:
    • 🍿 OpenAI: text-embedding-ada-002
    • πŸ” Jina: Jina-Embedding-base
TagModelParametersContext WindowMethodEmbeddingAvg. Accuracy ⬆️
🏐GPT-4-128K🏐 Vanilla-78.59
πŸŽΎπŸ”Yi-chat34B200K🎾 RAGπŸ” Jina63.81
🎾🍿Yi-chat34B200K🎾 RAG🍿 OpenAI63.56
🎾🍿Tutu2-DPO70B8K🎾 RAG🍿 OpenAI61.97
πŸŽΎπŸ”Tutu2-DPO70B8K🎾 RAGπŸ” Jina61.52
πŸŽΎπŸ”Qwen14B8K🎾 RAGπŸ” Jina58.12
🏐ChatGPT-16K🏐 Vanilla-57.37
🏐Yi-chat34B200K🏐 Vanilla-55.91
πŸŽΎπŸ”Beluga270B4K🎾 RAGπŸ” Jina55.72
🏐ChatGLM36B32K🏐 Vanilla-55.05
πŸŽΎπŸ”Zephyr7B32K🎾 RAGπŸ” Jina53.79
🎾🍿Qwen14B8K🎾 RAG🍿 OpenAI53.46
πŸ€Beluga270B4KπŸ€ PC-52.29
πŸŽΎπŸ”Mistral7B32K🎾 RAGπŸ” Jina52.04
🎾🍿Alfred40B8K🎾 RAG🍿 OpenAI51.35
πŸŽΎπŸ”Alfred40B8K🎾 RAGπŸ” Jina51.24
🎾🍿ChatGLM36B32K🎾 RAG🍿 OpenAI50.99
πŸŽΎπŸ”ChatGLM36B32K🎾 RAGπŸ” Jina50.60
🎾🍿Mistral7B32K🎾 RAG🍿 OpenAI50.18
🎾🍿Zephyr7B32K🎾 RAG🍿 OpenAI49.63
🏐Beluga270B4K🏐 Vanilla-49.51
πŸ€Yi34B200KπŸ€ PC-48.66
🎾🍿Beluga270B4K🎾 RAG🍿 OpenAI48.24
πŸ€ChatGLM36B32KπŸ€ PC-47.91
πŸ€Tulu2-DPO70B8KπŸ€ PC-46.56
πŸ€Qwen14B8KπŸ€ PC-44.12
🏐Mistral7B32K🏐 Vanilla-39.81
🏐Qwen14B8K🏐 Vanilla-39.27
πŸ€Alfred40B8KπŸ€ PC-38.82
🏐Zephyr7B32K🏐 Vanilla-37.97
🏐Tulu2-DPO7B8K🏐 Vanilla-37.92
πŸŽΎπŸ”Longchat13B16K🎾 RAGπŸ” Jina37.78
🏐Alfred40B8K🏐 Vanilla-37.31
πŸ€Mistral7B32KπŸ€ PC-37.01
🏐Longchat13B16K🏐 Vanilla-35.87
πŸ€Longchat13B16KπŸ€ PC-35.61
πŸ€Zephyr7B32KπŸ€ PC-30.23
🎾🍿Longchat13B16K🎾 RAG🍿 OpenAI29.95

Online Evaluation

Welcome to Marathon Race, online evaluation is now available at https://openbenchmark.online/marathon.

Answer File Format

The file should be a JSON file containing a list of dictionaries with a length of 1530. Each dictionary must include at least two fields: 'id' and 'answer'. Here is a sample answer file:

[
  {
    "id": "0",
    "answer": "C"
  },
  {
    "id": "1",
    "answer": "B"
  },
  {
    "id": "2",
    "answer": "B"
  },
  ...
   {
    "id": "1529",
    "answer": "C"
  }
]

Results File Format

The Results file is a JSON file that includes the accuracy of the LLM (Language Learning Model) in 6 tasks within the Marathon, as well as the average accuracy across all tasks. Here is a sample results file:

{
    "comprehension_and_reasoning": {
        "accuracy": 0.46218487394957986,
        "correct": 165,
        "total": 357
    },
    "multiple_information_retrieval": {
        "accuracy": 0.41935483870967744,
        "correct": 143,
        "total": 341
    },
    "timeline_reorder": {
        "accuracy": 0.2894736842105263,
        "correct": 44,
        "total": 152
    },
    "computation": {
        "accuracy": 0.23711340206185566,
        "correct": 23,
        "total": 97
    },
    "passage_retrieval": {
        "accuracy": 0.49666666666666665,
        "correct": 149,
        "total": 300
    },
    "shortdep_qa": {
        "accuracy": 0.4840989399293286,
        "correct": 137,
        "total": 283
    },
    "average": 0.39814873425460573
}

Citations

If you find our work useful, please cite us.

@article{zhang2023marathon,
  title={Marathon: A Race Through the Realm of Long Context with Large Language Models},
  author={Zhang, Lei and Li, Yunshui and Liu, Ziqiang and Liu, Junhao and Yang, Jiaxi and Yang, Min},
  url={https://huggingface.co/datasets/Lemoncoke/Marathon},
  year={2023}
}

When citing our work, please kindly consider citing the original dataset papers.

@misc{li2023loogle,
  title={Can Long-Context Language Models Understand Long Contexts?},
  author={ Li, Jiaqi and Wang, Mengmeng and Zheng, Zilong and Zhang, Muhan },
  url={https://github.com/bigai-nlco/LooGLE},
  year={2023}
}
@article{bai2023longbench,
  title={LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding},
  author={Bai, Yushi and Lv, Xin and Zhang, Jiajie and Lyu, Hongchang and Tang, Jiankai and Huang, Zhidian and Du, Zhengxiao and Liu, Xiao and Zeng, Aohan and Hou, Lei and Dong, Yuxiao and Tang, Jie and Li, Juanzi},
  journal={arXiv preprint arXiv:2308.14508},
  year={2023}
}

Contributors

LZ
Lei Zhang

7 commits

Hambaobao

1 commits