The Open RAG Benchmark is a unique, high-quality Retrieval-Augmented Generation (RAG) dataset constructed directly from arXiv PDF documents, specifically designed for evaluating RAG systems with a focus on multimodal PDF understanding. Unlike other datasets, Open RAG Benchmark emphasizes pure PDF content, meticulously extracting and generating queries on diverse modalities including text, tables, and images, even when they are intricately interwoven within a document.
This dataset is purpose-built to power the company's Open RAG Evaluation project, facilitating a holistic, end-to-end evaluation of RAG systems by offering:
The current draft version of the Arxiv dataset, as the first step in this multimodal RAG dataset collection, includes:
The dataset is organized similar to the BEIR dataset format within the official/pdf/arxiv/ directory.
official/
└── pdf
└── arxiv
├── answers.json
├── corpus
│ ├── {PAPER_ID_1}.json
│ ├── {PAPER_ID_2}.json
│ └── ...
├── pdf_urls.json
├── qrels.json
└── queries.json
Each file's format is detailed below:
pdf_urls.jsonThis file provides the original PDF links to the papers in this dataset for downloading purposes.
{
"Paper ID": "Paper URL",
...
}
corpus/This folder contains all processed papers in JSON format.
{
"title": "Paper Title",
"sections": [
{
"text": "Section text content with placeholders for tables/images",
"tables": {"table_id1": "markdown_table_string", ...},
"images": {"image_id1": "base64_encoded_string", ...},
},
...
],
"id": "Paper ID",
"authors": ["Author1", "Author2", ...],
"categories": ["Category1", "Category2", ...],
"abstract": "Abstract text",
"updated": "Updated date",
"published": "Published date"
}
queries.jsonThis file contains all generated queries.
{
"Query UUID": {
"query": "Query text",
"type": "Query type (abstractive/extractive)",
"source": "Generation source (text/text-image/text-table/text-table-image)"
},
...
}
qrels.jsonThis file contains the query-document-section relevance labels.
{
"Query UUID": {
"doc_id": "Paper ID",
"section_id": Section Index
},
...
}
answers.jsonThis file contains the answers for the generated queries.
{
"Query UUID": "Answer text",
...
}
The Open RAG Benchmark dataset is created through a systematic process involving document collection, processing, content segmentation, query generation, and quality filtering.
gpt-4o-mini) to generate retrieval queries for each section, handling multimodal content such as tables and images.gpt-4o-mini for query quality filtering.The code for reproducing and customizing the dataset generation process is available in the Open RAG Benchmark GitHub repository.
Several challenges are inherent in the current dataset development process:
The project aims for continuous improvement and expansion of the dataset, with key next steps including:
The Open RAG Benchmark project uses OpenAI's GPT models (specifically gpt-4o-mini) for query generation and evaluation. For post-filtering and retrieval filtering, the following embedding models, recognized for their outstanding performance on the MTEB Benchmark, were utilized:
The Open RAG Benchmark is a unique, high-quality Retrieval-Augmented Generation (RAG) dataset constructed directly from arXiv PDF documents, specifically designed for evaluating RAG systems with a focus on multimodal PDF understanding. Unlike other datasets, Open RAG Benchmark emphasizes pure PDF content, meticulously extracting and generating queries on diverse modalities including text, tables, and images, even when they are intricately interwoven within a document.
This dataset is purpose-built to power the company's Open RAG Evaluation project, facilitating a holistic, end-to-end evaluation of RAG systems by offering:
The current draft version of the Arxiv dataset, as the first step in this multimodal RAG dataset collection, includes:
The dataset is organized similar to the BEIR dataset format within the official/pdf/arxiv/ directory.
official/
└── pdf
└── arxiv
├── answers.json
├── corpus
│ ├── {PAPER_ID_1}.json
│ ├── {PAPER_ID_2}.json
│ └── ...
├── pdf_urls.json
├── qrels.json
└── queries.json
Each file's format is detailed below:
pdf_urls.jsonThis file provides the original PDF links to the papers in this dataset for downloading purposes.
{
"Paper ID": "Paper URL",
...
}
corpus/This folder contains all processed papers in JSON format.
{
"title": "Paper Title",
"sections": [
{
"text": "Section text content with placeholders for tables/images",
"tables": {"table_id1": "markdown_table_string", ...},
"images": {"image_id1": "base64_encoded_string", ...},
},
...
],
"id": "Paper ID",
"authors": ["Author1", "Author2", ...],
"categories": ["Category1", "Category2", ...],
"abstract": "Abstract text",
"updated": "Updated date",
"published": "Published date"
}
queries.jsonThis file contains all generated queries.
{
"Query UUID": {
"query": "Query text",
"type": "Query type (abstractive/extractive)",
"source": "Generation source (text/text-image/text-table/text-table-image)"
},
...
}
qrels.jsonThis file contains the query-document-section relevance labels.
{
"Query UUID": {
"doc_id": "Paper ID",
"section_id": Section Index
},
...
}
answers.jsonThis file contains the answers for the generated queries.
{
"Query UUID": "Answer text",
...
}
The Open RAG Benchmark dataset is created through a systematic process involving document collection, processing, content segmentation, query generation, and quality filtering.
gpt-4o-mini) to generate retrieval queries for each section, handling multimodal content such as tables and images.gpt-4o-mini for query quality filtering.The code for reproducing and customizing the dataset generation process is available in the Open RAG Benchmark GitHub repository.
Several challenges are inherent in the current dataset development process:
The project aims for continuous improvement and expansion of the dataset, with key next steps including:
The Open RAG Benchmark project uses OpenAI's GPT models (specifically gpt-4o-mini) for query generation and evaluation. For post-filtering and retrieval filtering, the following embedding models, recognized for their outstanding performance on the MTEB Benchmark, were utilized: