mmrag_train.json: Training set for model training.mmrag_dev.json: Validation set for hyperparameter tuning and development.mmrag_test.json: Test set for evaluation.processed_documents.json: The chunks used for retrieval.You can load and work with the mmRAG dataset using standard Python libraries like json. Below is a simple example of how to load and interact with the data files.
import json
# Load query datasets
with open("mmrag_train.json", "r", encoding="utf-8") as f:
train_data = json.load(f)
with open("mmrag_dev.json", "r", encoding="utf-8") as f:
dev_data = json.load(f)
with open("mmrag_test.json", "r", encoding="utf-8") as f:
test_data = json.load(f)
# Load document chunks
with open("processed_documents.json", "r", encoding="utf-8") as f:
documents = json.load(f)
# Load as dict if needed
documents = {doc["id"]: doc["text"] for doc in documents}
# Example query
query_example = train_data[0]
print("Query:", query_example["query"])
print("Answer:", query_example["answer"])
print("Relevant Chunks:", query_example["relevant_chunks"])
# Get the text of a relevant chunk
for chunk_id, relevance in query_example["relevant_chunks"].items():
if relevance > 0:
print(f"Chunk ID: {chunk_id}, Relevance label: {relevance}\nText: {documents[chunk_id]}")
The following example shows how to extract and sort the dataset_score field of a query to understand which dataset is most relevant to the query.
# Choose a query from the dataset
query_example = train_data[0]
print("Query:", query_example["query"])
print("Answer:", query_example["answer"])
# Get dataset routing scores
routing_scores = query_example["dataset_score"]
# Sort datasets by relevance score (descending)
sorted_routing = sorted(routing_scores.items(), key=lambda x: x[1], reverse=True)
print("\nRouting Results (sorted):")
for dataset, score in sorted_routing:
print(f"{dataset}: {score}")
mmrag_train.json, mmrag_dev.json, mmrag_test.jsonThe three files are all lists of dictionaries. Each dictionary contains the following fields:
idSourceDataset_queryIDinDataset.ott_144, means this query is picked from OTT-QA datasetquery"What is the capital of France?"answer"Paris"relevant_chunksjson{"ott_23573_2": 1, "ott_114_0": 2, "m.12345_0": 0}ori_context["ott_144"], means all chunk IDs start with "ott_114" is from the original document.dataset_score{"tat": 0, "triviaqa": 2, "ott": 4, "kg": 1, "nq": 0}, where 0 means there is no relevant chunks in the dataset. The higher the score is, the more relevant chunks the dataset have.processed_documents.jsonThis file is a list of chunks used for document retrieval, which contains the following fields:
iddataset_documentID_chunkIndex, equivalent to dataset_queryID_chunkIndexott_8075_0 (chunks from NQ, TriviaQA, OTT, TAT)m.0cpy1b_5 (chunks from documents of knowledge graph(Freebase))textA molecule editor is a computer program for creating and modifying representations of chemical structures.This dataset is licensed under the Apache License 2.0.
mmrag_train.json: Training set for model training.mmrag_dev.json: Validation set for hyperparameter tuning and development.mmrag_test.json: Test set for evaluation.processed_documents.json: The chunks used for retrieval.You can load and work with the mmRAG dataset using standard Python libraries like json. Below is a simple example of how to load and interact with the data files.
import json
# Load query datasets
with open("mmrag_train.json", "r", encoding="utf-8") as f:
train_data = json.load(f)
with open("mmrag_dev.json", "r", encoding="utf-8") as f:
dev_data = json.load(f)
with open("mmrag_test.json", "r", encoding="utf-8") as f:
test_data = json.load(f)
# Load document chunks
with open("processed_documents.json", "r", encoding="utf-8") as f:
documents = json.load(f)
# Load as dict if needed
documents = {doc["id"]: doc["text"] for doc in documents}
# Example query
query_example = train_data[0]
print("Query:", query_example["query"])
print("Answer:", query_example["answer"])
print("Relevant Chunks:", query_example["relevant_chunks"])
# Get the text of a relevant chunk
for chunk_id, relevance in query_example["relevant_chunks"].items():
if relevance > 0:
print(f"Chunk ID: {chunk_id}, Relevance label: {relevance}\nText: {documents[chunk_id]}")
The following example shows how to extract and sort the dataset_score field of a query to understand which dataset is most relevant to the query.
# Choose a query from the dataset
query_example = train_data[0]
print("Query:", query_example["query"])
print("Answer:", query_example["answer"])
# Get dataset routing scores
routing_scores = query_example["dataset_score"]
# Sort datasets by relevance score (descending)
sorted_routing = sorted(routing_scores.items(), key=lambda x: x[1], reverse=True)
print("\nRouting Results (sorted):")
for dataset, score in sorted_routing:
print(f"{dataset}: {score}")
mmrag_train.json, mmrag_dev.json, mmrag_test.jsonThe three files are all lists of dictionaries. Each dictionary contains the following fields:
idSourceDataset_queryIDinDataset.ott_144, means this query is picked from OTT-QA datasetquery"What is the capital of France?"answer"Paris"relevant_chunksjson{"ott_23573_2": 1, "ott_114_0": 2, "m.12345_0": 0}ori_context["ott_144"], means all chunk IDs start with "ott_114" is from the original document.dataset_score{"tat": 0, "triviaqa": 2, "ott": 4, "kg": 1, "nq": 0}, where 0 means there is no relevant chunks in the dataset. The higher the score is, the more relevant chunks the dataset have.processed_documents.jsonThis file is a list of chunks used for document retrieval, which contains the following fields:
iddataset_documentID_chunkIndex, equivalent to dataset_queryID_chunkIndexott_8075_0 (chunks from NQ, TriviaQA, OTT, TAT)m.0cpy1b_5 (chunks from documents of knowledge graph(Freebase))textA molecule editor is a computer program for creating and modifying representations of chemical structures.This dataset is licensed under the Apache License 2.0.