TriviaqQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaqQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions.
English.
An example of 'train' looks as follows.
An example of 'train' looks as follows.
An example of 'validation' looks as follows.
An example of 'train' looks as follows.
The data fields are the same among all splits.
question: a string feature.question_id: a string feature.question_source: a string feature.entity_pages: a dictionary feature containing:
doc_source: a string feature.filename: a string feature.title: a string feature.wiki_context: a string feature.search_results: a dictionary feature containing:
description: a string feature.filename: a string feature.rank: a int32 feature.title: a string feature.url: a string feature.search_context: a string feature.aliases: a list of string features.normalized_aliases: a list of string features.matched_wiki_entity_name: a string feature.normalized_matched_wiki_entity_name: a string feature.normalized_value: a string feature.type: a string feature.value: a string feature.question: a string feature.question_id: a string feature.question_source: a string feature.entity_pages: a dictionary feature containing:
doc_source: a string feature.filename: a string feature.title: a string feature.wiki_context: a string feature.search_results: a dictionary feature containing:
description: a string feature.filename: a string feature.rank: a int32 feature.title: a string feature.url: a string feature.search_context: a string feature.aliases: a list of string features.normalized_aliases: a list of string features.matched_wiki_entity_name: a string feature.normalized_matched_wiki_entity_name: a string feature.normalized_value: a string feature.type: a string feature.value: a string feature.question: a string feature.question_id: a string feature.question_source: a string feature.entity_pages: a dictionary feature containing:
doc_source: a string feature.filename: a string feature.title: a string feature.wiki_context: a string feature.search_results: a dictionary feature containing:
description: a string feature.filename: a string feature.rank: a int32 feature.title: a string feature.url: a string feature.search_context: a string feature.aliases: a list of string features.normalized_aliases: a list of string features.matched_wiki_entity_name: a string feature.normalized_matched_wiki_entity_name: a string feature.normalized_value: a string feature.type: a string feature.value: a string feature.question: a string feature.question_id: a string feature.question_source: a string feature.entity_pages: a dictionary feature containing:
doc_source: a string feature.filename: a string feature.title: a string feature.wiki_context: a string feature.search_results: a dictionary feature containing:
description: a string feature.filename: a string feature.rank: a int32 feature.title: a string feature.url: a string feature.search_context: a string feature.aliases: a list of string features.normalized_aliases: a list of string features.matched_wiki_entity_name: a string feature.normalized_matched_wiki_entity_name: a string feature.normalized_value: a string feature.type: a string feature.value: a string feature.| name | train | validation | test |
|---|---|---|---|
| rc | 138384 | 18669 | 17210 |
| rc.nocontext | 138384 | 18669 | 17210 |
| unfiltered | 87622 | 11313 | 10832 |
| unfiltered.nocontext | 87622 | 11313 | 10832 |
The University of Washington does not own the copyright of the questions and documents included in TriviaQA.
@article{2017arXivtriviaqa,
author = {{Joshi}, Mandar and {Choi}, Eunsol and {Weld},
Daniel and {Zettlemoyer}, Luke},
title = "{triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension}",
journal = {arXiv e-prints},
year = 2017,
eid = {arXiv:1705.03551},
pages = {arXiv:1705.03551},
archivePrefix = {arXiv},
eprint = {1705.03551},
}
Thanks to @thomwolf, @patrickvonplaten, @lewtun for adding this dataset.
TriviaqQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaqQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions.
English.
An example of 'train' looks as follows.
An example of 'train' looks as follows.
An example of 'validation' looks as follows.
An example of 'train' looks as follows.
The data fields are the same among all splits.
question: a string feature.question_id: a string feature.question_source: a string feature.entity_pages: a dictionary feature containing:
doc_source: a string feature.filename: a string feature.title: a string feature.wiki_context: a string feature.search_results: a dictionary feature containing:
description: a string feature.filename: a string feature.rank: a int32 feature.title: a string feature.url: a string feature.search_context: a string feature.aliases: a list of string features.normalized_aliases: a list of string features.matched_wiki_entity_name: a string feature.normalized_matched_wiki_entity_name: a string feature.normalized_value: a string feature.type: a string feature.value: a string feature.question: a string feature.question_id: a string feature.question_source: a string feature.entity_pages: a dictionary feature containing:
doc_source: a string feature.filename: a string feature.title: a string feature.wiki_context: a string feature.search_results: a dictionary feature containing:
description: a string feature.filename: a string feature.rank: a int32 feature.title: a string feature.url: a string feature.search_context: a string feature.aliases: a list of string features.normalized_aliases: a list of string features.matched_wiki_entity_name: a string feature.normalized_matched_wiki_entity_name: a string feature.normalized_value: a string feature.type: a string feature.value: a string feature.question: a string feature.question_id: a string feature.question_source: a string feature.entity_pages: a dictionary feature containing:
doc_source: a string feature.filename: a string feature.title: a string feature.wiki_context: a string feature.search_results: a dictionary feature containing:
description: a string feature.filename: a string feature.rank: a int32 feature.title: a string feature.url: a string feature.search_context: a string feature.aliases: a list of string features.normalized_aliases: a list of string features.matched_wiki_entity_name: a string feature.normalized_matched_wiki_entity_name: a string feature.normalized_value: a string feature.type: a string feature.value: a string feature.question: a string feature.question_id: a string feature.question_source: a string feature.entity_pages: a dictionary feature containing:
doc_source: a string feature.filename: a string feature.title: a string feature.wiki_context: a string feature.search_results: a dictionary feature containing:
description: a string feature.filename: a string feature.rank: a int32 feature.title: a string feature.url: a string feature.search_context: a string feature.aliases: a list of string features.normalized_aliases: a list of string features.matched_wiki_entity_name: a string feature.normalized_matched_wiki_entity_name: a string feature.normalized_value: a string feature.type: a string feature.value: a string feature.| name | train | validation | test |
|---|---|---|---|
| rc | 138384 | 18669 | 17210 |
| rc.nocontext | 138384 | 18669 | 17210 |
| unfiltered | 87622 | 11313 | 10832 |
| unfiltered.nocontext | 87622 | 11313 | 10832 |
The University of Washington does not own the copyright of the questions and documents included in TriviaQA.
@article{2017arXivtriviaqa,
author = {{Joshi}, Mandar and {Choi}, Eunsol and {Weld},
Daniel and {Zettlemoyer}, Luke},
title = "{triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension}",
journal = {arXiv e-prints},
year = 2017,
eid = {arXiv:1705.03551},
pages = {arXiv:1705.03551},
archivePrefix = {arXiv},
eprint = {1705.03551},
}
Thanks to @thomwolf, @patrickvonplaten, @lewtun for adding this dataset.