tokyotech-llm/JEMHopQA

Dataset

0

stars

41

commits

1

linked in READMEs

Aug 8, 2025

updated

README

JEMHopQA

このデータセットは SB Intuitions様が公開されている sbintuitions/JEMHopQA を,評価フレームワーク swallow-evaluation-instruct で用いるためにクローンしたものです.

出典

  • v1, v1.1, v1.2: aiishii/JEMHopQA on GitHub の複製.
  • v1.[1,2]-extended-answers: SB Intuitions 様が同義語や異表記の別解を追加したもの. 具体的には answer: stranswers: List[str] に変更され,オリジナルの正解および別解が answers に格納されている.

JEMHopQA

JEMHopQA (Japanese Explainable Multi-hop Question Answering) is a Japanese multi-hop QA dataset that can evaluate internal reasoning. It is a task that takes a question as input and generates an answer and derivations. Derivations are a set of derivation steps and is a semi-structured representation of relationships between entities. This dataset contains both compositional (linking information from two Wikipedia articles) and comparison (comparing information from two Wikipedia articles) questions.

Licensing Information

Creative Commons Attribution Share Alike 4.0 International

Citation Information

@inproceedings{ishii-etal-2024-jemhopqa-dataset,
    title = "{JEMH}op{QA}: Dataset for {J}apanese Explainable Multi-Hop Question Answering",
    author = "Ishii, Ai  and
      Inoue, Naoya  and
      Suzuki, Hisami  and
      Sekine, Satoshi",
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.831",
    pages = "9515--9525",
}

Subsets

v1

v1: JEMHopQA/corpus on GitHub

v1.1

v1.1: JEMHopQA/corpus_ver1.1 on GitHub

  • qid (str): Unique identifier for each entry in the dataset.
  • type (str): The category of the question ("comparison" or "compositional").
  • question (str): The text of the question.
  • answer (str): The correct answer to the question.
  • derivations (dict[str, list[str]]): Knowledge triples for reasoning used to arrive at the answer.
  • page_ids (list[str]): Identifiers for related Wikipedia pages.
  • time_dependent (bool): Indicates whether the question/answer is time-sensitive.

v1.1-extended-answers

  • v1.1 の answer に別解を加え、answers (list[str]) に拡張したもの
    • e.g., "カリフォルニア州クパチーノ" -> ["カリフォルニア州クパチーノ", "アメリカ合衆国カリフォルニア州クパチーノ", "アメリカ合衆国カリフォルニア州クパティーノ", "カリフォルニア州クパティーノ"]
  • split: validation のみ
  • question と answrers は (未 NFKC 正規化)

Contributors

teruo6939

21 commits

ryo0634

16 commits

s-mizuki-nlp

4 commits

tokyotech-llm/JEMHopQA

Dataset

0

stars

41

commits

1

linked in READMEs

Aug 8, 2025

updated

README

JEMHopQA

このデータセットは SB Intuitions様が公開されている sbintuitions/JEMHopQA を,評価フレームワーク swallow-evaluation-instruct で用いるためにクローンしたものです.

出典

  • v1, v1.1, v1.2: aiishii/JEMHopQA on GitHub の複製.
  • v1.[1,2]-extended-answers: SB Intuitions 様が同義語や異表記の別解を追加したもの. 具体的には answer: stranswers: List[str] に変更され,オリジナルの正解および別解が answers に格納されている.

JEMHopQA

JEMHopQA (Japanese Explainable Multi-hop Question Answering) is a Japanese multi-hop QA dataset that can evaluate internal reasoning. It is a task that takes a question as input and generates an answer and derivations. Derivations are a set of derivation steps and is a semi-structured representation of relationships between entities. This dataset contains both compositional (linking information from two Wikipedia articles) and comparison (comparing information from two Wikipedia articles) questions.

Licensing Information

Creative Commons Attribution Share Alike 4.0 International

Citation Information

@inproceedings{ishii-etal-2024-jemhopqa-dataset,
    title = "{JEMH}op{QA}: Dataset for {J}apanese Explainable Multi-Hop Question Answering",
    author = "Ishii, Ai  and
      Inoue, Naoya  and
      Suzuki, Hisami  and
      Sekine, Satoshi",
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.831",
    pages = "9515--9525",
}

Subsets

v1

v1: JEMHopQA/corpus on GitHub

v1.1

v1.1: JEMHopQA/corpus_ver1.1 on GitHub

  • qid (str): Unique identifier for each entry in the dataset.
  • type (str): The category of the question ("comparison" or "compositional").
  • question (str): The text of the question.
  • answer (str): The correct answer to the question.
  • derivations (dict[str, list[str]]): Knowledge triples for reasoning used to arrive at the answer.
  • page_ids (list[str]): Identifiers for related Wikipedia pages.
  • time_dependent (bool): Indicates whether the question/answer is time-sensitive.

v1.1-extended-answers

  • v1.1 の answer に別解を加え、answers (list[str]) に拡張したもの
    • e.g., "カリフォルニア州クパチーノ" -> ["カリフォルニア州クパチーノ", "アメリカ合衆国カリフォルニア州クパチーノ", "アメリカ合衆国カリフォルニア州クパティーノ", "カリフォルニア州クパティーノ"]
  • split: validation のみ
  • question と answrers は (未 NFKC 正規化)

Contributors

teruo6939

21 commits

ryo0634

16 commits

s-mizuki-nlp

4 commits