Writer/FailSafeQA

Dataset

Benchmark data introduced in the paper: \

10

22 commits

1 linked in READMEs

updated Feb 13, 2025

See the code

README

Benchmark data introduced in the paper:
Expect the Unexpected: FailSafeQA Long Context for Finance (https://arxiv.org/abs/2502.06329)

Dataset count: 220

{           
            "idx": int,
            "tokens": int,
            "context": string,
            "ocr_context": string,
            "answer": string,
            "query": string,
            "incomplete_query": string, 
            "out-of-domain_query": string, 
            "error_query": string, 
            "out-of-scope_query": string, 
            "citations": string,
            "citations_tokens": int
}

Tasks variety: Root verbs and their direct objects from the first sentence of each normalized query, the top 20 verbs and their top five direct object. Tasks types:

  • 83.0% question answering (QA)
  • 17.0% involving text generation (TG)

To cite FailSafeQA, please use:

@misc{kiran2024failsafeqa,
      title={Expect the Unexpected: FailSafeQA Long Context for Finance}, 
      author={Kiran Kamble and Melisa Russak and Dmytro Mozolevskyi and Muayad Ali and Mateusz Russak and Waseem AlShikh},
      year={2024},
      eprint={2502.06329},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}
finance
synthetic

Contributors

melisa

19 commits

nielsr

2 commits

wassemgtk

1 commits

Writer/FailSafeQA

Dataset

Benchmark data introduced in the paper: \

10

22 commits

1 linked in READMEs

updated Feb 13, 2025

See the code

README

Benchmark data introduced in the paper:
Expect the Unexpected: FailSafeQA Long Context for Finance (https://arxiv.org/abs/2502.06329)

Dataset count: 220

{           
            "idx": int,
            "tokens": int,
            "context": string,
            "ocr_context": string,
            "answer": string,
            "query": string,
            "incomplete_query": string, 
            "out-of-domain_query": string, 
            "error_query": string, 
            "out-of-scope_query": string, 
            "citations": string,
            "citations_tokens": int
}

Tasks variety: Root verbs and their direct objects from the first sentence of each normalized query, the top 20 verbs and their top five direct object. Tasks types:

  • 83.0% question answering (QA)
  • 17.0% involving text generation (TG)

To cite FailSafeQA, please use:

@misc{kiran2024failsafeqa,
      title={Expect the Unexpected: FailSafeQA Long Context for Finance}, 
      author={Kiran Kamble and Melisa Russak and Dmytro Mozolevskyi and Muayad Ali and Mateusz Russak and Waseem AlShikh},
      year={2024},
      eprint={2502.06329},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}
finance
synthetic

Contributors

melisa

19 commits

nielsr

2 commits

wassemgtk

1 commits