svea-klaus/Legal-Document-Summarization

3

13 commits

updated Apr 29, 2022

See the code

README

Legal-Document-Summarization

This corpus is for the paper "Summarizing Legal Regulatory Documents using Transformers" (https://doi.org/10.1145/3477495.3531872, doi will be active soon) accepted for SIGIR 2022. Please cite the paper as:

@inproceedings{10.1145/3477495.3531872, Author = {Svea Klaus, Ria Van Hecke, Kaweh Djafari Naini, Ismail Sengor Altingovde, Juan Bernabé-Moreno and Enrique Herrera-Viedma}, Title = {Summarizing Legal Regulatory Documents using Transformers}, year = {2022}, Booktitle = { {SIGIR} ’22: The 45th International {ACM} {SIGIR} Conference on Research and Development in Information Retrieval, Spain, July 11-15, 2022} url = {https://doi.org/10.1145/3477495.3531872} }

The corpus is named EUR-LexSum and contains 4595 curated European regulatory documents and their corresponding summaries passed by the European Union between July 2003 and February 2022. To this date, the vast majority of summarization research has been conducted on a limited amount of datasets, like the CNN/Daily Mail, NYT, DUC, or Gigaword. To address the growing need to summarize texts in domains other than news, we introduce a novel version of the EUR-Lex. To our knowledge, the data has not been provided for fine-tuning summarization models before. However, as aforementioned, the legal domain especially benefits from effective models because it entails a large amount of judicial terminology that lay summaries have to grasp as well. The texts are intended for a non-specialist audience and structured into 32 policy fields (such as Energy, Public Health, or Taxation). The very first section of a summary contains a hyperlink to the full and official version of the legal act. The following sections address the legal act’s target, e.g. replacing an older regulation, and it’s key points consisting of four parts: the legal act’s title along with the legal

Contributors

svea-klaus

13 commits

svea-klaus/Legal-Document-Summarization

3

13 commits

updated Apr 29, 2022

See the code

README

Legal-Document-Summarization

This corpus is for the paper "Summarizing Legal Regulatory Documents using Transformers" (https://doi.org/10.1145/3477495.3531872, doi will be active soon) accepted for SIGIR 2022. Please cite the paper as:

@inproceedings{10.1145/3477495.3531872, Author = {Svea Klaus, Ria Van Hecke, Kaweh Djafari Naini, Ismail Sengor Altingovde, Juan Bernabé-Moreno and Enrique Herrera-Viedma}, Title = {Summarizing Legal Regulatory Documents using Transformers}, year = {2022}, Booktitle = { {SIGIR} ’22: The 45th International {ACM} {SIGIR} Conference on Research and Development in Information Retrieval, Spain, July 11-15, 2022} url = {https://doi.org/10.1145/3477495.3531872} }

The corpus is named EUR-LexSum and contains 4595 curated European regulatory documents and their corresponding summaries passed by the European Union between July 2003 and February 2022. To this date, the vast majority of summarization research has been conducted on a limited amount of datasets, like the CNN/Daily Mail, NYT, DUC, or Gigaword. To address the growing need to summarize texts in domains other than news, we introduce a novel version of the EUR-Lex. To our knowledge, the data has not been provided for fine-tuning summarization models before. However, as aforementioned, the legal domain especially benefits from effective models because it entails a large amount of judicial terminology that lay summaries have to grasp as well. The texts are intended for a non-specialist audience and structured into 32 policy fields (such as Energy, Public Health, or Taxation). The very first section of a summary contains a hyperlink to the full and official version of the legal act. The following sections address the legal act’s target, e.g. replacing an older regulation, and it’s key points consisting of four parts: the legal act’s title along with the legal

Contributors

svea-klaus

13 commits