ShengbinYue/DISC-Law-SFT

Dataset

DISC-Law-SFT Dataset

179

stars

18

commits

2

linked in READMEs

May 22, 2025

updated

legal

README

DISC-Law-SFT Dataset

Legal Intelligent systems in Chinese require a combination of various abilities, including legal text understanding and generation. To achieve this, we have constructed a high-quality supervised fine-tuning dataset called DISC-Law-SFT, which covers different legal scenarios such as legal information extraction, legal judgment prediction, legal document summarization, and legal question answering. DISC-Law-SFT comprises two subsets, DISC-Law-SFT-Pair and DISC-Law-SFT-Triplet. The former aims to introduce legal reasoning abilities to the LLM, while the latter helps enhance the model's capability to utilize external legal knowledge. For more detailed information, please refer to our technical report or paper. The distribution of the dataset is:

DatasetTask/SourceSizeScenario
DISC-Law-SFT-PairLegal information extraction32KLegal professional assistant
Legal event detection27K
Legal case classification20K
Legal judgement prediction11K
Legal case matching8K
Legal text summarization9K
Judicial public opinion summarization6K
Legal question answering93KLegal consultation services
Legal reading comprehension38KJudicial examination assistant
Judicial examination12K
DISC-Law-SFT-TripleLegal judgement prediction16KLegal professional assistant
Legal question answering23KLegal consultation services
GeneralAlpaca-GPT448KGeneral scenarios
Firefly60K
Total403K

We currently open-source most of the DISC-Law-SFT Dataset.

More detail and news check our homepage !

Citation

If our project has been helpful for your research and work, please kindly cite our work as follows:

@misc{yue2023disclawllm,
    title={DISC-LawLLM: Fine-tuning Large Language Models for Intelligent Legal Services}, 
    author={Shengbin Yue and Wei Chen and Siyuan Wang and Bingxuan Li and Chenchen Shen and Shujun Liu and Yuxuan Zhou and Yao Xiao and Song Yun and Xuanjing Huang and Zhongyu Wei},
    year={2023},
    eprint={2309.11325},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

@inproceedings{yue2024lawllm,
  title={LawLLM: Intelligent Legal System with Legal Reasoning and Verifiable Retrieval},
  author={Yue, Shengbin and Liu, Shujun and Zhou, Yuxuan and Shen, Chenchen and Wang, Siyuan and Xiao, Yao and Li, Bingxuan and Song, Yun and Shen, Xiaoyu and Chen, Wei and others},
  booktitle={International Conference on Database Systems for Advanced Applications},
  pages={304--321},
  year={2024},
  organization={Springer}
}

Contributors

ShengbinYue

18 commits

ShengbinYue/DISC-Law-SFT

Dataset

DISC-Law-SFT Dataset

179

stars

18

commits

2

linked in READMEs

May 22, 2025

updated

legal

README

DISC-Law-SFT Dataset

Legal Intelligent systems in Chinese require a combination of various abilities, including legal text understanding and generation. To achieve this, we have constructed a high-quality supervised fine-tuning dataset called DISC-Law-SFT, which covers different legal scenarios such as legal information extraction, legal judgment prediction, legal document summarization, and legal question answering. DISC-Law-SFT comprises two subsets, DISC-Law-SFT-Pair and DISC-Law-SFT-Triplet. The former aims to introduce legal reasoning abilities to the LLM, while the latter helps enhance the model's capability to utilize external legal knowledge. For more detailed information, please refer to our technical report or paper. The distribution of the dataset is:

DatasetTask/SourceSizeScenario
DISC-Law-SFT-PairLegal information extraction32KLegal professional assistant
Legal event detection27K
Legal case classification20K
Legal judgement prediction11K
Legal case matching8K
Legal text summarization9K
Judicial public opinion summarization6K
Legal question answering93KLegal consultation services
Legal reading comprehension38KJudicial examination assistant
Judicial examination12K
DISC-Law-SFT-TripleLegal judgement prediction16KLegal professional assistant
Legal question answering23KLegal consultation services
GeneralAlpaca-GPT448KGeneral scenarios
Firefly60K
Total403K

We currently open-source most of the DISC-Law-SFT Dataset.

More detail and news check our homepage !

Citation

If our project has been helpful for your research and work, please kindly cite our work as follows:

@misc{yue2023disclawllm,
    title={DISC-LawLLM: Fine-tuning Large Language Models for Intelligent Legal Services}, 
    author={Shengbin Yue and Wei Chen and Siyuan Wang and Bingxuan Li and Chenchen Shen and Shujun Liu and Yuxuan Zhou and Yao Xiao and Song Yun and Xuanjing Huang and Zhongyu Wei},
    year={2023},
    eprint={2309.11325},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

@inproceedings{yue2024lawllm,
  title={LawLLM: Intelligent Legal System with Legal Reasoning and Verifiable Retrieval},
  author={Yue, Shengbin and Liu, Shujun and Zhou, Yuxuan and Shen, Chenchen and Wang, Siyuan and Xiao, Yao and Li, Bingxuan and Song, Yun and Shen, Xiaoyu and Chen, Wei and others},
  booktitle={International Conference on Database Systems for Advanced Applications},
  pages={304--321},
  year={2024},
  organization={Springer}
}

Contributors

ShengbinYue

18 commits