twnlp/ChineseErrorCorrector4-4B

Dataset

2

stars

13

commits

1

linked in READMEs

Jun 3, 2026

updated

cgec
chain-of-thought
csc
text-correction

README

ChineseErrorCorrect4 Data

This dataset is associated with the paper CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards.

It contains 340,000 Chain-of-Thought (CoT) reasoning samples designed for Chinese Grammatical Error Correction (CGEC) and Chinese Spelling Correction (CSC). These samples provide explicit error reasoning for diagnostic transparency, helping models internalize linguistic priors and improve edit efficiency.

Resources

Dataset Summary

The CSRP framework addresses challenges in Chinese text correction through a three-stage approach:

  1. Continual Pre-training (CPT): Using 5.9M balanced samples to internalize domain knowledge.
  2. Chain-of-Thought SFT: Utilizing this 340k dataset to provide explicit error reasoning.
  3. Reinforcement Learning (GRPO): Optimizing with Efficiency-Aware Rewards to mitigate over-correction bias.

Citation

If you use this dataset or the CSRP framework in your research, please cite:

@misc{tian2026csrpchainofthoughtreasoningchinese,
      title={CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards}, 
      author={Wei Tian and Yuhao Zhou and Man Lan},
      year={2026},
      eprint={2606.00020},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2606.00020}, 
}

Contributors

twnlp

12 commits

nielsr

1 commits

twnlp/ChineseErrorCorrector4-4B

Dataset

2

stars

13

commits

1

linked in READMEs

Jun 3, 2026

updated

cgec
chain-of-thought
csc
text-correction

README

ChineseErrorCorrect4 Data

This dataset is associated with the paper CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards.

It contains 340,000 Chain-of-Thought (CoT) reasoning samples designed for Chinese Grammatical Error Correction (CGEC) and Chinese Spelling Correction (CSC). These samples provide explicit error reasoning for diagnostic transparency, helping models internalize linguistic priors and improve edit efficiency.

Resources

Dataset Summary

The CSRP framework addresses challenges in Chinese text correction through a three-stage approach:

  1. Continual Pre-training (CPT): Using 5.9M balanced samples to internalize domain knowledge.
  2. Chain-of-Thought SFT: Utilizing this 340k dataset to provide explicit error reasoning.
  3. Reinforcement Learning (GRPO): Optimizing with Efficiency-Aware Rewards to mitigate over-correction bias.

Citation

If you use this dataset or the CSRP framework in your research, please cite:

@misc{tian2026csrpchainofthoughtreasoningchinese,
      title={CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards}, 
      author={Wei Tian and Yuhao Zhou and Man Lan},
      year={2026},
      eprint={2606.00020},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2606.00020}, 
}

Contributors

twnlp

12 commits

nielsr

1 commits