Multilingual Instruction Dataset Generation

11 repos

Methods and datasets for creating instruction-following training data across diverse languages through evolutionary algorithms. The cluster centers on the Evol-Instruct approach—a technique for generating and refining high-quality instruction-response pairs by iteratively applying transformations like constraints, deepening, and concretization. Repositories implement this methodology for Spanish, Hindi, Indonesian, German, Chinese, Korean, and other languages, enabling the development of multilingual large language models.