7 repos
Datasets and systems for training machine learning models on code generation tasks, including instructional code examples, benchmark datasets for evaluating code synthesis, and tools for generating or documenting code from natural language. The cluster centers on curated collections like CONALA, MBPP, and CodeInsight that pair code with natural language descriptions, enabling research in neural code generation, code-to-text tasks, and automated programming assistants.