13 repos
Synthetic and curated datasets for training large language models, with a focus on instruction-following, reasoning, and multilingual capabilities. The cluster centers on the Magpie project's approach to generating high-quality training data at scale across multiple model architectures (Qwen, Llama, Gemma). This area covers dataset creation methodologies, alignment techniques, and language model evaluation benchmarks for improving model reasoning and instruction adherence.