LLM Training Datasets & Synthetic Data

13 repos

Synthetic and curated datasets for training large language models, with a focus on instruction-following, reasoning, and multilingual capabilities. The cluster centers on the Magpie project's approach to generating high-quality training data at scale across multiple model architectures (Qwen, Llama, Gemma). This area covers dataset creation methodologies, alignment techniques, and language model evaluation benchmarks for improving model reasoning and instruction adherence.

Python · 1
llm ·899
dataset ·880
gemma ·880
llama2 ·880
llama3 ·880
alignment ·880
synthetic-data ·880
synthetic-dataset-generation ·880
nlp ·880
paper ·880