Egg-Hu/Awesome-Synthetic-Data-Generation

23

18 commits

updated Jan 7, 2026

See the code

README

Awesome Synthetic Data Generation

PRs Welcome Stars Forks

A comprehensive survey and curated collection of resources on synthetic data generation.


📚 Table of Contents (Click to Expand)

1. Methodologies

1.1 Generation-Based Synthesis

Synthesis from scratch

Synthesis from seeds

Synthesis from structure

Synthesis with evolution

1.2 Inversion-Based Synthesis

Data-space inversion

Latent-space inversion

1.3 Simulation-Based Synthesis

Agent-based simulation

Platform-based simulation

1.4 Augmentation-Based Synthesis

Rule-based augmentation

Generative augmentation


2. Applications

2.1 Data-centric AI

Data Accessibility

Zero/Few-shot learning

Federated learning

Data-free knowledge distillation

Data-free pruning/quantization

Data-free meta-learning

Data-free continual learning

Data Refinement

Dataset distillation

Dataset purification


2.2 Model-centric AI

General Model Enhancement

General ability

Domain Model Enhancement

Reasoning

Paper TitleYearConference/Journal
Absolute zero: Reinforced self-play reasoning with zero data2025arXiv
HS-STAR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation2025arXiv
Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis for Large Reasoning Models2025arXiv
Logictree: Improving complex reasoning of LLMs via instantiated multi-step synthetic logical data2025NeurIPS
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning2025arXiv
Seed-Coder: Let the Code Model Curate Data for Itself2025arXiv
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment2025ICLR
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models2025arXiv
Thinking LLMs: General Instruction Following with Thought Generation2025ICML
Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch2025arXiv
A Graph-Based Synthetic Data Pipeline for Scaling High-Quality Reasoning Instructions2024arXiv
Aligning teacher with student preferences for tailored training data generation2024arXiv
Autocoder: Enhancing code large language model with$\backslash$textsc $$AIEV-Instruct$$2024arXiv
Boosting reward model with preference-conditional multi-aspect synthetic data generation2024arXiv
ControlMath: Controllable Data Generation Promotes Math Generalist Models2024EMNLP
From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis2024EMNLP
HexaCoder: Secure Code Generation via Oracle-Guided Synthetic Training Data2024arXiv
Infinitymath: A scalable instruction tuning dataset in programmatic mathematical reasoning2024CIKM
Jiuzhang3. 0: Efficiently improving mathematical reasoning by training small data synthesis models2024NeurIPS
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning2024ICLR
Marco-o1: Towards open reasoning models for open-ended solutions2024arXiv
MathScale: Scaling Instruction Tuning for Mathematical Reasoning2024ICML
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models2024ICLR
Openmathinstruct-1: A 1.8 million math instruction tuning dataset2024NeurIPS
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking2024COLM
Refined direct preference optimization with synthetic data for behavioral alignment of llms2024arXiv
Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold2024NeurIPS
Self-Consistency Preference Optimization2024arXiv
Self-play with execution feedback: Improving instruction-following capabilities of large language models2024arXiv
Self-Rewarding Language Models2024arXiv
Small Language Models Need Strong Verifiers to Self-Correct Reasoning2024ACL
Strengthening multimodal large language model with bootstrapped preference optimization2024ECCV
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving2024ICLR
Tree-instruct: A preliminary study of the intrinsic relationship between complexity and alignment2024COLING
WizardLM: Empowering large pre-trained language models to follow complex instructions2024ICLR
Lamini-lm: A diverse herd of distilled models from large-scale instructions2023arXiv
Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning2023arXiv
Reflection-tuning: Recycling data for better instruction-tuning2023NeurIPS
What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning2023arXiv
Wizardcoder: Empowering code large language models with evol-instruct2023arXiv
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct2023arXiv
Star: Bootstrapping reasoning with reasoning2022NeurIPS

Code

Instruction following

Preference

In-context learning

Reinforcement Learning

Model Evaluation

Synthetic benchmark


2.3 Trustworthy AI

Privacy

Privacy-preserving learning

Model inversion attack

Safety & Security

Model stealing attack

Adversarial defense

Machine unlearning

Fairness

De-bias learning

Long-tail learning

Interpretability

Explainable AI

Governance

Data watermarking


2.4 Embodied AI

Perception

Visual sensing

Force sensing

Sensor fusion

Interaction

Trajectory synthesis

Environment synthesis

Human behavior synthesis

Generalization

Cross-embodiment training

Vision-language-action models

Sim-to-real transfer


3. Challenges & Future Directions

Model Collapse

Utility-Privacy Tradeoffs

Generation-Evaluation Bias

Active Data Synthesis

Synthetic Data Evaluation

Multi-Modal Data Synthesis


↑ Back to Top ↑

Contributors

Egg-Hu

6 commits

Egg-Hu/Awesome-Synthetic-Data-Generation

23

18 commits

updated Jan 7, 2026

See the code

README

Awesome Synthetic Data Generation

PRs Welcome Stars Forks

A comprehensive survey and curated collection of resources on synthetic data generation.


📚 Table of Contents (Click to Expand)

1. Methodologies

1.1 Generation-Based Synthesis

Synthesis from scratch

Synthesis from seeds

Synthesis from structure

Synthesis with evolution

1.2 Inversion-Based Synthesis

Data-space inversion

Latent-space inversion

1.3 Simulation-Based Synthesis

Agent-based simulation

Platform-based simulation

1.4 Augmentation-Based Synthesis

Rule-based augmentation

Generative augmentation


2. Applications

2.1 Data-centric AI

Data Accessibility

Zero/Few-shot learning

Federated learning

Data-free knowledge distillation

Data-free pruning/quantization

Data-free meta-learning

Data-free continual learning

Data Refinement

Dataset distillation

Dataset purification


2.2 Model-centric AI

General Model Enhancement

General ability

Domain Model Enhancement

Reasoning

Paper TitleYearConference/Journal
Absolute zero: Reinforced self-play reasoning with zero data2025arXiv
HS-STAR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation2025arXiv
Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis for Large Reasoning Models2025arXiv
Logictree: Improving complex reasoning of LLMs via instantiated multi-step synthetic logical data2025NeurIPS
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning2025arXiv
Seed-Coder: Let the Code Model Curate Data for Itself2025arXiv
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment2025ICLR
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models2025arXiv
Thinking LLMs: General Instruction Following with Thought Generation2025ICML
Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch2025arXiv
A Graph-Based Synthetic Data Pipeline for Scaling High-Quality Reasoning Instructions2024arXiv
Aligning teacher with student preferences for tailored training data generation2024arXiv
Autocoder: Enhancing code large language model with$\backslash$textsc $$AIEV-Instruct$$2024arXiv
Boosting reward model with preference-conditional multi-aspect synthetic data generation2024arXiv
ControlMath: Controllable Data Generation Promotes Math Generalist Models2024EMNLP
From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis2024EMNLP
HexaCoder: Secure Code Generation via Oracle-Guided Synthetic Training Data2024arXiv
Infinitymath: A scalable instruction tuning dataset in programmatic mathematical reasoning2024CIKM
Jiuzhang3. 0: Efficiently improving mathematical reasoning by training small data synthesis models2024NeurIPS
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning2024ICLR
Marco-o1: Towards open reasoning models for open-ended solutions2024arXiv
MathScale: Scaling Instruction Tuning for Mathematical Reasoning2024ICML
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models2024ICLR
Openmathinstruct-1: A 1.8 million math instruction tuning dataset2024NeurIPS
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking2024COLM
Refined direct preference optimization with synthetic data for behavioral alignment of llms2024arXiv
Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold2024NeurIPS
Self-Consistency Preference Optimization2024arXiv
Self-play with execution feedback: Improving instruction-following capabilities of large language models2024arXiv
Self-Rewarding Language Models2024arXiv
Small Language Models Need Strong Verifiers to Self-Correct Reasoning2024ACL
Strengthening multimodal large language model with bootstrapped preference optimization2024ECCV
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving2024ICLR
Tree-instruct: A preliminary study of the intrinsic relationship between complexity and alignment2024COLING
WizardLM: Empowering large pre-trained language models to follow complex instructions2024ICLR
Lamini-lm: A diverse herd of distilled models from large-scale instructions2023arXiv
Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning2023arXiv
Reflection-tuning: Recycling data for better instruction-tuning2023NeurIPS
What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning2023arXiv
Wizardcoder: Empowering code large language models with evol-instruct2023arXiv
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct2023arXiv
Star: Bootstrapping reasoning with reasoning2022NeurIPS

Code

Instruction following

Preference

In-context learning

Reinforcement Learning

Model Evaluation

Synthetic benchmark


2.3 Trustworthy AI

Privacy

Privacy-preserving learning

Model inversion attack

Safety & Security

Model stealing attack

Adversarial defense

Machine unlearning

Fairness

De-bias learning

Long-tail learning

Interpretability

Explainable AI

Governance

Data watermarking


2.4 Embodied AI

Perception

Visual sensing

Force sensing

Sensor fusion

Interaction

Trajectory synthesis

Environment synthesis

Human behavior synthesis

Generalization

Cross-embodiment training

Vision-language-action models

Sim-to-real transfer


3. Challenges & Future Directions

Model Collapse

Utility-Privacy Tradeoffs

Generation-Evaluation Bias

Active Data Synthesis

Synthetic Data Evaluation

Multi-Modal Data Synthesis


↑ Back to Top ↑

Contributors

Egg-Hu

6 commits