A comprehensive survey and curated collection of resources on synthetic data generation.
Synthesis from scratch
Synthesis from seeds
Synthesis from structure
Synthesis with evolution
Data-space inversion
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Reverse-Engineered Reasoning for Open-Ended Generation | 2025 | arXiv |
| Dreaming to distill: Data-free knowledge transfer via deepinversion | 2020 | CVPR |
Latent-space inversion
Agent-based simulation
Platform-based simulation
Rule-based augmentation
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Cutmix: Regularization strategy to train strong classifiers with localizable features | 2019 | ICCV |
| EDA: Easy data augmentation techniques for boosting performance on text classification tasks | 2019 | arXiv |
| mixup: Beyond empirical risk minimization | 2017 | arXiv |
Generative augmentation
Zero/Few-shot learning
Federated learning
Data-free knowledge distillation
Data-free pruning/quantization
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Sharpness-aware data generation for zero-shot quantization | 2024 | arXiv |
| Distilled Pruning: Using Synthetic Data to Win the Lottery | 2023 | arXiv |
Data-free meta-learning
Data-free continual learning
Dataset distillation
Dataset purification
General ability
Reasoning
Code
Instruction following
Preference
In-context learning
Reinforcement Learning
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Kimi K2: Open Agentic Intelligence | 2025 | arXiv |
| s1: Simple test-time scaling | 2025 | arXiv |
| Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use | 2025 | arXiv |
| Synthetic Data RL: Task Definition Is All You Need | 2025 | arXiv |
| RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold | 2024 | arXiv |
Synthetic benchmark
Privacy-preserving learning
Model inversion attack
Model stealing attack
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Unifying Multimodal Large Language Model Capabilities and Modalities via Model Merging | 2025 | arXiv |
| Data-free model extraction | 2021 | CVPR |
| Maze: Data-free model stealing attack using zeroth-order gradient estimation | 2021 | CVPR |
Adversarial defense
Machine unlearning
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Mitigating catastrophic forgetting in large language models with self-synthesized rehearsal | 2024 | arXiv |
De-bias learning
Long-tail learning
Explainable AI
Data watermarking
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Can watermarking large language models prevent copyrighted text generation and hide training data? | 2025 | AAAI |
| TimeWak: Temporal Chained-Hashing Watermark for Time Series Data | 2025 | arXiv |
Visual sensing
Force sensing
Sensor fusion
| Paper Title | Year | Conference/Journal |
|---|---|---|
| SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities | 2024 | CVPR |
| Embodiedgpt: Vision-language pre-training via embodied chain of thought | 2023 | NeurIPS |
| PaLM-E: An Embodied Multimodal Language Model | 2023 | arXiv |
| Rt-2: Vision-language-action models transfer web knowledge to robotic control | 2023 | CoRL |
Trajectory synthesis
Environment synthesis
Human behavior synthesis
Cross-embodiment training
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Droid: A large-scale in-the-wild robot manipulation dataset | 2024 | arXiv |
| Octo: An open-source generalist robot policy | 2024 | arXiv |
| Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0 | 2024 | ICRA |
| Openvla: An open-source vision-language-action model | 2024 | arXiv |
Vision-language-action models
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Diffusion forcing: Next-token prediction meets full-sequence diffusion | 2024 | NeurIPS |
| Learning universal policies via text-guided video generation | 2023 | NeurIPS |
| PaLM-E: An Embodied Multimodal Language Model | 2023 | arXiv |
| Rt-2: Vision-language-action models transfer web knowledge to robotic control | 2023 | CoRL |
Sim-to-real transfer
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Controlled training data generation with diffusion models | 2024 | arXiv |
| Llm see, llm do: Guiding data generation to target non-differentiable objectives | 2024 | arXiv |
| Paper Title | Year | Conference/Journal |
|---|---|---|
| A multi-faceted evaluation framework for assessing synthetic data generated by large language models | 2024 | arXiv |
12 commits
6 commits
A comprehensive survey and curated collection of resources on synthetic data generation.
Synthesis from scratch
Synthesis from seeds
Synthesis from structure
Synthesis with evolution
Data-space inversion
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Reverse-Engineered Reasoning for Open-Ended Generation | 2025 | arXiv |
| Dreaming to distill: Data-free knowledge transfer via deepinversion | 2020 | CVPR |
Latent-space inversion
Agent-based simulation
Platform-based simulation
Rule-based augmentation
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Cutmix: Regularization strategy to train strong classifiers with localizable features | 2019 | ICCV |
| EDA: Easy data augmentation techniques for boosting performance on text classification tasks | 2019 | arXiv |
| mixup: Beyond empirical risk minimization | 2017 | arXiv |
Generative augmentation
Zero/Few-shot learning
Federated learning
Data-free knowledge distillation
Data-free pruning/quantization
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Sharpness-aware data generation for zero-shot quantization | 2024 | arXiv |
| Distilled Pruning: Using Synthetic Data to Win the Lottery | 2023 | arXiv |
Data-free meta-learning
Data-free continual learning
Dataset distillation
Dataset purification
General ability
Reasoning
Code
Instruction following
Preference
In-context learning
Reinforcement Learning
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Kimi K2: Open Agentic Intelligence | 2025 | arXiv |
| s1: Simple test-time scaling | 2025 | arXiv |
| Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use | 2025 | arXiv |
| Synthetic Data RL: Task Definition Is All You Need | 2025 | arXiv |
| RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold | 2024 | arXiv |
Synthetic benchmark
Privacy-preserving learning
Model inversion attack
Model stealing attack
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Unifying Multimodal Large Language Model Capabilities and Modalities via Model Merging | 2025 | arXiv |
| Data-free model extraction | 2021 | CVPR |
| Maze: Data-free model stealing attack using zeroth-order gradient estimation | 2021 | CVPR |
Adversarial defense
Machine unlearning
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Mitigating catastrophic forgetting in large language models with self-synthesized rehearsal | 2024 | arXiv |
De-bias learning
Long-tail learning
Explainable AI
Data watermarking
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Can watermarking large language models prevent copyrighted text generation and hide training data? | 2025 | AAAI |
| TimeWak: Temporal Chained-Hashing Watermark for Time Series Data | 2025 | arXiv |
Visual sensing
Force sensing
Sensor fusion
| Paper Title | Year | Conference/Journal |
|---|---|---|
| SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities | 2024 | CVPR |
| Embodiedgpt: Vision-language pre-training via embodied chain of thought | 2023 | NeurIPS |
| PaLM-E: An Embodied Multimodal Language Model | 2023 | arXiv |
| Rt-2: Vision-language-action models transfer web knowledge to robotic control | 2023 | CoRL |
Trajectory synthesis
Environment synthesis
Human behavior synthesis
Cross-embodiment training
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Droid: A large-scale in-the-wild robot manipulation dataset | 2024 | arXiv |
| Octo: An open-source generalist robot policy | 2024 | arXiv |
| Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0 | 2024 | ICRA |
| Openvla: An open-source vision-language-action model | 2024 | arXiv |
Vision-language-action models
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Diffusion forcing: Next-token prediction meets full-sequence diffusion | 2024 | NeurIPS |
| Learning universal policies via text-guided video generation | 2023 | NeurIPS |
| PaLM-E: An Embodied Multimodal Language Model | 2023 | arXiv |
| Rt-2: Vision-language-action models transfer web knowledge to robotic control | 2023 | CoRL |
Sim-to-real transfer
| Paper Title | Year | Conference/Journal |
|---|---|---|
| Controlled training data generation with diffusion models | 2024 | arXiv |
| Llm see, llm do: Guiding data generation to target non-differentiable objectives | 2024 | arXiv |
| Paper Title | Year | Conference/Journal |
|---|---|---|
| A multi-faceted evaluation framework for assessing synthetic data generated by large language models | 2024 | arXiv |
12 commits
6 commits