A curated and automatically updated collection of LLM data research, covering data substrates, data creation & selection, and data ingestion strategies across the full LLM lifecycle.
See the codeA curated, operation-centered map of LLM data across the full lifecycle — from data substrates, to data creation and selection, to training- and inference-time information use.
Paper · Taxonomy · Auto Tracking · Resource List · Cross-Tier Synthesis · Contribute
LLM development is often organized by training stage (pretraining, SFT, RL) or by isolated techniques (synthetic data, selection, RAG, agents). This repository instead organizes the literature around the operation performed on information.
The accompanying survey asks three operational questions:
This perspective makes it easier to place methods that cross conventional boundaries—for example, model-aware selection, data mixing, rejection-sampling loops, executable agent environments, retrieval, test-time search, and verification.
| Area | Coverage |
|---|---|
| Data Substrates | Pretraining corpora, instruction/preference data, reasoning and code resources, safety data, agent trajectories, benchmarks, executable environments |
| Data Creation & Selection | Annotation, processing, prompting, distillation, synthetic data, back-translation, human–AI collaboration, symbolic generation, diversity/quality/model-aware selection |
| Data Ingestion Strategies | Packing, contextual conditioning, mixture weighting, curriculum, annealing, replay, multi-stage training, RAG, decoding, search, verification |
| Agentic Data | Tool-use data, multi-turn trajectories, environment synthesis, executable benchmarks, feedback-driven data loops |
| Cross-Tier Synthesis | Model-conditional data utility, human specification/verification, executable data-producing systems, runtime information acquisition |
If this repository is useful for your research, consider starring it. It helps other researchers discover the resource.
flowchart LR
A["Tier 1 · Data Substrates<br/>What information exists?"]
B["Tier 2 · Data Creation & Selection<br/>What supervision is created or changed?"]
C["Tier 3 · Data Ingestion Strategies<br/>How and when is information consumed?"]
M["Target Model<br/>Capability State"]
A --> B
A --> C
B --> M
C --> M
M -. "errors · loss · confidence · interaction feedback" .-> B
M -. "adaptive scheduling / retrieval" .-> C
Placement rule.
Tier 2 changes the supervision or data artifacts that exist; Tier 3 changes how available supervision or information is delivered to and consumed by the model.
The unit of analysis is the operation rather than the paper as a whole. A method can therefore span multiple tiers. For example, rejection-sampling fine-tuning may generate and filter trajectories in Tier 2, then immediately consume the retained trajectories during optimization in Tier 3.
| If you are interested in... | Recommended entry point |
|---|---|
| Building or auditing a pretraining corpus | Data Substrates → Pretrain, Data Processing, Data Selection, Sample Mixing |
| Designing instruction-tuning / alignment data | SFT, Synthesis, Selection, Capability Alignment |
| Improving reasoning / code data | Reasoning and Code, Symbolic Generation, Model-Aware Filtering |
| Building LLM agents / tool-use data | Agent and Tool Use, Agentic Data Pipelines |
| Studying curriculum, mixing, replay, or mid-training | Data Ingestion Strategies, Training Pipeline Optimization |
| Studying RAG / search / verification / test-time scaling | Truthfulness & Consistency, Test-Time Strategy |
| Looking for the survey's main cross-paper insights | Cross-Tier Synthesis |
To keep this repository useful beyond the survey's fixed literature cut-off, we provide an automated literature-tracking workflow built with GitHub Actions. The bot continuously discovers recent work related to the LLM data lifecycle, filters duplicates, assigns candidates to the operation-centered taxonomy, and proposes updates through pull requests for human review.
The automation is designed to assist curation rather than replace it: newly discovered papers are never auto-merged into the curated list.
Recent arXiv papers
↓
Topic-specific query families
↓
Date / category / LLM-data relevance filtering
↓
README + historical-database deduplication
↓
Operation-centered taxonomy classification
↓
Primary tier + optional cross-tier role
↓
Generated Markdown update
↓
GitHub Pull Request
↓
Human review → merge
The classifier follows the same placement principle as the survey:
Because the unit of analysis is the operation rather than the paper as a whole, the bot can also flag cross-tier papers. For example, a method may generate or filter supervision in Tier 2 and immediately consume that supervision during optimization in Tier 3.
| Mode | Purpose | Output |
|---|---|---|
| Weekly literature update | Discover newly released papers on a rolling window | Updates the bot-maintained recent-paper block and opens a PR |
| Historical backfill | Scan an exact date range, e.g. 2026-07-01 → 2026-08-17 | Generates an independent Markdown report under updates/ and opens a PR |
The historical mode is useful when the repository has not been updated for a period of time or when a new query family is introduced and earlier papers need to be recovered.
The implementation lives in:
.github/workflows/
├── update-papers.yml # scheduled / manual weekly update
└── backfill-papers.yml # exact-date historical scan
scripts/
└── paper_bot.py # fetch → filter → deduplicate → classify → render
config/
└── paper_bot.yaml # search queries, taxonomy keywords, thresholds
data/
└── papers.json # persistent bot-curated paper history
The automatically maintained README block is bounded by:
<!-- AUTO-LITERATURE:START -->
## Recent Automatically Discovered Papers
> This section is generated by the weekly literature bot and reviewed through pull requests before merge.
> Last generated: **2026-08-24** · Showing up to **30** recent papers.
### Data Substrates
- **[Benchmarking Patent Drafting from Inventor-Style Disclosures](http://arxiv.org/abs/2608.21249v1)** — Lekang Jiang et al. (2026). *Other*
- **[Enhancing LLMs in Predictive Political QA with Semi-Structured Data](http://arxiv.org/abs/2608.21218v1)** — Yinan Liu et al. (2026). *Other*
- **[Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment](http://arxiv.org/abs/2608.21057v1)** — Emma Granqvist et al. (2026). *Other*
- **[MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos](http://arxiv.org/abs/2608.20984v1)** — Fatima Haouari et al. (2026). *Other*
- **[VortexChat: An agentic framework for autonomous multi-objective integrated photonic design](http://arxiv.org/abs/2608.20688v1)** — Faqian Chong et al. (2026). *Other*
- **[ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection](http://arxiv.org/abs/2608.20637v1)** — Chunyi Wang et al. (2026). *Other*
- **[Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents](http://arxiv.org/abs/2608.20631v1)** — Quang Dao et al. (2026). *Other*
- **[ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models](http://arxiv.org/abs/2608.20338v1)** — Sahil Kale et al. (2026). *Other*
- **[From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation](http://arxiv.org/abs/2608.20195v1)** — Zhijun Gao et al. (2026). *Other*
- **[OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models](http://arxiv.org/abs/2608.20106v1)** — Nikita Khudov (2026). *Other*
### Data Creation and Selection
- **[Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models](http://arxiv.org/abs/2608.21019v1)** — Zhen Yang et al. (2026). *Data Selection*
- **[An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction](http://arxiv.org/abs/2608.20320v1)** — Narges Ahmadi et al. (2026). *Annotation / Processing*
### Data Ingestion Strategies
- **[Beyond Fault Localization: A Trajectory-Level Study of LLM Agents for Microservice Root Cause Analysis](http://arxiv.org/abs/2608.21310v1)** — Qisheng Lu et al. (2026). *Runtime Retrieval / Search / Verification*
- **[EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering](http://arxiv.org/abs/2608.21252v1)** — Xuanyu Meng et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration](http://arxiv.org/abs/2608.21208v1)** — Oleg Grynets et al. (2026). *Runtime Retrieval / Search / Verification*
- **[ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models](http://arxiv.org/abs/2608.21100v1)** — Wenzheng Jiang et al. (2026). *Runtime Retrieval / Search / Verification* · _Cross-tier: Data Substrates_
- **[Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems](http://arxiv.org/abs/2608.21095v1)** — Balkrishna Giri et al. (2026). *Runtime Retrieval / Search / Verification*
- **[$Z^2$-ACT: End-to-End Verifiable Agentic Intent Control for Open 6G RAN](http://arxiv.org/abs/2608.21049v1)** — Sunder Ali Khowaja et al. (2026). *Runtime Retrieval / Search / Verification*
- **[UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists](http://arxiv.org/abs/2608.20918v1)** — Ye Chen et al. (2026). *Curriculum / Mid-Training / Replay*
- **[TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding](http://arxiv.org/abs/2608.20844v1)** — Rohan Kumar et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation](http://arxiv.org/abs/2608.20794v1)** — Haodong Chen et al. (2026). *Curriculum / Mid-Training / Replay*
- **[Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation](http://arxiv.org/abs/2608.20756v1)** — Rujin Liang et al. (2026). *Runtime Retrieval / Search / Verification*
- **[AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification](http://arxiv.org/abs/2608.20711v1)** — Ji Liu et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Towards Faithful Simulation of Human Shopping Behavior](http://arxiv.org/abs/2608.20707v1)** — Jiakai Tang et al. (2026). *Runtime Retrieval / Search / Verification*
- **[JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification](http://arxiv.org/abs/2608.20607v1)** — Tianxin Zhou et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Terminal Agents: A Survey of AI Agents in Command-Line Environments](http://arxiv.org/abs/2608.20485v1)** — Yi Bin et al. (2026). *Runtime Retrieval / Search / Verification*
- **[MidTool: Mid-training Data Synthesis for Agentic Tool Use](http://arxiv.org/abs/2608.20314v1)** — Fengqing Jiang et al. (2026). *Curriculum / Mid-Training / Replay* · _Cross-tier: Data Creation and Selection_
- **[Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization](http://arxiv.org/abs/2608.20281v1)** — Qian Kou et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Decoding silent reading from non-invasive EEG](http://arxiv.org/abs/2608.20186v1)** — Ingo Marquardt et al. (2026). *Runtime Retrieval / Search / Verification*
- **[FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models](http://arxiv.org/abs/2608.20153v1)** — Dingzirui Wang et al. (2026). *Runtime Retrieval / Search / Verification*
<!-- AUTO-LITERATURE:END -->
so the workflow never rewrites the manually curated core of this repository.
Automation is intentionally conservative:
This makes the repository a living companion to the survey while preserving the quality and interpretability of the manually curated taxonomy.
Recent additions reflected in the survey/repository include:
Data-centric Artificial Intelligence: A Survey ACM Computing Surveys 2023
An Empirical Survey of Data Augmentation for Limited Data Learning in NLP TACL 2023.
Best Practices and Lessons Learned on Synthetic Data for Language Models Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zhang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, Andrew M. Dai. COLM 2024.
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, Haobo Wang. arXiv 2024.
Large Language Models for Data Annotation: A Survey Zhen Tan, Dawei Li, Song Wang, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng, Huan Liu. arXiv 2024.
A Survey on Data Synthesis and Augmentation for Large Language Models arXiv 2024.
Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application ACM Transactions on Intelligent Systems and Technology 2024.
Data Augmentation using LLMs: Data Perspectives, Learning Paradigms and Challenges arXiv 2024.
A Survey of Multimodal Large Language Model from A Data-centric Perspective arXiv 2024.
A Survey of LLM × DATA arXiv 2025.
AI Alignment: A Comprehensive Survey arXiv 2025.
A Survey of Post-Training Scaling in Large Language Models ACL 2025.
Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives arXiv 2025.
Agentic Tool Use in Large Language Models. arXiv 2026.
The LLM Data Auditor: A Survey on Data-Centric Metrics for Cross-Modal Large Language Models. arXiv 2026.
We organize the LLM data lifecycle through an operation-centered, three-tier taxonomy. The goal is not merely to separate training stages, but to distinguish what information-bearing objects exist, how their content or membership is changed, and how available information is presented to and consumed by the model.
| Tier | Operational question | Scope | Representative examples |
|---|---|---|---|
| Tier 1: Data Substrates | What information-bearing artifacts or environments are available? | Corpora, instruction/preference data, benchmarks, reasoning/code resources, trajectories, executable environments | Pretraining corpora, SFT/RL data, math/code benchmarks, tool-use trajectories, agent environments |
| Tier 2: Data Creation and Selection | What supervision exists, and how is it created or modified? | Annotation, synthesis, transformation, filtering, diversity/quality/model-aware selection | Distillation, self-instruct, back-translation, symbolic generation, data filtering, influential-data selection |
| Tier 3: Data Ingestion Strategies | How, when, how often, and under what context is information consumed? | Packing, contextual conditioning, mixture weighting, curriculum, annealing, replay, retrieval, search, decoding, verification | Sample construction, data mixing, mid-training, multi-stage curricula, RAG, test-time search |
Placement rule. Tier 2 changes the supervision or data artifacts that exist; Tier 3 changes how available supervision or information is delivered to and consumed by the model. For example, offline data selection is primarily Tier 2, whereas dynamic data mixing is Tier 3. An agent trajectory is a Tier-1 substrate; generating or filtering trajectories is Tier 2; using them in iterative optimization or test-time interaction is Tier 3.
Cross-tier methods. The unit of analysis is the operation rather than the paper as a whole. Methods such as rejection-sampling fine-tuning or error-driven preference optimization can span multiple tiers: they may generate/filter supervision in Tier 2 and immediately consume the retained supervision during optimization in Tier 3.
The substrate view highlights a shift from raw scale toward curated information density, and—especially for agents—from stored text/trajectories toward executable environments that can generate tasks, observations, failures, recovery paths, and verification signals on demand.
Across creation and selection methods, the survey highlights three recurring transitions: teacher distillation → autonomous/self-play generation, static quality filtering → model-conditional utility, and stored examples → verifier-grounded or environment-generated supervision. Selection and synthesis increasingly interact through closed loops rather than remaining independent preprocessing stages.
Cross-tier note: this section focuses on the ingestion/optimization-time role of sample construction. Methods that generate persistent new supervision (e.g., rejection sampling or targeted error injection) also have a Tier-2 creation component.
Data ingestion spans both training-time scheduling and runtime information use. Training-time methods determine how supervision is organized, weighted, ordered, replayed, or conditionally presented; runtime methods determine when self-generated or external evidence is retrieved, searched, decoded, or verified.
The lifecycle view exposes several recurring transitions that are less visible when datasets, synthesis, selection, training schedules, and agent systems are studied separately:
These trends motivate treating data engineering as the joint design of artifacts, transformations, and adaptive information-use policies.
If you use this repository or the taxonomy in your research, please cite the accompanying survey:
@article{rao2025datacentric,
title = {A Data-Centric Perspective on the Lifecycle of Large Language Models},
author = {Rao, Jun and Liu, Xuebo and Yan, Haotian and Shen, Junjie and Mo, Haosi and
Dong, Yanghaopeng and Yan, Zihao and Wang, Ziyi and Lin, Zepeng and Meng, Xiaojun and
Yu, Zixiong and Deng, Liqun and Wei, Jiansheng and Wang, Yunhe and Zhang, Min},
journal = {TechRxiv},
year = {2025},
doi = {10.36227/techrxiv.176620610.03288677/v1},
url = {https://doi.org/10.36227/techrxiv.176620610.03288677/v1}
}
For GitHub-native citation support, we also recommend adding a CITATION.cff file to the repository root.
Contributions are welcome. Please open an issue or pull request to add a paper, dataset, benchmark, correction, or missing link.
To keep the list consistent:
Title — venue/year — link.Suggested pull-request title:
[Add] <Paper / Dataset / Benchmark Name> → <Section>
We thank the authors of the papers, datasets, benchmarks, and open-source resources collected here, as well as community contributors who help keep the list accurate and up to date.
Data is not only what a model trains on — it is also what gets created, selected, scheduled, retrieved, verified, and regenerated throughout the model lifecycle.
Python
100.0%
A curated and automatically updated collection of LLM data research, covering data substrates, data creation & selection, and data ingestion strategies across the full LLM lifecycle.
See the codeA curated, operation-centered map of LLM data across the full lifecycle — from data substrates, to data creation and selection, to training- and inference-time information use.
Paper · Taxonomy · Auto Tracking · Resource List · Cross-Tier Synthesis · Contribute
LLM development is often organized by training stage (pretraining, SFT, RL) or by isolated techniques (synthetic data, selection, RAG, agents). This repository instead organizes the literature around the operation performed on information.
The accompanying survey asks three operational questions:
This perspective makes it easier to place methods that cross conventional boundaries—for example, model-aware selection, data mixing, rejection-sampling loops, executable agent environments, retrieval, test-time search, and verification.
| Area | Coverage |
|---|---|
| Data Substrates | Pretraining corpora, instruction/preference data, reasoning and code resources, safety data, agent trajectories, benchmarks, executable environments |
| Data Creation & Selection | Annotation, processing, prompting, distillation, synthetic data, back-translation, human–AI collaboration, symbolic generation, diversity/quality/model-aware selection |
| Data Ingestion Strategies | Packing, contextual conditioning, mixture weighting, curriculum, annealing, replay, multi-stage training, RAG, decoding, search, verification |
| Agentic Data | Tool-use data, multi-turn trajectories, environment synthesis, executable benchmarks, feedback-driven data loops |
| Cross-Tier Synthesis | Model-conditional data utility, human specification/verification, executable data-producing systems, runtime information acquisition |
If this repository is useful for your research, consider starring it. It helps other researchers discover the resource.
flowchart LR
A["Tier 1 · Data Substrates<br/>What information exists?"]
B["Tier 2 · Data Creation & Selection<br/>What supervision is created or changed?"]
C["Tier 3 · Data Ingestion Strategies<br/>How and when is information consumed?"]
M["Target Model<br/>Capability State"]
A --> B
A --> C
B --> M
C --> M
M -. "errors · loss · confidence · interaction feedback" .-> B
M -. "adaptive scheduling / retrieval" .-> C
Placement rule.
Tier 2 changes the supervision or data artifacts that exist; Tier 3 changes how available supervision or information is delivered to and consumed by the model.
The unit of analysis is the operation rather than the paper as a whole. A method can therefore span multiple tiers. For example, rejection-sampling fine-tuning may generate and filter trajectories in Tier 2, then immediately consume the retained trajectories during optimization in Tier 3.
| If you are interested in... | Recommended entry point |
|---|---|
| Building or auditing a pretraining corpus | Data Substrates → Pretrain, Data Processing, Data Selection, Sample Mixing |
| Designing instruction-tuning / alignment data | SFT, Synthesis, Selection, Capability Alignment |
| Improving reasoning / code data | Reasoning and Code, Symbolic Generation, Model-Aware Filtering |
| Building LLM agents / tool-use data | Agent and Tool Use, Agentic Data Pipelines |
| Studying curriculum, mixing, replay, or mid-training | Data Ingestion Strategies, Training Pipeline Optimization |
| Studying RAG / search / verification / test-time scaling | Truthfulness & Consistency, Test-Time Strategy |
| Looking for the survey's main cross-paper insights | Cross-Tier Synthesis |
To keep this repository useful beyond the survey's fixed literature cut-off, we provide an automated literature-tracking workflow built with GitHub Actions. The bot continuously discovers recent work related to the LLM data lifecycle, filters duplicates, assigns candidates to the operation-centered taxonomy, and proposes updates through pull requests for human review.
The automation is designed to assist curation rather than replace it: newly discovered papers are never auto-merged into the curated list.
Recent arXiv papers
↓
Topic-specific query families
↓
Date / category / LLM-data relevance filtering
↓
README + historical-database deduplication
↓
Operation-centered taxonomy classification
↓
Primary tier + optional cross-tier role
↓
Generated Markdown update
↓
GitHub Pull Request
↓
Human review → merge
The classifier follows the same placement principle as the survey:
Because the unit of analysis is the operation rather than the paper as a whole, the bot can also flag cross-tier papers. For example, a method may generate or filter supervision in Tier 2 and immediately consume that supervision during optimization in Tier 3.
| Mode | Purpose | Output |
|---|---|---|
| Weekly literature update | Discover newly released papers on a rolling window | Updates the bot-maintained recent-paper block and opens a PR |
| Historical backfill | Scan an exact date range, e.g. 2026-07-01 → 2026-08-17 | Generates an independent Markdown report under updates/ and opens a PR |
The historical mode is useful when the repository has not been updated for a period of time or when a new query family is introduced and earlier papers need to be recovered.
The implementation lives in:
.github/workflows/
├── update-papers.yml # scheduled / manual weekly update
└── backfill-papers.yml # exact-date historical scan
scripts/
└── paper_bot.py # fetch → filter → deduplicate → classify → render
config/
└── paper_bot.yaml # search queries, taxonomy keywords, thresholds
data/
└── papers.json # persistent bot-curated paper history
The automatically maintained README block is bounded by:
<!-- AUTO-LITERATURE:START -->
## Recent Automatically Discovered Papers
> This section is generated by the weekly literature bot and reviewed through pull requests before merge.
> Last generated: **2026-08-24** · Showing up to **30** recent papers.
### Data Substrates
- **[Benchmarking Patent Drafting from Inventor-Style Disclosures](http://arxiv.org/abs/2608.21249v1)** — Lekang Jiang et al. (2026). *Other*
- **[Enhancing LLMs in Predictive Political QA with Semi-Structured Data](http://arxiv.org/abs/2608.21218v1)** — Yinan Liu et al. (2026). *Other*
- **[Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment](http://arxiv.org/abs/2608.21057v1)** — Emma Granqvist et al. (2026). *Other*
- **[MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos](http://arxiv.org/abs/2608.20984v1)** — Fatima Haouari et al. (2026). *Other*
- **[VortexChat: An agentic framework for autonomous multi-objective integrated photonic design](http://arxiv.org/abs/2608.20688v1)** — Faqian Chong et al. (2026). *Other*
- **[ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection](http://arxiv.org/abs/2608.20637v1)** — Chunyi Wang et al. (2026). *Other*
- **[Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents](http://arxiv.org/abs/2608.20631v1)** — Quang Dao et al. (2026). *Other*
- **[ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models](http://arxiv.org/abs/2608.20338v1)** — Sahil Kale et al. (2026). *Other*
- **[From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation](http://arxiv.org/abs/2608.20195v1)** — Zhijun Gao et al. (2026). *Other*
- **[OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models](http://arxiv.org/abs/2608.20106v1)** — Nikita Khudov (2026). *Other*
### Data Creation and Selection
- **[Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models](http://arxiv.org/abs/2608.21019v1)** — Zhen Yang et al. (2026). *Data Selection*
- **[An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction](http://arxiv.org/abs/2608.20320v1)** — Narges Ahmadi et al. (2026). *Annotation / Processing*
### Data Ingestion Strategies
- **[Beyond Fault Localization: A Trajectory-Level Study of LLM Agents for Microservice Root Cause Analysis](http://arxiv.org/abs/2608.21310v1)** — Qisheng Lu et al. (2026). *Runtime Retrieval / Search / Verification*
- **[EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering](http://arxiv.org/abs/2608.21252v1)** — Xuanyu Meng et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration](http://arxiv.org/abs/2608.21208v1)** — Oleg Grynets et al. (2026). *Runtime Retrieval / Search / Verification*
- **[ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models](http://arxiv.org/abs/2608.21100v1)** — Wenzheng Jiang et al. (2026). *Runtime Retrieval / Search / Verification* · _Cross-tier: Data Substrates_
- **[Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems](http://arxiv.org/abs/2608.21095v1)** — Balkrishna Giri et al. (2026). *Runtime Retrieval / Search / Verification*
- **[$Z^2$-ACT: End-to-End Verifiable Agentic Intent Control for Open 6G RAN](http://arxiv.org/abs/2608.21049v1)** — Sunder Ali Khowaja et al. (2026). *Runtime Retrieval / Search / Verification*
- **[UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists](http://arxiv.org/abs/2608.20918v1)** — Ye Chen et al. (2026). *Curriculum / Mid-Training / Replay*
- **[TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding](http://arxiv.org/abs/2608.20844v1)** — Rohan Kumar et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation](http://arxiv.org/abs/2608.20794v1)** — Haodong Chen et al. (2026). *Curriculum / Mid-Training / Replay*
- **[Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation](http://arxiv.org/abs/2608.20756v1)** — Rujin Liang et al. (2026). *Runtime Retrieval / Search / Verification*
- **[AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification](http://arxiv.org/abs/2608.20711v1)** — Ji Liu et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Towards Faithful Simulation of Human Shopping Behavior](http://arxiv.org/abs/2608.20707v1)** — Jiakai Tang et al. (2026). *Runtime Retrieval / Search / Verification*
- **[JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification](http://arxiv.org/abs/2608.20607v1)** — Tianxin Zhou et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Terminal Agents: A Survey of AI Agents in Command-Line Environments](http://arxiv.org/abs/2608.20485v1)** — Yi Bin et al. (2026). *Runtime Retrieval / Search / Verification*
- **[MidTool: Mid-training Data Synthesis for Agentic Tool Use](http://arxiv.org/abs/2608.20314v1)** — Fengqing Jiang et al. (2026). *Curriculum / Mid-Training / Replay* · _Cross-tier: Data Creation and Selection_
- **[Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization](http://arxiv.org/abs/2608.20281v1)** — Qian Kou et al. (2026). *Runtime Retrieval / Search / Verification*
- **[Decoding silent reading from non-invasive EEG](http://arxiv.org/abs/2608.20186v1)** — Ingo Marquardt et al. (2026). *Runtime Retrieval / Search / Verification*
- **[FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models](http://arxiv.org/abs/2608.20153v1)** — Dingzirui Wang et al. (2026). *Runtime Retrieval / Search / Verification*
<!-- AUTO-LITERATURE:END -->
so the workflow never rewrites the manually curated core of this repository.
Automation is intentionally conservative:
This makes the repository a living companion to the survey while preserving the quality and interpretability of the manually curated taxonomy.
Recent additions reflected in the survey/repository include:
Data-centric Artificial Intelligence: A Survey ACM Computing Surveys 2023
An Empirical Survey of Data Augmentation for Limited Data Learning in NLP TACL 2023.
Best Practices and Lessons Learned on Synthetic Data for Language Models Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zhang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, Andrew M. Dai. COLM 2024.
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, Haobo Wang. arXiv 2024.
Large Language Models for Data Annotation: A Survey Zhen Tan, Dawei Li, Song Wang, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng, Huan Liu. arXiv 2024.
A Survey on Data Synthesis and Augmentation for Large Language Models arXiv 2024.
Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application ACM Transactions on Intelligent Systems and Technology 2024.
Data Augmentation using LLMs: Data Perspectives, Learning Paradigms and Challenges arXiv 2024.
A Survey of Multimodal Large Language Model from A Data-centric Perspective arXiv 2024.
A Survey of LLM × DATA arXiv 2025.
AI Alignment: A Comprehensive Survey arXiv 2025.
A Survey of Post-Training Scaling in Large Language Models ACL 2025.
Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives arXiv 2025.
Agentic Tool Use in Large Language Models. arXiv 2026.
The LLM Data Auditor: A Survey on Data-Centric Metrics for Cross-Modal Large Language Models. arXiv 2026.
We organize the LLM data lifecycle through an operation-centered, three-tier taxonomy. The goal is not merely to separate training stages, but to distinguish what information-bearing objects exist, how their content or membership is changed, and how available information is presented to and consumed by the model.
| Tier | Operational question | Scope | Representative examples |
|---|---|---|---|
| Tier 1: Data Substrates | What information-bearing artifacts or environments are available? | Corpora, instruction/preference data, benchmarks, reasoning/code resources, trajectories, executable environments | Pretraining corpora, SFT/RL data, math/code benchmarks, tool-use trajectories, agent environments |
| Tier 2: Data Creation and Selection | What supervision exists, and how is it created or modified? | Annotation, synthesis, transformation, filtering, diversity/quality/model-aware selection | Distillation, self-instruct, back-translation, symbolic generation, data filtering, influential-data selection |
| Tier 3: Data Ingestion Strategies | How, when, how often, and under what context is information consumed? | Packing, contextual conditioning, mixture weighting, curriculum, annealing, replay, retrieval, search, decoding, verification | Sample construction, data mixing, mid-training, multi-stage curricula, RAG, test-time search |
Placement rule. Tier 2 changes the supervision or data artifacts that exist; Tier 3 changes how available supervision or information is delivered to and consumed by the model. For example, offline data selection is primarily Tier 2, whereas dynamic data mixing is Tier 3. An agent trajectory is a Tier-1 substrate; generating or filtering trajectories is Tier 2; using them in iterative optimization or test-time interaction is Tier 3.
Cross-tier methods. The unit of analysis is the operation rather than the paper as a whole. Methods such as rejection-sampling fine-tuning or error-driven preference optimization can span multiple tiers: they may generate/filter supervision in Tier 2 and immediately consume the retained supervision during optimization in Tier 3.
The substrate view highlights a shift from raw scale toward curated information density, and—especially for agents—from stored text/trajectories toward executable environments that can generate tasks, observations, failures, recovery paths, and verification signals on demand.
Across creation and selection methods, the survey highlights three recurring transitions: teacher distillation → autonomous/self-play generation, static quality filtering → model-conditional utility, and stored examples → verifier-grounded or environment-generated supervision. Selection and synthesis increasingly interact through closed loops rather than remaining independent preprocessing stages.
Cross-tier note: this section focuses on the ingestion/optimization-time role of sample construction. Methods that generate persistent new supervision (e.g., rejection sampling or targeted error injection) also have a Tier-2 creation component.
Data ingestion spans both training-time scheduling and runtime information use. Training-time methods determine how supervision is organized, weighted, ordered, replayed, or conditionally presented; runtime methods determine when self-generated or external evidence is retrieved, searched, decoded, or verified.
The lifecycle view exposes several recurring transitions that are less visible when datasets, synthesis, selection, training schedules, and agent systems are studied separately:
These trends motivate treating data engineering as the joint design of artifacts, transformations, and adaptive information-use policies.
If you use this repository or the taxonomy in your research, please cite the accompanying survey:
@article{rao2025datacentric,
title = {A Data-Centric Perspective on the Lifecycle of Large Language Models},
author = {Rao, Jun and Liu, Xuebo and Yan, Haotian and Shen, Junjie and Mo, Haosi and
Dong, Yanghaopeng and Yan, Zihao and Wang, Ziyi and Lin, Zepeng and Meng, Xiaojun and
Yu, Zixiong and Deng, Liqun and Wei, Jiansheng and Wang, Yunhe and Zhang, Min},
journal = {TechRxiv},
year = {2025},
doi = {10.36227/techrxiv.176620610.03288677/v1},
url = {https://doi.org/10.36227/techrxiv.176620610.03288677/v1}
}
For GitHub-native citation support, we also recommend adding a CITATION.cff file to the repository root.
Contributions are welcome. Please open an issue or pull request to add a paper, dataset, benchmark, correction, or missing link.
To keep the list consistent:
Title — venue/year — link.Suggested pull-request title:
[Add] <Paper / Dataset / Benchmark Name> → <Section>
We thank the authors of the papers, datasets, benchmarks, and open-source resources collected here, as well as community contributors who help keep the list accurate and up to date.
Data is not only what a model trains on — it is also what gets created, selected, scheduled, retrieved, verified, and regenerated throughout the model lifecycle.
Python
100.0%