We maintain a structured collection of research on the interaction between bandit learning and large language models.
Corpus at a Glance · Search & Review Methodology · Taxonomy · Literature Navigation · Bibliography
We maintain this companion repository for our Bandit–LLM Interaction survey. We organize the literature in two directions: bandit methods that improve large language model systems, and LLM capabilities that augment bandit learning. We provide the complete classification in taxonomy.yaml and the corresponding citation records in references.bib.
153 unique studies · 2 directions · 8 stages · 18 components
We froze the survey corpus at August 15, 2026, with 153 unique studies. We may add newly released Bandit–LLM research after the survey cutoff as clearly marked post-survey updates.
We invite readers to use the Literature Navigation to browse papers by research stream and follow each reference to a verified arXiv record or official publication page. We also provide references.bib for citation management and taxonomy.yaml for reuse of our classification.
To provide broad and structured coverage of the rapidly evolving literature on Bandit–LLM interaction, we organized the literature search using the Population–Concept–Context (PCC) framework.
Rather than restricting retrieval to the components of our final taxonomy, we used PCC to define a broad search space around modern LLMs, genuine bandit methods, and their substantive technical interaction.
| PCC Element | Scope in This Review |
|---|---|
| Population | We consider modern large language models (LLMs) and LLM-based systems. |
| Concept | We focus on multi-armed bandits and related genuine bandit formulations, algorithms, and sequential decision mechanisms. |
| Context | We require substantive technical interaction between LLMs and bandits, either through bandit-based control of the LLM lifecycle or LLM-based augmentation of the bandit decision pipeline. |
We used the Population and Concept dimensions to construct broad retrieval queries, and we primarily applied the Context criterion during title/abstract screening and full-text assessment. This separation helped us preserve recall during literature identification without prematurely restricting retrieval to the component taxonomy that we developed later in the review.
We searched the literature published between January 1, 2022 and August 15, 2026. We use the latter date as the cutoff for our survey corpus and identify any later repository additions as Post-Survey Updates.
We searched Scopus, Web of Science Core Collection, ACM Digital Library, IEEE Xplore, and arXiv as our primary literature sources. We supplemented this search with Google Scholar, Semantic Scholar, backward citation tracing, and forward citation tracing. By combining these sources, we cover machine learning, natural language processing, information retrieval, recommender systems, data mining, operations research, and online learning.
We used the following canonical Population terms:
"large language model" OR "large language models" OR LLM OR LLMs
OR "language model" OR "language models"
We used the following canonical Concept terms:
bandit OR "multi-armed bandit" OR "multi armed bandit"
OR "contextual bandit" OR "combinatorial bandit" OR "linear bandit"
OR "bandit learning" OR "bandit algorithm" OR "Thompson sampling"
OR "upper confidence bound"
We combined the two groups using the following canonical search logic:
(Population terms) AND (Concept terms)
We adapted the syntax to each database interface. We present these concepts as a canonical reproducible strategy rather than as character-for-character historical queries. We did not require taxonomy-specific terms—such as prompting, retrieval, routing, caching, agent orchestration, reward estimation, exploration, and feedback interpretation—in the initial broad query; we applied them during screening and synthesis.
Core principle: We retained a study only when Bandit–LLM interaction formed a substantive part of its problem formulation, methodology, learning procedure, decision mechanism, or system design.
| We included | We excluded |
|---|---|
| We include studies in which modern LLMs or LLM-based systems form part of the method or studied environment. | We exclude studies in which LLMs or bandits appear only in background, introduction, related work, baselines, or incidental implementation components. |
| We include studies with a genuine bandit formulation, algorithm, exploration mechanism, or partial-feedback decision process. | We exclude studies that use “bandit” metaphorically. |
| We include studies in which bandits substantively control or adapt an LLM component, or LLMs substantively augment a bandit component. | We exclude studies that use generic reinforcement learning without a genuine bandit formulation or algorithm. |
| We include theoretical, methodological, empirical, systems, and negative-result studies. | We exclude older or generic language-model work included only because it is conceptually related. |
| We include simulated users, proxy tasks, synthetic data, and synthetic environments when the interaction is methodologically substantive. | We exclude reports with insufficient technical information to determine the substantive interaction. |
| We include peer-reviewed, accepted, forthcoming, and high-quality preprint studies, and we impose no venue restriction. | We do not count duplicate or superseded versions of the same substantive study independently. |
We preferred the final published version where available. We treated a later conference or journal publication and its earlier preprint as one substantive study unless they clearly constituted distinct technical contributions.
Operational scope. We define Bandit-Enhanced Large Language Models as work in which bandit methods adapt or control computational decisions across Pre-training → Post-training → Utilization → Evaluation. We define LLM-Enhanced Bandits as work in which LLMs augment Representation → Learning → Decision → Feedback. We assign a study to both directions when both interactions are methodologically substantive.
PCC Scope Definition
↓
Broad Literature Identification
↓
Backward / Forward Citation Expansion
↓
Deduplication and Version Consolidation
↓
Title / Abstract Screening
↓
Full-Text Eligibility Assessment
↓
Structured Evidence Extraction
↓
Component-Level Synthesis
↓
153 Unique Included Studies
We first screened candidate studies by title and abstract using the PCC scope. We retained ambiguous studies for full-text assessment rather than excluding them prematurely. For eligible studies, we consolidated multiple versions of the same substantive work and preferred the final published version where available.
We extracted evidence on the problem formulation, Bandit–LLM intervention mechanism, bandit formulation, LLM integration, theoretical analysis, experimental setting, empirical findings, comparisons and ablations, and reported limitations. We used this evidence to develop our component-level synthesis and bidirectional taxonomy.
We organize the literature according to where one technology intervenes in the computational process of the other. For Bandit-Enhanced Large Language Models, we follow the LLM lifecycle—Pre-training, Post-training, Utilization, and Evaluation. For LLM-Enhanced Bandits, we follow the bandit decision pipeline—Representation, Learning, Decision, and Feedback.
We allow multi-component and bidirectional studies to appear in multiple research streams. We provide our complete reusable classification in taxonomy.yaml.
We place a study in multiple research streams when it contains multiple substantive intervention mechanisms. We provide two complementary views: component-level tables for structural comparison and a detailed paper index for title-based browsing. Both views cover all 153 studies.
Choose a view: Component-Level Overview · Detailed Paper Index
We present each stage using the same four-column structure as our survey: Component, Research Stream, intervention mechanism, and References. We link every reference label directly to a verified arXiv record or, when no arXiv identifier is available, the official publication page.
| Component | Research Stream | Bandit Intervention | References |
|---|---|---|---|
| Pre-training | Adaptive data mixing | Allocate updates across data domains or sources | Albalak et al. (2023) |
| Pre-training configuration optimization | Adapt masking policies or training configurations | Urteaga et al. (2023) |
| Component | Research Stream | Bandit Intervention | References |
|---|---|---|---|
| Fine-tuning | Adaptive curriculum and training-data scheduling | Adapt datasets, examples, rollouts, or tasks to the evolving learning state | Do et al. (2026); Lu et al. (2026); McKenzie et al. (2026); Shin et al. (2026); Yang et al. (2026) |
| Online experience and skill control | Regulate newly generated experience, auxiliary skills, or reward-driven updates | Gönç et al. (2023); He et al. (2026); Hu et al. (2026) | |
| Bandit-guided policy learning and co-evolution | Train or control evolving decision policies under sequential reward feedback | Chen et al. (2025); Nie et al. (2025); Schmied et al. (2026); Xia et al. (2024) | |
| Alignment | Active preference acquisition | Allocate limited feedback to informative contexts, responses, or comparisons | Das et al. (2025); Dwaracherla et al. (2024); Ji et al. (2025); Mehta et al. (2023); Scheid et al. (2024) |
| Exploration-aware preference optimization | Expand response-space coverage through uncertainty-aware exploration | Bai et al. (2025); Xie et al. (2025); Xiong et al. (2024); Zhang et al. (2025) | |
| Policy-coupled online alignment | Adapt comparison collection and preference updates to the evolving policy | Li et al. (2025); Li & Yan (2025) | |
| Adaptive supervision and feedback control | Select and adapt reward signals, evidence, logged feedback, or response candidates | Duan et al. (2026); Azar et al. (2024); Kim et al. (2026); Lau et al. (2024); Liu et al. (2024); Nguyen et al. (2025) |
| Component | Research Stream | Bandit Intervention | References |
|---|---|---|---|
| Prompting | Fixed-pool prompt selection and structured sharing | Allocate evaluations across prompt candidates while sharing evidence through representations, features, preferences, or distributed statistics | Shi et al. (2024); Lin et al. (2024); Wu et al. (2024); Wang et al. (2025); Lu et al. (2025); Li et al. (2026); Lin et al. (2024); Wu et al. (2026) |
| Prompt generation and refinement | Adapt prompt-generation or modification strategies as the candidate space evolves | Ashizawa et al. (2025); Park et al. (2025); Kong et al. (2025); Hong et al. (2026) | |
| Contextual and deployment-time prompting | Select or adapt prompts according to users, queries, dialogue state, or interaction history | Chen et al. (2024); Cho et al. (2026); Ishikawa et al. (2026); Li et al. (2026); Monea et al. (2024); Nie et al. (2025); Ramesh et al. (2025) | |
| Joint prompt and system optimization | Optimize prompts jointly with retrieval, inference compute, logged feedback, or surrounding workflow decisions | Fu et al. (2024); Li et al. (2025); Mahmud et al. (2026); Kiyohara et al. (2025); Kiyohara et al. (2025); Young & Björner (2026) | |
| Retrieval | Retrieval strategy and configuration adaptation | Adapt retrieval depth, strategy, or RAG configuration according to request-level utility and cost | Tang et al. (2025); Dai et al. (2025); Fu et al. (2024) |
| Evidence allocation and selection | Allocate limited retrieval or context capacity across candidate evidence sources | Petcu et al. (2026); Tan et al. (2026); Du et al. (2026) | |
| Retrieval computation and dynamic memory control | Allocate retrieval-side computation or select useful information from evolving memory repositories | Pony et al. (2026); Zhang et al. (2026) | |
| Routing | Contextual model routing | Match requests to individual LLMs using contextual rewards, representations, priors, or auxiliary feedback | Hu et al. (2025); Nguyen et al. (2024); Tsai & Tran (2026); Chiang et al. (2025); Panda et al. (2025); Bao et al. (2026); Sridhar et al. (2026); Nguyen et al. (2026); Chadderwala (2025); Poon et al. (2026); Chu et al. (2026) |
| Resource-aware and constrained routing | Allocate models under quality–cost trade-offs, budgets, capacities, queues, or other operational constraints | Nguyen et al. (2024); Li (2025); Wei et al. (2025); Ziller et al. (2026); Dai et al. (2024); Huang et al. (2026); Wu et al. (2026); Bae et al. (2026); Taberner-Miller (2026); Zu et al. (2026); Zhang et al. (2026); Patra et al. (2026) | |
| Sequential and combinatorial routing | Select cascades, repeated attempts, subsets, or ensembles involving multiple serving options | Atalar (2026); Belloni et al. (2026); Hu et al. (2025); Liu et al. (2026); Rau et al. (2025); Xu et al. (2026) | |
| Composite serving-configuration routing | Route over speculative decoding, model–prompt–tool combinations, retrieval paths, compute budgets, or cached model states | Huang et al. (2024); Hou et al. (2025); Kim et al. (2026); Li et al. (2025); Ren et al. (2026); Tang et al. (2025); Huang et al. (2026); Li & Li (2026); Jadav et al. (2026) | |
| Adaptive routing under system change | Adapt routing as reward mappings, model pools, retrievers, services, or model quality change over time | Chen et al. (2024); Li & Li (2026); Wu & Lu (2026); Xia et al. (2024); Taberner-Miller (2026); Wang et al. (2025); Tang et al. (2025) | |
| Generation | Adaptive decoding and inference policies | Select decoding, speculative-inference, or inference-scaling configurations according to context and compute | Hou et al. (2025); Sridhar et al. (2025); Su et al. (2026); Huang et al. (2026); Mahmud et al. (2026) |
| Response- and token-level adaptive generation | Adapt response production or token-level decisions from sequential preference or reward feedback | Lau et al. (2024); Qu et al. (2025); Shin et al. (2025) | |
| Test-time compute and candidate allocation | Allocate additional generation across queries, evolving candidates, or evaluators | Zuo & Zhu (2025); Karlekar et al. (2026); Nguyen et al. (2025) | |
| Structured intermediate generation control | Select compact intermediate actions or optimization strategies that guide open-ended LLM generation | Song et al. (2025); Ran et al. (2025) | |
| Caching | Exact response-cache management | Learn retention, replacement, and reuse of exact query–response pairs under limited cache capacity | Yang et al. (2025) |
| Semantic caching | Reuse responses across semantically related requests while balancing inference cost against reuse mismatch | Liu et al. (2026); Atalar et al. (2026) | |
| Model-state caching | Jointly learn request routing and residency of reusable model states | Li & Li (2026) | |
| Agent Orchestration | Local agent and tool selection | Select reasoning modes, tools, specialists, executors, or local orchestration configurations during execution | Chadderwala (2025); Yu et al. (2026); Tang et al. (2026); Guan et al. (2026); Jin et al. (2026) |
| Workflow and topology control | Adapt communication structures, collaboration protocols, pipelines, or joint multi-agent configurations | Hoveyda et al. (2024); Chen et al. (2026); Jadav et al. (2026); Suntaxi et al. (2026); Atalar (2026); Dai et al. (2024) | |
| Adaptive computation allocation | Allocate additional LLM calls, iterations, branches, or search effort within an ongoing workflow | Belloni et al. (2026); Tang et al. (2024); Xing et al. (2026) | |
| Trust, verification, and integrity control | Adapt trust, validation, fallback, grounding, or integrity mechanisms during agentic execution | Xia et al. (2025); Young & Björner (2026) |
| Component | Research Stream | Bandit Intervention | References |
|---|---|---|---|
| Adaptive Evaluation | Best-model identification | Allocate evaluation budget toward candidate models that remain plausible winners while exploiting shared evaluation structure | Zhou et al. (2025); Tolochinsky et al. (2026); Lyu et al. (2026) |
| Ranking and Pareto identification | Allocate evaluations to resolve uncertain rankings or identify nondominated configurations under multiple objectives | Zouhar et al. (2026); Xue et al. (2026) | |
| Preference- and judge-based evaluation | Allocate pairwise comparisons or repeated judge calls according to information, cost, or evaluation uncertainty | Gharat et al. (2026); Saha et al. (2026) | |
| Adaptive diagnostic evaluation | Direct evaluation toward informative responses, context perturbations, behavioral probes, or evolving candidate solutions | Dai et al. (2026); Pan et al. (2026); Krishnamurthy et al. (2024); Karlekar et al. (2026) |
| Component | Research Stream | LLM Intervention | References |
|---|---|---|---|
| Context Representation | Semantic context encoding | Encode textual or prompt–response contexts into dense semantic features for reward prediction and exploration | Baheri & Alm (2023); Lin et al. (2024); Dwaracherla et al. (2024); Gönç et al. (2023) |
| Task-adapted context representation | Construct decision-specific representations that expose semantics relevant to routing, retrieval, rewriting, or supervision selection | Wang et al. (2025); Cho et al. (2026); Tan et al. (2026); Tang et al. (2025); Wu & Lu (2026) | |
| Stateful and trajectory-aware representation | Encode evolving dialogue histories, reasoning traces, observations, or open-world agent states for sequential decision making | Li et al. (2026); Yu et al. (2026); Tang et al. (2026) | |
| Objective- and modality-aware representation | Augment semantic context with decision-relevant structure such as safety, resource demand, inference configuration, or multimodal information | Huang et al. (2026); Zhang et al. (2026); Zhang et al. (2026) | |
| Action Modeling | Semantic action representation | Embed prompts, models, demonstrations, or inference configurations so feedback can generalize across related actions | Wu et al. (2024); Li et al. (2026); Chiang et al. (2025); Huang et al. (2026) |
| Consequence-based action similarity | Reuse logged feedback across actions through semantic similarity among their generated outcomes | Kiyohara et al. (2025); Kiyohara et al. (2025) | |
| Relational and structured action modeling | Organize actions through clusters, hierarchies, graphs, or compositional structure to support statistical sharing | Do et al. (2026); McKenzie et al. (2026); Hong et al. (2026) | |
| Dynamic action-space construction | Use LLMs to generate, revise, mutate, or expand candidate actions during learning | Ishikawa et al. (2026); He et al. (2026); Karlekar et al. (2026); Zhang et al. (2026) |
| Component | Research Stream | LLM Intervention | References |
|---|---|---|---|
| Warm Start | Synthetic-interaction pretraining | Generate pseudo-interactions or synthetic preferences to initialize reward estimates and uncertainty before substantial online feedback is available | Alamdari et al. (2024); Bayley et al. (2026) |
| Prior-based initialization | Encode LLM-derived semantic knowledge into statistical priors that are subsequently revised through online observations | Feng et al. (2026); Lee et al. (2026); Wu & Lu (2026) | |
| Guided initialization and early interaction | Use LLM signals to initialize action values, model parameters, or early decisions while preserving subsequent bandit exploration | Duan et al. (2026); Chen et al. (2024); Yu et al. (2026) | |
| Reward Estimation | LLM-based outcome modeling | Use LLMs to predict action-level rewards or reward distributions while the bandit retains control over uncertainty-aware exploration | Felicioni et al. (2024); Sun et al. (2026); Berdica et al. (2026) |
| Proxy augmentation and correction | Use LLM predictions as auxiliary or surrogate observations and correct their bias using real rewards, residuals, or selective audits | Alamdari et al. (2024); Pershin et al. (2026); Ma et al. (2026); Ao et al. (2026) | |
| Language-to-reward construction | Translate natural-language objectives or preferences into executable reward functions and aggregate competing criteria | Behari et al. (2024); Verma et al. (2025) | |
| Semantic reward surrogates | Construct operational reward signals from LLM judgments, likelihoods, or pairwise preferences when direct task utility is unavailable | Cho et al. (2026); Du et al. (2026); Karlekar et al. (2026) | |
| Environment Modeling | Language-mediated posterior modeling | Maintain and update language-based beliefs over latent environment hypotheses for posterior sampling and sequential exploration | Arumugam & Griffiths (2025) |
| Component | Research Stream | LLM Intervention | References |
|---|---|---|---|
| Exploration | Uncertainty-aware exploration | Use LLM reward predictions or predictive variability within explicit optimism, posterior-sampling, or randomized exploration mechanisms | Felicioni et al. (2024); Sun et al. (2026); Berdica et al. (2026) |
| Semantic action-space restriction | Use pretrained semantic knowledge to identify a tractable candidate region before conventional statistical exploration | Harris & Slivkins (2025) | |
| Direct exploration control | Delegate exploration schedules or history-dependent exploration policies directly to an LLM | Curtò et al. (2023); Chen et al. (2025) | |
| Language-mediated model-based exploration | Represent uncertainty over latent environments in language and use sampled hypotheses or information gain to guide exploration | Arumugam & Griffiths (2025) | |
| Action Selection | Direct LLM action selection | Use interaction history and semantic reasoning to let the LLM directly select the next action or preference candidate | Hazime & Farooq (2025); Harris & Slivkins (2025); Xia et al. (2025) |
| Confidence-gated and validated selection | Subject LLM recommendations to statistical confidence tests, validation, or fallback procedures before execution | Cao et al. (2026); Xia et al. (2025) | |
| Candidate generation and restricted selection | Use the LLM to construct or update a smaller candidate set while a downstream bandit determines the executed action | Harris & Slivkins (2025); Liu et al. (2024) | |
| Proxy- and diagnosis-guided selection | Use LLM-generated proxies, diagnoses, or intermediate analysis as auxiliary evidence for statistically controlled action allocation | Ma et al. (2026); Suntaxi et al. (2026) |
| Component | Research Stream | LLM Intervention | References |
|---|---|---|---|
| Feedback Interpretation | Scalar and binary judging | Convert generated or unstructured outcomes into scalar or binary observations for conventional bandit updates | Chu et al. (2026); Ramesh et al. (2025) |
| Pairwise preference interpretation | Convert comparative outputs into pairwise preference observations for dueling-bandit learning | Wu et al. (2026) | |
| Semantic feedback shaping and propagation | Use semantic feedback to bias future decisions or propagate observed evidence across related actions and comparisons | Nguyen et al. (2026); Wang et al. (2026) | |
| Structured diagnosis and attribution | Interpret complex simulation or execution outcomes while preserving diagnostic information and attribution to the action that produced them | Behari et al. (2024); Suntaxi et al. (2026) |
We also list every paper as a full bibliographic entry for title-based browsing. We retain the same Direction → Stage → Component → Research Stream organization and link each entry to arXiv or its official publication page.
We provide references.bib for the bibliographic records used in our survey. We include the 153-study corpus together with supporting background and methodological references, and we identify the included corpus in taxonomy.yaml.
.
├── README.md # Methodology and literature navigation
├── references.bib # Survey and supporting references
├── taxonomy.yaml # Reusable corpus classification
└── LICENSE
We freeze our Survey Corpus Snapshot at 153 studies through August 15, 2026. We may add later work under Post-Survey Updates, clearly separated from our frozen snapshot.
We welcome suggestions for missing or newly published work through issues or pull requests. We ask contributors to include an authoritative citation, a short explanation of the substantive Bandit–LLM interaction, and a proposed taxonomy location. We do not include work that mentions bandits or LLMs only as background.
If you use this collection, please cite our accompanying survey. We will add the complete publication metadata here when our paper is publicly available.
7 commits
TeX
100.0%
We maintain a structured collection of research on the interaction between bandit learning and large language models.
Corpus at a Glance · Search & Review Methodology · Taxonomy · Literature Navigation · Bibliography
We maintain this companion repository for our Bandit–LLM Interaction survey. We organize the literature in two directions: bandit methods that improve large language model systems, and LLM capabilities that augment bandit learning. We provide the complete classification in taxonomy.yaml and the corresponding citation records in references.bib.
153 unique studies · 2 directions · 8 stages · 18 components
We froze the survey corpus at August 15, 2026, with 153 unique studies. We may add newly released Bandit–LLM research after the survey cutoff as clearly marked post-survey updates.
We invite readers to use the Literature Navigation to browse papers by research stream and follow each reference to a verified arXiv record or official publication page. We also provide references.bib for citation management and taxonomy.yaml for reuse of our classification.
To provide broad and structured coverage of the rapidly evolving literature on Bandit–LLM interaction, we organized the literature search using the Population–Concept–Context (PCC) framework.
Rather than restricting retrieval to the components of our final taxonomy, we used PCC to define a broad search space around modern LLMs, genuine bandit methods, and their substantive technical interaction.
| PCC Element | Scope in This Review |
|---|---|
| Population | We consider modern large language models (LLMs) and LLM-based systems. |
| Concept | We focus on multi-armed bandits and related genuine bandit formulations, algorithms, and sequential decision mechanisms. |
| Context | We require substantive technical interaction between LLMs and bandits, either through bandit-based control of the LLM lifecycle or LLM-based augmentation of the bandit decision pipeline. |
We used the Population and Concept dimensions to construct broad retrieval queries, and we primarily applied the Context criterion during title/abstract screening and full-text assessment. This separation helped us preserve recall during literature identification without prematurely restricting retrieval to the component taxonomy that we developed later in the review.
We searched the literature published between January 1, 2022 and August 15, 2026. We use the latter date as the cutoff for our survey corpus and identify any later repository additions as Post-Survey Updates.
We searched Scopus, Web of Science Core Collection, ACM Digital Library, IEEE Xplore, and arXiv as our primary literature sources. We supplemented this search with Google Scholar, Semantic Scholar, backward citation tracing, and forward citation tracing. By combining these sources, we cover machine learning, natural language processing, information retrieval, recommender systems, data mining, operations research, and online learning.
We used the following canonical Population terms:
"large language model" OR "large language models" OR LLM OR LLMs
OR "language model" OR "language models"
We used the following canonical Concept terms:
bandit OR "multi-armed bandit" OR "multi armed bandit"
OR "contextual bandit" OR "combinatorial bandit" OR "linear bandit"
OR "bandit learning" OR "bandit algorithm" OR "Thompson sampling"
OR "upper confidence bound"
We combined the two groups using the following canonical search logic:
(Population terms) AND (Concept terms)
We adapted the syntax to each database interface. We present these concepts as a canonical reproducible strategy rather than as character-for-character historical queries. We did not require taxonomy-specific terms—such as prompting, retrieval, routing, caching, agent orchestration, reward estimation, exploration, and feedback interpretation—in the initial broad query; we applied them during screening and synthesis.
Core principle: We retained a study only when Bandit–LLM interaction formed a substantive part of its problem formulation, methodology, learning procedure, decision mechanism, or system design.
| We included | We excluded |
|---|---|
| We include studies in which modern LLMs or LLM-based systems form part of the method or studied environment. | We exclude studies in which LLMs or bandits appear only in background, introduction, related work, baselines, or incidental implementation components. |
| We include studies with a genuine bandit formulation, algorithm, exploration mechanism, or partial-feedback decision process. | We exclude studies that use “bandit” metaphorically. |
| We include studies in which bandits substantively control or adapt an LLM component, or LLMs substantively augment a bandit component. | We exclude studies that use generic reinforcement learning without a genuine bandit formulation or algorithm. |
| We include theoretical, methodological, empirical, systems, and negative-result studies. | We exclude older or generic language-model work included only because it is conceptually related. |
| We include simulated users, proxy tasks, synthetic data, and synthetic environments when the interaction is methodologically substantive. | We exclude reports with insufficient technical information to determine the substantive interaction. |
| We include peer-reviewed, accepted, forthcoming, and high-quality preprint studies, and we impose no venue restriction. | We do not count duplicate or superseded versions of the same substantive study independently. |
We preferred the final published version where available. We treated a later conference or journal publication and its earlier preprint as one substantive study unless they clearly constituted distinct technical contributions.
Operational scope. We define Bandit-Enhanced Large Language Models as work in which bandit methods adapt or control computational decisions across Pre-training → Post-training → Utilization → Evaluation. We define LLM-Enhanced Bandits as work in which LLMs augment Representation → Learning → Decision → Feedback. We assign a study to both directions when both interactions are methodologically substantive.
PCC Scope Definition
↓
Broad Literature Identification
↓
Backward / Forward Citation Expansion
↓
Deduplication and Version Consolidation
↓
Title / Abstract Screening
↓
Full-Text Eligibility Assessment
↓
Structured Evidence Extraction
↓
Component-Level Synthesis
↓
153 Unique Included Studies
We first screened candidate studies by title and abstract using the PCC scope. We retained ambiguous studies for full-text assessment rather than excluding them prematurely. For eligible studies, we consolidated multiple versions of the same substantive work and preferred the final published version where available.
We extracted evidence on the problem formulation, Bandit–LLM intervention mechanism, bandit formulation, LLM integration, theoretical analysis, experimental setting, empirical findings, comparisons and ablations, and reported limitations. We used this evidence to develop our component-level synthesis and bidirectional taxonomy.
We organize the literature according to where one technology intervenes in the computational process of the other. For Bandit-Enhanced Large Language Models, we follow the LLM lifecycle—Pre-training, Post-training, Utilization, and Evaluation. For LLM-Enhanced Bandits, we follow the bandit decision pipeline—Representation, Learning, Decision, and Feedback.
We allow multi-component and bidirectional studies to appear in multiple research streams. We provide our complete reusable classification in taxonomy.yaml.
We place a study in multiple research streams when it contains multiple substantive intervention mechanisms. We provide two complementary views: component-level tables for structural comparison and a detailed paper index for title-based browsing. Both views cover all 153 studies.
Choose a view: Component-Level Overview · Detailed Paper Index
We present each stage using the same four-column structure as our survey: Component, Research Stream, intervention mechanism, and References. We link every reference label directly to a verified arXiv record or, when no arXiv identifier is available, the official publication page.
| Component | Research Stream | Bandit Intervention | References |
|---|---|---|---|
| Pre-training | Adaptive data mixing | Allocate updates across data domains or sources | Albalak et al. (2023) |
| Pre-training configuration optimization | Adapt masking policies or training configurations | Urteaga et al. (2023) |
| Component | Research Stream | Bandit Intervention | References |
|---|---|---|---|
| Fine-tuning | Adaptive curriculum and training-data scheduling | Adapt datasets, examples, rollouts, or tasks to the evolving learning state | Do et al. (2026); Lu et al. (2026); McKenzie et al. (2026); Shin et al. (2026); Yang et al. (2026) |
| Online experience and skill control | Regulate newly generated experience, auxiliary skills, or reward-driven updates | Gönç et al. (2023); He et al. (2026); Hu et al. (2026) | |
| Bandit-guided policy learning and co-evolution | Train or control evolving decision policies under sequential reward feedback | Chen et al. (2025); Nie et al. (2025); Schmied et al. (2026); Xia et al. (2024) | |
| Alignment | Active preference acquisition | Allocate limited feedback to informative contexts, responses, or comparisons | Das et al. (2025); Dwaracherla et al. (2024); Ji et al. (2025); Mehta et al. (2023); Scheid et al. (2024) |
| Exploration-aware preference optimization | Expand response-space coverage through uncertainty-aware exploration | Bai et al. (2025); Xie et al. (2025); Xiong et al. (2024); Zhang et al. (2025) | |
| Policy-coupled online alignment | Adapt comparison collection and preference updates to the evolving policy | Li et al. (2025); Li & Yan (2025) | |
| Adaptive supervision and feedback control | Select and adapt reward signals, evidence, logged feedback, or response candidates | Duan et al. (2026); Azar et al. (2024); Kim et al. (2026); Lau et al. (2024); Liu et al. (2024); Nguyen et al. (2025) |
| Component | Research Stream | Bandit Intervention | References |
|---|---|---|---|
| Prompting | Fixed-pool prompt selection and structured sharing | Allocate evaluations across prompt candidates while sharing evidence through representations, features, preferences, or distributed statistics | Shi et al. (2024); Lin et al. (2024); Wu et al. (2024); Wang et al. (2025); Lu et al. (2025); Li et al. (2026); Lin et al. (2024); Wu et al. (2026) |
| Prompt generation and refinement | Adapt prompt-generation or modification strategies as the candidate space evolves | Ashizawa et al. (2025); Park et al. (2025); Kong et al. (2025); Hong et al. (2026) | |
| Contextual and deployment-time prompting | Select or adapt prompts according to users, queries, dialogue state, or interaction history | Chen et al. (2024); Cho et al. (2026); Ishikawa et al. (2026); Li et al. (2026); Monea et al. (2024); Nie et al. (2025); Ramesh et al. (2025) | |
| Joint prompt and system optimization | Optimize prompts jointly with retrieval, inference compute, logged feedback, or surrounding workflow decisions | Fu et al. (2024); Li et al. (2025); Mahmud et al. (2026); Kiyohara et al. (2025); Kiyohara et al. (2025); Young & Björner (2026) | |
| Retrieval | Retrieval strategy and configuration adaptation | Adapt retrieval depth, strategy, or RAG configuration according to request-level utility and cost | Tang et al. (2025); Dai et al. (2025); Fu et al. (2024) |
| Evidence allocation and selection | Allocate limited retrieval or context capacity across candidate evidence sources | Petcu et al. (2026); Tan et al. (2026); Du et al. (2026) | |
| Retrieval computation and dynamic memory control | Allocate retrieval-side computation or select useful information from evolving memory repositories | Pony et al. (2026); Zhang et al. (2026) | |
| Routing | Contextual model routing | Match requests to individual LLMs using contextual rewards, representations, priors, or auxiliary feedback | Hu et al. (2025); Nguyen et al. (2024); Tsai & Tran (2026); Chiang et al. (2025); Panda et al. (2025); Bao et al. (2026); Sridhar et al. (2026); Nguyen et al. (2026); Chadderwala (2025); Poon et al. (2026); Chu et al. (2026) |
| Resource-aware and constrained routing | Allocate models under quality–cost trade-offs, budgets, capacities, queues, or other operational constraints | Nguyen et al. (2024); Li (2025); Wei et al. (2025); Ziller et al. (2026); Dai et al. (2024); Huang et al. (2026); Wu et al. (2026); Bae et al. (2026); Taberner-Miller (2026); Zu et al. (2026); Zhang et al. (2026); Patra et al. (2026) | |
| Sequential and combinatorial routing | Select cascades, repeated attempts, subsets, or ensembles involving multiple serving options | Atalar (2026); Belloni et al. (2026); Hu et al. (2025); Liu et al. (2026); Rau et al. (2025); Xu et al. (2026) | |
| Composite serving-configuration routing | Route over speculative decoding, model–prompt–tool combinations, retrieval paths, compute budgets, or cached model states | Huang et al. (2024); Hou et al. (2025); Kim et al. (2026); Li et al. (2025); Ren et al. (2026); Tang et al. (2025); Huang et al. (2026); Li & Li (2026); Jadav et al. (2026) | |
| Adaptive routing under system change | Adapt routing as reward mappings, model pools, retrievers, services, or model quality change over time | Chen et al. (2024); Li & Li (2026); Wu & Lu (2026); Xia et al. (2024); Taberner-Miller (2026); Wang et al. (2025); Tang et al. (2025) | |
| Generation | Adaptive decoding and inference policies | Select decoding, speculative-inference, or inference-scaling configurations according to context and compute | Hou et al. (2025); Sridhar et al. (2025); Su et al. (2026); Huang et al. (2026); Mahmud et al. (2026) |
| Response- and token-level adaptive generation | Adapt response production or token-level decisions from sequential preference or reward feedback | Lau et al. (2024); Qu et al. (2025); Shin et al. (2025) | |
| Test-time compute and candidate allocation | Allocate additional generation across queries, evolving candidates, or evaluators | Zuo & Zhu (2025); Karlekar et al. (2026); Nguyen et al. (2025) | |
| Structured intermediate generation control | Select compact intermediate actions or optimization strategies that guide open-ended LLM generation | Song et al. (2025); Ran et al. (2025) | |
| Caching | Exact response-cache management | Learn retention, replacement, and reuse of exact query–response pairs under limited cache capacity | Yang et al. (2025) |
| Semantic caching | Reuse responses across semantically related requests while balancing inference cost against reuse mismatch | Liu et al. (2026); Atalar et al. (2026) | |
| Model-state caching | Jointly learn request routing and residency of reusable model states | Li & Li (2026) | |
| Agent Orchestration | Local agent and tool selection | Select reasoning modes, tools, specialists, executors, or local orchestration configurations during execution | Chadderwala (2025); Yu et al. (2026); Tang et al. (2026); Guan et al. (2026); Jin et al. (2026) |
| Workflow and topology control | Adapt communication structures, collaboration protocols, pipelines, or joint multi-agent configurations | Hoveyda et al. (2024); Chen et al. (2026); Jadav et al. (2026); Suntaxi et al. (2026); Atalar (2026); Dai et al. (2024) | |
| Adaptive computation allocation | Allocate additional LLM calls, iterations, branches, or search effort within an ongoing workflow | Belloni et al. (2026); Tang et al. (2024); Xing et al. (2026) | |
| Trust, verification, and integrity control | Adapt trust, validation, fallback, grounding, or integrity mechanisms during agentic execution | Xia et al. (2025); Young & Björner (2026) |
| Component | Research Stream | Bandit Intervention | References |
|---|---|---|---|
| Adaptive Evaluation | Best-model identification | Allocate evaluation budget toward candidate models that remain plausible winners while exploiting shared evaluation structure | Zhou et al. (2025); Tolochinsky et al. (2026); Lyu et al. (2026) |
| Ranking and Pareto identification | Allocate evaluations to resolve uncertain rankings or identify nondominated configurations under multiple objectives | Zouhar et al. (2026); Xue et al. (2026) | |
| Preference- and judge-based evaluation | Allocate pairwise comparisons or repeated judge calls according to information, cost, or evaluation uncertainty | Gharat et al. (2026); Saha et al. (2026) | |
| Adaptive diagnostic evaluation | Direct evaluation toward informative responses, context perturbations, behavioral probes, or evolving candidate solutions | Dai et al. (2026); Pan et al. (2026); Krishnamurthy et al. (2024); Karlekar et al. (2026) |
| Component | Research Stream | LLM Intervention | References |
|---|---|---|---|
| Context Representation | Semantic context encoding | Encode textual or prompt–response contexts into dense semantic features for reward prediction and exploration | Baheri & Alm (2023); Lin et al. (2024); Dwaracherla et al. (2024); Gönç et al. (2023) |
| Task-adapted context representation | Construct decision-specific representations that expose semantics relevant to routing, retrieval, rewriting, or supervision selection | Wang et al. (2025); Cho et al. (2026); Tan et al. (2026); Tang et al. (2025); Wu & Lu (2026) | |
| Stateful and trajectory-aware representation | Encode evolving dialogue histories, reasoning traces, observations, or open-world agent states for sequential decision making | Li et al. (2026); Yu et al. (2026); Tang et al. (2026) | |
| Objective- and modality-aware representation | Augment semantic context with decision-relevant structure such as safety, resource demand, inference configuration, or multimodal information | Huang et al. (2026); Zhang et al. (2026); Zhang et al. (2026) | |
| Action Modeling | Semantic action representation | Embed prompts, models, demonstrations, or inference configurations so feedback can generalize across related actions | Wu et al. (2024); Li et al. (2026); Chiang et al. (2025); Huang et al. (2026) |
| Consequence-based action similarity | Reuse logged feedback across actions through semantic similarity among their generated outcomes | Kiyohara et al. (2025); Kiyohara et al. (2025) | |
| Relational and structured action modeling | Organize actions through clusters, hierarchies, graphs, or compositional structure to support statistical sharing | Do et al. (2026); McKenzie et al. (2026); Hong et al. (2026) | |
| Dynamic action-space construction | Use LLMs to generate, revise, mutate, or expand candidate actions during learning | Ishikawa et al. (2026); He et al. (2026); Karlekar et al. (2026); Zhang et al. (2026) |
| Component | Research Stream | LLM Intervention | References |
|---|---|---|---|
| Warm Start | Synthetic-interaction pretraining | Generate pseudo-interactions or synthetic preferences to initialize reward estimates and uncertainty before substantial online feedback is available | Alamdari et al. (2024); Bayley et al. (2026) |
| Prior-based initialization | Encode LLM-derived semantic knowledge into statistical priors that are subsequently revised through online observations | Feng et al. (2026); Lee et al. (2026); Wu & Lu (2026) | |
| Guided initialization and early interaction | Use LLM signals to initialize action values, model parameters, or early decisions while preserving subsequent bandit exploration | Duan et al. (2026); Chen et al. (2024); Yu et al. (2026) | |
| Reward Estimation | LLM-based outcome modeling | Use LLMs to predict action-level rewards or reward distributions while the bandit retains control over uncertainty-aware exploration | Felicioni et al. (2024); Sun et al. (2026); Berdica et al. (2026) |
| Proxy augmentation and correction | Use LLM predictions as auxiliary or surrogate observations and correct their bias using real rewards, residuals, or selective audits | Alamdari et al. (2024); Pershin et al. (2026); Ma et al. (2026); Ao et al. (2026) | |
| Language-to-reward construction | Translate natural-language objectives or preferences into executable reward functions and aggregate competing criteria | Behari et al. (2024); Verma et al. (2025) | |
| Semantic reward surrogates | Construct operational reward signals from LLM judgments, likelihoods, or pairwise preferences when direct task utility is unavailable | Cho et al. (2026); Du et al. (2026); Karlekar et al. (2026) | |
| Environment Modeling | Language-mediated posterior modeling | Maintain and update language-based beliefs over latent environment hypotheses for posterior sampling and sequential exploration | Arumugam & Griffiths (2025) |
| Component | Research Stream | LLM Intervention | References |
|---|---|---|---|
| Exploration | Uncertainty-aware exploration | Use LLM reward predictions or predictive variability within explicit optimism, posterior-sampling, or randomized exploration mechanisms | Felicioni et al. (2024); Sun et al. (2026); Berdica et al. (2026) |
| Semantic action-space restriction | Use pretrained semantic knowledge to identify a tractable candidate region before conventional statistical exploration | Harris & Slivkins (2025) | |
| Direct exploration control | Delegate exploration schedules or history-dependent exploration policies directly to an LLM | Curtò et al. (2023); Chen et al. (2025) | |
| Language-mediated model-based exploration | Represent uncertainty over latent environments in language and use sampled hypotheses or information gain to guide exploration | Arumugam & Griffiths (2025) | |
| Action Selection | Direct LLM action selection | Use interaction history and semantic reasoning to let the LLM directly select the next action or preference candidate | Hazime & Farooq (2025); Harris & Slivkins (2025); Xia et al. (2025) |
| Confidence-gated and validated selection | Subject LLM recommendations to statistical confidence tests, validation, or fallback procedures before execution | Cao et al. (2026); Xia et al. (2025) | |
| Candidate generation and restricted selection | Use the LLM to construct or update a smaller candidate set while a downstream bandit determines the executed action | Harris & Slivkins (2025); Liu et al. (2024) | |
| Proxy- and diagnosis-guided selection | Use LLM-generated proxies, diagnoses, or intermediate analysis as auxiliary evidence for statistically controlled action allocation | Ma et al. (2026); Suntaxi et al. (2026) |
| Component | Research Stream | LLM Intervention | References |
|---|---|---|---|
| Feedback Interpretation | Scalar and binary judging | Convert generated or unstructured outcomes into scalar or binary observations for conventional bandit updates | Chu et al. (2026); Ramesh et al. (2025) |
| Pairwise preference interpretation | Convert comparative outputs into pairwise preference observations for dueling-bandit learning | Wu et al. (2026) | |
| Semantic feedback shaping and propagation | Use semantic feedback to bias future decisions or propagate observed evidence across related actions and comparisons | Nguyen et al. (2026); Wang et al. (2026) | |
| Structured diagnosis and attribution | Interpret complex simulation or execution outcomes while preserving diagnostic information and attribution to the action that produced them | Behari et al. (2024); Suntaxi et al. (2026) |
We also list every paper as a full bibliographic entry for title-based browsing. We retain the same Direction → Stage → Component → Research Stream organization and link each entry to arXiv or its official publication page.
We provide references.bib for the bibliographic records used in our survey. We include the 153-study corpus together with supporting background and methodological references, and we identify the included corpus in taxonomy.yaml.
.
├── README.md # Methodology and literature navigation
├── references.bib # Survey and supporting references
├── taxonomy.yaml # Reusable corpus classification
└── LICENSE
We freeze our Survey Corpus Snapshot at 153 studies through August 15, 2026. We may add later work under Post-Survey Updates, clearly separated from our frozen snapshot.
We welcome suggestions for missing or newly published work through issues or pull requests. We ask contributors to include an authoritative citation, a short explanation of the substantive Bandit–LLM interaction, and a proposed taxonomy location. We do not include work that mentions bandits or LLMs only as background.
If you use this collection, please cite our accompanying survey. We will add the complete publication metadata here when our paper is publicly available.
7 commits
TeX
100.0%