bucky1119/Awesome-BanditLLM-Interaction

TeX

0

7 commits

updated Sep 9, 2026

See the code

README

Awesome Bandit–LLM Interaction

We maintain a structured collection of research on the interaction between bandit learning and large language models.

Papers Directions Coverage License: MIT

Corpus at a Glance · Search & Review Methodology · Taxonomy · Literature Navigation · Bibliography

👋 About

We maintain this companion repository for our Bandit–LLM Interaction survey. We organize the literature in two directions: bandit methods that improve large language model systems, and LLM capabilities that augment bandit learning. We provide the complete classification in taxonomy.yaml and the corresponding citation records in references.bib.

📊 Corpus at a Glance

153 unique studies · 2 directions · 8 stages · 18 components

We froze the survey corpus at August 15, 2026, with 153 unique studies. We may add newly released Bandit–LLM research after the survey cutoff as clearly marked post-survey updates.

We invite readers to use the Literature Navigation to browse papers by research stream and follow each reference to a verified arXiv record or official publication page. We also provide references.bib for citation management and taxonomy.yaml for reuse of our classification.

🔍 Search & Review Methodology

🧩 PCC Search Framework

To provide broad and structured coverage of the rapidly evolving literature on Bandit–LLM interaction, we organized the literature search using the Population–Concept–Context (PCC) framework.

Rather than restricting retrieval to the components of our final taxonomy, we used PCC to define a broad search space around modern LLMs, genuine bandit methods, and their substantive technical interaction.

PCC ElementScope in This Review
PopulationWe consider modern large language models (LLMs) and LLM-based systems.
ConceptWe focus on multi-armed bandits and related genuine bandit formulations, algorithms, and sequential decision mechanisms.
ContextWe require substantive technical interaction between LLMs and bandits, either through bandit-based control of the LLM lifecycle or LLM-based augmentation of the bandit decision pipeline.

We used the Population and Concept dimensions to construct broad retrieval queries, and we primarily applied the Context criterion during title/abstract screening and full-text assessment. This separation helped us preserve recall during literature identification without prematurely restricting retrieval to the component taxonomy that we developed later in the review.

🔎 Search Strategy

We searched the literature published between January 1, 2022 and August 15, 2026. We use the latter date as the cutoff for our survey corpus and identify any later repository additions as Post-Survey Updates.

We searched Scopus, Web of Science Core Collection, ACM Digital Library, IEEE Xplore, and arXiv as our primary literature sources. We supplemented this search with Google Scholar, Semantic Scholar, backward citation tracing, and forward citation tracing. By combining these sources, we cover machine learning, natural language processing, information retrieval, recommender systems, data mining, operations research, and online learning.

We used the following canonical Population terms:

"large language model" OR "large language models" OR LLM OR LLMs
OR "language model" OR "language models"

We used the following canonical Concept terms:

bandit OR "multi-armed bandit" OR "multi armed bandit"
OR "contextual bandit" OR "combinatorial bandit" OR "linear bandit"
OR "bandit learning" OR "bandit algorithm" OR "Thompson sampling"
OR "upper confidence bound"

We combined the two groups using the following canonical search logic:

(Population terms) AND (Concept terms)

We adapted the syntax to each database interface. We present these concepts as a canonical reproducible strategy rather than as character-for-character historical queries. We did not require taxonomy-specific terms—such as prompting, retrieval, routing, caching, agent orchestration, reward estimation, exploration, and feedback interpretation—in the initial broad query; we applied them during screening and synthesis.

✅ Eligibility Criteria

Core principle: We retained a study only when Bandit–LLM interaction formed a substantive part of its problem formulation, methodology, learning procedure, decision mechanism, or system design.

We includedWe excluded
We include studies in which modern LLMs or LLM-based systems form part of the method or studied environment.We exclude studies in which LLMs or bandits appear only in background, introduction, related work, baselines, or incidental implementation components.
We include studies with a genuine bandit formulation, algorithm, exploration mechanism, or partial-feedback decision process.We exclude studies that use “bandit” metaphorically.
We include studies in which bandits substantively control or adapt an LLM component, or LLMs substantively augment a bandit component.We exclude studies that use generic reinforcement learning without a genuine bandit formulation or algorithm.
We include theoretical, methodological, empirical, systems, and negative-result studies.We exclude older or generic language-model work included only because it is conceptually related.
We include simulated users, proxy tasks, synthetic data, and synthetic environments when the interaction is methodologically substantive.We exclude reports with insufficient technical information to determine the substantive interaction.
We include peer-reviewed, accepted, forthcoming, and high-quality preprint studies, and we impose no venue restriction.We do not count duplicate or superseded versions of the same substantive study independently.

We preferred the final published version where available. We treated a later conference or journal publication and its earlier preprint as one substantive study unless they clearly constituted distinct technical contributions.

Operational scope. We define Bandit-Enhanced Large Language Models as work in which bandit methods adapt or control computational decisions across Pre-training → Post-training → Utilization → Evaluation. We define LLM-Enhanced Bandits as work in which LLMs augment Representation → Learning → Decision → Feedback. We assign a study to both directions when both interactions are methodologically substantive.

🔄 Review Workflow

PCC Scope Definition
        ↓
Broad Literature Identification
        ↓
Backward / Forward Citation Expansion
        ↓
Deduplication and Version Consolidation
        ↓
Title / Abstract Screening
        ↓
Full-Text Eligibility Assessment
        ↓
Structured Evidence Extraction
        ↓
Component-Level Synthesis
        ↓
153 Unique Included Studies

We first screened candidate studies by title and abstract using the PCC scope. We retained ambiguous studies for full-text assessment rather than excluding them prematurely. For eligible studies, we consolidated multiple versions of the same substantive work and preferred the final published version where available.

We extracted evidence on the problem formulation, Bandit–LLM intervention mechanism, bandit formulation, LLM integration, theoretical analysis, experimental setting, empirical findings, comparisons and ablations, and reported limitations. We used this evidence to develop our component-level synthesis and bidirectional taxonomy.

🧭 Taxonomy

We organize the literature according to where one technology intervenes in the computational process of the other. For Bandit-Enhanced Large Language Models, we follow the LLM lifecycle—Pre-training, Post-training, Utilization, and Evaluation. For LLM-Enhanced Bandits, we follow the bandit decision pipeline—Representation, Learning, Decision, and Feedback.

We allow multi-component and bidirectional studies to appear in multiple research streams. We provide our complete reusable classification in taxonomy.yaml.

📚 Literature Navigation

We place a study in multiple research streams when it contains multiple substantive intervention mechanisms. We provide two complementary views: component-level tables for structural comparison and a detailed paper index for title-based browsing. Both views cover all 153 studies.

Choose a view: Component-Level Overview · Detailed Paper Index

Component-Level Overview

We present each stage using the same four-column structure as our survey: Component, Research Stream, intervention mechanism, and References. We link every reference label directly to a verified arXiv record or, when no arXiv identifier is available, the official publication page.

Bandit-Enhanced Large Language Models

Pre-training — 2 research streams · 2 papers
ComponentResearch StreamBandit InterventionReferences
Pre-trainingAdaptive data mixingAllocate updates across data domains or sourcesAlbalak et al. (2023)
Pre-training configuration optimizationAdapt masking policies or training configurationsUrteaga et al. (2023)
Post-training — 7 research streams · 28 papers
ComponentResearch StreamBandit InterventionReferences
Fine-tuningAdaptive curriculum and training-data schedulingAdapt datasets, examples, rollouts, or tasks to the evolving learning stateDo et al. (2026); Lu et al. (2026); McKenzie et al. (2026); Shin et al. (2026); Yang et al. (2026)
Online experience and skill controlRegulate newly generated experience, auxiliary skills, or reward-driven updatesGönç et al. (2023); He et al. (2026); Hu et al. (2026)
Bandit-guided policy learning and co-evolutionTrain or control evolving decision policies under sequential reward feedbackChen et al. (2025); Nie et al. (2025); Schmied et al. (2026); Xia et al. (2024)
AlignmentActive preference acquisitionAllocate limited feedback to informative contexts, responses, or comparisonsDas et al. (2025); Dwaracherla et al. (2024); Ji et al. (2025); Mehta et al. (2023); Scheid et al. (2024)
Exploration-aware preference optimizationExpand response-space coverage through uncertainty-aware explorationBai et al. (2025); Xie et al. (2025); Xiong et al. (2024); Zhang et al. (2025)
Policy-coupled online alignmentAdapt comparison collection and preference updates to the evolving policyLi et al. (2025); Li & Yan (2025)
Adaptive supervision and feedback controlSelect and adapt reward signals, evidence, logged feedback, or response candidatesDuan et al. (2026); Azar et al. (2024); Kim et al. (2026); Lau et al. (2024); Liu et al. (2024); Nguyen et al. (2025)
Utilization — 23 research streams · 96 papers
ComponentResearch StreamBandit InterventionReferences
PromptingFixed-pool prompt selection and structured sharingAllocate evaluations across prompt candidates while sharing evidence through representations, features, preferences, or distributed statisticsShi et al. (2024); Lin et al. (2024); Wu et al. (2024); Wang et al. (2025); Lu et al. (2025); Li et al. (2026); Lin et al. (2024); Wu et al. (2026)
Prompt generation and refinementAdapt prompt-generation or modification strategies as the candidate space evolvesAshizawa et al. (2025); Park et al. (2025); Kong et al. (2025); Hong et al. (2026)
Contextual and deployment-time promptingSelect or adapt prompts according to users, queries, dialogue state, or interaction historyChen et al. (2024); Cho et al. (2026); Ishikawa et al. (2026); Li et al. (2026); Monea et al. (2024); Nie et al. (2025); Ramesh et al. (2025)
Joint prompt and system optimizationOptimize prompts jointly with retrieval, inference compute, logged feedback, or surrounding workflow decisionsFu et al. (2024); Li et al. (2025); Mahmud et al. (2026); Kiyohara et al. (2025); Kiyohara et al. (2025); Young & Björner (2026)
RetrievalRetrieval strategy and configuration adaptationAdapt retrieval depth, strategy, or RAG configuration according to request-level utility and costTang et al. (2025); Dai et al. (2025); Fu et al. (2024)
Evidence allocation and selectionAllocate limited retrieval or context capacity across candidate evidence sourcesPetcu et al. (2026); Tan et al. (2026); Du et al. (2026)
Retrieval computation and dynamic memory controlAllocate retrieval-side computation or select useful information from evolving memory repositoriesPony et al. (2026); Zhang et al. (2026)
RoutingContextual model routingMatch requests to individual LLMs using contextual rewards, representations, priors, or auxiliary feedbackHu et al. (2025); Nguyen et al. (2024); Tsai & Tran (2026); Chiang et al. (2025); Panda et al. (2025); Bao et al. (2026); Sridhar et al. (2026); Nguyen et al. (2026); Chadderwala (2025); Poon et al. (2026); Chu et al. (2026)
Resource-aware and constrained routingAllocate models under quality–cost trade-offs, budgets, capacities, queues, or other operational constraintsNguyen et al. (2024); Li (2025); Wei et al. (2025); Ziller et al. (2026); Dai et al. (2024); Huang et al. (2026); Wu et al. (2026); Bae et al. (2026); Taberner-Miller (2026); Zu et al. (2026); Zhang et al. (2026); Patra et al. (2026)
Sequential and combinatorial routingSelect cascades, repeated attempts, subsets, or ensembles involving multiple serving optionsAtalar (2026); Belloni et al. (2026); Hu et al. (2025); Liu et al. (2026); Rau et al. (2025); Xu et al. (2026)
Composite serving-configuration routingRoute over speculative decoding, model–prompt–tool combinations, retrieval paths, compute budgets, or cached model statesHuang et al. (2024); Hou et al. (2025); Kim et al. (2026); Li et al. (2025); Ren et al. (2026); Tang et al. (2025); Huang et al. (2026); Li & Li (2026); Jadav et al. (2026)
Adaptive routing under system changeAdapt routing as reward mappings, model pools, retrievers, services, or model quality change over timeChen et al. (2024); Li & Li (2026); Wu & Lu (2026); Xia et al. (2024); Taberner-Miller (2026); Wang et al. (2025); Tang et al. (2025)
GenerationAdaptive decoding and inference policiesSelect decoding, speculative-inference, or inference-scaling configurations according to context and computeHou et al. (2025); Sridhar et al. (2025); Su et al. (2026); Huang et al. (2026); Mahmud et al. (2026)
Response- and token-level adaptive generationAdapt response production or token-level decisions from sequential preference or reward feedbackLau et al. (2024); Qu et al. (2025); Shin et al. (2025)
Test-time compute and candidate allocationAllocate additional generation across queries, evolving candidates, or evaluatorsZuo & Zhu (2025); Karlekar et al. (2026); Nguyen et al. (2025)
Structured intermediate generation controlSelect compact intermediate actions or optimization strategies that guide open-ended LLM generationSong et al. (2025); Ran et al. (2025)
CachingExact response-cache managementLearn retention, replacement, and reuse of exact query–response pairs under limited cache capacityYang et al. (2025)
Semantic cachingReuse responses across semantically related requests while balancing inference cost against reuse mismatchLiu et al. (2026); Atalar et al. (2026)
Model-state cachingJointly learn request routing and residency of reusable model statesLi & Li (2026)
Agent OrchestrationLocal agent and tool selectionSelect reasoning modes, tools, specialists, executors, or local orchestration configurations during executionChadderwala (2025); Yu et al. (2026); Tang et al. (2026); Guan et al. (2026); Jin et al. (2026)
Workflow and topology controlAdapt communication structures, collaboration protocols, pipelines, or joint multi-agent configurationsHoveyda et al. (2024); Chen et al. (2026); Jadav et al. (2026); Suntaxi et al. (2026); Atalar (2026); Dai et al. (2024)
Adaptive computation allocationAllocate additional LLM calls, iterations, branches, or search effort within an ongoing workflowBelloni et al. (2026); Tang et al. (2024); Xing et al. (2026)
Trust, verification, and integrity controlAdapt trust, validation, fallback, grounding, or integrity mechanisms during agentic executionXia et al. (2025); Young & Björner (2026)
Evaluation — 4 research streams · 11 papers
ComponentResearch StreamBandit InterventionReferences
Adaptive EvaluationBest-model identificationAllocate evaluation budget toward candidate models that remain plausible winners while exploiting shared evaluation structureZhou et al. (2025); Tolochinsky et al. (2026); Lyu et al. (2026)
Ranking and Pareto identificationAllocate evaluations to resolve uncertain rankings or identify nondominated configurations under multiple objectivesZouhar et al. (2026); Xue et al. (2026)
Preference- and judge-based evaluationAllocate pairwise comparisons or repeated judge calls according to information, cost, or evaluation uncertaintyGharat et al. (2026); Saha et al. (2026)
Adaptive diagnostic evaluationDirect evaluation toward informative responses, context perturbations, behavioral probes, or evolving candidate solutionsDai et al. (2026); Pan et al. (2026); Krishnamurthy et al. (2024); Karlekar et al. (2026)

LLM-Enhanced Bandits

Representation — 8 research streams · 27 papers
ComponentResearch StreamLLM InterventionReferences
Context RepresentationSemantic context encodingEncode textual or prompt–response contexts into dense semantic features for reward prediction and explorationBaheri & Alm (2023); Lin et al. (2024); Dwaracherla et al. (2024); Gönç et al. (2023)
Task-adapted context representationConstruct decision-specific representations that expose semantics relevant to routing, retrieval, rewriting, or supervision selectionWang et al. (2025); Cho et al. (2026); Tan et al. (2026); Tang et al. (2025); Wu & Lu (2026)
Stateful and trajectory-aware representationEncode evolving dialogue histories, reasoning traces, observations, or open-world agent states for sequential decision makingLi et al. (2026); Yu et al. (2026); Tang et al. (2026)
Objective- and modality-aware representationAugment semantic context with decision-relevant structure such as safety, resource demand, inference configuration, or multimodal informationHuang et al. (2026); Zhang et al. (2026); Zhang et al. (2026)
Action ModelingSemantic action representationEmbed prompts, models, demonstrations, or inference configurations so feedback can generalize across related actionsWu et al. (2024); Li et al. (2026); Chiang et al. (2025); Huang et al. (2026)
Consequence-based action similarityReuse logged feedback across actions through semantic similarity among their generated outcomesKiyohara et al. (2025); Kiyohara et al. (2025)
Relational and structured action modelingOrganize actions through clusters, hierarchies, graphs, or compositional structure to support statistical sharingDo et al. (2026); McKenzie et al. (2026); Hong et al. (2026)
Dynamic action-space constructionUse LLMs to generate, revise, mutate, or expand candidate actions during learningIshikawa et al. (2026); He et al. (2026); Karlekar et al. (2026); Zhang et al. (2026)
Learning — 8 research streams · 20 papers
ComponentResearch StreamLLM InterventionReferences
Warm StartSynthetic-interaction pretrainingGenerate pseudo-interactions or synthetic preferences to initialize reward estimates and uncertainty before substantial online feedback is availableAlamdari et al. (2024); Bayley et al. (2026)
Prior-based initializationEncode LLM-derived semantic knowledge into statistical priors that are subsequently revised through online observationsFeng et al. (2026); Lee et al. (2026); Wu & Lu (2026)
Guided initialization and early interactionUse LLM signals to initialize action values, model parameters, or early decisions while preserving subsequent bandit explorationDuan et al. (2026); Chen et al. (2024); Yu et al. (2026)
Reward EstimationLLM-based outcome modelingUse LLMs to predict action-level rewards or reward distributions while the bandit retains control over uncertainty-aware explorationFelicioni et al. (2024); Sun et al. (2026); Berdica et al. (2026)
Proxy augmentation and correctionUse LLM predictions as auxiliary or surrogate observations and correct their bias using real rewards, residuals, or selective auditsAlamdari et al. (2024); Pershin et al. (2026); Ma et al. (2026); Ao et al. (2026)
Language-to-reward constructionTranslate natural-language objectives or preferences into executable reward functions and aggregate competing criteriaBehari et al. (2024); Verma et al. (2025)
Semantic reward surrogatesConstruct operational reward signals from LLM judgments, likelihoods, or pairwise preferences when direct task utility is unavailableCho et al. (2026); Du et al. (2026); Karlekar et al. (2026)
Environment ModelingLanguage-mediated posterior modelingMaintain and update language-based beliefs over latent environment hypotheses for posterior sampling and sequential explorationArumugam & Griffiths (2025)
Decision — 8 research streams · 13 papers
ComponentResearch StreamLLM InterventionReferences
ExplorationUncertainty-aware explorationUse LLM reward predictions or predictive variability within explicit optimism, posterior-sampling, or randomized exploration mechanismsFelicioni et al. (2024); Sun et al. (2026); Berdica et al. (2026)
Semantic action-space restrictionUse pretrained semantic knowledge to identify a tractable candidate region before conventional statistical explorationHarris & Slivkins (2025)
Direct exploration controlDelegate exploration schedules or history-dependent exploration policies directly to an LLMCurtò et al. (2023); Chen et al. (2025)
Language-mediated model-based explorationRepresent uncertainty over latent environments in language and use sampled hypotheses or information gain to guide explorationArumugam & Griffiths (2025)
Action SelectionDirect LLM action selectionUse interaction history and semantic reasoning to let the LLM directly select the next action or preference candidateHazime & Farooq (2025); Harris & Slivkins (2025); Xia et al. (2025)
Confidence-gated and validated selectionSubject LLM recommendations to statistical confidence tests, validation, or fallback procedures before executionCao et al. (2026); Xia et al. (2025)
Candidate generation and restricted selectionUse the LLM to construct or update a smaller candidate set while a downstream bandit determines the executed actionHarris & Slivkins (2025); Liu et al. (2024)
Proxy- and diagnosis-guided selectionUse LLM-generated proxies, diagnoses, or intermediate analysis as auxiliary evidence for statistically controlled action allocationMa et al. (2026); Suntaxi et al. (2026)
Feedback — 4 research streams · 7 papers
ComponentResearch StreamLLM InterventionReferences
Feedback InterpretationScalar and binary judgingConvert generated or unstructured outcomes into scalar or binary observations for conventional bandit updatesChu et al. (2026); Ramesh et al. (2025)
Pairwise preference interpretationConvert comparative outputs into pairwise preference observations for dueling-bandit learningWu et al. (2026)
Semantic feedback shaping and propagationUse semantic feedback to bias future decisions or propagate observed evidence across related actions and comparisonsNguyen et al. (2026); Wang et al. (2026)
Structured diagnosis and attributionInterpret complex simulation or execution outcomes while preserving diagnostic information and attribution to the action that produced themBehari et al. (2024); Suntaxi et al. (2026)

Detailed Paper Index

We also list every paper as a full bibliographic entry for title-based browsing. We retain the same Direction → Stage → Component → Research Stream organization and link each entry to arXiv or its official publication page.

Bandit-Enhanced Large Language Models

Pre-training — 2 research streams · 2 papers
Pre-training
Adaptive data mixing
  • Alon Albalak et al. Efficient Online Data Mixing For Language Model Pre-Training. arXiv, 2023. arXiv
Pre-training configuration optimization
  • Iñigo Urteaga et al. Multi-armed bandits for resource efficient, online optimization of language model pre-training: the use case of dynamic masking. Findings of ACL, 2023. arXiv
Post-training — 7 research streams · 29 papers
Fine-tuning
Adaptive curriculum and training-data scheduling
  • Van Dai Do et al. SPaCe: Unlocking Sample-Efficient Large Language Models Training With Self-Pace Curriculum Learning. Findings of ACL, 2026. Paper
  • Xiaodong Lu et al. Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards. arXiv, 2026. arXiv
  • Darrien M. McKenzie, Nicklas Hansen, and Xiaolong Wang. Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models. arXiv, 2026. arXiv
  • Haebin Shin et al. DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections. Findings of ACL, 2026. Paper
  • Zairun Yang et al. Distribution-Value Coevolution for Adaptive RLHF Data Scheduling. KDD, 2026. Paper
Online experience and skill control
  • Kaan Gönç et al. User Feedback-based Online Learning for Intent Classification. ICMI, 2023. Paper
  • Zelin He et al. ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL. arXiv, 2026. arXiv
  • Xiao Hu et al. Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective. arXiv, 2026. arXiv
Bandit-guided policy learning and co-evolution
  • Sanxing Chen et al. When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training. arXiv, 2025. arXiv
  • Allen Nie et al. EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration. ICML, 2025. arXiv
  • Thomas Schmied et al. LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities. ICLR, 2026. arXiv
  • Yu Xia et al. Which LLM to Play? Convergence-Aware Online Model Selection with Time-Increasing Bandits. The Web Conference, 2024. Paper
Alignment
Active preference acquisition
  • Nirjhar Das et al. Active Preference Optimization for Sample Efficient RLHF. ECML PKDD, 2025. Paper
  • Vikranth Dwaracherla et al. Efficient Exploration for LLMs. ICML, 2024. arXiv
  • Kaixuan Ji, Jiafan He, and Quanquan Gu. Reinforcement Learning from Human Feedback with Active Queries. TMLR, 2025. arXiv
  • Viraj Mehta et al. Sample Efficient Preference Alignment in LLMs via Active Exploration. arXiv, 2023. arXiv
  • Antoine Scheid et al. Optimal Design for Reward Modeling in RLHF. arXiv, 2024. arXiv
Exploration-aware preference optimization
  • Chenjia Bai et al. Online Preference Alignment for Language Models via Count-based Exploration. ICLR, 2025. arXiv
  • Tengyang Xie et al. Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF. ICLR, 2025. arXiv
  • Wei Xiong et al. Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint. ICML, 2024. arXiv
  • Shenao Zhang et al. Self-Exploring Language Models: Active Preference Elicitation for Online Alignment. TMLR, 2025. arXiv
Policy-coupled online alignment
  • Long-Fei Li et al. Provably Efficient Online RLHF with One-Pass Reward Modeling. NeurIPS, 2025. Paper
  • Gen Li and Yuling Yan. Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback. arXiv, 2025. arXiv
Adaptive supervision and feedback control
  • Shaohua Duan et al. Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization. ACL, 2026. arXiv
  • Mohammad Gheshlaghi Azar et al. A General Theoretical Paradigm to Understand Learning from Human Preferences. AISTATS, 2024. arXiv
  • Taesan Kim et al. Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target. arXiv, 2026. arXiv
  • Allison Lau et al. Personalized Adaptation via In-Context Preference Learning. arXiv, 2024. arXiv
  • Zichen Liu et al. Sample-Efficient Alignment for LLMs. arXiv, 2024. arXiv
  • Duy Nguyen et al. LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits. NeurIPS, 2025. arXiv
Utilization — 23 research streams · 96 papers
Prompting
Fixed-pool prompt selection and structured sharing
  • Chengshuai Shi et al. Efficient Prompt Optimization Through the Lens of Best Arm Identification. NeurIPS, 2024. Paper
  • Xiaoqiang Lin et al. Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers. ICML, 2024. arXiv
  • Zhaoxuan Wu et al. Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars. NeurIPS, 2024. arXiv
  • Shuyang Wang, Somayeh Moazeni, and Diego Klabjan. SOPL: A Sequential Optimal Learning Approach to Automated Prompt Engineering in Large Language Models. Findings of ACL, 2025. arXiv
  • Pingchen Lu et al. FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits. arXiv, 2025. arXiv
  • Donghao Li et al. Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits. arXiv, 2026. arXiv
  • Xiaoqiang Lin et al. Prompt Optimization with Human Feedback. arXiv, 2024. arXiv
  • Yuanchen Wu et al. LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization. Findings of ACL, 2026. arXiv
Prompt generation and refinement
  • Rin Ashizawa et al. Bandit-Based Prompt Design Strategy Selection Improves Prompt Optimizers. Findings of ACL, 2025. Paper
  • Young-Joon Park et al. TwinBandit Prompt Optimizer: Adaptive Prompt Optimization via Synergistic Dual MAB-Guided Feedback. CIKM, 2025. Paper
  • Mingze Kong et al. Meta-Prompt Optimization for LLM-Based Sequential Decision Making. arXiv, 2025. arXiv
  • Zhi Hong et al. MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks. arXiv, 2026. arXiv
Contextual and deployment-time prompting
  • Zekai Chen, Po-Yu Chen, and Francois Buet-Golfouse. Online Personalizing White-box LLMs Generation with Neural Bandits. ICAIF, 2024. Paper
  • Nicole Cho et al. No One Size Fits All: QueryBandits for Hallucination Mitigation. arXiv, 2026. arXiv
  • Shion Ishikawa et al. Progressive Content Refinement with Decaying Reward Joint LinUCB. arXiv, 2026. arXiv
  • Xiang Li et al. ALSO: Adversarial Online Strategy Optimization for Social Agents. arXiv, 2026. arXiv
  • Giovanni Monea et al. LLMs Are In-Context Bandit Reinforcement Learners. arXiv, 2024. arXiv
  • Allen Nie et al. EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration. ICML, 2025. arXiv
  • Aditya Ramesh et al. Efficient Jailbreak Attack sequences on Large Language Models via Multi-Armed Bandit-based Context switching. ICLR, 2025. Paper
Joint prompt and system optimization
  • Jia Fu et al. AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation. Findings of ACL, 2024. arXiv
  • Yixuan Li et al. Online Prompt Selection for Program Synthesis. AAAI, 2025. arXiv
  • Saaduddin Mahmud et al. Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models. AAAI, 2026. arXiv
  • Haruka Kiyohara et al. Prompt Optimization with Logged Bandit Data. arXiv, 2025. arXiv
  • Haruka Kiyohara et al. An Off-Policy Learning Approach for Steering Sentence Generation towards Personalization. RecSys, 2025. Paper
  • Halley Young and Nikolaj Björner. Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves. arXiv, 2026. arXiv
Retrieval
Retrieval strategy and configuration adaptation
  • Xiaqiang Tang et al. MBA-RAG: a Bandit Approach for Adaptive Retrieval-Augmented Generation through Question Complexity. Proceedings of the 31st International Conference on Computational Linguistics, COLING 2025, Abu Dhabi, UAE, January 19-24, 2025, 2025. arXiv
  • Yuhang Dai, Jing Li, and Bohan Li. Relative Performance Bandits: An Adaptive RAG Framework with Reward-Aware Exploration. 31th IEEE International Conference on Parallel and Distributed Systems, ICPADS 2025, Hefei, China, December 14-18, 2025, 2025. Paper
  • Jia Fu et al. AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation. Findings of ACL, 2024. arXiv
Evidence allocation and selection
  • Roxana Petcu et al. Query Decomposition for RAG: Balancing Exploration-Exploitation. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2026 - Volume 1: Long Papers, Rabat, Morocco, March 24-29, 2026, 2026. arXiv
  • Hanzhuo Tan et al. Prompt-Based Code Completion via Multi-Retrieval Augmented Generation. ACM TOSEM, 2026. Paper
  • Linfeng Du et al. Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization. ACL, 2026. arXiv
Retrieval computation and dynamic memory control
  • Roi Pony et al. Col-Bandit: Zero-Shot Query-Time Pruning for Late-Interaction Retrieval. arXiv, 2026. arXiv
  • Junke Zhang et al. Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking. arXiv, 2026. arXiv
Routing
Contextual model routing
  • Xiaoyan Hu, Ho-fung Leung, and Farzan Farnia. PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs. ICML, 2025. arXiv
  • Quang H. Nguyen et al. MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs. arXiv, 2024. arXiv
  • M. Tsai and Phat Tran. Reward-Based Online LLM Routing via NeuralUCB. arXiv, 2026. arXiv
  • Chao-Kai Chiang, Takashi Ishida, and Masashi Sugiyama. LLM Routing with Dueling Feedback. arXiv, 2025. arXiv
  • Pranoy Panda et al. Adaptive LLM Routing under Budget Constraints. Findings of ACL, 2025. arXiv
  • Zhenghua Bao et al. OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning. arXiv, 2026. arXiv
  • Ajay Narayanan Sridhar et al. Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing. arXiv, 2026. arXiv
  • Son Nguyen, Xinyuan Liu, and Ransalu Senanayake. CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM. arXiv, 2026. arXiv
  • Nihir Chadderwala. Optimizing Life Sciences Agents in Real-Time using Reinforcement Learning. arXiv, 2025. arXiv
  • Manhin Poon et al. Online Multi-LLM Selection via Contextual Bandits Under Unstructured Context Evolution. AAAI, 2026. arXiv
  • Kexin Chu, Dawei Xiang, and Wei Zhang. Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents. arXiv, 2026. arXiv
Resource-aware and constrained routing
  • Quang H. Nguyen et al. MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs. arXiv, 2024. arXiv
  • Yang Li. LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing. arXiv, 2025. arXiv
  • Wang Wei et al. Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs. arXiv, 2025. arXiv
  • Thomas Ziller et al. GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference. arXiv, 2026. arXiv
  • Xiangxiang Dai et al. Cost-Effective Online Multi-LLM Selection with Versatile Reward Models. arXiv, 2024. arXiv
  • Yin Huang, Qingsong Liu, and Jie Xu. Online LLM Selection via Constrained Bandits with Time-Varying Demand. arXiv, 2026. arXiv
  • Shanglin Wu, Saatvik Kher, and Padhraic Smyth. Learning to Assign Prediction Tasks to Agents with Capacity Constraints. arXiv, 2026. arXiv
  • Seoungbin Bae, Junyoung Son, and Dabeen Lee. Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits. arXiv, 2026. arXiv
  • Annette Taberner-Miller. ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving. arXiv, 2026. arXiv
  • Ling Zu, Xiyue Peng, and Xin Liu. BARouter: A Budget-adaptive Online Large Language Model Router Framework. The Web Conference, 2026. Paper
  • Xianzhi Zhang et al. Adapter-Augmented Bandits for Online Multi-Constrained Multi-Modal Inference Scheduling. arXiv, 2026. arXiv
  • P. Patra et al. Truthful Reverse Auctions for Adaptive Selection via Contextual Multi-Armed Bandits. Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026. Paper
Sequential and combinatorial routing
  • Baran Atalar. Neural Bandit Based Optimal LLM Selection for Pipeline of Tasks. ACM SIGMETRICS Performance Evaluation Review, 2026. arXiv
  • Alexandre Belloni, Yan Chen, and Yehua Wei. Online Pandora's Box for Contextual LLM Cascading. arXiv, 2026. arXiv
  • Xiaoyan Hu et al. PromptWise: Online Learning for Cost-Aware Prompt Assignment in Generative Models. arXiv, 2025. arXiv
  • Xutong Liu et al. Combinatorial Logistic Online Learning and Its Applications in Nonlinear Networked Systems. IEEE Trans. Netw., 2026. Paper
  • Jonathan Rau et al. CoCoMaMa: Contextual Combinatorial Multi-Armed Bandit Router for Multi-Agent Systems with Volatile Arms. Proceedings of the Second International Workshop on Hypermedia Multi-Agent Systems (HyperAgents 2025) co-located with 28th European Conference on Artificial Intelligence (ECAI 2025), Bologna, Italy, October 26, 2025, 2025. Paper
  • Jinkun Xu et al. CES: Combinatorial Experts Selection via Contextual Linear Bandits. KDD, 2026. Paper
Composite serving-configuration routing
  • Jerry Huang et al. Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models. EMNLP, 2024. arXiv
  • Yunlong Hou et al. BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms. ICML, 2025. arXiv
  • Taehyeon Kim, Hojung Jung, and Se-Young Yun. Multi-Drafter Speculative Decoding with Alignment Feedback. Findings of ACL, 2026. arXiv
  • Yixuan Li et al. Online Prompt Selection for Program Synthesis. AAAI, 2025. arXiv
  • Junxiao Ren et al. CH-RAG: Complexity-Guided Hybrid Retrieval-Augmented for Adaptive LLM Generation. 2026 29th International Conference on Computer Supported Cooperative Work in Design (CSCWD), 2026. Paper
  • Xiaqiang Tang et al. Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs. AAAI, 2025. arXiv
  • Kaiyu Huang et al. UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling. arXiv, 2026. arXiv
  • Shaoang Li and Jian Li. POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving. arXiv, 2026. arXiv
  • Vasanth Rao Jadav, Shalini Sudarsan, and Vikram Isanaka. Cost-Aware LLM Orchestration via Contextual Bandit Learning. 2026 International Conference on Artificial Intelligence, Systems, and Emerging Technologies (ICAISET), 2026. Paper
Adaptive routing under system change
  • Dingyang Chen, Qi Zhang, and Yinglun Zhu. Efficient Sequential Decision Making with Large Language Models. EMNLP, 2024. arXiv
  • Shaoang Li and Jian Li. Near-Optimal Online Deployment and Routing for Streaming LLMs. ICLR, 2026. arXiv
  • Xinle Wu and Yao Lu. Reward Model Routing in Alignment. ICLR, 2026. arXiv
  • Yu Xia et al. Which LLM to Play? Convergence-Aware Online Model Selection with Time-Increasing Bandits. The Web Conference, 2024. Paper
  • Annette Taberner-Miller. ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving. arXiv, 2026. arXiv
  • Xinyuan Wang et al. MixLLM: Dynamic Routing in Mixed Large Language Models. NAACL, 2025. arXiv
  • Xiaqiang Tang et al. Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs. AAAI, 2025. arXiv
Generation
Adaptive decoding and inference policies
  • Yunlong Hou et al. BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms. ICML, 2025. arXiv
  • Aditya Sridhar et al. TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding. arXiv, 2025. arXiv
  • Chloe Su et al. Learning Adaptive LLM Decoding. arXiv, 2026. arXiv
  • Kaiyu Huang et al. UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling. arXiv, 2026. arXiv
  • Saaduddin Mahmud et al. Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models. AAAI, 2026. arXiv
Response- and token-level adaptive generation
  • Allison Lau et al. Personalized Adaptation via In-Context Preference Learning. arXiv, 2024. arXiv
  • Zikun Qu et al. T-POP: Test-Time Personalization with Online Preference Feedback. arXiv, 2025. arXiv
  • Suho Shin et al. Tokenized Bandit for LLM Decoding and Alignment. ICML, 2025. arXiv
Test-time compute and candidate allocation
  • Bowen Zuo and Yinglun Zhu. Strategic Scaling of Test-Time Compute: A Bandit Learning Approach. arXiv, 2025. arXiv
  • Sweta Karlekar et al. Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. arXiv, 2026. arXiv
  • Duy Nguyen et al. LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits. NeurIPS, 2025. arXiv
Structured intermediate generation control
  • Haochen Song et al. Tailored Behavior-Change Messaging for Physical Activity: Integrating Contextual Bandits and Large Language Models. arXiv, 2025. arXiv
  • Dezhi Ran et al. KernelBand: Boosting LLM-based Kernel Optimization with a Hierarchical and Hardware-aware Multi-armed Bandit. arXiv, 2025. arXiv
Caching
Exact response-cache management
  • Hantao Yang et al. LLM Cache Bandit Revisited: Addressing Query Heterogeneity for Cost-Effective LLM Inference. arXiv, 2025. arXiv
Semantic caching
  • Xutong Liu et al. Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation. IEEE INFOCOM 2026 - IEEE Conference on Computer Communications, Tokyo, Japan, May 18-21, 2026, 2026. Paper
  • Baran Atalar et al. Continuous Semantic Caching for Low-Cost LLM Serving. arXiv, 2026. arXiv
Model-state caching
  • Shaoang Li and Jian Li. POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving. arXiv, 2026. arXiv
Agent Orchestration
Local agent and tool selection
  • Nihir Chadderwala. Optimizing Life Sciences Agents in Real-Time using Reinforcement Learning. arXiv, 2025. arXiv
  • Sheldon Yu et al. OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents. arXiv, 2026. arXiv
  • Yuqi Tang et al. SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition. arXiv, 2026. arXiv
  • Zhaoyang Guan et al. Symphony-Coord: Adaptive Routing for Multi-Agent LLM Systems. arXiv, 2026. arXiv
  • Dian Jin et al. Personalizing Large Language Model Agents with Small Policy Models. arXiv, 2026. arXiv
Workflow and topology control
  • Mohanna Hoveyda et al. AQA: Adaptive Question Answering in a Society of LLMs via Contextual Multi-Armed Bandit. arXiv, 2024. arXiv
  • Huan Chen et al. Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm. arXiv, 2026. arXiv
  • Vasanth Rao Jadav, Shalini Sudarsan, and Vikram Isanaka. Cost-Aware LLM Orchestration via Contextual Bandit Learning. 2026 International Conference on Artificial Intelligence, Systems, and Emerging Technologies (ICAISET), 2026. Paper
  • Geremy Loachamín Suntaxi et al. Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection. arXiv, 2026. arXiv
  • Baran Atalar. Neural Bandit Based Optimal LLM Selection for Pipeline of Tasks. ACM SIGMETRICS Performance Evaluation Review, 2026. arXiv
  • Xiangxiang Dai et al. Cost-Effective Online Multi-LLM Selection with Versatile Reward Models. arXiv, 2024. arXiv
Adaptive computation allocation
  • Alexandre Belloni, Yan Chen, and Yehua Wei. Online Pandora's Box for Contextual LLM Cascading. arXiv, 2026. arXiv
  • Hao Tang et al. Code Repair with LLMs gives an Exploration-Exploitation Tradeoff. NeurIPS, 2024. arXiv
  • Sixue Xing et al. Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits. arXiv, 2026. arXiv
Trust, verification, and integrity control
  • Fanzeng Xia et al. Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of ACL, 2025. Paper
  • Halley Young and Nikolaj Björner. Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves. arXiv, 2026. arXiv
Evaluation — 4 research streams · 11 papers
Adaptive Evaluation
Best-model identification
  • Jin Peng Zhou et al. On Speeding Up Language Model Evaluation. ICLR, 2025. arXiv
  • Elad Tolochinsky, Yaniv Tenzer, and Yaniv Romano. Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization. arXiv, 2026. arXiv
  • Zifan Lyu et al. Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm that Provably Exploits Model Similarity. arXiv, 2026. arXiv
Ranking and Pareto identification
  • Vilém Zouhar et al. Dynamically Allocating Evaluation Effort for Model Ranking. arXiv, 2026. arXiv
  • Bo Xue et al. Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation. arXiv, 2026. arXiv
Preference- and judge-based evaluation
  • Sarvesh Gharat, Nikhil Karamchandani, and Jayakrishnan Nair. Cost-Aware Best Arm Identification via Dueling Feedback with Applications to Large Language Models. Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026. Paper
  • Aadirupa Saha, A. Wagde, and B. Kveton. LLM-as-Judge on a Budget. arXiv, 2026. arXiv
Adaptive diagnostic evaluation
  • Xiangxiang Dai et al. A Multi-Agent Conversational Bandit Approach to Online Evaluation and Selection of User-Aligned LLM Responses. AAAI, 2026. Paper
  • Deng Pan et al. Context Attribution with Multi-Armed Bandit Optimization. Findings of ACL, 2026. arXiv
  • Akshay Krishnamurthy et al. Can large language models explore in-context? NeurIPS, 2024. arXiv
  • Sweta Karlekar et al. Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. arXiv, 2026. arXiv

LLM-Enhanced Bandits

Representation — 8 research streams · 27 papers
Context Representation
Semantic context encoding
  • Ali Baheri and Cecilia O. Alm. LLMs-augmented Contextual Bandit. arXiv, 2023. arXiv
  • Xiaoqiang Lin et al. Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers. ICML, 2024. arXiv
  • Vikranth Dwaracherla et al. Efficient Exploration for LLMs. ICML, 2024. arXiv
  • Kaan Gönç et al. User Feedback-based Online Learning for Intent Classification. ICMI, 2023. Paper
Task-adapted context representation
  • Xinyuan Wang et al. MixLLM: Dynamic Routing in Mixed Large Language Models. NAACL, 2025. arXiv
  • Nicole Cho et al. No One Size Fits All: QueryBandits for Hallucination Mitigation. arXiv, 2026. arXiv
  • Hanzhuo Tan et al. Prompt-Based Code Completion via Multi-Retrieval Augmented Generation. ACM TOSEM, 2026. Paper
  • Xiaqiang Tang et al. Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs. AAAI, 2025. arXiv
  • Xinle Wu and Yao Lu. Reward Model Routing in Alignment. ICLR, 2026. arXiv
Stateful and trajectory-aware representation
  • Xiang Li et al. ALSO: Adversarial Online Strategy Optimization for Social Agents. arXiv, 2026. arXiv
  • Sheldon Yu et al. OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents. arXiv, 2026. arXiv
  • Yuqi Tang et al. SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition. arXiv, 2026. arXiv
Objective- and modality-aware representation
  • Kaiyu Huang et al. UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling. arXiv, 2026. arXiv
  • Zeyu Zhang et al. Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing. arXiv, 2026. arXiv
  • Xianzhi Zhang et al. Adapter-Augmented Bandits for Online Multi-Constrained Multi-Modal Inference Scheduling. arXiv, 2026. arXiv
Action Modeling
Semantic action representation
  • Zhaoxuan Wu et al. Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars. NeurIPS, 2024. arXiv
  • Donghao Li et al. Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits. arXiv, 2026. arXiv
  • Chao-Kai Chiang, Takashi Ishida, and Masashi Sugiyama. LLM Routing with Dueling Feedback. arXiv, 2025. arXiv
  • Kaiyu Huang et al. UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling. arXiv, 2026. arXiv
Consequence-based action similarity
  • Haruka Kiyohara et al. Prompt Optimization with Logged Bandit Data. arXiv, 2025. arXiv
  • Haruka Kiyohara et al. An Off-Policy Learning Approach for Steering Sentence Generation towards Personalization. RecSys, 2025. Paper
Relational and structured action modeling
  • Van Dai Do et al. SPaCe: Unlocking Sample-Efficient Large Language Models Training With Self-Pace Curriculum Learning. Findings of ACL, 2026. Paper
  • Darrien M. McKenzie, Nicklas Hansen, and Xiaolong Wang. Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models. arXiv, 2026. arXiv
  • Zhi Hong et al. MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks. arXiv, 2026. arXiv
Dynamic action-space construction
  • Shion Ishikawa et al. Progressive Content Refinement with Decaying Reward Joint LinUCB. arXiv, 2026. arXiv
  • Zelin He et al. ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL. arXiv, 2026. arXiv
  • Sweta Karlekar et al. Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. arXiv, 2026. arXiv
  • Junke Zhang et al. Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking. arXiv, 2026. arXiv
Learning — 8 research streams · 20 papers
Warm Start
Synthetic-interaction pretraining
  • Parand Alamdari, Yanshuai Cao, and Kevin Wilson. Jump Starting Bandits with LLM-Generated Prior Knowledge. EMNLP, 2024. arXiv
  • Adam Bayley et al. Jump Start or False Start? A Theoretical and Empirical Evaluation of LLM-initialized Bandits. TMLR, 2026. arXiv
Prior-based initialization
  • Qing Feng et al. LLM-Informed Bayesian Content Exploration in Ultra-Recency Recommendation. SIGIR, 2026. Paper
  • E. Lee et al. LLM-Derived Priors for Thompson Sampling in Cold-Start Comment Recommendation. arXiv, 2026. arXiv
  • Xinle Wu and Yao Lu. Reward Model Routing in Alignment. ICLR, 2026. arXiv
Guided initialization and early interaction
  • Shaohua Duan et al. Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization. ACL, 2026. arXiv
  • Dingyang Chen, Qi Zhang, and Yinglun Zhu. Efficient Sequential Decision Making with Large Language Models. EMNLP, 2024. arXiv
  • Sheldon Yu et al. OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents. arXiv, 2026. arXiv
Reward Estimation
LLM-based outcome modeling
  • Nicolò Felicioni et al. On the Importance of Uncertainty in Decision-Making with Large Language Models. TMLR, 2024. arXiv
  • Jiahang Sun et al. Large Language Model-Enhanced Multi-Armed Bandits. ACL, 2026. arXiv
  • Uljad Berdica et al. When Do We Need LLMs? A Diagnostic for Language-Driven Bandits. arXiv, 2026. arXiv
Proxy augmentation and correction
  • Parand Alamdari, Yanshuai Cao, and Kevin Wilson. Jump Starting Bandits with LLM-Generated Prior Knowledge. EMNLP, 2024. arXiv
  • M.N. Pershin et al. Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits. arXiv, 2026. arXiv
  • Tianyi Ma et al. Best-Arm Identification with Generative Proxy. arXiv, 2026. arXiv
  • Ruicheng Ao et al. Best Arm Identification with LLM Judges and Limited Human. arXiv, 2026. arXiv
Language-to-reward construction
  • Nikhil Behari et al. A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health. NeurIPS, 2024. arXiv
  • Shresth Verma et al. Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards. GameSec, 2025. arXiv
Semantic reward surrogates
  • Nicole Cho et al. No One Size Fits All: QueryBandits for Hallucination Mitigation. arXiv, 2026. arXiv
  • Linfeng Du et al. Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization. ACL, 2026. arXiv
  • Sweta Karlekar et al. Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. arXiv, 2026. arXiv
Environment Modeling
Language-mediated posterior modeling
  • Dilip Arumugam and Thomas L. Griffiths. Toward Efficient Exploration by Large Language Model Agents. arXiv, 2025. arXiv
Decision — 8 research streams · 13 papers
Exploration
Uncertainty-aware exploration
  • Nicolò Felicioni et al. On the Importance of Uncertainty in Decision-Making with Large Language Models. TMLR, 2024. arXiv
  • Jiahang Sun et al. Large Language Model-Enhanced Multi-Armed Bandits. ACL, 2026. arXiv
  • Uljad Berdica et al. When Do We Need LLMs? A Diagnostic for Language-Driven Bandits. arXiv, 2026. arXiv
Semantic action-space restriction
  • Keegan Harris and Aleksandrs Slivkins. Should You Use Your Large Language Model to Explore or Exploit? arXiv, 2025. arXiv
Direct exploration control
  • J. de Curtò et al. LLM-Informed Multi-Armed Bandit Strategies for Non-Stationary Environments. Electronics, 2023. Paper
  • Sanxing Chen et al. When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training. arXiv, 2025. arXiv
Language-mediated model-based exploration
  • Dilip Arumugam and Thomas L. Griffiths. Toward Efficient Exploration by Large Language Model Agents. arXiv, 2025. arXiv
Action Selection
Direct LLM action selection
  • Jawad Hazime and Junaid Farooq. Evaluation of LLM Powered Agentic AI for Solving Multi-Arm Bandit Problems. IEEE COINS, 2025. Paper
  • Keegan Harris and Aleksandrs Slivkins. Should You Use Your Large Language Model to Explore or Exploit? arXiv, 2025. arXiv
  • Fanzeng Xia et al. Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of ACL, 2025. Paper
Confidence-gated and validated selection
  • Junyu Cao et al. LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning. arXiv, 2026. arXiv
  • Fanzeng Xia et al. Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of ACL, 2025. Paper
Candidate generation and restricted selection
  • Keegan Harris and Aleksandrs Slivkins. Should You Use Your Large Language Model to Explore or Exploit? arXiv, 2025. arXiv
  • Zichen Liu et al. Sample-Efficient Alignment for LLMs. arXiv, 2024. arXiv
Proxy- and diagnosis-guided selection
  • Tianyi Ma et al. Best-Arm Identification with Generative Proxy. arXiv, 2026. arXiv
  • Geremy Loachamín Suntaxi et al. Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection. arXiv, 2026. arXiv
Feedback — 4 research streams · 7 papers
Feedback Interpretation
Scalar and binary judging
  • Kexin Chu, Dawei Xiang, and Wei Zhang. Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents. arXiv, 2026. arXiv
  • Aditya Ramesh et al. Efficient Jailbreak Attack sequences on Large Language Models via Multi-Armed Bandit-based Context switching. ICLR, 2025. Paper
Pairwise preference interpretation
  • Yuanchen Wu et al. LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization. Findings of ACL, 2026. arXiv
Semantic feedback shaping and propagation
  • Son Nguyen, Xinyuan Liu, and Ransalu Senanayake. CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM. arXiv, 2026. arXiv
  • Shengbo Wang, Hong Sun, and Ke Li. Preference Is More than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback. AAAI, 2026. Paper
Structured diagnosis and attribution
  • Nikhil Behari et al. A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health. NeurIPS, 2024. arXiv
  • Geremy Loachamín Suntaxi et al. Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection. arXiv, 2026. arXiv

📄 Bibliography

We provide references.bib for the bibliographic records used in our survey. We include the 153-study corpus together with supporting background and methodological references, and we identify the included corpus in taxonomy.yaml.

🗂️ Repository Structure

.
├── README.md                # Methodology and literature navigation
├── references.bib           # Survey and supporting references
├── taxonomy.yaml            # Reusable corpus classification
└── LICENSE

🤝 Updates / Contributing

We freeze our Survey Corpus Snapshot at 153 studies through August 15, 2026. We may add later work under Post-Survey Updates, clearly separated from our frozen snapshot.

We welcome suggestions for missing or newly published work through issues or pull requests. We ask contributors to include an authoritative citation, a short explanation of the substantive Bandit–LLM interaction, and a proposed taxonomy location. We do not include work that mentions bandits or LLMs only as background.

📝 Citation

If you use this collection, please cite our accompanying survey. We will add the complete publication metadata here when our paper is publicly available.

Contributors

bucky1119

7 commits

bucky1119/Awesome-BanditLLM-Interaction

TeX

0

7 commits

updated Sep 9, 2026

See the code

README

Awesome Bandit–LLM Interaction

We maintain a structured collection of research on the interaction between bandit learning and large language models.

Papers Directions Coverage License: MIT

Corpus at a Glance · Search & Review Methodology · Taxonomy · Literature Navigation · Bibliography

👋 About

We maintain this companion repository for our Bandit–LLM Interaction survey. We organize the literature in two directions: bandit methods that improve large language model systems, and LLM capabilities that augment bandit learning. We provide the complete classification in taxonomy.yaml and the corresponding citation records in references.bib.

📊 Corpus at a Glance

153 unique studies · 2 directions · 8 stages · 18 components

We froze the survey corpus at August 15, 2026, with 153 unique studies. We may add newly released Bandit–LLM research after the survey cutoff as clearly marked post-survey updates.

We invite readers to use the Literature Navigation to browse papers by research stream and follow each reference to a verified arXiv record or official publication page. We also provide references.bib for citation management and taxonomy.yaml for reuse of our classification.

🔍 Search & Review Methodology

🧩 PCC Search Framework

To provide broad and structured coverage of the rapidly evolving literature on Bandit–LLM interaction, we organized the literature search using the Population–Concept–Context (PCC) framework.

Rather than restricting retrieval to the components of our final taxonomy, we used PCC to define a broad search space around modern LLMs, genuine bandit methods, and their substantive technical interaction.

PCC ElementScope in This Review
PopulationWe consider modern large language models (LLMs) and LLM-based systems.
ConceptWe focus on multi-armed bandits and related genuine bandit formulations, algorithms, and sequential decision mechanisms.
ContextWe require substantive technical interaction between LLMs and bandits, either through bandit-based control of the LLM lifecycle or LLM-based augmentation of the bandit decision pipeline.

We used the Population and Concept dimensions to construct broad retrieval queries, and we primarily applied the Context criterion during title/abstract screening and full-text assessment. This separation helped us preserve recall during literature identification without prematurely restricting retrieval to the component taxonomy that we developed later in the review.

🔎 Search Strategy

We searched the literature published between January 1, 2022 and August 15, 2026. We use the latter date as the cutoff for our survey corpus and identify any later repository additions as Post-Survey Updates.

We searched Scopus, Web of Science Core Collection, ACM Digital Library, IEEE Xplore, and arXiv as our primary literature sources. We supplemented this search with Google Scholar, Semantic Scholar, backward citation tracing, and forward citation tracing. By combining these sources, we cover machine learning, natural language processing, information retrieval, recommender systems, data mining, operations research, and online learning.

We used the following canonical Population terms:

"large language model" OR "large language models" OR LLM OR LLMs
OR "language model" OR "language models"

We used the following canonical Concept terms:

bandit OR "multi-armed bandit" OR "multi armed bandit"
OR "contextual bandit" OR "combinatorial bandit" OR "linear bandit"
OR "bandit learning" OR "bandit algorithm" OR "Thompson sampling"
OR "upper confidence bound"

We combined the two groups using the following canonical search logic:

(Population terms) AND (Concept terms)

We adapted the syntax to each database interface. We present these concepts as a canonical reproducible strategy rather than as character-for-character historical queries. We did not require taxonomy-specific terms—such as prompting, retrieval, routing, caching, agent orchestration, reward estimation, exploration, and feedback interpretation—in the initial broad query; we applied them during screening and synthesis.

✅ Eligibility Criteria

Core principle: We retained a study only when Bandit–LLM interaction formed a substantive part of its problem formulation, methodology, learning procedure, decision mechanism, or system design.

We includedWe excluded
We include studies in which modern LLMs or LLM-based systems form part of the method or studied environment.We exclude studies in which LLMs or bandits appear only in background, introduction, related work, baselines, or incidental implementation components.
We include studies with a genuine bandit formulation, algorithm, exploration mechanism, or partial-feedback decision process.We exclude studies that use “bandit” metaphorically.
We include studies in which bandits substantively control or adapt an LLM component, or LLMs substantively augment a bandit component.We exclude studies that use generic reinforcement learning without a genuine bandit formulation or algorithm.
We include theoretical, methodological, empirical, systems, and negative-result studies.We exclude older or generic language-model work included only because it is conceptually related.
We include simulated users, proxy tasks, synthetic data, and synthetic environments when the interaction is methodologically substantive.We exclude reports with insufficient technical information to determine the substantive interaction.
We include peer-reviewed, accepted, forthcoming, and high-quality preprint studies, and we impose no venue restriction.We do not count duplicate or superseded versions of the same substantive study independently.

We preferred the final published version where available. We treated a later conference or journal publication and its earlier preprint as one substantive study unless they clearly constituted distinct technical contributions.

Operational scope. We define Bandit-Enhanced Large Language Models as work in which bandit methods adapt or control computational decisions across Pre-training → Post-training → Utilization → Evaluation. We define LLM-Enhanced Bandits as work in which LLMs augment Representation → Learning → Decision → Feedback. We assign a study to both directions when both interactions are methodologically substantive.

🔄 Review Workflow

PCC Scope Definition
        ↓
Broad Literature Identification
        ↓
Backward / Forward Citation Expansion
        ↓
Deduplication and Version Consolidation
        ↓
Title / Abstract Screening
        ↓
Full-Text Eligibility Assessment
        ↓
Structured Evidence Extraction
        ↓
Component-Level Synthesis
        ↓
153 Unique Included Studies

We first screened candidate studies by title and abstract using the PCC scope. We retained ambiguous studies for full-text assessment rather than excluding them prematurely. For eligible studies, we consolidated multiple versions of the same substantive work and preferred the final published version where available.

We extracted evidence on the problem formulation, Bandit–LLM intervention mechanism, bandit formulation, LLM integration, theoretical analysis, experimental setting, empirical findings, comparisons and ablations, and reported limitations. We used this evidence to develop our component-level synthesis and bidirectional taxonomy.

🧭 Taxonomy

We organize the literature according to where one technology intervenes in the computational process of the other. For Bandit-Enhanced Large Language Models, we follow the LLM lifecycle—Pre-training, Post-training, Utilization, and Evaluation. For LLM-Enhanced Bandits, we follow the bandit decision pipeline—Representation, Learning, Decision, and Feedback.

We allow multi-component and bidirectional studies to appear in multiple research streams. We provide our complete reusable classification in taxonomy.yaml.

📚 Literature Navigation

We place a study in multiple research streams when it contains multiple substantive intervention mechanisms. We provide two complementary views: component-level tables for structural comparison and a detailed paper index for title-based browsing. Both views cover all 153 studies.

Choose a view: Component-Level Overview · Detailed Paper Index

Component-Level Overview

We present each stage using the same four-column structure as our survey: Component, Research Stream, intervention mechanism, and References. We link every reference label directly to a verified arXiv record or, when no arXiv identifier is available, the official publication page.

Bandit-Enhanced Large Language Models

Pre-training — 2 research streams · 2 papers
ComponentResearch StreamBandit InterventionReferences
Pre-trainingAdaptive data mixingAllocate updates across data domains or sourcesAlbalak et al. (2023)
Pre-training configuration optimizationAdapt masking policies or training configurationsUrteaga et al. (2023)
Post-training — 7 research streams · 28 papers
ComponentResearch StreamBandit InterventionReferences
Fine-tuningAdaptive curriculum and training-data schedulingAdapt datasets, examples, rollouts, or tasks to the evolving learning stateDo et al. (2026); Lu et al. (2026); McKenzie et al. (2026); Shin et al. (2026); Yang et al. (2026)
Online experience and skill controlRegulate newly generated experience, auxiliary skills, or reward-driven updatesGönç et al. (2023); He et al. (2026); Hu et al. (2026)
Bandit-guided policy learning and co-evolutionTrain or control evolving decision policies under sequential reward feedbackChen et al. (2025); Nie et al. (2025); Schmied et al. (2026); Xia et al. (2024)
AlignmentActive preference acquisitionAllocate limited feedback to informative contexts, responses, or comparisonsDas et al. (2025); Dwaracherla et al. (2024); Ji et al. (2025); Mehta et al. (2023); Scheid et al. (2024)
Exploration-aware preference optimizationExpand response-space coverage through uncertainty-aware explorationBai et al. (2025); Xie et al. (2025); Xiong et al. (2024); Zhang et al. (2025)
Policy-coupled online alignmentAdapt comparison collection and preference updates to the evolving policyLi et al. (2025); Li & Yan (2025)
Adaptive supervision and feedback controlSelect and adapt reward signals, evidence, logged feedback, or response candidatesDuan et al. (2026); Azar et al. (2024); Kim et al. (2026); Lau et al. (2024); Liu et al. (2024); Nguyen et al. (2025)
Utilization — 23 research streams · 96 papers
ComponentResearch StreamBandit InterventionReferences
PromptingFixed-pool prompt selection and structured sharingAllocate evaluations across prompt candidates while sharing evidence through representations, features, preferences, or distributed statisticsShi et al. (2024); Lin et al. (2024); Wu et al. (2024); Wang et al. (2025); Lu et al. (2025); Li et al. (2026); Lin et al. (2024); Wu et al. (2026)
Prompt generation and refinementAdapt prompt-generation or modification strategies as the candidate space evolvesAshizawa et al. (2025); Park et al. (2025); Kong et al. (2025); Hong et al. (2026)
Contextual and deployment-time promptingSelect or adapt prompts according to users, queries, dialogue state, or interaction historyChen et al. (2024); Cho et al. (2026); Ishikawa et al. (2026); Li et al. (2026); Monea et al. (2024); Nie et al. (2025); Ramesh et al. (2025)
Joint prompt and system optimizationOptimize prompts jointly with retrieval, inference compute, logged feedback, or surrounding workflow decisionsFu et al. (2024); Li et al. (2025); Mahmud et al. (2026); Kiyohara et al. (2025); Kiyohara et al. (2025); Young & Björner (2026)
RetrievalRetrieval strategy and configuration adaptationAdapt retrieval depth, strategy, or RAG configuration according to request-level utility and costTang et al. (2025); Dai et al. (2025); Fu et al. (2024)
Evidence allocation and selectionAllocate limited retrieval or context capacity across candidate evidence sourcesPetcu et al. (2026); Tan et al. (2026); Du et al. (2026)
Retrieval computation and dynamic memory controlAllocate retrieval-side computation or select useful information from evolving memory repositoriesPony et al. (2026); Zhang et al. (2026)
RoutingContextual model routingMatch requests to individual LLMs using contextual rewards, representations, priors, or auxiliary feedbackHu et al. (2025); Nguyen et al. (2024); Tsai & Tran (2026); Chiang et al. (2025); Panda et al. (2025); Bao et al. (2026); Sridhar et al. (2026); Nguyen et al. (2026); Chadderwala (2025); Poon et al. (2026); Chu et al. (2026)
Resource-aware and constrained routingAllocate models under quality–cost trade-offs, budgets, capacities, queues, or other operational constraintsNguyen et al. (2024); Li (2025); Wei et al. (2025); Ziller et al. (2026); Dai et al. (2024); Huang et al. (2026); Wu et al. (2026); Bae et al. (2026); Taberner-Miller (2026); Zu et al. (2026); Zhang et al. (2026); Patra et al. (2026)
Sequential and combinatorial routingSelect cascades, repeated attempts, subsets, or ensembles involving multiple serving optionsAtalar (2026); Belloni et al. (2026); Hu et al. (2025); Liu et al. (2026); Rau et al. (2025); Xu et al. (2026)
Composite serving-configuration routingRoute over speculative decoding, model–prompt–tool combinations, retrieval paths, compute budgets, or cached model statesHuang et al. (2024); Hou et al. (2025); Kim et al. (2026); Li et al. (2025); Ren et al. (2026); Tang et al. (2025); Huang et al. (2026); Li & Li (2026); Jadav et al. (2026)
Adaptive routing under system changeAdapt routing as reward mappings, model pools, retrievers, services, or model quality change over timeChen et al. (2024); Li & Li (2026); Wu & Lu (2026); Xia et al. (2024); Taberner-Miller (2026); Wang et al. (2025); Tang et al. (2025)
GenerationAdaptive decoding and inference policiesSelect decoding, speculative-inference, or inference-scaling configurations according to context and computeHou et al. (2025); Sridhar et al. (2025); Su et al. (2026); Huang et al. (2026); Mahmud et al. (2026)
Response- and token-level adaptive generationAdapt response production or token-level decisions from sequential preference or reward feedbackLau et al. (2024); Qu et al. (2025); Shin et al. (2025)
Test-time compute and candidate allocationAllocate additional generation across queries, evolving candidates, or evaluatorsZuo & Zhu (2025); Karlekar et al. (2026); Nguyen et al. (2025)
Structured intermediate generation controlSelect compact intermediate actions or optimization strategies that guide open-ended LLM generationSong et al. (2025); Ran et al. (2025)
CachingExact response-cache managementLearn retention, replacement, and reuse of exact query–response pairs under limited cache capacityYang et al. (2025)
Semantic cachingReuse responses across semantically related requests while balancing inference cost against reuse mismatchLiu et al. (2026); Atalar et al. (2026)
Model-state cachingJointly learn request routing and residency of reusable model statesLi & Li (2026)
Agent OrchestrationLocal agent and tool selectionSelect reasoning modes, tools, specialists, executors, or local orchestration configurations during executionChadderwala (2025); Yu et al. (2026); Tang et al. (2026); Guan et al. (2026); Jin et al. (2026)
Workflow and topology controlAdapt communication structures, collaboration protocols, pipelines, or joint multi-agent configurationsHoveyda et al. (2024); Chen et al. (2026); Jadav et al. (2026); Suntaxi et al. (2026); Atalar (2026); Dai et al. (2024)
Adaptive computation allocationAllocate additional LLM calls, iterations, branches, or search effort within an ongoing workflowBelloni et al. (2026); Tang et al. (2024); Xing et al. (2026)
Trust, verification, and integrity controlAdapt trust, validation, fallback, grounding, or integrity mechanisms during agentic executionXia et al. (2025); Young & Björner (2026)
Evaluation — 4 research streams · 11 papers
ComponentResearch StreamBandit InterventionReferences
Adaptive EvaluationBest-model identificationAllocate evaluation budget toward candidate models that remain plausible winners while exploiting shared evaluation structureZhou et al. (2025); Tolochinsky et al. (2026); Lyu et al. (2026)
Ranking and Pareto identificationAllocate evaluations to resolve uncertain rankings or identify nondominated configurations under multiple objectivesZouhar et al. (2026); Xue et al. (2026)
Preference- and judge-based evaluationAllocate pairwise comparisons or repeated judge calls according to information, cost, or evaluation uncertaintyGharat et al. (2026); Saha et al. (2026)
Adaptive diagnostic evaluationDirect evaluation toward informative responses, context perturbations, behavioral probes, or evolving candidate solutionsDai et al. (2026); Pan et al. (2026); Krishnamurthy et al. (2024); Karlekar et al. (2026)

LLM-Enhanced Bandits

Representation — 8 research streams · 27 papers
ComponentResearch StreamLLM InterventionReferences
Context RepresentationSemantic context encodingEncode textual or prompt–response contexts into dense semantic features for reward prediction and explorationBaheri & Alm (2023); Lin et al. (2024); Dwaracherla et al. (2024); Gönç et al. (2023)
Task-adapted context representationConstruct decision-specific representations that expose semantics relevant to routing, retrieval, rewriting, or supervision selectionWang et al. (2025); Cho et al. (2026); Tan et al. (2026); Tang et al. (2025); Wu & Lu (2026)
Stateful and trajectory-aware representationEncode evolving dialogue histories, reasoning traces, observations, or open-world agent states for sequential decision makingLi et al. (2026); Yu et al. (2026); Tang et al. (2026)
Objective- and modality-aware representationAugment semantic context with decision-relevant structure such as safety, resource demand, inference configuration, or multimodal informationHuang et al. (2026); Zhang et al. (2026); Zhang et al. (2026)
Action ModelingSemantic action representationEmbed prompts, models, demonstrations, or inference configurations so feedback can generalize across related actionsWu et al. (2024); Li et al. (2026); Chiang et al. (2025); Huang et al. (2026)
Consequence-based action similarityReuse logged feedback across actions through semantic similarity among their generated outcomesKiyohara et al. (2025); Kiyohara et al. (2025)
Relational and structured action modelingOrganize actions through clusters, hierarchies, graphs, or compositional structure to support statistical sharingDo et al. (2026); McKenzie et al. (2026); Hong et al. (2026)
Dynamic action-space constructionUse LLMs to generate, revise, mutate, or expand candidate actions during learningIshikawa et al. (2026); He et al. (2026); Karlekar et al. (2026); Zhang et al. (2026)
Learning — 8 research streams · 20 papers
ComponentResearch StreamLLM InterventionReferences
Warm StartSynthetic-interaction pretrainingGenerate pseudo-interactions or synthetic preferences to initialize reward estimates and uncertainty before substantial online feedback is availableAlamdari et al. (2024); Bayley et al. (2026)
Prior-based initializationEncode LLM-derived semantic knowledge into statistical priors that are subsequently revised through online observationsFeng et al. (2026); Lee et al. (2026); Wu & Lu (2026)
Guided initialization and early interactionUse LLM signals to initialize action values, model parameters, or early decisions while preserving subsequent bandit explorationDuan et al. (2026); Chen et al. (2024); Yu et al. (2026)
Reward EstimationLLM-based outcome modelingUse LLMs to predict action-level rewards or reward distributions while the bandit retains control over uncertainty-aware explorationFelicioni et al. (2024); Sun et al. (2026); Berdica et al. (2026)
Proxy augmentation and correctionUse LLM predictions as auxiliary or surrogate observations and correct their bias using real rewards, residuals, or selective auditsAlamdari et al. (2024); Pershin et al. (2026); Ma et al. (2026); Ao et al. (2026)
Language-to-reward constructionTranslate natural-language objectives or preferences into executable reward functions and aggregate competing criteriaBehari et al. (2024); Verma et al. (2025)
Semantic reward surrogatesConstruct operational reward signals from LLM judgments, likelihoods, or pairwise preferences when direct task utility is unavailableCho et al. (2026); Du et al. (2026); Karlekar et al. (2026)
Environment ModelingLanguage-mediated posterior modelingMaintain and update language-based beliefs over latent environment hypotheses for posterior sampling and sequential explorationArumugam & Griffiths (2025)
Decision — 8 research streams · 13 papers
ComponentResearch StreamLLM InterventionReferences
ExplorationUncertainty-aware explorationUse LLM reward predictions or predictive variability within explicit optimism, posterior-sampling, or randomized exploration mechanismsFelicioni et al. (2024); Sun et al. (2026); Berdica et al. (2026)
Semantic action-space restrictionUse pretrained semantic knowledge to identify a tractable candidate region before conventional statistical explorationHarris & Slivkins (2025)
Direct exploration controlDelegate exploration schedules or history-dependent exploration policies directly to an LLMCurtò et al. (2023); Chen et al. (2025)
Language-mediated model-based explorationRepresent uncertainty over latent environments in language and use sampled hypotheses or information gain to guide explorationArumugam & Griffiths (2025)
Action SelectionDirect LLM action selectionUse interaction history and semantic reasoning to let the LLM directly select the next action or preference candidateHazime & Farooq (2025); Harris & Slivkins (2025); Xia et al. (2025)
Confidence-gated and validated selectionSubject LLM recommendations to statistical confidence tests, validation, or fallback procedures before executionCao et al. (2026); Xia et al. (2025)
Candidate generation and restricted selectionUse the LLM to construct or update a smaller candidate set while a downstream bandit determines the executed actionHarris & Slivkins (2025); Liu et al. (2024)
Proxy- and diagnosis-guided selectionUse LLM-generated proxies, diagnoses, or intermediate analysis as auxiliary evidence for statistically controlled action allocationMa et al. (2026); Suntaxi et al. (2026)
Feedback — 4 research streams · 7 papers
ComponentResearch StreamLLM InterventionReferences
Feedback InterpretationScalar and binary judgingConvert generated or unstructured outcomes into scalar or binary observations for conventional bandit updatesChu et al. (2026); Ramesh et al. (2025)
Pairwise preference interpretationConvert comparative outputs into pairwise preference observations for dueling-bandit learningWu et al. (2026)
Semantic feedback shaping and propagationUse semantic feedback to bias future decisions or propagate observed evidence across related actions and comparisonsNguyen et al. (2026); Wang et al. (2026)
Structured diagnosis and attributionInterpret complex simulation or execution outcomes while preserving diagnostic information and attribution to the action that produced themBehari et al. (2024); Suntaxi et al. (2026)

Detailed Paper Index

We also list every paper as a full bibliographic entry for title-based browsing. We retain the same Direction → Stage → Component → Research Stream organization and link each entry to arXiv or its official publication page.

Bandit-Enhanced Large Language Models

Pre-training — 2 research streams · 2 papers
Pre-training
Adaptive data mixing
  • Alon Albalak et al. Efficient Online Data Mixing For Language Model Pre-Training. arXiv, 2023. arXiv
Pre-training configuration optimization
  • Iñigo Urteaga et al. Multi-armed bandits for resource efficient, online optimization of language model pre-training: the use case of dynamic masking. Findings of ACL, 2023. arXiv
Post-training — 7 research streams · 29 papers
Fine-tuning
Adaptive curriculum and training-data scheduling
  • Van Dai Do et al. SPaCe: Unlocking Sample-Efficient Large Language Models Training With Self-Pace Curriculum Learning. Findings of ACL, 2026. Paper
  • Xiaodong Lu et al. Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards. arXiv, 2026. arXiv
  • Darrien M. McKenzie, Nicklas Hansen, and Xiaolong Wang. Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models. arXiv, 2026. arXiv
  • Haebin Shin et al. DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections. Findings of ACL, 2026. Paper
  • Zairun Yang et al. Distribution-Value Coevolution for Adaptive RLHF Data Scheduling. KDD, 2026. Paper
Online experience and skill control
  • Kaan Gönç et al. User Feedback-based Online Learning for Intent Classification. ICMI, 2023. Paper
  • Zelin He et al. ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL. arXiv, 2026. arXiv
  • Xiao Hu et al. Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective. arXiv, 2026. arXiv
Bandit-guided policy learning and co-evolution
  • Sanxing Chen et al. When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training. arXiv, 2025. arXiv
  • Allen Nie et al. EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration. ICML, 2025. arXiv
  • Thomas Schmied et al. LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities. ICLR, 2026. arXiv
  • Yu Xia et al. Which LLM to Play? Convergence-Aware Online Model Selection with Time-Increasing Bandits. The Web Conference, 2024. Paper
Alignment
Active preference acquisition
  • Nirjhar Das et al. Active Preference Optimization for Sample Efficient RLHF. ECML PKDD, 2025. Paper
  • Vikranth Dwaracherla et al. Efficient Exploration for LLMs. ICML, 2024. arXiv
  • Kaixuan Ji, Jiafan He, and Quanquan Gu. Reinforcement Learning from Human Feedback with Active Queries. TMLR, 2025. arXiv
  • Viraj Mehta et al. Sample Efficient Preference Alignment in LLMs via Active Exploration. arXiv, 2023. arXiv
  • Antoine Scheid et al. Optimal Design for Reward Modeling in RLHF. arXiv, 2024. arXiv
Exploration-aware preference optimization
  • Chenjia Bai et al. Online Preference Alignment for Language Models via Count-based Exploration. ICLR, 2025. arXiv
  • Tengyang Xie et al. Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF. ICLR, 2025. arXiv
  • Wei Xiong et al. Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint. ICML, 2024. arXiv
  • Shenao Zhang et al. Self-Exploring Language Models: Active Preference Elicitation for Online Alignment. TMLR, 2025. arXiv
Policy-coupled online alignment
  • Long-Fei Li et al. Provably Efficient Online RLHF with One-Pass Reward Modeling. NeurIPS, 2025. Paper
  • Gen Li and Yuling Yan. Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback. arXiv, 2025. arXiv
Adaptive supervision and feedback control
  • Shaohua Duan et al. Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization. ACL, 2026. arXiv
  • Mohammad Gheshlaghi Azar et al. A General Theoretical Paradigm to Understand Learning from Human Preferences. AISTATS, 2024. arXiv
  • Taesan Kim et al. Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target. arXiv, 2026. arXiv
  • Allison Lau et al. Personalized Adaptation via In-Context Preference Learning. arXiv, 2024. arXiv
  • Zichen Liu et al. Sample-Efficient Alignment for LLMs. arXiv, 2024. arXiv
  • Duy Nguyen et al. LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits. NeurIPS, 2025. arXiv
Utilization — 23 research streams · 96 papers
Prompting
Fixed-pool prompt selection and structured sharing
  • Chengshuai Shi et al. Efficient Prompt Optimization Through the Lens of Best Arm Identification. NeurIPS, 2024. Paper
  • Xiaoqiang Lin et al. Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers. ICML, 2024. arXiv
  • Zhaoxuan Wu et al. Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars. NeurIPS, 2024. arXiv
  • Shuyang Wang, Somayeh Moazeni, and Diego Klabjan. SOPL: A Sequential Optimal Learning Approach to Automated Prompt Engineering in Large Language Models. Findings of ACL, 2025. arXiv
  • Pingchen Lu et al. FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits. arXiv, 2025. arXiv
  • Donghao Li et al. Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits. arXiv, 2026. arXiv
  • Xiaoqiang Lin et al. Prompt Optimization with Human Feedback. arXiv, 2024. arXiv
  • Yuanchen Wu et al. LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization. Findings of ACL, 2026. arXiv
Prompt generation and refinement
  • Rin Ashizawa et al. Bandit-Based Prompt Design Strategy Selection Improves Prompt Optimizers. Findings of ACL, 2025. Paper
  • Young-Joon Park et al. TwinBandit Prompt Optimizer: Adaptive Prompt Optimization via Synergistic Dual MAB-Guided Feedback. CIKM, 2025. Paper
  • Mingze Kong et al. Meta-Prompt Optimization for LLM-Based Sequential Decision Making. arXiv, 2025. arXiv
  • Zhi Hong et al. MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks. arXiv, 2026. arXiv
Contextual and deployment-time prompting
  • Zekai Chen, Po-Yu Chen, and Francois Buet-Golfouse. Online Personalizing White-box LLMs Generation with Neural Bandits. ICAIF, 2024. Paper
  • Nicole Cho et al. No One Size Fits All: QueryBandits for Hallucination Mitigation. arXiv, 2026. arXiv
  • Shion Ishikawa et al. Progressive Content Refinement with Decaying Reward Joint LinUCB. arXiv, 2026. arXiv
  • Xiang Li et al. ALSO: Adversarial Online Strategy Optimization for Social Agents. arXiv, 2026. arXiv
  • Giovanni Monea et al. LLMs Are In-Context Bandit Reinforcement Learners. arXiv, 2024. arXiv
  • Allen Nie et al. EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration. ICML, 2025. arXiv
  • Aditya Ramesh et al. Efficient Jailbreak Attack sequences on Large Language Models via Multi-Armed Bandit-based Context switching. ICLR, 2025. Paper
Joint prompt and system optimization
  • Jia Fu et al. AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation. Findings of ACL, 2024. arXiv
  • Yixuan Li et al. Online Prompt Selection for Program Synthesis. AAAI, 2025. arXiv
  • Saaduddin Mahmud et al. Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models. AAAI, 2026. arXiv
  • Haruka Kiyohara et al. Prompt Optimization with Logged Bandit Data. arXiv, 2025. arXiv
  • Haruka Kiyohara et al. An Off-Policy Learning Approach for Steering Sentence Generation towards Personalization. RecSys, 2025. Paper
  • Halley Young and Nikolaj Björner. Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves. arXiv, 2026. arXiv
Retrieval
Retrieval strategy and configuration adaptation
  • Xiaqiang Tang et al. MBA-RAG: a Bandit Approach for Adaptive Retrieval-Augmented Generation through Question Complexity. Proceedings of the 31st International Conference on Computational Linguistics, COLING 2025, Abu Dhabi, UAE, January 19-24, 2025, 2025. arXiv
  • Yuhang Dai, Jing Li, and Bohan Li. Relative Performance Bandits: An Adaptive RAG Framework with Reward-Aware Exploration. 31th IEEE International Conference on Parallel and Distributed Systems, ICPADS 2025, Hefei, China, December 14-18, 2025, 2025. Paper
  • Jia Fu et al. AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation. Findings of ACL, 2024. arXiv
Evidence allocation and selection
  • Roxana Petcu et al. Query Decomposition for RAG: Balancing Exploration-Exploitation. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2026 - Volume 1: Long Papers, Rabat, Morocco, March 24-29, 2026, 2026. arXiv
  • Hanzhuo Tan et al. Prompt-Based Code Completion via Multi-Retrieval Augmented Generation. ACM TOSEM, 2026. Paper
  • Linfeng Du et al. Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization. ACL, 2026. arXiv
Retrieval computation and dynamic memory control
  • Roi Pony et al. Col-Bandit: Zero-Shot Query-Time Pruning for Late-Interaction Retrieval. arXiv, 2026. arXiv
  • Junke Zhang et al. Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking. arXiv, 2026. arXiv
Routing
Contextual model routing
  • Xiaoyan Hu, Ho-fung Leung, and Farzan Farnia. PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs. ICML, 2025. arXiv
  • Quang H. Nguyen et al. MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs. arXiv, 2024. arXiv
  • M. Tsai and Phat Tran. Reward-Based Online LLM Routing via NeuralUCB. arXiv, 2026. arXiv
  • Chao-Kai Chiang, Takashi Ishida, and Masashi Sugiyama. LLM Routing with Dueling Feedback. arXiv, 2025. arXiv
  • Pranoy Panda et al. Adaptive LLM Routing under Budget Constraints. Findings of ACL, 2025. arXiv
  • Zhenghua Bao et al. OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning. arXiv, 2026. arXiv
  • Ajay Narayanan Sridhar et al. Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing. arXiv, 2026. arXiv
  • Son Nguyen, Xinyuan Liu, and Ransalu Senanayake. CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM. arXiv, 2026. arXiv
  • Nihir Chadderwala. Optimizing Life Sciences Agents in Real-Time using Reinforcement Learning. arXiv, 2025. arXiv
  • Manhin Poon et al. Online Multi-LLM Selection via Contextual Bandits Under Unstructured Context Evolution. AAAI, 2026. arXiv
  • Kexin Chu, Dawei Xiang, and Wei Zhang. Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents. arXiv, 2026. arXiv
Resource-aware and constrained routing
  • Quang H. Nguyen et al. MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs. arXiv, 2024. arXiv
  • Yang Li. LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing. arXiv, 2025. arXiv
  • Wang Wei et al. Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs. arXiv, 2025. arXiv
  • Thomas Ziller et al. GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference. arXiv, 2026. arXiv
  • Xiangxiang Dai et al. Cost-Effective Online Multi-LLM Selection with Versatile Reward Models. arXiv, 2024. arXiv
  • Yin Huang, Qingsong Liu, and Jie Xu. Online LLM Selection via Constrained Bandits with Time-Varying Demand. arXiv, 2026. arXiv
  • Shanglin Wu, Saatvik Kher, and Padhraic Smyth. Learning to Assign Prediction Tasks to Agents with Capacity Constraints. arXiv, 2026. arXiv
  • Seoungbin Bae, Junyoung Son, and Dabeen Lee. Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits. arXiv, 2026. arXiv
  • Annette Taberner-Miller. ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving. arXiv, 2026. arXiv
  • Ling Zu, Xiyue Peng, and Xin Liu. BARouter: A Budget-adaptive Online Large Language Model Router Framework. The Web Conference, 2026. Paper
  • Xianzhi Zhang et al. Adapter-Augmented Bandits for Online Multi-Constrained Multi-Modal Inference Scheduling. arXiv, 2026. arXiv
  • P. Patra et al. Truthful Reverse Auctions for Adaptive Selection via Contextual Multi-Armed Bandits. Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026. Paper
Sequential and combinatorial routing
  • Baran Atalar. Neural Bandit Based Optimal LLM Selection for Pipeline of Tasks. ACM SIGMETRICS Performance Evaluation Review, 2026. arXiv
  • Alexandre Belloni, Yan Chen, and Yehua Wei. Online Pandora's Box for Contextual LLM Cascading. arXiv, 2026. arXiv
  • Xiaoyan Hu et al. PromptWise: Online Learning for Cost-Aware Prompt Assignment in Generative Models. arXiv, 2025. arXiv
  • Xutong Liu et al. Combinatorial Logistic Online Learning and Its Applications in Nonlinear Networked Systems. IEEE Trans. Netw., 2026. Paper
  • Jonathan Rau et al. CoCoMaMa: Contextual Combinatorial Multi-Armed Bandit Router for Multi-Agent Systems with Volatile Arms. Proceedings of the Second International Workshop on Hypermedia Multi-Agent Systems (HyperAgents 2025) co-located with 28th European Conference on Artificial Intelligence (ECAI 2025), Bologna, Italy, October 26, 2025, 2025. Paper
  • Jinkun Xu et al. CES: Combinatorial Experts Selection via Contextual Linear Bandits. KDD, 2026. Paper
Composite serving-configuration routing
  • Jerry Huang et al. Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models. EMNLP, 2024. arXiv
  • Yunlong Hou et al. BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms. ICML, 2025. arXiv
  • Taehyeon Kim, Hojung Jung, and Se-Young Yun. Multi-Drafter Speculative Decoding with Alignment Feedback. Findings of ACL, 2026. arXiv
  • Yixuan Li et al. Online Prompt Selection for Program Synthesis. AAAI, 2025. arXiv
  • Junxiao Ren et al. CH-RAG: Complexity-Guided Hybrid Retrieval-Augmented for Adaptive LLM Generation. 2026 29th International Conference on Computer Supported Cooperative Work in Design (CSCWD), 2026. Paper
  • Xiaqiang Tang et al. Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs. AAAI, 2025. arXiv
  • Kaiyu Huang et al. UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling. arXiv, 2026. arXiv
  • Shaoang Li and Jian Li. POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving. arXiv, 2026. arXiv
  • Vasanth Rao Jadav, Shalini Sudarsan, and Vikram Isanaka. Cost-Aware LLM Orchestration via Contextual Bandit Learning. 2026 International Conference on Artificial Intelligence, Systems, and Emerging Technologies (ICAISET), 2026. Paper
Adaptive routing under system change
  • Dingyang Chen, Qi Zhang, and Yinglun Zhu. Efficient Sequential Decision Making with Large Language Models. EMNLP, 2024. arXiv
  • Shaoang Li and Jian Li. Near-Optimal Online Deployment and Routing for Streaming LLMs. ICLR, 2026. arXiv
  • Xinle Wu and Yao Lu. Reward Model Routing in Alignment. ICLR, 2026. arXiv
  • Yu Xia et al. Which LLM to Play? Convergence-Aware Online Model Selection with Time-Increasing Bandits. The Web Conference, 2024. Paper
  • Annette Taberner-Miller. ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving. arXiv, 2026. arXiv
  • Xinyuan Wang et al. MixLLM: Dynamic Routing in Mixed Large Language Models. NAACL, 2025. arXiv
  • Xiaqiang Tang et al. Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs. AAAI, 2025. arXiv
Generation
Adaptive decoding and inference policies
  • Yunlong Hou et al. BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms. ICML, 2025. arXiv
  • Aditya Sridhar et al. TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding. arXiv, 2025. arXiv
  • Chloe Su et al. Learning Adaptive LLM Decoding. arXiv, 2026. arXiv
  • Kaiyu Huang et al. UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling. arXiv, 2026. arXiv
  • Saaduddin Mahmud et al. Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models. AAAI, 2026. arXiv
Response- and token-level adaptive generation
  • Allison Lau et al. Personalized Adaptation via In-Context Preference Learning. arXiv, 2024. arXiv
  • Zikun Qu et al. T-POP: Test-Time Personalization with Online Preference Feedback. arXiv, 2025. arXiv
  • Suho Shin et al. Tokenized Bandit for LLM Decoding and Alignment. ICML, 2025. arXiv
Test-time compute and candidate allocation
  • Bowen Zuo and Yinglun Zhu. Strategic Scaling of Test-Time Compute: A Bandit Learning Approach. arXiv, 2025. arXiv
  • Sweta Karlekar et al. Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. arXiv, 2026. arXiv
  • Duy Nguyen et al. LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits. NeurIPS, 2025. arXiv
Structured intermediate generation control
  • Haochen Song et al. Tailored Behavior-Change Messaging for Physical Activity: Integrating Contextual Bandits and Large Language Models. arXiv, 2025. arXiv
  • Dezhi Ran et al. KernelBand: Boosting LLM-based Kernel Optimization with a Hierarchical and Hardware-aware Multi-armed Bandit. arXiv, 2025. arXiv
Caching
Exact response-cache management
  • Hantao Yang et al. LLM Cache Bandit Revisited: Addressing Query Heterogeneity for Cost-Effective LLM Inference. arXiv, 2025. arXiv
Semantic caching
  • Xutong Liu et al. Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation. IEEE INFOCOM 2026 - IEEE Conference on Computer Communications, Tokyo, Japan, May 18-21, 2026, 2026. Paper
  • Baran Atalar et al. Continuous Semantic Caching for Low-Cost LLM Serving. arXiv, 2026. arXiv
Model-state caching
  • Shaoang Li and Jian Li. POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving. arXiv, 2026. arXiv
Agent Orchestration
Local agent and tool selection
  • Nihir Chadderwala. Optimizing Life Sciences Agents in Real-Time using Reinforcement Learning. arXiv, 2025. arXiv
  • Sheldon Yu et al. OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents. arXiv, 2026. arXiv
  • Yuqi Tang et al. SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition. arXiv, 2026. arXiv
  • Zhaoyang Guan et al. Symphony-Coord: Adaptive Routing for Multi-Agent LLM Systems. arXiv, 2026. arXiv
  • Dian Jin et al. Personalizing Large Language Model Agents with Small Policy Models. arXiv, 2026. arXiv
Workflow and topology control
  • Mohanna Hoveyda et al. AQA: Adaptive Question Answering in a Society of LLMs via Contextual Multi-Armed Bandit. arXiv, 2024. arXiv
  • Huan Chen et al. Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm. arXiv, 2026. arXiv
  • Vasanth Rao Jadav, Shalini Sudarsan, and Vikram Isanaka. Cost-Aware LLM Orchestration via Contextual Bandit Learning. 2026 International Conference on Artificial Intelligence, Systems, and Emerging Technologies (ICAISET), 2026. Paper
  • Geremy Loachamín Suntaxi et al. Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection. arXiv, 2026. arXiv
  • Baran Atalar. Neural Bandit Based Optimal LLM Selection for Pipeline of Tasks. ACM SIGMETRICS Performance Evaluation Review, 2026. arXiv
  • Xiangxiang Dai et al. Cost-Effective Online Multi-LLM Selection with Versatile Reward Models. arXiv, 2024. arXiv
Adaptive computation allocation
  • Alexandre Belloni, Yan Chen, and Yehua Wei. Online Pandora's Box for Contextual LLM Cascading. arXiv, 2026. arXiv
  • Hao Tang et al. Code Repair with LLMs gives an Exploration-Exploitation Tradeoff. NeurIPS, 2024. arXiv
  • Sixue Xing et al. Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits. arXiv, 2026. arXiv
Trust, verification, and integrity control
  • Fanzeng Xia et al. Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of ACL, 2025. Paper
  • Halley Young and Nikolaj Björner. Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves. arXiv, 2026. arXiv
Evaluation — 4 research streams · 11 papers
Adaptive Evaluation
Best-model identification
  • Jin Peng Zhou et al. On Speeding Up Language Model Evaluation. ICLR, 2025. arXiv
  • Elad Tolochinsky, Yaniv Tenzer, and Yaniv Romano. Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization. arXiv, 2026. arXiv
  • Zifan Lyu et al. Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm that Provably Exploits Model Similarity. arXiv, 2026. arXiv
Ranking and Pareto identification
  • Vilém Zouhar et al. Dynamically Allocating Evaluation Effort for Model Ranking. arXiv, 2026. arXiv
  • Bo Xue et al. Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation. arXiv, 2026. arXiv
Preference- and judge-based evaluation
  • Sarvesh Gharat, Nikhil Karamchandani, and Jayakrishnan Nair. Cost-Aware Best Arm Identification via Dueling Feedback with Applications to Large Language Models. Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026. Paper
  • Aadirupa Saha, A. Wagde, and B. Kveton. LLM-as-Judge on a Budget. arXiv, 2026. arXiv
Adaptive diagnostic evaluation
  • Xiangxiang Dai et al. A Multi-Agent Conversational Bandit Approach to Online Evaluation and Selection of User-Aligned LLM Responses. AAAI, 2026. Paper
  • Deng Pan et al. Context Attribution with Multi-Armed Bandit Optimization. Findings of ACL, 2026. arXiv
  • Akshay Krishnamurthy et al. Can large language models explore in-context? NeurIPS, 2024. arXiv
  • Sweta Karlekar et al. Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. arXiv, 2026. arXiv

LLM-Enhanced Bandits

Representation — 8 research streams · 27 papers
Context Representation
Semantic context encoding
  • Ali Baheri and Cecilia O. Alm. LLMs-augmented Contextual Bandit. arXiv, 2023. arXiv
  • Xiaoqiang Lin et al. Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers. ICML, 2024. arXiv
  • Vikranth Dwaracherla et al. Efficient Exploration for LLMs. ICML, 2024. arXiv
  • Kaan Gönç et al. User Feedback-based Online Learning for Intent Classification. ICMI, 2023. Paper
Task-adapted context representation
  • Xinyuan Wang et al. MixLLM: Dynamic Routing in Mixed Large Language Models. NAACL, 2025. arXiv
  • Nicole Cho et al. No One Size Fits All: QueryBandits for Hallucination Mitigation. arXiv, 2026. arXiv
  • Hanzhuo Tan et al. Prompt-Based Code Completion via Multi-Retrieval Augmented Generation. ACM TOSEM, 2026. Paper
  • Xiaqiang Tang et al. Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs. AAAI, 2025. arXiv
  • Xinle Wu and Yao Lu. Reward Model Routing in Alignment. ICLR, 2026. arXiv
Stateful and trajectory-aware representation
  • Xiang Li et al. ALSO: Adversarial Online Strategy Optimization for Social Agents. arXiv, 2026. arXiv
  • Sheldon Yu et al. OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents. arXiv, 2026. arXiv
  • Yuqi Tang et al. SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition. arXiv, 2026. arXiv
Objective- and modality-aware representation
  • Kaiyu Huang et al. UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling. arXiv, 2026. arXiv
  • Zeyu Zhang et al. Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing. arXiv, 2026. arXiv
  • Xianzhi Zhang et al. Adapter-Augmented Bandits for Online Multi-Constrained Multi-Modal Inference Scheduling. arXiv, 2026. arXiv
Action Modeling
Semantic action representation
  • Zhaoxuan Wu et al. Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars. NeurIPS, 2024. arXiv
  • Donghao Li et al. Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits. arXiv, 2026. arXiv
  • Chao-Kai Chiang, Takashi Ishida, and Masashi Sugiyama. LLM Routing with Dueling Feedback. arXiv, 2025. arXiv
  • Kaiyu Huang et al. UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling. arXiv, 2026. arXiv
Consequence-based action similarity
  • Haruka Kiyohara et al. Prompt Optimization with Logged Bandit Data. arXiv, 2025. arXiv
  • Haruka Kiyohara et al. An Off-Policy Learning Approach for Steering Sentence Generation towards Personalization. RecSys, 2025. Paper
Relational and structured action modeling
  • Van Dai Do et al. SPaCe: Unlocking Sample-Efficient Large Language Models Training With Self-Pace Curriculum Learning. Findings of ACL, 2026. Paper
  • Darrien M. McKenzie, Nicklas Hansen, and Xiaolong Wang. Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models. arXiv, 2026. arXiv
  • Zhi Hong et al. MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks. arXiv, 2026. arXiv
Dynamic action-space construction
  • Shion Ishikawa et al. Progressive Content Refinement with Decaying Reward Joint LinUCB. arXiv, 2026. arXiv
  • Zelin He et al. ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL. arXiv, 2026. arXiv
  • Sweta Karlekar et al. Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. arXiv, 2026. arXiv
  • Junke Zhang et al. Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking. arXiv, 2026. arXiv
Learning — 8 research streams · 20 papers
Warm Start
Synthetic-interaction pretraining
  • Parand Alamdari, Yanshuai Cao, and Kevin Wilson. Jump Starting Bandits with LLM-Generated Prior Knowledge. EMNLP, 2024. arXiv
  • Adam Bayley et al. Jump Start or False Start? A Theoretical and Empirical Evaluation of LLM-initialized Bandits. TMLR, 2026. arXiv
Prior-based initialization
  • Qing Feng et al. LLM-Informed Bayesian Content Exploration in Ultra-Recency Recommendation. SIGIR, 2026. Paper
  • E. Lee et al. LLM-Derived Priors for Thompson Sampling in Cold-Start Comment Recommendation. arXiv, 2026. arXiv
  • Xinle Wu and Yao Lu. Reward Model Routing in Alignment. ICLR, 2026. arXiv
Guided initialization and early interaction
  • Shaohua Duan et al. Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization. ACL, 2026. arXiv
  • Dingyang Chen, Qi Zhang, and Yinglun Zhu. Efficient Sequential Decision Making with Large Language Models. EMNLP, 2024. arXiv
  • Sheldon Yu et al. OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents. arXiv, 2026. arXiv
Reward Estimation
LLM-based outcome modeling
  • Nicolò Felicioni et al. On the Importance of Uncertainty in Decision-Making with Large Language Models. TMLR, 2024. arXiv
  • Jiahang Sun et al. Large Language Model-Enhanced Multi-Armed Bandits. ACL, 2026. arXiv
  • Uljad Berdica et al. When Do We Need LLMs? A Diagnostic for Language-Driven Bandits. arXiv, 2026. arXiv
Proxy augmentation and correction
  • Parand Alamdari, Yanshuai Cao, and Kevin Wilson. Jump Starting Bandits with LLM-Generated Prior Knowledge. EMNLP, 2024. arXiv
  • M.N. Pershin et al. Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits. arXiv, 2026. arXiv
  • Tianyi Ma et al. Best-Arm Identification with Generative Proxy. arXiv, 2026. arXiv
  • Ruicheng Ao et al. Best Arm Identification with LLM Judges and Limited Human. arXiv, 2026. arXiv
Language-to-reward construction
  • Nikhil Behari et al. A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health. NeurIPS, 2024. arXiv
  • Shresth Verma et al. Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards. GameSec, 2025. arXiv
Semantic reward surrogates
  • Nicole Cho et al. No One Size Fits All: QueryBandits for Hallucination Mitigation. arXiv, 2026. arXiv
  • Linfeng Du et al. Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization. ACL, 2026. arXiv
  • Sweta Karlekar et al. Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. arXiv, 2026. arXiv
Environment Modeling
Language-mediated posterior modeling
  • Dilip Arumugam and Thomas L. Griffiths. Toward Efficient Exploration by Large Language Model Agents. arXiv, 2025. arXiv
Decision — 8 research streams · 13 papers
Exploration
Uncertainty-aware exploration
  • Nicolò Felicioni et al. On the Importance of Uncertainty in Decision-Making with Large Language Models. TMLR, 2024. arXiv
  • Jiahang Sun et al. Large Language Model-Enhanced Multi-Armed Bandits. ACL, 2026. arXiv
  • Uljad Berdica et al. When Do We Need LLMs? A Diagnostic for Language-Driven Bandits. arXiv, 2026. arXiv
Semantic action-space restriction
  • Keegan Harris and Aleksandrs Slivkins. Should You Use Your Large Language Model to Explore or Exploit? arXiv, 2025. arXiv
Direct exploration control
  • J. de Curtò et al. LLM-Informed Multi-Armed Bandit Strategies for Non-Stationary Environments. Electronics, 2023. Paper
  • Sanxing Chen et al. When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training. arXiv, 2025. arXiv
Language-mediated model-based exploration
  • Dilip Arumugam and Thomas L. Griffiths. Toward Efficient Exploration by Large Language Model Agents. arXiv, 2025. arXiv
Action Selection
Direct LLM action selection
  • Jawad Hazime and Junaid Farooq. Evaluation of LLM Powered Agentic AI for Solving Multi-Arm Bandit Problems. IEEE COINS, 2025. Paper
  • Keegan Harris and Aleksandrs Slivkins. Should You Use Your Large Language Model to Explore or Exploit? arXiv, 2025. arXiv
  • Fanzeng Xia et al. Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of ACL, 2025. Paper
Confidence-gated and validated selection
  • Junyu Cao et al. LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning. arXiv, 2026. arXiv
  • Fanzeng Xia et al. Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of ACL, 2025. Paper
Candidate generation and restricted selection
  • Keegan Harris and Aleksandrs Slivkins. Should You Use Your Large Language Model to Explore or Exploit? arXiv, 2025. arXiv
  • Zichen Liu et al. Sample-Efficient Alignment for LLMs. arXiv, 2024. arXiv
Proxy- and diagnosis-guided selection
  • Tianyi Ma et al. Best-Arm Identification with Generative Proxy. arXiv, 2026. arXiv
  • Geremy Loachamín Suntaxi et al. Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection. arXiv, 2026. arXiv
Feedback — 4 research streams · 7 papers
Feedback Interpretation
Scalar and binary judging
  • Kexin Chu, Dawei Xiang, and Wei Zhang. Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents. arXiv, 2026. arXiv
  • Aditya Ramesh et al. Efficient Jailbreak Attack sequences on Large Language Models via Multi-Armed Bandit-based Context switching. ICLR, 2025. Paper
Pairwise preference interpretation
  • Yuanchen Wu et al. LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization. Findings of ACL, 2026. arXiv
Semantic feedback shaping and propagation
  • Son Nguyen, Xinyuan Liu, and Ransalu Senanayake. CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM. arXiv, 2026. arXiv
  • Shengbo Wang, Hong Sun, and Ke Li. Preference Is More than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback. AAAI, 2026. Paper
Structured diagnosis and attribution
  • Nikhil Behari et al. A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health. NeurIPS, 2024. arXiv
  • Geremy Loachamín Suntaxi et al. Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection. arXiv, 2026. arXiv

📄 Bibliography

We provide references.bib for the bibliographic records used in our survey. We include the 153-study corpus together with supporting background and methodological references, and we identify the included corpus in taxonomy.yaml.

🗂️ Repository Structure

.
├── README.md                # Methodology and literature navigation
├── references.bib           # Survey and supporting references
├── taxonomy.yaml            # Reusable corpus classification
└── LICENSE

🤝 Updates / Contributing

We freeze our Survey Corpus Snapshot at 153 studies through August 15, 2026. We may add later work under Post-Survey Updates, clearly separated from our frozen snapshot.

We welcome suggestions for missing or newly published work through issues or pull requests. We ask contributors to include an authoritative citation, a short explanation of the substantive Bandit–LLM interaction, and a proposed taxonomy location. We do not include work that mentions bandits or LLMs only as background.

📝 Citation

If you use this collection, please cite our accompanying survey. We will add the complete publication metadata here when our paper is publicly available.

Contributors

bucky1119

7 commits

Languages

TeX

100.0%