Awesome-HCI (Ubiquitous, LLM, MLLM, Agent, RAG, Embodied-AI, RLHF)
Python
26
0 commits
updated Mar 8, 2026
A curated collection of research papers on HCI, LLM, MLLM, Agent, RAG, Agentic-RL, and Embodied AI (2021–present).
[Jan 2025] Added new sections: Agentic-RL and MLLM. Regular updates resumed.
python -m pip install -e . # Install CLI
paper add 2312.00752 LLM -t "llm, mamba" # Add paper
paper search transformer -t IMU # Search
paper stats # Statistics
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | Evaluating the Usage of African-American Vernacular English in Large Language Models | Deja Dunlap, et al. | LLM | cs.CL, cs.HC | 2026.02 | |
| arXiv(v1) 2026 | GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents | Shuofei Qiao, et al. | GUI agent, visual grounding, tool use, perception | cs.CV, cs.AI | 2026.01 | |
| arXiv(v1) 2025 (WACV 2026) | AFRAgent: An Adaptive Feature Renormalization Based High Resolution Aware GUI agent | Neeraj Anand, et al. | GUI agent, VLM, smartphone automation, multimodal | cs.CV | WACV 2026 | 2025.12 |
| arXiv(v2) 2025 | Mobile-Agent-v3: Fundamental Agents for GUI Automation | Junyang Wang, et al. | GUI agent, mobile, smartphone automation, VLM | cs.CV, cs.AI | 2025.08 | |
| arXiv(v6) 2025 | LLaVA-CoT: Let Vision Language Models Reason Step-by-Step | Guowei Xu, et al. | LLM, GUI agent, survey, computer use | cs.CV | 17 pages, ICCV 2025 | 2025.07 |
| arXiv(v3) 2025 | Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs | Kang Zhao, et al. | LLM, GUI agent, interface | cs.LG, cs.AI | 2025.06 | |
| arXiv(v3) 2025 | Lai Loss: A Novel Loss for Gradient Control | YuFei Lai | LLM, agent, interface, UI | cs.LG | The experiment in this article is not very rigorous and may require further testing for its effectiveness | 2025.05 |
| arXiv(v3) 2025 | SignLLM: Sign Language Production Large Language Models | Sen Fang, et al. | LLM, agent, user interface, HCI, interaction | cs.CV, cs.CL | website at https://signllm.github.io/ | 2025.04 |
| arXiv(v1) 2024 | ShowUI: One Vision-Language-Action Model for GUI Visual Agent | Kevin Qinghong Lin, et al. | LLM, UI, vision-language, GUI | cs.CV, cs.AI, cs.CL, cs.HC | Technical Report. Github: https://github.com/showlab/ShowUI | 2024.11 |
| arXiv(v1) 2024 | A Scalable Communication Protocol for Networks of Large Language Models | Samuele Marro, et al. | LLM, agent, GUI, computer use, interface | cs.AI, cs.LG | 2024.10 | |
| arXiv(v1) 2024 | In-Band Full-Duplex MIMO Systems for Simultaneous Communications and Sensing: Challenges, Methods, and Future Perspectives | Besma Smida, et al. | LLM, GUI agent, computer use | cs.IT, cs.ET, eess.SP | 12 pages, 5 figures, White Paper to appear at IEEE SPM | 2024.10 |
| arXiv(v2) 2024 | Massively parallel CMA-ES with increasing population | David Redon, et al. | LLM, agent, HCI, user interface | cs.DC | 2024.10 | |
| arXiv(v1) 2024 | OS-ATLAS: A Foundation Action Model for Generalist GUI Agents | Zhiyong Wu, et al. | LLM, agent, computer use, GUI, foundation model | cs.CL, cs.CV, cs.HC | 2024.10 | |
| arXiv(v1) 2024 | OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing Images | Jiaqi Zhao, et al. | LLM, agent, GUI, interface | cs.CV | The paper is accepted by IEEE Transactions on Geoscience and Remote Sensing (TGRS) | 2024.09 |
| arXiv(v1) 2024 | Unlocking the Power of Environment Assumptions for Unit Proofs | Siddharth Priya, et al. | LLM, GUI, agent, survey | cs.SE, cs.PL | SEFM 2024 | 2024.09 |
| arXiv(v2) 2024 (ACL 2025) | GUICourse: From General Vision Language Models to Versatile GUI Agents | Wentong Chen, et al. | GUI agent, VLM, training data, OCR, grounding | cs.CV, cs.CL, cs.HC | ACL 2025 | 2024.06 |
| arXiv(v1) 2024 | Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs | Keen You, et al. | ui, mllm, benchmark, any-resolution | cs.CV, cs.CL, cs.HC | 2024.04 | |
| arXiv(v1) 2024 | Are You Being Tracked? Discover the Power of Zero-Shot Trajectory Tracing with LLMs! | Huanqi Yang, et al. | iot, imu, cot, prompt | cs.CL, cs.AI, cs.HC, cs.LG | 2024.03 | |
| arXiv(v1) 2024 | Design2Code: How Far Are We From Automating Front-End Engineering? | Chenglei Si, et al. | llm, auto, google | cs.CL, cs.CV, cs.CY | 2024.03 | |
| arXiv(v2) 2023 | The Good, The Bad, and Why: Unveiling Emotions in Generative AI | Cheng Li, et al. | emotion, prompt, attack, decode | cs.AI, cs.CL, cs.HC | extension of Large language models understand and can be enhanced by emotional stimuli | 2023.12 |
| arXiv(v1) (NIPS23) | Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias | Yue Yu, et al. | synthetic data generation | cs.CL, cs.AI, cs.LG | arXiv(v2) 2023.10 | |
| arXiv(v1) 2023 | Multimodal Foundation Models: From Specialists to General-Purpose Assistants | Chunyuan Li, et al. | survey | cs.CV, cs.CL | 2023.09 | |
| arXiv(v7) 2023 | Attention Is All You Need | Ashish Vaswani, et al. | arxiv | cs.CL, cs.LG | 15 pages, 5 figures | 2023.08 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation | Zhengren Wang, et al. | RAG | cs.CV, cs.CL | 2026.02 | |
| arXiv(v1) 2026 | CLFEC: A New Task for Unified Linguistic and Factual Error Correction in paragraph-level Chinese Professional Writing | Jian Kai, et al. | RAG | cs.CL | 2026.02 | |
| arXiv(v1) 2026 | CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era | Zhengqing Yuan, et al. | RAG | cs.CL, cs.DL | 2026.02 | |
| arXiv(v1) 2026 | CiteLLM: An Agentic Platform for Trustworthy Scientific Reference Discovery | Mengze Hong, et al. | RAG | cs.CL, cs.IR | Accepted by TheWebConf 2026 Demo Track | 2026.02 |
| arXiv(v2) 2026 | MoDora: Tree-Based Semi-Structured Document Analysis System | Bangrui Xu, et al. | RAG | cs.IR, cs.AI, cs.CL, cs.DB, cs.LG | Extension of our SIGMOD 2026 paper. Please refer to source code available at https://github.com/weAIDB/MoDora | 2026.02 |
| arXiv(v1) 2026 | Search-P1: Path-Centric Reward Shaping for Stable and Efficient Agentic RAG Training | Tianle Xia, et al. | RAG | cs.CL, cs.IR, cs.LG | 2026.02 | |
| arXiv(v1) 2026 | TCM-DiffRAG: Personalized Syndrome Differentiation Reasoning Method for Traditional Chinese Medicine based on Knowledge Graph and Chain of Thought | Jianmin Li, et al. | RAG | cs.CL, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | TRIZ-RAGNER: A Retrieval-Augmented Large Language Model for TRIZ-Aware Named Entity Recognition in Patent-Based Contradiction Mining | Zitong Xu, et al. | RAG | cs.CL, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning | Chris Samarinas, et al. | RAG | cs.CL, cs.IR | 2026.02 | |
| arXiv(v1) 2026 | Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators | Zhengyang Su, et al. | RAG | cs.IR, cs.CL, cs.LG | 14 pages, 4 figures | 2026.02 |
| arXiv(v1) 2026 | L-RAG: Balancing Context and Retrieval with Entropy-Based Lazy Loading | Authors TBD | RAG, lazy loading, entropy, context | cs.CL, cs.IR | 2026.01 | |
| arXiv(v1) 2025 | RAGLens: Toward Faithful RAG with Sparse Autoencoders | Authors TBD | RAG, hallucination, faithfulness, detection | cs.CL | 2025.12 | |
| arXiv(v1) 2025 | CDTA: Cross-Document Topic-Aligned Chunking for RAG | Authors TBD | RAG, chunking, cross-document, topic alignment | cs.CL, cs.IR | 0.93 faithfulness on HotpotQA | 2025.11 |
| arXiv(v1) 2025 | Agentic RAG for Fintech: Design and Evaluation | Authors TBD | RAG, agentic, fintech, query reformulation | cs.CL, cs.IR | 2025.10 | |
| arXiv(v1) 2025 | Practical Code RAG at Scale: Task-Aware Retrieval Design | Authors TBD | RAG, code, retrieval, hybrid, dense | cs.CL, cs.IR | BM25 + dense hybrid | 2025.10 |
| arXiv(v1) 2025 | A Systematic Review of Key RAG Systems: Progress, Gaps, and Future Directions | Authors TBD | RAG, survey, systematic review, knowledge base | cs.CL, cs.IR | 2025.07 | |
| arXiv(v1) 2025 | Late Chunking: Contextual Chunk Embeddings for RAG | Authors TBD | RAG, chunking, embedding, contextual | cs.CL, cs.IR | Updated July 2025 | 2025.07 |
| arXiv(v1) 2025 | GraphRAG-Bench: When to Use Graphs in RAG | Authors TBD | RAG, graph, benchmark, evaluation | cs.CL, cs.IR | 2025.06 | |
| arXiv(v1) 2025 | RAG Survey: Architectures, Enhancements, and Robustness Frontiers | Authors TBD | RAG, survey, architecture, robustness | cs.CL, cs.IR | 2025.06 | |
| arXiv(v1) 2025 | Rethinking Chunk Size for Long-Document Retrieval: Multi-Dataset Analysis | Authors TBD | RAG, chunking, chunk size, retrieval | cs.CL, cs.IR | 64-1024 tokens optimal | 2025.05 |
| arXiv(v1) 2025 | A Survey of Multimodal Retrieval-Augmented Generation | Zihan Zhao, et al. | multimodal, RAG, retrieval, survey, vision-language | cs.CV, cs.CL | 2025.04 | |
| arXiv(v1) 2025 | RAG Evaluation in the Era of LLMs: A Comprehensive Survey | Authors TBD | RAG, evaluation, benchmark, LLM | cs.CL, cs.IR | 2025.04 | |
| arXiv(v1) 2025 | HiRAG: Retrieval-Augmented Generation with Hierarchical Knowledge | Authors TBD | RAG, hierarchical, knowledge, indexing | cs.CL, cs.IR | 2025.03 | |
| arXiv(v1) 2025 | Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation | Zihan Wang, et al. | multimodal, RAG, retrieval, survey, LLM | cs.CL, cs.IR | 2025.02 | |
| arXiv(v1) 2025 | GFM-RAG: Graph Foundation Model for Retrieval Augmented Generation | Authors TBD | RAG, graph, foundation model, knowledge graph | cs.CL, cs.IR | 8M params, 60 KGs, 14M triples | 2025.02 |
| arXiv(v1) 2025 | KG2RAG: Knowledge Graph-Guided Retrieval Augmented Generation | Authors TBD | RAG, knowledge graph, retrieval, fact-level | cs.CL, cs.IR | 2025.02 | |
| arXiv(v1) 2025 | RAG-Fusion: Query Expansion and Multi-Source Retrieval | Authors TBD | RAG, query expansion, multi-source, fusion | cs.CL, cs.IR | 2025.02 | |
| arXiv(v1) 2025 | Vendi-RAG: Adaptively Trading Off Diversity and Quality in RAG | Authors TBD | RAG, diversity, quality, adaptive | cs.CL, cs.IR | 2025.02 | |
| arXiv(v1) 2025 | Agentic RAG: A Survey | Authors TBD | RAG, agentic, survey, LLM agent | cs.CL, cs.IR | 2025.01 | |
| arXiv(v1) 2025 | CG-RAG: Citation Graph RAG for Research Question Answering | Authors TBD | RAG, citation graph, research QA | cs.CL, cs.IR | 2025.01 | |
| arXiv(v1) 2025 | ChunkRAG: Novel Context-Aware Chunking for RAG Systems | Authors TBD | RAG, chunking, context-aware, retrieval | cs.CL, cs.IR | 2025.01 | |
| arXiv(v1) 2024 | LLM-Augmented Retrieval: Enhancing Retrieval Models Through Language Models and Doc-Level Embedding | relevant query, doc-Level embedding, embedding-based retrieval, dense retrieval | 2024.04 | |||
| arXiv(v6) 2024 | Health-LLM: Personalized Retrieval-Augmented Disease Prediction System | Qinkai Yu, et al. | RAG, XGBoost, AutoML | cs.CL | 2024.03 | |
| arXiv(v1) 2024 | CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models | retrieval-augmented generation, large language models, evaluation | 2024.02 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | "Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems | Xinfeng Li, et al. | LLM, agent | cs.HC, cs.AI, cs.CR, cs.SI | 2026.02 | |
| arXiv(v1) 2026 | AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning | Yutong Wang, et al. | LLM, agent | cs.AI, cs.CL | 2026.02 | |
| arXiv(v1) 2026 | E3VA: Enhancing Emotional Expressiveness in Virtual Conversational Agents | Abhishek Kulkarni, et al. | LLM, agent | cs.HC | 5 pages | 2026.02 |
| arXiv(v1) 2026 | ESAA: Event Sourcing for Autonomous Agents in LLM-Based Software Engineering | Elzo Brito dos Santos Filho | LLM, agent | cs.AI | 13 pages, 1 figure, 4 tables. Includes 5 technical appendices | 2026.02 |
| arXiv(v1) 2026 | From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems | Yawen Wang, et al. | LLM, agent | cs.AI, cs.SE | 2026.02 | |
| arXiv(v1) 2026 | HotelQuEST: Balancing Quality and Efficiency in Agentic Search | Guy Hadad, et al. | LLM, agent | cs.IR, cs.AI | To be published in EACL 2026 | 2026.02 |
| arXiv(v1) 2026 | Hybrid LLM-Embedded Dialogue Agents for Learner Reflection: Designing Responsive and Theory-Driven Interactions | Paras Sharma, et al. | LLM, agent | cs.HC, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | ProductResearch: Training E-Commerce Deep Research Agents via Multi-Agent Synthetic Trajectory Distillation | Jiangyuan Wang, et al. | LLM, agent | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | PseudoAct: Leveraging Pseudocode Synthesis for Flexible Planning and Action Control in Large Language Model Agents | Yihan, et al. | LLM, agent | cs.AI, eess.SY | 2026.02 | |
| arXiv(v1) 2026 | SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems | Jialiang Fan, et al. | LLM, agent | cs.RO, cs.AI | 12 pages, 6 figures | 2026.02 |
| arXiv(v1) 2026 | The Auton Agentic AI Framework | Sheng Cao, et al. | LLM, agent | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training | Authors TBD | agent, web, GUI, training, environment | cs.CV, cs.AI | 600 tasks across websites | 2026.01 |
| arXiv(v1) 2026 | LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities | Authors TBD | code agent, software engineering, survey | cs.SE, cs.AI | 2026.01 | |
| arXiv(v1) 2025 | Beyond Task Completion: Assessment Framework for Agentic AI Systems | Authors TBD | LLM agent, evaluation, framework, benchmark | cs.AI, cs.CL | 2025.12 | |
| arXiv(v1) 2025 | DeepCode: Open Agentic Coding | Authors TBD | code agent, agentic coding, open source | cs.SE, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | LongVideoAgent: Multi-Agent Reasoning with Long Videos | Jingyi Zhang, et al. | multi-agent, video understanding, LLM, reasoning | cs.CV, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | MAR: Multi-Agent Reflexion Improves Reasoning Abilities in LLMs | Xinyuan Lu, et al. | multi-agent, LLM, reflexion, reasoning, debate | cs.CL, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | SWE-RL: Training Superintelligent Software Agents through Self-Play | Authors TBD | code agent, self-play, RL, software engineering | cs.SE, cs.LG | 2025.12 | |
| arXiv(v1) 2025 | Agent0: Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning | Authors TBD | LLM agent, self-evolving, tool use, reasoning | cs.AI, cs.CL | 18% improvement on math, 24% on general | 2025.11 |
| arXiv(v1) 2025 | Building Browser Agents: Architecture, Security, and Practical Solutions | Authors TBD | agent, browser, security, architecture | cs.AI, cs.CL | 2025.11 | |
| arXiv(v1) 2025 | ReMA: Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation | Zhiwei Zhang, et al. | multi-agent, LLM, reasoning, lazy agent, deliberation | cs.AI, cs.CL | 2025.11 | |
| arXiv(v1) 2025 | BrowserAgent: Web Agents with Human-Inspired Browsing Actions | Authors TBD | agent, browser, web, automation | cs.AI, cs.CL | 2025.10 | |
| arXiv(v1) 2025 | Dark Patterns Impact on LLM-Based Web Agents | Authors TBD | agent, web, dark patterns, decision making | cs.HC, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | Architecting Resilient LLM Agents | Authors TBD | LLM agent, resilience, architecture | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | A Survey on Code Generation with LLM-based Agents | Authors TBD | code agent, LLM, code generation, survey | cs.SE, cs.AI | 2025.08 | |
| arXiv(v1) 2025 | LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios | Bingxi Zhao, et al. | LLM, agent, reasoning, survey, framework | cs.AI, cs.CL | 2025.08 | |
| arXiv(v1) 2025 | MCP-Bench: Benchmarking Tool-Using LLM Agents | Authors TBD | agent, benchmark, tool use, MCP | cs.AI, cs.CL | 2025.08 | |
| arXiv(v4) 2025 | Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction | Harshitha Manoj, et al. | embodied AI, robot, LLM, VLM, autonomy, survey | cs.RO, cs.AI | 2025.08 | |
| arXiv(v1) 2025 | Understanding Tool-Integrated Reasoning | Heng Lin, et al. | LLM, agent, tool use, reasoning, TIR | cs.LG, cs.AI, stat.ML | 2025.08 | |
| arXiv(v2) 2025 | Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools | Junde Wu, et al. | LLM, agent, agentic reasoning, tool use, framework | cs.AI, cs.CL | ACL 2025 | 2025.07 |
| arXiv(v3) 2025 | Embodied AI Agents: Modeling the World | Pascale Fung, et al. | embodied AI, world model, VLM, robot, avatar | cs.AI | Meta AI | 2025.07 |
| arXiv(v1) 2025 | Routine: A Structural Planning Framework for LLM Agent System in Enterprise | Authors TBD | LLM agent, planning, enterprise, framework | cs.AI, cs.CL | 2025.07 | |
| arXiv(v1) 2025 | Toward a Theory of Agents as Tool-Use Decision-Makers | Authors TBD | LLM agent, tool use, theory, decision making | cs.AI, cs.CL | 2025.06 | |
| arXiv(v1) 2025 | WebRL: Training LLM Web Agents via Self-Evolving Curriculum RL | Authors TBD | agent, web, RL, curriculum, self-evolving | cs.LG, cs.AI | 2025.05 | |
| arXiv(v1) 2025 | AgentRewardBench: Benchmark for LLM Judges in Web Agent Evaluation | Authors TBD | agent, benchmark, web, evaluation, reward | cs.AI, cs.CL | 1,302 trajectories | 2025.04 |
| arXiv(v1) 2025 | From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review | Mohamed Amine Ferrag, et al. | LLM, agent, survey, autonomous, comprehensive review | cs.AI, cs.LG | 2025.04 | |
| arXiv(v1) 2024 | Simultaneous identification of the parameters in the plasticity function for power hardening materials : A Bayesian approach | Salih Tatar, et al. | LLM, agent, computer use, evaluation | math.NA, math.AP | 2024.12 | |
| arXiv(v2) 2024 | Feasibility Consistent Representation Learning for Safe Reinforcement Learning | Zhepeng Cen, et al. | LLM, agent, UI, LAUI, interface | cs.LG | ICML 2024 | 2024.06 |
| arXiv(v1) 2024 | DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows | Ajay Patel, et al. | pipeline, liarbry, generation | cs.CL, cs.LG | 2024.02 | |
| arXiv(v1) 2024 | More Agents Is All You Need | Junyou Li, et al. | multi agent, vote, task | cs.CL, cs.AI, cs.LG | 2024.02 | |
| arXiv(v1) (NIPS23) | HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face | Yongliang Shen, et al. | hugging face, API | cs.CL, cs.AI, cs.CV, cs.LG | 2023.12 | |
| arXiv(v1) (NIPS23) | CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society | Guohao Li, et al. | role play, autonomous, user&assistant | cs.AI, cs.CL, cs.CY, cs.LG, cs.MA | 2023.11(v2) | |
| arXiv(v2) 2023 | MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models | Dingyao Yu, et al. | pipeline | cs.CL, cs.MM, eess.AS | 2023.10 | |
| arXiv(v2) 2023 | VOYAGER: An Open-Ended Embodied Agent with Large Language Models | Guanzhi Wang, et al. | multi, autonomous, microcraft, game | cs.AI, cs.LG | 2023.10 | |
| arXiv(v3) 2023 | The Rise and Potential of Large Language Model Based Agents: A Survey | Zhiheng Xi, et al. | survey, github paper list | cs.AI, cs.CL | 2023.09 | |
| arXiv(v3) 2023 | A Survey on Large Language Model based Autonomous Agents | Lei Wang, et al. | survey, autonomous | cs.AI, cs.CL | Latest version is v4(2024.03), double columns. But v3(2023.09) single columns is easy to read. | |
| arXiv(v1) (ICLR24) | MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework | Sirui Hong, et al. | autonomous system, SOP, multi-agent, framework | cs.AI, cs.MA |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation | Weinan Dai, et al. | LLM, agentic-RL | cs.LG, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization | Zeyuan Liu, et al. | LLM, agentic-RL | cs.LG, cs.AI | Accepted to ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning | Zehao Li, et al. | LLM, agentic-RL | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | RF-Agent: Automated Reward Function Design via Language Agent Tree Search | Ning Gao, et al. | LLM, agentic-RL | cs.AI, cs.LG | 39 pages, 9 tables, 11 figures, Project page see https://github.com/deng-ai-lab/RF-Agent | 2026.02 |
| arXiv(v1) 2026 | RUMAD: Reinforcement-Unifying Multi-Agent Debate | Chao Wang, et al. | LLM, agentic-RL | cs.AI | 13 pages, 3 figures | 2026.02 |
| arXiv(v1) 2026 | Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance | Yanwei Ren, et al. | LLM, agentic-RL | cs.AI, cs.CL | 2026.02 | |
| arXiv(v2) 2026 | Regularized Online RLHF with Generalized Bilinear Preferences | Junghyun Lee, et al. | LLM, agentic-RL | cs.LG, stat.ML | 43 pages, 1 table (ver2: more colorful boxes, fixed some typos) | 2026.02 |
| arXiv(v1) 2026 | RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models | Daniel Yang, et al. | LLM, agentic-RL | cs.LG, cs.AI, cs.CL | 2026.02 | |
| arXiv(v1) 2026 | Agent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems Over Extended Interactions | Abhishek Rath | multi-agent, behavioral degradation, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v5) 2026 | AgentOrchestra: Orchestrating Multi-Agent Intelligence with the Tool-Environment-Agent(TEA) Protocol | Wentao Zhang, et al. | hierarchical multi-agent, task solving, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity | Doyoung Kim, et al. | evaluation, benchmark, API complexity, LLM, agent | cs.CL, cs.AI | 26 pages | 2026.01 |
| arXiv(v1) 2026 | Beyond Rule-Based Workflows: An Information-Flow-Orchestrated Multi-Agents Paradigm via Agent-to-Agent Communication from CORAL | Xinxing Ren, et al. | information flow, multi-agent communication, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v2) 2026 | Cochain: Balancing Insufficient and Excessive Collaboration in LLM Agent Workflows | Jiaxing Zhao, et al. | chain-of-collaboration, multi-agent, LLM, agent | cs.CL | 35 pages, 23 figures | 2026.01 |
| arXiv(v2) 2026 | Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes | Abhijnan Nath, et al. | alignment, multi-agent coordination, LLM, agent | cs.CL, cs.AI, cs.LG | This submission is a new version of arXiv:2509.05882v1. with a substantially revised experimental pipeline and new metrics. In particular, collaborator agents are now instantiated independently via separate API calls, rather than generated autoregressively by a single agent. All experimental results are new. Accepted as an extended abstract at AAMAS 2026 | 2026.01 |
| arXiv(v1) 2026 | DemMA: Dementia Multi-Turn Dialogue Agent with Expert-Guided Reasoning and Action Simulation | Yutong Song, et al. | dementia dialogue, multi-turn, medical, LLM, agent | cs.MA | 2026.01 | |
| arXiv(v1) 2026 | EduSim-LLM: An Educational Platform Integrating Large Language Models and Robotic Simulation for Beginners | Shenqi Lu, et al. | educational platform, robotics, simulation, LLM | cs.RO | 2026.01 | |
| arXiv(v1) 2026 | Game-Theoretic Lens on LLM-based Multi-Agent Systems | Jianing Hao, et al. | game theory, multi-agent, LLM, agent | cs.MA, cs.GT | 9 pages, 5 figures | 2026.01 |
| arXiv(v2) 2026 (Spotlight paper of NeurIPS 2025) | KARMA: Leveraging Multi-Agent LLMs for Automated Knowledge Graph Enrichment | Yuxing Lu, et al. | knowledge graph enrichment, multi-agent, LLM, agent | cs.CL, cs.AI, cs.CE, cs.DL | 24 pages, 3 figures, 2 tables | 2026.01 |
| arXiv(v1) 2026 | LLM-in-Sandbox Elicits General Agentic Intelligence | Daixuan Cheng, et al. | sandbox, agentic intelligence, LLM | cs.CL, cs.AI | Project Page: https://llm-in-sandbox.github.io | 2026.01 |
| arXiv(v2) 2026 | Lifelong Learning of Large Language Model based Agents: A Roadmap | Junhao Zheng, et al. | lifelong learning, continual learning, LLM, agent | cs.AI | Accepted to IEEE TPAMI | 2026.01 |
| arXiv(v1) 2026 | Nalar: An agent serving framework | Marco Laju, et al. | workflow serving, agent workflows, LLM, agent | cs.DC, cs.MA | 2026.01 | |
| arXiv(v1) 2026 | Orchestral AI: A Framework for Agent Orchestration | Alexander Roman, et al. | agent orchestration, framework, LLM, agent | cs.AI, astro-ph.IM, hep-ph | 17 pages, 3 figures. For more information visit https://orchestral-ai.com | 2026.01 |
| arXiv(v1) 2025 | PRL: Process Reward Learning Improves LLM Reasoning | Authors TBD | agentic RL, PRM, process reward, reasoning | cs.LG, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions | Aayush Gupta | reliability evaluation, stress testing, LLM, agent | cs.AI | 18 pages, 5 figures, 8 tables. Evaluates ReAct vs Reflexion across four tool-using domains with perturbation (epsilon) and fault-injection (lambda) stress testing; 1,280 total episodes | 2026.01 |
| arXiv(v2) 2026 | SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds | Jiawei Ren, et al. | simulation, autonomous agent, world model, LLM | cs.AI | 2026.01 | |
| arXiv(v2) 2026 | The Agentic Leash: Extracting Causal Feedback Fuzzy Cognitive Maps with LLMs | Akash Kumar Panda, et al. | causal extraction, fuzzy cognitive maps, LLM, agent | cs.AI, cs.CL, cs.HC, cs.IR | 15 figures | 2026.01 |
| arXiv(v2) 2026 | The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning | Qiguang Chen, et al. | chain-of-thought, reasoning topology, LLM | cs.CL, cs.AI | Preprint | 2026.01 |
| arXiv(v1) 2026 | The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs | Zibo Zhao, et al. | self-reflection, reinforcement learning, LLM, agent | cs.LG, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents | Alessio Buscemi, et al. | numerical coordination, multi-agent, LLM, agent | cs.MA, cs.AI | 2026.01 | |
| arXiv(v2) 2025 | DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems | Ming Ma, et al. | auto debugging, multi-agent, intervention, LLM, agent | cs.AI, cs.SE | 2025.12 | |
| arXiv(v1) 2025 | Enhancing Agentic RL with Progressive Reward Shaping and VSPO | Authors TBD | agentic RL, reward shaping, GRPO, tool use | cs.LG, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | From Word to World: Can Large Language Models be Implicit Text-based World Models? | Yixia Li, et al. | world model, text-based, LLM, agent | cs.CL | 2025.12 | |
| arXiv(v2) 2025 | ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving | Haoyuan Wu, et al. | tree-of-thought, reasoning, puzzle solving, LLM | cs.CL | 2025.12 | |
| arXiv(v2) 2025 | Understanding LLM Agent Behaviours via Game Theory: Strategy Recognition, Biases and Multi-Agent Dynamics | Trung-Kiet Huynh, et al. | game theory, agent behavior, LLM, agent | cs.MA, cs.AI, cs.GT, cs.LG, math.DS | 2025.12 | |
| arXiv(v2) 2025 | Large Language Model-based Data Science Agent: A Survey | Ke Chen, et al. | data science, agent, LLM | cs.AI | 2025.11 | |
| arXiv(v1) 2025 | MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism | Shulin Liu, et al. | multi-agent, agentic RL, reasoning, LLM, pipeline parallelism | cs.AI | 10 pages | 2025.11 |
| arXiv(v1) 2025 | AGENTRL: Scaling Agentic Reinforcement Learning | Authors TBD | agentic RL, RLHF, LLM agent, scaling | cs.LG, cs.AI | Outperforms GPT-5 and Claude-Sonnet-4 | 2025.10 |
| arXiv(v2) 2025 | BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks | Sagnik Anupam, et al. | web navigation, evaluation, real-world, LLM, agent | cs.AI, cs.LG | 2025.10 | |
| arXiv(v1) 2025 | Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning | Marwa Abdulhai, et al. | persona simulation, multi-turn RL, LLM, agent | cs.CL, cs.AI | 2025.10 | |
| arXiv(v3) 2025 | DS-STAR: Data Science Agent via Iterative Planning and Verification | Jaehyun Nam, et al. | data science, iterative planning, verification, LLM, agent | cs.AI | 2025.10 | |
| arXiv(v1) 2025 | Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments | Crystal Qian, et al. | social experiments, human-AI interaction, LLM, agent | cs.HC, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | Demystifying Reinforcement Learning in Agentic Reasoning | Zhaochen Yu, et al. | agentic RL, reasoning, LLM, agent | cs.CL | Code and models: https://github.com/Gen-Verse/Open-AgentRL | 2025.10 |
| arXiv(v3) 2025 | DoctorAgent-RL: A Multi-Agent Collaborative Reinforcement Learning System for Multi-Turn Clinical Dialogue | Yichun Feng, et al. | clinical dialogue, multi-agent RL, medical, LLM, agent | cs.CL | 2025.10 | |
| arXiv(v1) 2025 | GEM: A Gym for Agentic LLMs | Zichen Liu, et al. | training environment, agentic LLM, gym, agent | cs.LG, cs.AI, cs.CL | 2025.10 | |
| arXiv(v3) 2025 | MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research | Hui Chen, et al. | ML research, evaluation, benchmark, LLM, agent | cs.LG, cs.AI, cs.CL | 49 pages, 9 figures. Accepted by NeurIPS 2025 D&B Track | 2025.10 |
| arXiv(v1) 2025 | Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents | Reid T. Johnson, et al. | natural language tool calling, LLM, agent | cs.CL | 31 pages, 7 figures | 2025.10 |
| arXiv(v1) 2025 | On Designing Effective RL Reward at Training Time for LLM Reasoning | Authors TBD | agentic RL, reward design, reasoning, training | cs.LG, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | RLSR: Reinforcement Learning with Supervised Reward | Authors TBD | agentic RL, supervised reward, instruction following | cs.LG, cs.AI | 2025.10 | |
| arXiv(v2) 2025 | SimuRA: A World-Model-Driven Simulative Reasoning Architecture for General Goal-Oriented Agents | Mingkai Deng, et al. | world model, simulative reasoning, goal-oriented, LLM, agent | cs.AI, cs.CL, cs.LG, cs.RO | This submission has been updated to adjust the scope and presentation of the work | 2025.10 |
| arXiv(v2) 2025 (Am. Statist. (2025) 1-14) | A Survey on Large Language Model-based Agents for Statistics and Data Science | Maojun Sun, et al. | data science, statistics, survey, LLM, agent | cs.AI, cs.CL, cs.LG, stat.OT | 2025.09 | |
| arXiv(v1) 2025 | OPPO: Accelerating PPO-based RLHF via Pipeline Overlap | Authors TBD | agentic RL, PPO, RLHF, efficiency, overlap | cs.LG, cs.AI | 2025.09 | |
| arXiv(v1) 2025 | RL Foundations for Deep Research Systems: A Survey | Authors TBD | agentic RL, deep research, survey, post-DeepSeek | cs.LG, cs.AI | Post Feb 2025 papers | 2025.09 |
| arXiv(v1) 2025 | Reward Hacking Mitigation using Verifiable Composite Rewards | Authors TBD | RLHF, reward hacking, RLVR, verification | cs.LG, cs.AI | 2025.09 | |
| arXiv(v1) 2025 | SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection | Yubin Ge, et al. | self-learning, multi-level reflection, LLM, agent | cs.AI | Accepted at EMNLP 2025 Main Conference | 2025.09 |
| arXiv(v1) 2025 | Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning | Pulkit Verma, et al. | logical chain-of-thought, symbolic planning, LLM | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | The Landscape of Agentic Reinforcement Learning for LLMs: A Survey | Authors TBD | agentic RL, survey, LLM, POMDP | cs.LG, cs.AI | 2025.09 | |
| arXiv(v1) 2025 | Where LLM Agents Fail and How They can Learn From Failures | Kunlun Zhu, et al. | failure detection, learning from failures, LLM, agent | cs.AI | 2025.09 | |
| arXiv(v1) 2025 | iStar: Agentic Reinforcement Learning with Implicit Step Rewards | Authors TBD | agentic RL, credit assignment, implicit PRM | cs.LG, cs.AI | 2025.09 | |
| arXiv(v3) 2025 | MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement | Jaehyun Nam, et al. | ML engineering, code generation, search, LLM, agent | cs.LG | 2025.08 | |
| arXiv(v1) 2025 | Agent Safety Alignment via Reinforcement Learning | Zeyang Sha, et al. | safety alignment, reinforcement learning, LLM, agent | cs.AI, cs.CR | 2025.07 | |
| arXiv(v1) 2025 | AgentMesh: A Cooperative Multi-Agent Generative AI Framework for Software Development Automation | Sourena Khanzadeh | multi-agent, software development, code generation, LLM, agent | cs.SE, cs.AI | 2025.07 | |
| arXiv(v1) 2025 | Technical Survey of RL Techniques for Large Language Models | Authors TBD | agentic RL, survey, PPO, DPO, GRPO | cs.LG, cs.AI | 2025.07 | |
| arXiv(v2) 2025 | ToolACE: Winning the Points of LLM Function Calling | Weiwen Liu, et al. | function calling, tool use, LLM, agent | cs.LG, cs.AI, cs.CL | 21 pages, 22 figures | 2025.07 |
| arXiv(v1) 2025 | A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy | Henry Peng Zou, et al. | human-agent systems, collaboration, LLM, agent | cs.AI, cs.CL, cs.HC, cs.LG, cs.MA | 2025.06 | |
| arXiv(v2) 2025 | Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems | Bingyu Yan, et al. | multi-agent communication, survey, LLM, agent | cs.MA, cs.CL | 2025.06 | |
| arXiv(v1) 2025 | Enhancing Decision-Making of Large Language Models via Actor-Critic | Heng Dong, et al. | decision-making, actor-critic, LLM, agent | cs.CL, cs.AI | Forty-second International Conference on Machine Learning (ICML 2025) | 2025.06 |
| arXiv(v1) 2025 | GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation | Ning Gao, et al. | simulation, robotic manipulation, LLM, agent | cs.RO | 2025.06 | |
| arXiv(v1) 2025 | OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems | Xiaozhe Li, et al. | optimization, benchmark, evaluation, LLM, agent | cs.AI, cs.LG | 2025.06 | |
| arXiv(v1) 2025 | RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments | Yuchuan Fu, et al. | security evaluation, benchmark, LLM, agent | cs.CR, cs.AI | 12 pages, 8 figures | 2025.06 |
| arXiv(v1) 2025 | Sailing by the Stars: Survey on Reward Models and Learning Strategies | Authors TBD | agentic RL, reward model, survey, learning | cs.LG, cs.AI | 2025.06 | |
| arXiv(v1) 2025 | ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering | Zexi Liu, et al. | agentic RL, ML engineering, LLM, agent | cs.CL, cs.AI, cs.LG | 2025.05 | |
| arXiv(v1) 2025 | Multi-Agent Systems for Robotic Autonomy with LLMs | Junhong Chen, et al. | multi-agent, robotics, LLM, agent | cs.RO, cs.AI | 11 pages, 2 figures, 5 tables, submitted for publication | 2025.05 |
| arXiv(v1) 2025 | Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking | Yihan Chen, et al. | synthetic trajectories, self-reflection, training, LLM, agent | cs.CL | 2025.05 | |
| arXiv(v1) 2025 | Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning | Joykirat Singh, et al. | agentic RL, tool integration, reasoning, LLM | cs.AI | 2025.04 | |
| arXiv(v1) 2025 | Comprehensive Survey of Reward Models: Taxonomy and Applications | Authors TBD | agentic RL, reward model, survey, taxonomy | cs.LG, cs.AI | 2025.04 | |
| arXiv(v1) 2025 | DPO Meets PPO: Reinforced Token Optimization for RLHF | Han Zhong, et al. | agentic RL, DPO, PPO, RLHF, token-level | cs.LG, cs.AI | RTO framework | 2025.04 |
| arXiv(v1) 2025 | Hierarchical Multi-Step Reward Models for Enhanced Reasoning | Authors TBD | agentic RL, reward model, hierarchical, reasoning | cs.LG, cs.AI | 2025.03 | |
| arXiv(v1) 2025 | Look Before You Leap: Using Serialized State Machine for Language Conditioned Robotic Manipulation | Tong Mu, et al. | finite state machine, robotic manipulation, LLM | cs.RO, cs.AI | 7 pages, 4 figures | 2025.03 |
| arXiv(v1) 2025 | SafePlan: Leveraging Formal Logic and Chain-of-Thought Reasoning for Enhanced Safety in LLM-based Robotic Task Planning | Ike Obi, et al. | formal logic, chain-of-thought, safety, robotic planning, LLM | cs.RO | 2025.03 | |
| arXiv(v2) 2025 | Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation | Hyungjoo Chae, et al. | web navigation, environment dynamics, LLM, agent | cs.CL | ICLR 2025 | 2025.03 |
| arXiv(v1) 2025 | Every Software as an Agent: Blueprint and Case Study | Mengwei Xu | software agent, autonomous, LLM | cs.SE, cs.AI | 2025.02 | |
| arXiv(v2) 2025 | Flow: Modularized Agentic Workflow Automation | Boye Niu, et al. | workflow generation, modular, LLM, agent | cs.AI, cs.LG, cs.MA | 2025.02 | |
| arXiv(v1) 2025 | Policy Learning with a Natural Language Action Space: A Causal Approach | Bohan Zhang, et al. | policy learning, natural language action, LLM, agent | cs.CL | 2025.02 | |
| arXiv(v1) 2025 | Process Reward Models for LLM Agents: Practical Framework | Authors TBD | PRM, reward model, LLM agent, RLHF | cs.LG, cs.AI | InversePRM | 2025.02 |
| arXiv(v1) 2025 | Provably Efficient Online RLHF with One-Pass Reward Modeling | Authors TBD | online RLHF, reward modeling, efficiency | cs.LG, cs.AI | 2025.02 | |
| arXiv(v1) 2025 (Proceedings of the 2024 IEEE International Japan-Africa Conference on Electronics communications and Computations (JAC ECC)) | Guided Code Generation with LLMs: A Multi-Agent Framework for Complex Code Tasks | Amr Almorsi, et al. | multi-agent, code generation, LLM, agent | cs.AI | 4 pages, 3 figures | 2025.01 |
| arXiv(v1) 2025 | REINFORCE++: Critic-Free Policy Optimization with Global Advantage Normalization | Authors TBD | agentic RL, REINFORCE, critic-free, GRPO | cs.LG, cs.AI | Outperforms PPO | 2025.01 |
| arXiv(v4) 2024 | Planning with Multi-Constraints via Collaborative Language Agents | Cong Zhang, et al. | meta-task planning, multi-agent, LLM, agent | cs.AI, cs.CL, cs.LG | 2024.12 | |
| arXiv(v2) 2024 | AutoWebGLM: A Large Language Model-based Web Navigating Agent | Hanyu Lai, et al. | web navigation, browsing, LLM, agent | cs.CL | Accepted to KDD 2024 | 2024.10 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v2) 2026 | A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation | Weiayn Shi, et al. | VLM | cs.HC | Accepted at CHI 2026 Posters | 2026.02 |
| arXiv(v1) 2026 | Beyond Dominant Patches: Spatial Credit Redistribution For Grounded Vision-Language Models | Niamul Hassan Samin, et al. | VLM | cs.CV, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Beyond Static Artifacts: A Forensic Benchmark for Video Deepfake Reasoning in Vision Language Models | Zheyuan Gu, et al. | VLM | cs.CV, cs.AI | 16 pages, 9 figures. Submitted to CVPR 2026 | 2026.02 |
| arXiv(v1) 2026 | Causal Decoding for Hallucination-Resistant Multimodal Large Language Models | Shiwei Tan, et al. | VLM | cs.LG, cs.AI, cs.CV | Published in Transactions on Machine Learning Research (TMLR), 2026 | 2026.02 |
| arXiv(v1) 2026 | Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models | Jianghao Yin, et al. | VLM | cs.CV, cs.AI | Accepted by ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models | Xingyu Zhu, et al. | VLM | cs.CV, cs.MM | ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | HulluEdit: Single-Pass Evidence-Consistent Subspace Editing for Mitigating Hallucinations in Large Vision-Language Models | Yangguang Lin, et al. | VLM | cs.CV | accepted at CVPR 2026 | 2026.02 |
| arXiv(v1) 2026 | Large Multimodal Models as General In-Context Classifiers | Marco Garosi, et al. | VLM | cs.CV | CVPR Findings 2026. Project website at https://circle-lmm.github.io/ | 2026.02 |
| arXiv(v1) 2026 | Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation | Xingyu Zhu, et al. | VLM | cs.CV | ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | MediX-R1: Open Ended Medical Reinforcement Learning | Sahal Shaji Mullappilly, et al. | VLM | cs.CV | 2026.02 | |
| arXiv(v1) 2026 | NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors | Lingfeng Ren, et al. | VLM | cs.CV, cs.AI, cs.CL | Code: https://github.com/lingfengren/NoLan | 2026.02 |
| arXiv(v1) 2026 | See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs | Yongchang Zhang, et al. | VLM | cs.CV | CVPR2026 Accepted | 2026.02 |
| arXiv(v1) 2026 | Seeing Graphs Like Humans: Benchmarking Computational Measures and MLLMs for Similarity Assessment | Seokweon Jung, et al. | VLM | cs.HC | 21 pages including 1 page of appendix, 9 figures, 4 tables | 2026.02 |
| arXiv(v1) 2026 | Suppressing Prior-Comparison Hallucinations in Radiology Report Generation via Semantically Decoupled Latent Steering | Ao Li, et al. | VLM | cs.CV | 15 pages, 5 figures | 2026.02 |
| arXiv(v1) 2026 | SurGo-R1: Benchmarking and Modeling Contextual Reasoning for Operative Zone in Surgical Video | Guanyi Qin, et al. | VLM | cs.CV, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification | Vikash Singh, et al. | VLM | cs.CV, cs.AI, cs.CL, cs.LO | 2026.02 | |
| arXiv(v1) 2026 | VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation | Seongheon Park, et al. | VLM | cs.CV, cs.AI, cs.CL | 2026.02 | |
| arXiv(v1) 2026 | Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback | Authors TBD | MLLM, hallucination, caption feedback, grounding | cs.CV, cs.CL | 2026.01 | |
| arXiv(v1) 2026 | Innovator-VL: A Multimodal Large Language Model for Scientific Discovery | Zichen Wen, et al. | scientific discovery, MLLM, vision-language | cs.CV, cs.AI | Innovator-VL tech report | 2026.01 |
| arXiv(v2) 2026 | MGPC: Multimodal Network for Generalizable Point Cloud Completion With Modality Dropout and Progressive Decoding | Jiangyuan Liu, et al. | point cloud completion, multimodal, MLLM | cs.CV | Code and dataset are available at https://github.com/L-J-Yuan/MGPC | 2026.01 |
| arXiv(v1) 2026 | Multimodal In-context Learning for ASR of Low-resource Languages | Zhaolin Li, et al. | multimodal in-context learning, ASR, low-resource, MLLM | cs.CL, cs.AI | Under review | 2026.01 |
| arXiv(v3) 2026 | Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models | Umberto Cappellazzo, et al. | unified speech recognition, multimodal, MLLM | eess.AS, cs.CV, cs.SD | Accepted to IEEE ICASSP 2026 (camera-ready version). Project website (code and model weights): https://umbertocappellazzo.github.io/Omni-AVSR/ | 2026.01 |
| arXiv(v2) 2026 | Table as a Modality for Large Language Models | Liyao Li, et al. | table modality, structured data, MLLM | cs.CL, cs.AI | Accepted to NeurIPS 2025 | 2026.01 |
| arXiv(v1) 2026 | The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection | Wei Ai, et al. | fake news detection, vision-language, MLLM, survey | cs.AI, cs.CV | 2026.01 | |
| arXiv(v3) 2026 | UniVideo: Unified Understanding, Generation, and Editing for Videos | Cong Wei, et al. | video understanding, generation, editing, unified, MLLM | cs.CV | Project Website https://congwei1230.github.io/UniVideo/ | 2026.01 |
| arXiv(v1) 2026 | VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory | Shaoan Wang, et al. | embodied navigation, adaptive reasoning, MLLM, vision-language | cs.RO, cs.CV | Project page: https://wsakobe.github.io/VLingNav-web/ | 2026.01 |
| arXiv(v1) 2026 | VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding | Jiapeng Shi, et al. | video understanding, spatial-temporal, MLLM, vision-language | cs.CV | 2026.01 | |
| arXiv(v1) 2025 | A Medical Multimodal Diagnostic Framework Integrating Vision-Language Models and Logic Tree Reasoning | Zelin Zang, et al. | medical diagnosis, logic tree reasoning, vision-language, MLLM | cs.AI | 2025.12 | |
| arXiv(v1) 2025 | DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models | Zefeng He, et al. | generative reasoning, diffusion model, MLLM | cs.CV | Project page: https://diffthinker-project.github.io | 2025.12 |
| arXiv(v1) 2025 | From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs | Authors TBD | MLLM, spatial reasoning, benchmark, open world | cs.CV, cs.CL | 2025.12 | |
| arXiv(v1) 2025 | Kling-Omni Technical Report | Kling Team, et al. | video generation, multimodal synthesis, MLLM | cs.CV | Kling-Omni Technical Report | 2025.12 |
| arXiv(v1) 2025 | Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding | Yongyuan Liang, et al. | 3D understanding, point cloud, spatial understanding, MLLM | cs.CV, cs.AI | 2025.12 | |
| arXiv(v3) 2025 (Proc. 2025 IEEE 8th International Conference on Multimedia Information Processing and Retrieval (MIPR), pp. 456-462, 2025) | MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models | Philip R. Liu, et al. | multi-agent, medical diagnosis, MLLM | cs.MA, cs.AI, cs.CV, cs.LG | 2025.12 | |
| arXiv(v2) 2025 | TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning | Tao Wu, et al. | temporal understanding, reinforcement learning, MLLM, video | cs.CV | 2025.12 | |
| arXiv(v1) 2025 | MVU-Eval: Multi-Video Understanding Evaluation for MLLMs | Authors TBD | MLLM, multi-video, evaluation, benchmark | cs.CV, cs.CL | 2025.11 | |
| arXiv(v1) 2025 (Proceedings of the Conference on Language Modeling (COLM 2025)) | REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories | Jacob Thompson, et al. | embodied spatial reasoning, trajectory, MLLM | cs.LG, cs.AI, cs.CV | 2025.11 | |
| arXiv(v1) 2025 | Seeing is Believing: Rich-Context Hallucination Detection via Backward Visual Grounding | Authors TBD | MLLM, hallucination, detection, visual grounding | cs.CV, cs.CL | Outperforms GPT-4o | 2025.11 |
| arXiv(v1) 2025 | SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards | Authors TBD | MLLM, 3D reasoning, spatial, RL | cs.CV, cs.CL | Outperforms GPT-4o | 2025.11 |
| arXiv(v1) 2025 | MT-Video-Bench: Video Understanding Benchmark for MLLMs in Multi-Turn Dialogues | Authors TBD | MLLM, video understanding, benchmark, multi-turn | cs.CV, cs.CL | 2025.10 | |
| arXiv(v1) 2025 | MemVR: Memory-Space Visual Retracing for Hallucination Mitigation in MLLMs | Authors TBD | MLLM, hallucination, mitigation, memory | cs.CV, cs.CL | Plug-and-play | 2025.10 |
| arXiv(v2) 2025 | Revealing Multimodal Causality with Large Language Models | Jin Li, et al. | causal discovery, MLLM, multimodal causality | cs.LG, cs.AI | Accepted at NeurIPS 2025 | 2025.10 |
| arXiv(v1) 2025 | Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph | Wentao Wang, et al. | spatio-temporal reasoning, relation graph, MLLM, video | cs.AI | 2025.10 | |
| arXiv(v1) 2025 | Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs | Authors TBD | MLLM, hallucination, omission, fabrication | cs.CV, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | VIRAL: Visual Representation Alignment for MLLMs | Authors TBD | MLLM, visual alignment, fine-grained understanding | cs.CV, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents | Han Lin, et al. | diffusion model, patch-level CLIP, image generation, MLLM | cs.CV, cs.AI, cs.CL | Project Page: https://bifrost-1.github.io | 2025.08 |
| arXiv(v1) 2025 | Grounding the Ungrounded: Spectral-Graph Framework for Quantifying Hallucinations in MLLMs | Authors TBD | MLLM, hallucination, grounding, detection | cs.CV, cs.CL | 2025.08 | |
| arXiv(v1) 2025 | Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey | Authors TBD | MLLM, VLA, robotics, manipulation, survey | cs.RO, cs.CV | 2025.08 | |
| arXiv(v1) 2025 | Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting | Miaosen Luo, et al. | affective computing, emotion recognition, MLLM | cs.AI, cs.LG | 2025.08 | |
| arXiv(v1) 2025 | RynnEC: Bringing MLLMs into Embodied World | Authors TBD | MLLM, embodied, video, spatial reasoning | cs.CV, cs.RO | 2025.08 | |
| arXiv(v3) 2025 | SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding | Can Liu, et al. | visualization understanding, dataset, MLLM | cs.HC, cs.CV | 2025.07 | |
| arXiv(v2) 2025 | UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation | Yanzhe Chen, et al. | codebook, multimodal generation, MLLM | cs.CV, cs.MM | 19 pages, 5 figures | 2025.07 |
| arXiv(v1) 2025 | CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning | Kailing Li, et al. | embodied visual reasoning, cognitive map, MLLM | cs.CV, cs.AI, cs.CL | 2025.06 | |
| arXiv(v1) 2025 | Insight-V: Exploring Long-Chain Visual Reasoning with MLLMs | Authors TBD | MLLM, visual reasoning, long-chain, CVPR | cs.CV | CVPR 2025 | 2025.06 |
| arXiv(v2) 2025 | LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning | Zebin You, et al. | diffusion model, visual instruction tuning, MLLM | cs.LG, cs.CL, cs.CV | Project page and codes: \url{https://ml-gsai.github.io/LLaDA-V-demo/} | 2025.06 |
| arXiv(v2) 2025 | LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding | Hongyu Li, et al. | spatial-temporal understanding, MLLM, vision-language | cs.CV | Accepted by CVPR2025 | 2025.06 |
| arXiv(v1) 2025 (CVPR 2025) | LLaVA-ST: Multimodal LLM for Fine-Grained Spatial-Temporal Understanding | Authors TBD | MLLM, spatial-temporal, video, CVPR | cs.CV | CVPR 2025 | 2025.06 |
| arXiv(v1) 2025 | Manager: Aggregating Insights from Unimodal Experts in VLMs and MLLMs | Authors TBD | MLLM, VLM, unimodal experts, fusion | cs.CV, cs.CL | 2025.06 | |
| arXiv(v1) 2025 | MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis | Yuting Zhang, et al. | medical reasoning, diagnosis, MLLM | eess.IV, cs.CL, cs.CV, q-bio.QM | 2025.06 | |
| arXiv(v1) 2025 | Multimodal Tabular Reasoning with Privileged Structured Information | Jun-Peng Jiang, et al. | tabular reasoning, structured information, MLLM | cs.LG, cs.AI, cs.CL, cs.CV | 2025.06 | |
| arXiv(v1) 2025 | Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models | Hugues Thomas, et al. | 3D scene understanding, point cloud, token structure, MLLM | cs.CV | Main paper and appendix | 2025.06 |
| CVPR25 (CVPR 2025) | Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Reweighting | Authors TBD | MLLM, hallucination, attention, CVPR | cs.CV | CVPR 2025 | 2025.06 |
| arXiv(v3) 2025 | SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models | Wufei Ma, et al. | 3D-informed, spatial intelligence, MLLM | cs.CV | CVPR 2025 highlight | 2025.06 |
| arXiv(v1) 2025 | Structured Attention Matters to Multimodal LLMs in Document Understanding | Chang Liu, et al. | document understanding, structured attention, MLLM | cs.CL, cs.AI, cs.IR | 2025.06 | |
| arXiv(v2) 2025 | Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM | Zinuo Li, et al. | audio-visual-speech, multimodal, MLLM, video understanding | cs.CL | 2025.06 | |
| arXiv(v3) 2025 | Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning | NVIDIA, et al. | physical common sense, embodied reasoning, MLLM | cs.AI, cs.CV, cs.LG, cs.RO | 2025.05 | |
| arXiv(v1) 2025 | HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and Reasoning | Chuhao Zhou, et al. | multisensory, human sensing, reasoning, MLLM | cs.CV, cs.AI, cs.CL, cs.LG, cs.MM | 18 pages, 13 figures, 6 tables | 2025.05 |
| arXiv(v2) 2025 | Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation | Liu He, et al. | video generation, multi-agent collaboration, MLLM | cs.CV, cs.GR, cs.MM | Accepted by CVPR 2025 AI4CC Workshop | 2025.05 |
| arXiv(v1) 2025 | MedBridge: Bridging Foundation Vision-Language Models to Medical Image Diagnosis | Authors TBD | MLLM, medical, VLM, diagnosis | cs.CV | 2025.05 | |
| arXiv(v1) 2025 | VideoLLM Benchmarks and Evaluation: A Survey | Authors TBD | MLLM, video, benchmark, evaluation, survey | cs.CV, cs.CL | 2025.05 | |
| arXiv(v2) 2025 | Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark | Hanlei Zhang, et al. | multimodal language analysis, MLLM, semantics | cs.CL, cs.AI, cs.MM | 23 pages, 5 figures | 2025.04 |
| arXiv(v2) 2025 | Dual Diffusion for Unified Image Generation and Understanding | Zijie Li, et al. | dual diffusion, unified generation, understanding, MLLM | cs.CV, cs.AI, cs.LG | 2025.04 | |
| arXiv(v1) 2025 | Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents | Gavin Greif, et al. | OCR, historical documents, named entity recognition, MLLM | cs.CL, cs.AI, cs.DL | 2025.04 | |
| arXiv(v1) 2025 | Socratic Chart: Cooperating Multiple Agents for Robust SVG Chart Understanding | Yuyang Ji, et al. | chart understanding, SVG, multi-agent, MLLM | cs.CV | 2025.04 | |
| arXiv(v1) 2025 | VLM-R1: A Stable and Generalizable R1-Style Large VLM | Authors TBD | MLLM, VLM, reasoning, R1-style | cs.CV, cs.CL | 2025.04 | |
| arXiv(v1) 2025 | MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation | Authors TBD | MLLM, 3D, segmentation, reasoning | cs.CV | 2025.03 | |
| arXiv(v1) 2025 | MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs | Erik Daxberger, et al. | MLLM, 3D, spatial, understanding, benchmark | cs.CV, cs.CL | ICCV 2025 | 2025.03 |
| arXiv(v1) 2025 | Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis | Authors TBD | MLLM, medical, 3D, VLM, CT | cs.CV | 2025.03 | |
| arXiv(v1) 2025 | Mobile-VideoGPT: Fast and Accurate Video Understanding Language Model | Authors TBD | MLLM, video, mobile, efficient | cs.CV, cs.CL | 2025.03 | |
| arXiv(v1) 2025 | R1-Zero's Aha Moment in Visual Reasoning on a 2B Non-SFT Model | Authors TBD | MLLM, visual reasoning, R1-Zero, emergent | cs.CV, cs.CL | 2025.03 | |
| arXiv(v1) 2025 | SpaceVLLM: Endowing MLLM with Spatio-Temporal Video Grounding | Authors TBD | MLLM, video, spatio-temporal, grounding | cs.CV, cs.CL | 2025.03 | |
| arXiv(v1) 2025 | Vision-R1: Incentivizing Reasoning Capability in MLLMs | Authors TBD | MLLM, reasoning, visual reasoning | cs.CV, cs.CL | 2025.03 | |
| arXiv(v2) 2025 | Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning | Bohao Yang, et al. | table understanding, scientific data, MLLM | cs.CL | 2025.02 | |
| arXiv(v4) 2025 | LMFusion: Adapting Pretrained Language Models for Multimodal Generation | Weijia Shi, et al. | multimodal generation, LLM adaptation, MLLM | cs.CL, cs.AI, cs.CV, cs.LG | Name change: LlamaFusion to LMFusion | 2025.02 |
| arXiv(v1) 2025 | Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review | Pei Fu, et al. | text-rich image understanding, MLLM, vision-language | cs.CV | 2025.02 | |
| arXiv(v1) 2025 | Visual Perception Token for Multimodal Large Language Models | Authors TBD | MLLM, visual perception, token, autonomous control | cs.CV, cs.CL | 829k training samples | 2025.02 |
| arXiv(v1) 2025 | Weak Supervision Dynamic KL-Weighted Diffusion Models Guided by Large Language Models | Julian Perry, et al. | diffusion model, LLM guidance, weak supervision, image generation | cs.CL | 2025.02 | |
| arXiv(v1) 2025 | Bridging Visualization and Optimization: Multimodal Large Language Models on Graph-Structured Combinatorial Optimization | Jie Zhao, et al. | graph-structured optimization, combinatorial, MLLM | cs.AI, cs.LG | 2025.01 | |
| arXiv(v1) 2025 | Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding | Yun Li, et al. | temporal modeling, video understanding, MLLM | cs.CV, cs.CL | 2025.01 | |
| arXiv(v5) 2025 | Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation | Yuyang Ye, et al. | sequential recommendation, MLLM, multimodal | cs.IR, cs.AI | 2025.01 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | Contextual Memory Virtualisation: DAG-Based State Management and Structurally Lossless Trimming for LLM Agents | Cosmo Santoni | LLM, memory | cs.SE, cs.AI, cs.HC, cs.OS | 11 pages. 6 figures. Introduces a DAG-based state management system for LLM agents. Evaluation on 76 coding sessions shows up to 86% token reduction (mean 20%) while remaining economically viable under prompt caching. Includes reference implementation for Claude Code | 2026.02 |
| arXiv(v1) 2026 | A Dynamic Retrieval-Augmented Generation System with Selective Memory and Remembrance | Okan Bursa | dynamic RAG, selective memory, retrieval, LLM | cs.IR, cs.AI | 6 Pages, 2 figures | 2026.01 |
| arXiv(v1) 2026 | Active Context Compression: Autonomous Memory Management in LLM Agents | Authors TBD | memory, compression, context, autonomous, Focus | cs.CL, cs.AI | 22.7% token savings | 2026.01 |
| arXiv(v1) 2025 | Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management | Authors TBD | memory, unified, long-term, short-term, management | cs.AI, cs.CL | 2026.01 | |
| arXiv(v1) 2026 | Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents | Authors TBD | memory, temporal, semantic, personalized | cs.CL, cs.AI | Durative memory, Zep architecture | 2026.01 |
| arXiv(v1) 2026 | Beyond Static Summarization: Proactive Memory Extraction for LLM Agents | Chengyuan Yang, et al. | memory, extraction, summarization, agent | cs.CL, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Continuum Memory Architectures for Long-Horizon LLM Agents | Authors TBD | memory, continuum, long-horizon, consolidation | cs.AI, cs.CL | Episodic-to-semantic conversion | 2026.01 |
| arXiv(v2) 2026 | Cost and accuracy of long-term memory in Distributed Multi-Agent Systems based on Large Language Models | Benedict Wolff, et al. | graph memory, distributed multi-agent, LLM | cs.IR | 23 pages, 4 figures, 7 tables | 2026.01 |
| arXiv(v1) 2026 | Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration | Sen Wang, et al. | memory, benchmark, multimodal, RL, MLLM | cs.AI, cs.CV | Our dataset and code will be released at our \href{https://wangsen99.github.io/papers/lmee/}{website} | 2026.01 |
| arXiv(v1) 2026 | Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory Management | Weitao Ma, et al. | memory, management, feedback, alignment, agent | cs.CL | 18 pages, 5 figures | 2026.01 |
| arXiv(v3) 2026 | HaluMem: Evaluating Hallucinations in Memory Systems of Agents | Ding Chen, et al. | memory, hallucination, evaluation, agent | cs.CL | 2026.01 | |
| arXiv(v1) 2026 | HiMeS: Hippocampus-inspired Memory System for Personalized AI Assistants | Hailong Li, et al. | memory, personalized assistant, hippocampus, long-term | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | HiMem: Hierarchical Long-Term Memory for LLM Long-Horizon Agents | Ningning Zhang, et al. | memory, long-term, hierarchical, agent, LLM | cs.AI | 2026.01 | |
| arXiv(v2) 2026 | Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory | Sizhe Yuen, et al. | multi-agent, structured memory, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents | Davide Baldelli, et al. | working memory, evaluation, agent, LLM | cs.CL | 2026.01 | |
| arXiv(v2) 2026 | LiCoMemory: Lightweight and Cognitive Agentic Memory for Efficient Long-Term Reasoning | Zhengjun Huang, et al. | cognitive memory, long-term reasoning, LLM, agent | cs.IR | 2026.01 | |
| arXiv(v1) 2026 | MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents | Dongming Jiang, et al. | multi-graph, agentic memory, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents | Yuanchen Bei, et al. | memory, multimodal, conversational, MLLM benchmark | cs.CL, cs.AI | 34 pages, 18 figures | 2026.01 |
| arXiv(v1) 2026 | MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards | Authors TBD | memory, long-term, construction, dense reward | cs.CL, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction | Haonan Bian, et al. | memory, benchmark, interaction, agent | cs.CL, cs.AI | 2026.01 | |
| arXiv(v3) 2026 | SimpleMem: Efficient Lifelong Memory for LLM Agents | Jiaqi Liu, et al. | lifelong memory, efficient, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | SwiftMem: Fast Agentic Memory via Query-aware Indexing | Anxin Tian, et al. | memory, retrieval, indexing, agent, LLM | cs.CL, cs.AI | 2026.01 | |
| arXiv(v3) 2026 | TeleMem: Building Long-Term and Multimodal Memory for Agentic AI | Chunliang Chen, et al. | memory, multimodal, long-term, agent, LLM | cs.CL, cs.AI, cs.CV | 2026.01 | |
| arXiv(v1) 2025 | Tool-Memory Conflicts in Tool-Augmented LLMs | Authors TBD | memory, tool use, conflict, LLM | cs.CL, cs.AI | 2026.01 | |
| arXiv(v1) (NeurIPS25) | A-MEM: Agentic Memory for LLM Agents (NeurIPS) | Authors TBD | memory, agentic, Zettelkasten, NeurIPS | cs.AI, cs.CL | NeurIPS 2025 publication | 2025.12 |
| arXiv(v1) 2025 | AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents | Jiafeng Liang, et al. | memory, cognitive, survey, agent | cs.CL, cs.AI, cs.CV | 57 pages, 5 figures | 2025.12 |
| arXiv(v1) 2025 | Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Rewards, Experience Synthesis, and Continual Memory | Ken Huang, et al. | memory, continual, self-improvement, agent | cs.CR, cs.AI | 11 pages, 4 figures. Includes a complete runnable reference implementation and audit logging framework | 2025.12 |
| arXiv(v1) 2025 | Beyond Heuristics: A Decision-Theoretic Framework for Agent Memory Management | Changzhi Sun, et al. | memory, agent, decision-theoretic, management | cs.CL | 2025.12 | |
| arXiv(v1) 2025 | Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs | Ngoc Bui, et al. | memory, KV cache, retention, long-context | cs.LG, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | Context as a Tool: Context Management for Long-Horizon SWE-Agents | Shukai Liu, et al. | memory, context management, long-horizon, SWE-agent | cs.CL | 2025.12 | |
| arXiv(v2) 2025 | Evaluating Long-Term Memory for Long-Context Question Answering | Alessandra Terranova, et al. | memory, long-context, evaluation, QA | cs.CL | Accepted as a poster at Metacognition in Generative AI EurIPS workshop | 2025.12 |
| arXiv(v1) 2025 | Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects | Authors TBD | memory, hindsight, reflection, retention | cs.AI, cs.CL | Retain-Recall-Reflect framework | 2025.12 |
| arXiv(v1) 2025 | Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement | Saman Forouzandeh, et al. | memory, procedural, hierarchical, agent | cs.LG, cs.AI | Accepted at The 25th International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS 2026). 21 pages including references, with 7 figures and 8 tables. Code is publicly available at the authors GitHub repository: https://github.com/S-Forouzandeh/MACLA-LLM-Agents-AAMAS-Conference | 2025.12 |
| arXiv(v2) 2025 | MMAG: Mixed Memory-Augmented Generation for Large Language Models Applications | Stefano Zeppieri | memory, memory-augmented generation, RAG, LLM | cs.CL, cs.IR | 2025.12 | |
| arXiv(v1) 2025 | MemEvolve: Meta-Evolution of Agent Memory Systems | Guibin Zhang, et al. | memory, evolution, meta-learning, agent | cs.CL, cs.MA | 2025.12 | |
| arXiv(v1) 2025 | MemR3: Memory Retrieval via Reflective Reasoning for LLM Agents | Authors TBD | memory, retrieval, reflection, reasoning | cs.AI, cs.CL | Router + evidence-gap tracker | 2025.12 |
| arXiv(v2) 2025 | Memento 2: Learning by Stateful Reflective Memory | Jun Wang | memory, agent, reflection, stateful | cs.AI, cs.CV, cs.LG | 35 pages, four figures | 2025.12 |
| arXiv(v1) 2025 | Memory in the Age of AI Agents | Yuyang Hu, et al. | memory, survey, LLM, agent | cs.CL, cs.AI | Comprehensive survey on agent memory | 2025.12 |
| arXiv(v1) 2025 | MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval | Saksham Sahai Srivastava, et al. | memory, security, agent, attack | cs.CR, cs.AI, cs.LG | 14 pages, 1 figure, includes appendix | 2025.12 |
| arXiv(v3) 2025 | O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents | Piaohong Wang, et al. | memory, long-horizon, self-evolving, agent | cs.CL | 2025.12 | |
| arXiv(v1) 2025 | R-Debater: Retrieval-Augmented Debate Generation through Argumentative Memory | Maoyuan Li, et al. | memory, retrieval, debate, RAG | cs.CL, cs.AI | Accepteed by AAMAS 2026 full paper | 2025.12 |
| arXiv(v2) 2025 | Significant Other AI: Identity, Memory, and Emotional Regulation as Long-Term Relational Intelligence | Sung Park | memory, identity, relational, long-term | cs.HC, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | Adaptive Focus Memory for Language Models | Authors TBD | memory, adaptive, focus, compression, AFM | cs.CL, cs.AI | 2/3 token reduction | 2025.11 |
| arXiv(v1) 2025 | BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing | Authors TBD | memory, selective, budget, efficient | cs.CL, cs.AI | Learned gating + BM25 | 2025.11 |
| arXiv(v1) 2025 | CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing | Guihang Hong, et al. | RAG, edge computing, optimization, LLM | cs.DC | Accepted by RTSS 2025 (Real-Time Systems Symposium, 2025) | 2025.11 |
| arXiv(v1) 2025 | EMem: Event-Centric Memory for Long-Term Conversational Agents | Authors TBD | memory, event-centric, conversation, neo-Davidsonian | cs.CL, cs.AI | Event-like propositions | 2025.11 |
| arXiv(v1) 2025 | Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory | Tianxin Wei, et al. | memory, benchmark, test-time learning, agent | cs.CL, cs.AI | 2025.11 | |
| arXiv(v1) 2025 | G-KV: Decoding-Time KV Cache Eviction with Global Attention | Mengqi Liao, et al. | memory, KV cache, eviction, efficiency | cs.CL, cs.AI | 2025.11 | |
| arXiv(v1) 2025 | GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory | Jeong Hun Yeo, et al. | memory, episodic, video, MLLM | cs.CV, cs.AI | 2025.11 | |
| arXiv(v1) 2025 | Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks | Yicong Zheng, et al. | memory, compression, long-context, search | cs.CL, cs.AI, cs.LG | 2025.11 | |
| arXiv(v1) 2025 | KVzip: Memory Compression for LLM Chatbots via KV Cache Optimization | Authors TBD | memory, compression, KV cache, chatbot | cs.CL, cs.AI | 3-4x compression, 170K tokens | 2025.11 |
| MobiCom25 | Poster: MemAura: Persistent Personalized Context Memory for LLM Services in Smart Environments | Siyuan Liu, et al. | LLM, memory, personalization, smart environment, context | 2025.11 | ||
| arXiv(v1) 2025 | Trainable Graph Memory for LLM Agents: From Experience to Strategy | Authors TBD | memory, graph, trainable, strategy | cs.AI, cs.CL | Utility assessment mechanism | 2025.11 |
| arXiv(v1) 2025 | WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance | Genglin Liu, et al. | memory, web agent, cross-session, self-evolving | cs.AI, cs.CL | 18 pages; work in progress | 2025.11 |
| arXiv(v1) 2025 | A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents | Zhipeng Liao, et al. | memory-efficient, RAG, wearable, medical, LLM, agent | cs.AR | Accepted by BioCAS2025 | 2025.10 |
| arXiv(v1) 2025 | Acon: Optimizing Context Compression for Long-horizon LLM Agents | Authors TBD | memory, compression, context, optimization | cs.CL, cs.AI | 26-54% memory reduction | 2025.10 |
| arXiv(v1) 2025 | Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs | Mohammad Tavakoli, et al. | memory, long-context, benchmark, long-term | cs.CL, cs.AI, cs.IR | 2025.10 | |
| arXiv(v1) 2025 | CAM: Contextual Augmentation Memory for LLM Agents | Authors TBD | memory, contextual, augmentation, agent | cs.AI, cs.CL | 2025.10 | |
| arXiv(v1) 2025 | Dynamic Affective Memory Management for Personalized LLM Agents | Junfeng Lu, et al. | memory, affective, personalization, agent | cs.CL | 12 pasges, 8 figures | 2025.10 |
| arXiv(v1) 2025 | Enabling Personalized Long-term Interactions in LLM-based Agents through Persistent Memory | Authors TBD | memory, personalized, long-term, user profile | cs.CL, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | LightMem: Lightweight Memory for Efficient LLM Agents | Authors TBD | memory, lightweight, efficient, agent | cs.AI, cs.CL | 2025.10 | |
| arXiv(v1) 2025 | MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments | Darshan Deshpande, et al. | memory, benchmark, state tracking, agent | cs.AI, cs.CL | Accepted to NeurIPS 2025 SEA Workshop | 2025.10 |
| arXiv(v1) 2025 | Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy Games | Runnan Qi, et al. | memory, prompting, state machine, agent | cs.AI | 10 pages, 4 figures, 1 table, 1 algorithm. Submitted to conference | 2025.10 |
| arXiv(v1) 2025 | Pre-Storage Reasoning for Episodic Memory in LLM Agents | Authors TBD | memory, episodic, pre-storage, reasoning | cs.AI, cs.CL | 2025.10 | |
| arXiv(v2) 2025 | Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions | Yuanzhe Hu, et al. | memory evaluation, multi-turn, LLM, agent | cs.CL, cs.AI | Y. Hu and Y. Wang contribute equally | 2025.09 |
| arXiv(v1) 2025 | HopRAG: Multi-Hop Reasoning with Graph Memory for LLM Agents | Authors TBD | memory, multi-hop, graph, reasoning, RAG | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Mem-α: Learning Memory Construction via Reinforcement Learning | Yu Wang, et al. | memory, RL, construction, agent | cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Mem-α: Memory with Adaptive Forgetting for LLM Agents | Authors TBD | memory, forgetting, adaptive, agent | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Memory in LLM-based Multi-agent Systems: Mechanisms, Challenges, and Collective | Authors TBD | memory, multi-agent, collective, survey | cs.AI, cs.MA | 2025.09 | |
| arXiv(v1) 2025 | Multiple Memory Systems for Enhancing Long-term Memory of LLM Agents | Authors TBD | memory, multiple systems, long-term, enhancement | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Nemori: Neural Memory Organization for LLM Agents | Authors TBD | memory, neural, organization, agent | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | SGMem: Sentence Graph Memory for Long-Term Conversational Agents | Yaxiong Wu, et al. | memory, graph, conversation, retrieval | cs.CL, cs.IR | 19 pages, 6 figures, 1 table | 2025.09 |
| arXiv(v1) 2025 | SGMem: Structured Graph Memory for LLM Agents | Authors TBD | memory, graph, structured, agent | cs.AI, cs.CL | 2025.09 | |
| IJCAI25 (IJCAI 2025) | AriGraph: Learning Knowledge Graph World Models with Episodic Memory | Authors TBD | memory, knowledge graph, episodic, world model | cs.AI | IJCAI 2025 | 2025.08 |
| arXiv(v1) 2025 | Cognitive Workspace: Active Memory Management for LLMs - Functional Infinite Context | Authors TBD | memory, cognitive, workspace, active management | cs.CL, cs.AI | Metacognitive control | 2025.08 |
| arXiv(v1) 2025 | Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework | Zeyu Zhang, et al. | adaptive memory, optimization, LLM, agent | cs.LG, cs.AI, cs.CL, cs.IR | 17 pages, 4 figures, 5 tables | 2025.08 |
| arXiv(v1) 2025 | Memory-Augmented Transformers: A Systematic Review | Authors TBD | memory, transformer, survey, augmented | cs.CL, cs.AI | Systematic review | 2025.08 |
| arXiv(v1) 2025 | Memory-R1: Enhancing LLM Agents to Manage and Utilize Memories via RL | Authors TBD | memory, RL, memory manager, ADD/UPDATE/DELETE | cs.AI, cs.LG | Memory Manager + Answer Agent | 2025.08 |
| arXiv(v1) 2025 | Recursive Summarization for Long-Term Dialogue Memory in LLMs | Authors TBD | memory, summarization, dialogue, recursive | cs.CL, cs.AI | Updated 2025 | 2025.08 |
| arXiv(v2) 2025 | In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents | Zhen Tan, et al. | reflective memory, dialogue agent, personalization, LLM | cs.CL, cs.AI | Accepted to ACL 2025 | 2025.07 |
| ACL25 (ACL 2025) | Pretraining Context Compressor for LLMs with Embedding-Based Memory | Authors TBD | memory, compression, embedding, context | cs.CL | PCC framework | 2025.07 |
| arXiv(v1) 2025 | Cross-Attention Networks for Memory Retrieval in Generative Agents | Authors TBD | memory, retrieval, cross-attention, generative | cs.AI, cs.CL | Frontiers in Psychology | 2025.04 |
| arXiv(v1) 2025 | From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs | Authors TBD | memory, survey, episodic, semantic, working memory | cs.CL, cs.AI | Personal/system, parametric/non-parametric | 2025.04 |
| arXiv(v3) 2025 | LIFT: Improving Long Context Understanding of Large Language Models through Long Input Fine-Tuning | Yansheng Mao, et al. | long context, fine-tuning, LLM | cs.CL | 2025.04 | |
| arXiv(v1) 2025 | Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory | Authors TBD | memory, long-term, scalable, production | cs.AI, cs.CL | 26% improvement over OpenAI | 2025.04 |
| arXiv(v1) 2025 | In Prospect and Retrospect: Reflective Memory Management for Long-term Dialogue Agents | Authors TBD | memory, reflective, dialogue, long-term | cs.CL, cs.AI | 2025.03 | |
| arXiv(v1) 2025 | Tuning LLMs by RAG Principles: Towards LLM-native Memory | Jiale Wei, et al. | RAG, fine-tuning, optimization, LLM | cs.CL, cs.AI, cs.IR | 2025.03 | |
| arXiv(v1) 2025 | A-MEM: Agentic Memory for LLM Agents | Authors TBD | memory, agentic, self-organizing, Zettelkasten | cs.AI, cs.CL | Dynamic memory organization | 2025.02 |
| arXiv(v1) 2025 | Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents | Authors TBD | memory, episodic, long-term, position paper | cs.AI, cs.CL | Encoding and retrieval | 2025.02 |
| arXiv(v1) 2025 | Zep: Temporal Knowledge Graph Architecture for Agent Memory | Authors TBD | memory, temporal, knowledge graph, agent | cs.AI, cs.CL | Episodic + semantic + community | 2025.02 |
| arXiv(v1) 2024 | Memory-Augmented Agent Training for Business Document Understanding | Jiale Liu, et al. | memory, agent, training, document | cs.CL, cs.AI | 11 pages, 8 figures | 2024.12 |
| arXiv(v1) 2024 | On the Structural Memory of LLM Agents | Ruihong Zeng, et al. | memory, agent, analysis, structural | cs.CL, cs.AI | 2024.12 | |
| arXiv(v1) 2024 | XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference | Weizhuo Li, et al. | memory, KV cache, long-context, personalization | cs.LG, cs.CL | 2024.12 | |
| arXiv(v1) 2024 | MELODI: Exploring Memory Compression for Long Contexts | Yinpeng Chen, et al. | memory compression, long context, LLM | cs.LG, cs.AI | 2024.10 | |
| arXiv(v1) 2024 | A Survey on the Memory Mechanism of Large Language Model based Agents | Zeyu Zhang, et al. | memory, survey, LLM, agent | cs.AI | ACM TOIS, 39 pages | 2024.04 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | DeepInterestGR: Mining Deep Multi-Interest Using Multi-Modal LLMs for Generative Recommendation | Yangchen Zeng | LLM, personalization | cs.LG, cs.CV, cs.CY | 2026.02 | |
| arXiv(v1) 2026 | Dynamic Personality Adaptation in Large Language Models via State Machines | Leon Pielage, et al. | LLM, personalization | cs.CL, cs.HC, cs.LG | 22 pages, 5 figures, submitted to ICPR 2026 | 2026.02 |
| arXiv(v1) 2026 | Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs | Wenqiu Tang, et al. | LLM, personalization | cs.CL | Accepted in PAKDD 2026 special session on Data Science :Foundation and Applications | 2026.02 |
| arXiv(v1) 2026 | InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation | Yu Li, et al. | LLM, personalization | cs.CL, cs.AI, cs.CY | 2026.02 | |
| arXiv(v1) 2026 | Learning to Reason for Multi-Step Retrieval of Personal Context in Personalized Question Answering | Maryam Amirizaniani, et al. | LLM, personalization | cs.CL, cs.AI, cs.IR | 2026.02 | |
| arXiv(v1) 2026 | Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization | Shangding Gu | LLM, personalization | cs.LG, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Multi-Agent Large Language Model Based Emotional Detoxification Through Personalized Intensity Control for Consumer Protection | Keito Inoshita | LLM, personalization | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Offline Reasoning for Efficient Recommendation: LLM-Empowered Persona-Profiled Item Indexing | Deogyong Kim, et al. | LLM, personalization | cs.IR, cs.LG | Under review | 2026.02 |
| arXiv(v1) 2026 | PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra | Xiachong Feng, et al. | LLM, personalization | cs.AI | ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | PRECTR-V2:Unified Relevance-CTR Framework with Cross-User Preference Mining, Exposure Bias Correction, and LLM-Distilled Encoder Optimization | Shuzhi Cao, et al. | LLM, personalization | cs.IR, cs.AI | arXiv admin note: text overlap with arXiv:2503.18395 | 2026.02 |
| arXiv(v1) 2026 | Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History | Serin Kim, et al. | LLM, personalization | cs.CL, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Personalized Graph-Empowered Large Language Model for Proactive Information Access | Chia Cheng Chang, et al. | LLM, personalization | cs.CL | 2026.02 | |
| arXiv(v1) 2026 | Personalized Prediction of Perceived Message Effectiveness Using Large Language Model Based Digital Twins | Jasmin Han, et al. | LLM, personalization | cs.CL, stat.AP | 31 pages, 5 figures, submitted to Journal of the American Medical Informatics Association (JAMIA). Drs. Chen and Thrul share last authorship | 2026.02 |
| arXiv(v1) 2026 | Sydney Telling Fables on AI and Humans: A Corpus Tracing Memetic Transfer of Persona between LLMs | Jiří Milička, et al. | LLM, personalization | cs.CL, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions | Yue Xu, et al. | LLM, personalization | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs | Liang Wang, et al. | arxiv | cs.CL | 2026.01 | |
| arXiv(v1) 2026 | Enriching Semantic Profiles into Knowledge Graph for Recommender Systems Using Large Language Models | Seokho Ahn, et al. | knowledge graph, semantic profiles, recommendation, LLM | cs.IR, cs.AI, cs.LG | Accepted at KDD 2026 | 2026.01 |
| ICLR 2026 | FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents | ICLR,OpenReview | Mobile Agent, LLM Agent, GUI, Proactive Agent, Personalization | OpenReview ID: n3iFV0gLMc | 2026.01 | |
| arXiv(v1) 2026 | HumanLLM: Towards Personalized Understanding and Simulation of Human Nature | Yuxuan Lei, et al. | arxiv | cs.CL | 12 pages, 5 figures, 7 tables, to be published in KDD 2026 | 2026.01 |
| arXiv(v1) 2026 | Improving User Privacy in Personalized Generation: Client-Side Retrieval-Augmented Modification of Server-Side Generated Speculations | Alireza Salemi, et al. | arxiv | cs.CL, cs.AI, cs.CR, cs.IR | 2026.01 | |
| arXiv(v2) 2026 | Linear Personality Probing and Steering in LLMs: A Big Five Study | Michel Frising, et al. | personalization, personality, probing, steering | cs.CL | 29 pages, 6 figures | 2026.01 |
| arXiv(v1) 2026 | Me-Agent: A Personalized Mobile Agent with Two-Level User Habit Learning for Enhanced Interaction | Shuoxin Wang, et al. | arxiv | cs.CL | 2026.01 | |
| ICLR 2026 | Meta-Router: Bridging Gold-standard and Preference-based Evaluations in LLM Routing | ICLR,OpenReview | Causal learning, Meta-learner, Large Language Model, query routing | OpenReview ID: r0BFucF2dH | 2026.01 | |
| ICLR 2026 | NextQuill: Causal Preference Modeling for Enhancing LLM Personalization | ICLR,OpenReview | Personalized text generation, Large Language Models, LLM Personalization | OpenReview ID: xYpVlKMFqv | 2026.01 | |
| arXiv(v1) 2026 | One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment | Hongru Cai, et al. | meta reward modeling, alignment, personalization, LLM | cs.CL, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization | Linfeng Du, et al. | arxiv | cs.CL, cs.IR | 2026.01 | |
| arXiv(v1) 2026 | PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation | Junho Park, et al. | few-shot, privacy-safe, lightweight adaptation, personalization, LLM | cs.CL, cs.AI, cs.LG | 16 pages, 9 figures | 2026.01 |
| arXiv(v1) 2026 | PersonaDual: Balancing Personalization and Objectivity via Adaptive Reasoning | Xiaoyou Liu, et al. | personalization, persona, reasoning, LLM | cs.AI | 2026.01 | |
| ICLR 2026 | Preference Leakage: A Contamination Problem in LLM-as-a-judge | ICLR,OpenReview | LLM-as-a-judge, Preference Leakage, Data Contamination | OpenReview ID: grIvSXVJ65 | 2026.01 | |
| ICLR 2026 | ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation | ICLR,OpenReview | Benchmark, Agent Simulation, Personalization, Proactivity | OpenReview ID: RV2aeCgxdB | 2026.01 | |
| arXiv(v1) 2026 | SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation | Seoyeon Kim, et al. | personalization, continual, retrieval, parametric adaptation, LLM | cs.AI, cs.CL | under review, 23 pages | 2026.01 |
| ICLR 2026 | Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents | ICLR,OpenReview | LLM-based Agents, Process Supervision, Curriculum Learning | OpenReview ID: s8usvGHYlk | 2026.01 | |
| arXiv(v1) 2026 | Structured Personality Control and Adaptation for LLM Agents | Jinpeng Wang, et al. | personalization, personality, control, agent, LLM | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Styles + Persona-plug = Customized LLMs | Yutong Song, et al. | style customization, persona-plug, LLM | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models | Christina Lu, et al. | persona, default persona, alignment, LLM | cs.CL | 2026.01 | |
| arXiv(v2) 2026 | The Reward Model Selection Crisis in Personalized Alignment | Fady Rezk, et al. | personalization, alignment, reward model, RLHF | cs.AI, cs.LG | 2026.01 | |
| ICLR 2026 | Towards Understanding Valuable Preference Data for Large Language Model Alignment | ICLR,OpenReview | Large language model alignment, preference data, influence function | OpenReview ID: FUp0KeEEBs | 2026.01 | |
| ICLR 2026 | Verification and Co-Alignment via Heterogeneous Consistency for Preference-Aligned LLM Annotations | ICLR,OpenReview | Verification, Co-Alignment, Preference-Aligned LLM Annotations, Reference-Free Metric | OpenReview ID: jugY302BAh | 2026.01 | |
| arXiv(v1) 2026 | When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs | Zhongxiang Sun, et al. | personalization, hallucination, safety, evaluation | cs.CL, cs.AI | 20 pages, 15 figures | 2026.01 |
| arXiv(v1) 2025 | Agentic Multi-Persona Framework for Evidence-Aware Fake News Detection | Roopa Bukke, et al. | persona, multi-agent, fake news, LLM | cs.IR, cs.LG | 12 pages, 8 tables, 2 figures | 2025.12 |
| arXiv(v1) 2025 | Interpolative Decoding: Exploring the Spectrum of Personality Traits in LLMs | Eric Yeh, et al. | personalization, personality, decoding, control | cs.AI | 20 pages, 5 figures | 2025.12 |
| arXiv(v1) 2025 | LLM Personas as a Substitute for Field Experiments in Method Benchmarking | Enoch Hyunwook Kang | persona, evaluation, methodology, LLM | cs.AI, cs.LG, econ.EM | 2025.12 | |
| arXiv(v1) 2025 | Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI | Samarth Sarin, et al. | personalization, memory, conversational AI, agent | cs.AI, cs.CL | Paper accepted at 5th International Conference of AIML Systems 2025, Bangalore, India | 2025.12 |
| arXiv(v1) 2025 | PILAR: Personalizing Augmented Reality Interactions with LLM-based Human-Centric and Trustworthy Explanations for Daily Use Cases | Ripan Kumar Kundu, et al. | personalization, AR, explanations, LLM | cs.HC, cs.AI | Published in the 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct) | 2025.12 |
| arXiv(v1) 2025 | PRISM: A Personality-Driven Multi-Agent Framework for Social Media Simulation | Zhixiang Lu, et al. | personality, multi-agent, social simulation, LLM | cs.CL | 2025.12 | |
| arXiv(v1) 2025 | PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas | Authors TBD | personalization, persona, implicit, user modeling | cs.CL, cs.AI | 1000 personas, 20k preferences, 128k context | 2025.12 |
| arXiv(v1) 2025 | Personalized Multimodal Large Language Models: A Survey | Authors TBD | personalization, MLLM, survey, multimodal | cs.CV, cs.CL | Comprehensive MLLM personalization survey | 2025.12 |
| arXiv(v1) 2025 | PrefGen: Multimodal Preference Learning for Image Generation | Authors TBD | personalization, MLLM, image generation, preference | cs.CV, cs.CL | User-specific conditioning | 2025.12 |
| arXiv(v2) 2025 | ProEx: A Unified Framework Leveraging Large Language Model with Profile Extrapolation for Recommendation | Yi Zhang, et al. | profile extrapolation, recommendation, LLM | cs.IR | Accepted by KDD 2026 (First Cycle) | 2025.12 |
| arXiv(v1) 2025 | SPARK: Search Personalization via Agent-Driven Retrieval and Knowledge-sharing | Gaurab Chhetri, et al. | personalization, search, agent, retrieval | cs.AI | Accepted to WEB&GRAPH 2026 (WSDM 2026 workshop) | 2025.12 |
| arXiv(v1) 2025 | TAME: Long-Context MLLM Personalization with Double Memories | Authors TBD | personalization, MLLM, memory, training-free | cs.CV, cs.CL | RA2G recipe | 2025.12 |
| arXiv(v1) 2025 | The Mental World of Large Language Models in Recommendation: A Benchmark on Association, Personalization, and Knowledgeability | Guangneng Hu | personalization, recommendation, benchmark, LLM | cs.IR | 21 pages, 13 figures, 27 tables, submission to KDD 2025 | 2025.12 |
| arXiv(v1) 2025 | Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues | Xiaotian Zhang, et al. | proactive personalization, profile customization, dialogue, LLM | cs.CL | 2025.12 | |
| TiiS25 | User Perceptions of Personalized and Generic Explanations in LLM-Driven Recommender Systems | Ítallo De Sousa Silva, et al. | LLM, recommender system, personalization, explanations, user study | 2025.12 | ||
| arXiv(v1) 2025 | Fixed-Persona SLMs with Modular Memory: Scalable NPC Dialogue on Consumer Hardware | Martin Braas, et al. | personalization, persona, modular memory, NPC | cs.AI, cs.IR | 2025.11 | |
| arXiv(v2) 2025 | Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction | Zhaopei Huang, et al. | personalization, dialogue, memory, long-term | cs.CL | Accepted by AAAI 2026 (Oral) | 2025.11 |
| arXiv(v1) 2025 | PLUM: Learning to Remember User Conversations for Personalization | Authors TBD | personalization, memory, conversation, LoRA | cs.CL, cs.AI | Parameter-efficient | 2025.11 |
| arXiv(v1) 2025 | PersonaAgent with GraphRAG: Community-Aware KG for Personalized LLM | Authors TBD | personalization, GraphRAG, knowledge graph, agent | cs.CL, cs.AI | 11.1% F1 improvement on LaMP | 2025.11 |
| arXiv(v1) 2025 | PersonalizedRouter: Personalized LLM Routing via Graph-based User Preference Modeling | Authors TBD | personalization, routing, GNN, user preference | cs.CL, cs.AI | 2025.11 | |
| arXiv(v1) 2025 | Profile-LLM: Dynamic Profile Optimization for Realistic Personality Expression | Authors TBD | personalization, profile, personality, dynamic | cs.CL, cs.AI | Education, therapy, entertainment | 2025.11 |
| arXiv(v1) 2025 | LLMDiRec: LLM-Enhanced Intent Diffusion for Sequential Recommendation | Bo-Chian Chen, et al. | intent diffusion, sequential recommendation, LLM | cs.IR | Under review | 2025.10 |
| arXiv(v1) 2025 | MemWeaver: A Hierarchical Memory from Textual Interactive Behaviors for Personalized Generation | Shuo Yu, et al. | personalization, memory, user behavior, generation | cs.CL | 12 pages, 8 figures | 2025.10 |
| arXiv(v1) 2025 | P2P: Instant Personalized LLM Adaptation via Hypernetwork | Authors TBD | personalization, hypernetwork, instant adaptation | cs.CL, cs.AI | Single-pass generation | 2025.10 |
| arXiv(v1) 2025 | Preference-Aware Memory Update for Long-Term LLM Agents | Haoran Sun, et al. | personalization, preference, memory update, long-term | cs.CL, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | RGMem: Renormalization Group-based Memory Evolution for Language Agent User Profile | Ao Tian, et al. | personalization, user profile, memory, agent | cs.AI | 11 pages,3 figures | 2025.10 |
| arXiv(v1) 2025 | Real-Time Personalization for LLM-based Recommendation with Customized ICL | Authors TBD | personalization, recommendation, ICL, real-time, online | cs.IR, cs.AI | No model update needed | 2025.10 |
| arXiv(v2) 2025 | CoPL: Collaborative Preference Learning for Personalizing LLMs | Youngbin Choi, et al. | collaborative preference learning, personalization, LLM | cs.LG, cs.AI, cs.IR | 19pages, 13 figures, 11 tables | 2025.09 |
| arXiv(v1) 2025 | DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models | Honghui Xu, et al. | federated learning, privacy-enhanced, on-device LLM, fine-tuning | cs.CR, cs.AI | 2025.09 | |
| arXiv(v1) 2025 | HumAIne-Chatbot: Real-Time Personalized Conversational AI via RL | Authors TBD | personalization, chatbot, RL, real-time, industry | cs.CL, cs.AI | Production deployment | 2025.09 |
| arXiv(v1) 2025 | MMPB: Multi-Modal Personalization Benchmark for VLMs | Authors TBD | personalization, MLLM, benchmark, VLM | cs.CV, cs.CL | First MLLM personalization benchmark | 2025.09 |
| arXiv(v1) 2025 | Personalized Reasoning: Just-In-Time Personalization and Why LLMs Fail At It | Shuyue Stella Li, et al. | just-in-time personalization, LLM, user preference | cs.CL, cs.AI | 57 pages, 6 figures | 2025.09 |
| RecSys25 | Revisiting Prompt Engineering: A Comprehensive Evaluation for LLM-based Personalized Recommendation | Genki Kusano, et al. | LLM, recommendation, personalization, prompt engineering | 2025.09 | ||
| arXiv(v1) 2025 | T-POP: Test-Time Personalization with Online Preference Feedback | Zikun Qu, et al. | test-time personalization, online preference, LLM | cs.LG, cs.AI | Preprint | 2025.09 |
| arXiv(v1) 2025 | DGDPO: Diagnostic-Guided Dynamic Profile Optimization for User Simulators | Authors TBD | personalization, user simulation, profile, dynamic | cs.IR, cs.AI | Bidirectional evolution | 2025.08 |
| arXiv(v1) 2025 | End-to-End Personalization: Unifying Recommender Systems with Large Language Models | Danial Ebrat, et al. | recommendation system, unifying, LLM | cs.IR, cs.LG | Second Workshop on Generative AI for Recommender Systems and Personalization at the ACM Conference on Knowledge Discovery and Data Mining (GenAIRecP@KDD 2025) | 2025.08 |
| arXiv(v1) 2025 | MLLMRec: MLLMs in Recommender Systems | Authors TBD | personalization, MLLM, recommendation, visual | cs.CV, cs.IR | Visual attribute extraction | 2025.08 |
| arXiv(v1) 2025 | MM-R1: Unified MLLMs for Personalized Image Generation | Authors TBD | personalization, MLLM, image generation, GRPO | cs.CV, cs.CL | X-CoT reasoning | 2025.08 |
| arXiv(v1) 2025 | MSPA: Multimodal Self-Corrective Preference Alignment for Recommendation | Authors TBD | personalization, MLLM, recommendation, self-corrective | cs.CV, cs.IR | 4D multimodal signals | 2025.08 |
| arXiv(v2) 2025 | Personalized LLM for Generating Customized Responses to the Same Query from Different Users | Hang Zeng, et al. | personalization, response generation, user-specific, LLM | cs.CL | Accepted by CIKM'25 | 2025.08 |
| arXiv(v1) 2025 | RLHF Fine-Tuning of LLMs for Alignment with Implicit User Feedback in Conversational Recommenders | Authors TBD | personalization, RLHF, implicit feedback, recommender | cs.IR, cs.AI | Dwell time, sentiment signals | 2025.08 |
| arXiv(v1) (SIGIR25) | CoT-Rec: Enhancing LLM-Based Recommendations Through Personalized Reasoning | Authors TBD | personalization, recommendation, CoT, reasoning | cs.IR, cs.AI | SIGIR 2025 | 2025.07 |
| arXiv(v1) 2025 | Comprehensive Review on LLMs for Recommender Systems | Authors TBD | personalization, recommendation, survey, LLM | cs.IR, cs.AI | Hybrid RAG approaches | 2025.07 |
| arXiv(v1) 2025 | DEP: Latent Inter-User Difference Modeling for LLM Personalization | Authors TBD | personalization, latent, embedding, difference-aware | cs.CL, cs.AI | Sparse autoencoder | 2025.07 |
| arXiv(v1) 2025 | PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training | Authors TBD | personalization, inference-time, alignment, preference | cs.CL, cs.AI | No reward model needed | 2025.07 |
| arXiv(v1) 2025 | PLUS: Learning to Summarize User Information for Personalized RLHF | Authors TBD | personalization, RLHF, user summary, preference | cs.LG, cs.AI | 11-77% reward model improvement | 2025.07 |
| arXiv(v1) 2025 | PRIME: LLM Personalization with Cognitive Memory and Thought Processes | Authors TBD | personalization, memory, episodic, semantic, cognitive | cs.CL, cs.AI | Dual-memory model | 2025.07 |
| arXiv(v1) 2025 | PURE: LLM-based User Profile Management for Recommender System | Authors TBD | personalization, user profile, recommendation, management | cs.IR, cs.AI | Profile extraction and updating | 2025.07 |
| arXiv(v1) 2025 | Personalization of Large Language Models: A Survey | Zhehao Zhang, et al. | personalization, survey, LLM, user profile | cs.CL, cs.AI | Comprehensive taxonomy | 2025.07 |
| arXiv(v2) 2025 | Comparison-based Active Preference Learning for Multi-dimensional Personalization | Minhyeon Oh, et al. | active preference learning, multi-dimensional, personalization, LLM | cs.LG | 2025.06 | |
| arXiv(v1) 2025 | PersonalAI: KG Storage and Retrieval for Personalized LLM Agents | Authors TBD | personalization, knowledge graph, memory, agent | cs.AI, cs.CL | Hybrid graph with hyperedges | 2025.06 |
| UMAP25 | Personalizing LLM Responses to Combat Political Misinformation | Adiba Proma, et al. | LLM, personalization, misinformation, user modeling | 2025.06 | ||
| arXiv(v1) 2025 | ProfiLLM: LLM-Based Framework for Implicit User Profiling | Authors TBD | personalization, profiling, implicit, chatbot | cs.CL, cs.AI | IT/cybersecurity domain | 2025.06 |
| arXiv(v1) 2025 | SEAL: Self-Adapting Language Models | Authors TBD | personalization, self-adaptation, online learning | cs.CL, cs.AI | Self-generated finetuning data | 2025.06 |
| arXiv(v3) 2025 | Drift: Decoding-time Personalized Alignments with Implicit User Preferences | Minbeom Kim, et al. | decoding-time alignment, implicit preferences, personalization, LLM | cs.CL | 19 pages, 6 figures | 2025.05 |
| arXiv(v2) 2025 | HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation | Cristina Garbacea, et al. | alignment, hypothesis generation, personalization, LLM | cs.CL | 2025.05 | |
| arXiv(v1) 2025 | MAP: Memory Assisted LLM for Personalized Recommendation System | Authors TBD | personalization, recommendation, memory, history | cs.IR, cs.AI | 2025.05 | |
| arXiv(v1) 2025 | PROSE: Aligning LLMs by Predicting Preferences from User Writing Samples | Authors TBD | personalization, preference prediction, writing samples | cs.CL, cs.AI | Iterative refinement | 2025.05 |
| arXiv(v1) 2025 | Privacy-preserving Prompt Personalization in Federated Learning for Multimodal Large Language Models | Sizai Hou, et al. | federated learning, privacy-preserving, prompt personalization, MLLM | cs.CR | Under Review | 2025.05 |
| arXiv(v1) 2025 | RLPA: Teaching LLMs to Evolve with Users via Dynamic Profile Modeling | Authors TBD | personalization, RLHF, dynamic profile, user evolution | cs.CL, cs.AI | Outperforms Claude-3.5, DeepSeek-V3 | 2025.05 |
| arXiv(v1) 2025 | Steerable Chatbots: Personalizing LLMs with Preference-Based Activation Steering | Authors TBD | personalization, chatbot, activation steering, inference | cs.CL, cs.AI | Training-free | 2025.05 |
| arXiv(v1) 2025 | Towards Explainable Temporal User Profiling with LLMs | Milad Sabouri, et al. | temporal user profiling, explainable, LLM | cs.IR, cs.AI | 2025.05 | |
| arXiv(v1) 2025 | Towards a unified user modeling language for engineering human centered AI systems | Aaron Conrardy, et al. | user modeling, LLM, personalization | cs.SE | Accepted at the Third Workshop on Engineering Interactive Systems Embedding AI Technologies (EISEAIT workshop at EICS 2025) | 2025.05 |
| arXiv(v1) 2025 | A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models | Authors TBD | personalization, preference alignment, survey, pluralistic | cs.CL, cs.AI | Training and inference-time methods | 2025.04 |
| arXiv(v2) 2025 | Differential Privacy Personalized Federated Learning Based on Dynamically Sparsified Client Updates | Chuanyin Wang, et al. | differential privacy, federated learning, personalization, LLM | cs.LG, cs.CR | 10 pages,2 figures | 2025.04 |
| arXiv(v1) 2025 | Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling | Authors TBD | personalization, benchmark, user profiling, dynamic | cs.CL, cs.AI | 2025.04 | |
| arXiv(v1) 2025 | LoRe: Personalizing LLMs via Low-Rank Reward Modeling | Avinandan Bose, et al. | low-rank reward modeling, personalization, LLM, RLHF | cs.LG, cs.AI, cs.CL | 2025.04 | |
| arXiv(v1) 2025 | PaRT: Enhancing Proactive Social Chatbots with Personalized Real-Time Retrieval | Authors TBD | personalization, chatbot, retrieval, real-time | cs.CL, cs.AI | 21.77% dialogue improvement, production deployed | 2025.04 |
| arXiv(v3) 2025 | Tuning-Free Personalized Alignment via Trial-Error-Explain In-Context Learning | Hyundong Cho, et al. | in-context learning, trial-error-explain, personalization, LLM | cs.CL, cs.AI | NAACL 2025 Findings | 2025.04 |
| arXiv(v1) 2025 | User Feedback Alignment for LLM-powered Exploration in Large-scale Recommendation | Authors TBD | personalization, recommendation, feedback, exploration | cs.IR, cs.AI | Click and dwell time signals | 2025.04 |
| arXiv(v1) 2025 | A Shared Low-Rank Adaptation Approach to Personalized RLHF | Renpu Liu, et al. | RLHF, low-rank adaptation, personalization, LLM | cs.LG, cs.AI | Published as a conference paper at AISTATS 2025 | 2025.03 |
| arXiv(v1) 2025 | Agentic Recommender Systems in the Era of Multimodal LLMs: Survey | Authors TBD | personalization, recommendation, agent, MLLM, survey | cs.IR, cs.AI | User agent simulation | 2025.03 |
| arXiv(v1) 2025 | BAHE: LLM-Enhanced CTR Prediction in Long Textual User Behaviors | Authors TBD | personalization, CTR, user behavior, industry | cs.IR, cs.AI | Deployed on 50M daily data | 2025.03 |
| arXiv(v1) 2025 | Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Online Shopping | Authors TBD | personalization, agent, user simulation, behavior | cs.AI, cs.HC | 31,865 shopping sessions | 2025.03 |
| arXiv(v1) 2025 | Language Model Personalization via Reward Factorization | Idan Shenfeld, et al. | reward factorization, RLHF, personalization, LLM | cs.LG | 2025.03 | |
| arXiv(v1) 2025 | Measuring What Makes You Unique: Difference-Aware User Modeling for LLM Personalization | Authors TBD | personalization, user modeling, difference-aware | cs.CL, cs.AI | 2025.03 | |
| arXiv(v7) 2025 | PAD: Personalized Alignment of LLMs at Decoding-Time | Ruizhe Chen, et al. | decoding-time alignment, personalization, LLM | cs.CL, cs.AI | ICLR 2025 | 2025.03 |
| arXiv(v1) 2025 | PersonaX: A Recommendation Agent Oriented User Modeling Framework | Authors TBD | personalization, recommendation, user modeling, agent | cs.IR, cs.AI | 3-11% improvement on AgentCF | 2025.03 |
| arXiv(v1) 2025 | A Survey of Personalized Large Language Models: Progress and Future Directions | Authors TBD | personalization, survey, LLM, prompting, finetuning | cs.CL, cs.AI | Input/model/objective level | 2025.02 |
| arXiv(v1) 2025 | FSPO: Few-Shot Preference Optimization of Synthetic Preference Data in LLMs Elicits Effective Personalization to Real Users | Anikait Singh, et al. | few-shot, preference optimization, personalization, LLM | cs.LG, cs.AI, cs.CL, cs.HC, stat.ML | Website: https://fewshot-preference-optimization.github.io/ | 2025.02 |
| arXiv(v1) 2025 | LoCoMo: Evaluating Very Long-Term Conversational Memory of LLM Agents | Authors TBD | personalization, memory, conversation, benchmark | cs.CL, cs.AI | 300 turns, 9K tokens, 35 sessions | 2025.02 |
| arXiv(v1) 2025 | PrefEval: Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following | Authors TBD | personalization, benchmark, preference, evaluation | cs.CL | 3000 preference-query pairs, 20 topics | 2025.02 |
| arXiv(v3) 2025 | Privacy-Preserving Personalized Federated Prompt Learning for Multimodal Large Language Models | Linh Tran, et al. | federated learning, privacy-preserving, personalized prompt, MLLM | cs.LG | 2025.02 | |
| arXiv(v1) 2025 | RLTHF: Targeted Human Feedback for LLM Alignment | Authors TBD | personalization, RLHF, targeted feedback, efficient | cs.CL, cs.AI | 6-7% human annotation effort | 2025.02 |
| arXiv(v3) 2025 | SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World | Jiaqi Zhang, et al. | personalization, agent, user modeling, embodied | cs.AI | 2025.02 | |
| arXiv(v1) 2025 | User Profile Construction and Updating with LLMs: Benchmark | Authors TBD | personalization, user profile, construction, updating | cs.CL, cs.AI | Static and dynamic profiling | 2025.02 |
| arXiv(v1) 2025 | When Personalization Meets Reality: Multi-Faceted Analysis of Personalized Preference Learning | Authors TBD | personalization, preference learning, fairness, evaluation | cs.CL, cs.AI | 2025.02 | |
| arXiv(v1) 2025 | Advancing Personalized Federated Learning: Integrative Approaches with AI for Enhanced Privacy and Customization | Kevin Cooper, et al. | federated learning, privacy, personalization, LLM | cs.LG, eess.SP | arXiv admin note: substantial text overlap with arXiv:2501.16758 | 2025.01 |
| arXiv(v2) 2025 | Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering | Rumi Allbert, et al. | personalization, personality, activation engineering, steering | cs.CL, cs.AI | 2025.01 | |
| arXiv(v1) 2025 | PerRecBench: Can LLMs Understand Preferences in Personalized Recommendation? | Authors TBD | personalization, benchmark, recommendation, preference | cs.IR, cs.CL | 19 LLMs evaluated | 2025.01 |
| arXiv(v2) 2025 | PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health | Huy Vu, et al. | personalization, personality, traits, mental health | cs.AI, cs.CL | 2025.01 | |
| arXiv(v1) 2024 | AI PERSONA: Towards Life-long Personalization of LLMs | Tiannan Wang, et al. | personalization, persona, lifelong, LLM | cs.CL, cs.AI | Work in progress | 2024.12 |
| arXiv(v1) 2024 | Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations | Sayantan Pal, et al. | personalization, persona, personality modeling, long-term | cs.CL, cs.AI | Accepted in COLING 2025 | 2024.12 |
| arXiv(v1) 2024 | Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits | Bohan Li, et al. | personalization, personality, dataset, MBTI | cs.CL, cs.CY | Accepted by COLING 2025. 28 papges, 20 figures, 10 tables | 2024.12 |
| arXiv(v1) 2024 | Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment | Jianfei Zhang, et al. | personalization, preference alignment, efficiency, LLM | cs.CL, cs.AI | Coling 2025 | 2024.12 |
| arXiv(v2) 2024 | Molar: Multimodal LLMs with Collaborative Filtering Alignment for Enhanced Sequential Recommendation | Yucong Luo, et al. | collaborative filtering, sequential recommendation, MLLM | cs.IR, cs.AI | 2024.12 | |
| arXiv(v1) 2024 | Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization | Guanghan Li, et al. | recommendation system, alignment, behavioral semantic, LLM | cs.IR, cs.AI, cs.CL | 7 pages, 3 figures, AAAI 2025 | 2024.12 |
| arXiv(v1) 2024 | ULMRec: User-centric Large Language Model for Sequential Recommendation | Minglai Shao, et al. | personalization, recommendation, sequential, LLM | cs.IR | 2024.12 | |
| arXiv(v1) 2024 | Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning | Sriyash Poddar, et al. | RLHF, variational preference learning, personalization, LLM | cs.LG, cs.AI, cs.CL, cs.RO | weirdlabuw.github.io/vpl | 2024.08 |
| arXiv(v2) 2024 | RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation | Chanwoo Park, et al. | RLHF, heterogeneous feedback, preference aggregation, personalization, LLM | cs.AI, cs.LG | Added experiments | 2024.05 |
Python
100.0%
Awesome-HCI (Ubiquitous, LLM, MLLM, Agent, RAG, Embodied-AI, RLHF)
Python
26
0 commits
updated Mar 8, 2026
A curated collection of research papers on HCI, LLM, MLLM, Agent, RAG, Agentic-RL, and Embodied AI (2021–present).
[Jan 2025] Added new sections: Agentic-RL and MLLM. Regular updates resumed.
python -m pip install -e . # Install CLI
paper add 2312.00752 LLM -t "llm, mamba" # Add paper
paper search transformer -t IMU # Search
paper stats # Statistics
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | Evaluating the Usage of African-American Vernacular English in Large Language Models | Deja Dunlap, et al. | LLM | cs.CL, cs.HC | 2026.02 | |
| arXiv(v1) 2026 | GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents | Shuofei Qiao, et al. | GUI agent, visual grounding, tool use, perception | cs.CV, cs.AI | 2026.01 | |
| arXiv(v1) 2025 (WACV 2026) | AFRAgent: An Adaptive Feature Renormalization Based High Resolution Aware GUI agent | Neeraj Anand, et al. | GUI agent, VLM, smartphone automation, multimodal | cs.CV | WACV 2026 | 2025.12 |
| arXiv(v2) 2025 | Mobile-Agent-v3: Fundamental Agents for GUI Automation | Junyang Wang, et al. | GUI agent, mobile, smartphone automation, VLM | cs.CV, cs.AI | 2025.08 | |
| arXiv(v6) 2025 | LLaVA-CoT: Let Vision Language Models Reason Step-by-Step | Guowei Xu, et al. | LLM, GUI agent, survey, computer use | cs.CV | 17 pages, ICCV 2025 | 2025.07 |
| arXiv(v3) 2025 | Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs | Kang Zhao, et al. | LLM, GUI agent, interface | cs.LG, cs.AI | 2025.06 | |
| arXiv(v3) 2025 | Lai Loss: A Novel Loss for Gradient Control | YuFei Lai | LLM, agent, interface, UI | cs.LG | The experiment in this article is not very rigorous and may require further testing for its effectiveness | 2025.05 |
| arXiv(v3) 2025 | SignLLM: Sign Language Production Large Language Models | Sen Fang, et al. | LLM, agent, user interface, HCI, interaction | cs.CV, cs.CL | website at https://signllm.github.io/ | 2025.04 |
| arXiv(v1) 2024 | ShowUI: One Vision-Language-Action Model for GUI Visual Agent | Kevin Qinghong Lin, et al. | LLM, UI, vision-language, GUI | cs.CV, cs.AI, cs.CL, cs.HC | Technical Report. Github: https://github.com/showlab/ShowUI | 2024.11 |
| arXiv(v1) 2024 | A Scalable Communication Protocol for Networks of Large Language Models | Samuele Marro, et al. | LLM, agent, GUI, computer use, interface | cs.AI, cs.LG | 2024.10 | |
| arXiv(v1) 2024 | In-Band Full-Duplex MIMO Systems for Simultaneous Communications and Sensing: Challenges, Methods, and Future Perspectives | Besma Smida, et al. | LLM, GUI agent, computer use | cs.IT, cs.ET, eess.SP | 12 pages, 5 figures, White Paper to appear at IEEE SPM | 2024.10 |
| arXiv(v2) 2024 | Massively parallel CMA-ES with increasing population | David Redon, et al. | LLM, agent, HCI, user interface | cs.DC | 2024.10 | |
| arXiv(v1) 2024 | OS-ATLAS: A Foundation Action Model for Generalist GUI Agents | Zhiyong Wu, et al. | LLM, agent, computer use, GUI, foundation model | cs.CL, cs.CV, cs.HC | 2024.10 | |
| arXiv(v1) 2024 | OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing Images | Jiaqi Zhao, et al. | LLM, agent, GUI, interface | cs.CV | The paper is accepted by IEEE Transactions on Geoscience and Remote Sensing (TGRS) | 2024.09 |
| arXiv(v1) 2024 | Unlocking the Power of Environment Assumptions for Unit Proofs | Siddharth Priya, et al. | LLM, GUI, agent, survey | cs.SE, cs.PL | SEFM 2024 | 2024.09 |
| arXiv(v2) 2024 (ACL 2025) | GUICourse: From General Vision Language Models to Versatile GUI Agents | Wentong Chen, et al. | GUI agent, VLM, training data, OCR, grounding | cs.CV, cs.CL, cs.HC | ACL 2025 | 2024.06 |
| arXiv(v1) 2024 | Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs | Keen You, et al. | ui, mllm, benchmark, any-resolution | cs.CV, cs.CL, cs.HC | 2024.04 | |
| arXiv(v1) 2024 | Are You Being Tracked? Discover the Power of Zero-Shot Trajectory Tracing with LLMs! | Huanqi Yang, et al. | iot, imu, cot, prompt | cs.CL, cs.AI, cs.HC, cs.LG | 2024.03 | |
| arXiv(v1) 2024 | Design2Code: How Far Are We From Automating Front-End Engineering? | Chenglei Si, et al. | llm, auto, google | cs.CL, cs.CV, cs.CY | 2024.03 | |
| arXiv(v2) 2023 | The Good, The Bad, and Why: Unveiling Emotions in Generative AI | Cheng Li, et al. | emotion, prompt, attack, decode | cs.AI, cs.CL, cs.HC | extension of Large language models understand and can be enhanced by emotional stimuli | 2023.12 |
| arXiv(v1) (NIPS23) | Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias | Yue Yu, et al. | synthetic data generation | cs.CL, cs.AI, cs.LG | arXiv(v2) 2023.10 | |
| arXiv(v1) 2023 | Multimodal Foundation Models: From Specialists to General-Purpose Assistants | Chunyuan Li, et al. | survey | cs.CV, cs.CL | 2023.09 | |
| arXiv(v7) 2023 | Attention Is All You Need | Ashish Vaswani, et al. | arxiv | cs.CL, cs.LG | 15 pages, 5 figures | 2023.08 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation | Zhengren Wang, et al. | RAG | cs.CV, cs.CL | 2026.02 | |
| arXiv(v1) 2026 | CLFEC: A New Task for Unified Linguistic and Factual Error Correction in paragraph-level Chinese Professional Writing | Jian Kai, et al. | RAG | cs.CL | 2026.02 | |
| arXiv(v1) 2026 | CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era | Zhengqing Yuan, et al. | RAG | cs.CL, cs.DL | 2026.02 | |
| arXiv(v1) 2026 | CiteLLM: An Agentic Platform for Trustworthy Scientific Reference Discovery | Mengze Hong, et al. | RAG | cs.CL, cs.IR | Accepted by TheWebConf 2026 Demo Track | 2026.02 |
| arXiv(v2) 2026 | MoDora: Tree-Based Semi-Structured Document Analysis System | Bangrui Xu, et al. | RAG | cs.IR, cs.AI, cs.CL, cs.DB, cs.LG | Extension of our SIGMOD 2026 paper. Please refer to source code available at https://github.com/weAIDB/MoDora | 2026.02 |
| arXiv(v1) 2026 | Search-P1: Path-Centric Reward Shaping for Stable and Efficient Agentic RAG Training | Tianle Xia, et al. | RAG | cs.CL, cs.IR, cs.LG | 2026.02 | |
| arXiv(v1) 2026 | TCM-DiffRAG: Personalized Syndrome Differentiation Reasoning Method for Traditional Chinese Medicine based on Knowledge Graph and Chain of Thought | Jianmin Li, et al. | RAG | cs.CL, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | TRIZ-RAGNER: A Retrieval-Augmented Large Language Model for TRIZ-Aware Named Entity Recognition in Patent-Based Contradiction Mining | Zitong Xu, et al. | RAG | cs.CL, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning | Chris Samarinas, et al. | RAG | cs.CL, cs.IR | 2026.02 | |
| arXiv(v1) 2026 | Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators | Zhengyang Su, et al. | RAG | cs.IR, cs.CL, cs.LG | 14 pages, 4 figures | 2026.02 |
| arXiv(v1) 2026 | L-RAG: Balancing Context and Retrieval with Entropy-Based Lazy Loading | Authors TBD | RAG, lazy loading, entropy, context | cs.CL, cs.IR | 2026.01 | |
| arXiv(v1) 2025 | RAGLens: Toward Faithful RAG with Sparse Autoencoders | Authors TBD | RAG, hallucination, faithfulness, detection | cs.CL | 2025.12 | |
| arXiv(v1) 2025 | CDTA: Cross-Document Topic-Aligned Chunking for RAG | Authors TBD | RAG, chunking, cross-document, topic alignment | cs.CL, cs.IR | 0.93 faithfulness on HotpotQA | 2025.11 |
| arXiv(v1) 2025 | Agentic RAG for Fintech: Design and Evaluation | Authors TBD | RAG, agentic, fintech, query reformulation | cs.CL, cs.IR | 2025.10 | |
| arXiv(v1) 2025 | Practical Code RAG at Scale: Task-Aware Retrieval Design | Authors TBD | RAG, code, retrieval, hybrid, dense | cs.CL, cs.IR | BM25 + dense hybrid | 2025.10 |
| arXiv(v1) 2025 | A Systematic Review of Key RAG Systems: Progress, Gaps, and Future Directions | Authors TBD | RAG, survey, systematic review, knowledge base | cs.CL, cs.IR | 2025.07 | |
| arXiv(v1) 2025 | Late Chunking: Contextual Chunk Embeddings for RAG | Authors TBD | RAG, chunking, embedding, contextual | cs.CL, cs.IR | Updated July 2025 | 2025.07 |
| arXiv(v1) 2025 | GraphRAG-Bench: When to Use Graphs in RAG | Authors TBD | RAG, graph, benchmark, evaluation | cs.CL, cs.IR | 2025.06 | |
| arXiv(v1) 2025 | RAG Survey: Architectures, Enhancements, and Robustness Frontiers | Authors TBD | RAG, survey, architecture, robustness | cs.CL, cs.IR | 2025.06 | |
| arXiv(v1) 2025 | Rethinking Chunk Size for Long-Document Retrieval: Multi-Dataset Analysis | Authors TBD | RAG, chunking, chunk size, retrieval | cs.CL, cs.IR | 64-1024 tokens optimal | 2025.05 |
| arXiv(v1) 2025 | A Survey of Multimodal Retrieval-Augmented Generation | Zihan Zhao, et al. | multimodal, RAG, retrieval, survey, vision-language | cs.CV, cs.CL | 2025.04 | |
| arXiv(v1) 2025 | RAG Evaluation in the Era of LLMs: A Comprehensive Survey | Authors TBD | RAG, evaluation, benchmark, LLM | cs.CL, cs.IR | 2025.04 | |
| arXiv(v1) 2025 | HiRAG: Retrieval-Augmented Generation with Hierarchical Knowledge | Authors TBD | RAG, hierarchical, knowledge, indexing | cs.CL, cs.IR | 2025.03 | |
| arXiv(v1) 2025 | Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation | Zihan Wang, et al. | multimodal, RAG, retrieval, survey, LLM | cs.CL, cs.IR | 2025.02 | |
| arXiv(v1) 2025 | GFM-RAG: Graph Foundation Model for Retrieval Augmented Generation | Authors TBD | RAG, graph, foundation model, knowledge graph | cs.CL, cs.IR | 8M params, 60 KGs, 14M triples | 2025.02 |
| arXiv(v1) 2025 | KG2RAG: Knowledge Graph-Guided Retrieval Augmented Generation | Authors TBD | RAG, knowledge graph, retrieval, fact-level | cs.CL, cs.IR | 2025.02 | |
| arXiv(v1) 2025 | RAG-Fusion: Query Expansion and Multi-Source Retrieval | Authors TBD | RAG, query expansion, multi-source, fusion | cs.CL, cs.IR | 2025.02 | |
| arXiv(v1) 2025 | Vendi-RAG: Adaptively Trading Off Diversity and Quality in RAG | Authors TBD | RAG, diversity, quality, adaptive | cs.CL, cs.IR | 2025.02 | |
| arXiv(v1) 2025 | Agentic RAG: A Survey | Authors TBD | RAG, agentic, survey, LLM agent | cs.CL, cs.IR | 2025.01 | |
| arXiv(v1) 2025 | CG-RAG: Citation Graph RAG for Research Question Answering | Authors TBD | RAG, citation graph, research QA | cs.CL, cs.IR | 2025.01 | |
| arXiv(v1) 2025 | ChunkRAG: Novel Context-Aware Chunking for RAG Systems | Authors TBD | RAG, chunking, context-aware, retrieval | cs.CL, cs.IR | 2025.01 | |
| arXiv(v1) 2024 | LLM-Augmented Retrieval: Enhancing Retrieval Models Through Language Models and Doc-Level Embedding | relevant query, doc-Level embedding, embedding-based retrieval, dense retrieval | 2024.04 | |||
| arXiv(v6) 2024 | Health-LLM: Personalized Retrieval-Augmented Disease Prediction System | Qinkai Yu, et al. | RAG, XGBoost, AutoML | cs.CL | 2024.03 | |
| arXiv(v1) 2024 | CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models | retrieval-augmented generation, large language models, evaluation | 2024.02 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | "Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems | Xinfeng Li, et al. | LLM, agent | cs.HC, cs.AI, cs.CR, cs.SI | 2026.02 | |
| arXiv(v1) 2026 | AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning | Yutong Wang, et al. | LLM, agent | cs.AI, cs.CL | 2026.02 | |
| arXiv(v1) 2026 | E3VA: Enhancing Emotional Expressiveness in Virtual Conversational Agents | Abhishek Kulkarni, et al. | LLM, agent | cs.HC | 5 pages | 2026.02 |
| arXiv(v1) 2026 | ESAA: Event Sourcing for Autonomous Agents in LLM-Based Software Engineering | Elzo Brito dos Santos Filho | LLM, agent | cs.AI | 13 pages, 1 figure, 4 tables. Includes 5 technical appendices | 2026.02 |
| arXiv(v1) 2026 | From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems | Yawen Wang, et al. | LLM, agent | cs.AI, cs.SE | 2026.02 | |
| arXiv(v1) 2026 | HotelQuEST: Balancing Quality and Efficiency in Agentic Search | Guy Hadad, et al. | LLM, agent | cs.IR, cs.AI | To be published in EACL 2026 | 2026.02 |
| arXiv(v1) 2026 | Hybrid LLM-Embedded Dialogue Agents for Learner Reflection: Designing Responsive and Theory-Driven Interactions | Paras Sharma, et al. | LLM, agent | cs.HC, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | ProductResearch: Training E-Commerce Deep Research Agents via Multi-Agent Synthetic Trajectory Distillation | Jiangyuan Wang, et al. | LLM, agent | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | PseudoAct: Leveraging Pseudocode Synthesis for Flexible Planning and Action Control in Large Language Model Agents | Yihan, et al. | LLM, agent | cs.AI, eess.SY | 2026.02 | |
| arXiv(v1) 2026 | SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems | Jialiang Fan, et al. | LLM, agent | cs.RO, cs.AI | 12 pages, 6 figures | 2026.02 |
| arXiv(v1) 2026 | The Auton Agentic AI Framework | Sheng Cao, et al. | LLM, agent | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training | Authors TBD | agent, web, GUI, training, environment | cs.CV, cs.AI | 600 tasks across websites | 2026.01 |
| arXiv(v1) 2026 | LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities | Authors TBD | code agent, software engineering, survey | cs.SE, cs.AI | 2026.01 | |
| arXiv(v1) 2025 | Beyond Task Completion: Assessment Framework for Agentic AI Systems | Authors TBD | LLM agent, evaluation, framework, benchmark | cs.AI, cs.CL | 2025.12 | |
| arXiv(v1) 2025 | DeepCode: Open Agentic Coding | Authors TBD | code agent, agentic coding, open source | cs.SE, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | LongVideoAgent: Multi-Agent Reasoning with Long Videos | Jingyi Zhang, et al. | multi-agent, video understanding, LLM, reasoning | cs.CV, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | MAR: Multi-Agent Reflexion Improves Reasoning Abilities in LLMs | Xinyuan Lu, et al. | multi-agent, LLM, reflexion, reasoning, debate | cs.CL, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | SWE-RL: Training Superintelligent Software Agents through Self-Play | Authors TBD | code agent, self-play, RL, software engineering | cs.SE, cs.LG | 2025.12 | |
| arXiv(v1) 2025 | Agent0: Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning | Authors TBD | LLM agent, self-evolving, tool use, reasoning | cs.AI, cs.CL | 18% improvement on math, 24% on general | 2025.11 |
| arXiv(v1) 2025 | Building Browser Agents: Architecture, Security, and Practical Solutions | Authors TBD | agent, browser, security, architecture | cs.AI, cs.CL | 2025.11 | |
| arXiv(v1) 2025 | ReMA: Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation | Zhiwei Zhang, et al. | multi-agent, LLM, reasoning, lazy agent, deliberation | cs.AI, cs.CL | 2025.11 | |
| arXiv(v1) 2025 | BrowserAgent: Web Agents with Human-Inspired Browsing Actions | Authors TBD | agent, browser, web, automation | cs.AI, cs.CL | 2025.10 | |
| arXiv(v1) 2025 | Dark Patterns Impact on LLM-Based Web Agents | Authors TBD | agent, web, dark patterns, decision making | cs.HC, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | Architecting Resilient LLM Agents | Authors TBD | LLM agent, resilience, architecture | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | A Survey on Code Generation with LLM-based Agents | Authors TBD | code agent, LLM, code generation, survey | cs.SE, cs.AI | 2025.08 | |
| arXiv(v1) 2025 | LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios | Bingxi Zhao, et al. | LLM, agent, reasoning, survey, framework | cs.AI, cs.CL | 2025.08 | |
| arXiv(v1) 2025 | MCP-Bench: Benchmarking Tool-Using LLM Agents | Authors TBD | agent, benchmark, tool use, MCP | cs.AI, cs.CL | 2025.08 | |
| arXiv(v4) 2025 | Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction | Harshitha Manoj, et al. | embodied AI, robot, LLM, VLM, autonomy, survey | cs.RO, cs.AI | 2025.08 | |
| arXiv(v1) 2025 | Understanding Tool-Integrated Reasoning | Heng Lin, et al. | LLM, agent, tool use, reasoning, TIR | cs.LG, cs.AI, stat.ML | 2025.08 | |
| arXiv(v2) 2025 | Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools | Junde Wu, et al. | LLM, agent, agentic reasoning, tool use, framework | cs.AI, cs.CL | ACL 2025 | 2025.07 |
| arXiv(v3) 2025 | Embodied AI Agents: Modeling the World | Pascale Fung, et al. | embodied AI, world model, VLM, robot, avatar | cs.AI | Meta AI | 2025.07 |
| arXiv(v1) 2025 | Routine: A Structural Planning Framework for LLM Agent System in Enterprise | Authors TBD | LLM agent, planning, enterprise, framework | cs.AI, cs.CL | 2025.07 | |
| arXiv(v1) 2025 | Toward a Theory of Agents as Tool-Use Decision-Makers | Authors TBD | LLM agent, tool use, theory, decision making | cs.AI, cs.CL | 2025.06 | |
| arXiv(v1) 2025 | WebRL: Training LLM Web Agents via Self-Evolving Curriculum RL | Authors TBD | agent, web, RL, curriculum, self-evolving | cs.LG, cs.AI | 2025.05 | |
| arXiv(v1) 2025 | AgentRewardBench: Benchmark for LLM Judges in Web Agent Evaluation | Authors TBD | agent, benchmark, web, evaluation, reward | cs.AI, cs.CL | 1,302 trajectories | 2025.04 |
| arXiv(v1) 2025 | From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review | Mohamed Amine Ferrag, et al. | LLM, agent, survey, autonomous, comprehensive review | cs.AI, cs.LG | 2025.04 | |
| arXiv(v1) 2024 | Simultaneous identification of the parameters in the plasticity function for power hardening materials : A Bayesian approach | Salih Tatar, et al. | LLM, agent, computer use, evaluation | math.NA, math.AP | 2024.12 | |
| arXiv(v2) 2024 | Feasibility Consistent Representation Learning for Safe Reinforcement Learning | Zhepeng Cen, et al. | LLM, agent, UI, LAUI, interface | cs.LG | ICML 2024 | 2024.06 |
| arXiv(v1) 2024 | DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows | Ajay Patel, et al. | pipeline, liarbry, generation | cs.CL, cs.LG | 2024.02 | |
| arXiv(v1) 2024 | More Agents Is All You Need | Junyou Li, et al. | multi agent, vote, task | cs.CL, cs.AI, cs.LG | 2024.02 | |
| arXiv(v1) (NIPS23) | HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face | Yongliang Shen, et al. | hugging face, API | cs.CL, cs.AI, cs.CV, cs.LG | 2023.12 | |
| arXiv(v1) (NIPS23) | CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society | Guohao Li, et al. | role play, autonomous, user&assistant | cs.AI, cs.CL, cs.CY, cs.LG, cs.MA | 2023.11(v2) | |
| arXiv(v2) 2023 | MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models | Dingyao Yu, et al. | pipeline | cs.CL, cs.MM, eess.AS | 2023.10 | |
| arXiv(v2) 2023 | VOYAGER: An Open-Ended Embodied Agent with Large Language Models | Guanzhi Wang, et al. | multi, autonomous, microcraft, game | cs.AI, cs.LG | 2023.10 | |
| arXiv(v3) 2023 | The Rise and Potential of Large Language Model Based Agents: A Survey | Zhiheng Xi, et al. | survey, github paper list | cs.AI, cs.CL | 2023.09 | |
| arXiv(v3) 2023 | A Survey on Large Language Model based Autonomous Agents | Lei Wang, et al. | survey, autonomous | cs.AI, cs.CL | Latest version is v4(2024.03), double columns. But v3(2023.09) single columns is easy to read. | |
| arXiv(v1) (ICLR24) | MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework | Sirui Hong, et al. | autonomous system, SOP, multi-agent, framework | cs.AI, cs.MA |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation | Weinan Dai, et al. | LLM, agentic-RL | cs.LG, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization | Zeyuan Liu, et al. | LLM, agentic-RL | cs.LG, cs.AI | Accepted to ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning | Zehao Li, et al. | LLM, agentic-RL | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | RF-Agent: Automated Reward Function Design via Language Agent Tree Search | Ning Gao, et al. | LLM, agentic-RL | cs.AI, cs.LG | 39 pages, 9 tables, 11 figures, Project page see https://github.com/deng-ai-lab/RF-Agent | 2026.02 |
| arXiv(v1) 2026 | RUMAD: Reinforcement-Unifying Multi-Agent Debate | Chao Wang, et al. | LLM, agentic-RL | cs.AI | 13 pages, 3 figures | 2026.02 |
| arXiv(v1) 2026 | Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance | Yanwei Ren, et al. | LLM, agentic-RL | cs.AI, cs.CL | 2026.02 | |
| arXiv(v2) 2026 | Regularized Online RLHF with Generalized Bilinear Preferences | Junghyun Lee, et al. | LLM, agentic-RL | cs.LG, stat.ML | 43 pages, 1 table (ver2: more colorful boxes, fixed some typos) | 2026.02 |
| arXiv(v1) 2026 | RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models | Daniel Yang, et al. | LLM, agentic-RL | cs.LG, cs.AI, cs.CL | 2026.02 | |
| arXiv(v1) 2026 | Agent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems Over Extended Interactions | Abhishek Rath | multi-agent, behavioral degradation, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v5) 2026 | AgentOrchestra: Orchestrating Multi-Agent Intelligence with the Tool-Environment-Agent(TEA) Protocol | Wentao Zhang, et al. | hierarchical multi-agent, task solving, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity | Doyoung Kim, et al. | evaluation, benchmark, API complexity, LLM, agent | cs.CL, cs.AI | 26 pages | 2026.01 |
| arXiv(v1) 2026 | Beyond Rule-Based Workflows: An Information-Flow-Orchestrated Multi-Agents Paradigm via Agent-to-Agent Communication from CORAL | Xinxing Ren, et al. | information flow, multi-agent communication, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v2) 2026 | Cochain: Balancing Insufficient and Excessive Collaboration in LLM Agent Workflows | Jiaxing Zhao, et al. | chain-of-collaboration, multi-agent, LLM, agent | cs.CL | 35 pages, 23 figures | 2026.01 |
| arXiv(v2) 2026 | Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes | Abhijnan Nath, et al. | alignment, multi-agent coordination, LLM, agent | cs.CL, cs.AI, cs.LG | This submission is a new version of arXiv:2509.05882v1. with a substantially revised experimental pipeline and new metrics. In particular, collaborator agents are now instantiated independently via separate API calls, rather than generated autoregressively by a single agent. All experimental results are new. Accepted as an extended abstract at AAMAS 2026 | 2026.01 |
| arXiv(v1) 2026 | DemMA: Dementia Multi-Turn Dialogue Agent with Expert-Guided Reasoning and Action Simulation | Yutong Song, et al. | dementia dialogue, multi-turn, medical, LLM, agent | cs.MA | 2026.01 | |
| arXiv(v1) 2026 | EduSim-LLM: An Educational Platform Integrating Large Language Models and Robotic Simulation for Beginners | Shenqi Lu, et al. | educational platform, robotics, simulation, LLM | cs.RO | 2026.01 | |
| arXiv(v1) 2026 | Game-Theoretic Lens on LLM-based Multi-Agent Systems | Jianing Hao, et al. | game theory, multi-agent, LLM, agent | cs.MA, cs.GT | 9 pages, 5 figures | 2026.01 |
| arXiv(v2) 2026 (Spotlight paper of NeurIPS 2025) | KARMA: Leveraging Multi-Agent LLMs for Automated Knowledge Graph Enrichment | Yuxing Lu, et al. | knowledge graph enrichment, multi-agent, LLM, agent | cs.CL, cs.AI, cs.CE, cs.DL | 24 pages, 3 figures, 2 tables | 2026.01 |
| arXiv(v1) 2026 | LLM-in-Sandbox Elicits General Agentic Intelligence | Daixuan Cheng, et al. | sandbox, agentic intelligence, LLM | cs.CL, cs.AI | Project Page: https://llm-in-sandbox.github.io | 2026.01 |
| arXiv(v2) 2026 | Lifelong Learning of Large Language Model based Agents: A Roadmap | Junhao Zheng, et al. | lifelong learning, continual learning, LLM, agent | cs.AI | Accepted to IEEE TPAMI | 2026.01 |
| arXiv(v1) 2026 | Nalar: An agent serving framework | Marco Laju, et al. | workflow serving, agent workflows, LLM, agent | cs.DC, cs.MA | 2026.01 | |
| arXiv(v1) 2026 | Orchestral AI: A Framework for Agent Orchestration | Alexander Roman, et al. | agent orchestration, framework, LLM, agent | cs.AI, astro-ph.IM, hep-ph | 17 pages, 3 figures. For more information visit https://orchestral-ai.com | 2026.01 |
| arXiv(v1) 2025 | PRL: Process Reward Learning Improves LLM Reasoning | Authors TBD | agentic RL, PRM, process reward, reasoning | cs.LG, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions | Aayush Gupta | reliability evaluation, stress testing, LLM, agent | cs.AI | 18 pages, 5 figures, 8 tables. Evaluates ReAct vs Reflexion across four tool-using domains with perturbation (epsilon) and fault-injection (lambda) stress testing; 1,280 total episodes | 2026.01 |
| arXiv(v2) 2026 | SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds | Jiawei Ren, et al. | simulation, autonomous agent, world model, LLM | cs.AI | 2026.01 | |
| arXiv(v2) 2026 | The Agentic Leash: Extracting Causal Feedback Fuzzy Cognitive Maps with LLMs | Akash Kumar Panda, et al. | causal extraction, fuzzy cognitive maps, LLM, agent | cs.AI, cs.CL, cs.HC, cs.IR | 15 figures | 2026.01 |
| arXiv(v2) 2026 | The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning | Qiguang Chen, et al. | chain-of-thought, reasoning topology, LLM | cs.CL, cs.AI | Preprint | 2026.01 |
| arXiv(v1) 2026 | The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs | Zibo Zhao, et al. | self-reflection, reinforcement learning, LLM, agent | cs.LG, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents | Alessio Buscemi, et al. | numerical coordination, multi-agent, LLM, agent | cs.MA, cs.AI | 2026.01 | |
| arXiv(v2) 2025 | DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems | Ming Ma, et al. | auto debugging, multi-agent, intervention, LLM, agent | cs.AI, cs.SE | 2025.12 | |
| arXiv(v1) 2025 | Enhancing Agentic RL with Progressive Reward Shaping and VSPO | Authors TBD | agentic RL, reward shaping, GRPO, tool use | cs.LG, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | From Word to World: Can Large Language Models be Implicit Text-based World Models? | Yixia Li, et al. | world model, text-based, LLM, agent | cs.CL | 2025.12 | |
| arXiv(v2) 2025 | ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving | Haoyuan Wu, et al. | tree-of-thought, reasoning, puzzle solving, LLM | cs.CL | 2025.12 | |
| arXiv(v2) 2025 | Understanding LLM Agent Behaviours via Game Theory: Strategy Recognition, Biases and Multi-Agent Dynamics | Trung-Kiet Huynh, et al. | game theory, agent behavior, LLM, agent | cs.MA, cs.AI, cs.GT, cs.LG, math.DS | 2025.12 | |
| arXiv(v2) 2025 | Large Language Model-based Data Science Agent: A Survey | Ke Chen, et al. | data science, agent, LLM | cs.AI | 2025.11 | |
| arXiv(v1) 2025 | MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism | Shulin Liu, et al. | multi-agent, agentic RL, reasoning, LLM, pipeline parallelism | cs.AI | 10 pages | 2025.11 |
| arXiv(v1) 2025 | AGENTRL: Scaling Agentic Reinforcement Learning | Authors TBD | agentic RL, RLHF, LLM agent, scaling | cs.LG, cs.AI | Outperforms GPT-5 and Claude-Sonnet-4 | 2025.10 |
| arXiv(v2) 2025 | BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks | Sagnik Anupam, et al. | web navigation, evaluation, real-world, LLM, agent | cs.AI, cs.LG | 2025.10 | |
| arXiv(v1) 2025 | Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning | Marwa Abdulhai, et al. | persona simulation, multi-turn RL, LLM, agent | cs.CL, cs.AI | 2025.10 | |
| arXiv(v3) 2025 | DS-STAR: Data Science Agent via Iterative Planning and Verification | Jaehyun Nam, et al. | data science, iterative planning, verification, LLM, agent | cs.AI | 2025.10 | |
| arXiv(v1) 2025 | Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments | Crystal Qian, et al. | social experiments, human-AI interaction, LLM, agent | cs.HC, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | Demystifying Reinforcement Learning in Agentic Reasoning | Zhaochen Yu, et al. | agentic RL, reasoning, LLM, agent | cs.CL | Code and models: https://github.com/Gen-Verse/Open-AgentRL | 2025.10 |
| arXiv(v3) 2025 | DoctorAgent-RL: A Multi-Agent Collaborative Reinforcement Learning System for Multi-Turn Clinical Dialogue | Yichun Feng, et al. | clinical dialogue, multi-agent RL, medical, LLM, agent | cs.CL | 2025.10 | |
| arXiv(v1) 2025 | GEM: A Gym for Agentic LLMs | Zichen Liu, et al. | training environment, agentic LLM, gym, agent | cs.LG, cs.AI, cs.CL | 2025.10 | |
| arXiv(v3) 2025 | MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research | Hui Chen, et al. | ML research, evaluation, benchmark, LLM, agent | cs.LG, cs.AI, cs.CL | 49 pages, 9 figures. Accepted by NeurIPS 2025 D&B Track | 2025.10 |
| arXiv(v1) 2025 | Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents | Reid T. Johnson, et al. | natural language tool calling, LLM, agent | cs.CL | 31 pages, 7 figures | 2025.10 |
| arXiv(v1) 2025 | On Designing Effective RL Reward at Training Time for LLM Reasoning | Authors TBD | agentic RL, reward design, reasoning, training | cs.LG, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | RLSR: Reinforcement Learning with Supervised Reward | Authors TBD | agentic RL, supervised reward, instruction following | cs.LG, cs.AI | 2025.10 | |
| arXiv(v2) 2025 | SimuRA: A World-Model-Driven Simulative Reasoning Architecture for General Goal-Oriented Agents | Mingkai Deng, et al. | world model, simulative reasoning, goal-oriented, LLM, agent | cs.AI, cs.CL, cs.LG, cs.RO | This submission has been updated to adjust the scope and presentation of the work | 2025.10 |
| arXiv(v2) 2025 (Am. Statist. (2025) 1-14) | A Survey on Large Language Model-based Agents for Statistics and Data Science | Maojun Sun, et al. | data science, statistics, survey, LLM, agent | cs.AI, cs.CL, cs.LG, stat.OT | 2025.09 | |
| arXiv(v1) 2025 | OPPO: Accelerating PPO-based RLHF via Pipeline Overlap | Authors TBD | agentic RL, PPO, RLHF, efficiency, overlap | cs.LG, cs.AI | 2025.09 | |
| arXiv(v1) 2025 | RL Foundations for Deep Research Systems: A Survey | Authors TBD | agentic RL, deep research, survey, post-DeepSeek | cs.LG, cs.AI | Post Feb 2025 papers | 2025.09 |
| arXiv(v1) 2025 | Reward Hacking Mitigation using Verifiable Composite Rewards | Authors TBD | RLHF, reward hacking, RLVR, verification | cs.LG, cs.AI | 2025.09 | |
| arXiv(v1) 2025 | SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection | Yubin Ge, et al. | self-learning, multi-level reflection, LLM, agent | cs.AI | Accepted at EMNLP 2025 Main Conference | 2025.09 |
| arXiv(v1) 2025 | Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning | Pulkit Verma, et al. | logical chain-of-thought, symbolic planning, LLM | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | The Landscape of Agentic Reinforcement Learning for LLMs: A Survey | Authors TBD | agentic RL, survey, LLM, POMDP | cs.LG, cs.AI | 2025.09 | |
| arXiv(v1) 2025 | Where LLM Agents Fail and How They can Learn From Failures | Kunlun Zhu, et al. | failure detection, learning from failures, LLM, agent | cs.AI | 2025.09 | |
| arXiv(v1) 2025 | iStar: Agentic Reinforcement Learning with Implicit Step Rewards | Authors TBD | agentic RL, credit assignment, implicit PRM | cs.LG, cs.AI | 2025.09 | |
| arXiv(v3) 2025 | MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement | Jaehyun Nam, et al. | ML engineering, code generation, search, LLM, agent | cs.LG | 2025.08 | |
| arXiv(v1) 2025 | Agent Safety Alignment via Reinforcement Learning | Zeyang Sha, et al. | safety alignment, reinforcement learning, LLM, agent | cs.AI, cs.CR | 2025.07 | |
| arXiv(v1) 2025 | AgentMesh: A Cooperative Multi-Agent Generative AI Framework for Software Development Automation | Sourena Khanzadeh | multi-agent, software development, code generation, LLM, agent | cs.SE, cs.AI | 2025.07 | |
| arXiv(v1) 2025 | Technical Survey of RL Techniques for Large Language Models | Authors TBD | agentic RL, survey, PPO, DPO, GRPO | cs.LG, cs.AI | 2025.07 | |
| arXiv(v2) 2025 | ToolACE: Winning the Points of LLM Function Calling | Weiwen Liu, et al. | function calling, tool use, LLM, agent | cs.LG, cs.AI, cs.CL | 21 pages, 22 figures | 2025.07 |
| arXiv(v1) 2025 | A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy | Henry Peng Zou, et al. | human-agent systems, collaboration, LLM, agent | cs.AI, cs.CL, cs.HC, cs.LG, cs.MA | 2025.06 | |
| arXiv(v2) 2025 | Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems | Bingyu Yan, et al. | multi-agent communication, survey, LLM, agent | cs.MA, cs.CL | 2025.06 | |
| arXiv(v1) 2025 | Enhancing Decision-Making of Large Language Models via Actor-Critic | Heng Dong, et al. | decision-making, actor-critic, LLM, agent | cs.CL, cs.AI | Forty-second International Conference on Machine Learning (ICML 2025) | 2025.06 |
| arXiv(v1) 2025 | GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation | Ning Gao, et al. | simulation, robotic manipulation, LLM, agent | cs.RO | 2025.06 | |
| arXiv(v1) 2025 | OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems | Xiaozhe Li, et al. | optimization, benchmark, evaluation, LLM, agent | cs.AI, cs.LG | 2025.06 | |
| arXiv(v1) 2025 | RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments | Yuchuan Fu, et al. | security evaluation, benchmark, LLM, agent | cs.CR, cs.AI | 12 pages, 8 figures | 2025.06 |
| arXiv(v1) 2025 | Sailing by the Stars: Survey on Reward Models and Learning Strategies | Authors TBD | agentic RL, reward model, survey, learning | cs.LG, cs.AI | 2025.06 | |
| arXiv(v1) 2025 | ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering | Zexi Liu, et al. | agentic RL, ML engineering, LLM, agent | cs.CL, cs.AI, cs.LG | 2025.05 | |
| arXiv(v1) 2025 | Multi-Agent Systems for Robotic Autonomy with LLMs | Junhong Chen, et al. | multi-agent, robotics, LLM, agent | cs.RO, cs.AI | 11 pages, 2 figures, 5 tables, submitted for publication | 2025.05 |
| arXiv(v1) 2025 | Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking | Yihan Chen, et al. | synthetic trajectories, self-reflection, training, LLM, agent | cs.CL | 2025.05 | |
| arXiv(v1) 2025 | Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning | Joykirat Singh, et al. | agentic RL, tool integration, reasoning, LLM | cs.AI | 2025.04 | |
| arXiv(v1) 2025 | Comprehensive Survey of Reward Models: Taxonomy and Applications | Authors TBD | agentic RL, reward model, survey, taxonomy | cs.LG, cs.AI | 2025.04 | |
| arXiv(v1) 2025 | DPO Meets PPO: Reinforced Token Optimization for RLHF | Han Zhong, et al. | agentic RL, DPO, PPO, RLHF, token-level | cs.LG, cs.AI | RTO framework | 2025.04 |
| arXiv(v1) 2025 | Hierarchical Multi-Step Reward Models for Enhanced Reasoning | Authors TBD | agentic RL, reward model, hierarchical, reasoning | cs.LG, cs.AI | 2025.03 | |
| arXiv(v1) 2025 | Look Before You Leap: Using Serialized State Machine for Language Conditioned Robotic Manipulation | Tong Mu, et al. | finite state machine, robotic manipulation, LLM | cs.RO, cs.AI | 7 pages, 4 figures | 2025.03 |
| arXiv(v1) 2025 | SafePlan: Leveraging Formal Logic and Chain-of-Thought Reasoning for Enhanced Safety in LLM-based Robotic Task Planning | Ike Obi, et al. | formal logic, chain-of-thought, safety, robotic planning, LLM | cs.RO | 2025.03 | |
| arXiv(v2) 2025 | Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation | Hyungjoo Chae, et al. | web navigation, environment dynamics, LLM, agent | cs.CL | ICLR 2025 | 2025.03 |
| arXiv(v1) 2025 | Every Software as an Agent: Blueprint and Case Study | Mengwei Xu | software agent, autonomous, LLM | cs.SE, cs.AI | 2025.02 | |
| arXiv(v2) 2025 | Flow: Modularized Agentic Workflow Automation | Boye Niu, et al. | workflow generation, modular, LLM, agent | cs.AI, cs.LG, cs.MA | 2025.02 | |
| arXiv(v1) 2025 | Policy Learning with a Natural Language Action Space: A Causal Approach | Bohan Zhang, et al. | policy learning, natural language action, LLM, agent | cs.CL | 2025.02 | |
| arXiv(v1) 2025 | Process Reward Models for LLM Agents: Practical Framework | Authors TBD | PRM, reward model, LLM agent, RLHF | cs.LG, cs.AI | InversePRM | 2025.02 |
| arXiv(v1) 2025 | Provably Efficient Online RLHF with One-Pass Reward Modeling | Authors TBD | online RLHF, reward modeling, efficiency | cs.LG, cs.AI | 2025.02 | |
| arXiv(v1) 2025 (Proceedings of the 2024 IEEE International Japan-Africa Conference on Electronics communications and Computations (JAC ECC)) | Guided Code Generation with LLMs: A Multi-Agent Framework for Complex Code Tasks | Amr Almorsi, et al. | multi-agent, code generation, LLM, agent | cs.AI | 4 pages, 3 figures | 2025.01 |
| arXiv(v1) 2025 | REINFORCE++: Critic-Free Policy Optimization with Global Advantage Normalization | Authors TBD | agentic RL, REINFORCE, critic-free, GRPO | cs.LG, cs.AI | Outperforms PPO | 2025.01 |
| arXiv(v4) 2024 | Planning with Multi-Constraints via Collaborative Language Agents | Cong Zhang, et al. | meta-task planning, multi-agent, LLM, agent | cs.AI, cs.CL, cs.LG | 2024.12 | |
| arXiv(v2) 2024 | AutoWebGLM: A Large Language Model-based Web Navigating Agent | Hanyu Lai, et al. | web navigation, browsing, LLM, agent | cs.CL | Accepted to KDD 2024 | 2024.10 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v2) 2026 | A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation | Weiayn Shi, et al. | VLM | cs.HC | Accepted at CHI 2026 Posters | 2026.02 |
| arXiv(v1) 2026 | Beyond Dominant Patches: Spatial Credit Redistribution For Grounded Vision-Language Models | Niamul Hassan Samin, et al. | VLM | cs.CV, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Beyond Static Artifacts: A Forensic Benchmark for Video Deepfake Reasoning in Vision Language Models | Zheyuan Gu, et al. | VLM | cs.CV, cs.AI | 16 pages, 9 figures. Submitted to CVPR 2026 | 2026.02 |
| arXiv(v1) 2026 | Causal Decoding for Hallucination-Resistant Multimodal Large Language Models | Shiwei Tan, et al. | VLM | cs.LG, cs.AI, cs.CV | Published in Transactions on Machine Learning Research (TMLR), 2026 | 2026.02 |
| arXiv(v1) 2026 | Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models | Jianghao Yin, et al. | VLM | cs.CV, cs.AI | Accepted by ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models | Xingyu Zhu, et al. | VLM | cs.CV, cs.MM | ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | HulluEdit: Single-Pass Evidence-Consistent Subspace Editing for Mitigating Hallucinations in Large Vision-Language Models | Yangguang Lin, et al. | VLM | cs.CV | accepted at CVPR 2026 | 2026.02 |
| arXiv(v1) 2026 | Large Multimodal Models as General In-Context Classifiers | Marco Garosi, et al. | VLM | cs.CV | CVPR Findings 2026. Project website at https://circle-lmm.github.io/ | 2026.02 |
| arXiv(v1) 2026 | Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation | Xingyu Zhu, et al. | VLM | cs.CV | ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | MediX-R1: Open Ended Medical Reinforcement Learning | Sahal Shaji Mullappilly, et al. | VLM | cs.CV | 2026.02 | |
| arXiv(v1) 2026 | NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors | Lingfeng Ren, et al. | VLM | cs.CV, cs.AI, cs.CL | Code: https://github.com/lingfengren/NoLan | 2026.02 |
| arXiv(v1) 2026 | See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs | Yongchang Zhang, et al. | VLM | cs.CV | CVPR2026 Accepted | 2026.02 |
| arXiv(v1) 2026 | Seeing Graphs Like Humans: Benchmarking Computational Measures and MLLMs for Similarity Assessment | Seokweon Jung, et al. | VLM | cs.HC | 21 pages including 1 page of appendix, 9 figures, 4 tables | 2026.02 |
| arXiv(v1) 2026 | Suppressing Prior-Comparison Hallucinations in Radiology Report Generation via Semantically Decoupled Latent Steering | Ao Li, et al. | VLM | cs.CV | 15 pages, 5 figures | 2026.02 |
| arXiv(v1) 2026 | SurGo-R1: Benchmarking and Modeling Contextual Reasoning for Operative Zone in Surgical Video | Guanyi Qin, et al. | VLM | cs.CV, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification | Vikash Singh, et al. | VLM | cs.CV, cs.AI, cs.CL, cs.LO | 2026.02 | |
| arXiv(v1) 2026 | VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation | Seongheon Park, et al. | VLM | cs.CV, cs.AI, cs.CL | 2026.02 | |
| arXiv(v1) 2026 | Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback | Authors TBD | MLLM, hallucination, caption feedback, grounding | cs.CV, cs.CL | 2026.01 | |
| arXiv(v1) 2026 | Innovator-VL: A Multimodal Large Language Model for Scientific Discovery | Zichen Wen, et al. | scientific discovery, MLLM, vision-language | cs.CV, cs.AI | Innovator-VL tech report | 2026.01 |
| arXiv(v2) 2026 | MGPC: Multimodal Network for Generalizable Point Cloud Completion With Modality Dropout and Progressive Decoding | Jiangyuan Liu, et al. | point cloud completion, multimodal, MLLM | cs.CV | Code and dataset are available at https://github.com/L-J-Yuan/MGPC | 2026.01 |
| arXiv(v1) 2026 | Multimodal In-context Learning for ASR of Low-resource Languages | Zhaolin Li, et al. | multimodal in-context learning, ASR, low-resource, MLLM | cs.CL, cs.AI | Under review | 2026.01 |
| arXiv(v3) 2026 | Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models | Umberto Cappellazzo, et al. | unified speech recognition, multimodal, MLLM | eess.AS, cs.CV, cs.SD | Accepted to IEEE ICASSP 2026 (camera-ready version). Project website (code and model weights): https://umbertocappellazzo.github.io/Omni-AVSR/ | 2026.01 |
| arXiv(v2) 2026 | Table as a Modality for Large Language Models | Liyao Li, et al. | table modality, structured data, MLLM | cs.CL, cs.AI | Accepted to NeurIPS 2025 | 2026.01 |
| arXiv(v1) 2026 | The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection | Wei Ai, et al. | fake news detection, vision-language, MLLM, survey | cs.AI, cs.CV | 2026.01 | |
| arXiv(v3) 2026 | UniVideo: Unified Understanding, Generation, and Editing for Videos | Cong Wei, et al. | video understanding, generation, editing, unified, MLLM | cs.CV | Project Website https://congwei1230.github.io/UniVideo/ | 2026.01 |
| arXiv(v1) 2026 | VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory | Shaoan Wang, et al. | embodied navigation, adaptive reasoning, MLLM, vision-language | cs.RO, cs.CV | Project page: https://wsakobe.github.io/VLingNav-web/ | 2026.01 |
| arXiv(v1) 2026 | VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding | Jiapeng Shi, et al. | video understanding, spatial-temporal, MLLM, vision-language | cs.CV | 2026.01 | |
| arXiv(v1) 2025 | A Medical Multimodal Diagnostic Framework Integrating Vision-Language Models and Logic Tree Reasoning | Zelin Zang, et al. | medical diagnosis, logic tree reasoning, vision-language, MLLM | cs.AI | 2025.12 | |
| arXiv(v1) 2025 | DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models | Zefeng He, et al. | generative reasoning, diffusion model, MLLM | cs.CV | Project page: https://diffthinker-project.github.io | 2025.12 |
| arXiv(v1) 2025 | From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs | Authors TBD | MLLM, spatial reasoning, benchmark, open world | cs.CV, cs.CL | 2025.12 | |
| arXiv(v1) 2025 | Kling-Omni Technical Report | Kling Team, et al. | video generation, multimodal synthesis, MLLM | cs.CV | Kling-Omni Technical Report | 2025.12 |
| arXiv(v1) 2025 | Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding | Yongyuan Liang, et al. | 3D understanding, point cloud, spatial understanding, MLLM | cs.CV, cs.AI | 2025.12 | |
| arXiv(v3) 2025 (Proc. 2025 IEEE 8th International Conference on Multimedia Information Processing and Retrieval (MIPR), pp. 456-462, 2025) | MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models | Philip R. Liu, et al. | multi-agent, medical diagnosis, MLLM | cs.MA, cs.AI, cs.CV, cs.LG | 2025.12 | |
| arXiv(v2) 2025 | TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning | Tao Wu, et al. | temporal understanding, reinforcement learning, MLLM, video | cs.CV | 2025.12 | |
| arXiv(v1) 2025 | MVU-Eval: Multi-Video Understanding Evaluation for MLLMs | Authors TBD | MLLM, multi-video, evaluation, benchmark | cs.CV, cs.CL | 2025.11 | |
| arXiv(v1) 2025 (Proceedings of the Conference on Language Modeling (COLM 2025)) | REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories | Jacob Thompson, et al. | embodied spatial reasoning, trajectory, MLLM | cs.LG, cs.AI, cs.CV | 2025.11 | |
| arXiv(v1) 2025 | Seeing is Believing: Rich-Context Hallucination Detection via Backward Visual Grounding | Authors TBD | MLLM, hallucination, detection, visual grounding | cs.CV, cs.CL | Outperforms GPT-4o | 2025.11 |
| arXiv(v1) 2025 | SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards | Authors TBD | MLLM, 3D reasoning, spatial, RL | cs.CV, cs.CL | Outperforms GPT-4o | 2025.11 |
| arXiv(v1) 2025 | MT-Video-Bench: Video Understanding Benchmark for MLLMs in Multi-Turn Dialogues | Authors TBD | MLLM, video understanding, benchmark, multi-turn | cs.CV, cs.CL | 2025.10 | |
| arXiv(v1) 2025 | MemVR: Memory-Space Visual Retracing for Hallucination Mitigation in MLLMs | Authors TBD | MLLM, hallucination, mitigation, memory | cs.CV, cs.CL | Plug-and-play | 2025.10 |
| arXiv(v2) 2025 | Revealing Multimodal Causality with Large Language Models | Jin Li, et al. | causal discovery, MLLM, multimodal causality | cs.LG, cs.AI | Accepted at NeurIPS 2025 | 2025.10 |
| arXiv(v1) 2025 | Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph | Wentao Wang, et al. | spatio-temporal reasoning, relation graph, MLLM, video | cs.AI | 2025.10 | |
| arXiv(v1) 2025 | Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs | Authors TBD | MLLM, hallucination, omission, fabrication | cs.CV, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | VIRAL: Visual Representation Alignment for MLLMs | Authors TBD | MLLM, visual alignment, fine-grained understanding | cs.CV, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents | Han Lin, et al. | diffusion model, patch-level CLIP, image generation, MLLM | cs.CV, cs.AI, cs.CL | Project Page: https://bifrost-1.github.io | 2025.08 |
| arXiv(v1) 2025 | Grounding the Ungrounded: Spectral-Graph Framework for Quantifying Hallucinations in MLLMs | Authors TBD | MLLM, hallucination, grounding, detection | cs.CV, cs.CL | 2025.08 | |
| arXiv(v1) 2025 | Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey | Authors TBD | MLLM, VLA, robotics, manipulation, survey | cs.RO, cs.CV | 2025.08 | |
| arXiv(v1) 2025 | Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting | Miaosen Luo, et al. | affective computing, emotion recognition, MLLM | cs.AI, cs.LG | 2025.08 | |
| arXiv(v1) 2025 | RynnEC: Bringing MLLMs into Embodied World | Authors TBD | MLLM, embodied, video, spatial reasoning | cs.CV, cs.RO | 2025.08 | |
| arXiv(v3) 2025 | SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding | Can Liu, et al. | visualization understanding, dataset, MLLM | cs.HC, cs.CV | 2025.07 | |
| arXiv(v2) 2025 | UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation | Yanzhe Chen, et al. | codebook, multimodal generation, MLLM | cs.CV, cs.MM | 19 pages, 5 figures | 2025.07 |
| arXiv(v1) 2025 | CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning | Kailing Li, et al. | embodied visual reasoning, cognitive map, MLLM | cs.CV, cs.AI, cs.CL | 2025.06 | |
| arXiv(v1) 2025 | Insight-V: Exploring Long-Chain Visual Reasoning with MLLMs | Authors TBD | MLLM, visual reasoning, long-chain, CVPR | cs.CV | CVPR 2025 | 2025.06 |
| arXiv(v2) 2025 | LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning | Zebin You, et al. | diffusion model, visual instruction tuning, MLLM | cs.LG, cs.CL, cs.CV | Project page and codes: \url{https://ml-gsai.github.io/LLaDA-V-demo/} | 2025.06 |
| arXiv(v2) 2025 | LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding | Hongyu Li, et al. | spatial-temporal understanding, MLLM, vision-language | cs.CV | Accepted by CVPR2025 | 2025.06 |
| arXiv(v1) 2025 (CVPR 2025) | LLaVA-ST: Multimodal LLM for Fine-Grained Spatial-Temporal Understanding | Authors TBD | MLLM, spatial-temporal, video, CVPR | cs.CV | CVPR 2025 | 2025.06 |
| arXiv(v1) 2025 | Manager: Aggregating Insights from Unimodal Experts in VLMs and MLLMs | Authors TBD | MLLM, VLM, unimodal experts, fusion | cs.CV, cs.CL | 2025.06 | |
| arXiv(v1) 2025 | MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis | Yuting Zhang, et al. | medical reasoning, diagnosis, MLLM | eess.IV, cs.CL, cs.CV, q-bio.QM | 2025.06 | |
| arXiv(v1) 2025 | Multimodal Tabular Reasoning with Privileged Structured Information | Jun-Peng Jiang, et al. | tabular reasoning, structured information, MLLM | cs.LG, cs.AI, cs.CL, cs.CV | 2025.06 | |
| arXiv(v1) 2025 | Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models | Hugues Thomas, et al. | 3D scene understanding, point cloud, token structure, MLLM | cs.CV | Main paper and appendix | 2025.06 |
| CVPR25 (CVPR 2025) | Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Reweighting | Authors TBD | MLLM, hallucination, attention, CVPR | cs.CV | CVPR 2025 | 2025.06 |
| arXiv(v3) 2025 | SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models | Wufei Ma, et al. | 3D-informed, spatial intelligence, MLLM | cs.CV | CVPR 2025 highlight | 2025.06 |
| arXiv(v1) 2025 | Structured Attention Matters to Multimodal LLMs in Document Understanding | Chang Liu, et al. | document understanding, structured attention, MLLM | cs.CL, cs.AI, cs.IR | 2025.06 | |
| arXiv(v2) 2025 | Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM | Zinuo Li, et al. | audio-visual-speech, multimodal, MLLM, video understanding | cs.CL | 2025.06 | |
| arXiv(v3) 2025 | Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning | NVIDIA, et al. | physical common sense, embodied reasoning, MLLM | cs.AI, cs.CV, cs.LG, cs.RO | 2025.05 | |
| arXiv(v1) 2025 | HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and Reasoning | Chuhao Zhou, et al. | multisensory, human sensing, reasoning, MLLM | cs.CV, cs.AI, cs.CL, cs.LG, cs.MM | 18 pages, 13 figures, 6 tables | 2025.05 |
| arXiv(v2) 2025 | Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation | Liu He, et al. | video generation, multi-agent collaboration, MLLM | cs.CV, cs.GR, cs.MM | Accepted by CVPR 2025 AI4CC Workshop | 2025.05 |
| arXiv(v1) 2025 | MedBridge: Bridging Foundation Vision-Language Models to Medical Image Diagnosis | Authors TBD | MLLM, medical, VLM, diagnosis | cs.CV | 2025.05 | |
| arXiv(v1) 2025 | VideoLLM Benchmarks and Evaluation: A Survey | Authors TBD | MLLM, video, benchmark, evaluation, survey | cs.CV, cs.CL | 2025.05 | |
| arXiv(v2) 2025 | Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark | Hanlei Zhang, et al. | multimodal language analysis, MLLM, semantics | cs.CL, cs.AI, cs.MM | 23 pages, 5 figures | 2025.04 |
| arXiv(v2) 2025 | Dual Diffusion for Unified Image Generation and Understanding | Zijie Li, et al. | dual diffusion, unified generation, understanding, MLLM | cs.CV, cs.AI, cs.LG | 2025.04 | |
| arXiv(v1) 2025 | Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents | Gavin Greif, et al. | OCR, historical documents, named entity recognition, MLLM | cs.CL, cs.AI, cs.DL | 2025.04 | |
| arXiv(v1) 2025 | Socratic Chart: Cooperating Multiple Agents for Robust SVG Chart Understanding | Yuyang Ji, et al. | chart understanding, SVG, multi-agent, MLLM | cs.CV | 2025.04 | |
| arXiv(v1) 2025 | VLM-R1: A Stable and Generalizable R1-Style Large VLM | Authors TBD | MLLM, VLM, reasoning, R1-style | cs.CV, cs.CL | 2025.04 | |
| arXiv(v1) 2025 | MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation | Authors TBD | MLLM, 3D, segmentation, reasoning | cs.CV | 2025.03 | |
| arXiv(v1) 2025 | MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs | Erik Daxberger, et al. | MLLM, 3D, spatial, understanding, benchmark | cs.CV, cs.CL | ICCV 2025 | 2025.03 |
| arXiv(v1) 2025 | Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis | Authors TBD | MLLM, medical, 3D, VLM, CT | cs.CV | 2025.03 | |
| arXiv(v1) 2025 | Mobile-VideoGPT: Fast and Accurate Video Understanding Language Model | Authors TBD | MLLM, video, mobile, efficient | cs.CV, cs.CL | 2025.03 | |
| arXiv(v1) 2025 | R1-Zero's Aha Moment in Visual Reasoning on a 2B Non-SFT Model | Authors TBD | MLLM, visual reasoning, R1-Zero, emergent | cs.CV, cs.CL | 2025.03 | |
| arXiv(v1) 2025 | SpaceVLLM: Endowing MLLM with Spatio-Temporal Video Grounding | Authors TBD | MLLM, video, spatio-temporal, grounding | cs.CV, cs.CL | 2025.03 | |
| arXiv(v1) 2025 | Vision-R1: Incentivizing Reasoning Capability in MLLMs | Authors TBD | MLLM, reasoning, visual reasoning | cs.CV, cs.CL | 2025.03 | |
| arXiv(v2) 2025 | Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning | Bohao Yang, et al. | table understanding, scientific data, MLLM | cs.CL | 2025.02 | |
| arXiv(v4) 2025 | LMFusion: Adapting Pretrained Language Models for Multimodal Generation | Weijia Shi, et al. | multimodal generation, LLM adaptation, MLLM | cs.CL, cs.AI, cs.CV, cs.LG | Name change: LlamaFusion to LMFusion | 2025.02 |
| arXiv(v1) 2025 | Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review | Pei Fu, et al. | text-rich image understanding, MLLM, vision-language | cs.CV | 2025.02 | |
| arXiv(v1) 2025 | Visual Perception Token for Multimodal Large Language Models | Authors TBD | MLLM, visual perception, token, autonomous control | cs.CV, cs.CL | 829k training samples | 2025.02 |
| arXiv(v1) 2025 | Weak Supervision Dynamic KL-Weighted Diffusion Models Guided by Large Language Models | Julian Perry, et al. | diffusion model, LLM guidance, weak supervision, image generation | cs.CL | 2025.02 | |
| arXiv(v1) 2025 | Bridging Visualization and Optimization: Multimodal Large Language Models on Graph-Structured Combinatorial Optimization | Jie Zhao, et al. | graph-structured optimization, combinatorial, MLLM | cs.AI, cs.LG | 2025.01 | |
| arXiv(v1) 2025 | Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding | Yun Li, et al. | temporal modeling, video understanding, MLLM | cs.CV, cs.CL | 2025.01 | |
| arXiv(v5) 2025 | Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation | Yuyang Ye, et al. | sequential recommendation, MLLM, multimodal | cs.IR, cs.AI | 2025.01 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | Contextual Memory Virtualisation: DAG-Based State Management and Structurally Lossless Trimming for LLM Agents | Cosmo Santoni | LLM, memory | cs.SE, cs.AI, cs.HC, cs.OS | 11 pages. 6 figures. Introduces a DAG-based state management system for LLM agents. Evaluation on 76 coding sessions shows up to 86% token reduction (mean 20%) while remaining economically viable under prompt caching. Includes reference implementation for Claude Code | 2026.02 |
| arXiv(v1) 2026 | A Dynamic Retrieval-Augmented Generation System with Selective Memory and Remembrance | Okan Bursa | dynamic RAG, selective memory, retrieval, LLM | cs.IR, cs.AI | 6 Pages, 2 figures | 2026.01 |
| arXiv(v1) 2026 | Active Context Compression: Autonomous Memory Management in LLM Agents | Authors TBD | memory, compression, context, autonomous, Focus | cs.CL, cs.AI | 22.7% token savings | 2026.01 |
| arXiv(v1) 2025 | Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management | Authors TBD | memory, unified, long-term, short-term, management | cs.AI, cs.CL | 2026.01 | |
| arXiv(v1) 2026 | Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents | Authors TBD | memory, temporal, semantic, personalized | cs.CL, cs.AI | Durative memory, Zep architecture | 2026.01 |
| arXiv(v1) 2026 | Beyond Static Summarization: Proactive Memory Extraction for LLM Agents | Chengyuan Yang, et al. | memory, extraction, summarization, agent | cs.CL, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Continuum Memory Architectures for Long-Horizon LLM Agents | Authors TBD | memory, continuum, long-horizon, consolidation | cs.AI, cs.CL | Episodic-to-semantic conversion | 2026.01 |
| arXiv(v2) 2026 | Cost and accuracy of long-term memory in Distributed Multi-Agent Systems based on Large Language Models | Benedict Wolff, et al. | graph memory, distributed multi-agent, LLM | cs.IR | 23 pages, 4 figures, 7 tables | 2026.01 |
| arXiv(v1) 2026 | Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration | Sen Wang, et al. | memory, benchmark, multimodal, RL, MLLM | cs.AI, cs.CV | Our dataset and code will be released at our \href{https://wangsen99.github.io/papers/lmee/}{website} | 2026.01 |
| arXiv(v1) 2026 | Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory Management | Weitao Ma, et al. | memory, management, feedback, alignment, agent | cs.CL | 18 pages, 5 figures | 2026.01 |
| arXiv(v3) 2026 | HaluMem: Evaluating Hallucinations in Memory Systems of Agents | Ding Chen, et al. | memory, hallucination, evaluation, agent | cs.CL | 2026.01 | |
| arXiv(v1) 2026 | HiMeS: Hippocampus-inspired Memory System for Personalized AI Assistants | Hailong Li, et al. | memory, personalized assistant, hippocampus, long-term | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | HiMem: Hierarchical Long-Term Memory for LLM Long-Horizon Agents | Ningning Zhang, et al. | memory, long-term, hierarchical, agent, LLM | cs.AI | 2026.01 | |
| arXiv(v2) 2026 | Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory | Sizhe Yuen, et al. | multi-agent, structured memory, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents | Davide Baldelli, et al. | working memory, evaluation, agent, LLM | cs.CL | 2026.01 | |
| arXiv(v2) 2026 | LiCoMemory: Lightweight and Cognitive Agentic Memory for Efficient Long-Term Reasoning | Zhengjun Huang, et al. | cognitive memory, long-term reasoning, LLM, agent | cs.IR | 2026.01 | |
| arXiv(v1) 2026 | MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents | Dongming Jiang, et al. | multi-graph, agentic memory, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents | Yuanchen Bei, et al. | memory, multimodal, conversational, MLLM benchmark | cs.CL, cs.AI | 34 pages, 18 figures | 2026.01 |
| arXiv(v1) 2026 | MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards | Authors TBD | memory, long-term, construction, dense reward | cs.CL, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction | Haonan Bian, et al. | memory, benchmark, interaction, agent | cs.CL, cs.AI | 2026.01 | |
| arXiv(v3) 2026 | SimpleMem: Efficient Lifelong Memory for LLM Agents | Jiaqi Liu, et al. | lifelong memory, efficient, LLM, agent | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | SwiftMem: Fast Agentic Memory via Query-aware Indexing | Anxin Tian, et al. | memory, retrieval, indexing, agent, LLM | cs.CL, cs.AI | 2026.01 | |
| arXiv(v3) 2026 | TeleMem: Building Long-Term and Multimodal Memory for Agentic AI | Chunliang Chen, et al. | memory, multimodal, long-term, agent, LLM | cs.CL, cs.AI, cs.CV | 2026.01 | |
| arXiv(v1) 2025 | Tool-Memory Conflicts in Tool-Augmented LLMs | Authors TBD | memory, tool use, conflict, LLM | cs.CL, cs.AI | 2026.01 | |
| arXiv(v1) (NeurIPS25) | A-MEM: Agentic Memory for LLM Agents (NeurIPS) | Authors TBD | memory, agentic, Zettelkasten, NeurIPS | cs.AI, cs.CL | NeurIPS 2025 publication | 2025.12 |
| arXiv(v1) 2025 | AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents | Jiafeng Liang, et al. | memory, cognitive, survey, agent | cs.CL, cs.AI, cs.CV | 57 pages, 5 figures | 2025.12 |
| arXiv(v1) 2025 | Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Rewards, Experience Synthesis, and Continual Memory | Ken Huang, et al. | memory, continual, self-improvement, agent | cs.CR, cs.AI | 11 pages, 4 figures. Includes a complete runnable reference implementation and audit logging framework | 2025.12 |
| arXiv(v1) 2025 | Beyond Heuristics: A Decision-Theoretic Framework for Agent Memory Management | Changzhi Sun, et al. | memory, agent, decision-theoretic, management | cs.CL | 2025.12 | |
| arXiv(v1) 2025 | Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs | Ngoc Bui, et al. | memory, KV cache, retention, long-context | cs.LG, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | Context as a Tool: Context Management for Long-Horizon SWE-Agents | Shukai Liu, et al. | memory, context management, long-horizon, SWE-agent | cs.CL | 2025.12 | |
| arXiv(v2) 2025 | Evaluating Long-Term Memory for Long-Context Question Answering | Alessandra Terranova, et al. | memory, long-context, evaluation, QA | cs.CL | Accepted as a poster at Metacognition in Generative AI EurIPS workshop | 2025.12 |
| arXiv(v1) 2025 | Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects | Authors TBD | memory, hindsight, reflection, retention | cs.AI, cs.CL | Retain-Recall-Reflect framework | 2025.12 |
| arXiv(v1) 2025 | Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement | Saman Forouzandeh, et al. | memory, procedural, hierarchical, agent | cs.LG, cs.AI | Accepted at The 25th International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS 2026). 21 pages including references, with 7 figures and 8 tables. Code is publicly available at the authors GitHub repository: https://github.com/S-Forouzandeh/MACLA-LLM-Agents-AAMAS-Conference | 2025.12 |
| arXiv(v2) 2025 | MMAG: Mixed Memory-Augmented Generation for Large Language Models Applications | Stefano Zeppieri | memory, memory-augmented generation, RAG, LLM | cs.CL, cs.IR | 2025.12 | |
| arXiv(v1) 2025 | MemEvolve: Meta-Evolution of Agent Memory Systems | Guibin Zhang, et al. | memory, evolution, meta-learning, agent | cs.CL, cs.MA | 2025.12 | |
| arXiv(v1) 2025 | MemR3: Memory Retrieval via Reflective Reasoning for LLM Agents | Authors TBD | memory, retrieval, reflection, reasoning | cs.AI, cs.CL | Router + evidence-gap tracker | 2025.12 |
| arXiv(v2) 2025 | Memento 2: Learning by Stateful Reflective Memory | Jun Wang | memory, agent, reflection, stateful | cs.AI, cs.CV, cs.LG | 35 pages, four figures | 2025.12 |
| arXiv(v1) 2025 | Memory in the Age of AI Agents | Yuyang Hu, et al. | memory, survey, LLM, agent | cs.CL, cs.AI | Comprehensive survey on agent memory | 2025.12 |
| arXiv(v1) 2025 | MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval | Saksham Sahai Srivastava, et al. | memory, security, agent, attack | cs.CR, cs.AI, cs.LG | 14 pages, 1 figure, includes appendix | 2025.12 |
| arXiv(v3) 2025 | O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents | Piaohong Wang, et al. | memory, long-horizon, self-evolving, agent | cs.CL | 2025.12 | |
| arXiv(v1) 2025 | R-Debater: Retrieval-Augmented Debate Generation through Argumentative Memory | Maoyuan Li, et al. | memory, retrieval, debate, RAG | cs.CL, cs.AI | Accepteed by AAMAS 2026 full paper | 2025.12 |
| arXiv(v2) 2025 | Significant Other AI: Identity, Memory, and Emotional Regulation as Long-Term Relational Intelligence | Sung Park | memory, identity, relational, long-term | cs.HC, cs.AI | 2025.12 | |
| arXiv(v1) 2025 | Adaptive Focus Memory for Language Models | Authors TBD | memory, adaptive, focus, compression, AFM | cs.CL, cs.AI | 2/3 token reduction | 2025.11 |
| arXiv(v1) 2025 | BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing | Authors TBD | memory, selective, budget, efficient | cs.CL, cs.AI | Learned gating + BM25 | 2025.11 |
| arXiv(v1) 2025 | CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing | Guihang Hong, et al. | RAG, edge computing, optimization, LLM | cs.DC | Accepted by RTSS 2025 (Real-Time Systems Symposium, 2025) | 2025.11 |
| arXiv(v1) 2025 | EMem: Event-Centric Memory for Long-Term Conversational Agents | Authors TBD | memory, event-centric, conversation, neo-Davidsonian | cs.CL, cs.AI | Event-like propositions | 2025.11 |
| arXiv(v1) 2025 | Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory | Tianxin Wei, et al. | memory, benchmark, test-time learning, agent | cs.CL, cs.AI | 2025.11 | |
| arXiv(v1) 2025 | G-KV: Decoding-Time KV Cache Eviction with Global Attention | Mengqi Liao, et al. | memory, KV cache, eviction, efficiency | cs.CL, cs.AI | 2025.11 | |
| arXiv(v1) 2025 | GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory | Jeong Hun Yeo, et al. | memory, episodic, video, MLLM | cs.CV, cs.AI | 2025.11 | |
| arXiv(v1) 2025 | Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks | Yicong Zheng, et al. | memory, compression, long-context, search | cs.CL, cs.AI, cs.LG | 2025.11 | |
| arXiv(v1) 2025 | KVzip: Memory Compression for LLM Chatbots via KV Cache Optimization | Authors TBD | memory, compression, KV cache, chatbot | cs.CL, cs.AI | 3-4x compression, 170K tokens | 2025.11 |
| MobiCom25 | Poster: MemAura: Persistent Personalized Context Memory for LLM Services in Smart Environments | Siyuan Liu, et al. | LLM, memory, personalization, smart environment, context | 2025.11 | ||
| arXiv(v1) 2025 | Trainable Graph Memory for LLM Agents: From Experience to Strategy | Authors TBD | memory, graph, trainable, strategy | cs.AI, cs.CL | Utility assessment mechanism | 2025.11 |
| arXiv(v1) 2025 | WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance | Genglin Liu, et al. | memory, web agent, cross-session, self-evolving | cs.AI, cs.CL | 18 pages; work in progress | 2025.11 |
| arXiv(v1) 2025 | A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents | Zhipeng Liao, et al. | memory-efficient, RAG, wearable, medical, LLM, agent | cs.AR | Accepted by BioCAS2025 | 2025.10 |
| arXiv(v1) 2025 | Acon: Optimizing Context Compression for Long-horizon LLM Agents | Authors TBD | memory, compression, context, optimization | cs.CL, cs.AI | 26-54% memory reduction | 2025.10 |
| arXiv(v1) 2025 | Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs | Mohammad Tavakoli, et al. | memory, long-context, benchmark, long-term | cs.CL, cs.AI, cs.IR | 2025.10 | |
| arXiv(v1) 2025 | CAM: Contextual Augmentation Memory for LLM Agents | Authors TBD | memory, contextual, augmentation, agent | cs.AI, cs.CL | 2025.10 | |
| arXiv(v1) 2025 | Dynamic Affective Memory Management for Personalized LLM Agents | Junfeng Lu, et al. | memory, affective, personalization, agent | cs.CL | 12 pasges, 8 figures | 2025.10 |
| arXiv(v1) 2025 | Enabling Personalized Long-term Interactions in LLM-based Agents through Persistent Memory | Authors TBD | memory, personalized, long-term, user profile | cs.CL, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | LightMem: Lightweight Memory for Efficient LLM Agents | Authors TBD | memory, lightweight, efficient, agent | cs.AI, cs.CL | 2025.10 | |
| arXiv(v1) 2025 | MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments | Darshan Deshpande, et al. | memory, benchmark, state tracking, agent | cs.AI, cs.CL | Accepted to NeurIPS 2025 SEA Workshop | 2025.10 |
| arXiv(v1) 2025 | Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy Games | Runnan Qi, et al. | memory, prompting, state machine, agent | cs.AI | 10 pages, 4 figures, 1 table, 1 algorithm. Submitted to conference | 2025.10 |
| arXiv(v1) 2025 | Pre-Storage Reasoning for Episodic Memory in LLM Agents | Authors TBD | memory, episodic, pre-storage, reasoning | cs.AI, cs.CL | 2025.10 | |
| arXiv(v2) 2025 | Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions | Yuanzhe Hu, et al. | memory evaluation, multi-turn, LLM, agent | cs.CL, cs.AI | Y. Hu and Y. Wang contribute equally | 2025.09 |
| arXiv(v1) 2025 | HopRAG: Multi-Hop Reasoning with Graph Memory for LLM Agents | Authors TBD | memory, multi-hop, graph, reasoning, RAG | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Mem-α: Learning Memory Construction via Reinforcement Learning | Yu Wang, et al. | memory, RL, construction, agent | cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Mem-α: Memory with Adaptive Forgetting for LLM Agents | Authors TBD | memory, forgetting, adaptive, agent | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Memory in LLM-based Multi-agent Systems: Mechanisms, Challenges, and Collective | Authors TBD | memory, multi-agent, collective, survey | cs.AI, cs.MA | 2025.09 | |
| arXiv(v1) 2025 | Multiple Memory Systems for Enhancing Long-term Memory of LLM Agents | Authors TBD | memory, multiple systems, long-term, enhancement | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | Nemori: Neural Memory Organization for LLM Agents | Authors TBD | memory, neural, organization, agent | cs.AI, cs.CL | 2025.09 | |
| arXiv(v1) 2025 | SGMem: Sentence Graph Memory for Long-Term Conversational Agents | Yaxiong Wu, et al. | memory, graph, conversation, retrieval | cs.CL, cs.IR | 19 pages, 6 figures, 1 table | 2025.09 |
| arXiv(v1) 2025 | SGMem: Structured Graph Memory for LLM Agents | Authors TBD | memory, graph, structured, agent | cs.AI, cs.CL | 2025.09 | |
| IJCAI25 (IJCAI 2025) | AriGraph: Learning Knowledge Graph World Models with Episodic Memory | Authors TBD | memory, knowledge graph, episodic, world model | cs.AI | IJCAI 2025 | 2025.08 |
| arXiv(v1) 2025 | Cognitive Workspace: Active Memory Management for LLMs - Functional Infinite Context | Authors TBD | memory, cognitive, workspace, active management | cs.CL, cs.AI | Metacognitive control | 2025.08 |
| arXiv(v1) 2025 | Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework | Zeyu Zhang, et al. | adaptive memory, optimization, LLM, agent | cs.LG, cs.AI, cs.CL, cs.IR | 17 pages, 4 figures, 5 tables | 2025.08 |
| arXiv(v1) 2025 | Memory-Augmented Transformers: A Systematic Review | Authors TBD | memory, transformer, survey, augmented | cs.CL, cs.AI | Systematic review | 2025.08 |
| arXiv(v1) 2025 | Memory-R1: Enhancing LLM Agents to Manage and Utilize Memories via RL | Authors TBD | memory, RL, memory manager, ADD/UPDATE/DELETE | cs.AI, cs.LG | Memory Manager + Answer Agent | 2025.08 |
| arXiv(v1) 2025 | Recursive Summarization for Long-Term Dialogue Memory in LLMs | Authors TBD | memory, summarization, dialogue, recursive | cs.CL, cs.AI | Updated 2025 | 2025.08 |
| arXiv(v2) 2025 | In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents | Zhen Tan, et al. | reflective memory, dialogue agent, personalization, LLM | cs.CL, cs.AI | Accepted to ACL 2025 | 2025.07 |
| ACL25 (ACL 2025) | Pretraining Context Compressor for LLMs with Embedding-Based Memory | Authors TBD | memory, compression, embedding, context | cs.CL | PCC framework | 2025.07 |
| arXiv(v1) 2025 | Cross-Attention Networks for Memory Retrieval in Generative Agents | Authors TBD | memory, retrieval, cross-attention, generative | cs.AI, cs.CL | Frontiers in Psychology | 2025.04 |
| arXiv(v1) 2025 | From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs | Authors TBD | memory, survey, episodic, semantic, working memory | cs.CL, cs.AI | Personal/system, parametric/non-parametric | 2025.04 |
| arXiv(v3) 2025 | LIFT: Improving Long Context Understanding of Large Language Models through Long Input Fine-Tuning | Yansheng Mao, et al. | long context, fine-tuning, LLM | cs.CL | 2025.04 | |
| arXiv(v1) 2025 | Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory | Authors TBD | memory, long-term, scalable, production | cs.AI, cs.CL | 26% improvement over OpenAI | 2025.04 |
| arXiv(v1) 2025 | In Prospect and Retrospect: Reflective Memory Management for Long-term Dialogue Agents | Authors TBD | memory, reflective, dialogue, long-term | cs.CL, cs.AI | 2025.03 | |
| arXiv(v1) 2025 | Tuning LLMs by RAG Principles: Towards LLM-native Memory | Jiale Wei, et al. | RAG, fine-tuning, optimization, LLM | cs.CL, cs.AI, cs.IR | 2025.03 | |
| arXiv(v1) 2025 | A-MEM: Agentic Memory for LLM Agents | Authors TBD | memory, agentic, self-organizing, Zettelkasten | cs.AI, cs.CL | Dynamic memory organization | 2025.02 |
| arXiv(v1) 2025 | Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents | Authors TBD | memory, episodic, long-term, position paper | cs.AI, cs.CL | Encoding and retrieval | 2025.02 |
| arXiv(v1) 2025 | Zep: Temporal Knowledge Graph Architecture for Agent Memory | Authors TBD | memory, temporal, knowledge graph, agent | cs.AI, cs.CL | Episodic + semantic + community | 2025.02 |
| arXiv(v1) 2024 | Memory-Augmented Agent Training for Business Document Understanding | Jiale Liu, et al. | memory, agent, training, document | cs.CL, cs.AI | 11 pages, 8 figures | 2024.12 |
| arXiv(v1) 2024 | On the Structural Memory of LLM Agents | Ruihong Zeng, et al. | memory, agent, analysis, structural | cs.CL, cs.AI | 2024.12 | |
| arXiv(v1) 2024 | XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference | Weizhuo Li, et al. | memory, KV cache, long-context, personalization | cs.LG, cs.CL | 2024.12 | |
| arXiv(v1) 2024 | MELODI: Exploring Memory Compression for Long Contexts | Yinpeng Chen, et al. | memory compression, long context, LLM | cs.LG, cs.AI | 2024.10 | |
| arXiv(v1) 2024 | A Survey on the Memory Mechanism of Large Language Model based Agents | Zeyu Zhang, et al. | memory, survey, LLM, agent | cs.AI | ACM TOIS, 39 pages | 2024.04 |
| Source | Title (Link) | Authors | Tag | Subjects | Additional info | Date |
|---|---|---|---|---|---|---|
| arXiv(v1) 2026 | DeepInterestGR: Mining Deep Multi-Interest Using Multi-Modal LLMs for Generative Recommendation | Yangchen Zeng | LLM, personalization | cs.LG, cs.CV, cs.CY | 2026.02 | |
| arXiv(v1) 2026 | Dynamic Personality Adaptation in Large Language Models via State Machines | Leon Pielage, et al. | LLM, personalization | cs.CL, cs.HC, cs.LG | 22 pages, 5 figures, submitted to ICPR 2026 | 2026.02 |
| arXiv(v1) 2026 | Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs | Wenqiu Tang, et al. | LLM, personalization | cs.CL | Accepted in PAKDD 2026 special session on Data Science :Foundation and Applications | 2026.02 |
| arXiv(v1) 2026 | InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation | Yu Li, et al. | LLM, personalization | cs.CL, cs.AI, cs.CY | 2026.02 | |
| arXiv(v1) 2026 | Learning to Reason for Multi-Step Retrieval of Personal Context in Personalized Question Answering | Maryam Amirizaniani, et al. | LLM, personalization | cs.CL, cs.AI, cs.IR | 2026.02 | |
| arXiv(v1) 2026 | Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization | Shangding Gu | LLM, personalization | cs.LG, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Multi-Agent Large Language Model Based Emotional Detoxification Through Personalized Intensity Control for Consumer Protection | Keito Inoshita | LLM, personalization | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Offline Reasoning for Efficient Recommendation: LLM-Empowered Persona-Profiled Item Indexing | Deogyong Kim, et al. | LLM, personalization | cs.IR, cs.LG | Under review | 2026.02 |
| arXiv(v1) 2026 | PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra | Xiachong Feng, et al. | LLM, personalization | cs.AI | ICLR 2026 | 2026.02 |
| arXiv(v1) 2026 | PRECTR-V2:Unified Relevance-CTR Framework with Cross-User Preference Mining, Exposure Bias Correction, and LLM-Distilled Encoder Optimization | Shuzhi Cao, et al. | LLM, personalization | cs.IR, cs.AI | arXiv admin note: text overlap with arXiv:2503.18395 | 2026.02 |
| arXiv(v1) 2026 | Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History | Serin Kim, et al. | LLM, personalization | cs.CL, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Personalized Graph-Empowered Large Language Model for Proactive Information Access | Chia Cheng Chang, et al. | LLM, personalization | cs.CL | 2026.02 | |
| arXiv(v1) 2026 | Personalized Prediction of Perceived Message Effectiveness Using Large Language Model Based Digital Twins | Jasmin Han, et al. | LLM, personalization | cs.CL, stat.AP | 31 pages, 5 figures, submitted to Journal of the American Medical Informatics Association (JAMIA). Drs. Chen and Thrul share last authorship | 2026.02 |
| arXiv(v1) 2026 | Sydney Telling Fables on AI and Humans: A Corpus Tracing Memetic Transfer of Persona between LLMs | Jiří Milička, et al. | LLM, personalization | cs.CL, cs.AI | 2026.02 | |
| arXiv(v1) 2026 | Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions | Yue Xu, et al. | LLM, personalization | cs.AI | 2026.02 | |
| arXiv(v1) 2026 | CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs | Liang Wang, et al. | arxiv | cs.CL | 2026.01 | |
| arXiv(v1) 2026 | Enriching Semantic Profiles into Knowledge Graph for Recommender Systems Using Large Language Models | Seokho Ahn, et al. | knowledge graph, semantic profiles, recommendation, LLM | cs.IR, cs.AI, cs.LG | Accepted at KDD 2026 | 2026.01 |
| ICLR 2026 | FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents | ICLR,OpenReview | Mobile Agent, LLM Agent, GUI, Proactive Agent, Personalization | OpenReview ID: n3iFV0gLMc | 2026.01 | |
| arXiv(v1) 2026 | HumanLLM: Towards Personalized Understanding and Simulation of Human Nature | Yuxuan Lei, et al. | arxiv | cs.CL | 12 pages, 5 figures, 7 tables, to be published in KDD 2026 | 2026.01 |
| arXiv(v1) 2026 | Improving User Privacy in Personalized Generation: Client-Side Retrieval-Augmented Modification of Server-Side Generated Speculations | Alireza Salemi, et al. | arxiv | cs.CL, cs.AI, cs.CR, cs.IR | 2026.01 | |
| arXiv(v2) 2026 | Linear Personality Probing and Steering in LLMs: A Big Five Study | Michel Frising, et al. | personalization, personality, probing, steering | cs.CL | 29 pages, 6 figures | 2026.01 |
| arXiv(v1) 2026 | Me-Agent: A Personalized Mobile Agent with Two-Level User Habit Learning for Enhanced Interaction | Shuoxin Wang, et al. | arxiv | cs.CL | 2026.01 | |
| ICLR 2026 | Meta-Router: Bridging Gold-standard and Preference-based Evaluations in LLM Routing | ICLR,OpenReview | Causal learning, Meta-learner, Large Language Model, query routing | OpenReview ID: r0BFucF2dH | 2026.01 | |
| ICLR 2026 | NextQuill: Causal Preference Modeling for Enhancing LLM Personalization | ICLR,OpenReview | Personalized text generation, Large Language Models, LLM Personalization | OpenReview ID: xYpVlKMFqv | 2026.01 | |
| arXiv(v1) 2026 | One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment | Hongru Cai, et al. | meta reward modeling, alignment, personalization, LLM | cs.CL, cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization | Linfeng Du, et al. | arxiv | cs.CL, cs.IR | 2026.01 | |
| arXiv(v1) 2026 | PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation | Junho Park, et al. | few-shot, privacy-safe, lightweight adaptation, personalization, LLM | cs.CL, cs.AI, cs.LG | 16 pages, 9 figures | 2026.01 |
| arXiv(v1) 2026 | PersonaDual: Balancing Personalization and Objectivity via Adaptive Reasoning | Xiaoyou Liu, et al. | personalization, persona, reasoning, LLM | cs.AI | 2026.01 | |
| ICLR 2026 | Preference Leakage: A Contamination Problem in LLM-as-a-judge | ICLR,OpenReview | LLM-as-a-judge, Preference Leakage, Data Contamination | OpenReview ID: grIvSXVJ65 | 2026.01 | |
| ICLR 2026 | ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation | ICLR,OpenReview | Benchmark, Agent Simulation, Personalization, Proactivity | OpenReview ID: RV2aeCgxdB | 2026.01 | |
| arXiv(v1) 2026 | SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation | Seoyeon Kim, et al. | personalization, continual, retrieval, parametric adaptation, LLM | cs.AI, cs.CL | under review, 23 pages | 2026.01 |
| ICLR 2026 | Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents | ICLR,OpenReview | LLM-based Agents, Process Supervision, Curriculum Learning | OpenReview ID: s8usvGHYlk | 2026.01 | |
| arXiv(v1) 2026 | Structured Personality Control and Adaptation for LLM Agents | Jinpeng Wang, et al. | personalization, personality, control, agent, LLM | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | Styles + Persona-plug = Customized LLMs | Yutong Song, et al. | style customization, persona-plug, LLM | cs.AI | 2026.01 | |
| arXiv(v1) 2026 | The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models | Christina Lu, et al. | persona, default persona, alignment, LLM | cs.CL | 2026.01 | |
| arXiv(v2) 2026 | The Reward Model Selection Crisis in Personalized Alignment | Fady Rezk, et al. | personalization, alignment, reward model, RLHF | cs.AI, cs.LG | 2026.01 | |
| ICLR 2026 | Towards Understanding Valuable Preference Data for Large Language Model Alignment | ICLR,OpenReview | Large language model alignment, preference data, influence function | OpenReview ID: FUp0KeEEBs | 2026.01 | |
| ICLR 2026 | Verification and Co-Alignment via Heterogeneous Consistency for Preference-Aligned LLM Annotations | ICLR,OpenReview | Verification, Co-Alignment, Preference-Aligned LLM Annotations, Reference-Free Metric | OpenReview ID: jugY302BAh | 2026.01 | |
| arXiv(v1) 2026 | When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs | Zhongxiang Sun, et al. | personalization, hallucination, safety, evaluation | cs.CL, cs.AI | 20 pages, 15 figures | 2026.01 |
| arXiv(v1) 2025 | Agentic Multi-Persona Framework for Evidence-Aware Fake News Detection | Roopa Bukke, et al. | persona, multi-agent, fake news, LLM | cs.IR, cs.LG | 12 pages, 8 tables, 2 figures | 2025.12 |
| arXiv(v1) 2025 | Interpolative Decoding: Exploring the Spectrum of Personality Traits in LLMs | Eric Yeh, et al. | personalization, personality, decoding, control | cs.AI | 20 pages, 5 figures | 2025.12 |
| arXiv(v1) 2025 | LLM Personas as a Substitute for Field Experiments in Method Benchmarking | Enoch Hyunwook Kang | persona, evaluation, methodology, LLM | cs.AI, cs.LG, econ.EM | 2025.12 | |
| arXiv(v1) 2025 | Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI | Samarth Sarin, et al. | personalization, memory, conversational AI, agent | cs.AI, cs.CL | Paper accepted at 5th International Conference of AIML Systems 2025, Bangalore, India | 2025.12 |
| arXiv(v1) 2025 | PILAR: Personalizing Augmented Reality Interactions with LLM-based Human-Centric and Trustworthy Explanations for Daily Use Cases | Ripan Kumar Kundu, et al. | personalization, AR, explanations, LLM | cs.HC, cs.AI | Published in the 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct) | 2025.12 |
| arXiv(v1) 2025 | PRISM: A Personality-Driven Multi-Agent Framework for Social Media Simulation | Zhixiang Lu, et al. | personality, multi-agent, social simulation, LLM | cs.CL | 2025.12 | |
| arXiv(v1) 2025 | PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas | Authors TBD | personalization, persona, implicit, user modeling | cs.CL, cs.AI | 1000 personas, 20k preferences, 128k context | 2025.12 |
| arXiv(v1) 2025 | Personalized Multimodal Large Language Models: A Survey | Authors TBD | personalization, MLLM, survey, multimodal | cs.CV, cs.CL | Comprehensive MLLM personalization survey | 2025.12 |
| arXiv(v1) 2025 | PrefGen: Multimodal Preference Learning for Image Generation | Authors TBD | personalization, MLLM, image generation, preference | cs.CV, cs.CL | User-specific conditioning | 2025.12 |
| arXiv(v2) 2025 | ProEx: A Unified Framework Leveraging Large Language Model with Profile Extrapolation for Recommendation | Yi Zhang, et al. | profile extrapolation, recommendation, LLM | cs.IR | Accepted by KDD 2026 (First Cycle) | 2025.12 |
| arXiv(v1) 2025 | SPARK: Search Personalization via Agent-Driven Retrieval and Knowledge-sharing | Gaurab Chhetri, et al. | personalization, search, agent, retrieval | cs.AI | Accepted to WEB&GRAPH 2026 (WSDM 2026 workshop) | 2025.12 |
| arXiv(v1) 2025 | TAME: Long-Context MLLM Personalization with Double Memories | Authors TBD | personalization, MLLM, memory, training-free | cs.CV, cs.CL | RA2G recipe | 2025.12 |
| arXiv(v1) 2025 | The Mental World of Large Language Models in Recommendation: A Benchmark on Association, Personalization, and Knowledgeability | Guangneng Hu | personalization, recommendation, benchmark, LLM | cs.IR | 21 pages, 13 figures, 27 tables, submission to KDD 2025 | 2025.12 |
| arXiv(v1) 2025 | Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues | Xiaotian Zhang, et al. | proactive personalization, profile customization, dialogue, LLM | cs.CL | 2025.12 | |
| TiiS25 | User Perceptions of Personalized and Generic Explanations in LLM-Driven Recommender Systems | Ítallo De Sousa Silva, et al. | LLM, recommender system, personalization, explanations, user study | 2025.12 | ||
| arXiv(v1) 2025 | Fixed-Persona SLMs with Modular Memory: Scalable NPC Dialogue on Consumer Hardware | Martin Braas, et al. | personalization, persona, modular memory, NPC | cs.AI, cs.IR | 2025.11 | |
| arXiv(v2) 2025 | Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction | Zhaopei Huang, et al. | personalization, dialogue, memory, long-term | cs.CL | Accepted by AAAI 2026 (Oral) | 2025.11 |
| arXiv(v1) 2025 | PLUM: Learning to Remember User Conversations for Personalization | Authors TBD | personalization, memory, conversation, LoRA | cs.CL, cs.AI | Parameter-efficient | 2025.11 |
| arXiv(v1) 2025 | PersonaAgent with GraphRAG: Community-Aware KG for Personalized LLM | Authors TBD | personalization, GraphRAG, knowledge graph, agent | cs.CL, cs.AI | 11.1% F1 improvement on LaMP | 2025.11 |
| arXiv(v1) 2025 | PersonalizedRouter: Personalized LLM Routing via Graph-based User Preference Modeling | Authors TBD | personalization, routing, GNN, user preference | cs.CL, cs.AI | 2025.11 | |
| arXiv(v1) 2025 | Profile-LLM: Dynamic Profile Optimization for Realistic Personality Expression | Authors TBD | personalization, profile, personality, dynamic | cs.CL, cs.AI | Education, therapy, entertainment | 2025.11 |
| arXiv(v1) 2025 | LLMDiRec: LLM-Enhanced Intent Diffusion for Sequential Recommendation | Bo-Chian Chen, et al. | intent diffusion, sequential recommendation, LLM | cs.IR | Under review | 2025.10 |
| arXiv(v1) 2025 | MemWeaver: A Hierarchical Memory from Textual Interactive Behaviors for Personalized Generation | Shuo Yu, et al. | personalization, memory, user behavior, generation | cs.CL | 12 pages, 8 figures | 2025.10 |
| arXiv(v1) 2025 | P2P: Instant Personalized LLM Adaptation via Hypernetwork | Authors TBD | personalization, hypernetwork, instant adaptation | cs.CL, cs.AI | Single-pass generation | 2025.10 |
| arXiv(v1) 2025 | Preference-Aware Memory Update for Long-Term LLM Agents | Haoran Sun, et al. | personalization, preference, memory update, long-term | cs.CL, cs.AI | 2025.10 | |
| arXiv(v1) 2025 | RGMem: Renormalization Group-based Memory Evolution for Language Agent User Profile | Ao Tian, et al. | personalization, user profile, memory, agent | cs.AI | 11 pages,3 figures | 2025.10 |
| arXiv(v1) 2025 | Real-Time Personalization for LLM-based Recommendation with Customized ICL | Authors TBD | personalization, recommendation, ICL, real-time, online | cs.IR, cs.AI | No model update needed | 2025.10 |
| arXiv(v2) 2025 | CoPL: Collaborative Preference Learning for Personalizing LLMs | Youngbin Choi, et al. | collaborative preference learning, personalization, LLM | cs.LG, cs.AI, cs.IR | 19pages, 13 figures, 11 tables | 2025.09 |
| arXiv(v1) 2025 | DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models | Honghui Xu, et al. | federated learning, privacy-enhanced, on-device LLM, fine-tuning | cs.CR, cs.AI | 2025.09 | |
| arXiv(v1) 2025 | HumAIne-Chatbot: Real-Time Personalized Conversational AI via RL | Authors TBD | personalization, chatbot, RL, real-time, industry | cs.CL, cs.AI | Production deployment | 2025.09 |
| arXiv(v1) 2025 | MMPB: Multi-Modal Personalization Benchmark for VLMs | Authors TBD | personalization, MLLM, benchmark, VLM | cs.CV, cs.CL | First MLLM personalization benchmark | 2025.09 |
| arXiv(v1) 2025 | Personalized Reasoning: Just-In-Time Personalization and Why LLMs Fail At It | Shuyue Stella Li, et al. | just-in-time personalization, LLM, user preference | cs.CL, cs.AI | 57 pages, 6 figures | 2025.09 |
| RecSys25 | Revisiting Prompt Engineering: A Comprehensive Evaluation for LLM-based Personalized Recommendation | Genki Kusano, et al. | LLM, recommendation, personalization, prompt engineering | 2025.09 | ||
| arXiv(v1) 2025 | T-POP: Test-Time Personalization with Online Preference Feedback | Zikun Qu, et al. | test-time personalization, online preference, LLM | cs.LG, cs.AI | Preprint | 2025.09 |
| arXiv(v1) 2025 | DGDPO: Diagnostic-Guided Dynamic Profile Optimization for User Simulators | Authors TBD | personalization, user simulation, profile, dynamic | cs.IR, cs.AI | Bidirectional evolution | 2025.08 |
| arXiv(v1) 2025 | End-to-End Personalization: Unifying Recommender Systems with Large Language Models | Danial Ebrat, et al. | recommendation system, unifying, LLM | cs.IR, cs.LG | Second Workshop on Generative AI for Recommender Systems and Personalization at the ACM Conference on Knowledge Discovery and Data Mining (GenAIRecP@KDD 2025) | 2025.08 |
| arXiv(v1) 2025 | MLLMRec: MLLMs in Recommender Systems | Authors TBD | personalization, MLLM, recommendation, visual | cs.CV, cs.IR | Visual attribute extraction | 2025.08 |
| arXiv(v1) 2025 | MM-R1: Unified MLLMs for Personalized Image Generation | Authors TBD | personalization, MLLM, image generation, GRPO | cs.CV, cs.CL | X-CoT reasoning | 2025.08 |
| arXiv(v1) 2025 | MSPA: Multimodal Self-Corrective Preference Alignment for Recommendation | Authors TBD | personalization, MLLM, recommendation, self-corrective | cs.CV, cs.IR | 4D multimodal signals | 2025.08 |
| arXiv(v2) 2025 | Personalized LLM for Generating Customized Responses to the Same Query from Different Users | Hang Zeng, et al. | personalization, response generation, user-specific, LLM | cs.CL | Accepted by CIKM'25 | 2025.08 |
| arXiv(v1) 2025 | RLHF Fine-Tuning of LLMs for Alignment with Implicit User Feedback in Conversational Recommenders | Authors TBD | personalization, RLHF, implicit feedback, recommender | cs.IR, cs.AI | Dwell time, sentiment signals | 2025.08 |
| arXiv(v1) (SIGIR25) | CoT-Rec: Enhancing LLM-Based Recommendations Through Personalized Reasoning | Authors TBD | personalization, recommendation, CoT, reasoning | cs.IR, cs.AI | SIGIR 2025 | 2025.07 |
| arXiv(v1) 2025 | Comprehensive Review on LLMs for Recommender Systems | Authors TBD | personalization, recommendation, survey, LLM | cs.IR, cs.AI | Hybrid RAG approaches | 2025.07 |
| arXiv(v1) 2025 | DEP: Latent Inter-User Difference Modeling for LLM Personalization | Authors TBD | personalization, latent, embedding, difference-aware | cs.CL, cs.AI | Sparse autoencoder | 2025.07 |
| arXiv(v1) 2025 | PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training | Authors TBD | personalization, inference-time, alignment, preference | cs.CL, cs.AI | No reward model needed | 2025.07 |
| arXiv(v1) 2025 | PLUS: Learning to Summarize User Information for Personalized RLHF | Authors TBD | personalization, RLHF, user summary, preference | cs.LG, cs.AI | 11-77% reward model improvement | 2025.07 |
| arXiv(v1) 2025 | PRIME: LLM Personalization with Cognitive Memory and Thought Processes | Authors TBD | personalization, memory, episodic, semantic, cognitive | cs.CL, cs.AI | Dual-memory model | 2025.07 |
| arXiv(v1) 2025 | PURE: LLM-based User Profile Management for Recommender System | Authors TBD | personalization, user profile, recommendation, management | cs.IR, cs.AI | Profile extraction and updating | 2025.07 |
| arXiv(v1) 2025 | Personalization of Large Language Models: A Survey | Zhehao Zhang, et al. | personalization, survey, LLM, user profile | cs.CL, cs.AI | Comprehensive taxonomy | 2025.07 |
| arXiv(v2) 2025 | Comparison-based Active Preference Learning for Multi-dimensional Personalization | Minhyeon Oh, et al. | active preference learning, multi-dimensional, personalization, LLM | cs.LG | 2025.06 | |
| arXiv(v1) 2025 | PersonalAI: KG Storage and Retrieval for Personalized LLM Agents | Authors TBD | personalization, knowledge graph, memory, agent | cs.AI, cs.CL | Hybrid graph with hyperedges | 2025.06 |
| UMAP25 | Personalizing LLM Responses to Combat Political Misinformation | Adiba Proma, et al. | LLM, personalization, misinformation, user modeling | 2025.06 | ||
| arXiv(v1) 2025 | ProfiLLM: LLM-Based Framework for Implicit User Profiling | Authors TBD | personalization, profiling, implicit, chatbot | cs.CL, cs.AI | IT/cybersecurity domain | 2025.06 |
| arXiv(v1) 2025 | SEAL: Self-Adapting Language Models | Authors TBD | personalization, self-adaptation, online learning | cs.CL, cs.AI | Self-generated finetuning data | 2025.06 |
| arXiv(v3) 2025 | Drift: Decoding-time Personalized Alignments with Implicit User Preferences | Minbeom Kim, et al. | decoding-time alignment, implicit preferences, personalization, LLM | cs.CL | 19 pages, 6 figures | 2025.05 |
| arXiv(v2) 2025 | HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation | Cristina Garbacea, et al. | alignment, hypothesis generation, personalization, LLM | cs.CL | 2025.05 | |
| arXiv(v1) 2025 | MAP: Memory Assisted LLM for Personalized Recommendation System | Authors TBD | personalization, recommendation, memory, history | cs.IR, cs.AI | 2025.05 | |
| arXiv(v1) 2025 | PROSE: Aligning LLMs by Predicting Preferences from User Writing Samples | Authors TBD | personalization, preference prediction, writing samples | cs.CL, cs.AI | Iterative refinement | 2025.05 |
| arXiv(v1) 2025 | Privacy-preserving Prompt Personalization in Federated Learning for Multimodal Large Language Models | Sizai Hou, et al. | federated learning, privacy-preserving, prompt personalization, MLLM | cs.CR | Under Review | 2025.05 |
| arXiv(v1) 2025 | RLPA: Teaching LLMs to Evolve with Users via Dynamic Profile Modeling | Authors TBD | personalization, RLHF, dynamic profile, user evolution | cs.CL, cs.AI | Outperforms Claude-3.5, DeepSeek-V3 | 2025.05 |
| arXiv(v1) 2025 | Steerable Chatbots: Personalizing LLMs with Preference-Based Activation Steering | Authors TBD | personalization, chatbot, activation steering, inference | cs.CL, cs.AI | Training-free | 2025.05 |
| arXiv(v1) 2025 | Towards Explainable Temporal User Profiling with LLMs | Milad Sabouri, et al. | temporal user profiling, explainable, LLM | cs.IR, cs.AI | 2025.05 | |
| arXiv(v1) 2025 | Towards a unified user modeling language for engineering human centered AI systems | Aaron Conrardy, et al. | user modeling, LLM, personalization | cs.SE | Accepted at the Third Workshop on Engineering Interactive Systems Embedding AI Technologies (EISEAIT workshop at EICS 2025) | 2025.05 |
| arXiv(v1) 2025 | A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models | Authors TBD | personalization, preference alignment, survey, pluralistic | cs.CL, cs.AI | Training and inference-time methods | 2025.04 |
| arXiv(v2) 2025 | Differential Privacy Personalized Federated Learning Based on Dynamically Sparsified Client Updates | Chuanyin Wang, et al. | differential privacy, federated learning, personalization, LLM | cs.LG, cs.CR | 10 pages,2 figures | 2025.04 |
| arXiv(v1) 2025 | Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling | Authors TBD | personalization, benchmark, user profiling, dynamic | cs.CL, cs.AI | 2025.04 | |
| arXiv(v1) 2025 | LoRe: Personalizing LLMs via Low-Rank Reward Modeling | Avinandan Bose, et al. | low-rank reward modeling, personalization, LLM, RLHF | cs.LG, cs.AI, cs.CL | 2025.04 | |
| arXiv(v1) 2025 | PaRT: Enhancing Proactive Social Chatbots with Personalized Real-Time Retrieval | Authors TBD | personalization, chatbot, retrieval, real-time | cs.CL, cs.AI | 21.77% dialogue improvement, production deployed | 2025.04 |
| arXiv(v3) 2025 | Tuning-Free Personalized Alignment via Trial-Error-Explain In-Context Learning | Hyundong Cho, et al. | in-context learning, trial-error-explain, personalization, LLM | cs.CL, cs.AI | NAACL 2025 Findings | 2025.04 |
| arXiv(v1) 2025 | User Feedback Alignment for LLM-powered Exploration in Large-scale Recommendation | Authors TBD | personalization, recommendation, feedback, exploration | cs.IR, cs.AI | Click and dwell time signals | 2025.04 |
| arXiv(v1) 2025 | A Shared Low-Rank Adaptation Approach to Personalized RLHF | Renpu Liu, et al. | RLHF, low-rank adaptation, personalization, LLM | cs.LG, cs.AI | Published as a conference paper at AISTATS 2025 | 2025.03 |
| arXiv(v1) 2025 | Agentic Recommender Systems in the Era of Multimodal LLMs: Survey | Authors TBD | personalization, recommendation, agent, MLLM, survey | cs.IR, cs.AI | User agent simulation | 2025.03 |
| arXiv(v1) 2025 | BAHE: LLM-Enhanced CTR Prediction in Long Textual User Behaviors | Authors TBD | personalization, CTR, user behavior, industry | cs.IR, cs.AI | Deployed on 50M daily data | 2025.03 |
| arXiv(v1) 2025 | Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Online Shopping | Authors TBD | personalization, agent, user simulation, behavior | cs.AI, cs.HC | 31,865 shopping sessions | 2025.03 |
| arXiv(v1) 2025 | Language Model Personalization via Reward Factorization | Idan Shenfeld, et al. | reward factorization, RLHF, personalization, LLM | cs.LG | 2025.03 | |
| arXiv(v1) 2025 | Measuring What Makes You Unique: Difference-Aware User Modeling for LLM Personalization | Authors TBD | personalization, user modeling, difference-aware | cs.CL, cs.AI | 2025.03 | |
| arXiv(v7) 2025 | PAD: Personalized Alignment of LLMs at Decoding-Time | Ruizhe Chen, et al. | decoding-time alignment, personalization, LLM | cs.CL, cs.AI | ICLR 2025 | 2025.03 |
| arXiv(v1) 2025 | PersonaX: A Recommendation Agent Oriented User Modeling Framework | Authors TBD | personalization, recommendation, user modeling, agent | cs.IR, cs.AI | 3-11% improvement on AgentCF | 2025.03 |
| arXiv(v1) 2025 | A Survey of Personalized Large Language Models: Progress and Future Directions | Authors TBD | personalization, survey, LLM, prompting, finetuning | cs.CL, cs.AI | Input/model/objective level | 2025.02 |
| arXiv(v1) 2025 | FSPO: Few-Shot Preference Optimization of Synthetic Preference Data in LLMs Elicits Effective Personalization to Real Users | Anikait Singh, et al. | few-shot, preference optimization, personalization, LLM | cs.LG, cs.AI, cs.CL, cs.HC, stat.ML | Website: https://fewshot-preference-optimization.github.io/ | 2025.02 |
| arXiv(v1) 2025 | LoCoMo: Evaluating Very Long-Term Conversational Memory of LLM Agents | Authors TBD | personalization, memory, conversation, benchmark | cs.CL, cs.AI | 300 turns, 9K tokens, 35 sessions | 2025.02 |
| arXiv(v1) 2025 | PrefEval: Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following | Authors TBD | personalization, benchmark, preference, evaluation | cs.CL | 3000 preference-query pairs, 20 topics | 2025.02 |
| arXiv(v3) 2025 | Privacy-Preserving Personalized Federated Prompt Learning for Multimodal Large Language Models | Linh Tran, et al. | federated learning, privacy-preserving, personalized prompt, MLLM | cs.LG | 2025.02 | |
| arXiv(v1) 2025 | RLTHF: Targeted Human Feedback for LLM Alignment | Authors TBD | personalization, RLHF, targeted feedback, efficient | cs.CL, cs.AI | 6-7% human annotation effort | 2025.02 |
| arXiv(v3) 2025 | SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World | Jiaqi Zhang, et al. | personalization, agent, user modeling, embodied | cs.AI | 2025.02 | |
| arXiv(v1) 2025 | User Profile Construction and Updating with LLMs: Benchmark | Authors TBD | personalization, user profile, construction, updating | cs.CL, cs.AI | Static and dynamic profiling | 2025.02 |
| arXiv(v1) 2025 | When Personalization Meets Reality: Multi-Faceted Analysis of Personalized Preference Learning | Authors TBD | personalization, preference learning, fairness, evaluation | cs.CL, cs.AI | 2025.02 | |
| arXiv(v1) 2025 | Advancing Personalized Federated Learning: Integrative Approaches with AI for Enhanced Privacy and Customization | Kevin Cooper, et al. | federated learning, privacy, personalization, LLM | cs.LG, eess.SP | arXiv admin note: substantial text overlap with arXiv:2501.16758 | 2025.01 |
| arXiv(v2) 2025 | Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering | Rumi Allbert, et al. | personalization, personality, activation engineering, steering | cs.CL, cs.AI | 2025.01 | |
| arXiv(v1) 2025 | PerRecBench: Can LLMs Understand Preferences in Personalized Recommendation? | Authors TBD | personalization, benchmark, recommendation, preference | cs.IR, cs.CL | 19 LLMs evaluated | 2025.01 |
| arXiv(v2) 2025 | PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health | Huy Vu, et al. | personalization, personality, traits, mental health | cs.AI, cs.CL | 2025.01 | |
| arXiv(v1) 2024 | AI PERSONA: Towards Life-long Personalization of LLMs | Tiannan Wang, et al. | personalization, persona, lifelong, LLM | cs.CL, cs.AI | Work in progress | 2024.12 |
| arXiv(v1) 2024 | Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations | Sayantan Pal, et al. | personalization, persona, personality modeling, long-term | cs.CL, cs.AI | Accepted in COLING 2025 | 2024.12 |
| arXiv(v1) 2024 | Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits | Bohan Li, et al. | personalization, personality, dataset, MBTI | cs.CL, cs.CY | Accepted by COLING 2025. 28 papges, 20 figures, 10 tables | 2024.12 |
| arXiv(v1) 2024 | Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment | Jianfei Zhang, et al. | personalization, preference alignment, efficiency, LLM | cs.CL, cs.AI | Coling 2025 | 2024.12 |
| arXiv(v2) 2024 | Molar: Multimodal LLMs with Collaborative Filtering Alignment for Enhanced Sequential Recommendation | Yucong Luo, et al. | collaborative filtering, sequential recommendation, MLLM | cs.IR, cs.AI | 2024.12 | |
| arXiv(v1) 2024 | Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization | Guanghan Li, et al. | recommendation system, alignment, behavioral semantic, LLM | cs.IR, cs.AI, cs.CL | 7 pages, 3 figures, AAAI 2025 | 2024.12 |
| arXiv(v1) 2024 | ULMRec: User-centric Large Language Model for Sequential Recommendation | Minglai Shao, et al. | personalization, recommendation, sequential, LLM | cs.IR | 2024.12 | |
| arXiv(v1) 2024 | Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning | Sriyash Poddar, et al. | RLHF, variational preference learning, personalization, LLM | cs.LG, cs.AI, cs.CL, cs.RO | weirdlabuw.github.io/vpl | 2024.08 |
| arXiv(v2) 2024 | RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation | Chanwoo Park, et al. | RLHF, heterogeneous feedback, preference aggregation, personalization, LLM | cs.AI, cs.LG | Added experiments | 2024.05 |
Python
100.0%