Awesome Generative AI in Search, Recommendation, Personalization
Generative AI and LLMs for Search, Recommender, Personalization Engines
The goal of this repository is to survey and review generative AI and LLM-based methods for building large-scale search and recommender engines.
see also LLM Evaluation methods repository
Table of Content
π LLM, Search & Recommender Engines

π Foundations
- Search Surveys β broad reviews of AI-driven search
- Recommender Engine Surveys β reviews of LLM-powered recommendation methods
- Conferences, Workshops β SIGIR, RecSys, KDD, WWW and other venues, top conferences for search and recommendation
- Tutorials β Tutorials
- Software, Libraries, Frameworks β open-source deep research and AI search/recommender tools
- Blog Posts, Whitepapers β industry write-ups from Netflix, Pinterest, Anthropic and other industrial blogs
π€ Agentic & Conversational Search
π― Core Capabilities
π§© RAG & Retrieval
π Ranking & Embeddings
π§ Response Generation & Deep Research
π¬ Recommender Systems
β
Evaluation
π’ Verticals
Search Surveys
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications, Oct 2025, arxiv
- A Survey on AI Search with Large Language Models, July 2025, preprints, not peer reviewed
- A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval, Mar 2025, arxiv
- A Survey on Knowledge-Oriented Retrieval-Augmented Generation, Mar 2025, arxiv
- From Matching to Generation: A Survey on Generative Information Retrieval, Feb 2025, Journal Version, ACM Transaction on Information Systems, Feb 2025
- Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions, Jan 2025, IEEE
- Improving Recommendation Systems & Search in the Age of LLMs by Eugene Neyan, Mar 2025, blog post
- A Survey of Model Architectures in Information Retrieval, Jan 2025, arxiv
- A Survey of Conversational Search, Oct 2024, arxiv
- Large language models for generative information extraction: a survey, 2024, Front Comp Sci
- From Matching to Generation: A Survey on Generative Information Retrieval, Apr 2024, arxiv
- Dense Text Retrieval Based on Pretrained Language Models: A Survey, Feb 2024, ACM
- Retrieval-Augmented Generation for Large Language Models: A Survey, 2023, simg
- Large Language Models for Information Retrieval: A Survey, Aug 2023, arxiv
Recommender Engine Surveys
- A comprehensive review of recommender systems: Transitioning from theory to practice, Feb 2026, Computer Science Review Feb 2026
- A survey on sequential recommendation, Nov 2025, Frontiers of Computer Science 2025
- A Comprehensive Review on Harnessing Large Language Models to Overcome Recommender System Challenges, Jul 2025, arxiv
- A Comprehensive Survey on Cross-Domain Recommendation: Taxonomy, Progress, and Prospects, Mar 2025. arxiv
- A Survey on LLM-powered Agents for Recommender Systems, Feb 2025, arxiv: RE: based on DeepSeek-R methods for training reasoning and interleaved LLMs calling search as a tool.
- How Can Recommender Systems Benefit from Large Language Models: A Survey, ACM Transactions on Information Systems 2025
- Graph Foundation Models for Recommendation: A Comprehensive Survey, Feb 2025, arxiv
- Recommender Systems in the Era of Large Language Models (LLMs), TKDE Nov 2024 by subscription
- Multimodal Pretraining, Adaptation, and Generation for Recommendation: A Survey, Jul 2024, arxiv
- A Review of Modern Recommender Systems Using Generative Models (Gen-RecSys), KDD 2024 pdf
- A Comprehensive Survey on Retrieval Methods in Recommender Systems, Jul 2024, arxiv
- A Survey of Generative Search and Recommendation in the Era of Large Language Models, Apr 2024, arxiv
- A survey on large language models for recommendation, WWW 2024 Springer
- Towards Next-Generation LLM-based Recommender Systems: A Survey and Beyond, Oct 2024, arxiv
- Exploring the Impact of Large Language Models on Recommender Systems: An Extensive Review, Feb 2024, arxiv
- Pre-train, Prompt, and Recommendation: A Comprehensive Survey of Language Modeling Paradigm Adaptations in Recommender Systems , Dec 2023, MIT
Conferences, Workshops
Industrial conferences
Tutorials
Software, libraries, frameworks
Agentic Search
also see Evaluation Agentic Search
- Question's Gambit: The First Move Matters in Agentic Deep Search, Sep 2026, arxiv
- Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems, Sep 2026, arxiv
- Iris: Climbing to the Search Frontier, Sep 2026, arxiv
- ITER: Interaction-Aware Retrieval for Agentic Search, Aug 2026, arxiv
- Beyond Document Retrieval: Architectural Challenges When LLM Agents Query Structured Enterprise Data, Aug 2026, arxiv
- Deep Agentic Search for Repository-Level Code Question Answering: An Empirical Study, Aug 2026, arxiv
- S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents, Jun 2026, arxiv
- SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions, Mar 2026, arxiv
- AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems, Jun 2025, arxiv
- Inference-Time Budget Control for LLM Search Agents, May 2026, arxiv
- Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction, May 2026, arxiv
- LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, May 2026, arxiv
- LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG, May 2026, arxiv
- Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval, May 2026, arxiv
- Rethinking Recommendation Paradigms: From Pipelines to Agentic Recommender Systems, Apr 2026, arxiv
- Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests, Jan 2026, arxiv
- LRAS: Advanced Legal Reasoning with Agentic Search, Jan 2026, arxiv
- Agentic-R: Learning to Retrieve for Agentic Search, Jan 2026, Baidu, arxiv
- Can Small Agent Collaboration Beat a Single Big LLM?, Jan 2026, arxiv
- Instructed Retriever: Unlocking System-Level Reasoning in Search Agents, Databricks, Jan 2026. databricks
- Dr. Zero: Self-Evolving Search Agents without Training Data, Jan 2026, arxiv
- A Systematic Framework for Enterprise Knowledge Retrieval: Leveraging LLM-Generated Metadata to Enhance RAG Systems, Dec 2025, arxiv
- Laser: Governing Long-Horizon Agentic Search via Structured Protocol and Context Register, Dec 2025, arxiv
- Search-o1: Agentic Search-Enhanced Large Reasoning Models, Nov 2025, arxiv
- SafeSearch: Do Not Trade Safety for Utility in LLM Search Agents, Oct 2025, arxiv
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications, Oct 2025, arxiv
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection, Oct 2025, arxiv
- Towards Agentic Self-Learning LLMs in Search Environment, Oct 2025, arxiv
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play, Sep 2025, arxiv
- DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling, Sep 2025, arxiv
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL, Aug 2025, arxiv
- Towards AI Search Paradigm, Jun 2025, arxiv
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge, Jun 2025, arxiv
- R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning, Jun 2025, arxiv
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge, Jun 2025, arxiv
- MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability, May 2025, arxiv
- Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers, May 2025, arxiv
- Open Deep Search: Democratizing Search with Open-source Reasoning Agents, Mar 2025, arxiv
- Synergizing RAG and Reasoning: A Systematic Review, Apr 2025, arxiv
- Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook, Mar 2025, arxiv
- A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval, Mar 2025, arxiv
- Search-o1: Agentic Search-Enhanced Large Reasoning Models, Jan 2025, arxiv
- Plan*RAG: Efficient Test-Time Planning for Retrieval Augmented Generation, Oct 2024, arxiv
- Agentic Information Retrieval, Oct 2024, arxiv
- MindSearch: Mimicking Human Minds Elicits Deep AI Searcher, Jul 2024, arxiv
Heterogeneous and Enterprise Agentic Search
- Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration, Sep 2026, arxiv
- SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs, September 2026, arxiv
- Agent-Enhanced Heterogeneous Graph RAG for Academic Question Answering, September 2026, arxiv
- Benchmarking Hybrid Deep Research Across Database Querying and Web Search, September 2026, arxiv
- Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems, September 2026, arxiv
- VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge, August 2026, arxiv
- Introducing Agentic Search, Mistral AI, August 2026, Mistral
- SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs, P July 2026, ACL Anthology
- Unlocking Dependable Responses with Gemini Enterprise Agent Platformβs Agentic RAG, Google Research, June 2026, Google Research
- AgenticRAG: Agentic Retrieval for Enterprise Knowledge Bases, Microsoft Research, May 2026, arxiv
- Knowledge Graph RAG: Agentic Crawling and Graph Construction in Enterprise Documents, April 2026, arxiv
FreshLLM and similar architectures (LLM and large scale search)
The section should be rewritten, there are a lot of changes from FreshLLM time
- Open Deep Search: Democratizing Search with Open-source Reasoning Agents, Mar 2025, arxiv
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching, Alibaba, May 2025, arxiv
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning, Mar 2025, arxiv
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Mar 2025, arxiv
- When Search Engine Services Meet Large Language Models: Visions and Challenges, Dec 2024, IEEE
- Long-form factuality in large language models, Mar 2024, arxiv
- When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively, Apr 2024, arxiv
- Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, May 2024, arxiv
- FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation, Oct 2023. arxiv
- Gorilla: Large Language Model Connected with Massive APIs, May 2023, arxiv
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions, Dec 2022, arxiv
Conversational Search
- Agentic Conversational Search with Contextualized Reasoning via Reinforcement Learning, Jan 2026, arxiv
- CTR-Guided Generative Query Suggestion in Conversational Search, EMNLP 2025, ACL
- A Survey of Conversational Search, Sep 2025, ACM
- Learning Contextual Retrieval for Robust Conversational Search, EMNLP 2025, ACL
- A Survey of Conversational Search, Oct 2024, arxiv
- Engineering Conversational Search Systems: A Review of Applications, Architectures, and Functional Components, Jul 2024, arxiv
- ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense Retrieval, Apr 2024, arxiv
- CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models, Feb 2024, arxiv
- Generalizing Conversational Dense Retrieval via LLM-Cognition Data Augmentation, Feb 2024, arxiv
- History-Aware Conversational Dense Retrieval, Jan 24, arxiv
- Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search, Findings of EMNLP 2023
- Improving Conversational Passage Re-ranking with View Ensemble, SIGIR 23 Short paper
- ConvGQR: Generative Query Reformulation for Conversational Search, May 2023, arxiv ACL 2023
- Phrase Retrieval for Open-Domain Conversational Question Answering with Conversational Dependency Modeling via Contrastive Learning, arxiv Findings of ACL 2023
- Curriculum Contrastive Context Denoising for Few-shot Conversational Dense Retrieval, SIGIR 2022
- Open-Retrieval Conversational Question Answering, SIGIR 2020
Search Assistance
autocomplete/autosuggest and other search assistance tasks, search clarification, query recommendation and other techniques guiding users in search
- Conversational Query Reformulation with the Guidance of Retrieved Documents, 2026, paper
- From βPeople Also Askβ to Clarifying Questions for Conversational Search Using Parameter-Efficient Fine-Tuning and Prompt Engineering, Jun 2026, paper
- Query Refinement in Dense Retrieval Using LLM-Driven Relevance Feedback, May 2026, paper
- Generating Multi-Aspect Queries for Conversational Search, EACL 2026, ACL Anthology
- Query-guided expansion and contraction of document sets, Feb 2026, paper
- Query Suggestion for Retrieval-Augmented Generation via Dynamic In-Context Learning, Jan 2026, arxiv
- SmartSearch: Process Reward-Guided Query Refinement for Search Agents, Jan 2026, arxiv
- In-Browser Agents for Search Assistance, CHIIR 2026, Mar 2026, paper
- OneSug: The Unified End-to-End Generative Framework for E-commerce Query Suggestion, AAAI 2026 AAAI 2026
- LLM-based Search Assistant with Holistically Guided MCTS for Intricate Information Seeking, 2025, SIGIR 2025
- Evaluating auto-complete ranking for diversity and relevance, ECIR 2025
- Enhancing Discoverability in Enterprise Conversational Systems with Proactive Question Suggestions, Dec 2024, arxiv
- DiAL: Diversity aware listwise ranking for query auto-complete, EMNLP 2024
- Evaluation and Continual Improvement for an Enterprise AI Assistant, Jun 2024, arxiv
- Generating Query Recommendations via LLMs, May 2024, arxiv
- Towards Asking Clarification Questions for Information Seeking on Task-Oriented Dialogues, May 2023, axiv
- Asking Clarification Questions to Handle Ambiguity in Open-Domain QA, May 2023, arxiv
- Asking Clarifying Questions in Open-Domain Information-Seeking Conversations, SIGIR 2019
Multi Turn
- Conversational Query Reformulation with the Guidance of Retrieved Documents, Jul 2026, paper
- Improving Ad-hoc Search Effectiveness for Conversational Information Retrieval via Model Merging, SIGIR 2026, Jul 2026, arxiv
- RECOR: Reasoning-focused Multi-turn Conversational Retrieval Benchmark, ACL 2026, Jul 2026, ACL Anthology
- Agentic Conversational Search with Contextualized Reasoning via Reinforcement Learning, Jan 2026, arxiv
- Beyond Whole Dialogue Modeling: Contextual Disentanglement for Conversational Recommendation, SIGIR 2025, arxiv
- Proactive Guidance of Multi-Turn Conversation in Industrial Search, Baidu, May 2025, arxiv
- LLMs Get Lost In Multi-Turn Conversation, May 2025, arxiv
- Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey, Mar 2025, arxiv
- Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding, Feb 2025, arxiv
- A Survey on Multi-Turn Interaction Capabilities of Large Language Models, Jan 2025, arxiv
- MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems, Jan 2025, arxiv
- CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search, Jun 2024, arxiv
- Generating Multi-turn Clarification for Web Information Seeking, WWW 2024
- An Empirical Analysis on Multi-turn Conversational Recommender Systems, SIGIR 2024
- Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search, CIKM 2024, CIML 2024
- Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search, Oct 2023, arxiv
- ConvGQR: Generative Query Reformulation for Conversational Search, Mar 2023, arxiv
- Few-Shot Conversational Dense Retrieval, SIGIR 2021
Task Solving
- Investigating Users' Search Behavior and Outcome with ChatGPT in Learning-oriented Search Tasks, SIGIR 2024
- Mind2Web: Towards a Generalist Agent for the Web, NeurIPS 2023
Personalization
- Bridging Personalization and Control in Scientific Personalized Search, SIGIR 2025, arxiv
- User-LLM: Efficient LLM Contextualization with User Embeddings, WWW 2025, ACM
- A Survey of Personalization: From RAG to Agent, apr 2025, arxiv
- Can Large Language Models Understand Preferences in Personalized Recommendation?, Jan 2025, arxiv
- Unified Embedding Based Personalized Retrieval in Etsy Search, Sep 2024, arxiv
- IntentRec: Predicting User Session Intent with Hierarchical Multi-Task Learning, Jul 2024 Netflix, arxiv
- LLM-based Medical Assistant Personalization with Short- and Long-Term Memory Coordination, NAACL 2024
Multi modal
- Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation, Jul 2026, ACL
- TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval, Jul 2026, ACL
- VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval, Jul 2026, ACL
- Zero-Shot Multimodal Retrieval with Multi-Scale Contextual Representations, Jul 2026, ACL
- Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering, Jul 2026, ACL
- DORA: A Dual-Objective Reinforcement Learning Framework for Effective and Efficient Multimodal Agentic Search, Jul 2026, ACL
- MMSearch-R1: Incentivizing LMMs to Search, Jul 2026, ACL
- Region-R1: Reinforcing Query-Side Region Cropping for Multi-Modal Re-Ranking, Jul 2026, ACL
- MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment, Jul 2026, ACL Findings
- Towards Long-horizon Agentic Multimodal Search, Apr 2026, arxiv
- VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning, Mar 2026, arxiv
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding, Jul 2026, ACL
- RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation, Mar 2026, ACM
- Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy, Oct 2025, arxiv
- MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion, May 2025, SIGIR 2025, arxiv
- Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions, Jan 2025, IEEEhttps://ieeexplore.ieee.org/abstract/document/10843094?casa_token=oXnLMUJ8EaoAAAAA:bLPPXHI2Sypz5wdjPLTZG965RDQ0jbp6lwbfKi2U3n70i3RWqwBUjHRmxriYp5H2InizkfA40sRs
- RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering, Jan 2025, arxiv
- EA-VTR: Event-Aware Video-Text Retrieval, ECCV 2024, ECCV 2024
- ColPali: Efficient Document Retrieval with Vision Language Models, Jun 2024, arxiv useful practical info vespa blog
- UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain Adaptation, MM 2024
- Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond, Feb 2024, arxiv
- Listen, Think, and Understand, OpenAQA dataset, May 2023, arxiv
- Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering, Apr 2022, arxiv
Multi Lingual
Cross Language Information Retrieval (CLIR) and other technique to make search engines multi lingual
Also, see several multi lingual benchmarks and multi lingual embedding models in the other parts of this survey
- Evaluating Large Language Models for Cross-Lingual Retrieval, Sep 2025, arxiv
- A Comprehensive Evaluation of Embedding Models and LLMs for IR and QA Across English and Italian, May 2025, Advances in Natural Language Processing and Text Mining May 2025
- The Cross-Lingual Cost: Retrieval Biases in RAG over Arabic-English Corpora, Jul 2025, arxiv
- XRAG: Cross-lingual Retrieval-Augmented Generation, Amazon, Heidelberg, May 2025, arxiv
- CLIRudit: Cross-Lingual Information Retrieval of Scientific Documents, Apr 2025, arxiv
- Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems, apr 2024, Google, DeepMind, UoE, arxiv
- Cl2cm: Improving cross-lingual cross-modal retrieval via cross-lingual knowledge transfer, Alibaba, AAAI 2024, AAAI
- Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval, MM 2024
- Cross-Lingual Cross-Modal Retrieval With Noise-Robust Fine-Tuning, IEEE 2024
Question Answering
- CoReQA: Uncovering Potentials of Language Models in Code Repository Question Answering, Jan 2025, arxiv
- Unveiling the power of language models in chemical research question answering, Jan 2025, Nature
- Toward expert-level medical question answering with large language models, Jan 2025, Nature
- LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language Models, Jan 2025, arxiv
- Harnessing Large Language Models for Knowledge Graph Question Answering via Adaptive Multi-Aspect Retrieval-Augmentation, Jan 2025, arxiv
- Assessing The Potential Of Mid-Sized Language Models For Clinical QA, apr 2024, arxiv
- Listen, Think, and Understand, OpenAQA dataset, May 2023, arxiv
- Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering, Apr 2022, arxiv
- A Survey on Employing Large Language Models for Text-to-SQL Tasks, May 2025, ACM Computing Surveys
- Querying Databases with Function Calling, Jan 2025, arxiv
- Large language model for table processing: a survey, Jan 2025, Frontiers of Computer Science
- A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?, Aug 2024, arxiv
- Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL, Jun 2024, arxiv
RAG
- Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation, Jan 2026, arxiv
- How We Built a Semantic Highlight Model To Save Token Cost for RAG, Jan 2026, HuggingFace
- InfoGain-RAG: Boosting Retrieval-Augmented Generation through Document Information Gain-based Reranking and Filtering, Nov 2025, EMNLP 2025
- Each to Their Own: Exploring the Optimal Embedding in RAG, Jul 2025, arxiv
- Synergizing RAG and Reasoning: A Systematic Review, Apr 2025, arxiv
- A Survey on Knowledge-Oriented Retrieval-Augmented Generation, Mar 2025, arxiv
- Sufficient Context: A New Lens on Retrieval Augmented Generation Systems, Google Research, ICLR 2025, Google Research
- Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG, Jan 2025, arxiv
- In Defense of RAG in the Era of Long-Context Language Models, Sep 2024, arxiv
- RAFT: Adapting Language Model to Domain Specific RAG, Jul 2024, open review
- RAGAs: Automated Evaluation of Retrieval Augmented Generation, EACL Demo 2024
- Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?, Jun 2024, arxiv
- RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement, Dec 2024, arxiv
- A Survey on Retrieval-Augmented Text Generation for Large Language Models, Apr 2024, arxiv
- RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation, Mar 2024, arxiv
Knowledge Graphs and RAG
- Is GraphRAG Needed? From Basic RAG to Graph-/Agentic Solutions with Context Optimization, Jun 2026, arxiv
- Millions of GeAR-s: Extending GraphRAG to Millions of Documents, jul 2025, arxiv
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation, Jun 2025, [arxiv(https://arxiv.org/abs/2506.05690)
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement Learning, Jul 2025, arxiv
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation, Jun 2025, arxiv
- RAG vs. GraphRAG: A Systematic Evaluation and Key Insights, Feb 2025, arxiv
- A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models, Jan 2025, arxiv
- Retrieval-Augmented Generation with Graphs (GraphRAG), Dec 2024, arxiv
Industrial blog articles
- We use MongoDB as a graph database to discover deep connections between disparate documents using an LLMβs inherent power to work with structured data, mongodb
- AutoKnow: self-driving knowledge collection for products of thousands of types, Amazon 2020, Amazon Science
- Algolia's Knowledge graphs and ontologies β Adding knowledge to keyword search. algolia
- Building Airbnb Categories with ML and Human-in-the-Loop, airbnb
- Scaling Knowledge Access and Retrieval at Airbnb, airbnb
- Contextualizing Airbnb by Building Knowledge Graph, airbnb
- ebay's Explainable Reasoning over Knowledge Graphs for Recommendation, ebay
- amazon's https://innovation.ebayinc.com/stories/explainable-reasoning-over-knowledge-graphs-for-recommendation, amazon
- Interest Taxonomy: A knowledge graph management system for content understanding at Pinterest, pinterest
- walmart Retail Graph β Walmartβs Product Knowledge Graph, walmart
- Food Discovery with Uber Eats: Using Graph Learning to Power Recommendations, uber
Retrieval
- Connected Content Retriever: Dense Graph Edge Features Powering Pre-Ranking at LinkedIn, Sep 2026, arxiv
- Guiding the coarse levels of semantic IDs makes the fine levels learnable, Sep 2026, arxiv
- Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale, Jul 2026, arxiv
- Scaling Laws for Embedding Dimension in Information Retrieval, Feb 2026, arxiv
- Agentic-R: Learning to Retrieve for Agentic Search, Jan 2026, Baidu, arxiv
- ExpandR: Teaching Dense Retrievers Beyond Queries with LLM Guidance, Nov 2025, EMNLP 2025
- CoEvo: Coevolution of LLM and Retrieval Model for Domain-Specific Information Retrieval, Nov 2025, ACL, 2025 Conference on Empirical Methods in Natural Language Processing
- Large Scale Retrieval for the LinkedIn Feed using Causal Language Models, Oct 2025, arxiv
- DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers, Meta, univ of Waterloo, Feb 2025, arxiv
- CAME: Competitively Learning a Mixture-of-Experts Model for First-stage Retrieval, Jan 2025, ACM
- On the Robustness of Generative Information Retrieval Models: An Out-of-Distribution Perspective, Jan 2025, link
- Fine-Tuning LLaMA for Multi-Stage Text Retrieval, Oct 2023, arxiv
- How Does Generative Retrieval Scale to Millions of Passages?, Google Research, May 2023 arxiv
- How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?, July 2024, arxiv
Ranking for Search
and Recommendations
- Adaptive Re-Ranking, Jun 2026, arxiv
- Joint Optimization of Relevance and Engagement in Multi-Task Ranking for E-Commerce with Efficient LLM Supervision, May 2026, arxiv
- Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design, Meta, Ranking architecture of Meta Ads, Feb 2026, arxiv
- DeepMTL2R: A Library for Deep Multi-task Learning to Rank, Feb 2026, Amazon. arxiv
- Deep Learning to Rank in Industrial Search Engines, Recommender Systems and Online Advertising: An Overview and New Perspectives, ACM, Review, Jan 2026, ACM
- Multi-Objective Recommendation in the Era of Generative AI: A Survey of Recent Progress and Future Prospects, Jun 2025, arxiv
- RankElectra: Semi-supervised Pre-training of Learning-to-Rank Electra for Web-scale Search, KDD 2025, ACM
- Language Model Re-rankers are Fooled by Lexical Similarities, Fact Extraction and VERification Workshop FEVER, Jul 2025, ACL
- MA4DIV: Multi-Agent Reinforcement Learning for Search Result Diversification, WWW 2025, ACM
- MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion, May 2025, SIGIR 2025, arxiv
- A Generative Re-ranking Model for List-level Multi-objective Optimization at Taobao, May 2025, arxiv
- Rank-K: Test-Time Reasoning for Listwise Reranking, May 2025, arxiv
- HIT Model: A Hierarchical Interaction-Enhanced Two-Tower Model for Pre-Ranking Systems, Tencent, May 2025, arxiv
- InteractRank: Personalized Web-Scale Search Pre-Ranking with Cross Interaction Features, PInterest, April 2025, arxiv
- Multi-objective contextual bandits in recommendation systems for smart tourism, Apr 2025, Nature
- Graph-Based Re-ranking: Emerging Techniques, Limitations, and Opportunities, Mar 2025, arxiv
- Cross-Encoder Rediscovers a Semantic Variant of BM25, Feb 2025, arxiv
- Orbit: A framework for designing and evaluating multi-objective rankers, ACM conf on intelligence user interfaces 2025
- Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation, Mar 2025, arxiv
- DISKCO: Disentangling knowledge from cross-encoder to bi-encoder, WWW 2024
- Adaptive Neural Ranking Framework: Toward Maximized Business Goal for Cascade Ranking Systems, WWW 2024
- Multi-objective Learning to Rank by Model Distillation, AirBnB, July 2024, [arxiv](Multi-objective Learning to Rank by Model Distillation)
- RankTower: A Synergistic Framework for Enhancing Two-Tower Pre-Ranking Model, Jul 2024, arxiv
- Bi-CAT: Improving robustness of LLM-based text rankers to conditional distribution shifts, Amazon Science, WWW 2024 workshop
- Fine-Tuning LLaMA for Multi-Stage Text Retrieval, SIGIR 2024
- Unbiased Learning to Rank Meets Reality: Lessons from Baidu's Large-Scale Search Dataset, Apr 2024, arxiv
- RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze!, Dec 2023
- Improving Training Stability for Multitask Ranking Models in Recommender Systems, Google Research, Feb 2023, arxiv
- Multi-Objective Ranking to Boost Navigational Suggestions in eCommerce AutoComplete, WWW 2023
- Multi-objective relevance ranking via constrained optimization, Amazon Science, 2020, amazon
- Multi-objective ranking optimization for product search using stochastic label aggregation Amazon Science 2020, amazon science
Classical bi-encoder and cross encoder ranking,
bert based ranking, hybrid encoder based ranking
- ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models, Feb 2026, arxiv
- A Thorough Comparison of Cross-Encoders and LLMs for Reranking SPLADE, Mar 2024
- Rankt5: Fine-tuning t5 for text ranking with ranking losses, 2023, SIGIR 2023
- ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction, 2022, arxiv
- Pretrained Transformers for Text Ranking: BERT and Beyond, 2021, ACM
- Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering, 2021, arxiv
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT, 2020, arxiv
- Dense Passage Retrieval for Open-Domain Question Answering, 2020, arxiv
- Understanding the Behaviors of BERT in Ranking, 2019, arxiv
- Passage Re-ranking with BERT, 2019, arxiv
Reranking
- Revisiting Text Ranking in Deep Research, SIGIR 2026, Jul 2026, arxiv
- Very Efficient Listwise Multimodal Reranking for Long Documents, May 2026, arxiv
- ResRank: Unifying Retrieval and Listwise Reranking via End-to-End Joint Training with Residual Passage Compression, Apr 2026, arxiv
- EviRerank: Adaptive Evidence Construction for Long-Document LLM Reranking, Jun 2026, arxiv
- Reproducing Adaptive Reranking for Reasoning-Intensive IR, SIGIR 2026, Jul 2026, paper
- APR: Adaptive Personalised Reranking for Conversational Search, SIGIR 2026, Jul 2026, paper
- uva-irlab-conv at SemEval-2026 Task 8: Multi-Turn RAG with Learned Sparse Retrieval and Listwise Reranking, Jun 2026, arxiv
- Rich-Media Re-Ranker: A User Satisfaction-Driven LLM Re-ranking Framework for Rich-Media Search, Baidu, Feb 2026, arxiv
- LANCER: LLM Reranking for Nugget Coverage, Jan 2026, arxiv
- Re-Rankers as Relevance Judges, Jan 2026, arxiv
- RankLLM: A Python Package for Reranking with LLMs, SIGIR 2025, ACM
- Accelerating Listwise Reranking: Reproducing and Enhancing FIRST, SIGIR 2025, ACM
- Supervised Fine-Tuning or Contrastive Learning? Towards Better Multimodal LLM Reranking, Oct 2025, arxiv
- Distillation versus Contrastive Learning: How to Train Your Rerankers, Jul 2025, arxiv
- Rank-K: Test-Time Reasoning for Listwise Reranking, May 2025, arxiv
- Rank1: Test-Time Compute for Reranking in Information Retrieval, Feb 2025, arxiv
Query Understanding
- Scaling Intent Understanding: A Framework for Classification with Clarification using Lightweight LLMs, Mar 2026, EACL
- What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance, Mar 2026, EACL
- Generating Multi-Aspect Queries for Conversational Search, Mar 2026, EACL
- QueStER: Query Specification for Generative Keyword-Based Retrieval, Mar 2026, EACL
- Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning, NVidia Oct 2025, arxiv
- ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning, NVidia, Aug 2025, arxiv
- Powering Job Search at Scale: LLM-Enhanced Query Understanding in Job Matching Systems, Aug 2025, Linkedin, arxiv
- Query Attribute Modeling: Improving search relevance with Semantic Search and Meta Data Filtering, Aug 2025, arxiv
- Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment, Jul 2025, arxiv
- Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion, Apr 2025, arxiv
- Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling. Apr 2025, arxiv
- LLM-Based Query Expansion with Gaussian Kernel Semantic Enhancement for Dense Retrieval, Mar 2025, mdpi
- LLM-QE: Improving Query Expansion by Aligning Large Language Models with Ranking Preferences, Feb 2025, arxiv
- Two Heads Are Better Than One: Improving Search Effectiveness Through LLM-Generated Query Variants, 2025, RMIT University
- Large Language Model based Long-tail Query Rewriting in Taobao Search, WWW 2024
- Near-duplicate question detection, WWW 2024
- Hierarchical query classification in e-commerce search, WWW 2024
- Query Understanding in the Age of Large Language Models, Jun 2023, arxiv
- Decomposing Complex Queries for Tip-of-the-tongue Retrieval, May 2023, arxiv
- ConvGQR: Generative Query Reformulation for Conversational Search, May 2023, arxiv
- Query Rewriting in Retrieval-Augmented Large Language Models, EMNLP 2023
- Query2doc: Query Expansion with Large Language Models, Mar 2023, arxiv
- Few-Shot Generative Conversational Query Rewriting, SIGIR 2020
Time Aware Search
- TimelyRAG: Semantic-Temporal Hybrid Retrieval for Time-Critical Question Answering in Overlapping-Evolving Documents, Sep 2026, arxiv
- Temporal Evidence Chain for Temporal Knowledge Graph Question Answering with Large Language Models, Jul 2026, ACL Anthology
- TDRΒ²A: Time-sensitive decomposition-retrieval-reorganization agent for temporal knowledge graph question answering, Jun 2026, ScienceDirect
- Temporal-spatial reasoning over hypergraph knowledge structures for multimodal retrieval-augmented generation, Aug 2026, Scientific Reports
- SAR: A Structure-Aligned Reasoning Framework for Temporal Knowledge Graph Question Answering, Mar 2026, AAAI
- Time-Aware Complex Question Answering over Temporal Knowledge Graph, Jan 2026, ScienceDirect
- TempQA: An LLM-based framework for temporal knowledge graph question answering, Jan 2026, ScienceDirect
- It's High Time: A Survey of Temporal Question Answering, Jul 2026, ACL Anthology
- MTRM: Multi-Granularity Trend-Aware Retrieval and Modeling for Temporal Knowledge Graph Extrapolation, Jul 2026, IEEE
- Temporal Knowledge Graph Question Answering via Sub-Question Decomposition and Cross-Checking, Aug 2026, IOS Press
- Right Answer at the Right Time - Temporal Retrieval-Augmented Generation via Graph Summarization, Oct 2025, arxiv
- It's High Time: A Survey of Temporal Question Answering, Aug 2025, arxiv
- TimeR4 : Time-aware Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering, Nov 2024, ACL EMNLP 2024
- Time-Sensitve Retrieval-Augmented Generation for Question Answering, Ot 2024, Semantic Scholar
Embedding models
- UEmbed: Unified Sparse and Dense Multimodal Embeddings, Aug 2026, arxiv
- DREAM: Dense Retrieval Embeddings via Autoregressive Modeling, Jun 2026, arxiv
- BitNet Text Embeddings, Jun 2026, arxiv
- Text Embeddings Inference, from Hugging Face inference layer for embeddings, Feb 2026 (hugging face)(https://github.com/huggingface/text-embeddings-inference)
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation, Feb 2026 arxiv, jina-embeddings-v5-text: New SOTA Small Multilingual Embeddings, blog post feb 2026
- What Actually Makes Embedding Model Inference Fast?, Jan 2026, blog post
- Tarka Embedding V1, blog post
- EmbeddingGemma: Powerful and Lightweight Text Representations, Nov 2025, SOA open weight embedding model from Google, 300M paramers, arxiv
- CSMF: Cascaded Selective Mask Fine-Tuning for Multi-Objective Embedding-Based Retrieval, Apr 2025, arxiv
- Granite Embedding Models (multi-lingual embedding models from IBM), Feb 2025 arxiv
- mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data, Feb 2025, arxiv
- Arctic-Embed 2.0: Multilingual Retrieval Without Compromise, Dec 2024, arxiv
- SFR-Embedding from Salesforce in Salesforce blog Oct 2024
- jina-embeddings-v3: Multilingual Embeddings With Task LoRA, Sep 2024, arxiv
- BGE-en-ICL, BGE-ICL embedding model, Making Text Embedders Few-Shot Learners, Sep 2024, arxiv
- OmniSearchSage: Multi-Task Multi-Entity Embeddings for Pinterest Search, Apr 2024, arxiv
- Multilingual E5 Text Embeddings: A Technical Report, Feb 2024, arxiv
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models, May 2024, from Nvidia arxiv
- BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation, Feb 2024, arxiv
- E5-Mistral embeddings from Microsoft in Improving Text Embeddings with Large Language Models Dec 2023
Embedding models evaluation
- Evaluating Large Language Models for Cross-Lingual Retrieval, Sep 2025, arxiv
- A Comprehensive Evaluation of Embedding Models and LLMs for IR and QA Across English and Italian, May 2025, Advances in Natural Language Processing and Text Mining May 2025
- The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks, Apr 2025, arxiv
- MMTEB: Massive Multilingual Text Embedding Benchmark, Feb 2025, arxiv
- Beyond Benchmarks: Evaluating Embedding Model Similarity for Retrieval Augmented Generation Systems, jul 2024, arxiv
- MTEB: Massive Text Embedding Benchmark Oct 2022 [arxiv](https://arxiv.org/abs/2210.07316 Leaderboard) Leaderboard
- Marqo embedding benchmark for eCommerce at Huggingface, text to image and category to image tasks
- The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding, openreview pdf
- MMTEB: Community driven extension to MTEB repository
- Chinese MTEB C-MTEB repository
- French MTEB repository
Optimization of embedding models
- The Future is Sparse: Embedding Compression for Scalable Retrieval in Recommender Systems, May 2025, arxiv
- A Universal Framework for Compressing Embeddings in CTR Prediction, Feb 2025. arxiv
Embedding training
- KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model, Jun 2025, arxiv
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models, Jun 2025, arxiv
- CSMF: Cascaded Selective Mask Fine-Tuning for Multi-Objective Embedding-Based Retrieval, Apr 2025, arxiv
Finetuning embedding models
- Resource-Efficient Adaptation of Large Language Models for Text Embeddings via Prompt Engineering and Contrastive Fine-tuning, July 2025, arxiv
- Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data, May 2025, arxiv
- Pre-training vs. Fine-tuning: A Reproducibility Study on Dense Retrieval Knowledge Acquisition, May 2025, arxiv
- DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers, Feb 2025, arxiv
- Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation, Feb 2025, arxiv
- REFINE on Scarce Data: Retrieval Enhancement through Fine-Tuning via Model Fusion of Embedding Models, Oct 2024, arxiv
Document understanding
- Query-aware index pruning for retrieval under budget constraints, Amazon, Should be document or text in the inddex, Sep 2026, Amazon science
- LongDA: Benchmarking LLM Agents for Long-Document Data Analysis, Jan 2026, arxiv
- Small Language Models for Phishing Website Detection: Cost, Performance, and Privacy Trade-Offs, Nov 2025, arxiv
- SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion, Mar 2025, arxiv
- Qwen2.5-VL Technical Report, see, 3.3.2 Document Understanding and OCR at Feb 2025 arxiv
- ColPali: Efficient Document Retrieval with Vision Language Models, Jun 2024, arxiv
Response Generation
- Improving Generative Ad Text on Facebook using Reinforcement Learning, Jul 2025, arxiv
- Neural headline generation: A comprehensive survey, Mar 2025, Neurocomputing
- Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents, Mar 2025, arxiv
- Beyond Relevant Documents: A Knowledge-Intensive Approach for Query-Focused Summarization using Large Language Models, Aug 2024, arxiv
Deep Research
and Deep Search
Also see Evaluation Deep Research
- Question's Gambit: The First Move Matters in Agentic Deep Search, Sep 2026, arxiv
- BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents, May 2026, arxiv
- DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute, Apr 2026, AllenAI, University of Maryland, arxiv
- Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination, Apr 2026, Salesfore AI, arxiv
- Self-Optimizing Multi-Agent Systems for Deep Research. Apr 2026, arxiv
- LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent, Apr 2026, arxiv
- MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification, Mar 2026, arxiv
- Marco DeepResearch: Unlocking Efficient Deep Research Agents via Verification-Centric Design, Mar 2026, arxiv
- AgentIR: Reasoning-Aware Retrieval for Deep Research Agents, Mar 2026, arxiv
- SAGE: Benchmarking and Improving Retrieval for Deep Research Agents, Feb 2026, arxiv
- How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1, Feb 2026. arxiv
- W&D:Scaling Parallel Tool Calling for Efficient Deep Research Agents, Feb 2026, arxiv
- MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents, Jan 2026, arxiv
- Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models, Jan 2026, arxiv
- SAGE: Steerable Agentic Data Generation for Deep Search with Execution Feedback, Jan 2026, arxiv
- DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering, Jan 2026, arxiv
- RESEARCHRUBRICS: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents, Nov 2025, arxiv
- DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping, Oct 2025, arxiv
- Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics, Oct 2025, arxiv
- A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers, Oct 2025, arxiv
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL, Sep 2025, arxiv
- GraphSearch: An Agentic Deep Searching Workflow for Graph Retrieval-Augmented Generation, Sep 2025, arxiv
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent, Sep 2025, arxiv
- Open Data Synthesis For Deep Research, Aug 2025, arxiv
- A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges, Aug 2025, arxiv
- A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications, jun 2025, arxiv
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents, jun 2025, arxiv
- ManuSearch: Democratizing Deep Search in Large Language Models with a Transparent and Open Multi-Agent Framework, May 2025, arxiv
- SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis, May 2025, arxiv
- WebThinker: Empowering Large Reasoning Models with Deep Research Capability, Apr 2025, arxiv
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments, Apr 2025, arxiv
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search, Apr 2025, arxiv @DeepResearch
- Towards Scientific Intelligence: A Survey of LLM-based Scientific Agents, Mar 2025, arxiv
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools, Feb 2025, arxiv
Prevention of hallucinations in Deep Research
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence, ICLR 2026, ICLR 2026
- DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality, ACL 2026, ACL 2026
- Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision, ACL 2026, ACL 2026
- Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification, ACL 2026, ACL 2026
- Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards, ACL 2026, ACL 2026
- DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments, Jul 2026, arxiv
- Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents, Apr 2026, arxiv
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents, Jun 2025, arXiv
- Removal of Hallucination on Hallucination: Debate-Augmented RAG, May 2025, arXiv
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments, Apr 2025, arXiv
- Monitoring Decoding: Mitigating Hallucination via Evaluating the Factuality of Partial Response during Generation, Mar 2025, arXiv
- FACTCHECKMATE: Preemptively Detecting and Mitigating Hallucinations in LMs, Nov 2025, ACL
Hybrid search vs vector search
- Modernizing Facebook Scoped Search: Keyword and Embedding Hybrid Retrieval with LLM Evaluation, Sep 2025, arxiv
- Efficient Knowledge Graph Construction and Retrieval from Unstructured Text for Large-Scale RAG Systems, Jul 2025, SAP, arxiv
- Deep Retrieval at CheckThat! 2025: Identifying Scientific Papers from Implicit Social Media Mentions via Hybrid Retrieval and Re-Ranking, May 2025, arxiv
- Domain-specific Question Answering with Hybrid Search, Dec 2024, arxiv
- COS-Mix: Cosine Similarity and Distance Fusion for Improved Information Retrieval, Jun 2024, arxiv
- Sparse Meets Dense: A Hybrid Approach to Enhance Scientific Document Retrieval, Jan 2024, arxiv
- Azure AI Search: Outperforming vector search with hybrid retrieval and reranking, Jan 2024, Azure tech blog
- Hybrid Hierarchical Retrieval for Open-Domain Question Answering, Jul 2023, ACL 2023
Recommender Engines
TODO to classify
- Beyond Raw Engagement: A Counterfactual Observability Framework for Recommender Systems at Netflix, Sep 2026, arxiv
- Inherit4Rec: Parameter Inheritance for Efficient Scaling of Recommendation Models, Sep 2026, arxiv
- MuSeR: Scalable Long-sequence Recommendation with Multi-interest Modeling, Sep 2026, arxiv
- Explainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery, Sep 2026, arxiv
- What Makes a Good Semantic ID for Generative Recommendation? A Reproducibility Study, Sep 2026, arxiv
- RecGPT: A User Intent-Centric Next-Generation LLM-Powered Recommender System in Industrial Practice, Aug 2026, ACM Transactions on Information Systems
- Who Are We Recommending To? Recommender Systems in the Agentic Web, Jul 2026, arxiv
- OneLoc: Geo-Aware Generative Recommender Systems for Local Life Service, Feb 2026, WSDM 2026
- Rethinking Recommendation Paradigms: From Pipelines to Agentic Recommender Systems, Apr 2026, arxiv
- Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning, Oct 2025, arxiv
- Rethinking Group Recommender Systems in the Era of Generative AI: From One-Shot Recommendations to Agentic Group Decision Support, Jul 2025, arxiv
- RecGPT: LLM-Driven Intent-Centric Recommender Systems at Industrial Scale, Technical Report,Jul 2025, Alibaba, arxiv
- EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration, Feb 2025, arxiv
- 360Brew: A Decoder-only Foundation Model for Personalized Ranking and Recommendation, Jan 2025, arxiv
- Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations, Baidu, Mar 2025, arxiv
- Personalised outfit recommendation via history-aware transformers, Amazon Science, WSDM 2025
- Representation Learning with Large Language Models for Recommendation, [WWW 2024](WWW 2024)
- Llmrec: Large language models with graph augmentation for recommendation, WSDM 2024
- Improved Estimation of Ranks for Learning Item Recommenders with Negative Sampling, Google CIKM 2024
- Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach, May 2023, arxiv
- Towards Open-World Recommendation with Knowledge Augmentation from Large Language Models, RecSys 2024
- Data-efficient Fine-tuning for LLM-based Recommendation, SIGIR 2024
- LLMRec: Large Language Models with Graph Augmentation for Recommendation, WSDM 2024
- DiffKG: Knowledge Graph Diffusion Model for Recommendation, WSDM 2024
- Bridging Language and Items for Retrieval and Recommendation, Mar 2024 arxiv BLAIR paper
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations, Feb 2024, arxiv
- Trinity: Syncretizing Multi-/Long-tail/Long-term Interests All in One, BydeDance , Feb 2024, arxiv
- Tapping the Potential of Large Language Models as Recommender Systems: A Comprehensive Framework and Empirical Analysis, Jan 2024 arxiv
- Leveraging Large Language Models for Sequential Recommendation, RecSys 2023
- Text Is All You Need: Learning Language Representations for Sequential Recommendation, May 2023, arxiv
- On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models, Sep 2022, arxiv
- Recommendation as Language Processing (RLP): A Unified Pretrain, Personalized Prompt & Predict Paradigm (P5), Mar 2022, arxiv
- Augmenting Netflix Search with In-Session Adapted Recommendations, RecSys 2022
- Sampling-Bias-Corrected Neural Modeling for Large Corpus Item Recommendations, Google RecSys 2019
Sequential Recommendation
- Unsupervised Graph Embeddings for Session-based Recommendation with Item Features, Feb 2025, arxiv
- TagRec: Temporal-Aware Graph Contrastive Learning with Theoretical Augmentation for Sequential Recommendation, IEEE KDE 2025, IEEE KDE
- LLMCDSR: Enhancing Cross-Domain Sequential Recommendation with Large Language Models, Large Language Models Cross-Domain Sequential Recommendation, ACM Transaction on Information Systems 2025
- Plug-In Diffusion Model for Sequential Recommendation, AAAI AI 2024
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations, Feb 2024, arxiv
- EAGER: Two-Stream Generative Recommender with Behavior-Semantic Collaboration, (EAGER, a two-strEAm GEnerative Recommender ) KDD 2024, KDD 2024 arxiv
- Mamba4Rec: Towards Efficient Sequential Recommendation with Selective State Space Models, Mar 2024, arxiv
- Leveraging Large Language Models for Sequential Recommendation, RecSys 2023
- Recommender Systems with Generative Retrieval, (Transformer Index for GEnerative Recommenders TIGER) NeurIPS 2023, NeurIPS 2023
- Efficient On-Device Session-Based Recommendation, ACM Transaction on Information Systems 2023
- XLNet4Rec: Recommendations Based on Users' Long-Term and Short-Term Interests Using Transformer, ICMLA 2023
- How to Index Item IDs for Recommendation Foundation Models, P5, SIGIR 2023, SIGIR 2023
- Text Is All You Need: Learning Language Representations for Sequential Recommendation, May 2023, arxiv
- Multi-Behavior Sequential Transformer Recommender, SIGIR 2024, arxiv
- Recommendation as Language Processing (RLP): A Unified Pretrain, Personalized Prompt & Predict Paradigm (P5), RecSys 2022
- Transformers4Rec: Bridging the Gap between NLP and Sequential / Session-Based Recommendation, RecSys 2021
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer, CIKM 2019, CIKM 2019
- Self-Attentive Sequential Recommendation, (SAS4REC), 2018, IEEE Explore
- Session-based Recommendations with Recurrent Neural Networks, (GRU4EC) 2015, arxiv
Discovery
- LLMs for User Interest Exploration in Large-scale Recommendation Systems, Generative AI and Recommender Systems Workshop at KDD 2024, work by Google, pdf
Unclassified
methods (unclassified. TODO classify). methods used in search engines
- Translational Generative Retrieval via Potential Query Generation, ICASSP 2025
- RouteLLM: Learning to Route LLMs with Preference Data, Jun 2024, arxiv
- INTERS: Unlocking the Power of Large Language Models in Search with Instruction Tuning, Jan 2024, arxiv
- A Comprehensive Study of Knowledge Editing for Large Language Models, Jan 2024, arxiv
- Zero-shot Document Retrieval with Hybrid Pseudo-document Retriever, ICASSP 2025
- Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach, Dec 2024, ACM
- Representation Learning with Large Language Models for Recommendation, WWW 2024
Recommender Rankers
- Large Language Models are Zero-Shot Rankers for Recommender Systems, Mar 2024, LLMRank, Springer
Industrial approaches
- ZeroEntropy π - Advanced AI Search Over Complex Documents launch doc
Evaluation of Search engines
- Q2D-Web: Evaluating First-Stage Retrievers at Scale, Sep 2026, Perplexity, blog post
- WildSEEK: Evaluating Language Models for Information-Seeking, Aug 2026, Stanford, IT University of Copenhagen etc arxiv
- As It Was: Aligning LLM Search Evaluation with Historical User Preferences, SIGIR 2026, Jul 2026, arxiv
- BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search Agents, Jul 2026, ACL 2026
- Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems, Jul 2026, ACL 2026
- AgentSearchBench: A Benchmark for AI Agent Search in the Wild, Apr 2026, arxiv
- Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests, SIGIR 2026
- WideSearch: Benchmarking Agentic Broad Info-Seeking, project page, Apr 2026, project page at github arxiv
- Evaluating the Search Agent in a Parallel World, Mar 2026, arxiv
- DRBench: A Realistic Benchmark for Enterprise Deep Research, ICLR 2026, ICLR 2026
- SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval, May 2026, arxiv
- Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents, May 2026, source verification, arxiv
- AgentIR: Reasoning-Aware Retrieval for Deep Research Agents, Mar 2026, arxiv
- SAGE: Benchmarking and Improving Retrieval for Deep Research Agents, Feb 2026, arxiv
- CLUE: Using Large Language Models for Judging Document Usefulness in Web Search Evaluation, CIKM 2025
- Harnessing the Power of Interleaving and Counterfactual Evaluation for Airbnb Search Ranking, Aug 2025, arxiv
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge, Jun 2025, arxiv
- Search Arena & What Weβre Learning About Human Preference, blog post of LMArena
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge, Jun 2025, arxiv
- LLM-Driven Usefulness Judgment for Web Search Evaluation, Apr 2025, arxiv
- FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents, Apr 2025, arxiv
- Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation, DeepMind, Mar 2025, arxiv
- LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?, Jan 2025, arxiv
- AI Search Has A Citation Problem, Mar 2025, CJR Columbia Journalism Review
- Large Language Models for Relevance Judgment in Product Search, Jul 2024, arxiv
- STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases, Apr 2024, arxiv
- Report on the 1st Workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) at SIGIR 2024, SIGIR
Evaluation of RAG
and Question Answering
and knowledge assistants and information seeking LLM based systems
- RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition, Jun 2025, arxiv
- MMTEB: Massive Multilingual Text Embedding Benchmark, Feb 2025, hugging face, leaderboard Brief: 1043 languages in total, primarily in Bitext mining (text pairing), but also 255 in classification, 209 in clustering, and 142 in Retrieval., 550 tasks, anything from sentiment analysis, question-answering reranking, to long-document retrieval. 17 domains, like legal, religious, programming, web, social, medical, blog, academic, etc. Across this collection of tasks, we subdivide into a lot of separate benchmarks, like MTEB(eng, v2), MTEB(Multilingual, v1), MTEB(Law, v1). Our new MTEB(eng, v2) is much smaller and faster than the original English MTEB, making submissions much cheaper and simpler. from Tom Aarsen's linkedin
- MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems, Jan 2025, arxiv
- RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues, Sep 2024, arrxiv
- IRSC: A Zero-shot Evaluation Benchmark for Information Retrieval through Semantic Comprehension in Retrieval-Augmented Generation Scenarios, Sep 2024, arxiv
- Evaluating Retrieval Quality in Retrieval-Augmented Generation, Apr 2024, arxiv
- Evaluation of Retrieval-Augmented Generation: A Survey, May 2024, arxiv
- RAGAS: Automated Evaluation of Retrieval Augmented Generation Jul 23, arxiv
- ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems Nov 23, arxiv
- TREC iKAT 2023: A Test Collection for Evaluating Conversational and Interactive Knowledge Assistants, SIGIR 2024
- MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries, Jan 2024, arxiv
- FaithDial: A Faithful Benchmark for Information-Seeking Dialogue , Dec 2022, MIT Press
- Open-Retrieval Conversational Question Answering, SIGIR 2020
- XOR QA: Cross-lingual Open-Retrieval Question Answering, Oct 2020, arxiv
Evaluation Deep Research
- DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality, Mar 2026, arxiv
- DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent, Mar 2026, arxiv
- Deep Research Arena: The First Exam of LLMsβ Research Abilities via Seminar-Grounded Tasks, Mar 2026, AAAI
- InnovatorBench: Evaluating Agentsβ Ability to Conduct Innovative LLM Research, Oct 2025 arxiv
- Deep Research Agents: Major Breakthrough or Incremental Progress for Medical AI?, Mar 2026, JMIR
- DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis, Aug 2025, arxiv
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent, Aug 2025, arxiv
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents, Jun 2025, arxiv
- DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research, May 2025, arxiv
- AstaBench (from AllenAI), Benchmark at Guthub
- FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks, May 2025, arxiv
- GAIA: a benchmark for General AI Assistants, Nov 2023, arxiv
Evaluation Agentic Search
- SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs, Sep 2026, arxivv
- CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information, Sep 2026, arxiv
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications, Oct 2025, arxiv
- WideSearch: Benchmarking Agentic Broad Info-Seeking, Aug 2025, arxiv
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent, Aug 2025, arxiv
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge, Jun 2025, arxiv
- DeepShop: A Benchmark for Deep Research Shopping Agents, Jun 2025, arxiv
- Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks, May 2025, arxiv
- InfoDeepDeek emergntmind
- WebArena: A Realistic Web Environment for Building Autonomous Agents, Apr 2024, arxiv
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents, Dec 2024, arxiv
Evaluation Reasoning and RAG
- R2MED: A Benchmark for Reasoning-Driven Medical Retrieval, May 2025, arxiv
- GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation, Jun 2025, arxiv
- MR2-BENCH: GOING BEYOND MATCHING TO REA-SONING IN MULTIMODAL RETRIEVAL, Sep 2025, arxiv
- BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval, Jul 2024, arxiv
QA Benchmarks
QA is used in many vertical domains, see Vertical section bellow
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines, Mar 2025, arxiv
- CoReQA: Uncovering Potentials of Language Models in Code Repository Question Answering, Jan 2025, arxiv
- Unveiling the power of language models in chemical research question answering, Jan 2025, Nature, communication chemistry ScholarChemQA Dataset
- Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses, Oct 2024, Salesforce, arxiv Answer Engine (RAG) Evaluation Repository
- HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly, Oct 2024, arxiv
- Introducing SimpleQA, OpenAI, Oct 2024 OpenAI
- NovelQA: A Benchmark for Long-Range Novel Question Answering, Mar 2024, arxiv
- NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens, Mar 2024, arxiv
- Are Large Language Models Consistent over Value-laden Questions?, Jul 2024, arxiv
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Aug 2023, arxiv
- L-Eval: Instituting Standardized Evaluation for Long Context Language Models, Jul 2023. arxiv
- A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers, QASPER, May 2021, arxiv
- MultiDoc2Dial: Modeling Dialogues Grounded in Multiple Documents, EMNLP 2021, ACL
- CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge, Jun 2019, ACL
- Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering, Sep 2018, arxiv OpenBookQA dataset at AllenAI
- Jin, Di, et al. "What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams., 2020, arxiv MedQA
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, 2018, arxiv ARC Easy dataset ARC dataset
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions, 2019, arxiv BoolQ dataset
- BookQA: Stories of Challenges and Opportunities, Oct 2019, arxiv
- HellaSwag, HellaSwag: Can a Machine Really Finish Your Sentence? 2019, arxiv Paper + code + dataset https://rowanzellers.com/hellaswag/
- PIQA: Reasoning about Physical Commonsense in Natural Language, Nov 2019, arxiv
PIQA dataset
- Crowdsourcing Multiple Choice Science Questions arxiv SciQ dataset
- The NarrativeQA Reading Comprehension Challenge, Dec 2017, arxiv dataset at deepmind
- WinoGrande: An Adversarial Winograd Schema Challenge at Scale, 2017, arxiv Winogrande dataset
- TruthfulQA: Measuring How Models Mimic Human Falsehoods, Sep 2021, arxiv
- TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages, 2020, arxiv data
- Natural Questions: A Benchmark for Question Answering Research, Transactions ACL 2019
AI Scientists for Search
- AI Co-Scientist for Ranking: Discovering Novel Search Ranking Models alongside LLM-based AI Agents with Cloud Computing Access, Mar 2026, arxiv
Blog posts, whitepapers
- Increase web search accuracy and efficiency with dynamic filtering, Feb 2026, Anthropic blog
- Pinterest Feb 2026 Serving two tower models using GPU
-- Beyond 2 tower model for ranking
-- Re-designing ads serving systems
-- Non-engagement signals
-- Offsite content understanding
- Perplexity, firmly Build Merchant Network to Power GenAI Commerce, Mar 2025, press release
- Adobe Analytics: Traffic to U.S. retail websites from Generative AI sources jumps 1,200 percent, Mar 2025, adobe
- Foundation Model for Personalized Recommendation by Netflix, Mar 2025, Netflix blog
- Improving Recommendation Systems & Search in the Age of LLMs by Eugene Neyan, Mar 2025, blog post
- A Coding Implementation to Build a Conversational Research Assistant with FAISS, Langchain, Pypdf, and TinyLlama-1.1B-Chat-v1.0, Mar 2025, marktechpost
- Investigating ChatGPT Search: Insights from 80 Million Clickstream Records, Feb 2025, SemRush blog
- Query Expansion with LLMs: Searching Better by Saying More, Feb 2025, Jina ai
- Transformers in music recommendation, Google on how transformers are used for music recommendation at youtube, google research blog
- Scaling the Instagram Explore Recommendations Systems, Meta 08 2023
- PDF Retrieval with Vision Language Models, about ColPali and using it for document search from Vespa
- Evaluating search relevance part 2 - Phi-3 as relevance judge, a series of articles from ElasticSearch, practical experience on using Phi-3 llm family for relevance evaluation elasticsearch
Verticals
Product Search
- Beyond Relevance: Structured Semantic Supervision for Product Search with LLM-Augmented Annotations, Sep 2026, arxiv
- Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce, Sep 2026, arxiv
- From Clicks to Clips: A Multimodal Retrieval System for E-Commerce Video Recommendations, Sep 2026, Springer
- ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants, Jul 2026, ACL Anthology
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders, Jul 2026, ACL Anthology
- AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce, Jul 2026, ACL Anthology
- Enhancing Multimodal Product Retrieval in E-Commerce by Reversing Typographic Attacks, Jun 2026, PMLR
- Semantic Retrieval for Product Search in E-Commerce, May 2026, arxiv
- SHOPPER: Semantic, Historical, Order-aware, and Product-conditioned Pooling for E-commerce Recommendation, 2026, AdKDD
- ARCHER: Shooting Straight in Multimodal E-Commerce Search at Alibaba with Progressive Alignment, Apr 2026, ACM WWW 2026
- Hierarchical Agentic RAG Framework for Intelligent Product Discovery in Distributed E-Commerce Microservices Using MCP and Dynamic Context-Aware Vector Intelligence, Jun 2026, IJCTT
- Generative Retrieval and Alignment Model: A New Paradigm for E-commerce Retrieval, apr 2025, arxiv
- Personalised outfit recommendation via history-aware transformers, Amazon Science, WSDM 2025
- Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents, Mar 2025, arxiv
- Automated Query-Product Relevance Labeling using Large Language Models for E-commerce Search, Feb 2025, arxiv
- Behavior Modeling Space Reconstruction for E-Commerce Search, Jan 2025, arxiv
- Enhancing Relevance of Embedding-based Retrieval at Walmart, Oct 2024, CIKM 2024
- Towards translating objective product attributes into customer language, Amazon Science
- Manipulating Large Language Models to Increase Product Visibility, Sep 2024, arxiv
- Hierarchical query classification in e-commerce search, WWW 2024
- An interpretable ensemble of graph and language models for improving search relevance in e-commerce, WWW 2024
- Web-Scale Semantic Product Search with Large Language Models, May 2023, KDDM 2023
- Web-scale semantic product search with large language models, Amazon Science, PAKDD 2023
- Behavior-driven query similarity prediction based on pre-trained language models for e-commerce search, Amazon Science, SIGIR 2023 eCommerce workshop
- Rethinking E-Commerce Search, Instacart, 2023, arxiv
- Overview of the TREC 2023 Product Product Search Track, TREC 2023
Location Aware (Maps, real estate, local. travel)
- Spatial-RAG: Spatial Retrieval Augmented Generation for Real-World Geospatial Reasoning Questions, Jul 2026, ACL Anthology
- Reasoning Over Space: Enabling Geographic Reasoning for LLM-Based Generative Next POI Recommendation, Jul 2026, ACL Anthology
- CompassLLM: A Multi-Agent Approach toward Geo-Spatial Reasoning for Popular Path Query, Jul 2026, ACL Anthology
- MFC4POI: Multi-factor collaboration for next point-of-interest recommendation using large language models, Nov 2026, ScienceDirect
- Graph-Enhanced Large Language Models for Spatial Search, Jun 2026, arxiv
- Revisiting General Map Search via Generative Point-of-Interest Retrieval, May 2026, arxiv
- GeoAgentic-RAG: A Multi-Agent framework for autonomous geospatial reasoning and visual insight generation with LLM, 2026, ScienceDirect
- Intelligent Multimodal Retrieval and Reasoning for Geospatial Knowledge Discovery on the I-GUIDE Platform, Jun 2026, arxiv
- Towards language-based retrieval of complex geospatial data: A case study on UK National geographic database, 2026, ScienceDirect
- RALLM-POI: Retrieval-Augmented LLM for Zero-Shot Next POI Recommendation with Geographical Reranking, Apr 2026, Springer
- Spatial RAG for Big Data-Driven urban itinerary recommendation, 2026, CNR
- TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning, Mar 2026, AAAI
- Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection, Jun 2026, arxiv
- Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation, Aug 2026, ISPRS Archives
- Learning to Rank for Maps at Airbnb, KDD 2024
- Transforming Location Retrieval at Airbnb: A Journey from Heuristics to Reinforcement Learning, CIKM 2024
- Optimizing Airbnb Search Journey with Multi-task Learning SIGKDD 2023
- Learning To Rank Diversely At Airbnb, CIKM 2023
- Improving Deep Learning for Airbnb Search, KDD 2020
- Real-time Personalization using Embeddings for Search Ranking at Airbnb, KDD 2018
Ads / advertisement
- Improving Generative Ad Text on Facebook using Reinforcement Learning, Jul 2025, arxiv
- TeamCMU at TouchΓ©: Adversarial Co-Evolution for Advertisement Integration and Detection in Conversational Search, Jul 2025, arxiv
- Set-based state estimation of nonlinear discrete-time systems using constrained zonotopes and polyhedral relaxations, Mar 2025, arxiv
- Semantic Ads Retrieval at Walmart eCommerce with Language Models Progressively Trained on Multiple Knowledge Domains, Mar 2025, arxiv
- Applying Deep Learning to Ads Conversion Prediction in Last Mile Delivery Marketplace, Feb 2025, DoorDash, arxiv
- Scaling Laws for Online Advertisement Retrieval, Nov 2024, arxiv
Real estate
- Beyond Relevance: A Demand Balancer Model for Rental Platforms with Single-Unit Inventory, WSDM 2025
Healthcare
- MedExpQA: Multilingual benchmarking of Large Language Models for Medical Question Answering, Sep 2024, AI in Medicine
- Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries, WWW 2024
- JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability, Jun 2024, arxiv
- BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains, Feb 2024, arxiv
- DISC-MedLLM: Bridging General Large Language Models and Real-World Medical Consultation, Aug 2023, arxiv
- Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding, May 2023, arxiv
- MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data, Apr 2023, arxiv
Science
- Reimagining research papers as interactive and reliable AI agents, Sep 2026, Nature
- IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review, Jul 2026, ACL Anthology
- SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration, Jul 2026, ACL Anthology
- Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework, Jul 2026, ACL Anthology
- Agentic-R: Learning to Retrieve for Agentic Search, Jul 2026, ACL Anthology
- A multi-agent system for automating scientific discovery, May 2026, Nature
- Accelerating scientific discovery with Co-Scientist, May 2026, Nature
- An AI system to help scientists write expert-level empirical software, May 2026, Nature
- MAKES-QA: A multi-agent framework for knowledge graph construction and enrichment over scientific literature for question answering, Jul 2026, ScienceDirect
- Empowering biomedical evidence exploration and synthesis with deep knowledge graph research, Jul 2026, Nature Machine Intelligence
- Bridging data and discovery: a survey on knowledge graphs in AI for science, Mar 2026, National Science Review
- Synthesizing scientific literature with retrieval-augmented language models, Feb 2026, Nature
- Towards end-to-end automation of AI research, 2026, Nature
- Adaptive Mining of Scientific Knowledge Graphs via Reinforcement Learning, 2026, ScienceDirect
- Mapping scholarly knowledge: A systematic review of Knowledge Graphs for academic papers, 2026, ScienceDirect
- A Survey of Large Language Model-Based Search Agents, Jul 2026, ACL Anthology
- PaSa: An LLM Agent for Comprehensive Academic Paper Search, Jan 2025, arxiv
Finance
- FinSight: Towards Real-World Financial Deep Research, Jul 2026, ACL Anthology
- FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering, Jul 2026, ACL Anthology
- FinMRAGBench: A Realistic and Complex Benchmark for Multi-Modal RAG in Financial Document Analysis, Jul 2026, ACL Anthology
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning, Apr 2026, ICLR
- MimirRAG: A Multi-Agent RAG Framework for Financial Data Retrieval with Metadata Integration, May 2026, arxiv
- Agentic Retrieval-Augmented Generation for Financial Document Question Answering, May 2026, arxiv
- Measuring and Closing the Retrieval Gap in Financial Question Answering, May 2026, PMLR
- An Auditable LLM-RAG Architecture for Financial Document Intelligence and Decision Support, May 2026, MDPI
- A Chinese financial event knowledge graph-based retrieval-augmented generation framework for financial question answering, Jul 2026, ScienceDirect
- HybridRAG-Finance: Agent-Guided Information Summarization with Vector and Graph Retrieval for Financial Documents, Jun 2026, SSRN
- AgenticRAG: Agentic Retrieval for Enterprise Knowledge Bases, May 2026, Microsoft Research
- Sustainable Hybrid Document-Routed Retrieval for financial RAG: Resolving the robustness-precision trade-off, Dec 2026, ScienceDirect
Legal
- Legalyze: An Open-Source Structure-Aware Agentic RAG Platform for High-Precision Legal Intelligence, Sep 2026, Wiley
- Agentic RAG for Legal Question Answering in Civil Law: Evidence From the Korean Bar Examination, Aug 2026, IEEE
- LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal Reasoning, Jul 2026, ACL Anthology
- LLM Agents in Law: Taxonomy, Applications, and Challenges, Jul 2026, ACL Anthology
- Leverage Knowledge Graph and Large Language Model for law article recommendation: A case study of Chinese criminal law, Apr 2026, ScienceDirect
- LEXA: Legal case retrieval via graph contrastive learning with contextualised LLM embeddings, Mar 2026, Springer
- Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents, Mar 2026, arxiv
- Legal RAG Bench: an end-to-end benchmark for legal RAG, Mar 2026, arxiv
- Benchmarking Legal RAG: The Promise and Limits of AI Statutory Surveys, Feb 2026, arxiv
- Automated data synthesis and retrieval-augmented generation for legal large language models, Jul 2026, ScienceDirect
- Conversational vs Traditional: Comparing Search Behavior and Outcome in Legal Case Retrieval, SIGIR 21 short paper
Search Engine Optimization
Search Engine Optimization, adversial
- Dynamics of Adversarial Attacks on Large Language Model-Based Search Engines, Jan 2025, arxiv
- Adversarial Search Engine Optimization for Large Language Models, Jul 2024, arxiv
- Stealthy Attack on Large Language Model based Recommendation, Feb 2024, arxiv
Job
- Powering Job Search at Scale: LLM-Enhanced Query Understanding in Job Matching Systems, Aug 2025, Linkedin, arxiv
Maintenance, Repair, Manufacturing
- Beyond Maintenance Manual Multimodal RAG: Suggesting What Tool, Sep 2026, arxiv
- Agentic AI for maintenance management: a process-centric review and a staged framework for industrial adoption, Oct 2026, ScienceDirect
- Reducing Technician Search Burden: A Multimodal RAG for Cessna 172 Maintenance Manual, Aug 2026, arxiv
- A Multi-Agent and synergistic Knowledge Graph retrieval-augmented generation framework for intelligent maintenance, Apr 2026, ScienceDirect
- Automating Information Extraction and Retrieval for Industrial Spare Parts Pooling, Jun 2026, arxiv
- A topological-graph and regulation-constrained retrieval-augmented generation method for refinery and petrochemical pipeline maintenance plan generation, 2026, ScienceDirect
- Hierarchical multi-agent reinforcement learning for retrieval-augmented industrial document question answering, Mar 2026, Nature
- Retrieval-augmented generation enhanced LLM for industrial anomaly detection, May 2026, Springer
- A collaborative approach based on large language model and knowledge graphs for information integration towards smart manufacturing, 2026, ScienceDirect
- Large language models in manufacturing: a comprehensive review, Jul 2026, Springer
- Cache-augmented multimodal generative AI for energy-aware predictive maintenance, Aug 2026, ScienceDirect
- Automated Structuring and Analysis of Unstructured Equipment Maintenance Text Data in Manufacturing Using Generative AI Models, Feb 2026, MDPI
- Generative AI in Manufacturing and Industrial Contexts: A Systematic Review of Applications, Challenges, and Future Directions, Aug 2026, MDPI
- A Compliance-Preserving Retrieval System for Aircraft MRO Task Search, Nov 2025, arxiv
- Bridging Industrial Expertise and XR with LLM-Powered Conversational Agents, Apr 2025, arxiv
- Prescriptive Agents based on RAG for Automated Maintenance (PARAM), Jul 2025, arxiv
- Enhancing Manufacturing Knowledge Access with LLMs and Context-aware Prompting, Jul 2025, arxiv
- MetalMind: A knowledge graph-driven human-centric knowledge system for metal additive manufacturing, Jun 2025, Nature NPJ Advanced manufacturing
- Optimizing Aerospace Product Maintenance A Novel Multi-Modal Knowledge Graph and LLM Approach for Enhanced Decision Support, Jul 2024, ESWC conference
OLD TOC