A Survey on Large Language Model-Based Game Agents (ACM CSUR)
🔥 Must-read papers for LLM-based Game agents.
📘 Our survey has been accepted by ACM Computing Surveys (CSUR). We are preparing the camera ready. Feel free to reach out if you find missing reference.
💫 We continuously update the GitHub list on a weekly basis.
📝 If you discover any papers that are suitable but not yet included, please open an issue or submit a pull request.
[2026/01] MineNPC-Task: Task Suite for Memory-Aware Minecraft AgentsarXiv [paper] #minecraft
[2025/12] Synergizing Code Coverage and Gameplay Intent: Coverage-Aware Game Playtesting with LLM-Guided Reinforcement LearningarXiv [paper] #minecraft#training
[2025/11] Knowledge Graph-enhanced Large Language Model for Incremental Game PlayTestingIEICE Transactions on Information and Systems 2025 [paper] #minecraft#memory
[2025/09] PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments2025 IEEE Conference on Games (CoG) 2025 [paper] #minecraft#multi-agent#self-improvement
[2025/09] Experience-based Knowledge Correction for Robust Planning in MinecraftICLR 2026 Poster [paper] #minecraft#planning
[2025/08] CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative TasksFindings of EMNLP 2025 [paper] #minecraft#planning#multi-agent
[2025/08] Vistawise: Building Cost-effective Agent with Cross-modal Knowledge Graph for MinecraftEMNLP 2025 [paper] #minecraft#memory#tool-use
[2025/07] VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning SystemsarXiv [paper] #minecraft#vlm
[2025/07] Referential ambiguity and clarification requests: comparing human and LLM behaviourProceedings of the Eighth Workshop on Computational Models of Reference, Anaphora and Coreference 2025 [paper] #minecraft
[2025/06] Optimus-3: Towards Generalist Multimodal Minecraft Agents with Scalable Task ExpertsarXiv [paper][code] #minecraft#planning
[2025/06] Matrix-Game: Interactive World Foundation ModelarXiv [paper][code] #minecraft
[2025/06] GuessBench: Sensemaking Multimodal Creativity in the WildarXiv [paper] #minecraft#training#vlm
[2025/05] Don’t Just Follow MLLM Plans: Robust and Efficient Planning for Open-World AgentsarXiv [paper] #minecraft#planning#vlm
[2025/05] Knowledge Retrieval in LLM Gaming: A Shift from Entity-Centric to Goal-Oriented GraphsKnowledge-Based Systems 2025 [paper] #minecraft#planning#memory
[2025/05] BeliefNest: A Joint Action Simulator for Embodied Agents with Theory of MindarXiv [paper] #minecraft
[2025/05] MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural LearningNeurIPS 2025 poster [paper] #minecraft#training
[2025/05] WALL-E: World Alignment by NeuroSymbolic Learning improves World Model-based LLM AgentsNeurIPS 2025 poster [paper] #minecraft#planning#world-model#prompting
[2025/04] Collaborating Action by Action: A Multi-agent LLM Framework for Embodied ReasoningarXiv [paper][code] #minecraft#multi-agent#training
[2025/04] WALL-E 2.0: World Alignment by NeuroSymbolic Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model#prompting
[2025/03] Uncertainty in Action: Confidence Elicitation in Embodied AgentsarXiv [paper] #minecraft
[2025/03] Parallelized Planning-Acting for Efficient LLM-based Multi-Agent SystemsarXiv [paper] #minecraft#planning#multi-agent#tool-use
[2025/03] NeSyC: A Neuro-symbolic Continual Learner For Complex Embodied Tasks In Open DomainsICLR 2025 Poster [paper] #minecraft
[2025/03] Word2Minecraft: Generating 3D Game Levels through Large Language ModelsarXiv [paper][code] #minecraft#generation
[2025/03] Plancraft: an evaluation dataset for planning with LLM agentsCOLM 2025 [paper] #minecraft#planning#memory#tool-use
[2025/02] GATE: Graph-based Adaptive Tool Evolution Across Diverse TasksarXiv [paper][code] #minecraft
[2025/02] Optimus-2: Multimodal World Model for Open-World Minecraft AgentsCVPR 2025 [paper][code] #minecraft#planning#world-model#vlm
[2025/01] LARM: Large Auto-Regressive Model for Long-Horizon Embodied IntelligenceICML 2025 poster [paper] #minecraft#training
[2024/12] TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in MinecraftarXiv [paper][code] #minecraft#multi-agent#training#vlm
[2024/11] MrSteve: Instruction-Following Agents with What-Where-When MemoryICLR 2025 [paper][code] #minecraft#memory
[2024/10] WALL-E: World Alignment by Rule Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model
[2024/10] ADAM: An Embodied Causal Agent in Open-World EnvironmentsICLR 2025 [paper][code] #minecraft#planning
[2024/09] MrSteve: Instruction-Following Agents in Minecraft with What-Where-When MemoryICLR 2025 Poster [paper] #minecraft#memory
[2024/07] Odyssey: Empowering Agents with Open-World Skills.IJCAI 2024 [paper][code] #minecraft#planning#tool-use
[2024/07] OmniJARVIS: Omni-Modal Open-World Agents in MinecraftNeurIPS 2024 [paper][code] #minecraft#training#vlm
[2024/06] VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in MinecraftFindings of ACL 2024 [paper][code] #minecraft#planning#multi-agent
[2024/03] MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World ControlarXiv [paper][code] #minecraft
[2024/03] MineLand: Simulating Large-Scale Multi-Agent Interactions with Limited Multimodal Senses and Physical NeedsarXiv [paper][code] #minecraft#multi-agent#vlm
[2024/03] Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation (HAS)ICLR 2024 Workshop [paper] #minecraft#planning#multi-agent#vlm
[2023/12] MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active PerceptionCVPR 2024 [paper][code] #minecraft#planning#vlm
[2023/12] Auto MC-Reward: Automated Dense Reward Design with Large Language Models for MinecraftCVPR 2023 [paper] #minecraft#training
[2023/12] Creative Agents: Empowering Agents with Imagination for Creative TasksUAI 2023 [paper][code] #minecraft#planning
[2023/11] JARVIS-1: Open-world Multi-task Agents with Memory-Augmented Multimodal Language ModelsTPAMI 2023 [paper][code] #minecraft#planning#memory
[2023/11] See and Think: Embodied Agent in Virtual EnvironmentECCV 2023 [paper][code] #minecraft#memory
[2023/10] LLaMA Rider: Spurring Large Language Models to Explore the Open WorldNAACL 2023 [paper][code] #minecraft#planning#training
[2023/10] MCU: A Task-centric Framework for Open-ended Agent Evaluation in MinecraftICML 2023 [paper][code] #minecraft
[2023/10] Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open WorldsICLR 2024 [paper] #minecraft#vlm
[2023/05] Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and MemoryarXiv [paper] #minecraft#training
[2023/05] VOYAGER: An Open-Ended Embodied Agent with Large Language ModelsFMDM@NeurIPS2023 [paper][code] #minecraft#tool-use#training
[2023/03] Plan4MC: Skill Reinforcement Learning and Planning for Open-World Minecraft TasksFMDM@NeurIPS2023 [paper][code] #minecraft#planning#training
[2023/02] Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task AgentsNeurIPS 2023 [paper][code] #minecraft#planning#prompting
[2022/07] Craft an Iron Sword: Dynamically Generating Interactive Game Characters by Prompting Large Language Models Tuned on CodeWordplay@ACL 2022 [paper] #minecraft
text-adventure
[2026/06] SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent TrainingarXiv [paper][code] #text-adventure#memory#training
[2026/06] SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at ScalearXiv [paper] #text-adventure#memory#self-improvement
[2026/06] LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM AgentsarXiv [paper] #text-adventure
[2026/06] From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM AgentsarXiv [paper] #text-adventure#planning
[2026/06] Cross-Environment Neural Reranking for Sample-Efficient Action Selection in Text-Based AgentsarXiv [paper] #text-adventure#training
[2026/06] Unified Context Evolution for LLM AgentsarXiv [paper] #text-adventure
[2026/06] When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent TrainingarXiv [paper] #text-adventure#training
[2026/05] T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement LearningarXiv [paper][code] #text-adventure#training
[2026/04] Aligning Progress and Feasibility: A Neuro-Symbolic Dual Memory Framework for Long-Horizon LLM AgentsarXiv [paper] #text-adventure#planning#memory
[2026/04] Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM AgentsarXiv [paper][code] #text-adventure#training
[2026/04] DORA Explorer: Improving the Exploration Ability of LLMs Without TrainingarXiv [paper] #text-adventure#planning#prompting
[2026/04] From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM AgentsarXiv [paper] #text-adventure#planning
[2026/04] Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent WorkflowsarXiv [paper] #text-adventure#planning#memory#multi-agent
[2026/04] ReDAct: Uncertainty-Aware Deferral for LLM AgentsarXiv [paper] #text-adventure
[2026/04] GraSP: Graph-Structured Skill Compositions for LLM AgentsarXiv [paper] #text-adventure#planning#memory#self-improvement
[2026/04] DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based AgentsarXiv [paper][code] #text-adventure#training
[2026/03] Hindsight Credit Assignment for Long-Horizon LLM AgentsarXiv [paper] #text-adventure#training
[2026/03] How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game EnvironmentarXiv [paper] #text-adventure#multi-agent#training
[2026/03] RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic FeedbackarXiv [paper] #text-adventure#memory#training#self-improvement
[2026/03] Reward Prediction with Factorized World StatesarXiv [paper] #text-adventure#planning#prompting
[2026/02] MemSkill: Learning and Evolving Memory Skills for Self-Evolving AgentsarXiv [paper] #text-adventure#self-improvement
[2026/02] Active Epistemic Control for Query-Efficient Verified PlanningarXiv [paper] #text-adventure#planning
[2026/02] Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/02] TAPE: Tool-Guided Adaptive Planning and Constrained Execution in Language Model AgentsarXiv [paper] #text-adventure#planning
[2026/02] CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesarXiv [paper] #text-adventure#planning#world-model#training
[2026/02] Reinforcement World Model Learning for LLM-based AgentsarXiv [paper] #text-adventure#world-model
[2026/01] Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/01] Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient UpdatesarXiv [paper][code] #text-adventure#training
[2025/12] Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive RefinementProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#memory
[2025/12] GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment SimulatorsarXiv [paper] #text-adventure
[2025/10] Fine-tuning with RAG for Improving LLM Learning of New SkillsarXiv [paper] #text-adventure#planning#memory#training
[2025/10] GenQuest: An LLM-based Text Adventure Game for Language LearnersarXiv [paper] #text-adventure#vlm#generation
[2025/10] Constrained Natural Language Action Planning for Resilient Embodied SystemsarXiv [paper] #text-adventure#planning#prompting
[2025/10] The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with SchemasarXiv [paper] #text-adventure#planning
[2025/10] Graph-Enhanced Policy Optimization in LLM Agent TrainingarXiv [paper] #text-adventure#planning#training
[2025/10] SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory GrapharXiv [paper] #text-adventure#training
[2025/09] World Model Implanting for Test-time Adaptation of Embodied AgentsICML 2025 poster [paper] #text-adventure#memory#world-model#prompting
[2025/09] Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM AgentarXiv [paper] #text-adventure#planning#multi-agent#self-improvement
[2025/09] Code Driven Planning with Domain-Adaptive CriticarXiv [paper] #text-adventure#planning#self-improvement
[2025/09] Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationICLR 2026 Poster [paper] #text-adventure#memory#training
[2025/09] DreamPhase: Offline Imagination and Uncertainty-Guided Planning for Large-Language-Model AgentsICLR 2026 Poster [paper] #text-adventure#planning#world-model
[2025/09] Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningICLR 2026 Poster [paper] #text-adventure#tool-use#training
[2025/08] Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World SuccessProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#planning#training#vlm
[2025/07] CoEx -- Co-evolving World-model and ExplorationEMNLP 2025 [paper] #text-adventure#planning#world-model
[2025/06] Enhancing Decision-Making of Large Language Models via Actor-CriticICML 2025 poster [paper] #text-adventure#training
[2025/06] Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead SearcharXiv [paper] #text-adventure#planning#world-model#training
[2025/06] StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi TurnsarXiv [paper] #text-adventure#memory
[2025/06] OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic ReflectionsarXiv [paper] #text-adventure#planning#training#self-improvement
[2025/06] KnowMap: Efficient Knowledge-Driven Task Adaptation for LLMsarXiv [paper] #text-adventure#training
[2025/05] STORY2GAME: Generating (Almost) Everything in an Interactive Fiction GamearXiv [paper] #text-adventure#generation
[2025/05] LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMsarXiv [paper][code] #text-adventure#planning
[2025/05] Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningICML 2025 poster [paper] #text-adventure#planning#training
[2025/05] ActiveVOO: Value of Observation Guided Active Knowledge Acquisition for Open-World Embodied Lifted Regression PlanningNeurIPS 2025 poster [paper] #text-adventure#planning#prompting#vlm
[2025/03] Haunted House: A text-based game for comparing the flexibility of mental models in humans and LLMsarXiv [paper] #text-adventure
[2025/03] GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow NetworksCVPR 2025 [paper] #text-adventure#planning#training#vlm
[2025/02] TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning.arXiv [paper] #text-adventure#self-improvement
[2025/02] Process Reward Models for LLM Agents: Practical Framework and DirectionsarXiv [paper][code] #text-adventure#training
[2024/12] Fine-tuning large vision-language models as decision-making agents via reinforcement learningNeurIPS 2024 [paper][code] #text-adventure#training#vlm
[2024/09] Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackICLR 2025 Poster [paper] #text-adventure#planning#training#self-improvement
[2024/07] AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding AgentsACL 2024 [paper][code] #text-adventure#tool-use
[2024/07] Arigraph: Learning knowledge graph world models with episodic memory for llm agentsIJCAI 2024 [paper] #text-adventure#planning#memory
[2024/06] Watch Every Step! LLM Agent Learning via Iterative Step-Level Process RefinementEMNLP 2024 [paper][code] #text-adventure#self-improvement
[2024/06] STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language ModelsACL 2024 [paper][code] #text-adventure#training
[2024/05] Agent Planning with World Knowledge ModelNeurIPS 2024 [paper][code] #text-adventure#planning#memory#world-model
[2024/05] THREAD: Thinking Deeper with Recursive SpawningNAACL 2024 [paper] #text-adventure#prompting
[2024/05] Policy Improvement using Language Feedback ModelsNeurIPS 2024 poster [paper] #text-adventure#training
[2024/05] AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental LearningNeurIPS 2024 poster [paper][code] #text-adventure#planning
[2024/04] Learning From Failure: Integrating Negative Examples When Fine-tuning Large Language Models as AgentarXiv [paper][code] #text-adventure#tool-use#training
[2024/04] ReAct Meets ActRe: When Language Agents Enjoy Training Data AutonomyarXiv [paper] #text-adventure#planning#training#self-improvement
[2024/03] KnowAgent: Knowledge-Augmented Planning for LLM-Based AgentsNAACL 2024 [paper][code] #text-adventure#planning
[2024/03] Language Guided Exploration for RL Agents in Text EnvironmentsNAACL 2024 [paper][code] #text-adventure#training
[2024/03] Trial and Error: Exploration-Based Trajectory Optimization for LLM AgentsACL 2024 [paper][code] #text-adventure#training#self-improvement
[2024/03] O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language ModelsCOLM [paper] #text-adventure#training
[2024/03] ReAct Meets ActRe: Autonomous Annotation of Agent Trajectories for Contrastive Self-TrainingCOLM [paper] #text-adventure#planning#training#self-improvement
[2024/03] StateFlow: Enhancing LLM Task-Solving through State-Driven WorkflowsCOLM [paper] #text-adventure#planning#self-improvement
[2024/02] Soft Self-Consistency Improves Language Model AgentsarXiv [paper][code] #text-adventure#self-improvement
[2024/02] Empowering Large Language Model Agents through Action LearningCOLM [paper][code] #text-adventure#planning#self-improvement
[2023/11] ADaPT: As-Needed Decomposition and Planning with Language ModelsNAACL 2023 [paper][code] #text-adventure#planning
[2023/10] FireAct: Toward Language Agent Fine-tuningarXiv [paper][code] #text-adventure#training
[2023/10] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language ModelsICML 2024 [paper][code] #text-adventure#planning#training
[2023/05] SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksNeurIPS 2023 [paper][code] #text-adventure#planning
[2023/04] Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open QuestionsarXiv [paper] #text-adventure#world-model
[2023/03] Reflexion: Language Agents with Verbal Reinforcement LearningNeurIPS 2023 [paper][code] #text-adventure#training#self-improvement
[2022/10] ReAct: Synergizing Reasoning and Acting in Language ModelsICLR 2023 [paper][code] #text-adventure#planning#training
[2022/03] ScienceWorld: Is your Agent Smarter than a 5th Grader?EMNLP 2022 [paper][code] #text-adventure
[2020/10] ALFWorld: Aligning Text and Embodied Environments for Interactive LearningICLR 2021 [paper][code] #text-adventure#planning
[2019/09] Interactive Fiction Games: A Colossal AdventureAAAI 2020 [paper][code] #text-adventure
communication
[2026/05] Evaluating Large Language Models in a Complex Hidden Role GamearXiv [paper] #communication#planning
[2026/05] QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction AgentsarXiv [paper][code] #communication
[2026/05] MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMsarXiv [paper] #communication#multi-agent
[2026/05] Playing with Words, Improving with Rewards: Training Language Models for Creative AssociationarXiv [paper] #communication#training
[2026/04] Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM AgentsarXiv [paper] #communication
[2026/03] Enhancing Consistency of Werewolf AI through Dialogue Summarization and Persona InformationarXiv [paper] #communication#role-play
[2026/03] Deception and Communication in Autonomous Multi-Agent Systems: An Experimental Study with Among UsProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #communication#multi-agent
[2026/01] Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction GamesInternational Conference on Agents 2027 [paper] #communication#multi-agent
[2026/01] Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction GamearXiv [paper] #communication
[2025/12] WOLF: Werewolf-based Observations for LLM Deception and FalsehoodsarXiv [paper] #communication#multi-agent
[2025/12] Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in DiplomacyarXiv [paper] #communication#training
[2025/11] CSP4SDG: Constraint and Information-Theory Based Role Identification in Social Deduction Games with LLM-Enhanced InferenceAAAI 2025 [paper] #communication
[2025/11] Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal ReasoningarXiv [paper][code] #communication#multi-agent
[2025/10] Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned StrategiesarXiv [paper] #communication#multi-agent#self-improvement
[2025/08] Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press DiplomacyAAAI 2025 [paper] #communication#training
[2025/08] What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup PuzzlesAAAI 2025 [paper] #communication
[2025/08] Ethical Considerations of Large Language Models in Game PlayingarXiv [paper] #communication
[2025/07] CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language GamesACL 2025 [paper][code] #communication#multi-agent
[2025/07] Strategy Adaptation in Large Language Model Werewolf AgentsarXiv [paper] #communication#prompting
[2025/06] WereWolf-Plus: An Update of Werewolf Game setting Based on DSGBencharXiv [paper][code] #communication#multi-agent
[2025/06] DipLLM: Fine-Tuning LLM for Strategic Decision-making in DiplomacyICML 2025 poster [paper] #communication#training
[2025/05] Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RLNeurIPS 2025 poster [paper] #communication#planning#tool-use#training
[2025/01] DVM: Towards Controllable LLM Agents in Social Deduction GamesIEEE International Conference on Acoustics, Speech, and Signal Processing 2025 [paper] #communication#training
[2024/12] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 [paper][code] #communication#self-improvement
[2024/09] Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchICLR 2025 Poster [paper] #communication#planning#multi-agent#training
[2024/06] PLAYER: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery GamesarXiv [paper] #communication#multi-agent
[2024/05] Learning to Discuss Strategically: A Case Study on One Night Ultimate WerewolfNeurIPS 2024 poster [paper] #communication#training
[2024/05] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 poster [paper] #communication#planning#multi-agent#self-improvement
[2024/04] Self-playing Adversarial Language Game Enhances LLM ReasoningNeurIPS 2024 [paper][code] #communication#training#self-improvement
[2024/03] Helmsman of the Masses? Evaluate the Opinion Leadership of Large Language Models in the Werewolf GameCOLM [paper] #communication#multi-agent
[2024/02] Enhance Reasoning for Large Language Models in the Game WerewolfarXiv [paper] #communication#training
[2024/02] What if LLMs Have Different World Views: Simulating Alien Civilizations with LLM-based AgentsarXiv [paper] #communication
[2024/02] Can Large Language Model Agents Simulate Human Trust Behaviors?NeurIPS 2024 [paper] #communication
[2024/02] Large Language Models Fall Short: Understanding Complex Relationships in Detective NarrativesACL 2024 [paper] #communication
[2023/12] Can Large Language Models Serve as Rational Players in Game Theory? A Systematic AnalysisAAAI 2024 [paper] #communication
[2023/12] Cooperation on the Fly: Exploring Language Agents for Ad Hoc Teamwork in the Avalon GamearXiv [paper] #communication#multi-agent
[2023/12] Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery GamesACL 2023 [paper] #communication#multi-agent#prompting
[2023/11] War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World WarsarXiv [paper][code] #communication#multi-agent
[2023/11] clembench: Systematic Evaluation of Chat-Optimized Language Models as Conversational AgentsEMNLP 2023 [paper] #communication
[2023/10] Language Agents with Reinforcement Learning for Strategic Play in the Werewolf GameICML 2023 [paper] #communication#training
[2023/10] Avalon's Game of Thoughts: Battle Against Deception through Recursive ContemplationarXiv [paper] #communication#training
[2023/10] LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon GameplayEMNLP 2023 [paper] #communication#multi-agent
[2023/10] Leveraging Word Guessing Games to Assess the Intelligence of Large Language ModelsarXiv [paper][code] #communication#multi-agent
[2023/10] AvalonBench: Evaluating LLMs Playing the Game of AvalonFMDM@NeurIPS2023 [paper][code] #communication
[2023/09] Exploring Large Language Models for Communication Games: An Empirical Study on WerewolfarXiv [paper] #communication#memory
[2023/08] GameEval: Evaluating LLMs on Conversational GamesarXiv [paper][code] #communication
[2022/12] Human-Level Play in the Game of Diplomacy by Combining Language Models with Strategic ReasoningScience [paper] #communication
competition
[2026/06] Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM AgentsarXiv [paper] #competition
[2026/06] SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language ModelsarXiv [paper] #competition#multi-agent
[2026/05] GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM CollectivesarXiv [paper] #competition#multi-agent#prompting
[2026/05] Watermarking Game-Playing Agents in Perfect-Information Extensive-Form GamesarXiv [paper] #competition
[2026/05] Generalization or Memorization? Brittleness Testing for Chess-Trained Language ModelsarXiv [paper] #competition#training
[2026/05] PokerSkill: LLMs Can Play Expert-Level Poker without Training or SolversarXiv [paper][code] #competition#tool-use#prompting
[2026/04] Readable Minds: Emergent Theory-of-Mind-Like Behavior in LLM Poker AgentsarXiv [paper] #competition
[2026/04] MARL-GPT: Foundation Model for Multi-Agent Reinforcement LearningProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #competition#multi-agent#training
[2026/03] Grounded Chess Reasoning in Language Models via Master DistillationarXiv [paper] #competition#planning#training
[2026/03] Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy ScenariosarXiv [paper] #competition#planning#memory#multi-agent
[2026/02] World Models for Policy Refinement in StarCraft IIarXiv [paper] #competition#world-model#prompting
[2026/02] VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case StudyarXiv [paper] #competition#training
[2025/12] LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through ChessarXiv [paper] #competition
[2025/12] Beyond Accuracy: A Geometric Stability Analysis of Large Language Models in Chess EvaluationarXiv [paper] #competition
[2025/10] Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy GamesarXiv [paper] #competition#memory
[2025/10] ChessQA: Evaluating Large Language Models for Chess UnderstandingarXiv [paper] #competition
[2025/10] Out-of-distribution Tests Reveal Compositionality in Chess TransformersarXiv [paper] #competition#planning
[2025/09] HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-MakingarXiv [paper] #competition#multi-agent#training
[2025/09] Speculative Actions: A Lossless Framework for Faster AI AgentsICLR 2026 Oral [paper] #competition#tool-use
[2025/09] SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningICLR 2026 Poster [paper] #competition#planning#multi-agent#training
[2025/08] Tracking World States with Language Models: State-Based Evaluation Using ChessarXiv [paper] #competition
[2025/08] SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making TasksarXiv [paper] #competition#planning#training#self-improvement
[2025/07] Learning to Imitate with Less: Efficient Individual Behavior Modeling in ChessarXiv [paper] #competition
[2025/05] Enfoque Odychess: Un método dialéctico, constructivista y adaptativo para la enseñanza del ajedrez con inteligencias artificiales generativasarXiv [paper] #competition#training
[2025/05] Can Large Language Models Master Complex Card Games?NeurIPS 2025 poster [paper][code] #competition#training
[2025/04] Explore the Reasoning Capability of LLMs in the Chess TestbedNAACL 2025 [paper] #competition
[2025/04] ZeroSumEval: Scaling LLM Evaluation with Inter-Model CompetitionarXiv [paper][code] #competition#planning
[2025/04] LLM-PySC2: Starcraft II learning environment for Large Language ModelsNeurIPS 2025 poster [paper] #competition#planning#multi-agent#vlm
[2025/04] The PokeAgent Challenge: Competitive and Long Context Learning at ScaleNeurIPS Competition Track 2025 [paper] #competition
[2025/03] Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic ReasoningCOLM 2025 [paper] #competition#multi-agent#training
[2025/02] Hierarchical Expert Prompt for Large-Language-Model: An Approach Defeat Elite AI in TextStarCraft II for the First TimearXiv [paper][code] #competition
[2025/02] Implicit Search via Discrete Diffusion: A Study on ChessICLR 2025 Poster [paper][code] #competition#planning
[2025/01] POKERBENCH: Training Large Language Models to become Professional Poker PlayersAAAI 2025 [paper] #competition#planning#training
[2025/01] Complete Chess Games Enable LLM Become A Chess MasterNAACL 2025 [paper] #competition#training
[2025/01] Mastering Board Games by External and Internal Planning with Language ModelsICML 2025 spotlightposter [paper] #competition#planning
[2025/01] Language Models as Implicit Tree SearchICML 2025 poster [paper] #competition#planning#training
[2024/10] PokéChamp: An Expert-level Minimax Language AgentICML 2025 [paper][code] #competition#planning
[2024/08] Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) 2024 [paper] #competition#planning#multi-agent#training
[2024/05] Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game ModelsNeurIPS 2024 poster [paper] #competition
[2024/05] Reflective Multi-Agent Collaboration based on Large Language ModelsNeurIPS 2024 poster [paper] #competition#planning#multi-agent#training
[2024/03] Embodied LLM Agents Learn to Cooperate in Organized TeamsIEEE Transactions on Computational Social Systems 2024 [paper] #competition#planning#multi-agent
[2024/02] PokéLLMon: A Human-Parity Agent for Pokémon Battles with Large Language ModelsTOIT 2025 [paper][code] #competition#training
[2024/02] Agent-Pro: Learning to Evolve via Policy-Level Reflection and OptimizationACL 2024 [paper][code] #competition#training
[2024/01] PokerGPT: An End-to-End Lightweight Solver for Multi-Player Texas Hold'em via Large Language ModelarXiv [paper] #competition#training
[2024/01] SwarmBrain: Embodied agent for real-time strategy game StarCraft II via large language modelsarXiv [paper] #competition#training
[2023/12] Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization ApproachNeurIPS 2024 poster [paper][code] #competition#planning
[2023/09] Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4COLM 2024 [paper] #competition#planning#memory#prompting
[2023/08] Are ChatGPT and GPT-4 Good Poker Players?--A Pre-Flop AnalysisarXiv [paper] #competition
[2023/06] ChessGPT: Bridging Policy Learning and Language ModelingNeurIPS 2023 [paper][code] #competition
[2022/10] Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskICLR 2023 [paper] #competition
cooperation
[2026/04] Don't Make the LLM Read the Graph: Make the Graph ThinkarXiv [paper] #cooperation#multi-agent
[2026/03] On the Strengths and Weaknesses of Data for Open-set Embodied AssistancearXiv [paper] #cooperation#training
[2025/10] LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration GamearXiv [paper] #cooperation#multi-agent
[2025/06] PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language ModelsarXiv [paper] #cooperation#planning#multi-agent
[2024/05] Towards Efficient LLM Grounding for Embodied Multi-Agent CollaborationACL 2024 [paper][code] #cooperation#planning#multi-agent#training
[2024/03] Can LLM-Augmented Autonomous Agents Cooperate?, An Evaluation of Their Cooperative Capabilities through Melting PotIEEE Transactions on Artificial Intelligence 2024 [paper] #cooperation#multi-agent
[2024/03] ProAgent: Building Proactive Cooperative Agents with Large Language ModelsAAAI 2024 [paper] #cooperation
[2024/02] S-Agents: Self-organizing Agents in Open-ended EnvironmentsarXiv [paper] #cooperation
[2023/12] LLM-Powered Hierarchical Language Agent for Real-time Human-AI CoordinationAAMAS 2023 [paper] #cooperation
[2023/10] Evaluating Multi-agent Coordination Abilities in Large Language ModelsNAACL 2023 [paper] #cooperation#planning#multi-agent#prompting
[2023/10] Theory of Mind for Multi-Agent Collaboration via Large Language ModelsEMNLP 2023 [paper][code] #cooperation#planning#multi-agent#training
[2026/04] LLM-Agent-based Social Simulation for Attitude DiffusionarXiv [paper] #sim-social#memory
[2026/04] Restoring Heterogeneity in LLM-based Social Simulation: An Audience Segmentation ApproacharXiv [paper] #sim-social#role-play
[2026/04] SLALOM: Simulation Lifecycle Analysis via Longitudinal Observation Metrics for Social SimulationarXiv [paper] #sim-social
[2026/04] RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing AgentsarXiv [paper] #sim-social#planning
[2026/04] Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing AgentsarXiv [paper] #sim-social
[2026/04] Auditing Support Strategies in LLMs through Grounded Multi-Turn Social SimulationarXiv [paper] #sim-social
[2026/03] PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy OptimizationWWW 2026 [paper] #sim-social#training
[2026/03] Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationWWW 2026 [paper][code] #sim-social#multi-agent
[2026/02] AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent ProfilesarXiv [paper] #sim-social#planning
[2026/02] Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent CommunityarXiv [paper] #sim-social
[2026/02] Does Socialization Emerge in AI Agent Society? A Case Study of MoltbookProceedings of the ACM Conference on AI and Agentic Systems 2026 [paper] #sim-social
[2026/02] The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI SocietiesarXiv [paper] #sim-social#multi-agent#self-improvement
[2026/01] When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered AgentsarXiv [paper] #sim-social#multi-agent
[2026/01] MARO: Learning Stronger Reasoning from Social InteractionarXiv [paper] #sim-social#multi-agent
[2026/01] HumanLLM: Towards Personalized Understanding and Simulation of Human NaturearXiv [paper] #sim-social#training
[2025/12] EZYer: A simulacrum of high school with generative agentarXiv [paper] #sim-social#memory
[2025/12] Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMsarXiv [paper] #sim-social#multi-agent
[2025/10] Multimodal Safety Evaluation in Generative Agent Social SimulationsarXiv [paper] #sim-social#planning#vlm
[2025/10] Alita-G: Self-Evolving Generative Agent for Agent GenerationarXiv [paper] #sim-social#memory#self-improvement
[2025/10] Doing Things with Words: Rethinking Theory of Mind Simulation in Large Language ModelsComputational Linguistics 2025 [paper] #sim-social
[2025/10] Social Simulations with Large Language Model Risk Utopian IllusionarXiv [paper] #sim-social#multi-agent
[2025/10] Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information OperationsWWW 2025 [paper] #sim-social
[2025/09] Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd SimulationsEMNLP 2025 [paper] #sim-social#role-play
[2025/09] The Emergence of Altruism in Large-Language-Model Agents SocietyarXiv [paper] #sim-social
[2025/07] LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative SimulacraarXiv [paper][code] #sim-social#multi-agent#training#role-play
[2025/07] Too Human to Model:The Uncanny Valley of LLMs in Social Simulation -- When Generative Language Agents Misalign with Modelling PrinciplesarXiv [paper] #sim-social
[2025/07] Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel PredictionsAnnual Meeting of the Cognitive Science Society 2025 [paper] #sim-social#role-play
[2025/06] Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based ModelingKDD 2025 [paper] #sim-social#training
[2025/06] IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent EnvironmentEMNLP 2025 [paper] #sim-social#multi-agent
[2025/06] AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System NeedarXiv [paper][code] #sim-social#planning#multi-agent
[2025/06] Infected Smallville: How Disease Threat Shapes Sociality in LLM AgentsarXiv [paper] #sim-social
[2025/05] EcoLANG: Efficient and Effective Agent Communication Language Induction for Social SimulationEMNLP 2025 [paper] #sim-social#role-play
[2025/04] SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulationNAACL 2025 [paper] #sim-social#planning
[2025/04] BookWorld: From Novels to Interactive Agent Societies for Creative Story GenerationarXiv [paper] #sim-social#multi-agent#generation
[2025/04] MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model FrameworkNeurIPS 2025 poster [paper] #sim-social#planning#training
[2025/03] The Impact of Big Five Personality Traits on AI Agent Decision-Making in Public Spaces: A Social Simulation StudyarXiv [paper] #sim-social
[2025/02] Investigating and Extending Homans' Social Exchange Theory with Large Language Model based AgentsACL 2025 [paper][code] #sim-social
[2025/01] Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical InteractionsarXiv [paper] #sim-social
[2024/10] Project Sid: Many-agent simulations toward AI civilizationarXiv [paper] #sim-social
[2024/06] Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract TheoryarXiv [paper] #sim-social#multi-agent
[2024/05] Agent hospital: A simulacrum of hospital with evolvable medical agentsarXiv [paper] #sim-social#planning
[2024/03] SOTOPIA-$\pi$: Interactive Learning of Socially Intelligent Language AgentsACL 2024 [paper][code] #sim-social#training
[2023/10] Lyfe Agents: Generative agents for low-cost real-time social interactionsarXiv [paper] #sim-social#memory#multi-agent#self-improvement
[2023/10] SOTOPIA: Interactive Evaluation for Social Intelligence in Language AgentsICLR 2023 [paper][code] #sim-social#role-play
[2023/08] AgentSims: An Open-Source Sandbox for Large Language Model EvaluationarXiv [paper][code] #sim-social#planning#tool-use
[2023/07] S3: Social-network Simulation System with Large Language Model-Empowered AgentsarXiv [paper] #sim-social#prompting
[2023/04] Generative Agents: Interactive Simulacra of Human BehaviorUIST 2023 [paper][code] #sim-social
sim-embodied
[2026/05] Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State UncertaintyarXiv [paper] #sim-embodied#training
[2026/05] Probing Embodied LLMs: When Higher Observation Fidelity Hurts Problem SolvingarXiv [paper] #sim-embodied
[2026/02] To Move or Not to Move: Constraint-based Planning Enables Zero-Shot Generalization for Interactive NavigationarXiv [paper] #sim-embodied#planning#prompting#vlm
[2025/12] Emergence: Overcoming Privileged Information Bias in Asymmetric Embodied Agents via Active QueryingarXiv [paper] #sim-embodied
[2025/12] ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement LearningarXiv [paper] #sim-embodied#planning#memory#training
[2025/12] HELP: Hierarchical Embodied Language Planner for Household TasksarXiv [paper] #sim-embodied#planning
[2025/11] DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent CollaborationarXiv [paper] #sim-embodied#planning#multi-agent#world-model
[2025/11] MADRA: Multi-Agent Debate for Risk-Aware Embodied PlanningarXiv [paper] #sim-embodied#planning#multi-agent#prompting
[2024/09] Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-Based Decision-Making SystemsICLR 2025 Poster [paper] #sim-embodied#training
[2024/09] BadRobot: Jailbreaking Embodied LLM Agents in the Physical WorldICLR 2025 Poster [paper] #sim-embodied#planning#vlm
[2024/09] GameGen-X: Interactive Open-world Game Video GenerationICLR 2025 Poster [paper][code] #sim-embodied#vlm
[2024/01] True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement LearningarXiv [paper][code] #sim-embodied#training
[2025/06] Automated Skill Discovery for Language Agents through Exploration and Iterative FeedbackarXiv [paper] #crafter#self-improvement
[2025/02] LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative PlanningarXiv [paper] #crafter#planning#memory#multi-agent
[2024/10] Mars: Situated Inductive Reasoning in an Open-World EnvironmentNeurIPS 2024 [paper] #crafter
[2024/07] Enhancing Agent Learning through World Dynamics ModelingEMNLP 2024 [paper] #crafter
[2024/04] AgentKit: Flow Engineering with Graphs, not CodingarXiv [paper][code] #crafter#planning
[2024/04] World Models with Hints of Large Language Models for Goal AchievingNAACL 2024 [paper] #crafter#training#vlm
[2024/03] EnvGen: Generating and Adapting Environments via LLMs for Training Embodied AgentsCOLM [paper] #crafter#training
[2024/03] AgentKit: Structured LLM Reasoning with Dynamic GraphsCOLM [paper] #crafter#planning
[2023/09] AdaRefiner: Refining Decisions of Language Models with Adaptive FeedbackNAACL 2023 [paper] #crafter#training
[2023/06] OMNI: Open-endedness via Models of human Notions of InterestingnessarXiv [paper][code] #crafter
[2023/05] SPRING: Studying Papers and Reasoning to play GamesNeurIPS 2023 [paper] #crafter
[2023/02] Guiding Pretraining in Reinforcement Learning with Large Language ModelsICML 2023 [paper] #crafter#training
action
[2026/05] Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement LearningarXiv [paper] #action#training#vlm
[2026/05] ANO: A Principled Approach to Robust Policy OptimizationarXiv [paper] #action#training
[2026/05] Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplayarXiv [paper] #action#planning#training#vlm
[2026/04] Playing DOOM with 1.3M Parameters: Specialized Small Models vs Large Language Models for Real-Time Game ControlarXiv [paper] #action
[2026/04] PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language ModelsarXiv [paper] #action#planning#vlm
[2025/05] Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme OnearXiv [paper] #action#training
[2025/05] Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsNeurIPS 2025 poster [paper] #action#training
[2025/05] LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language ModelsNeurIPS 2025 spotlight [paper][code] #action#training
[2025/05] PoE-World: Compositional World Modeling with Products of Programmatic ExpertsNeurIPS 2025 spotlight [paper] #action#planning#world-model
[2025/04] Better Decisions through the Right Causal World ModelarXiv [paper] #action#planning#world-model#training
[2025/01] LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal DemonstrationsICML 2025 poster [paper] #action#planning#training
[2024/10] Unbounded: A Generative Infinite Game of Character Life SimulationICLR 2024 [paper] #action
[2024/09] Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study CasearXiv [paper][code] #action#planning#training
[2024/09] BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesICLR 2025 Poster [paper] #action#planning#training#vlm
[2024/08] Atari-GPT: Investigating the Capabilities of Multimodal Large Language Models as Low-Level Policies for Atari GamesarXiv [paper] #action#planning#training
[2024/07] Baba Is AI: Break the Rules to Beat the BenchmarkICML 2024 [paper] #action#vlm
[2024/03] Will GPT-4 Run DOOM?IEEE Transactions on Games 2024 [paper][code] #action#planning#training
[2024/03] Evaluate LLMs in Real Time with Street Fighter IIIGitHub [paper][code] #action
[2023/02] Grounding Large Language Models in Interactive Environments with Online Reinforcement LearningICML 2023 [paper][code] #action#training
video-adventure
[2024/03] Cradle: Empowering Foundation Agents Towards General Computer ControlICML 2024 [paper][code] #video-adventure#planning
[2024/03] Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents2024 IEEE Conference on Games (CoG) 2024 [paper][code] #video-adventure#planning#prompting
[2025/05] lmgame-Bench: How Good are LLMs at Playing Games?."ICLR 2026 Poster [paper][code] #benchmark#planning#training
[2025/05] Is Your LLM Really Mastering the Concept? A Multi-Agent BenchmarkarXiv [paper][code] #benchmark#multi-agent
other
[2026/04] GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game AgentsarXiv [paper] #planning#vlm
[2026/04] Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language ModelsarXiv [paper] #tool-use#training#self-improvement
[2026/04] From World-Gen to Quest-Line: A Dependency-Driven Prompt Pipeline for Coherent RPG GenerationarXiv [paper] #planning#generation
[2026/03] Sensi: Learn One Thing at a Time -- Curriculum-Based Test-Time Learning for LLM Game AgentsarXiv [paper]
[2026/01] SNAP: A Plan-Driven Framework for Controllable Interactive Narrative GenerationarXiv [paper] #planning#generation
[2025/10] ROBOPSY PL[AI]: Using Role-Play to Investigate how LLMs Present Collective MemoryarXiv [paper] #role-play
[2025/08] All Stories Are One Story: Emotional Arc Guided Procedural Game Level GenerationarXiv [paper] #generation
[2025/08] CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMsarXiv [paper] #memory
[2025/06] The Decrypto Benchmark for Multi-Agent Reasoning and Theory of MindarXiv [paper] #multi-agent#training
[2025/04] PAYADOR: A Minimalist Approach to Grounding Language Models on Structured Data for Interactive Storytelling and Role-playing GamesarXiv [paper]
[2025/04] Enhancing Player Enjoyment with a Two-Tier DRL and LLM-Based Agent System for Fighting GamesarXiv [paper] #training
[2025/03] Playing games with Large language models: Randomness and strategyarXiv [paper] #multi-agent
[2025/03] Collaborative Storytelling and LLM: A Linguistic Analysis of Automatically-Generated Role-Playing Game SessionsarXiv [paper]
[2025/03] Cultivating Game Sense for Yourself: Making VLMs Gaming ExpertsarXiv [paper] #vlm
[2025/02] RPGBENCH: Evaluating Large Language Models as Role-Playing Game EnginesarXiv [paper]
[2025/02] Hybrid Voting-Based Task Assignment in Role-Playing GamesarXiv [paper] #planning
[2024/07] What if Red Can Talk? Dynamic Dialogue Generation Using Large Language Models.arXiv [paper] #generation
[2023/10] Language as reality: a co-creative storytelling game experience in 1001 nights using generative AI.AAAI 2023 [paper] #generation
By Mechanism
The same papers as above, re-grouped by agent design. A paper with multiple Mechanism tags appears in each relevant section.
planning
[2026/06] Can LLM Agents Sustain Long-Horizon Organizational Dynamics?arXiv [paper] #sim-social#planning#multi-agent
[2026/06] From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM AgentsarXiv [paper] #text-adventure#planning
[2026/05] From History to State: Constant-Context Skill Learning for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/05] PriorZero: Bridging Language Priors and World Models for Decision MakingarXiv [paper][code] #text-adventure#planning#world-model#training
[2026/05] Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplayarXiv [paper] #action#planning#training#vlm
[2026/03] Reward Prediction with Factorized World StatesarXiv [paper] #text-adventure#planning#prompting
[2026/03] Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy ScenariosarXiv [paper] #competition#planning#memory#multi-agent
[2026/02] Active Epistemic Control for Query-Efficient Verified PlanningarXiv [paper] #text-adventure#planning
[2026/02] AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent ProfilesarXiv [paper] #sim-social#planning
[2026/02] Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/02] TAPE: Tool-Guided Adaptive Planning and Constrained Execution in Language Model AgentsarXiv [paper] #text-adventure#planning
[2026/02] To Move or Not to Move: Constraint-based Planning Enables Zero-Shot Generalization for Interactive NavigationarXiv [paper] #sim-embodied#planning#prompting#vlm
[2026/02] CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesarXiv [paper] #text-adventure#planning#world-model#training
[2026/01] SNAP: A Plan-Driven Framework for Controllable Interactive Narrative GenerationarXiv [paper] #planning#generation
[2026/01] Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2025/12] ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement LearningarXiv [paper] #sim-embodied#planning#memory#training
[2025/12] HELP: Hierarchical Embodied Language Planner for Household TasksarXiv [paper] #sim-embodied#planning
[2025/11] DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent CollaborationarXiv [paper] #sim-embodied#planning#multi-agent#world-model
[2025/11] MADRA: Multi-Agent Debate for Risk-Aware Embodied PlanningarXiv [paper] #sim-embodied#planning#multi-agent#prompting
[2025/10] Fine-tuning with RAG for Improving LLM Learning of New SkillsarXiv [paper] #text-adventure#planning#memory#training
[2025/10] Constrained Natural Language Action Planning for Resilient Embodied SystemsarXiv [paper] #text-adventure#planning#prompting
[2025/10] The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with SchemasarXiv [paper] #text-adventure#planning
[2025/10] Multimodal Safety Evaluation in Generative Agent Social SimulationsarXiv [paper] #sim-social#planning#vlm
[2025/10] Graph-Enhanced Policy Optimization in LLM Agent TrainingarXiv [paper] #text-adventure#planning#training
[2025/10] Out-of-distribution Tests Reveal Compositionality in Chess TransformersarXiv [paper] #competition#planning
[2025/09] Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM AgentarXiv [paper] #text-adventure#planning#multi-agent#self-improvement
[2025/09] Code Driven Planning with Domain-Adaptive CriticarXiv [paper] #text-adventure#planning#self-improvement
[2025/09] Goal-Guided Efficient Exploration via Large Language Model in Reinforcement LearningarXiv [paper] #crafter#planning#training
[2025/09] Natural Language PDDL (NL-PDDL) for Open-world Goal-oriented Commonsense Regression Planning in Embodied AIICLR 2026 Poster [paper] #text-adventure#planning#vlm
[2025/09] Experience-based Knowledge Correction for Robust Planning in MinecraftICLR 2026 Poster [paper] #minecraft#planning
[2025/09] SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningICLR 2026 Poster [paper] #competition#planning#multi-agent#training
[2025/09] DreamPhase: Offline Imagination and Uncertainty-Guided Planning for Large-Language-Model AgentsICLR 2026 Poster [paper] #text-adventure#planning#world-model
[2025/08] CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative TasksFindings of EMNLP 2025 [paper] #minecraft#planning#multi-agent
[2025/08] Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World SuccessProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#planning#training#vlm
[2025/08] SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making TasksarXiv [paper] #competition#planning#training#self-improvement
[2025/07] CoEx -- Co-evolving World-model and ExplorationEMNLP 2025 [paper] #text-adventure#planning#world-model
[2025/06] Optimus-3: Towards Generalist Multimodal Minecraft Agents with Scalable Task ExpertsarXiv [paper][code] #minecraft#planning
[2025/06] Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead SearcharXiv [paper] #text-adventure#planning#world-model#training
[2025/06] OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic ReflectionsarXiv [paper] #text-adventure#planning#training#self-improvement
[2025/06] AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System NeedarXiv [paper][code] #sim-social#planning#multi-agent
[2025/06] PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language ModelsarXiv [paper] #cooperation#planning#multi-agent
[2025/05] Don’t Just Follow MLLM Plans: Robust and Efficient Planning for Open-World AgentsarXiv [paper] #minecraft#planning#vlm
[2025/05] Knowledge Retrieval in LLM Gaming: A Shift from Entity-Centric to Goal-Oriented GraphsKnowledge-Based Systems 2025 [paper] #minecraft#planning#memory
[2025/05] lmgame-Bench: How Good are LLMs at Playing Games?."ICLR 2026 Poster [paper][code] #benchmark#planning#training
[2025/05] LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMsarXiv [paper][code] #text-adventure#planning
[2025/05] Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningICML 2025 poster [paper] #text-adventure#planning#training
[2025/05] ActiveVOO: Value of Observation Guided Active Knowledge Acquisition for Open-World Embodied Lifted Regression PlanningNeurIPS 2025 poster [paper] #text-adventure#planning#prompting#vlm
[2025/05] Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RLNeurIPS 2025 poster [paper] #communication#planning#tool-use#training
[2025/05] PoE-World: Compositional World Modeling with Products of Programmatic ExpertsNeurIPS 2025 spotlight [paper] #action#planning#world-model
[2025/05] WALL-E: World Alignment by NeuroSymbolic Learning improves World Model-based LLM AgentsNeurIPS 2025 poster [paper] #minecraft#planning#world-model#prompting
[2025/04] WALL-E 2.0: World Alignment by NeuroSymbolic Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model#prompting
[2025/04] Better Decisions through the Right Causal World ModelarXiv [paper] #action#planning#world-model#training
[2025/04] ZeroSumEval: Scaling LLM Evaluation with Inter-Model CompetitionarXiv [paper][code] #competition#planning
[2025/04] SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulationNAACL 2025 [paper] #sim-social#planning
[2025/04] Monte Carlo Planning with Large Language Model for Text-Based Game AgentsICLR 2025 Poster [paper] #text-adventure#planning#training
[2025/04] LLM-PySC2: Starcraft II learning environment for Large Language ModelsNeurIPS 2025 poster [paper] #competition#planning#multi-agent#vlm
[2025/04] MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model FrameworkNeurIPS 2025 poster [paper] #sim-social#planning#training
[2025/03] Parallelized Planning-Acting for Efficient LLM-based Multi-Agent SystemsarXiv [paper] #minecraft#planning#multi-agent#tool-use
[2025/03] GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow NetworksCVPR 2025 [paper] #text-adventure#planning#training#vlm
[2025/03] Plancraft: an evaluation dataset for planning with LLM agentsCOLM 2025 [paper] #minecraft#planning#memory#tool-use
[2025/02] Optimus-2: Multimodal World Model for Open-World Minecraft AgentsCVPR 2025 [paper][code] #minecraft#planning#world-model#vlm
[2025/02] LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative PlanningarXiv [paper] #crafter#planning#memory#multi-agent
[2025/02] Hybrid Voting-Based Task Assignment in Role-Playing GamesarXiv [paper] #planning
[2025/02] Implicit Search via Discrete Diffusion: A Study on ChessICLR 2025 Poster [paper][code] #competition#planning
[2025/01] POKERBENCH: Training Large Language Models to become Professional Poker PlayersAAAI 2025 [paper] #competition#planning#training
[2025/01] Mastering Board Games by External and Internal Planning with Language ModelsICML 2025 spotlightposter [paper] #competition#planning
[2025/01] LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal DemonstrationsICML 2025 poster [paper] #action#planning#training
[2025/01] Language Models as Implicit Tree SearchICML 2025 poster [paper] #competition#planning#training
[2024/10] WALL-E: World Alignment by Rule Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model
[2024/10] ADAM: An Embodied Causal Agent in Open-World EnvironmentsICLR 2025 [paper][code] #minecraft#planning
[2024/10] PokéChamp: An Expert-level Minimax Language AgentICML 2025 [paper][code] #competition#planning
[2024/09] Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study CasearXiv [paper][code] #action#planning#training
[2024/09] Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchICLR 2025 Poster [paper] #communication#planning#multi-agent#training
[2024/09] Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackICLR 2025 Poster [paper] #text-adventure#planning#training#self-improvement
[2024/09] BadRobot: Jailbreaking Embodied LLM Agents in the Physical WorldICLR 2025 Poster [paper] #sim-embodied#planning#vlm
[2024/09] BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesICLR 2025 Poster [paper] #action#planning#training#vlm
[2024/08] Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) 2024 [paper] #competition#planning#multi-agent#training
[2024/08] Atari-GPT: Investigating the Capabilities of Multimodal Large Language Models as Low-Level Policies for Atari GamesarXiv [paper] #action#planning#training
[2024/07] Arigraph: Learning knowledge graph world models with episodic memory for llm agentsIJCAI 2024 [paper] #text-adventure#planning#memory
[2024/07] Odyssey: Empowering Agents with Open-World Skills.IJCAI 2024 [paper][code] #minecraft#planning#tool-use
[2024/06] VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in MinecraftFindings of ACL 2024 [paper][code] #minecraft#planning#multi-agent
[2024/05] Agent Planning with World Knowledge ModelNeurIPS 2024 [paper][code] #text-adventure#planning#memory#world-model
[2024/05] Agent hospital: A simulacrum of hospital with evolvable medical agentsarXiv [paper] #sim-social#planning
[2024/05] Towards Efficient LLM Grounding for Embodied Multi-Agent CollaborationACL 2024 [paper][code] #cooperation#planning#multi-agent#training
[2024/05] Reflective Multi-Agent Collaboration based on Large Language ModelsNeurIPS 2024 poster [paper] #competition#planning#multi-agent#training
[2024/05] AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental LearningNeurIPS 2024 poster [paper][code] #text-adventure#planning
[2024/05] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 poster [paper] #communication#planning#multi-agent#self-improvement
[2024/04] ReAct Meets ActRe: When Language Agents Enjoy Training Data AutonomyarXiv [paper] #text-adventure#planning#training#self-improvement
[2024/04] AgentKit: Flow Engineering with Graphs, not CodingarXiv [paper][code] #crafter#planning
[2024/03] KnowAgent: Knowledge-Augmented Planning for LLM-Based AgentsNAACL 2024 [paper][code] #text-adventure#planning
[2024/03] Cradle: Empowering Foundation Agents Towards General Computer ControlICML 2024 [paper][code] #video-adventure#planning
[2024/03] Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents2024 IEEE Conference on Games (CoG) 2024 [paper][code] #video-adventure#planning#prompting
[2024/03] Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation (HAS)ICLR 2024 Workshop [paper] #minecraft#planning#multi-agent#vlm
[2024/03] Embodied LLM Agents Learn to Cooperate in Organized TeamsIEEE Transactions on Computational Social Systems 2024 [paper] #competition#planning#multi-agent
[2024/03] Will GPT-4 Run DOOM?IEEE Transactions on Games 2024 [paper][code] #action#planning#training
[2024/03] ReAct Meets ActRe: Autonomous Annotation of Agent Trajectories for Contrastive Self-TrainingCOLM [paper] #text-adventure#planning#training#self-improvement
[2024/03] StateFlow: Enhancing LLM Task-Solving through State-Driven WorkflowsCOLM [paper] #text-adventure#planning#self-improvement
[2024/03] AgentKit: Structured LLM Reasoning with Dynamic GraphsCOLM [paper] #crafter#planning
[2024/02] Empowering Large Language Model Agents through Action LearningCOLM [paper][code] #text-adventure#planning#self-improvement
[2023/12] MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active PerceptionCVPR 2024 [paper][code] #minecraft#planning#vlm
[2023/12] Creative Agents: Empowering Agents with Imagination for Creative TasksUAI 2023 [paper][code] #minecraft#planning
[2023/12] Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization ApproachNeurIPS 2024 poster [paper][code] #competition#planning
[2023/11] ADaPT: As-Needed Decomposition and Planning with Language ModelsNAACL 2023 [paper][code] #text-adventure#planning
[2023/11] JARVIS-1: Open-world Multi-task Agents with Memory-Augmented Multimodal Language ModelsTPAMI 2023 [paper][code] #minecraft#planning#memory
[2023/10] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language ModelsICML 2024 [paper][code] #text-adventure#planning#training
[2023/10] LLaMA Rider: Spurring Large Language Models to Explore the Open WorldNAACL 2023 [paper][code] #minecraft#planning#training
[2023/08] AgentSims: An Open-Source Sandbox for Large Language Model EvaluationarXiv [paper][code] #sim-social#planning#tool-use
[2023/07] Building Cooperative Embodied Agents Modularly with Large Language ModelsICLR 2024 [paper][code] #cooperation#planning#multi-agent#training
[2023/05] Language Models Meet World Models: Embodied Experiences Enhance Language ModelsNeurIPS 2023 [paper][code] #sim-embodied#planning#world-model
[2023/05] SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksNeurIPS 2023 [paper][code] #text-adventure#planning
[2023/03] Plan4MC: Skill Reinforcement Learning and Planning for Open-World Minecraft TasksFMDM@NeurIPS2023 [paper][code] #minecraft#planning#training
[2023/02] Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task AgentsNeurIPS 2023 [paper][code] #minecraft#planning#prompting
[2022/12] LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsICCV 2023 [paper] #sim-embodied#planning#prompting
[2022/10] ReAct: Synergizing Reasoning and Acting in Language ModelsICLR 2023 [paper][code] #text-adventure#planning#training
[2020/10] ALFWorld: Aligning Text and Embodied Environments for Interactive LearningICLR 2021 [paper][code] #text-adventure#planning
memory
[2026/06] SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent TrainingarXiv [paper][code] #text-adventure#memory#training
[2026/06] SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at ScalearXiv [paper] #text-adventure#memory#self-improvement
[2026/05] Belief Memory: Agent Memory Under Partial ObservabilityarXiv [paper] #text-adventure#memory
[2026/05] GASim: A Graph-Accelerated Hybrid Framework for Social SimulationarXiv [paper][code] #sim-social#memory#multi-agent
[2026/05] ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic ProfilesarXiv [paper] #sim-social#memory#multi-agent
[2026/05] Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM AgentsarXiv [paper] #text-adventure#memory#tool-use
[2026/05] SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill GraphsarXiv [paper] #text-adventure#memory#tool-use#training
[2026/05] Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM AgentsarXiv [paper] #text-adventure#memory
[2026/03] RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic FeedbackarXiv [paper] #text-adventure#memory#training#self-improvement
[2026/03] Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy ScenariosarXiv [paper] #competition#planning#memory#multi-agent
[2025/12] EZYer: A simulacrum of high school with generative agentarXiv [paper] #sim-social#memory
[2025/12] ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement LearningarXiv [paper] #sim-embodied#planning#memory#training
[2025/12] Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive RefinementProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#memory
[2025/11] Knowledge Graph-enhanced Large Language Model for Incremental Game PlayTestingIEICE Transactions on Information and Systems 2025 [paper] #minecraft#memory
[2025/10] Fine-tuning with RAG for Improving LLM Learning of New SkillsarXiv [paper] #text-adventure#planning#memory#training
[2025/10] Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy GamesarXiv [paper] #competition#memory
[2025/10] Alita-G: Self-Evolving Generative Agent for Agent GenerationarXiv [paper] #sim-social#memory#self-improvement
[2025/09] World Model Implanting for Test-time Adaptation of Embodied AgentsICML 2025 poster [paper] #text-adventure#memory#world-model#prompting
[2025/09] Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationICLR 2026 Poster [paper] #text-adventure#memory#training
[2025/08] Vistawise: Building Cost-effective Agent with Cross-modal Knowledge Graph for MinecraftEMNLP 2025 [paper] #minecraft#memory#tool-use
[2025/08] CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMsarXiv [paper] #memory
[2025/06] StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi TurnsarXiv [paper] #text-adventure#memory
[2025/05] Knowledge Retrieval in LLM Gaming: A Shift from Entity-Centric to Goal-Oriented GraphsKnowledge-Based Systems 2025 [paper] #minecraft#planning#memory
[2025/03] Plancraft: an evaluation dataset for planning with LLM agentsCOLM 2025 [paper] #minecraft#planning#memory#tool-use
[2025/02] LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative PlanningarXiv [paper] #crafter#planning#memory#multi-agent
[2024/11] MrSteve: Instruction-Following Agents with What-Where-When MemoryICLR 2025 [paper][code] #minecraft#memory
[2024/09] MrSteve: Instruction-Following Agents in Minecraft with What-Where-When MemoryICLR 2025 Poster [paper] #minecraft#memory
[2024/07] Arigraph: Learning knowledge graph world models with episodic memory for llm agentsIJCAI 2024 [paper] #text-adventure#planning#memory
[2024/05] Agent Planning with World Knowledge ModelNeurIPS 2024 [paper][code] #text-adventure#planning#memory#world-model
[2023/11] JARVIS-1: Open-world Multi-task Agents with Memory-Augmented Multimodal Language ModelsTPAMI 2023 [paper][code] #minecraft#planning#memory
[2023/11] See and Think: Embodied Agent in Virtual EnvironmentECCV 2023 [paper][code] #minecraft#memory
[2023/10] Lyfe Agents: Generative agents for low-cost real-time social interactionsarXiv [paper] #sim-social#memory#multi-agent#self-improvement
[2023/09] Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4COLM 2024 [paper] #competition#planning#memory#prompting
[2023/09] Exploring Large Language Models for Communication Games: An Empirical Study on WerewolfarXiv [paper] #communication#memory
multi-agent
[2026/06] Can LLM Agents Sustain Long-Horizon Organizational Dynamics?arXiv [paper] #sim-social#planning#multi-agent
[2026/06] Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social SimulationarXiv [paper] #sim-social#multi-agent
[2026/06] SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language ModelsarXiv [paper] #competition#multi-agent
[2026/05] GASim: A Graph-Accelerated Hybrid Framework for Social SimulationarXiv [paper][code] #sim-social#memory#multi-agent
[2026/05] ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic ProfilesarXiv [paper] #sim-social#memory#multi-agent
[2026/05] GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM CollectivesarXiv [paper] #competition#multi-agent#prompting
[2026/05] ALSO: Adversarial Online Strategy Optimization for Social AgentsarXiv [paper] #sim-social#multi-agent#training
[2026/05] MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMsarXiv [paper] #communication#multi-agent
[2026/05] PatchBoard: Schema-Grounded State Mutation for Reliable and Auditable LLM Multi-Agent CollaborationarXiv [paper] #text-adventure#multi-agent
[2026/04] MARL-GPT: Foundation Model for Multi-Agent Reinforcement LearningProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #competition#multi-agent#training
[2026/04] Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent WorkflowsarXiv [paper] #text-adventure#planning#memory#multi-agent
[2026/04] Don't Make the LLM Read the Graph: Make the Graph ThinkarXiv [paper] #cooperation#multi-agent
[2026/04] Gated Coordination for Efficient Multi-Agent Collaboration in Minecraft GamearXiv [paper] #minecraft#memory#multi-agent#vlm
[2026/03] How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game EnvironmentarXiv [paper] #text-adventure#multi-agent#training
[2026/03] Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy ScenariosarXiv [paper] #competition#planning#memory#multi-agent
[2026/03] Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationWWW 2026 [paper][code] #sim-social#multi-agent
[2026/03] Deception and Communication in Autonomous Multi-Agent Systems: An Experimental Study with Among UsProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #communication#multi-agent
[2026/02] The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI SocietiesarXiv [paper] #sim-social#multi-agent#self-improvement
[2026/01] When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered AgentsarXiv [paper] #sim-social#multi-agent
[2026/01] Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction GamesInternational Conference on Agents 2027 [paper] #communication#multi-agent
[2026/01] MARO: Learning Stronger Reasoning from Social InteractionarXiv [paper] #sim-social#multi-agent
[2025/12] WOLF: Werewolf-based Observations for LLM Deception and FalsehoodsarXiv [paper] #communication#multi-agent
[2025/12] Robust Agents in Open-Ended WorldsarXiv [paper] #action#multi-agent#training#generation
[2025/12] Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMsarXiv [paper] #sim-social#multi-agent
[2025/11] DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent CollaborationarXiv [paper] #sim-embodied#planning#multi-agent#world-model
[2025/11] Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal ReasoningarXiv [paper][code] #communication#multi-agent
[2025/11] MADRA: Multi-Agent Debate for Risk-Aware Embodied PlanningarXiv [paper] #sim-embodied#planning#multi-agent#prompting
[2025/10] LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration GamearXiv [paper] #cooperation#multi-agent
[2025/10] Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned StrategiesarXiv [paper] #communication#multi-agent#self-improvement
[2025/10] Social Simulations with Large Language Model Risk Utopian IllusionarXiv [paper] #sim-social#multi-agent
[2025/09] Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM AgentarXiv [paper] #text-adventure#planning#multi-agent#self-improvement
[2025/09] PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments2025 IEEE Conference on Games (CoG) 2025 [paper] #minecraft#multi-agent#self-improvement
[2025/09] HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-MakingarXiv [paper] #competition#multi-agent#training
[2025/09] SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningICLR 2026 Poster [paper] #competition#planning#multi-agent#training
[2025/08] CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative TasksFindings of EMNLP 2025 [paper] #minecraft#planning#multi-agent
[2025/08] A Multi-Agent Pokemon Tournament for Evaluating Strategic Reasoning of Large Language ModelsarXiv [paper] #action#multi-agent
[2025/07] LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative SimulacraarXiv [paper][code] #sim-social#multi-agent#training#role-play
[2025/07] CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language GamesACL 2025 [paper][code] #communication#multi-agent
[2025/06] IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent EnvironmentEMNLP 2025 [paper] #sim-social#multi-agent
[2025/06] WereWolf-Plus: An Update of Werewolf Game setting Based on DSGBencharXiv [paper][code] #communication#multi-agent
[2025/06] The Decrypto Benchmark for Multi-Agent Reasoning and Theory of MindarXiv [paper] #multi-agent#training
[2025/06] AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System NeedarXiv [paper][code] #sim-social#planning#multi-agent
[2025/06] PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language ModelsarXiv [paper] #cooperation#planning#multi-agent
[2025/05] Is Your LLM Really Mastering the Concept? A Multi-Agent BenchmarkarXiv [paper][code] #benchmark#multi-agent
[2025/04] Collaborating Action by Action: A Multi-agent LLM Framework for Embodied ReasoningarXiv [paper][code] #minecraft#multi-agent#training
[2025/04] BookWorld: From Novels to Interactive Agent Societies for Creative Story GenerationarXiv [paper] #sim-social#multi-agent#generation
[2025/04] LLM-PySC2: Starcraft II learning environment for Large Language ModelsNeurIPS 2025 poster [paper] #competition#planning#multi-agent#vlm
[2025/03] Parallelized Planning-Acting for Efficient LLM-based Multi-Agent SystemsarXiv [paper] #minecraft#planning#multi-agent#tool-use
[2025/03] Playing games with Large language models: Randomness and strategyarXiv [paper] #multi-agent
[2025/03] Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic ReasoningCOLM 2025 [paper] #competition#multi-agent#training
[2025/02] LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative PlanningarXiv [paper] #crafter#planning#memory#multi-agent
[2024/12] TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in MinecraftarXiv [paper][code] #minecraft#multi-agent#training#vlm
[2024/09] Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchICLR 2025 Poster [paper] #communication#planning#multi-agent#training
[2024/08] Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) 2024 [paper] #competition#planning#multi-agent#training
[2024/06] VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in MinecraftFindings of ACL 2024 [paper][code] #minecraft#planning#multi-agent
[2024/06] Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract TheoryarXiv [paper] #sim-social#multi-agent
[2024/06] PLAYER: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery GamesarXiv [paper] #communication#multi-agent
[2024/05] Towards Efficient LLM Grounding for Embodied Multi-Agent CollaborationACL 2024 [paper][code] #cooperation#planning#multi-agent#training
[2024/05] Reflective Multi-Agent Collaboration based on Large Language ModelsNeurIPS 2024 poster [paper] #competition#planning#multi-agent#training
[2024/05] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 poster [paper] #communication#planning#multi-agent#self-improvement
[2024/03] MineLand: Simulating Large-Scale Multi-Agent Interactions with Limited Multimodal Senses and Physical NeedsarXiv [paper][code] #minecraft#multi-agent#vlm
[2024/03] Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation (HAS)ICLR 2024 Workshop [paper] #minecraft#planning#multi-agent#vlm
[2024/03] Embodied LLM Agents Learn to Cooperate in Organized TeamsIEEE Transactions on Computational Social Systems 2024 [paper] #competition#planning#multi-agent
[2024/03] Can LLM-Augmented Autonomous Agents Cooperate?, An Evaluation of Their Cooperative Capabilities through Melting PotIEEE Transactions on Artificial Intelligence 2024 [paper] #cooperation#multi-agent
[2024/03] Helmsman of the Masses? Evaluate the Opinion Leadership of Large Language Models in the Werewolf GameCOLM [paper] #communication#multi-agent
[2023/12] Cooperation on the Fly: Exploring Language Agents for Ad Hoc Teamwork in the Avalon GamearXiv [paper] #communication#multi-agent
[2023/12] Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery GamesACL 2023 [paper] #communication#multi-agent#prompting
[2023/11] War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World WarsarXiv [paper][code] #communication#multi-agent
[2023/10] Lyfe Agents: Generative agents for low-cost real-time social interactionsarXiv [paper] #sim-social#memory#multi-agent#self-improvement
[2023/10] Evaluating Multi-agent Coordination Abilities in Large Language ModelsNAACL 2023 [paper] #cooperation#planning#multi-agent#prompting
[2023/10] Theory of Mind for Multi-Agent Collaboration via Large Language ModelsEMNLP 2023 [paper][code] #cooperation#planning#multi-agent#training
[2023/10] LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon GameplayEMNLP 2023 [paper] #communication#multi-agent
[2023/10] Leveraging Word Guessing Games to Assess the Intelligence of Large Language ModelsarXiv [paper][code] #communication#multi-agent
[2023/07] Building Cooperative Embodied Agents Modularly with Large Language ModelsICLR 2024 [paper][code] #cooperation#planning#multi-agent#training
world-model
[2026/05] PriorZero: Bridging Language Priors and World Models for Decision MakingarXiv [paper][code] #text-adventure#planning#world-model#training
[2026/02] World Models for Policy Refinement in StarCraft IIarXiv [paper] #competition#world-model#prompting
[2026/02] CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesarXiv [paper] #text-adventure#planning#world-model#training
[2026/02] Reinforcement World Model Learning for LLM-based AgentsarXiv [paper] #text-adventure#world-model
[2025/11] DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent CollaborationarXiv [paper] #sim-embodied#planning#multi-agent#world-model
[2025/09] World Model Implanting for Test-time Adaptation of Embodied AgentsICML 2025 poster [paper] #text-adventure#memory#world-model#prompting
[2025/09] DreamPhase: Offline Imagination and Uncertainty-Guided Planning for Large-Language-Model AgentsICLR 2026 Poster [paper] #text-adventure#planning#world-model
[2025/07] CoEx -- Co-evolving World-model and ExplorationEMNLP 2025 [paper] #text-adventure#planning#world-model
[2025/06] Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead SearcharXiv [paper] #text-adventure#planning#world-model#training
[2025/05] PoE-World: Compositional World Modeling with Products of Programmatic ExpertsNeurIPS 2025 spotlight [paper] #action#planning#world-model
[2025/05] WALL-E: World Alignment by NeuroSymbolic Learning improves World Model-based LLM AgentsNeurIPS 2025 poster [paper] #minecraft#planning#world-model#prompting
[2025/04] WALL-E 2.0: World Alignment by NeuroSymbolic Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model#prompting
[2025/04] Better Decisions through the Right Causal World ModelarXiv [paper] #action#planning#world-model#training
[2025/02] Optimus-2: Multimodal World Model for Open-World Minecraft AgentsCVPR 2025 [paper][code] #minecraft#planning#world-model#vlm
[2024/10] WALL-E: World Alignment by Rule Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model
[2024/05] Agent Planning with World Knowledge ModelNeurIPS 2024 [paper][code] #text-adventure#planning#memory#world-model
[2023/05] Language Models Meet World Models: Embodied Experiences Enhance Language ModelsNeurIPS 2023 [paper][code] #sim-embodied#planning#world-model
[2023/04] Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open QuestionsarXiv [paper] #text-adventure#world-model
tool-use
[2026/05] Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement LearningarXiv [paper] #text-adventure#tool-use#training
[2026/05] Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM AgentsarXiv [paper] #text-adventure#memory#tool-use
[2026/05] PokerSkill: LLMs Can Play Expert-Level Poker without Training or SolversarXiv [paper][code] #competition#tool-use#prompting
[2026/05] SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill GraphsarXiv [paper] #text-adventure#memory#tool-use#training
[2026/04] Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language ModelsarXiv [paper] #tool-use#training#self-improvement
[2025/09] Speculative Actions: A Lossless Framework for Faster AI AgentsICLR 2026 Oral [paper] #competition#tool-use
[2025/09] Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningICLR 2026 Poster [paper] #text-adventure#tool-use#training
[2025/08] Vistawise: Building Cost-effective Agent with Cross-modal Knowledge Graph for MinecraftEMNLP 2025 [paper] #minecraft#memory#tool-use
[2025/05] Agent-Environment Alignment via Automated Interface GenerationarXiv [paper][code] #text-adventure#tool-use
[2025/05] Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RLNeurIPS 2025 poster [paper] #communication#planning#tool-use#training
[2025/03] Parallelized Planning-Acting for Efficient LLM-based Multi-Agent SystemsarXiv [paper] #minecraft#planning#multi-agent#tool-use
[2025/03] Plancraft: an evaluation dataset for planning with LLM agentsCOLM 2025 [paper] #minecraft#planning#memory#tool-use
[2024/07] AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding AgentsACL 2024 [paper][code] #text-adventure#tool-use
[2024/07] Odyssey: Empowering Agents with Open-World Skills.IJCAI 2024 [paper][code] #minecraft#planning#tool-use
[2024/04] Learning From Failure: Integrating Negative Examples When Fine-tuning Large Language Models as AgentarXiv [paper][code] #text-adventure#tool-use#training
[2023/08] AgentSims: An Open-Source Sandbox for Large Language Model EvaluationarXiv [paper][code] #sim-social#planning#tool-use
[2023/05] VOYAGER: An Open-Ended Embodied Agent with Large Language ModelsFMDM@NeurIPS2023 [paper][code] #minecraft#tool-use#training
training
[2026/06] SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent TrainingarXiv [paper][code] #text-adventure#memory#training
[2026/06] Cross-Environment Neural Reranking for Sample-Efficient Action Selection in Text-Based AgentsarXiv [paper] #text-adventure#training
[2026/06] When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent TrainingarXiv [paper] #text-adventure#training
[2026/05] Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement LearningarXiv [paper] #action#training#vlm
[2026/05] T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement LearningarXiv [paper][code] #text-adventure#training
[2026/05] ANO: A Principled Approach to Robust Policy OptimizationarXiv [paper] #action#training
[2026/05] From History to State: Constant-Context Skill Learning for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/05] Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State UncertaintyarXiv [paper] #sim-embodied#training
[2026/05] PriorZero: Bridging Language Priors and World Models for Decision MakingarXiv [paper][code] #text-adventure#planning#world-model#training
[2026/05] Generalization or Memorization? Brittleness Testing for Chess-Trained Language ModelsarXiv [paper] #competition#training
[2026/05] ALSO: Adversarial Online Strategy Optimization for Social AgentsarXiv [paper] #sim-social#multi-agent#training
[2026/05] Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplayarXiv [paper] #action#planning#training#vlm
[2026/05] What and When to Distill: Selective Hindsight Distillation for Multi-Turn AgentsarXiv [paper] #text-adventure#training
[2026/05] GROW: Aligning GRPO with State-Action Modeling for Open-World VLM AgentsarXiv [paper] #minecraft#training#vlm
[2026/05] SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit AssignmentarXiv [paper] #text-adventure#training
[2026/05] SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill GraphsarXiv [paper] #text-adventure#memory#tool-use#training
[2026/05] R2V Agent: Teaching SLMs When to Ask for HelparXiv [paper] #text-adventure#training
[2026/04] MARL-GPT: Foundation Model for Multi-Agent Reinforcement LearningProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #competition#multi-agent#training
[2026/04] Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM AgentsarXiv [paper][code] #text-adventure#training
[2026/04] Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language ModelsarXiv [paper] #tool-use#training#self-improvement
[2026/04] DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based AgentsarXiv [paper][code] #text-adventure#training
[2026/03] On the Strengths and Weaknesses of Data for Open-set Embodied AssistancearXiv [paper] #cooperation#training
[2026/03] Hindsight Credit Assignment for Long-Horizon LLM AgentsarXiv [paper] #text-adventure#training
[2026/03] How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game EnvironmentarXiv [paper] #text-adventure#multi-agent#training
[2026/03] PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy OptimizationWWW 2026 [paper] #sim-social#training
[2026/03] Grounded Chess Reasoning in Language Models via Master DistillationarXiv [paper] #competition#planning#training
[2026/03] RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic FeedbackarXiv [paper] #text-adventure#memory#training#self-improvement
[2026/02] Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker EnvironmentsarXiv [paper] #action#training
[2026/02] Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/02] VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case StudyarXiv [paper] #competition#training
[2026/02] CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesarXiv [paper] #text-adventure#planning#world-model#training
[2026/01] NitroGen: An Open Foundation Model for Generalist Gaming AgentsarXiv [paper] #benchmark#training
[2026/01] Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/01] Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient UpdatesarXiv [paper][code] #text-adventure#training
[2026/01] HumanLLM: Towards Personalized Understanding and Simulation of Human NaturearXiv [paper] #sim-social#training
[2025/12] Synergizing Code Coverage and Gameplay Intent: Coverage-Aware Game Playtesting with LLM-Guided Reinforcement LearningarXiv [paper] #minecraft#training
[2025/12] ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement LearningarXiv [paper] #sim-embodied#planning#memory#training
[2025/12] Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in DiplomacyarXiv [paper] #communication#training
[2025/12] Robust Agents in Open-Ended WorldsarXiv [paper] #action#multi-agent#training#generation
[2025/10] Fine-tuning with RAG for Improving LLM Learning of New SkillsarXiv [paper] #text-adventure#planning#memory#training
[2025/10] Graph-Enhanced Policy Optimization in LLM Agent TrainingarXiv [paper] #text-adventure#planning#training
[2025/10] SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory GrapharXiv [paper] #text-adventure#training
[2025/09] HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-MakingarXiv [paper] #competition#multi-agent#training
[2025/09] Exploration with Foundation Models: Capabilities, Limitations, and Hybrid ApproachesarXiv [paper] #action#training#vlm
[2025/09] Goal-Guided Efficient Exploration via Large Language Model in Reinforcement LearningarXiv [paper] #crafter#planning#training
[2025/09] Reward Is Enough: LLMs Are In-Context Reinforcement LearnersICLR 2026 Poster [paper] #text-adventure#training#self-improvement
[2025/09] Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful DemonstrationsICLR 2026 Poster [paper] #text-adventure#training
[2025/09] SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningICLR 2026 Poster [paper] #competition#planning#multi-agent#training
[2025/09] Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationICLR 2026 Poster [paper] #text-adventure#memory#training
[2025/09] Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningICLR 2026 Poster [paper] #text-adventure#tool-use#training
[2025/08] Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World SuccessProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#planning#training#vlm
[2025/08] Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press DiplomacyAAAI 2025 [paper] #communication#training
[2025/08] Learning Game-Playing Agents with Generative Code OptimizationarXiv [paper] #action#training#self-improvement
[2025/08] SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making TasksarXiv [paper] #competition#planning#training#self-improvement
[2025/07] LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative SimulacraarXiv [paper][code] #sim-social#multi-agent#training#role-play
[2025/06] Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video GamesarXiv [paper][code] #benchmark#training
[2025/05] Enfoque Odychess: Un método dialéctico, constructivista y adaptativo para la enseñanza del ajedrez con inteligencias artificiales generativasarXiv [paper] #competition#training
[2025/05] Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme OnearXiv [paper] #action#training
[2025/05] Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningICML 2025 poster [paper] #text-adventure#planning#training
[2025/05] Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsNeurIPS 2025 poster [paper] #action#training
[2025/05] Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RLNeurIPS 2025 poster [paper] #communication#planning#tool-use#training
[2025/05] Can Large Language Models Master Complex Card Games?NeurIPS 2025 poster [paper][code] #competition#training
[2025/05] MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural LearningNeurIPS 2025 poster [paper] #minecraft#training
[2025/05] LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language ModelsNeurIPS 2025 spotlight [paper][code] #action#training
[2025/04] Collaborating Action by Action: A Multi-agent LLM Framework for Embodied ReasoningarXiv [paper][code] #minecraft#multi-agent#training
[2025/04] Better Decisions through the Right Causal World ModelarXiv [paper] #action#planning#world-model#training
[2025/04] Enhancing Player Enjoyment with a Two-Tier DRL and LLM-Based Agent System for Fighting GamesarXiv [paper] #training
[2025/04] Monte Carlo Planning with Large Language Model for Text-Based Game AgentsICLR 2025 Poster [paper] #text-adventure#planning#training
[2025/04] MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model FrameworkNeurIPS 2025 poster [paper] #sim-social#planning#training
[2025/03] GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow NetworksCVPR 2025 [paper] #text-adventure#planning#training#vlm
[2025/03] Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic ReasoningCOLM 2025 [paper] #competition#multi-agent#training
[2025/02] Process Reward Models for LLM Agents: Practical Framework and DirectionsarXiv [paper][code] #text-adventure#training
[2025/01] POKERBENCH: Training Large Language Models to become Professional Poker PlayersAAAI 2025 [paper] #competition#planning#training
[2025/01] DVM: Towards Controllable LLM Agents in Social Deduction GamesIEEE International Conference on Acoustics, Speech, and Signal Processing 2025 [paper] #communication#training
[2025/01] Complete Chess Games Enable LLM Become A Chess MasterNAACL 2025 [paper] #competition#training
[2025/01] LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal DemonstrationsICML 2025 poster [paper] #action#planning#training
[2025/01] Language Models as Implicit Tree SearchICML 2025 poster [paper] #competition#planning#training
[2025/01] LARM: Large Auto-Regressive Model for Long-Horizon Embodied IntelligenceICML 2025 poster [paper] #minecraft#training
[2024/12] TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in MinecraftarXiv [paper][code] #minecraft#multi-agent#training#vlm
[2024/12] Fine-tuning large vision-language models as decision-making agents via reinforcement learningNeurIPS 2024 [paper][code] #text-adventure#training#vlm
[2024/09] Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study CasearXiv [paper][code] #action#planning#training
[2024/09] Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchICLR 2025 Poster [paper] #communication#planning#multi-agent#training
[2024/09] Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-Based Decision-Making SystemsICLR 2025 Poster [paper] #sim-embodied#training
[2024/09] Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackICLR 2025 Poster [paper] #text-adventure#planning#training#self-improvement
[2024/09] BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesICLR 2025 Poster [paper] #action#planning#training#vlm
[2024/08] Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) 2024 [paper] #competition#planning#multi-agent#training
[2024/08] Atari-GPT: Investigating the Capabilities of Multimodal Large Language Models as Low-Level Policies for Atari GamesarXiv [paper] #action#planning#training
[2024/07] OmniJARVIS: Omni-Modal Open-World Agents in MinecraftNeurIPS 2024 [paper][code] #minecraft#training#vlm
[2024/06] STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language ModelsACL 2024 [paper][code] #text-adventure#training
[2024/05] Towards Efficient LLM Grounding for Embodied Multi-Agent CollaborationACL 2024 [paper][code] #cooperation#planning#multi-agent#training
[2024/05] Policy Improvement using Language Feedback ModelsNeurIPS 2024 poster [paper] #text-adventure#training
[2024/05] Learning to Discuss Strategically: A Case Study on One Night Ultimate WerewolfNeurIPS 2024 poster [paper] #communication#training
[2024/05] Reflective Multi-Agent Collaboration based on Large Language ModelsNeurIPS 2024 poster [paper] #competition#planning#multi-agent#training
[2024/04] Learning From Failure: Integrating Negative Examples When Fine-tuning Large Language Models as AgentarXiv [paper][code] #text-adventure#tool-use#training
[2024/04] ReAct Meets ActRe: When Language Agents Enjoy Training Data AutonomyarXiv [paper] #text-adventure#planning#training#self-improvement
[2024/04] World Models with Hints of Large Language Models for Goal AchievingNAACL 2024 [paper] #crafter#training#vlm
[2024/04] Self-playing Adversarial Language Game Enhances LLM ReasoningNeurIPS 2024 [paper][code] #communication#training#self-improvement
[2024/03] Language Guided Exploration for RL Agents in Text EnvironmentsNAACL 2024 [paper][code] #text-adventure#training
[2024/03] Trial and Error: Exploration-Based Trajectory Optimization for LLM AgentsACL 2024 [paper][code] #text-adventure#training#self-improvement
[2024/03] EnvGen: Generating and Adapting Environments via LLMs for Training Embodied AgentsCOLM [paper] #crafter#training
[2024/03] SOTOPIA-$\pi$: Interactive Learning of Socially Intelligent Language AgentsACL 2024 [paper][code] #sim-social#training
[2024/03] Will GPT-4 Run DOOM?IEEE Transactions on Games 2024 [paper][code] #action#planning#training
[2024/03] O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language ModelsCOLM [paper] #text-adventure#training
[2024/03] ReAct Meets ActRe: Autonomous Annotation of Agent Trajectories for Contrastive Self-TrainingCOLM [paper] #text-adventure#planning#training#self-improvement
[2024/02] PokéLLMon: A Human-Parity Agent for Pokémon Battles with Large Language ModelsTOIT 2025 [paper][code] #competition#training
[2024/02] Agent-Pro: Learning to Evolve via Policy-Level Reflection and OptimizationACL 2024 [paper][code] #competition#training
[2024/02] Enhance Reasoning for Large Language Models in the Game WerewolfarXiv [paper] #communication#training
[2024/01] True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement LearningarXiv [paper][code] #sim-embodied#training
[2024/01] PokerGPT: An End-to-End Lightweight Solver for Multi-Player Texas Hold'em via Large Language ModelarXiv [paper] #competition#training
[2024/01] SwarmBrain: Embodied agent for real-time strategy game StarCraft II via large language modelsarXiv [paper] #competition#training
[2023/12] Auto MC-Reward: Automated Dense Reward Design with Large Language Models for MinecraftCVPR 2023 [paper] #minecraft#training
[2023/10] FireAct: Toward Language Agent Fine-tuningarXiv [paper][code] #text-adventure#training
[2023/10] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language ModelsICML 2024 [paper][code] #text-adventure#planning#training
[2023/10] LLaMA Rider: Spurring Large Language Models to Explore the Open WorldNAACL 2023 [paper][code] #minecraft#planning#training
[2023/09] AdaRefiner: Refining Decisions of Language Models with Adaptive FeedbackNAACL 2023 [paper] #crafter#training
[2023/07] Building Cooperative Embodied Agents Modularly with Large Language ModelsICLR 2024 [paper][code] #cooperation#planning#multi-agent#training
[2023/05] Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and MemoryarXiv [paper] #minecraft#training
[2023/05] VOYAGER: An Open-Ended Embodied Agent with Large Language ModelsFMDM@NeurIPS2023 [paper][code] #minecraft#tool-use#training
[2023/03] Plan4MC: Skill Reinforcement Learning and Planning for Open-World Minecraft TasksFMDM@NeurIPS2023 [paper][code] #minecraft#planning#training
[2023/03] Reflexion: Language Agents with Verbal Reinforcement LearningNeurIPS 2023 [paper][code] #text-adventure#training#self-improvement
[2023/02] Guiding Pretraining in Reinforcement Learning with Large Language ModelsICML 2023 [paper] #crafter#training
[2023/02] Grounding Large Language Models in Interactive Environments with Online Reinforcement LearningICML 2023 [paper][code] #action#training
[2022/10] ReAct: Synergizing Reasoning and Acting in Language ModelsICLR 2023 [paper][code] #text-adventure#planning#training
self-improvement
[2026/06] SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at ScalearXiv [paper] #text-adventure#memory#self-improvement
[2026/05] Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within AgentsarXiv [paper] #text-adventure#training#self-improvement
[2026/05] Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM AgentsarXiv [paper] #minecraft#self-improvement
[2026/04] Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language ModelsarXiv [paper] #tool-use#training#self-improvement
[2026/04] GraSP: Graph-Structured Skill Compositions for LLM AgentsarXiv [paper] #text-adventure#planning#memory#self-improvement
[2026/03] RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic FeedbackarXiv [paper] #text-adventure#memory#training#self-improvement
[2026/02] MemSkill: Learning and Evolving Memory Skills for Self-Evolving AgentsarXiv [paper] #text-adventure#self-improvement
[2026/02] The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI SocietiesarXiv [paper] #sim-social#multi-agent#self-improvement
[2025/10] Alita-G: Self-Evolving Generative Agent for Agent GenerationarXiv [paper] #sim-social#memory#self-improvement
[2025/10] Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned StrategiesarXiv [paper] #communication#multi-agent#self-improvement
[2025/09] Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM AgentarXiv [paper] #text-adventure#planning#multi-agent#self-improvement
[2025/09] PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments2025 IEEE Conference on Games (CoG) 2025 [paper] #minecraft#multi-agent#self-improvement
[2025/09] Code Driven Planning with Domain-Adaptive CriticarXiv [paper] #text-adventure#planning#self-improvement
[2025/09] Reward Is Enough: LLMs Are In-Context Reinforcement LearnersICLR 2026 Poster [paper] #text-adventure#training#self-improvement
[2025/08] Learning Game-Playing Agents with Generative Code OptimizationarXiv [paper] #action#training#self-improvement
[2025/08] SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making TasksarXiv [paper] #competition#planning#training#self-improvement
[2025/06] Automated Skill Discovery for Language Agents through Exploration and Iterative FeedbackarXiv [paper] #crafter#self-improvement
[2025/06] OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic ReflectionsarXiv [paper] #text-adventure#planning#training#self-improvement
[2025/02] TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning.arXiv [paper] #text-adventure#self-improvement
[2024/12] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 [paper][code] #communication#self-improvement
[2024/09] Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackICLR 2025 Poster [paper] #text-adventure#planning#training#self-improvement
[2024/06] Watch Every Step! LLM Agent Learning via Iterative Step-Level Process RefinementEMNLP 2024 [paper][code] #text-adventure#self-improvement
[2024/05] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 poster [paper] #communication#planning#multi-agent#self-improvement
[2024/04] ReAct Meets ActRe: When Language Agents Enjoy Training Data AutonomyarXiv [paper] #text-adventure#planning#training#self-improvement
[2024/04] Self-playing Adversarial Language Game Enhances LLM ReasoningNeurIPS 2024 [paper][code] #communication#training#self-improvement
[2024/03] Trial and Error: Exploration-Based Trajectory Optimization for LLM AgentsACL 2024 [paper][code] #text-adventure#training#self-improvement
[2024/03] ReAct Meets ActRe: Autonomous Annotation of Agent Trajectories for Contrastive Self-TrainingCOLM [paper] #text-adventure#planning#training#self-improvement
[2024/03] StateFlow: Enhancing LLM Task-Solving through State-Driven WorkflowsCOLM [paper] #text-adventure#planning#self-improvement
[2024/02] Soft Self-Consistency Improves Language Model AgentsarXiv [paper][code] #text-adventure#self-improvement
[2024/02] Empowering Large Language Model Agents through Action LearningCOLM [paper][code] #text-adventure#planning#self-improvement
[2023/10] Lyfe Agents: Generative agents for low-cost real-time social interactionsarXiv [paper] #sim-social#memory#multi-agent#self-improvement
[2023/03] Reflexion: Language Agents with Verbal Reinforcement LearningNeurIPS 2023 [paper][code] #text-adventure#training#self-improvement
prompting
[2026/05] GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM CollectivesarXiv [paper] #competition#multi-agent#prompting
[2026/05] PokerSkill: LLMs Can Play Expert-Level Poker without Training or SolversarXiv [paper][code] #competition#tool-use#prompting
[2026/04] DORA Explorer: Improving the Exploration Ability of LLMs Without TrainingarXiv [paper] #text-adventure#planning#prompting
A Survey on Large Language Model-Based Game Agents (ACM CSUR)
🔥 Must-read papers for LLM-based Game agents.
📘 Our survey has been accepted by ACM Computing Surveys (CSUR). We are preparing the camera ready. Feel free to reach out if you find missing reference.
💫 We continuously update the GitHub list on a weekly basis.
📝 If you discover any papers that are suitable but not yet included, please open an issue or submit a pull request.
[2026/01] MineNPC-Task: Task Suite for Memory-Aware Minecraft AgentsarXiv [paper] #minecraft
[2025/12] Synergizing Code Coverage and Gameplay Intent: Coverage-Aware Game Playtesting with LLM-Guided Reinforcement LearningarXiv [paper] #minecraft#training
[2025/11] Knowledge Graph-enhanced Large Language Model for Incremental Game PlayTestingIEICE Transactions on Information and Systems 2025 [paper] #minecraft#memory
[2025/09] PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments2025 IEEE Conference on Games (CoG) 2025 [paper] #minecraft#multi-agent#self-improvement
[2025/09] Experience-based Knowledge Correction for Robust Planning in MinecraftICLR 2026 Poster [paper] #minecraft#planning
[2025/08] CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative TasksFindings of EMNLP 2025 [paper] #minecraft#planning#multi-agent
[2025/08] Vistawise: Building Cost-effective Agent with Cross-modal Knowledge Graph for MinecraftEMNLP 2025 [paper] #minecraft#memory#tool-use
[2025/07] VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning SystemsarXiv [paper] #minecraft#vlm
[2025/07] Referential ambiguity and clarification requests: comparing human and LLM behaviourProceedings of the Eighth Workshop on Computational Models of Reference, Anaphora and Coreference 2025 [paper] #minecraft
[2025/06] Optimus-3: Towards Generalist Multimodal Minecraft Agents with Scalable Task ExpertsarXiv [paper][code] #minecraft#planning
[2025/06] Matrix-Game: Interactive World Foundation ModelarXiv [paper][code] #minecraft
[2025/06] GuessBench: Sensemaking Multimodal Creativity in the WildarXiv [paper] #minecraft#training#vlm
[2025/05] Don’t Just Follow MLLM Plans: Robust and Efficient Planning for Open-World AgentsarXiv [paper] #minecraft#planning#vlm
[2025/05] Knowledge Retrieval in LLM Gaming: A Shift from Entity-Centric to Goal-Oriented GraphsKnowledge-Based Systems 2025 [paper] #minecraft#planning#memory
[2025/05] BeliefNest: A Joint Action Simulator for Embodied Agents with Theory of MindarXiv [paper] #minecraft
[2025/05] MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural LearningNeurIPS 2025 poster [paper] #minecraft#training
[2025/05] WALL-E: World Alignment by NeuroSymbolic Learning improves World Model-based LLM AgentsNeurIPS 2025 poster [paper] #minecraft#planning#world-model#prompting
[2025/04] Collaborating Action by Action: A Multi-agent LLM Framework for Embodied ReasoningarXiv [paper][code] #minecraft#multi-agent#training
[2025/04] WALL-E 2.0: World Alignment by NeuroSymbolic Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model#prompting
[2025/03] Uncertainty in Action: Confidence Elicitation in Embodied AgentsarXiv [paper] #minecraft
[2025/03] Parallelized Planning-Acting for Efficient LLM-based Multi-Agent SystemsarXiv [paper] #minecraft#planning#multi-agent#tool-use
[2025/03] NeSyC: A Neuro-symbolic Continual Learner For Complex Embodied Tasks In Open DomainsICLR 2025 Poster [paper] #minecraft
[2025/03] Word2Minecraft: Generating 3D Game Levels through Large Language ModelsarXiv [paper][code] #minecraft#generation
[2025/03] Plancraft: an evaluation dataset for planning with LLM agentsCOLM 2025 [paper] #minecraft#planning#memory#tool-use
[2025/02] GATE: Graph-based Adaptive Tool Evolution Across Diverse TasksarXiv [paper][code] #minecraft
[2025/02] Optimus-2: Multimodal World Model for Open-World Minecraft AgentsCVPR 2025 [paper][code] #minecraft#planning#world-model#vlm
[2025/01] LARM: Large Auto-Regressive Model for Long-Horizon Embodied IntelligenceICML 2025 poster [paper] #minecraft#training
[2024/12] TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in MinecraftarXiv [paper][code] #minecraft#multi-agent#training#vlm
[2024/11] MrSteve: Instruction-Following Agents with What-Where-When MemoryICLR 2025 [paper][code] #minecraft#memory
[2024/10] WALL-E: World Alignment by Rule Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model
[2024/10] ADAM: An Embodied Causal Agent in Open-World EnvironmentsICLR 2025 [paper][code] #minecraft#planning
[2024/09] MrSteve: Instruction-Following Agents in Minecraft with What-Where-When MemoryICLR 2025 Poster [paper] #minecraft#memory
[2024/07] Odyssey: Empowering Agents with Open-World Skills.IJCAI 2024 [paper][code] #minecraft#planning#tool-use
[2024/07] OmniJARVIS: Omni-Modal Open-World Agents in MinecraftNeurIPS 2024 [paper][code] #minecraft#training#vlm
[2024/06] VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in MinecraftFindings of ACL 2024 [paper][code] #minecraft#planning#multi-agent
[2024/03] MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World ControlarXiv [paper][code] #minecraft
[2024/03] MineLand: Simulating Large-Scale Multi-Agent Interactions with Limited Multimodal Senses and Physical NeedsarXiv [paper][code] #minecraft#multi-agent#vlm
[2024/03] Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation (HAS)ICLR 2024 Workshop [paper] #minecraft#planning#multi-agent#vlm
[2023/12] MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active PerceptionCVPR 2024 [paper][code] #minecraft#planning#vlm
[2023/12] Auto MC-Reward: Automated Dense Reward Design with Large Language Models for MinecraftCVPR 2023 [paper] #minecraft#training
[2023/12] Creative Agents: Empowering Agents with Imagination for Creative TasksUAI 2023 [paper][code] #minecraft#planning
[2023/11] JARVIS-1: Open-world Multi-task Agents with Memory-Augmented Multimodal Language ModelsTPAMI 2023 [paper][code] #minecraft#planning#memory
[2023/11] See and Think: Embodied Agent in Virtual EnvironmentECCV 2023 [paper][code] #minecraft#memory
[2023/10] LLaMA Rider: Spurring Large Language Models to Explore the Open WorldNAACL 2023 [paper][code] #minecraft#planning#training
[2023/10] MCU: A Task-centric Framework for Open-ended Agent Evaluation in MinecraftICML 2023 [paper][code] #minecraft
[2023/10] Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open WorldsICLR 2024 [paper] #minecraft#vlm
[2023/05] Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and MemoryarXiv [paper] #minecraft#training
[2023/05] VOYAGER: An Open-Ended Embodied Agent with Large Language ModelsFMDM@NeurIPS2023 [paper][code] #minecraft#tool-use#training
[2023/03] Plan4MC: Skill Reinforcement Learning and Planning for Open-World Minecraft TasksFMDM@NeurIPS2023 [paper][code] #minecraft#planning#training
[2023/02] Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task AgentsNeurIPS 2023 [paper][code] #minecraft#planning#prompting
[2022/07] Craft an Iron Sword: Dynamically Generating Interactive Game Characters by Prompting Large Language Models Tuned on CodeWordplay@ACL 2022 [paper] #minecraft
text-adventure
[2026/06] SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent TrainingarXiv [paper][code] #text-adventure#memory#training
[2026/06] SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at ScalearXiv [paper] #text-adventure#memory#self-improvement
[2026/06] LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM AgentsarXiv [paper] #text-adventure
[2026/06] From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM AgentsarXiv [paper] #text-adventure#planning
[2026/06] Cross-Environment Neural Reranking for Sample-Efficient Action Selection in Text-Based AgentsarXiv [paper] #text-adventure#training
[2026/06] Unified Context Evolution for LLM AgentsarXiv [paper] #text-adventure
[2026/06] When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent TrainingarXiv [paper] #text-adventure#training
[2026/05] T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement LearningarXiv [paper][code] #text-adventure#training
[2026/04] Aligning Progress and Feasibility: A Neuro-Symbolic Dual Memory Framework for Long-Horizon LLM AgentsarXiv [paper] #text-adventure#planning#memory
[2026/04] Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM AgentsarXiv [paper][code] #text-adventure#training
[2026/04] DORA Explorer: Improving the Exploration Ability of LLMs Without TrainingarXiv [paper] #text-adventure#planning#prompting
[2026/04] From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM AgentsarXiv [paper] #text-adventure#planning
[2026/04] Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent WorkflowsarXiv [paper] #text-adventure#planning#memory#multi-agent
[2026/04] ReDAct: Uncertainty-Aware Deferral for LLM AgentsarXiv [paper] #text-adventure
[2026/04] GraSP: Graph-Structured Skill Compositions for LLM AgentsarXiv [paper] #text-adventure#planning#memory#self-improvement
[2026/04] DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based AgentsarXiv [paper][code] #text-adventure#training
[2026/03] Hindsight Credit Assignment for Long-Horizon LLM AgentsarXiv [paper] #text-adventure#training
[2026/03] How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game EnvironmentarXiv [paper] #text-adventure#multi-agent#training
[2026/03] RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic FeedbackarXiv [paper] #text-adventure#memory#training#self-improvement
[2026/03] Reward Prediction with Factorized World StatesarXiv [paper] #text-adventure#planning#prompting
[2026/02] MemSkill: Learning and Evolving Memory Skills for Self-Evolving AgentsarXiv [paper] #text-adventure#self-improvement
[2026/02] Active Epistemic Control for Query-Efficient Verified PlanningarXiv [paper] #text-adventure#planning
[2026/02] Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/02] TAPE: Tool-Guided Adaptive Planning and Constrained Execution in Language Model AgentsarXiv [paper] #text-adventure#planning
[2026/02] CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesarXiv [paper] #text-adventure#planning#world-model#training
[2026/02] Reinforcement World Model Learning for LLM-based AgentsarXiv [paper] #text-adventure#world-model
[2026/01] Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/01] Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient UpdatesarXiv [paper][code] #text-adventure#training
[2025/12] Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive RefinementProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#memory
[2025/12] GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment SimulatorsarXiv [paper] #text-adventure
[2025/10] Fine-tuning with RAG for Improving LLM Learning of New SkillsarXiv [paper] #text-adventure#planning#memory#training
[2025/10] GenQuest: An LLM-based Text Adventure Game for Language LearnersarXiv [paper] #text-adventure#vlm#generation
[2025/10] Constrained Natural Language Action Planning for Resilient Embodied SystemsarXiv [paper] #text-adventure#planning#prompting
[2025/10] The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with SchemasarXiv [paper] #text-adventure#planning
[2025/10] Graph-Enhanced Policy Optimization in LLM Agent TrainingarXiv [paper] #text-adventure#planning#training
[2025/10] SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory GrapharXiv [paper] #text-adventure#training
[2025/09] World Model Implanting for Test-time Adaptation of Embodied AgentsICML 2025 poster [paper] #text-adventure#memory#world-model#prompting
[2025/09] Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM AgentarXiv [paper] #text-adventure#planning#multi-agent#self-improvement
[2025/09] Code Driven Planning with Domain-Adaptive CriticarXiv [paper] #text-adventure#planning#self-improvement
[2025/09] Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationICLR 2026 Poster [paper] #text-adventure#memory#training
[2025/09] DreamPhase: Offline Imagination and Uncertainty-Guided Planning for Large-Language-Model AgentsICLR 2026 Poster [paper] #text-adventure#planning#world-model
[2025/09] Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningICLR 2026 Poster [paper] #text-adventure#tool-use#training
[2025/08] Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World SuccessProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#planning#training#vlm
[2025/07] CoEx -- Co-evolving World-model and ExplorationEMNLP 2025 [paper] #text-adventure#planning#world-model
[2025/06] Enhancing Decision-Making of Large Language Models via Actor-CriticICML 2025 poster [paper] #text-adventure#training
[2025/06] Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead SearcharXiv [paper] #text-adventure#planning#world-model#training
[2025/06] StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi TurnsarXiv [paper] #text-adventure#memory
[2025/06] OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic ReflectionsarXiv [paper] #text-adventure#planning#training#self-improvement
[2025/06] KnowMap: Efficient Knowledge-Driven Task Adaptation for LLMsarXiv [paper] #text-adventure#training
[2025/05] STORY2GAME: Generating (Almost) Everything in an Interactive Fiction GamearXiv [paper] #text-adventure#generation
[2025/05] LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMsarXiv [paper][code] #text-adventure#planning
[2025/05] Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningICML 2025 poster [paper] #text-adventure#planning#training
[2025/05] ActiveVOO: Value of Observation Guided Active Knowledge Acquisition for Open-World Embodied Lifted Regression PlanningNeurIPS 2025 poster [paper] #text-adventure#planning#prompting#vlm
[2025/03] Haunted House: A text-based game for comparing the flexibility of mental models in humans and LLMsarXiv [paper] #text-adventure
[2025/03] GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow NetworksCVPR 2025 [paper] #text-adventure#planning#training#vlm
[2025/02] TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning.arXiv [paper] #text-adventure#self-improvement
[2025/02] Process Reward Models for LLM Agents: Practical Framework and DirectionsarXiv [paper][code] #text-adventure#training
[2024/12] Fine-tuning large vision-language models as decision-making agents via reinforcement learningNeurIPS 2024 [paper][code] #text-adventure#training#vlm
[2024/09] Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackICLR 2025 Poster [paper] #text-adventure#planning#training#self-improvement
[2024/07] AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding AgentsACL 2024 [paper][code] #text-adventure#tool-use
[2024/07] Arigraph: Learning knowledge graph world models with episodic memory for llm agentsIJCAI 2024 [paper] #text-adventure#planning#memory
[2024/06] Watch Every Step! LLM Agent Learning via Iterative Step-Level Process RefinementEMNLP 2024 [paper][code] #text-adventure#self-improvement
[2024/06] STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language ModelsACL 2024 [paper][code] #text-adventure#training
[2024/05] Agent Planning with World Knowledge ModelNeurIPS 2024 [paper][code] #text-adventure#planning#memory#world-model
[2024/05] THREAD: Thinking Deeper with Recursive SpawningNAACL 2024 [paper] #text-adventure#prompting
[2024/05] Policy Improvement using Language Feedback ModelsNeurIPS 2024 poster [paper] #text-adventure#training
[2024/05] AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental LearningNeurIPS 2024 poster [paper][code] #text-adventure#planning
[2024/04] Learning From Failure: Integrating Negative Examples When Fine-tuning Large Language Models as AgentarXiv [paper][code] #text-adventure#tool-use#training
[2024/04] ReAct Meets ActRe: When Language Agents Enjoy Training Data AutonomyarXiv [paper] #text-adventure#planning#training#self-improvement
[2024/03] KnowAgent: Knowledge-Augmented Planning for LLM-Based AgentsNAACL 2024 [paper][code] #text-adventure#planning
[2024/03] Language Guided Exploration for RL Agents in Text EnvironmentsNAACL 2024 [paper][code] #text-adventure#training
[2024/03] Trial and Error: Exploration-Based Trajectory Optimization for LLM AgentsACL 2024 [paper][code] #text-adventure#training#self-improvement
[2024/03] O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language ModelsCOLM [paper] #text-adventure#training
[2024/03] ReAct Meets ActRe: Autonomous Annotation of Agent Trajectories for Contrastive Self-TrainingCOLM [paper] #text-adventure#planning#training#self-improvement
[2024/03] StateFlow: Enhancing LLM Task-Solving through State-Driven WorkflowsCOLM [paper] #text-adventure#planning#self-improvement
[2024/02] Soft Self-Consistency Improves Language Model AgentsarXiv [paper][code] #text-adventure#self-improvement
[2024/02] Empowering Large Language Model Agents through Action LearningCOLM [paper][code] #text-adventure#planning#self-improvement
[2023/11] ADaPT: As-Needed Decomposition and Planning with Language ModelsNAACL 2023 [paper][code] #text-adventure#planning
[2023/10] FireAct: Toward Language Agent Fine-tuningarXiv [paper][code] #text-adventure#training
[2023/10] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language ModelsICML 2024 [paper][code] #text-adventure#planning#training
[2023/05] SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksNeurIPS 2023 [paper][code] #text-adventure#planning
[2023/04] Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open QuestionsarXiv [paper] #text-adventure#world-model
[2023/03] Reflexion: Language Agents with Verbal Reinforcement LearningNeurIPS 2023 [paper][code] #text-adventure#training#self-improvement
[2022/10] ReAct: Synergizing Reasoning and Acting in Language ModelsICLR 2023 [paper][code] #text-adventure#planning#training
[2022/03] ScienceWorld: Is your Agent Smarter than a 5th Grader?EMNLP 2022 [paper][code] #text-adventure
[2020/10] ALFWorld: Aligning Text and Embodied Environments for Interactive LearningICLR 2021 [paper][code] #text-adventure#planning
[2019/09] Interactive Fiction Games: A Colossal AdventureAAAI 2020 [paper][code] #text-adventure
communication
[2026/05] Evaluating Large Language Models in a Complex Hidden Role GamearXiv [paper] #communication#planning
[2026/05] QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction AgentsarXiv [paper][code] #communication
[2026/05] MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMsarXiv [paper] #communication#multi-agent
[2026/05] Playing with Words, Improving with Rewards: Training Language Models for Creative AssociationarXiv [paper] #communication#training
[2026/04] Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM AgentsarXiv [paper] #communication
[2026/03] Enhancing Consistency of Werewolf AI through Dialogue Summarization and Persona InformationarXiv [paper] #communication#role-play
[2026/03] Deception and Communication in Autonomous Multi-Agent Systems: An Experimental Study with Among UsProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #communication#multi-agent
[2026/01] Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction GamesInternational Conference on Agents 2027 [paper] #communication#multi-agent
[2026/01] Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction GamearXiv [paper] #communication
[2025/12] WOLF: Werewolf-based Observations for LLM Deception and FalsehoodsarXiv [paper] #communication#multi-agent
[2025/12] Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in DiplomacyarXiv [paper] #communication#training
[2025/11] CSP4SDG: Constraint and Information-Theory Based Role Identification in Social Deduction Games with LLM-Enhanced InferenceAAAI 2025 [paper] #communication
[2025/11] Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal ReasoningarXiv [paper][code] #communication#multi-agent
[2025/10] Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned StrategiesarXiv [paper] #communication#multi-agent#self-improvement
[2025/08] Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press DiplomacyAAAI 2025 [paper] #communication#training
[2025/08] What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup PuzzlesAAAI 2025 [paper] #communication
[2025/08] Ethical Considerations of Large Language Models in Game PlayingarXiv [paper] #communication
[2025/07] CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language GamesACL 2025 [paper][code] #communication#multi-agent
[2025/07] Strategy Adaptation in Large Language Model Werewolf AgentsarXiv [paper] #communication#prompting
[2025/06] WereWolf-Plus: An Update of Werewolf Game setting Based on DSGBencharXiv [paper][code] #communication#multi-agent
[2025/06] DipLLM: Fine-Tuning LLM for Strategic Decision-making in DiplomacyICML 2025 poster [paper] #communication#training
[2025/05] Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RLNeurIPS 2025 poster [paper] #communication#planning#tool-use#training
[2025/01] DVM: Towards Controllable LLM Agents in Social Deduction GamesIEEE International Conference on Acoustics, Speech, and Signal Processing 2025 [paper] #communication#training
[2024/12] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 [paper][code] #communication#self-improvement
[2024/09] Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchICLR 2025 Poster [paper] #communication#planning#multi-agent#training
[2024/06] PLAYER: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery GamesarXiv [paper] #communication#multi-agent
[2024/05] Learning to Discuss Strategically: A Case Study on One Night Ultimate WerewolfNeurIPS 2024 poster [paper] #communication#training
[2024/05] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 poster [paper] #communication#planning#multi-agent#self-improvement
[2024/04] Self-playing Adversarial Language Game Enhances LLM ReasoningNeurIPS 2024 [paper][code] #communication#training#self-improvement
[2024/03] Helmsman of the Masses? Evaluate the Opinion Leadership of Large Language Models in the Werewolf GameCOLM [paper] #communication#multi-agent
[2024/02] Enhance Reasoning for Large Language Models in the Game WerewolfarXiv [paper] #communication#training
[2024/02] What if LLMs Have Different World Views: Simulating Alien Civilizations with LLM-based AgentsarXiv [paper] #communication
[2024/02] Can Large Language Model Agents Simulate Human Trust Behaviors?NeurIPS 2024 [paper] #communication
[2024/02] Large Language Models Fall Short: Understanding Complex Relationships in Detective NarrativesACL 2024 [paper] #communication
[2023/12] Can Large Language Models Serve as Rational Players in Game Theory? A Systematic AnalysisAAAI 2024 [paper] #communication
[2023/12] Cooperation on the Fly: Exploring Language Agents for Ad Hoc Teamwork in the Avalon GamearXiv [paper] #communication#multi-agent
[2023/12] Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery GamesACL 2023 [paper] #communication#multi-agent#prompting
[2023/11] War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World WarsarXiv [paper][code] #communication#multi-agent
[2023/11] clembench: Systematic Evaluation of Chat-Optimized Language Models as Conversational AgentsEMNLP 2023 [paper] #communication
[2023/10] Language Agents with Reinforcement Learning for Strategic Play in the Werewolf GameICML 2023 [paper] #communication#training
[2023/10] Avalon's Game of Thoughts: Battle Against Deception through Recursive ContemplationarXiv [paper] #communication#training
[2023/10] LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon GameplayEMNLP 2023 [paper] #communication#multi-agent
[2023/10] Leveraging Word Guessing Games to Assess the Intelligence of Large Language ModelsarXiv [paper][code] #communication#multi-agent
[2023/10] AvalonBench: Evaluating LLMs Playing the Game of AvalonFMDM@NeurIPS2023 [paper][code] #communication
[2023/09] Exploring Large Language Models for Communication Games: An Empirical Study on WerewolfarXiv [paper] #communication#memory
[2023/08] GameEval: Evaluating LLMs on Conversational GamesarXiv [paper][code] #communication
[2022/12] Human-Level Play in the Game of Diplomacy by Combining Language Models with Strategic ReasoningScience [paper] #communication
competition
[2026/06] Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM AgentsarXiv [paper] #competition
[2026/06] SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language ModelsarXiv [paper] #competition#multi-agent
[2026/05] GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM CollectivesarXiv [paper] #competition#multi-agent#prompting
[2026/05] Watermarking Game-Playing Agents in Perfect-Information Extensive-Form GamesarXiv [paper] #competition
[2026/05] Generalization or Memorization? Brittleness Testing for Chess-Trained Language ModelsarXiv [paper] #competition#training
[2026/05] PokerSkill: LLMs Can Play Expert-Level Poker without Training or SolversarXiv [paper][code] #competition#tool-use#prompting
[2026/04] Readable Minds: Emergent Theory-of-Mind-Like Behavior in LLM Poker AgentsarXiv [paper] #competition
[2026/04] MARL-GPT: Foundation Model for Multi-Agent Reinforcement LearningProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #competition#multi-agent#training
[2026/03] Grounded Chess Reasoning in Language Models via Master DistillationarXiv [paper] #competition#planning#training
[2026/03] Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy ScenariosarXiv [paper] #competition#planning#memory#multi-agent
[2026/02] World Models for Policy Refinement in StarCraft IIarXiv [paper] #competition#world-model#prompting
[2026/02] VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case StudyarXiv [paper] #competition#training
[2025/12] LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through ChessarXiv [paper] #competition
[2025/12] Beyond Accuracy: A Geometric Stability Analysis of Large Language Models in Chess EvaluationarXiv [paper] #competition
[2025/10] Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy GamesarXiv [paper] #competition#memory
[2025/10] ChessQA: Evaluating Large Language Models for Chess UnderstandingarXiv [paper] #competition
[2025/10] Out-of-distribution Tests Reveal Compositionality in Chess TransformersarXiv [paper] #competition#planning
[2025/09] HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-MakingarXiv [paper] #competition#multi-agent#training
[2025/09] Speculative Actions: A Lossless Framework for Faster AI AgentsICLR 2026 Oral [paper] #competition#tool-use
[2025/09] SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningICLR 2026 Poster [paper] #competition#planning#multi-agent#training
[2025/08] Tracking World States with Language Models: State-Based Evaluation Using ChessarXiv [paper] #competition
[2025/08] SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making TasksarXiv [paper] #competition#planning#training#self-improvement
[2025/07] Learning to Imitate with Less: Efficient Individual Behavior Modeling in ChessarXiv [paper] #competition
[2025/05] Enfoque Odychess: Un método dialéctico, constructivista y adaptativo para la enseñanza del ajedrez con inteligencias artificiales generativasarXiv [paper] #competition#training
[2025/05] Can Large Language Models Master Complex Card Games?NeurIPS 2025 poster [paper][code] #competition#training
[2025/04] Explore the Reasoning Capability of LLMs in the Chess TestbedNAACL 2025 [paper] #competition
[2025/04] ZeroSumEval: Scaling LLM Evaluation with Inter-Model CompetitionarXiv [paper][code] #competition#planning
[2025/04] LLM-PySC2: Starcraft II learning environment for Large Language ModelsNeurIPS 2025 poster [paper] #competition#planning#multi-agent#vlm
[2025/04] The PokeAgent Challenge: Competitive and Long Context Learning at ScaleNeurIPS Competition Track 2025 [paper] #competition
[2025/03] Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic ReasoningCOLM 2025 [paper] #competition#multi-agent#training
[2025/02] Hierarchical Expert Prompt for Large-Language-Model: An Approach Defeat Elite AI in TextStarCraft II for the First TimearXiv [paper][code] #competition
[2025/02] Implicit Search via Discrete Diffusion: A Study on ChessICLR 2025 Poster [paper][code] #competition#planning
[2025/01] POKERBENCH: Training Large Language Models to become Professional Poker PlayersAAAI 2025 [paper] #competition#planning#training
[2025/01] Complete Chess Games Enable LLM Become A Chess MasterNAACL 2025 [paper] #competition#training
[2025/01] Mastering Board Games by External and Internal Planning with Language ModelsICML 2025 spotlightposter [paper] #competition#planning
[2025/01] Language Models as Implicit Tree SearchICML 2025 poster [paper] #competition#planning#training
[2024/10] PokéChamp: An Expert-level Minimax Language AgentICML 2025 [paper][code] #competition#planning
[2024/08] Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) 2024 [paper] #competition#planning#multi-agent#training
[2024/05] Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game ModelsNeurIPS 2024 poster [paper] #competition
[2024/05] Reflective Multi-Agent Collaboration based on Large Language ModelsNeurIPS 2024 poster [paper] #competition#planning#multi-agent#training
[2024/03] Embodied LLM Agents Learn to Cooperate in Organized TeamsIEEE Transactions on Computational Social Systems 2024 [paper] #competition#planning#multi-agent
[2024/02] PokéLLMon: A Human-Parity Agent for Pokémon Battles with Large Language ModelsTOIT 2025 [paper][code] #competition#training
[2024/02] Agent-Pro: Learning to Evolve via Policy-Level Reflection and OptimizationACL 2024 [paper][code] #competition#training
[2024/01] PokerGPT: An End-to-End Lightweight Solver for Multi-Player Texas Hold'em via Large Language ModelarXiv [paper] #competition#training
[2024/01] SwarmBrain: Embodied agent for real-time strategy game StarCraft II via large language modelsarXiv [paper] #competition#training
[2023/12] Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization ApproachNeurIPS 2024 poster [paper][code] #competition#planning
[2023/09] Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4COLM 2024 [paper] #competition#planning#memory#prompting
[2023/08] Are ChatGPT and GPT-4 Good Poker Players?--A Pre-Flop AnalysisarXiv [paper] #competition
[2023/06] ChessGPT: Bridging Policy Learning and Language ModelingNeurIPS 2023 [paper][code] #competition
[2022/10] Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskICLR 2023 [paper] #competition
cooperation
[2026/04] Don't Make the LLM Read the Graph: Make the Graph ThinkarXiv [paper] #cooperation#multi-agent
[2026/03] On the Strengths and Weaknesses of Data for Open-set Embodied AssistancearXiv [paper] #cooperation#training
[2025/10] LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration GamearXiv [paper] #cooperation#multi-agent
[2025/06] PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language ModelsarXiv [paper] #cooperation#planning#multi-agent
[2024/05] Towards Efficient LLM Grounding for Embodied Multi-Agent CollaborationACL 2024 [paper][code] #cooperation#planning#multi-agent#training
[2024/03] Can LLM-Augmented Autonomous Agents Cooperate?, An Evaluation of Their Cooperative Capabilities through Melting PotIEEE Transactions on Artificial Intelligence 2024 [paper] #cooperation#multi-agent
[2024/03] ProAgent: Building Proactive Cooperative Agents with Large Language ModelsAAAI 2024 [paper] #cooperation
[2024/02] S-Agents: Self-organizing Agents in Open-ended EnvironmentsarXiv [paper] #cooperation
[2023/12] LLM-Powered Hierarchical Language Agent for Real-time Human-AI CoordinationAAMAS 2023 [paper] #cooperation
[2023/10] Evaluating Multi-agent Coordination Abilities in Large Language ModelsNAACL 2023 [paper] #cooperation#planning#multi-agent#prompting
[2023/10] Theory of Mind for Multi-Agent Collaboration via Large Language ModelsEMNLP 2023 [paper][code] #cooperation#planning#multi-agent#training
[2026/04] LLM-Agent-based Social Simulation for Attitude DiffusionarXiv [paper] #sim-social#memory
[2026/04] Restoring Heterogeneity in LLM-based Social Simulation: An Audience Segmentation ApproacharXiv [paper] #sim-social#role-play
[2026/04] SLALOM: Simulation Lifecycle Analysis via Longitudinal Observation Metrics for Social SimulationarXiv [paper] #sim-social
[2026/04] RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing AgentsarXiv [paper] #sim-social#planning
[2026/04] Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing AgentsarXiv [paper] #sim-social
[2026/04] Auditing Support Strategies in LLMs through Grounded Multi-Turn Social SimulationarXiv [paper] #sim-social
[2026/03] PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy OptimizationWWW 2026 [paper] #sim-social#training
[2026/03] Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationWWW 2026 [paper][code] #sim-social#multi-agent
[2026/02] AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent ProfilesarXiv [paper] #sim-social#planning
[2026/02] Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent CommunityarXiv [paper] #sim-social
[2026/02] Does Socialization Emerge in AI Agent Society? A Case Study of MoltbookProceedings of the ACM Conference on AI and Agentic Systems 2026 [paper] #sim-social
[2026/02] The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI SocietiesarXiv [paper] #sim-social#multi-agent#self-improvement
[2026/01] When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered AgentsarXiv [paper] #sim-social#multi-agent
[2026/01] MARO: Learning Stronger Reasoning from Social InteractionarXiv [paper] #sim-social#multi-agent
[2026/01] HumanLLM: Towards Personalized Understanding and Simulation of Human NaturearXiv [paper] #sim-social#training
[2025/12] EZYer: A simulacrum of high school with generative agentarXiv [paper] #sim-social#memory
[2025/12] Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMsarXiv [paper] #sim-social#multi-agent
[2025/10] Multimodal Safety Evaluation in Generative Agent Social SimulationsarXiv [paper] #sim-social#planning#vlm
[2025/10] Alita-G: Self-Evolving Generative Agent for Agent GenerationarXiv [paper] #sim-social#memory#self-improvement
[2025/10] Doing Things with Words: Rethinking Theory of Mind Simulation in Large Language ModelsComputational Linguistics 2025 [paper] #sim-social
[2025/10] Social Simulations with Large Language Model Risk Utopian IllusionarXiv [paper] #sim-social#multi-agent
[2025/10] Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information OperationsWWW 2025 [paper] #sim-social
[2025/09] Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd SimulationsEMNLP 2025 [paper] #sim-social#role-play
[2025/09] The Emergence of Altruism in Large-Language-Model Agents SocietyarXiv [paper] #sim-social
[2025/07] LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative SimulacraarXiv [paper][code] #sim-social#multi-agent#training#role-play
[2025/07] Too Human to Model:The Uncanny Valley of LLMs in Social Simulation -- When Generative Language Agents Misalign with Modelling PrinciplesarXiv [paper] #sim-social
[2025/07] Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel PredictionsAnnual Meeting of the Cognitive Science Society 2025 [paper] #sim-social#role-play
[2025/06] Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based ModelingKDD 2025 [paper] #sim-social#training
[2025/06] IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent EnvironmentEMNLP 2025 [paper] #sim-social#multi-agent
[2025/06] AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System NeedarXiv [paper][code] #sim-social#planning#multi-agent
[2025/06] Infected Smallville: How Disease Threat Shapes Sociality in LLM AgentsarXiv [paper] #sim-social
[2025/05] EcoLANG: Efficient and Effective Agent Communication Language Induction for Social SimulationEMNLP 2025 [paper] #sim-social#role-play
[2025/04] SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulationNAACL 2025 [paper] #sim-social#planning
[2025/04] BookWorld: From Novels to Interactive Agent Societies for Creative Story GenerationarXiv [paper] #sim-social#multi-agent#generation
[2025/04] MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model FrameworkNeurIPS 2025 poster [paper] #sim-social#planning#training
[2025/03] The Impact of Big Five Personality Traits on AI Agent Decision-Making in Public Spaces: A Social Simulation StudyarXiv [paper] #sim-social
[2025/02] Investigating and Extending Homans' Social Exchange Theory with Large Language Model based AgentsACL 2025 [paper][code] #sim-social
[2025/01] Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical InteractionsarXiv [paper] #sim-social
[2024/10] Project Sid: Many-agent simulations toward AI civilizationarXiv [paper] #sim-social
[2024/06] Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract TheoryarXiv [paper] #sim-social#multi-agent
[2024/05] Agent hospital: A simulacrum of hospital with evolvable medical agentsarXiv [paper] #sim-social#planning
[2024/03] SOTOPIA-$\pi$: Interactive Learning of Socially Intelligent Language AgentsACL 2024 [paper][code] #sim-social#training
[2023/10] Lyfe Agents: Generative agents for low-cost real-time social interactionsarXiv [paper] #sim-social#memory#multi-agent#self-improvement
[2023/10] SOTOPIA: Interactive Evaluation for Social Intelligence in Language AgentsICLR 2023 [paper][code] #sim-social#role-play
[2023/08] AgentSims: An Open-Source Sandbox for Large Language Model EvaluationarXiv [paper][code] #sim-social#planning#tool-use
[2023/07] S3: Social-network Simulation System with Large Language Model-Empowered AgentsarXiv [paper] #sim-social#prompting
[2023/04] Generative Agents: Interactive Simulacra of Human BehaviorUIST 2023 [paper][code] #sim-social
sim-embodied
[2026/05] Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State UncertaintyarXiv [paper] #sim-embodied#training
[2026/05] Probing Embodied LLMs: When Higher Observation Fidelity Hurts Problem SolvingarXiv [paper] #sim-embodied
[2026/02] To Move or Not to Move: Constraint-based Planning Enables Zero-Shot Generalization for Interactive NavigationarXiv [paper] #sim-embodied#planning#prompting#vlm
[2025/12] Emergence: Overcoming Privileged Information Bias in Asymmetric Embodied Agents via Active QueryingarXiv [paper] #sim-embodied
[2025/12] ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement LearningarXiv [paper] #sim-embodied#planning#memory#training
[2025/12] HELP: Hierarchical Embodied Language Planner for Household TasksarXiv [paper] #sim-embodied#planning
[2025/11] DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent CollaborationarXiv [paper] #sim-embodied#planning#multi-agent#world-model
[2025/11] MADRA: Multi-Agent Debate for Risk-Aware Embodied PlanningarXiv [paper] #sim-embodied#planning#multi-agent#prompting
[2024/09] Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-Based Decision-Making SystemsICLR 2025 Poster [paper] #sim-embodied#training
[2024/09] BadRobot: Jailbreaking Embodied LLM Agents in the Physical WorldICLR 2025 Poster [paper] #sim-embodied#planning#vlm
[2024/09] GameGen-X: Interactive Open-world Game Video GenerationICLR 2025 Poster [paper][code] #sim-embodied#vlm
[2024/01] True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement LearningarXiv [paper][code] #sim-embodied#training
[2025/06] Automated Skill Discovery for Language Agents through Exploration and Iterative FeedbackarXiv [paper] #crafter#self-improvement
[2025/02] LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative PlanningarXiv [paper] #crafter#planning#memory#multi-agent
[2024/10] Mars: Situated Inductive Reasoning in an Open-World EnvironmentNeurIPS 2024 [paper] #crafter
[2024/07] Enhancing Agent Learning through World Dynamics ModelingEMNLP 2024 [paper] #crafter
[2024/04] AgentKit: Flow Engineering with Graphs, not CodingarXiv [paper][code] #crafter#planning
[2024/04] World Models with Hints of Large Language Models for Goal AchievingNAACL 2024 [paper] #crafter#training#vlm
[2024/03] EnvGen: Generating and Adapting Environments via LLMs for Training Embodied AgentsCOLM [paper] #crafter#training
[2024/03] AgentKit: Structured LLM Reasoning with Dynamic GraphsCOLM [paper] #crafter#planning
[2023/09] AdaRefiner: Refining Decisions of Language Models with Adaptive FeedbackNAACL 2023 [paper] #crafter#training
[2023/06] OMNI: Open-endedness via Models of human Notions of InterestingnessarXiv [paper][code] #crafter
[2023/05] SPRING: Studying Papers and Reasoning to play GamesNeurIPS 2023 [paper] #crafter
[2023/02] Guiding Pretraining in Reinforcement Learning with Large Language ModelsICML 2023 [paper] #crafter#training
action
[2026/05] Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement LearningarXiv [paper] #action#training#vlm
[2026/05] ANO: A Principled Approach to Robust Policy OptimizationarXiv [paper] #action#training
[2026/05] Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplayarXiv [paper] #action#planning#training#vlm
[2026/04] Playing DOOM with 1.3M Parameters: Specialized Small Models vs Large Language Models for Real-Time Game ControlarXiv [paper] #action
[2026/04] PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language ModelsarXiv [paper] #action#planning#vlm
[2025/05] Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme OnearXiv [paper] #action#training
[2025/05] Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsNeurIPS 2025 poster [paper] #action#training
[2025/05] LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language ModelsNeurIPS 2025 spotlight [paper][code] #action#training
[2025/05] PoE-World: Compositional World Modeling with Products of Programmatic ExpertsNeurIPS 2025 spotlight [paper] #action#planning#world-model
[2025/04] Better Decisions through the Right Causal World ModelarXiv [paper] #action#planning#world-model#training
[2025/01] LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal DemonstrationsICML 2025 poster [paper] #action#planning#training
[2024/10] Unbounded: A Generative Infinite Game of Character Life SimulationICLR 2024 [paper] #action
[2024/09] Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study CasearXiv [paper][code] #action#planning#training
[2024/09] BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesICLR 2025 Poster [paper] #action#planning#training#vlm
[2024/08] Atari-GPT: Investigating the Capabilities of Multimodal Large Language Models as Low-Level Policies for Atari GamesarXiv [paper] #action#planning#training
[2024/07] Baba Is AI: Break the Rules to Beat the BenchmarkICML 2024 [paper] #action#vlm
[2024/03] Will GPT-4 Run DOOM?IEEE Transactions on Games 2024 [paper][code] #action#planning#training
[2024/03] Evaluate LLMs in Real Time with Street Fighter IIIGitHub [paper][code] #action
[2023/02] Grounding Large Language Models in Interactive Environments with Online Reinforcement LearningICML 2023 [paper][code] #action#training
video-adventure
[2024/03] Cradle: Empowering Foundation Agents Towards General Computer ControlICML 2024 [paper][code] #video-adventure#planning
[2024/03] Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents2024 IEEE Conference on Games (CoG) 2024 [paper][code] #video-adventure#planning#prompting
[2025/05] lmgame-Bench: How Good are LLMs at Playing Games?."ICLR 2026 Poster [paper][code] #benchmark#planning#training
[2025/05] Is Your LLM Really Mastering the Concept? A Multi-Agent BenchmarkarXiv [paper][code] #benchmark#multi-agent
other
[2026/04] GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game AgentsarXiv [paper] #planning#vlm
[2026/04] Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language ModelsarXiv [paper] #tool-use#training#self-improvement
[2026/04] From World-Gen to Quest-Line: A Dependency-Driven Prompt Pipeline for Coherent RPG GenerationarXiv [paper] #planning#generation
[2026/03] Sensi: Learn One Thing at a Time -- Curriculum-Based Test-Time Learning for LLM Game AgentsarXiv [paper]
[2026/01] SNAP: A Plan-Driven Framework for Controllable Interactive Narrative GenerationarXiv [paper] #planning#generation
[2025/10] ROBOPSY PL[AI]: Using Role-Play to Investigate how LLMs Present Collective MemoryarXiv [paper] #role-play
[2025/08] All Stories Are One Story: Emotional Arc Guided Procedural Game Level GenerationarXiv [paper] #generation
[2025/08] CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMsarXiv [paper] #memory
[2025/06] The Decrypto Benchmark for Multi-Agent Reasoning and Theory of MindarXiv [paper] #multi-agent#training
[2025/04] PAYADOR: A Minimalist Approach to Grounding Language Models on Structured Data for Interactive Storytelling and Role-playing GamesarXiv [paper]
[2025/04] Enhancing Player Enjoyment with a Two-Tier DRL and LLM-Based Agent System for Fighting GamesarXiv [paper] #training
[2025/03] Playing games with Large language models: Randomness and strategyarXiv [paper] #multi-agent
[2025/03] Collaborative Storytelling and LLM: A Linguistic Analysis of Automatically-Generated Role-Playing Game SessionsarXiv [paper]
[2025/03] Cultivating Game Sense for Yourself: Making VLMs Gaming ExpertsarXiv [paper] #vlm
[2025/02] RPGBENCH: Evaluating Large Language Models as Role-Playing Game EnginesarXiv [paper]
[2025/02] Hybrid Voting-Based Task Assignment in Role-Playing GamesarXiv [paper] #planning
[2024/07] What if Red Can Talk? Dynamic Dialogue Generation Using Large Language Models.arXiv [paper] #generation
[2023/10] Language as reality: a co-creative storytelling game experience in 1001 nights using generative AI.AAAI 2023 [paper] #generation
By Mechanism
The same papers as above, re-grouped by agent design. A paper with multiple Mechanism tags appears in each relevant section.
planning
[2026/06] Can LLM Agents Sustain Long-Horizon Organizational Dynamics?arXiv [paper] #sim-social#planning#multi-agent
[2026/06] From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM AgentsarXiv [paper] #text-adventure#planning
[2026/05] From History to State: Constant-Context Skill Learning for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/05] PriorZero: Bridging Language Priors and World Models for Decision MakingarXiv [paper][code] #text-adventure#planning#world-model#training
[2026/05] Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplayarXiv [paper] #action#planning#training#vlm
[2026/03] Reward Prediction with Factorized World StatesarXiv [paper] #text-adventure#planning#prompting
[2026/03] Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy ScenariosarXiv [paper] #competition#planning#memory#multi-agent
[2026/02] Active Epistemic Control for Query-Efficient Verified PlanningarXiv [paper] #text-adventure#planning
[2026/02] AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent ProfilesarXiv [paper] #sim-social#planning
[2026/02] Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/02] TAPE: Tool-Guided Adaptive Planning and Constrained Execution in Language Model AgentsarXiv [paper] #text-adventure#planning
[2026/02] To Move or Not to Move: Constraint-based Planning Enables Zero-Shot Generalization for Interactive NavigationarXiv [paper] #sim-embodied#planning#prompting#vlm
[2026/02] CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesarXiv [paper] #text-adventure#planning#world-model#training
[2026/01] SNAP: A Plan-Driven Framework for Controllable Interactive Narrative GenerationarXiv [paper] #planning#generation
[2026/01] Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2025/12] ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement LearningarXiv [paper] #sim-embodied#planning#memory#training
[2025/12] HELP: Hierarchical Embodied Language Planner for Household TasksarXiv [paper] #sim-embodied#planning
[2025/11] DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent CollaborationarXiv [paper] #sim-embodied#planning#multi-agent#world-model
[2025/11] MADRA: Multi-Agent Debate for Risk-Aware Embodied PlanningarXiv [paper] #sim-embodied#planning#multi-agent#prompting
[2025/10] Fine-tuning with RAG for Improving LLM Learning of New SkillsarXiv [paper] #text-adventure#planning#memory#training
[2025/10] Constrained Natural Language Action Planning for Resilient Embodied SystemsarXiv [paper] #text-adventure#planning#prompting
[2025/10] The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with SchemasarXiv [paper] #text-adventure#planning
[2025/10] Multimodal Safety Evaluation in Generative Agent Social SimulationsarXiv [paper] #sim-social#planning#vlm
[2025/10] Graph-Enhanced Policy Optimization in LLM Agent TrainingarXiv [paper] #text-adventure#planning#training
[2025/10] Out-of-distribution Tests Reveal Compositionality in Chess TransformersarXiv [paper] #competition#planning
[2025/09] Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM AgentarXiv [paper] #text-adventure#planning#multi-agent#self-improvement
[2025/09] Code Driven Planning with Domain-Adaptive CriticarXiv [paper] #text-adventure#planning#self-improvement
[2025/09] Goal-Guided Efficient Exploration via Large Language Model in Reinforcement LearningarXiv [paper] #crafter#planning#training
[2025/09] Natural Language PDDL (NL-PDDL) for Open-world Goal-oriented Commonsense Regression Planning in Embodied AIICLR 2026 Poster [paper] #text-adventure#planning#vlm
[2025/09] Experience-based Knowledge Correction for Robust Planning in MinecraftICLR 2026 Poster [paper] #minecraft#planning
[2025/09] SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningICLR 2026 Poster [paper] #competition#planning#multi-agent#training
[2025/09] DreamPhase: Offline Imagination and Uncertainty-Guided Planning for Large-Language-Model AgentsICLR 2026 Poster [paper] #text-adventure#planning#world-model
[2025/08] CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative TasksFindings of EMNLP 2025 [paper] #minecraft#planning#multi-agent
[2025/08] Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World SuccessProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#planning#training#vlm
[2025/08] SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making TasksarXiv [paper] #competition#planning#training#self-improvement
[2025/07] CoEx -- Co-evolving World-model and ExplorationEMNLP 2025 [paper] #text-adventure#planning#world-model
[2025/06] Optimus-3: Towards Generalist Multimodal Minecraft Agents with Scalable Task ExpertsarXiv [paper][code] #minecraft#planning
[2025/06] Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead SearcharXiv [paper] #text-adventure#planning#world-model#training
[2025/06] OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic ReflectionsarXiv [paper] #text-adventure#planning#training#self-improvement
[2025/06] AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System NeedarXiv [paper][code] #sim-social#planning#multi-agent
[2025/06] PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language ModelsarXiv [paper] #cooperation#planning#multi-agent
[2025/05] Don’t Just Follow MLLM Plans: Robust and Efficient Planning for Open-World AgentsarXiv [paper] #minecraft#planning#vlm
[2025/05] Knowledge Retrieval in LLM Gaming: A Shift from Entity-Centric to Goal-Oriented GraphsKnowledge-Based Systems 2025 [paper] #minecraft#planning#memory
[2025/05] lmgame-Bench: How Good are LLMs at Playing Games?."ICLR 2026 Poster [paper][code] #benchmark#planning#training
[2025/05] LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMsarXiv [paper][code] #text-adventure#planning
[2025/05] Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningICML 2025 poster [paper] #text-adventure#planning#training
[2025/05] ActiveVOO: Value of Observation Guided Active Knowledge Acquisition for Open-World Embodied Lifted Regression PlanningNeurIPS 2025 poster [paper] #text-adventure#planning#prompting#vlm
[2025/05] Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RLNeurIPS 2025 poster [paper] #communication#planning#tool-use#training
[2025/05] PoE-World: Compositional World Modeling with Products of Programmatic ExpertsNeurIPS 2025 spotlight [paper] #action#planning#world-model
[2025/05] WALL-E: World Alignment by NeuroSymbolic Learning improves World Model-based LLM AgentsNeurIPS 2025 poster [paper] #minecraft#planning#world-model#prompting
[2025/04] WALL-E 2.0: World Alignment by NeuroSymbolic Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model#prompting
[2025/04] Better Decisions through the Right Causal World ModelarXiv [paper] #action#planning#world-model#training
[2025/04] ZeroSumEval: Scaling LLM Evaluation with Inter-Model CompetitionarXiv [paper][code] #competition#planning
[2025/04] SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulationNAACL 2025 [paper] #sim-social#planning
[2025/04] Monte Carlo Planning with Large Language Model for Text-Based Game AgentsICLR 2025 Poster [paper] #text-adventure#planning#training
[2025/04] LLM-PySC2: Starcraft II learning environment for Large Language ModelsNeurIPS 2025 poster [paper] #competition#planning#multi-agent#vlm
[2025/04] MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model FrameworkNeurIPS 2025 poster [paper] #sim-social#planning#training
[2025/03] Parallelized Planning-Acting for Efficient LLM-based Multi-Agent SystemsarXiv [paper] #minecraft#planning#multi-agent#tool-use
[2025/03] GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow NetworksCVPR 2025 [paper] #text-adventure#planning#training#vlm
[2025/03] Plancraft: an evaluation dataset for planning with LLM agentsCOLM 2025 [paper] #minecraft#planning#memory#tool-use
[2025/02] Optimus-2: Multimodal World Model for Open-World Minecraft AgentsCVPR 2025 [paper][code] #minecraft#planning#world-model#vlm
[2025/02] LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative PlanningarXiv [paper] #crafter#planning#memory#multi-agent
[2025/02] Hybrid Voting-Based Task Assignment in Role-Playing GamesarXiv [paper] #planning
[2025/02] Implicit Search via Discrete Diffusion: A Study on ChessICLR 2025 Poster [paper][code] #competition#planning
[2025/01] POKERBENCH: Training Large Language Models to become Professional Poker PlayersAAAI 2025 [paper] #competition#planning#training
[2025/01] Mastering Board Games by External and Internal Planning with Language ModelsICML 2025 spotlightposter [paper] #competition#planning
[2025/01] LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal DemonstrationsICML 2025 poster [paper] #action#planning#training
[2025/01] Language Models as Implicit Tree SearchICML 2025 poster [paper] #competition#planning#training
[2024/10] WALL-E: World Alignment by Rule Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model
[2024/10] ADAM: An Embodied Causal Agent in Open-World EnvironmentsICLR 2025 [paper][code] #minecraft#planning
[2024/10] PokéChamp: An Expert-level Minimax Language AgentICML 2025 [paper][code] #competition#planning
[2024/09] Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study CasearXiv [paper][code] #action#planning#training
[2024/09] Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchICLR 2025 Poster [paper] #communication#planning#multi-agent#training
[2024/09] Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackICLR 2025 Poster [paper] #text-adventure#planning#training#self-improvement
[2024/09] BadRobot: Jailbreaking Embodied LLM Agents in the Physical WorldICLR 2025 Poster [paper] #sim-embodied#planning#vlm
[2024/09] BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesICLR 2025 Poster [paper] #action#planning#training#vlm
[2024/08] Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) 2024 [paper] #competition#planning#multi-agent#training
[2024/08] Atari-GPT: Investigating the Capabilities of Multimodal Large Language Models as Low-Level Policies for Atari GamesarXiv [paper] #action#planning#training
[2024/07] Arigraph: Learning knowledge graph world models with episodic memory for llm agentsIJCAI 2024 [paper] #text-adventure#planning#memory
[2024/07] Odyssey: Empowering Agents with Open-World Skills.IJCAI 2024 [paper][code] #minecraft#planning#tool-use
[2024/06] VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in MinecraftFindings of ACL 2024 [paper][code] #minecraft#planning#multi-agent
[2024/05] Agent Planning with World Knowledge ModelNeurIPS 2024 [paper][code] #text-adventure#planning#memory#world-model
[2024/05] Agent hospital: A simulacrum of hospital with evolvable medical agentsarXiv [paper] #sim-social#planning
[2024/05] Towards Efficient LLM Grounding for Embodied Multi-Agent CollaborationACL 2024 [paper][code] #cooperation#planning#multi-agent#training
[2024/05] Reflective Multi-Agent Collaboration based on Large Language ModelsNeurIPS 2024 poster [paper] #competition#planning#multi-agent#training
[2024/05] AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental LearningNeurIPS 2024 poster [paper][code] #text-adventure#planning
[2024/05] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 poster [paper] #communication#planning#multi-agent#self-improvement
[2024/04] ReAct Meets ActRe: When Language Agents Enjoy Training Data AutonomyarXiv [paper] #text-adventure#planning#training#self-improvement
[2024/04] AgentKit: Flow Engineering with Graphs, not CodingarXiv [paper][code] #crafter#planning
[2024/03] KnowAgent: Knowledge-Augmented Planning for LLM-Based AgentsNAACL 2024 [paper][code] #text-adventure#planning
[2024/03] Cradle: Empowering Foundation Agents Towards General Computer ControlICML 2024 [paper][code] #video-adventure#planning
[2024/03] Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents2024 IEEE Conference on Games (CoG) 2024 [paper][code] #video-adventure#planning#prompting
[2024/03] Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation (HAS)ICLR 2024 Workshop [paper] #minecraft#planning#multi-agent#vlm
[2024/03] Embodied LLM Agents Learn to Cooperate in Organized TeamsIEEE Transactions on Computational Social Systems 2024 [paper] #competition#planning#multi-agent
[2024/03] Will GPT-4 Run DOOM?IEEE Transactions on Games 2024 [paper][code] #action#planning#training
[2024/03] ReAct Meets ActRe: Autonomous Annotation of Agent Trajectories for Contrastive Self-TrainingCOLM [paper] #text-adventure#planning#training#self-improvement
[2024/03] StateFlow: Enhancing LLM Task-Solving through State-Driven WorkflowsCOLM [paper] #text-adventure#planning#self-improvement
[2024/03] AgentKit: Structured LLM Reasoning with Dynamic GraphsCOLM [paper] #crafter#planning
[2024/02] Empowering Large Language Model Agents through Action LearningCOLM [paper][code] #text-adventure#planning#self-improvement
[2023/12] MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active PerceptionCVPR 2024 [paper][code] #minecraft#planning#vlm
[2023/12] Creative Agents: Empowering Agents with Imagination for Creative TasksUAI 2023 [paper][code] #minecraft#planning
[2023/12] Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization ApproachNeurIPS 2024 poster [paper][code] #competition#planning
[2023/11] ADaPT: As-Needed Decomposition and Planning with Language ModelsNAACL 2023 [paper][code] #text-adventure#planning
[2023/11] JARVIS-1: Open-world Multi-task Agents with Memory-Augmented Multimodal Language ModelsTPAMI 2023 [paper][code] #minecraft#planning#memory
[2023/10] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language ModelsICML 2024 [paper][code] #text-adventure#planning#training
[2023/10] LLaMA Rider: Spurring Large Language Models to Explore the Open WorldNAACL 2023 [paper][code] #minecraft#planning#training
[2023/08] AgentSims: An Open-Source Sandbox for Large Language Model EvaluationarXiv [paper][code] #sim-social#planning#tool-use
[2023/07] Building Cooperative Embodied Agents Modularly with Large Language ModelsICLR 2024 [paper][code] #cooperation#planning#multi-agent#training
[2023/05] Language Models Meet World Models: Embodied Experiences Enhance Language ModelsNeurIPS 2023 [paper][code] #sim-embodied#planning#world-model
[2023/05] SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksNeurIPS 2023 [paper][code] #text-adventure#planning
[2023/03] Plan4MC: Skill Reinforcement Learning and Planning for Open-World Minecraft TasksFMDM@NeurIPS2023 [paper][code] #minecraft#planning#training
[2023/02] Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task AgentsNeurIPS 2023 [paper][code] #minecraft#planning#prompting
[2022/12] LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsICCV 2023 [paper] #sim-embodied#planning#prompting
[2022/10] ReAct: Synergizing Reasoning and Acting in Language ModelsICLR 2023 [paper][code] #text-adventure#planning#training
[2020/10] ALFWorld: Aligning Text and Embodied Environments for Interactive LearningICLR 2021 [paper][code] #text-adventure#planning
memory
[2026/06] SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent TrainingarXiv [paper][code] #text-adventure#memory#training
[2026/06] SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at ScalearXiv [paper] #text-adventure#memory#self-improvement
[2026/05] Belief Memory: Agent Memory Under Partial ObservabilityarXiv [paper] #text-adventure#memory
[2026/05] GASim: A Graph-Accelerated Hybrid Framework for Social SimulationarXiv [paper][code] #sim-social#memory#multi-agent
[2026/05] ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic ProfilesarXiv [paper] #sim-social#memory#multi-agent
[2026/05] Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM AgentsarXiv [paper] #text-adventure#memory#tool-use
[2026/05] SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill GraphsarXiv [paper] #text-adventure#memory#tool-use#training
[2026/05] Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM AgentsarXiv [paper] #text-adventure#memory
[2026/03] RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic FeedbackarXiv [paper] #text-adventure#memory#training#self-improvement
[2026/03] Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy ScenariosarXiv [paper] #competition#planning#memory#multi-agent
[2025/12] EZYer: A simulacrum of high school with generative agentarXiv [paper] #sim-social#memory
[2025/12] ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement LearningarXiv [paper] #sim-embodied#planning#memory#training
[2025/12] Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive RefinementProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#memory
[2025/11] Knowledge Graph-enhanced Large Language Model for Incremental Game PlayTestingIEICE Transactions on Information and Systems 2025 [paper] #minecraft#memory
[2025/10] Fine-tuning with RAG for Improving LLM Learning of New SkillsarXiv [paper] #text-adventure#planning#memory#training
[2025/10] Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy GamesarXiv [paper] #competition#memory
[2025/10] Alita-G: Self-Evolving Generative Agent for Agent GenerationarXiv [paper] #sim-social#memory#self-improvement
[2025/09] World Model Implanting for Test-time Adaptation of Embodied AgentsICML 2025 poster [paper] #text-adventure#memory#world-model#prompting
[2025/09] Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationICLR 2026 Poster [paper] #text-adventure#memory#training
[2025/08] Vistawise: Building Cost-effective Agent with Cross-modal Knowledge Graph for MinecraftEMNLP 2025 [paper] #minecraft#memory#tool-use
[2025/08] CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMsarXiv [paper] #memory
[2025/06] StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi TurnsarXiv [paper] #text-adventure#memory
[2025/05] Knowledge Retrieval in LLM Gaming: A Shift from Entity-Centric to Goal-Oriented GraphsKnowledge-Based Systems 2025 [paper] #minecraft#planning#memory
[2025/03] Plancraft: an evaluation dataset for planning with LLM agentsCOLM 2025 [paper] #minecraft#planning#memory#tool-use
[2025/02] LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative PlanningarXiv [paper] #crafter#planning#memory#multi-agent
[2024/11] MrSteve: Instruction-Following Agents with What-Where-When MemoryICLR 2025 [paper][code] #minecraft#memory
[2024/09] MrSteve: Instruction-Following Agents in Minecraft with What-Where-When MemoryICLR 2025 Poster [paper] #minecraft#memory
[2024/07] Arigraph: Learning knowledge graph world models with episodic memory for llm agentsIJCAI 2024 [paper] #text-adventure#planning#memory
[2024/05] Agent Planning with World Knowledge ModelNeurIPS 2024 [paper][code] #text-adventure#planning#memory#world-model
[2023/11] JARVIS-1: Open-world Multi-task Agents with Memory-Augmented Multimodal Language ModelsTPAMI 2023 [paper][code] #minecraft#planning#memory
[2023/11] See and Think: Embodied Agent in Virtual EnvironmentECCV 2023 [paper][code] #minecraft#memory
[2023/10] Lyfe Agents: Generative agents for low-cost real-time social interactionsarXiv [paper] #sim-social#memory#multi-agent#self-improvement
[2023/09] Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4COLM 2024 [paper] #competition#planning#memory#prompting
[2023/09] Exploring Large Language Models for Communication Games: An Empirical Study on WerewolfarXiv [paper] #communication#memory
multi-agent
[2026/06] Can LLM Agents Sustain Long-Horizon Organizational Dynamics?arXiv [paper] #sim-social#planning#multi-agent
[2026/06] Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social SimulationarXiv [paper] #sim-social#multi-agent
[2026/06] SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language ModelsarXiv [paper] #competition#multi-agent
[2026/05] GASim: A Graph-Accelerated Hybrid Framework for Social SimulationarXiv [paper][code] #sim-social#memory#multi-agent
[2026/05] ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic ProfilesarXiv [paper] #sim-social#memory#multi-agent
[2026/05] GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM CollectivesarXiv [paper] #competition#multi-agent#prompting
[2026/05] ALSO: Adversarial Online Strategy Optimization for Social AgentsarXiv [paper] #sim-social#multi-agent#training
[2026/05] MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMsarXiv [paper] #communication#multi-agent
[2026/05] PatchBoard: Schema-Grounded State Mutation for Reliable and Auditable LLM Multi-Agent CollaborationarXiv [paper] #text-adventure#multi-agent
[2026/04] MARL-GPT: Foundation Model for Multi-Agent Reinforcement LearningProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #competition#multi-agent#training
[2026/04] Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent WorkflowsarXiv [paper] #text-adventure#planning#memory#multi-agent
[2026/04] Don't Make the LLM Read the Graph: Make the Graph ThinkarXiv [paper] #cooperation#multi-agent
[2026/04] Gated Coordination for Efficient Multi-Agent Collaboration in Minecraft GamearXiv [paper] #minecraft#memory#multi-agent#vlm
[2026/03] How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game EnvironmentarXiv [paper] #text-adventure#multi-agent#training
[2026/03] Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy ScenariosarXiv [paper] #competition#planning#memory#multi-agent
[2026/03] Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationWWW 2026 [paper][code] #sim-social#multi-agent
[2026/03] Deception and Communication in Autonomous Multi-Agent Systems: An Experimental Study with Among UsProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #communication#multi-agent
[2026/02] The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI SocietiesarXiv [paper] #sim-social#multi-agent#self-improvement
[2026/01] When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered AgentsarXiv [paper] #sim-social#multi-agent
[2026/01] Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction GamesInternational Conference on Agents 2027 [paper] #communication#multi-agent
[2026/01] MARO: Learning Stronger Reasoning from Social InteractionarXiv [paper] #sim-social#multi-agent
[2025/12] WOLF: Werewolf-based Observations for LLM Deception and FalsehoodsarXiv [paper] #communication#multi-agent
[2025/12] Robust Agents in Open-Ended WorldsarXiv [paper] #action#multi-agent#training#generation
[2025/12] Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMsarXiv [paper] #sim-social#multi-agent
[2025/11] DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent CollaborationarXiv [paper] #sim-embodied#planning#multi-agent#world-model
[2025/11] Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal ReasoningarXiv [paper][code] #communication#multi-agent
[2025/11] MADRA: Multi-Agent Debate for Risk-Aware Embodied PlanningarXiv [paper] #sim-embodied#planning#multi-agent#prompting
[2025/10] LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration GamearXiv [paper] #cooperation#multi-agent
[2025/10] Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned StrategiesarXiv [paper] #communication#multi-agent#self-improvement
[2025/10] Social Simulations with Large Language Model Risk Utopian IllusionarXiv [paper] #sim-social#multi-agent
[2025/09] Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM AgentarXiv [paper] #text-adventure#planning#multi-agent#self-improvement
[2025/09] PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments2025 IEEE Conference on Games (CoG) 2025 [paper] #minecraft#multi-agent#self-improvement
[2025/09] HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-MakingarXiv [paper] #competition#multi-agent#training
[2025/09] SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningICLR 2026 Poster [paper] #competition#planning#multi-agent#training
[2025/08] CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative TasksFindings of EMNLP 2025 [paper] #minecraft#planning#multi-agent
[2025/08] A Multi-Agent Pokemon Tournament for Evaluating Strategic Reasoning of Large Language ModelsarXiv [paper] #action#multi-agent
[2025/07] LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative SimulacraarXiv [paper][code] #sim-social#multi-agent#training#role-play
[2025/07] CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language GamesACL 2025 [paper][code] #communication#multi-agent
[2025/06] IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent EnvironmentEMNLP 2025 [paper] #sim-social#multi-agent
[2025/06] WereWolf-Plus: An Update of Werewolf Game setting Based on DSGBencharXiv [paper][code] #communication#multi-agent
[2025/06] The Decrypto Benchmark for Multi-Agent Reasoning and Theory of MindarXiv [paper] #multi-agent#training
[2025/06] AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System NeedarXiv [paper][code] #sim-social#planning#multi-agent
[2025/06] PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language ModelsarXiv [paper] #cooperation#planning#multi-agent
[2025/05] Is Your LLM Really Mastering the Concept? A Multi-Agent BenchmarkarXiv [paper][code] #benchmark#multi-agent
[2025/04] Collaborating Action by Action: A Multi-agent LLM Framework for Embodied ReasoningarXiv [paper][code] #minecraft#multi-agent#training
[2025/04] BookWorld: From Novels to Interactive Agent Societies for Creative Story GenerationarXiv [paper] #sim-social#multi-agent#generation
[2025/04] LLM-PySC2: Starcraft II learning environment for Large Language ModelsNeurIPS 2025 poster [paper] #competition#planning#multi-agent#vlm
[2025/03] Parallelized Planning-Acting for Efficient LLM-based Multi-Agent SystemsarXiv [paper] #minecraft#planning#multi-agent#tool-use
[2025/03] Playing games with Large language models: Randomness and strategyarXiv [paper] #multi-agent
[2025/03] Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic ReasoningCOLM 2025 [paper] #competition#multi-agent#training
[2025/02] LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative PlanningarXiv [paper] #crafter#planning#memory#multi-agent
[2024/12] TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in MinecraftarXiv [paper][code] #minecraft#multi-agent#training#vlm
[2024/09] Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchICLR 2025 Poster [paper] #communication#planning#multi-agent#training
[2024/08] Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) 2024 [paper] #competition#planning#multi-agent#training
[2024/06] VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in MinecraftFindings of ACL 2024 [paper][code] #minecraft#planning#multi-agent
[2024/06] Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract TheoryarXiv [paper] #sim-social#multi-agent
[2024/06] PLAYER: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery GamesarXiv [paper] #communication#multi-agent
[2024/05] Towards Efficient LLM Grounding for Embodied Multi-Agent CollaborationACL 2024 [paper][code] #cooperation#planning#multi-agent#training
[2024/05] Reflective Multi-Agent Collaboration based on Large Language ModelsNeurIPS 2024 poster [paper] #competition#planning#multi-agent#training
[2024/05] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 poster [paper] #communication#planning#multi-agent#self-improvement
[2024/03] MineLand: Simulating Large-Scale Multi-Agent Interactions with Limited Multimodal Senses and Physical NeedsarXiv [paper][code] #minecraft#multi-agent#vlm
[2024/03] Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation (HAS)ICLR 2024 Workshop [paper] #minecraft#planning#multi-agent#vlm
[2024/03] Embodied LLM Agents Learn to Cooperate in Organized TeamsIEEE Transactions on Computational Social Systems 2024 [paper] #competition#planning#multi-agent
[2024/03] Can LLM-Augmented Autonomous Agents Cooperate?, An Evaluation of Their Cooperative Capabilities through Melting PotIEEE Transactions on Artificial Intelligence 2024 [paper] #cooperation#multi-agent
[2024/03] Helmsman of the Masses? Evaluate the Opinion Leadership of Large Language Models in the Werewolf GameCOLM [paper] #communication#multi-agent
[2023/12] Cooperation on the Fly: Exploring Language Agents for Ad Hoc Teamwork in the Avalon GamearXiv [paper] #communication#multi-agent
[2023/12] Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery GamesACL 2023 [paper] #communication#multi-agent#prompting
[2023/11] War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World WarsarXiv [paper][code] #communication#multi-agent
[2023/10] Lyfe Agents: Generative agents for low-cost real-time social interactionsarXiv [paper] #sim-social#memory#multi-agent#self-improvement
[2023/10] Evaluating Multi-agent Coordination Abilities in Large Language ModelsNAACL 2023 [paper] #cooperation#planning#multi-agent#prompting
[2023/10] Theory of Mind for Multi-Agent Collaboration via Large Language ModelsEMNLP 2023 [paper][code] #cooperation#planning#multi-agent#training
[2023/10] LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon GameplayEMNLP 2023 [paper] #communication#multi-agent
[2023/10] Leveraging Word Guessing Games to Assess the Intelligence of Large Language ModelsarXiv [paper][code] #communication#multi-agent
[2023/07] Building Cooperative Embodied Agents Modularly with Large Language ModelsICLR 2024 [paper][code] #cooperation#planning#multi-agent#training
world-model
[2026/05] PriorZero: Bridging Language Priors and World Models for Decision MakingarXiv [paper][code] #text-adventure#planning#world-model#training
[2026/02] World Models for Policy Refinement in StarCraft IIarXiv [paper] #competition#world-model#prompting
[2026/02] CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesarXiv [paper] #text-adventure#planning#world-model#training
[2026/02] Reinforcement World Model Learning for LLM-based AgentsarXiv [paper] #text-adventure#world-model
[2025/11] DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent CollaborationarXiv [paper] #sim-embodied#planning#multi-agent#world-model
[2025/09] World Model Implanting for Test-time Adaptation of Embodied AgentsICML 2025 poster [paper] #text-adventure#memory#world-model#prompting
[2025/09] DreamPhase: Offline Imagination and Uncertainty-Guided Planning for Large-Language-Model AgentsICLR 2026 Poster [paper] #text-adventure#planning#world-model
[2025/07] CoEx -- Co-evolving World-model and ExplorationEMNLP 2025 [paper] #text-adventure#planning#world-model
[2025/06] Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead SearcharXiv [paper] #text-adventure#planning#world-model#training
[2025/05] PoE-World: Compositional World Modeling with Products of Programmatic ExpertsNeurIPS 2025 spotlight [paper] #action#planning#world-model
[2025/05] WALL-E: World Alignment by NeuroSymbolic Learning improves World Model-based LLM AgentsNeurIPS 2025 poster [paper] #minecraft#planning#world-model#prompting
[2025/04] WALL-E 2.0: World Alignment by NeuroSymbolic Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model#prompting
[2025/04] Better Decisions through the Right Causal World ModelarXiv [paper] #action#planning#world-model#training
[2025/02] Optimus-2: Multimodal World Model for Open-World Minecraft AgentsCVPR 2025 [paper][code] #minecraft#planning#world-model#vlm
[2024/10] WALL-E: World Alignment by Rule Learning Improves World Model-based LLM AgentsarXiv [paper][code] #minecraft#planning#world-model
[2024/05] Agent Planning with World Knowledge ModelNeurIPS 2024 [paper][code] #text-adventure#planning#memory#world-model
[2023/05] Language Models Meet World Models: Embodied Experiences Enhance Language ModelsNeurIPS 2023 [paper][code] #sim-embodied#planning#world-model
[2023/04] Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open QuestionsarXiv [paper] #text-adventure#world-model
tool-use
[2026/05] Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement LearningarXiv [paper] #text-adventure#tool-use#training
[2026/05] Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM AgentsarXiv [paper] #text-adventure#memory#tool-use
[2026/05] PokerSkill: LLMs Can Play Expert-Level Poker without Training or SolversarXiv [paper][code] #competition#tool-use#prompting
[2026/05] SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill GraphsarXiv [paper] #text-adventure#memory#tool-use#training
[2026/04] Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language ModelsarXiv [paper] #tool-use#training#self-improvement
[2025/09] Speculative Actions: A Lossless Framework for Faster AI AgentsICLR 2026 Oral [paper] #competition#tool-use
[2025/09] Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningICLR 2026 Poster [paper] #text-adventure#tool-use#training
[2025/08] Vistawise: Building Cost-effective Agent with Cross-modal Knowledge Graph for MinecraftEMNLP 2025 [paper] #minecraft#memory#tool-use
[2025/05] Agent-Environment Alignment via Automated Interface GenerationarXiv [paper][code] #text-adventure#tool-use
[2025/05] Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RLNeurIPS 2025 poster [paper] #communication#planning#tool-use#training
[2025/03] Parallelized Planning-Acting for Efficient LLM-based Multi-Agent SystemsarXiv [paper] #minecraft#planning#multi-agent#tool-use
[2025/03] Plancraft: an evaluation dataset for planning with LLM agentsCOLM 2025 [paper] #minecraft#planning#memory#tool-use
[2024/07] AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding AgentsACL 2024 [paper][code] #text-adventure#tool-use
[2024/07] Odyssey: Empowering Agents with Open-World Skills.IJCAI 2024 [paper][code] #minecraft#planning#tool-use
[2024/04] Learning From Failure: Integrating Negative Examples When Fine-tuning Large Language Models as AgentarXiv [paper][code] #text-adventure#tool-use#training
[2023/08] AgentSims: An Open-Source Sandbox for Large Language Model EvaluationarXiv [paper][code] #sim-social#planning#tool-use
[2023/05] VOYAGER: An Open-Ended Embodied Agent with Large Language ModelsFMDM@NeurIPS2023 [paper][code] #minecraft#tool-use#training
training
[2026/06] SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent TrainingarXiv [paper][code] #text-adventure#memory#training
[2026/06] Cross-Environment Neural Reranking for Sample-Efficient Action Selection in Text-Based AgentsarXiv [paper] #text-adventure#training
[2026/06] When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent TrainingarXiv [paper] #text-adventure#training
[2026/05] Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement LearningarXiv [paper] #action#training#vlm
[2026/05] T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement LearningarXiv [paper][code] #text-adventure#training
[2026/05] ANO: A Principled Approach to Robust Policy OptimizationarXiv [paper] #action#training
[2026/05] From History to State: Constant-Context Skill Learning for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/05] Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State UncertaintyarXiv [paper] #sim-embodied#training
[2026/05] PriorZero: Bridging Language Priors and World Models for Decision MakingarXiv [paper][code] #text-adventure#planning#world-model#training
[2026/05] Generalization or Memorization? Brittleness Testing for Chess-Trained Language ModelsarXiv [paper] #competition#training
[2026/05] ALSO: Adversarial Online Strategy Optimization for Social AgentsarXiv [paper] #sim-social#multi-agent#training
[2026/05] Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplayarXiv [paper] #action#planning#training#vlm
[2026/05] What and When to Distill: Selective Hindsight Distillation for Multi-Turn AgentsarXiv [paper] #text-adventure#training
[2026/05] GROW: Aligning GRPO with State-Action Modeling for Open-World VLM AgentsarXiv [paper] #minecraft#training#vlm
[2026/05] SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit AssignmentarXiv [paper] #text-adventure#training
[2026/05] SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill GraphsarXiv [paper] #text-adventure#memory#tool-use#training
[2026/05] R2V Agent: Teaching SLMs When to Ask for HelparXiv [paper] #text-adventure#training
[2026/04] MARL-GPT: Foundation Model for Multi-Agent Reinforcement LearningProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2026 [paper] #competition#multi-agent#training
[2026/04] Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM AgentsarXiv [paper][code] #text-adventure#training
[2026/04] Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language ModelsarXiv [paper] #tool-use#training#self-improvement
[2026/04] DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based AgentsarXiv [paper][code] #text-adventure#training
[2026/03] On the Strengths and Weaknesses of Data for Open-set Embodied AssistancearXiv [paper] #cooperation#training
[2026/03] Hindsight Credit Assignment for Long-Horizon LLM AgentsarXiv [paper] #text-adventure#training
[2026/03] How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game EnvironmentarXiv [paper] #text-adventure#multi-agent#training
[2026/03] PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy OptimizationWWW 2026 [paper] #sim-social#training
[2026/03] Grounded Chess Reasoning in Language Models via Master DistillationarXiv [paper] #competition#planning#training
[2026/03] RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic FeedbackarXiv [paper] #text-adventure#memory#training#self-improvement
[2026/02] Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker EnvironmentsarXiv [paper] #action#training
[2026/02] Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/02] VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case StudyarXiv [paper] #competition#training
[2026/02] CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesarXiv [paper] #text-adventure#planning#world-model#training
[2026/01] NitroGen: An Open Foundation Model for Generalist Gaming AgentsarXiv [paper] #benchmark#training
[2026/01] Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM AgentsarXiv [paper] #text-adventure#planning#training
[2026/01] Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient UpdatesarXiv [paper][code] #text-adventure#training
[2026/01] HumanLLM: Towards Personalized Understanding and Simulation of Human NaturearXiv [paper] #sim-social#training
[2025/12] Synergizing Code Coverage and Gameplay Intent: Coverage-Aware Game Playtesting with LLM-Guided Reinforcement LearningarXiv [paper] #minecraft#training
[2025/12] ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement LearningarXiv [paper] #sim-embodied#planning#memory#training
[2025/12] Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in DiplomacyarXiv [paper] #communication#training
[2025/12] Robust Agents in Open-Ended WorldsarXiv [paper] #action#multi-agent#training#generation
[2025/10] Fine-tuning with RAG for Improving LLM Learning of New SkillsarXiv [paper] #text-adventure#planning#memory#training
[2025/10] Graph-Enhanced Policy Optimization in LLM Agent TrainingarXiv [paper] #text-adventure#planning#training
[2025/10] SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory GrapharXiv [paper] #text-adventure#training
[2025/09] HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-MakingarXiv [paper] #competition#multi-agent#training
[2025/09] Exploration with Foundation Models: Capabilities, Limitations, and Hybrid ApproachesarXiv [paper] #action#training#vlm
[2025/09] Goal-Guided Efficient Exploration via Large Language Model in Reinforcement LearningarXiv [paper] #crafter#planning#training
[2025/09] Reward Is Enough: LLMs Are In-Context Reinforcement LearnersICLR 2026 Poster [paper] #text-adventure#training#self-improvement
[2025/09] Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful DemonstrationsICLR 2026 Poster [paper] #text-adventure#training
[2025/09] SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningICLR 2026 Poster [paper] #competition#planning#multi-agent#training
[2025/09] Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationICLR 2026 Poster [paper] #text-adventure#memory#training
[2025/09] Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningICLR 2026 Poster [paper] #text-adventure#tool-use#training
[2025/08] Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World SuccessProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems 2025 [paper] #text-adventure#planning#training#vlm
[2025/08] Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press DiplomacyAAAI 2025 [paper] #communication#training
[2025/08] Learning Game-Playing Agents with Generative Code OptimizationarXiv [paper] #action#training#self-improvement
[2025/08] SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making TasksarXiv [paper] #competition#planning#training#self-improvement
[2025/07] LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative SimulacraarXiv [paper][code] #sim-social#multi-agent#training#role-play
[2025/06] Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video GamesarXiv [paper][code] #benchmark#training
[2025/05] Enfoque Odychess: Un método dialéctico, constructivista y adaptativo para la enseñanza del ajedrez con inteligencias artificiales generativasarXiv [paper] #competition#training
[2025/05] Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme OnearXiv [paper] #action#training
[2025/05] Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningICML 2025 poster [paper] #text-adventure#planning#training
[2025/05] Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsNeurIPS 2025 poster [paper] #action#training
[2025/05] Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RLNeurIPS 2025 poster [paper] #communication#planning#tool-use#training
[2025/05] Can Large Language Models Master Complex Card Games?NeurIPS 2025 poster [paper][code] #competition#training
[2025/05] MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural LearningNeurIPS 2025 poster [paper] #minecraft#training
[2025/05] LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language ModelsNeurIPS 2025 spotlight [paper][code] #action#training
[2025/04] Collaborating Action by Action: A Multi-agent LLM Framework for Embodied ReasoningarXiv [paper][code] #minecraft#multi-agent#training
[2025/04] Better Decisions through the Right Causal World ModelarXiv [paper] #action#planning#world-model#training
[2025/04] Enhancing Player Enjoyment with a Two-Tier DRL and LLM-Based Agent System for Fighting GamesarXiv [paper] #training
[2025/04] Monte Carlo Planning with Large Language Model for Text-Based Game AgentsICLR 2025 Poster [paper] #text-adventure#planning#training
[2025/04] MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model FrameworkNeurIPS 2025 poster [paper] #sim-social#planning#training
[2025/03] GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow NetworksCVPR 2025 [paper] #text-adventure#planning#training#vlm
[2025/03] Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic ReasoningCOLM 2025 [paper] #competition#multi-agent#training
[2025/02] Process Reward Models for LLM Agents: Practical Framework and DirectionsarXiv [paper][code] #text-adventure#training
[2025/01] POKERBENCH: Training Large Language Models to become Professional Poker PlayersAAAI 2025 [paper] #competition#planning#training
[2025/01] DVM: Towards Controllable LLM Agents in Social Deduction GamesIEEE International Conference on Acoustics, Speech, and Signal Processing 2025 [paper] #communication#training
[2025/01] Complete Chess Games Enable LLM Become A Chess MasterNAACL 2025 [paper] #competition#training
[2025/01] LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal DemonstrationsICML 2025 poster [paper] #action#planning#training
[2025/01] Language Models as Implicit Tree SearchICML 2025 poster [paper] #competition#planning#training
[2025/01] LARM: Large Auto-Regressive Model for Long-Horizon Embodied IntelligenceICML 2025 poster [paper] #minecraft#training
[2024/12] TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in MinecraftarXiv [paper][code] #minecraft#multi-agent#training#vlm
[2024/12] Fine-tuning large vision-language models as decision-making agents via reinforcement learningNeurIPS 2024 [paper][code] #text-adventure#training#vlm
[2024/09] Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study CasearXiv [paper][code] #action#planning#training
[2024/09] Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchICLR 2025 Poster [paper] #communication#planning#multi-agent#training
[2024/09] Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-Based Decision-Making SystemsICLR 2025 Poster [paper] #sim-embodied#training
[2024/09] Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackICLR 2025 Poster [paper] #text-adventure#planning#training#self-improvement
[2024/09] BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesICLR 2025 Poster [paper] #action#planning#training#vlm
[2024/08] Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT) 2024 [paper] #competition#planning#multi-agent#training
[2024/08] Atari-GPT: Investigating the Capabilities of Multimodal Large Language Models as Low-Level Policies for Atari GamesarXiv [paper] #action#planning#training
[2024/07] OmniJARVIS: Omni-Modal Open-World Agents in MinecraftNeurIPS 2024 [paper][code] #minecraft#training#vlm
[2024/06] STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language ModelsACL 2024 [paper][code] #text-adventure#training
[2024/05] Towards Efficient LLM Grounding for Embodied Multi-Agent CollaborationACL 2024 [paper][code] #cooperation#planning#multi-agent#training
[2024/05] Policy Improvement using Language Feedback ModelsNeurIPS 2024 poster [paper] #text-adventure#training
[2024/05] Learning to Discuss Strategically: A Case Study on One Night Ultimate WerewolfNeurIPS 2024 poster [paper] #communication#training
[2024/05] Reflective Multi-Agent Collaboration based on Large Language ModelsNeurIPS 2024 poster [paper] #competition#planning#multi-agent#training
[2024/04] Learning From Failure: Integrating Negative Examples When Fine-tuning Large Language Models as AgentarXiv [paper][code] #text-adventure#tool-use#training
[2024/04] ReAct Meets ActRe: When Language Agents Enjoy Training Data AutonomyarXiv [paper] #text-adventure#planning#training#self-improvement
[2024/04] World Models with Hints of Large Language Models for Goal AchievingNAACL 2024 [paper] #crafter#training#vlm
[2024/04] Self-playing Adversarial Language Game Enhances LLM ReasoningNeurIPS 2024 [paper][code] #communication#training#self-improvement
[2024/03] Language Guided Exploration for RL Agents in Text EnvironmentsNAACL 2024 [paper][code] #text-adventure#training
[2024/03] Trial and Error: Exploration-Based Trajectory Optimization for LLM AgentsACL 2024 [paper][code] #text-adventure#training#self-improvement
[2024/03] EnvGen: Generating and Adapting Environments via LLMs for Training Embodied AgentsCOLM [paper] #crafter#training
[2024/03] SOTOPIA-$\pi$: Interactive Learning of Socially Intelligent Language AgentsACL 2024 [paper][code] #sim-social#training
[2024/03] Will GPT-4 Run DOOM?IEEE Transactions on Games 2024 [paper][code] #action#planning#training
[2024/03] O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language ModelsCOLM [paper] #text-adventure#training
[2024/03] ReAct Meets ActRe: Autonomous Annotation of Agent Trajectories for Contrastive Self-TrainingCOLM [paper] #text-adventure#planning#training#self-improvement
[2024/02] PokéLLMon: A Human-Parity Agent for Pokémon Battles with Large Language ModelsTOIT 2025 [paper][code] #competition#training
[2024/02] Agent-Pro: Learning to Evolve via Policy-Level Reflection and OptimizationACL 2024 [paper][code] #competition#training
[2024/02] Enhance Reasoning for Large Language Models in the Game WerewolfarXiv [paper] #communication#training
[2024/01] True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement LearningarXiv [paper][code] #sim-embodied#training
[2024/01] PokerGPT: An End-to-End Lightweight Solver for Multi-Player Texas Hold'em via Large Language ModelarXiv [paper] #competition#training
[2024/01] SwarmBrain: Embodied agent for real-time strategy game StarCraft II via large language modelsarXiv [paper] #competition#training
[2023/12] Auto MC-Reward: Automated Dense Reward Design with Large Language Models for MinecraftCVPR 2023 [paper] #minecraft#training
[2023/10] FireAct: Toward Language Agent Fine-tuningarXiv [paper][code] #text-adventure#training
[2023/10] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language ModelsICML 2024 [paper][code] #text-adventure#planning#training
[2023/10] LLaMA Rider: Spurring Large Language Models to Explore the Open WorldNAACL 2023 [paper][code] #minecraft#planning#training
[2023/09] AdaRefiner: Refining Decisions of Language Models with Adaptive FeedbackNAACL 2023 [paper] #crafter#training
[2023/07] Building Cooperative Embodied Agents Modularly with Large Language ModelsICLR 2024 [paper][code] #cooperation#planning#multi-agent#training
[2023/05] Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and MemoryarXiv [paper] #minecraft#training
[2023/05] VOYAGER: An Open-Ended Embodied Agent with Large Language ModelsFMDM@NeurIPS2023 [paper][code] #minecraft#tool-use#training
[2023/03] Plan4MC: Skill Reinforcement Learning and Planning for Open-World Minecraft TasksFMDM@NeurIPS2023 [paper][code] #minecraft#planning#training
[2023/03] Reflexion: Language Agents with Verbal Reinforcement LearningNeurIPS 2023 [paper][code] #text-adventure#training#self-improvement
[2023/02] Guiding Pretraining in Reinforcement Learning with Large Language ModelsICML 2023 [paper] #crafter#training
[2023/02] Grounding Large Language Models in Interactive Environments with Online Reinforcement LearningICML 2023 [paper][code] #action#training
[2022/10] ReAct: Synergizing Reasoning and Acting in Language ModelsICLR 2023 [paper][code] #text-adventure#planning#training
self-improvement
[2026/06] SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at ScalearXiv [paper] #text-adventure#memory#self-improvement
[2026/05] Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within AgentsarXiv [paper] #text-adventure#training#self-improvement
[2026/05] Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM AgentsarXiv [paper] #minecraft#self-improvement
[2026/04] Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language ModelsarXiv [paper] #tool-use#training#self-improvement
[2026/04] GraSP: Graph-Structured Skill Compositions for LLM AgentsarXiv [paper] #text-adventure#planning#memory#self-improvement
[2026/03] RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic FeedbackarXiv [paper] #text-adventure#memory#training#self-improvement
[2026/02] MemSkill: Learning and Evolving Memory Skills for Self-Evolving AgentsarXiv [paper] #text-adventure#self-improvement
[2026/02] The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI SocietiesarXiv [paper] #sim-social#multi-agent#self-improvement
[2025/10] Alita-G: Self-Evolving Generative Agent for Agent GenerationarXiv [paper] #sim-social#memory#self-improvement
[2025/10] Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned StrategiesarXiv [paper] #communication#multi-agent#self-improvement
[2025/09] Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM AgentarXiv [paper] #text-adventure#planning#multi-agent#self-improvement
[2025/09] PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments2025 IEEE Conference on Games (CoG) 2025 [paper] #minecraft#multi-agent#self-improvement
[2025/09] Code Driven Planning with Domain-Adaptive CriticarXiv [paper] #text-adventure#planning#self-improvement
[2025/09] Reward Is Enough: LLMs Are In-Context Reinforcement LearnersICLR 2026 Poster [paper] #text-adventure#training#self-improvement
[2025/08] Learning Game-Playing Agents with Generative Code OptimizationarXiv [paper] #action#training#self-improvement
[2025/08] SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making TasksarXiv [paper] #competition#planning#training#self-improvement
[2025/06] Automated Skill Discovery for Language Agents through Exploration and Iterative FeedbackarXiv [paper] #crafter#self-improvement
[2025/06] OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic ReflectionsarXiv [paper] #text-adventure#planning#training#self-improvement
[2025/02] TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning.arXiv [paper] #text-adventure#self-improvement
[2024/12] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 [paper][code] #communication#self-improvement
[2024/09] Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackICLR 2025 Poster [paper] #text-adventure#planning#training#self-improvement
[2024/06] Watch Every Step! LLM Agent Learning via Iterative Step-Level Process RefinementEMNLP 2024 [paper][code] #text-adventure#self-improvement
[2024/05] Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyNeurIPS 2024 poster [paper] #communication#planning#multi-agent#self-improvement
[2024/04] ReAct Meets ActRe: When Language Agents Enjoy Training Data AutonomyarXiv [paper] #text-adventure#planning#training#self-improvement
[2024/04] Self-playing Adversarial Language Game Enhances LLM ReasoningNeurIPS 2024 [paper][code] #communication#training#self-improvement
[2024/03] Trial and Error: Exploration-Based Trajectory Optimization for LLM AgentsACL 2024 [paper][code] #text-adventure#training#self-improvement
[2024/03] ReAct Meets ActRe: Autonomous Annotation of Agent Trajectories for Contrastive Self-TrainingCOLM [paper] #text-adventure#planning#training#self-improvement
[2024/03] StateFlow: Enhancing LLM Task-Solving through State-Driven WorkflowsCOLM [paper] #text-adventure#planning#self-improvement
[2024/02] Soft Self-Consistency Improves Language Model AgentsarXiv [paper][code] #text-adventure#self-improvement
[2024/02] Empowering Large Language Model Agents through Action LearningCOLM [paper][code] #text-adventure#planning#self-improvement
[2023/10] Lyfe Agents: Generative agents for low-cost real-time social interactionsarXiv [paper] #sim-social#memory#multi-agent#self-improvement
[2023/03] Reflexion: Language Agents with Verbal Reinforcement LearningNeurIPS 2023 [paper][code] #text-adventure#training#self-improvement
prompting
[2026/05] GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM CollectivesarXiv [paper] #competition#multi-agent#prompting
[2026/05] PokerSkill: LLMs Can Play Expert-Level Poker without Training or SolversarXiv [paper][code] #competition#tool-use#prompting
[2026/04] DORA Explorer: Improving the Exploration Ability of LLMs Without TrainingarXiv [paper] #text-adventure#planning#prompting