NUS-Project/Landmark-of-medical-agent

HTML

186

193 commits

updated Jun 8, 2026

See the code

README

🚀 The Landscape of Medical Agents: A Survey

🚀 MedMASLab:A Framework for Multimodal Medical Multi-Agent Systems

Overall Landscape

Overall Landscape

🌟 Overview

This is the official repository for the survey paper: The Landscape of Medical Agents. This repository is a comprehensive and systematic research resource library for medical agents, dedicated to organizing and tracking the latest research progress, application practices, and technological developments of AI intelligent agents in the medical and health field. This investigative project covers the entire ecosystem from basic technical capabilities to clinical actual deployment, providing an authoritative research map for medical AI researchers, clinical practitioners, and system developers.

🔥 News

[2026/3/10] We release :A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems! Click here to visit: medmaslab

🔥 Add Your Paper in our Repo and Survey!!!!! Join our group and post your agent work!!!!

[-] You are welcome to give us an issue or PR for your medical agent work !!!!!

[-] Note that: Due to the huge paper in Arxiv, we are sorry to cover all in our survey. You can directly present a PR into this repo and we will record it for next version update of our survey.

[-] Our survey will be updated in 2026.3.

🤝 Thanks

If you think this project is useful and inspiring, we would greatly appreciate it if you could give us a Star to show your support! Your support is of great significance to us, as it encourages us to continue improving and developing this project.

📖 Keywords

Medical Agents, Clinical Workflows, Safety, Governance and Evaluation

🌟 Contributing

We will try to keep this list updated. If you find any errors or any missed paper, please don't hesitate to open issues or pull request.Please follow the instruction in CONTRIBUTING.md if you want to make one. Additionally, if you want to have any other issue, please add this wechat group.

🤝 Main Contacts

Citation

 @article{hu2025landscape,
  title={The Landscape of Medical Agents: A Survey},
  author={Hu, Xiaobin and Qian, Yunhang and Yu, Jiaquan and Liu, Jingjing and Tang, Peng and Ji, Xiaozhong and Xu, Chengming and Liu, Jiawei and Yan, Xiaoxiao and Yu, Xinlei and others},
  journal={Authorea Preprints},
  year={2025},
  publisher={Authorea}
}
@misc{qian2026medmaslabunifiedorchestrationframework,
      title={MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems}, 
      author={Yunhang Qian and Xiaobin Hu and Jiaquan Yu and Siyang Xin and Xiaokun Chen and Jiangning Zhang and Peng-Tao Jiang and Jiawei Liu and Hongwei Bran Li},
      year={2026},
      eprint={2603.09909},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2603.09909}, 
}

🌟 Table of Contents

✨ Latest Papers

🚀 Year-2026

January

TitlePaper-linkSections
EvoClinician: A Self-Evolving Agent for Multi-Turn Medical Diagnosis via Test-Time Evolutionary LearningPaperbenchmark
Structure-constrained Language-informed Diffusion Model for Unpaired Low-dose Computed Tomography Angiography ReconstructionPaperapplication
Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement LearningPaperbenchmark
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive DiagnosisPaperbenchmark
DEEPMED: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & InferencePaperbenchmark
Bayesian Multiple Testing for Suicide Risk in Pharmacoepidemiology: Leveraging Co-Prescription PatternsPaperframework
AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent ReasoningPaperbenchmark
Automated Rubrics for Reliable Evaluation of Medical Dialogue SystemsPaperframework
Query-Efficient Agentic Graph Extraction Attacks on GraphRAG SystemsPaperbenchmark
HyperWalker: Dynamic Hypergraph-Based Deep Diagnosis for Multi-Hop Clinical Modeling across EHR and X-Ray in Medical VLMsPaperframework
AgentEHR: Advancing Autonomous Clinical Decision-Making via Retrospective SummarizationPaperbenchmark
Towards Efficient and Robust Linguistic Emotion Diagnosis for Mental Health via Multi-Agent Instruction RefinementPaperbenchmark
Improving the Safety and Trustworthiness of Medical AI via Multi-Agent Evaluation LoopsPaperframework
MedConsultBench: A Full-Cycle, Fine-Grained, Process-Aware Benchmark for Medical Consultation AgentsPaperbenchmark
Knowing When to Abstain: Medical LLMs Under Clinical UncertaintyPaperbenchmark
A multitask framework for automated interpretation of multi-frame right upper quadrant ultrasound in clinical decision supportPaperbenchmark
Medication counseling with large language models: balancing flexibility and rigidityPaperbenchmark
MMedExpert-R1: Strengthening Multimodal Medical Reasoning via Domain-Specific Adaptation and Clinical Guideline ReinforcementPaperbenchmark
Japanese AI Agent System on Human Papillomavirus Vaccination: System DesignPaperframework
ART: Action-based Reasoning Task Benchmarking for Medical AI AgentsPaperbenchmark
Route, Retrieve, Reflect, Repair: Self-Improving Agentic Framework for Visual Detection and Linguistic Reasoning in Medical ImagingPaperframework
MEDVISTAGYM: A Scalable Training Environment for Thinking with Medical Images via Tool-Integrated Reinforcement LearningPaperbenchmark
MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential DiagnosisPaperbenchmark
Modeling Descriptive Norms in Multi-Agent Systems: An Auto-Aggregation PDE Framework with Adaptive Perception KernelsPaperbenchmark
Value of Information: A Framework for Human-Agent CommunicationPaperframework
DemMA: Dementia Multi-Turn Dialogue Agent with Expert-Guided Reasoning and Action SimulationPaperframework
Tool-MAD: A Multi-Agent Debate Framework for Fact Verification with Diverse Tool Augmentation and Adaptive RetrievalPaperbenchmark
Staged Voxel-Level Deep Reinforcement Learning for 3D Medical Image Segmentation with Noisy AnnotationsPaperbenchmark
RadDiff: Describing Differences in Radiology Image Sets with Natural LanguagePaperbenchmark
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and SegmentationPaperbenchmark
MedDialogRubrics: A Comprehensive Benchmark and Evaluation Framework for Multi-turn Medical Consultations in Large Language ModelsPaperbenchmark
Causal-Enhanced AI Agents for Medical Research ScreeningPapermedical
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate ForgettingPapermedical
Bayesian Orchestration of Multi-LLM Agents for Cost-Aware Sequential Decision-MakingPaperframework
An Explainable Agentic AI Framework for Uncertainty-Aware and Abstention-Enabled Acute Ischemic Stroke Imaging DecisionsPaperbenchmark

February

TitlePaper-linkSections
Evaluating Stochasticity in Deep Research AgentsPaperframework
Agency and Architectural Limits: Why Optimization-Based Systems Cannot Be Norm-ResponsivePapermedical
Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot StudyPaperbenchmark
Which Tool Response Should I Trust? Tool-Expertise-Aware Chest X-ray Agent with Multimodal Agentic LearningPaperframework
ALPACA: A Reinforcement Learning Environment for Medication Repurposing and Treatment Optimization in Alzheimer's DiseasePapermedical
LAMMI-Pathology: A Tool-Centric Bottom-Up LVLM-Agent Framework for Molecularly Informed Medical Intelligence in PathologyPaperframework
NutriOrion: A Hierarchical Multi-Agent Framework for Personalized Nutrition Intervention Grounded in Clinical GuidelinesPaperframework
4D-UNet improves clutter rejection in human transcranial contrast enhanced ultrasoundPaperbenchmark
INSURE-Dial: A Phase-Aware Conversational Dataset & Benchmark for Compliance Verification and Phase DetectionPaperbenchmark
3DMedAgent: Unified Perception-to-Understanding for 3D Medical AnalysisPaperbenchmark
Agentic Unlearning: When LLM Agent Meets Machine UnlearningPaperbenchmark
MedClarify: An information-seeking AI agent for medical diagnosis with case-specific follow-up questionsPapermedical
Agentic AI, Medical Morality, and the Transformation of the Patient-Physician RelationshipPaperframework
A Multi-Agent Framework for Medical AI: Leveraging Fine-Tuned GPT, LLaMA, and DeepSeek R1 for Evidence-Based and Bias-Aware Clinical Query ProcessingPaperbenchmark
MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool CallingPaperbenchmark
Implicit Bias in LLMs for Transgender PopulationsPaperapplication
TRACE: Temporal Reasoning via Agentic Context Evolution for Streaming Electronic Health Records (EHRs)Paperframework
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMsPaperbenchmark
Advancing AI Trustworthiness Through Patient Simulation: Risk Assessment of Conversational Agents for Antidepressant SelectionPaperframework
LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric EvaluationPaperbenchmark
Closing Reasoning Gaps in Clinical Agents with Differential Reasoning LearningPaperbenchmark
CoMMa: Contribution-Aware Medical Multi-Agents From A Game-Theoretic PerspectivePaperbenchmark
SynthAgent: A Multi-Agent LLM Framework for Realistic Patient Simulation -- A Case Study in Obesity with Mental Health ComorbiditiesPaperframework
MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive RegulationPaperbenchmark
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional TasksPaperbenchmark
Do LLMs Act Like Rational Agents? Measuring Belief Coherence in Probabilistic Decision MakingPaperapplication
Pruning Minimal Reasoning Graphs for Efficient Retrieval-Augmented GenerationPaperbenchmark
Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based AgentsPaperbenchmark
MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement LearningPaperbenchmark
Privasis: Synthesizing the Largest "Public" Private Dataset from ScratchPaperbenchmark
Perfusion Imaging and Single Material Reconstruction in Polychromatic Photon Counting CTPapermedical
RE-MCDF: Closed-Loop Multi-Expert LLM Reasoning for Knowledge-Grounded Clinical DiagnosisPaperbenchmark
MedBeads: An Agent-Native, Immutable Data Substrate for Trustworthy Medical AIPaperframework
ExperienceWeaver: Optimizing Small-sample Experience Learning for LLM-based Clinical Text ImprovementPaperbenchmark
Enhancing Imaging Depth and Sensitivity in Reflectance Mode Near Infrared Optical Imaging with Scatter Reducing AgentsPaperapplication

March

TitlePaper-linkSections
Symphony for Medical Coding: A Next-Generation Agentic System for Scalable and Explainable Medical CodingPaperbenchmark
Knowledge database development by large language models for countermeasures against viruses and marine toxinsPapermedical
Towards a Medical AI ScientistPaperframework
FeDMRA: Federated Incremental Learning with Dynamic Memory Replay AllocationPaperbenchmark
Improving Clinical Diagnosis with Counterfactual Multi-Agent ReasoningPaperbenchmark
MediHive: A Decentralized Agent Collective for Medical ReasoningPaperbenchmark
Autonomous Agent-Orchestrated Digital Twins (AADT): Leveraging the OpenClaw Framework for State Synchronization in Rare Genetic DisordersPaperframework
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AIPaperbenchmark
Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy VideosPaperbenchmark
OMIND: Framework for Knowledge Grounded Finetuning and Multi-Turn Dialogue Benchmark for Mental Health LLMsPaperbenchmark
Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationPaperbenchmark
MedOpenClaw: Auditable Medical Imaging Agents Reasoning over Uncurated Full StudiesPaperbenchmark
Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQAPaperbenchmark
RVLM: Recursive Vision-Language Models with Adaptive DepthPaperframework
CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in HealthcarePaperbenchmark
Dialogue to Question Generation for Evidence-based Medical Guideline Agent DevelopmentPaperbenchmark
3D-LLDM: Label-Guided 3D Latent Diffusion Model for Improving High-Resolution Synthetic MR Imaging in Hepatic Structure SegmentationPaperbenchmark
From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLMPaperframework
Training a Large Language Model for Medical Coding Using Privacy-Preserving Synthetic Clinical DataPapermedical
Privacy-Preserving EHR Data Transformation via Geometric Operators: A Human-AI Co-Design Technical ReportPaperframework
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical DatabasesPaperbenchmark
Cerebra: A Multidisciplinary AI Board for Multimodal Dementia Characterization and Risk AssessmentPaperbenchmark
Agentic Automation of BT-RADS Scoring: End-to-End Multi-Agent System for Standardized Brain Tumor Follow-up AssessmentPapermedical
ARYA: A Physics-Constrained Composable & Deterministic World Model ArchitecturePaperbenchmark
Anatomical Prior-Driven Framework for Autonomous Robotic Cardiac Ultrasound Standard View AcquisitionPaperbenchmark
TuLaBM: Tumor-Biased Latent Bridge Matching for Contrast-Enhanced MRI SynthesisPaperbenchmark
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective IntelligencePaperbenchmark
MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question AnsweringPaperbenchmark
Towards Equitable Robotic Furnishing Agents for Aging-in-Place: ADL-Grounded Design ExplorationPaperother
EviAgent: Evidence-Driven Agent for Radiology Report GenerationPaperbenchmark
OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance DatasetPaperbenchmark
Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in SpacePaperbenchmark
Six Interventions for the Responsible and Ethical Implementation of Medical AI AgentsPaperbenchmark
TheraAgent: Multi-Agent Framework with Self-Evolving Memory and Evidence-Calibrated Reasoning for PET TheranosticsPaperframework
Increasing intelligence in AI agents can worsen collective outcomesPapermedical
A Semi-Decentralized Approach to Multiagent ControlPaperbenchmark
When OpenClaw Meets Hospital: Toward an Agentic Operating System for Dynamic Clinical WorkflowsPaperbenchmark
UAV-MARL: Multi-Agent Reinforcement Learning for Time-Critical and Dynamic Medical Supply DeliveryPaperbenchmark
Human-AI Co-reasoning for Clinical Diagnosis with Evidence-Integrated Language AgentPaperbenchmark
MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent SystemsPaperbenchmark
Meissa: Multi-modal Medical Agentic IntelligencePaperbenchmark
RexDrug: Reliable Multi-Drug Combination Extraction through Reasoning-Enhanced LLMsPaperbenchmark
Empowering Locally Deployable Medical Agent via State Enhanced Logical Skills for FHIR-based Clinical TasksPaperbenchmark
Computational Pathology in the Era of Emerging Foundation and Agentic AI -- International Expert Perspectives on Clinical Integration and Translational ReadinessPaperbenchmark
Shifting Adaptation from Weight Space to Memory Space: A Memory-Augmented Agent for Medical Image SegmentationPaperbenchmark
Evolving Medical Imaging Agents via Experience-driven Self-skill DiscoveryPaperbenchmark
MedCoRAG: Interpretable Hepatology Diagnosis via Hybrid Evidence Retrieval and Multispecialty ConsensusPaperframework
Model Medicine: A Clinical Framework for Understanding, Diagnosing, and Treating AI ModelsPaperframework
Do Mixed-Vendor Multi-Agent LLMs Improve Clinical Diagnosis?Paperframework
A Multi-Agent Framework for Interpreting Multivariate Physiological Time SeriesPapermedical
MIND: Unified Inquiry and Diagnosis RL with Criteria Grounded Clinical Supports for Psychiatric ConsultationPaperframework
From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAGPaperbenchmark
ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical DialoguePaperbenchmark
CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic FrameworkPaperbenchmark
TARSE: Test-Time Adaptation via Retrieval of Skills and Experience for Reasoning AgentsPaperbenchmark
MedCollab: Causal-Driven Multi-Agent Collaboration for Full-Cycle Clinical Diagnosis via IBIS-Structured ArgumentationPaperbenchmark
DUCX: Decomposing Unfairness in Tool-Using Chest X-ray AgentsPaperframework
OPGAgent: An Agent for Auditable Dental Panoramic X-ray InterpretationPaperbenchmark

April

🚀 Year-2025

TitleGitHubSections
MedEyes: Learning Dynamic Visual Focus for Medical Progressive DiagnosisGitHubapplication
Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic RetrievalGitHubapplication
Reinventing Clinical Dialogue: Agentic Paradigms for LLM‑Enabled Healthcare CommunicationGitHubsurvey
Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal ReasoningGitHubevaluation
3mdbench: Medical multimodal multi-agent dialogue benchmarkGitHubevaluation
A co-evolving agentic AI system for medical imaging analysisGitHubother
A multimodal AI agent for clinical decision support in ophthalmologyNot Availableapplication
MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic WorkflowGitHubapplication
A dual-agent collaboration framework based on llms for nursing robots to perform bimanual coordination tasksNot Availablecapability
A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisationNot Availablesafety, evaluation
A hybrid reinforcement learning and knowledge graph framework for financial risk optimization in healthcare systemsNot Availableapplication
A Multi-Agent Approach to Neurological Clinical ReasoningNot Availablecapability, other
A Multimodal Multi-Agent Framework for Radiology Report GenerationNot Availabletask, evaluation, other
A Proposed LLM-Based Supported Treatment Framework for Intracerebral HemorrhageNot Availableintro, capability, application
A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language ModelsNot Availablecapability, application
A two-stage proactive dialogue generator for efficient clinical information collection using large language modelNot Availabletask, application
Actions speak louder than words: Agent decisions reveal implicit biases in language modelsNot Availablesafety
Adagent: Llm agent for alzheimer’s disease analysis with collaborative coordinatorNot Availablecapability, other
Agent-Based Uncertainty Awareness Improves Automated Radiology Report Labeling with an Open-Source Large Language ModelNot Availablecapability, other
Agentic AI for Clinical Decision Support: Real-Time Diagnosis, Triage, and Treatment PlanningNot Availableintro
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical KnowledgeGitHubcapability, task, application
Agentic Workflows in Healthcare: Advancing Clinical Efficiency through AI IntegrationNot Availableintro
Agentic-AI Healthcare: Multilingual, Privacy-First Framework with {MCP} AgentsGitHubcapability, application, safety
AGENTiGraph: A Multi-Agent Knowledge Graph Framework for Interactive, Domain-Specific LLM ChatbotsGitHubtask
Agentmd: Empowering language agents for risk prediction with large-scale clinical tool learningGitHubcapability, task, application
AI Agents in Clinical Medicine: A Systematic ReviewNot Availableevaluation
AI chatbots as professional service agents: developing a professional identityNot Availablecapability, application
AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question AnsweringGitHubcapability, task, other
AMQA: An Adversarial Dataset for Benchmarking Bias of LLMs in Medicine and HealthcareGitHubevaluation
An active inference strategy for prompting reliable responses from large language models in medical practiceNot Availablecapability
An Adaptive Multi-Agent LLM-Based Clinical Decision Support System Integrating Biomedical RAG and Web IntelligenceGitHubapplication
An Agentic Model Context Protocol Framework for Medical Concept StandardizationGitHubtask
An agentic system for rare disease diagnosis with traceable reasoningNot Availabletask, application, other
Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQAGitHubtask, other
ASTRID--An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering SystemsNot Availablesafety
At-cxr: Uncertainty-aware agentic triage for chest x-raysGitHubintro, other
Audited Reasoning Refinement: Fine-Tuning Language Models via LLM-Guided Step-Wise Evaluation and CorrectionNot Availableevaluation
AURA: A Multi-modal Medical Agent for Understanding, Reasoning and AnnotationGitHubother
Autonomous Multi-Modal LLM Agents for Treatment Planning in Focused Ultrasound Ablation SurgeryGitHubcapability, task, application
Autonomous Radiotherapy Treatment Planning Using DOLA: A Privacy-Preserving, LLM-Based Optimization AgentNot Availablesafety, other
Balancing Fairness and Performance in Healthcare {AI}: A Gradient Reconciliation ApproachNot Availablesafety
Benchmarking Automatic Speech Recognition coupled LLM Modules for Medical DiagnosticsNot Availableapplication
Beyond Benchmarks: Dynamic, Automatic and Systematic Red-Teaming Agents for Trustworthy Medical Language ModelsGitHubevaluation
Beyond Benchmarks: Evaluating Generalist Medical Artificial Intelligence With PsychometricsNot Availableevaluation
Bridging Clinical Narratives and ACR Appropriateness Guidelines: A Multi-Agent RAG System for Medical Imaging DecisionsGitHubtask
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False PresuppositionsGitHubother
CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement EstimationGitHubintro, capability, application, other
CARE-AD: a multi-agent large language model framework for Alzheimer’s disease prediction using longitudinal clinical notesNot Availablecapability, task, other
CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery PlanningNot Availablecapability, application, other
Chatbot To Help Patients Understand Their HealthGitHubapplication
ChatMyopia: An AI Agent for Pre-consultation Education in Primary Eye Care SettingsNot Availablecapability, application, other
Cod, towards an interpretable medical agent using chain of diagnosisGitHubcapability, application, safety, evaluation, other
Code Like Humans: A Multi-Agent Solution for Medical CodingGitHubapplication
Conversational health agents: a personalized large language model-powered agent frameworkGitHubsafety
CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question AnsweringNot Availableapplication
Data Overdose? Time for a Quadruple Shot: Knowledge Graph Construction Using Enhanced Triple ExtractionNot Availabletask
Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat AnalysisNot Availablesafety
Developing an Artificial Intelligence Tool for Personalized Breast Cancer Treatment Plans based on the NCCN GuidelinesNot Availableintro, capability, other
Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncologyNot Availableapplication, other
Differential privacy for medical deep learning: methods, tradeoffs, and deployment implicationsNot Availablesafety
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology ReasoningNot Availableevaluation
DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical servicesNot Availabletask, other
Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement LearningGitHubcapability, application
Doctoragent-rl: A multi-agent collaborative reinforcement learning system for multi-turn clinical dialogueGitHubtask
Dr. Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in RomanianNot Availablecapability, task
Drugagent: Multi-agent large language model-based reasoning for drug-target interaction predictionNot Availabletask
EH-Benchmark: Ophthalmic hallucination benchmark and agent-driven top-down traceable reasoning workflowGitHubother
Ehr-mcp: Real-world evaluation of clinical information retrieval by large language models via model context protocolNot Availableevaluation
Emerging cyber attack risks of medical ai agentsNot Availablesafety
Enhancing diagnostic capability with multi-agents conversational large language modelsGitHubcapability, task, application, other
Enhancing Medical Lung X-Ray Diagnosis Through Multi-Agent Vision-Language Model CollaborationNot Availablecapability, application
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in MedicineNot Availableevaluation
Evaluating the accuracy of a state-of-the-art large language model for prediction of admissions from the emergency roomNot Availableevaluation
Evaluating transparency in AI/ML model characteristics for FDA-reviewed medical devicesNot Availablesafety
Explainable AI for medical data: Current methods, limitations, and future directionsNot Availablesafety
Eyecaregpt: Boosting comprehensive ophthalmology understanding with tailored dataset, benchmark and modelGitHubapplication, other
Feat: A multi-agent forensic ai system with domain-adapted large language model for automated cause-of-death analysisGitHubcapability
Fine-tuning vision language models with graph-based knowledge for explainable medical image analysisNot Availablecapability
FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research InsightsGitHubother
GEMA-Score: Granular Explainable Multi-Agent Score for Radiology Report EvaluationGitHubother
Geometry-preserving encoder/decoder in latent generative modelsGitHubevaluation
GMAT: Grounded Multi-agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image ClassificationNot Availabletask
Haibu Mathematical-Medical Intelligent Agent: Enhancing Large Language Model Reliability in Medical Tasks via Verifiable Reasoning ChainsNot Availablecapability
Demo: Healthcare Agent Orchestrator (HAO) for Patient Summarization in Molecular Tumor BoardsNot Availablesections: task, other
Healthcare Agent: Eliciting the Power of Large Language Models for Medical ConsultationNot Availablecapability, application
Image Segmentation Using Only" Better or Worse" Expert FeedbackNot Availabletask
Improving Interactive Diagnostic Ability of a Large Language Model Agent Through Clinical Experience LearningGitHubcapability, application
In-Basket Message Volume in Primary Care: A Cross-sectional Analysis by Gender and SpecialtyNot Availableintro
KERAP: A Knowledge-Enhanced Reasoning Approach for Accurate Zero-shot Diagnosis Prediction Using Multi-agent LLMsGitHubcapability
Large language models in real-world clinical workflows: a systematic review of applications and implementationNot Availableevaluation
Learning to be a doctor: Searching for effective medical agent architecturesNot Availablecapability, application, other
Lessons Learned from Evaluation of LLM based Multi-agents in Safer Therapy RecommendationGitHubtask, other
LINS: A general medical Q&A framework for enhancing the quality and credibility of LLM-generated responsesGitHubcapability
Llms can simulate standardized patients via agent coevolutionGitHubtask, application
M3Builder: A Multi-Agent System for Automated Machine Learning in Medical ImagingGitHubtask
Magnetic Milli-Spinner for Robotic Endovascular SurgeryNot Availabletask
MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized CollaborationGitHubapplication, other
Mdteamgpt: A self-evolving llm-based multi-agent framework for multi-disciplinary team medical consultationGitHubtask, application
Measurement to Meaning: A Validity-Centered Framework for AI EvaluationNot Availableevaluation
Med-TAMARA: Trust-Aware Multi-Agent Risk Assessment in Medical AI DialogueNot Availablecapability, application
Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced AgentsGitHubcapability, application
MedAgent-Pro: Towards Evidence-Based Multi-Modal Medical Diagnosis via Reasoning Agentic WorkflowNot Availablesafety, evaluation, other
MedAgentAudit: Diagnosing and Quantifying Collaborative Failure Modes in Medical Multi-Agent SystemsGitHubcapability, safety, evaluation
MedAgentBench: Dataset for Benchmarking LLMs as AgentsGitHubevaluation
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical TasksGitHubevaluation
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical ReasoningGitHubevaluation
MedAgentSim: Self-evolving Multi-agent Simulations for Realistic Clinical InteractionsGitHubtask, other
MedBrowseComp: Benchmarking Medical Deep Research and Computer UseGitHubevaluation
Medchat: A multi-agent framework for multimodal diagnosis with large language modelsGitHubother
MedCoAct: Confidence-Aware Multi-Agent Collaboration for Complete Clinical DecisionNot Availablecapability, application
Meddxagent: A unified modular agent framework for explainable automatic differential diagnosisGitHubother
MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-Checking of LLM ResponsesGitHubevaluation
Medhallu: A comprehensive benchmark for detecting medical hallucinations in large language modelsGitHubsafety
Mediator-guided multi-agent collaboration among open-source models for medical decision-makingNot Availabletask, other
Medical AI Consensus: A Multi-Agent Framework for Radiology Report Generation and EvaluationNot Availablecapability, task
Medical hallucinations in foundation models and their impact on healthcareGitHubsafety
Position: Medical large language model benchmarks should prioritize construct validityNot Availableevaluation
MedicalOS: An {LLM} Agent based Operating System for Digital HealthcareNot Availablecapability, application
MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge GraphGitHubtask
MedKGEval: A Knowledge Graph-Based Multi-Turn Evaluation Framework for Open-Ended Patient Interactions with Clinical LLMsNot Availabletask
MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language ModelsGitHubapplication
Medmmv: A controllable multimodal multi-agent framework for reliable and verifiable clinical reasoningNot Availabletask, safety, evaluation, other
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible ExtensibilityNot Availabletask, other
MedPAO: A Protocol-Driven Agent for Structuring Medical ReportsGitHubcapability, task, application, other
Medrax: Medical reasoning agent for chest x-rayGitHubcapability, application, other
MedRepBench: A Comprehensive Benchmark for Medical Report InterpretationNot Availablecapability, evaluation
Medresearcher-r1: Expert-level medical deep researcher via a knowledge-informed trajectory synthesis frameworkGitHubevaluation
MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question AnsweringNot Availablecapability, task
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingGitHubevaluation
MEGA-RAG: a retrieval-augmented generation framework with multi-evidence guided answer refinement for mitigating hallucinations of LLMs in public healthNot Availablesafety
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical ReasoningNot Availabletask, other
MoMA: A Mixture-of-Multimodal-Agents Architecture for Enhancing Clinical Prediction ModellingNot Availablecapability, task, other
MRGAgents: A Multi-Agent Framework for Improved Medical Report Generation with Med-LVLMsNot Availabletask, other
Multi agent based medical assistant for edge devicesGitHubsafety
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationNot Availableevaluation
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in {LMIC}sNot Availablecapability, application
Multimodal Models in Healthcare: Methods, Challenges, and Future Directions for Enhanced Clinical Decision SupportNot Availableintro
NurseLLM: The First Specialized Language Model for NursingNot Availablecapability
OAAgent: Multimodal LLM Agent for Predicting Knee Osteoarthritis ProgressionNot Availablecapability
OpenLens AI: Fully Autonomous Research Agent for Health InfomaticsGitHubcapability, application
PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray ReasoningGitHubother
Pathfinder: A multi-modal multi-agent system for medical diagnostic decision-making applied to histopathologyGitHubtask
Patient-Zero: A Unified Framework for Real-Record-Free Patient Agent GenerationNot Availablecapability, application, safety
Performance of Retrieval-Augmented Generation Large Language Models in Guideline-Concordant Prostate-Specific Antigen Testing: Comparative Study With Junior CliniciansNot Availableevaluation
Privacy in action: Towards realistic privacy mitigation and evaluation for llm-powered agentsGitHubsafety
Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language ModelsGitHubcapability
Program Synthesis Dialog Agents for Interactive Decision-MakingNot Availablesafety
Proof-of-TBI--Fine-Tuned Vision Language Model Consortium and OpenAI-o3 Reasoning LLM-Based Medical Diagnosis Support System for Mild Traumatic Brain Injury (TBI) PredictionNot Availablecapability, application
Rapidly benchmarking large language models for diagnosing comorbid patients: comparative study leveraging the LLM-as-a-judge methodNot Availableevaluation
Real-World Evaluation of Large Language Models in Healthcare (RWE-LLM): A New Realm of AI Safety & ValidationNot Availableevaluation
Red-teaming llm multi-agent systems via communication attacksNot Availablesafety
Reducing Hallucinations and Trade-Offs in Responses in Generative AI Chatbots for Cancer Information: Development and Evaluation StudyNot Availablesafety
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical AgentsGitHubcapability, application
Resilient Multi-Agent Negotiation for Medical Supply Chains: Integrating LLMs and Blockchain for Transparent CoordinationNot Availabletask
RxLens: Multi-Agent LLM-powered Scan and Order for PharmacyNot Availablecapability, application, other
SCOPE: Speech-Guided COllaborative PErception Framework for Surgical Scene SegmentationNot Availabletask, other
Self-Assessment of Content, Pedagogy, and Technology Knowledge among Higher Education Academics in BahrainNot Availablecapability
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision MakingNot Availablecapability
SmartState: An Automated Research Protocol Adherence SystemGitHubcapability, application
SOLVE-Med: Specialized Orchestration for Leading Vertical Experts across Medical SpecialtiesGitHubtask
Standard Applicability Judgment and Cross-jurisdictional Reasoning: A RAG-based Framework for Medical Device ComplianceNot Availablecapability
Surgraw: Multi-agent workflow with chain-of-thought reasoning for surgical intelligenceGitHubtask
Survey and improvement strategies for gene prioritization with large language modelsNot Availabletask, other
Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QAGitHubcapability
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM SystemsNot Availablesafety
The evaluation illusion of large language models in medicineNot Availableevaluation
Tiered Agentic Oversight: A Hierarchical Multi-Agent System for AI Safety in HealthcareGitHubcapability, safety
Tool learning with large language models: A surveyGitHubintro
Towards conversational diagnostic artificial intelligenceNot Availabletask
Towards interpretable radiology report generation via concept bottlenecks using a multi-agentic ragGitHubtask
Towards safe ai clinicians: A comprehensive study on large language model jailbreaking in healthcareNot Availablesafety
Transforming healthcare delivery with conversational AI platformsNot Availablesafety
Tree-based RAG-Agent Recommendation System: A Case Study in Medical Test DataNot Availablecapability, task, application
Tree-of-Reasoning: Towards Complex Medical Diagnosis via Multi-Agent Reasoning with Evidence TreeGitHubsafety
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought ProcessesNot Availableevaluation
TxAgent: An AI agent for therapeutic reasoning across a universe of toolsGitHubintro, task, other
Using large language models for enhanced fraud analysis and detection in blockchain based health insurance claimsGitHubcapability, application
Vision-language model for report generation and outcome prediction in CT pulmonary angiogramGitHubcapability, application
Visual-Conversational Interface for Evidence-Based Explanation of Diabetes Risk PredictionGitHubintro
When Avatars Have Personality: Effects on Engagement and Communication in Immersive Medical TrainingNot Availableapplication
World Model for AI Autonomous Navigation in Mechanical ThrombectomyGitHubcapability, task, application
Zero-Shot Large Language Model Agents for Fully Automated Radiotherapy Treatment PlanningNot Availablecapability, application, other
Bias-Aware Agent: Enhancing Fairness in AI-Driven Knowledge RetrievalGitHubsafety
EMR-AGENT: Automating Cohort and Feature Extraction from EMR DatabasesGitHubcapability, task, application
Colacare: Enhancing electronic health record modeling through large language model-driven multi-agent collaborationGitHubcapability, application
A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-makingGitHubcapability, application
MeNTi: Bridging medical calculator and LLM agent with nested tool callingGitHubcapability, application
Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAGGitHubapplication, other
How Can We Diagnose and Treat Bias in Large Language Models for Clinical Decision-Making?Not Availableevaluation

🚀 Year-2024

TitleGitHubSections
A demonstration of adaptive collaboration of large language models for medical decision-makingGitHubcapability, application
A survey on large language model based autonomous agentsNot Availableintro
Achieving health equity through conversational AI: A roadmap for design and implementation of inclusive chatbots in healthcareNot Availablesafety
Adaptive Reasoning and Acting in Medical Language AgentsNot Availablecapability, application
Adversarial attacks on large language models in medicineNot Availablesafety
Agent Hospital: A Simulacrum of Hospital with Evolvable Medical AgentsGitHubcapability, task, application
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environmentsGitHubintro, capability, application, evaluation, other
Agentic llm workflows for generating patient-friendly medical reportsGitHubcapability, task, application, other
Agentigraph: An interactive knowledge graph platform for llm-based chatbots utilizing private dataGitHubcapability, application
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction SimulatorGitHubcapability, task, application, evaluation
Aligning Medical LLMs for Counterfactual FairnessGitHubsafety
ArgMed-Agents: explainable clinical decision reasoning with LLM disscusion via argumentation schemesNot Availablecapability, application
Autohealth: Advanced llm-empowered wearable personalized medical butler for parkinson’s disease managementNot Availableother
Autonomous artificial intelligence agents for clinical decision making in oncologyNot Availablecapability, other
Benchmarking Large Language Models on Communicative Medical Coaching: A Dataset and a Novel SystemGitHubcapability, application
Beyond direct diagnosis: LLM-based multi-specialist agent consultation for automatic diagnosisNot Availablecapability
Chatdev: Communicative agents for software developmentGitHubtask
ClinicalAgent: Clinical Trial Multi-Agent System with Large Language Model-based ReasoningGitHubintro, capability, application, other
ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real WorldGitHubcapability, evaluation
Cxr-agent: Vision-language models for chest x-ray interpretation with uncertainty aware radiology reportingNot Availablecapability, application, other
Development of a Large Language Model-based Multi-Agent Clinical Decision Support System for Korean Triage and Acuity Scale (KTAS)-Based Triage and Treatment Planning in Emergency DepartmentsGitHubcapability, task, other
Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health recordsGitHubcapability, task, application
Enhancing diagnostic accuracy through multi-agent conversations: using large language models to mitigate cognitive biasNot Availableapplication
Enhancing llms for impression generation in radiology reports through a multi-agent systemGitHubcapability, task, application, other
Ethical and regulatory challenges of large language models in medicineNot Availablesafety
Evaluating Large Language Models as Agents in the ClinicNot Availablecapability, evaluation
Exploring llm multi-agents for icd codingNot Availableapplication
Exploring LLM-based Data Annotation Strategies for Medical Dialogue Preference AlignmentNot Availablecapability
Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative AINot Availablesafety
GuidelineGuard: An Agentic Framework for Medical Note Evaluation with Guideline AdherenceNot Availablecapability, application
Imas: A comprehensive agentic approach to rural healthcare deliveryGitHubcapability, application
Improving Clinical Documentation with AI: A Comparative Study of Sporo AI Scribe and GPT-4o miniNot Availablecapability, task, application
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical ReasoningGitHubcapability, application
Integration of multi-source medical data for medical diagnosis question answeringGitHubcapability, application
Iryonlp at mediqa-corr 2024: Tackling the medical error detection & correction task on the shoulders of medical agentsNot Availablecapability, task, application
KG4Diagnosis: A Hierarchical Multi-Agent LLM Framework with Knowledge Graph Enhancement for Medical DiagnosisNot Availablecapability, application
Knowledge-infused llm-powered conversational health agent: A case study for diabetes patientsNot Availablecapability
Large Language Model-Enhanced Interactive Agent for Public Education on Newborn Auricular DeformitiesNot Availablecapability, application
Llm-based framework for administrative task automation in healthcareNot Availablecapability, application
Llm-medqa: Enhancing medical question answering through case studies in large language modelsNot Availablecapability, application
MAGDA: Multi-agent guideline-driven diagnostic assistanceNot Availablecapability, task, application
MALADE: orchestration of LLM-powered agents with retrieval augmented generation for pharmacovigilanceGitHubintro, capability, application, other
Mdagents: An adaptive collaboration of llms for medical decision-makingGitHubcapability, task, application
MedAgents: Large Language Models as Collaborators for Zero-shot Medical ReasoningGitHubcapability, task, other
MedAide: Towards an Omni Medical Aide via Specialized {LLM}-based Multi-Agent CollaborationNot Availablecapability
Medco: Medical education copilots based on a multi-agent frameworkNot Availablecapability, task, application
MedChain: Bridging the Gap Between LLM Agents and Clinical Practice through Interactive Sequential BenchmarkingGitHubcapability, application, evaluation
MedGen: An Explainable Multi-Agent Architecture for Clinical Decision Support through Multisource Knowledge FusionNot Availablecapability, application
Medhalu: Hallucinations in responses to healthcare queries by large language modelsNot Availablesafety
MedQA-CS: Benchmarking Large Language Models’ Clinical Skills Using an AI-SCE FrameworkGitHubevaluation
Mitigating cognitive biases in clinical decision-making through multi-agent conversations using large language models: simulation studyNot Availablecapability, application
Mitigating hallucinations in large language models via self-refinement-enhanced knowledge retrievalNot Availablesafety
Mmedagent: Learning to use medical tools with multi-modal agentGitHubcapability, application, other
MMLU-Pro: A More Robust Benchmark for Multi-Task Language UnderstandingGitHubevaluation
Natural Language Programming in Medicine: Administering Evidence Based Clinical Workflows with Autonomous Agents Powered by Generative Large Language ModelsNot Availablecapability, application
On protecting the data privacy of large language models (llms): A surveyNot Availablesafety
On the resilience of llm-based multi-agent collaboration with faulty agentsNot Availablesafety
Piors: Personalized intelligent outpatient reception based on large language model with multi-agents medical scenario simulationGitHubcapability, task, application
Polaris: A safety-focused llm constellation architecture for healthcareNot Availablecapability, application, evaluation
Privacy-Preserving Large Language Models: MechanismsNot Availablesafety
Advancing healthcare automation: Multi-agent system for medical necessity justificationNot Availablecapability, task, application
RareAgents: Advancing Rare Disease Care through LLM-Empowered Multi-disciplinary TeamNot Availablecapability, application
RareAgents: Autonomous Multi-disciplinary Team for Rare Disease Diagnosis and TreatmentNot Availabletask, other
Regulator-manufacturer AI agents modeling: Mathematical feedback-driven multi-agent LLM frameworkNot availablecapability, task, application
Remoni: An autonomous system integrating wearables and multimodal large language models for enhanced remote health monitoringNot availablecapability, application
Rx strategist: Prescription verification using llm agents systemNot availablecapability, task, application, other
Simulated patient systems are intelligent when powered by large language model-based AI agentsNot availablecapability, application
Smile: Single-turn to multi-turn inclusive language expansion via chatgpt for mental health supportGithubtask
Society of medical simplifiersGitHubcapability, task, application
Surgbox: Agent-driven operating room sandbox with surgery copilotNot availablecapability, task, application
T-agent: A term-aware agent for medical dialogue generationNot availablecapability, application
The role of explainability in AI-supported medical decision-makingNot availablesafety
Towards anatomy education with generative AI-based virtual assistants in immersive virtual reality environmentsNot availableapplication
Towards Automatic Evaluation for {LLM}s' Clinical Capabilities: Metric, Data, and AlgorithmNot availablecapability, application, evaluation
Towards next-generation medical agent: How o1 is reshaping decision-making in medical scenariosNot availablecapability, other
A trustworthy AI reality-check: the lack of transparency of artificial intelligence products in healthcareNot availablesafety
TWIN-GPT: digital twins for clinical trials via large language modelNot availableother
UMass-BioNLP at MEDIQA-M3G 2024: DermPrompt--A Systematic Exploration of Prompt Engineering with GPT-4V for Dermatological DiagnosisGitHubcapability, other
Zodiac: A cardiologist-level llm framework for multi-agent diagnosticsNot availablecapability
OpenAI o1 System CardGitHubcapability
FedAgentBench: Towards Automating Real-World Federated Medical Image Analysis with Server–Client LLM AgentsNot availableevaluation
Evaluating large language models as agents in the clinicNot available

🚀 Year-2023

TitleGitHubSections
A reinforcement learning approach for VQA validation: An application to diabetic macular edema gradingNot availablecapability
Adaptive multi-agent deep reinforcement learning for timely healthcare interventionsNot availablecapability, application
Asynchronous decentralized federated lifelong learning for landmark localization in medical imagingNot availablecapability
Beyond memorization: Violating privacy via inference with large language modelsNot availablesafety
Camel: Communicative agents for" mind" exploration of large language model societyNot availabletask
Clinically-inspired multi-agent transformers for disease trajectory forecasting from multimodal dataGitHubcapability, application
Cognitive architectures for language agentsNot availableintro
Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media ForumNot availableevaluation
Deep Imitation Learning for Automated Drop-In Gamma Probe ManipulationNot availabletask
Diaggpt: An llm-based and multi-agent dialogue system with automatic topic management for flexible task-oriented dialogueGitHubapplication
Dspy: Compiling declarative language model calls into self-improving pipelinesGitHubtask
Federated machine learning, privacy-enhancing technologies, and data protection laws in medical research: scoping reviewNot availablesafety
Generative agents: Interactive simulacra of human behaviorGitHubintro
Interactive medical image segmentation with self-adaptive confidence calibrationGitHubtask
Large language models as agents in the clinicNot availableapplication
MetaGPT: Meta programming for a multi-agent collaborative frameworkGitHubtask
Navigation Through Endoluminal Channels Using Q-LearningNot availablecapability, task
PROFSA: SELF-SUPERVISED POCKET PRETRAINING VIA PROTEIN FRAGMENT-SURROUNDINGS ALIGNGitHubcapability
Reflexion: Language agents with verbal reinforcement learningGitHubcapability
Td-mpc2: Scalable, robust world models for continuous controlGitHubtask
Temporally-extended prompts optimization for sam in interactive medical image segmentationNot availablecapability, task
The NCI Imaging Data Commons as a platform for reproducible research in computational pathologyNot availableintro
Towards Causality-Aware Inferring: A Sequential Discriminative Approach for Medical DiagnosisNot availablecapability, application

🚀 Earlier

TitleGitHubSections
"My Nose is Running." "Are You Also Coughing?": Building a Medical Diagnosis Agent with Interpretable Inquiry LogicsGitHubcapability, application
A Flexible Schema-Guided Dialogue Management Framework: From Friendly Peer to Virtual Standardized Cancer PatientGitHubcapability, application
Building an {ASR} Error Robust Spoken Virtual Patient System in a Highly Class-Imbalanced Scenario Without Speech DataNot availablecapability, application
Constitutional ai: Harmlessness from ai feedbackGitHubcapability
MedDG: An Entity-Centric Medical Consultation Dataset for Entity-Aware Medical Dialogue GenerationGitHubevaluation
MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoningNot availableintro
Multi-agent searching system for medical informationNot availablecapability
React: Synergizing reasoning and acting in language modelsGitHubintro
Scalable Online Disease Diagnosis via Multi-Model-Fused Actor-Critic Reinforcement LearningNot availablecapability, application
MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical Domain Question AnsweringGitHubcapability, evaluation
A grounded well-being conversational agent with multiple interaction modes: Preliminary resultsNot availablecapability, application
Adaptable image quality assessment using meta-reinforcement learning of task amenabilityGitHubcapability
An Edge Based Multi-Agent Model for Improving Hospital Bed ManagementNot availabletask
Cross Modality 3D Navigation Using Reinforcement Learning and Neural Style TransferNot availablecapability
Extracting Training Data from Large Language ModelsGitHubsafety
Human-AI collaboration in healthcare: A review and research agendaNot availabletask
Levels of autonomy and safety assurance for AI-Based clinical decision systemsNot availablesafety
Measuring Massive Multitask Language UnderstandingGitHubevaluation
Autonomous systems and artificial intelligence in healthcare transformation to 5P medicine--ethical challengesNot availableintro
Boundary-aware supervoxel-level iteratively refined interactive 3d image segmentation with multi-agent reinforcement learningNot availabletask
MedDialog: A Large-scale Medical Dialogue DatasetNot availableevaluation
Medical visual question answering via conditional reasoningGitHubtask
PathVQA: 30000+ Questions for Medical Visual Question AnsweringGitHubevaluation
PubMedQA: A Dataset for Biomedical Research Question AnsweringGitHubevaluation
A Dataset of Clinically Generated Visual Questions and Answers About Radiology ImagesNot availableevaluation
Modeling irregularly sampled clinical time seriesGitHubintro
The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive careGitHubtask
Medical robotics—Regulatory, ethical, and legal considerations for increasing levels of autonomyNot availabletask
Tethered to the EHR: primary care physician workload assessment using EHR event log data and time-motion observationsNot availableintro
Allocation of physician time in ambulatory practice: a time and motion study in 4 specialtiesNot availableintro
Assessing electronic note quality using the physician documentation quality instrument (PDQI-9)Not availabletask
Privacy by design: The 7 foundational principlesNot availablesafety
Upper processing stages of the perception--action cycleNot availableintro
What Disease Does This Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical ExamsGitHubevaluation

✨ Papers by Category

🚀 1. Capability

1.1 Planning

TitleGitHubYear
Evaluating Large Language Models as Agents in the ClinicNot Available2024
MedPAO: A Protocol-Driven Agent for Structuring Medical ReportsGitHub2025
Adaptable image quality assessment using meta-reinforcement learning of task amenabilityGitHub2023
World Model for AI Autonomous Navigation in Mechanical ThrombectomyGitHub2025
Polaris: A safety-focused llm constellation architecture for healthcareNot Available2024
MedChain: Bridging the Gap Between LLM Agents and Clinical Practice through Interactive Sequential BenchmarkingGitHub2024
Rx strategist: Prescription verification using llm agents systemNot Available2024
A Flexible Schema-Guided Dialogue Management Framework: From Friendly Peer to Virtual Standardized Cancer PatientGitHub2023
Surgbox: Agent-driven operating room sandbox with surgery copilotNot Available2024
MedicalOS: An {LLM} Agent based Operating System for Digital HealthcareNot Available2025
"My Nose is Running." "Are You Also Coughing?": Building a Medical Diagnosis Agent with Interpretable Inquiry LogicsGitHub2023
Cross Modality 3D Navigation Using Reinforcement Learning and Neural Style TransferNot Available2023
Scalable Online Disease Diagnosis via Multi-Model-Fused Actor-Critic Reinforcement LearningNot Available2023
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction SimulatorGitHub2024
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical ReasoningGitHub2024
Llm-medqa: Enhancing medical question answering through case studies in large language modelsNot Available2024
Medco: Medical education copilots based on a multi-agent frameworkNot Available2024
Mitigating cognitive biases in clinical decision-making through multi-agent conversations using large language models: simulation studyNot Available2024
Regulator-manufacturer AI agents modeling: Mathematical feedback-driven multi-agent LLM frameworkNot Available2024
Society of medical simplifiersGitHub2024
Towards Automatic Evaluation for {LLM}s' Clinical Capabilities: Metric, Data, and AlgorithmNot Available2024
Towards next-generation medical agent: How o1 is reshaping decision-making in medical scenariosNot Available2024
A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-makingGitHub2025
A dual-agent collaboration framework based on llms for nursing robots to perform bimanual coordination tasksNot Available2025
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical KnowledgeGitHub2025
Agentic-AI Healthcare: Multilingual, Privacy-First Framework with {MCP} AgentsGitHub2025
Autonomous Multi-Modal LLM Agents for Treatment Planning in Focused Ultrasound Ablation SurgeryGitHub2025
Colacare: Enhancing electronic health record modeling through large language model-driven multi-agent collaborationGitHub2025
Haibu Mathematical-Medical Intelligent Agent: Enhancing Large Language Model Reliability in Medical Tasks via Verifiable Reasoning ChainsNot Available2025
Improving Interactive Diagnostic Ability of a Large Language Model Agent Through Clinical Experience LearningGitHub2025
Learning to be a doctor: Searching for effective medical agent architecturesNot Available2025
MeNTi: Bridging medical calculator and LLM agent with nested tool callingGitHub2025
OpenLens AI: Fully Autonomous Research Agent for Health InfomaticsGitHub2025
Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QAGitHub2025
Tree-based RAG-Agent Recommendation System: A Case Study in Medical Test DataNot Available2025
AI chatbots as professional service agents: developing a professional identityNot Available2025
A reinforcement learning approach for VQA validation: An application to diabetic macular edema gradingNot Available2023
Building an {ASR} Error Robust Spoken Virtual Patient System in a Highly Class-Imbalanced Scenario Without Speech DataNot Available2023
MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical Domain Question AnsweringGitHub2023
Multi-agent searching system for medical informationNot Available2023
A demonstration of adaptive collaboration of large language models for medical decision-makingGitHub2024
Agentic llm workflows for generating patient-friendly medical reportsGitHub2024
Agentigraph: An interactive knowledge graph platform for llm-based chatbots utilizing private dataGitHub2024
Improving Clinical Documentation with AI: A Comparative Study of Sporo AI Scribe and GPT-4o miniNot Available2024
Piors: Personalized intelligent outpatient reception based on large language model with multi-agents medical scenario simulationGitHub2024
SmartState: An Automated Research Protocol Adherence SystemGitHub2025

1.2 Tool Use

TitleGitHubYear
ClinicalAgent: Clinical Trial Multi-Agent System with Large Language Model-based ReasoningGitHub2024
Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement LearningGitHub2025
KERAP: A Knowledge-Enhanced Reasoning Approach for Accurate Zero-shot Diagnosis Prediction Using Multi-agent LLMsGitHub2025
Iryonlp at mediqa-corr 2024: Tackling the medical error detection & correction task on the shoulders of medical agentsNot Available2024
Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language ModelsGitHub2025
Cxr-agent: Vision-language models for chest x-ray interpretation with uncertainty aware radiology reportingNot Available2024
Enhancing llms for impression generation in radiology reports through a multi-agent systemGitHub2024
GuidelineGuard: An Agentic Framework for Medical Note Evaluation with Guideline AdherenceNot Available2024
Medrax: Medical reasoning agent for chest x-rayGitHub2025
Mmedagent: Learning to use medical tools with multi-modal agentGitHub2024
Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced AgentsGitHub2025
MedAgents: Large Language Models as Collaborators for Zero-shot Medical ReasoningGitHub2024
AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question AnsweringGitHub2025
NurseLLM: The First Specialized Language Model for NursingNot Available2025
A multimodal AI agent for clinical decision support in ophthalmologyNot Available2025
Autonomous artificial intelligence agents for clinical decision making in oncologyNot Available2024
ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real WorldGitHub2024
Llm-based framework for administrative task automation in healthcareNot Available2024
A Multi-Agent Approach to Neurological Clinical ReasoningNot Available2025
Adagent: Llm agent for alzheimer’s disease analysis with collaborative coordinatorNot Available2025
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical AgentsGitHub2025
Development of a Large Language Model-based Multi-Agent Clinical Decision Support System for Korean Triage and Acuity Scale (KTAS)-Based Triage and Treatment Planning in Emergency DepartmentsGitHub2024
Large Language Model-Enhanced Interactive Agent for Public Education on Newborn Auricular DeformitiesNot Available2024
Natural Language Programming in Medicine: Administering Evidence Based Clinical Workflows with Autonomous Agents Powered by Generative Large Language ModelsNot Available2024
Agentmd: Empowering language agents for risk prediction with large-scale clinical tool learningGitHub2025
MoMA: A Mixture-of-Multimodal-Agents Architecture for Enhancing Clinical Prediction ModellingNot Available2025
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in {LMIC}sNot Available2025

1.3 Memory

1.4 Self-Improvement

1.5 Reasoning

TitleGitHubYear
Proof-of-TBI--Fine-Tuned Vision Language Model Consortium and OpenAI-o3 Reasoning LLM-Based Medical Diagnosis Support System for Mild Traumatic Brain Injury (TBI) PredictionNot Available2025
OpenAI o1 System CardGitHub2024
Beyond direct diagnosis: LLM-based multi-specialist agent consultation for automatic diagnosisNot Available2024
MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic WorkflowGitHub2025
MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question AnsweringNot Available2025
KG4Diagnosis: A Hierarchical Multi-Agent LLM Framework with Knowledge Graph Enhancement for Medical DiagnosisNot Available2024
MAGDA: Multi-agent guideline-driven diagnostic assistanceNot Available2024
Zero-Shot Large Language Model Agents for Fully Automated Radiotherapy Treatment PlanningNot Available2025
Cod, towards an interpretable medical agent using chain of diagnosisGitHub2025
Asynchronous decentralized federated lifelong learning for landmark localization in medical imagingNot Available2023
Temporally-extended prompts optimization for sam in interactive medical image segmentationNot Available2023
Integration of multi-source medical data for medical diagnosis question answeringGitHub2024
MedGen: An Explainable Multi-Agent Architecture for Clinical Decision Support through Multisource Knowledge FusionNot Available2024
Zodiac: A cardiologist-level llm framework for multi-agent diagnosticsNot Available2024
An active inference strategy for prompting reliable responses from large language models in medical practiceNot Available2025
CARE-AD: a multi-agent large language model framework for Alzheimer’s disease prediction using longitudinal clinical notesNot Available2025
Feat: A multi-agent forensic ai system with domain-adapted large language model for automated cause-of-death analysisGitHub2025
Fine-tuning vision language models with graph-based knowledge for explainable medical image analysisNot Available2025
Advancing healthcare automation: Multi-agent system for medical necessity justificationNot Available2024
ArgMed-Agents: explainable clinical decision reasoning with LLM disscusion via argumentation schemesNot Available2024
Imas: A comprehensive agentic approach to rural healthcare deliveryGitHub2024
Simulated patient systems are intelligent when powered by large language model-based AI agentsNot Available2024
UMass-BioNLP at MEDIQA-M3G 2024: DermPrompt--A Systematic Exploration of Prompt Engineering with GPT-4V for Dermatological DiagnosisGitHub2024
CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery PlanningNot Available2025
Dr. Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in RomanianNot Available2025
Enhancing Medical Lung X-Ray Diagnosis Through Multi-Agent Vision-Language Model CollaborationNot Available2025
Enhancing diagnostic capability with multi-agents conversational large language modelsGitHub2025
LINS: A general medical Q&A framework for enhancing the quality and credibility of LLM-generated responsesGitHub2025
OAAgent: Multimodal LLM Agent for Predicting Knee Osteoarthritis ProgressionNot Available2025
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision MakingNot Available2025
Standard Applicability Judgment and Cross-jurisdictional Reasoning: A RAG-based Framework for Medical Device ComplianceNot Available2025

1.6 Perception

1.7 Others (continual learning, uncertainty)

🚀 2. Atomic Function

2.1 Basic Technology Empowerment

2.2 Core Diagnostic & Therapeutic Assistance

2.3 Workflow & Documentation Optimization

🚀 3. Application

3.1 Intake & Clinical Dialogue

3.2 Virtual MDT Teams & Multimodal Reasoning

3.3 Treatment Procedures

3.4 Chronic Disease Management & Prescription Safety

3.5 Documentation, Coding & Knowledge Infrastructure

3.6 Simulation & Support Systems

3.7 Regulation, Payer Workflows & Administrative Automation

🚀 4. Safety

4.1 Medical Hallucination

4.2 Privacy & Data-Security

4.3 Explainability & Transparency

4.4 Adversarial Security & Threat Modeling

4.5 AI Governance & Systemic Safety

4.6 Bias, Fairness & Accessibility

🚀 5. Evaluation

5.1 Benchmarks

TitleGitHubYear
Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal ReasoningGitHub2025
A Dataset of Clinically Generated Visual Questions and Answers About Radiology ImagesNot Available2023
Measuring Massive Multitask Language UnderstandingGitHub2023
MedDG: An Entity-Centric Medical Consultation Dataset for Entity-Aware Medical Dialogue GenerationGitHub2023
MedDialog: A Large-scale Medical Dialogue DatasetNot Available2023
PathVQA: 30000+ Questions for Medical Visual Question AnsweringGitHub2023
PubMedQA: A Dataset for Biomedical Research Question AnsweringGitHub2023
FedAgentBench: Towards Automating Real-World Federated Medical Image Analysis with Server–Client LLM AgentsNot Available2024
MMLU-Pro: A More Robust Benchmark for Multi-Task Language UnderstandingGitHub2024
MedQA-CS: Benchmarking Large Language Models’ Clinical Skills Using an AI-SCE FrameworkGitHub2024
3mdbench: Medical multimodal multi-agent dialogue benchmarkGitHub2025
AMQA: An Adversarial Dataset for Benchmarking Bias of LLMs in Medicine and HealthcareGitHub2025
Beyond Benchmarks: Dynamic, Automatic and Systematic Red-Teaming Agents for Trustworthy Medical Language ModelsGitHub2025
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical TasksGitHub2025
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical ReasoningGitHub2025
MedBrowseComp: Benchmarking Medical Deep Research and Computer UseGitHub2025
MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-Checking of LLM ResponsesGitHub2025
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingGitHub2025
Medresearcher-r1: Expert-level medical deep researcher via a knowledge-informed trajectory synthesis frameworkGitHub2025

5.2 Metrics

TitleGitHubYear
Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media ForumNot Available2023
What Disease Does This Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical ExamsGitHub2023
AI Agents in Clinical Medicine: A Systematic ReviewNot Available2025
Audited Reasoning Refinement: Fine-Tuning Language Models via LLM-Guided Step-Wise Evaluation and CorrectionNot Available2025
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology ReasoningNot Available2025
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in MedicineNot Available2025
Evaluating the accuracy of a state-of-the-art large language model for prediction of admissions from the emergency roomNot Available2025
Geometry-preserving encoder/decoder in latent generative modelsGitHub2025
How Can We Diagnose and Treat Bias in Large Language Models for Clinical Decision-Making?Not Available2025
Large language models in real-world clinical workflows: a systematic review of applications and implementationNot Available2025
MedAgentBench: Dataset for Benchmarking LLMs as AgentsGitHub2025
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationNot Available2025
Performance of Retrieval-Augmented Generation Large Language Models in Guideline-Concordant Prostate-Specific Antigen Testing: Comparative Study With Junior CliniciansNot Available2025
Real-World Evaluation of Large Language Models in Healthcare (RWE-LLM): A New Realm of AI Safety & ValidationNot Available2025
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought ProcessesNot Available2025

5.3 Challenge & Discussion

🚀 6. Communication & Collaboration Mechanisms

🚀 7. Others

TitleGitHubYear
Drugagent: Multi-agent large language model-based reasoning for drug-target interaction predictionNot Available2025
An Edge Based Multi-Agent Model for Improving Hospital Bed ManagementNot Available2023
DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical servicesNot Available2025
Lessons Learned from Evaluation of LLM based Multi-agents in Safer Therapy RecommendationGitHub2025
M3Builder: A Multi-Agent System for Automated Machine Learning in Medical ImagingGitHub2025
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical ReasoningNot Available2025
Resilient Multi-Agent Negotiation for Medical Supply Chains: Integrating LLMs and Blockchain for Transparent CoordinationNot Available2025
Survey and improvement strategies for gene prioritization with large language modelsNot Available2025
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible ExtensibilityNot Available2025
Demo: Healthcare Agent Orchestrator (HAO) for Patient Summarization in Molecular Tumor BoardsNot Available2025
Mediator-guided multi-agent collaboration among open-source models for medical decision-makingNot Available2025
A co-evolving agentic AI system for medical imaging analysisGitHub2025
FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research InsightsGitHub2025
Meddxagent: A unified modular agent framework for explainable automatic differential diagnosisGitHub2025
AURA: A Multi-modal Medical Agent for Understanding, Reasoning and AnnotationGitHub2025
GEMA-Score: Granular Explainable Multi-Agent Score for Radiology Report EvaluationGitHub2025
PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray ReasoningGitHub2025
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False PresuppositionsGitHub2025
EH-Benchmark: Ophthalmic hallucination benchmark and agent-driven top-down traceable reasoning workflowGitHub2025
Medchat: A multi-agent framework for multimodal diagnosis with large language modelsGitHub2025
Autohealth: Advanced llm-empowered wearable personalized medical butler for parkinson’s disease managementNot Available2024
TWIN-GPT: digital twins for clinical trials via large language modelNot Available2024
A Proposed LLM-Based Supported Treatment Framework for Intracerebral HemorrhageNot Available2025
Agentic AI for Clinical Decision Support: Real-Time Diagnosis, Triage, and Treatment PlanningNot Available2025
Agentic Workflows in Healthcare: Advancing Clinical Efficiency through AI IntegrationNot Available2025
Developing an Artificial Intelligence Tool for Personalized Breast Cancer Treatment Plans based on the NCCN GuidelinesNot Available2025
TxAgent: An AI agent for therapeutic reasoning across a universe of toolsGitHub2025
Allocation of physician time in ambulatory practice: a time and motion study in 4 specialtiesNot Available2023
Tethered to the EHR: primary care physician workload assessment using EHR event log data and time-motion observationsNot Available2023
The NCI Imaging Data Commons as a platform for reproducible research in computational pathologyNot Available2023
At-cxr: Uncertainty-aware agentic triage for chest x-raysGitHub2025
In-Basket Message Volume in Primary Care: A Cross-sectional Analysis by Gender and SpecialtyNot Available2025
Modeling irregularly sampled clinical time seriesGitHub2023
Multimodal Models in Healthcare: Methods, Challenges, and Future Directions for Enhanced Clinical Decision SupportNot Available2025
Visual-Conversational Interface for Evidence-Based Explanation of Diabetes Risk PredictionGitHub2025
Autonomous systems and artificial intelligence in healthcare transformation to 5P medicine--ethical challengesNot Available2023
Cognitive architectures for language agentsNot Available2023
Generative agents: Interactive simulacra of human behaviorGitHub2023
MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoningNot Available2023
React: Synergizing reasoning and acting in language modelsGitHub2023
Upper processing stages of the perception--action cycleNot Available2023
A survey on large language model based autonomous agentsNot Available2024
Tool learning with large language models: A surveyGitHub2025

Contributors

NUS-Project

183 commits

eolasdfjnas

5 commits

Data-Designer

2 commits

yunhang8658

2 commits

NUS-Project/Landmark-of-medical-agent

HTML

186

193 commits

updated Jun 8, 2026

See the code

README

🚀 The Landscape of Medical Agents: A Survey

🚀 MedMASLab:A Framework for Multimodal Medical Multi-Agent Systems

Overall Landscape

Overall Landscape

🌟 Overview

This is the official repository for the survey paper: The Landscape of Medical Agents. This repository is a comprehensive and systematic research resource library for medical agents, dedicated to organizing and tracking the latest research progress, application practices, and technological developments of AI intelligent agents in the medical and health field. This investigative project covers the entire ecosystem from basic technical capabilities to clinical actual deployment, providing an authoritative research map for medical AI researchers, clinical practitioners, and system developers.

🔥 News

[2026/3/10] We release :A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems! Click here to visit: medmaslab

🔥 Add Your Paper in our Repo and Survey!!!!! Join our group and post your agent work!!!!

[-] You are welcome to give us an issue or PR for your medical agent work !!!!!

[-] Note that: Due to the huge paper in Arxiv, we are sorry to cover all in our survey. You can directly present a PR into this repo and we will record it for next version update of our survey.

[-] Our survey will be updated in 2026.3.

🤝 Thanks

If you think this project is useful and inspiring, we would greatly appreciate it if you could give us a Star to show your support! Your support is of great significance to us, as it encourages us to continue improving and developing this project.

📖 Keywords

Medical Agents, Clinical Workflows, Safety, Governance and Evaluation

🌟 Contributing

We will try to keep this list updated. If you find any errors or any missed paper, please don't hesitate to open issues or pull request.Please follow the instruction in CONTRIBUTING.md if you want to make one. Additionally, if you want to have any other issue, please add this wechat group.

🤝 Main Contacts

Citation

 @article{hu2025landscape,
  title={The Landscape of Medical Agents: A Survey},
  author={Hu, Xiaobin and Qian, Yunhang and Yu, Jiaquan and Liu, Jingjing and Tang, Peng and Ji, Xiaozhong and Xu, Chengming and Liu, Jiawei and Yan, Xiaoxiao and Yu, Xinlei and others},
  journal={Authorea Preprints},
  year={2025},
  publisher={Authorea}
}
@misc{qian2026medmaslabunifiedorchestrationframework,
      title={MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems}, 
      author={Yunhang Qian and Xiaobin Hu and Jiaquan Yu and Siyang Xin and Xiaokun Chen and Jiangning Zhang and Peng-Tao Jiang and Jiawei Liu and Hongwei Bran Li},
      year={2026},
      eprint={2603.09909},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2603.09909}, 
}

🌟 Table of Contents

✨ Latest Papers

🚀 Year-2026

January

TitlePaper-linkSections
EvoClinician: A Self-Evolving Agent for Multi-Turn Medical Diagnosis via Test-Time Evolutionary LearningPaperbenchmark
Structure-constrained Language-informed Diffusion Model for Unpaired Low-dose Computed Tomography Angiography ReconstructionPaperapplication
Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement LearningPaperbenchmark
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive DiagnosisPaperbenchmark
DEEPMED: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & InferencePaperbenchmark
Bayesian Multiple Testing for Suicide Risk in Pharmacoepidemiology: Leveraging Co-Prescription PatternsPaperframework
AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent ReasoningPaperbenchmark
Automated Rubrics for Reliable Evaluation of Medical Dialogue SystemsPaperframework
Query-Efficient Agentic Graph Extraction Attacks on GraphRAG SystemsPaperbenchmark
HyperWalker: Dynamic Hypergraph-Based Deep Diagnosis for Multi-Hop Clinical Modeling across EHR and X-Ray in Medical VLMsPaperframework
AgentEHR: Advancing Autonomous Clinical Decision-Making via Retrospective SummarizationPaperbenchmark
Towards Efficient and Robust Linguistic Emotion Diagnosis for Mental Health via Multi-Agent Instruction RefinementPaperbenchmark
Improving the Safety and Trustworthiness of Medical AI via Multi-Agent Evaluation LoopsPaperframework
MedConsultBench: A Full-Cycle, Fine-Grained, Process-Aware Benchmark for Medical Consultation AgentsPaperbenchmark
Knowing When to Abstain: Medical LLMs Under Clinical UncertaintyPaperbenchmark
A multitask framework for automated interpretation of multi-frame right upper quadrant ultrasound in clinical decision supportPaperbenchmark
Medication counseling with large language models: balancing flexibility and rigidityPaperbenchmark
MMedExpert-R1: Strengthening Multimodal Medical Reasoning via Domain-Specific Adaptation and Clinical Guideline ReinforcementPaperbenchmark
Japanese AI Agent System on Human Papillomavirus Vaccination: System DesignPaperframework
ART: Action-based Reasoning Task Benchmarking for Medical AI AgentsPaperbenchmark
Route, Retrieve, Reflect, Repair: Self-Improving Agentic Framework for Visual Detection and Linguistic Reasoning in Medical ImagingPaperframework
MEDVISTAGYM: A Scalable Training Environment for Thinking with Medical Images via Tool-Integrated Reinforcement LearningPaperbenchmark
MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential DiagnosisPaperbenchmark
Modeling Descriptive Norms in Multi-Agent Systems: An Auto-Aggregation PDE Framework with Adaptive Perception KernelsPaperbenchmark
Value of Information: A Framework for Human-Agent CommunicationPaperframework
DemMA: Dementia Multi-Turn Dialogue Agent with Expert-Guided Reasoning and Action SimulationPaperframework
Tool-MAD: A Multi-Agent Debate Framework for Fact Verification with Diverse Tool Augmentation and Adaptive RetrievalPaperbenchmark
Staged Voxel-Level Deep Reinforcement Learning for 3D Medical Image Segmentation with Noisy AnnotationsPaperbenchmark
RadDiff: Describing Differences in Radiology Image Sets with Natural LanguagePaperbenchmark
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and SegmentationPaperbenchmark
MedDialogRubrics: A Comprehensive Benchmark and Evaluation Framework for Multi-turn Medical Consultations in Large Language ModelsPaperbenchmark
Causal-Enhanced AI Agents for Medical Research ScreeningPapermedical
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate ForgettingPapermedical
Bayesian Orchestration of Multi-LLM Agents for Cost-Aware Sequential Decision-MakingPaperframework
An Explainable Agentic AI Framework for Uncertainty-Aware and Abstention-Enabled Acute Ischemic Stroke Imaging DecisionsPaperbenchmark

February

TitlePaper-linkSections
Evaluating Stochasticity in Deep Research AgentsPaperframework
Agency and Architectural Limits: Why Optimization-Based Systems Cannot Be Norm-ResponsivePapermedical
Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot StudyPaperbenchmark
Which Tool Response Should I Trust? Tool-Expertise-Aware Chest X-ray Agent with Multimodal Agentic LearningPaperframework
ALPACA: A Reinforcement Learning Environment for Medication Repurposing and Treatment Optimization in Alzheimer's DiseasePapermedical
LAMMI-Pathology: A Tool-Centric Bottom-Up LVLM-Agent Framework for Molecularly Informed Medical Intelligence in PathologyPaperframework
NutriOrion: A Hierarchical Multi-Agent Framework for Personalized Nutrition Intervention Grounded in Clinical GuidelinesPaperframework
4D-UNet improves clutter rejection in human transcranial contrast enhanced ultrasoundPaperbenchmark
INSURE-Dial: A Phase-Aware Conversational Dataset & Benchmark for Compliance Verification and Phase DetectionPaperbenchmark
3DMedAgent: Unified Perception-to-Understanding for 3D Medical AnalysisPaperbenchmark
Agentic Unlearning: When LLM Agent Meets Machine UnlearningPaperbenchmark
MedClarify: An information-seeking AI agent for medical diagnosis with case-specific follow-up questionsPapermedical
Agentic AI, Medical Morality, and the Transformation of the Patient-Physician RelationshipPaperframework
A Multi-Agent Framework for Medical AI: Leveraging Fine-Tuned GPT, LLaMA, and DeepSeek R1 for Evidence-Based and Bias-Aware Clinical Query ProcessingPaperbenchmark
MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool CallingPaperbenchmark
Implicit Bias in LLMs for Transgender PopulationsPaperapplication
TRACE: Temporal Reasoning via Agentic Context Evolution for Streaming Electronic Health Records (EHRs)Paperframework
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMsPaperbenchmark
Advancing AI Trustworthiness Through Patient Simulation: Risk Assessment of Conversational Agents for Antidepressant SelectionPaperframework
LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric EvaluationPaperbenchmark
Closing Reasoning Gaps in Clinical Agents with Differential Reasoning LearningPaperbenchmark
CoMMa: Contribution-Aware Medical Multi-Agents From A Game-Theoretic PerspectivePaperbenchmark
SynthAgent: A Multi-Agent LLM Framework for Realistic Patient Simulation -- A Case Study in Obesity with Mental Health ComorbiditiesPaperframework
MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive RegulationPaperbenchmark
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional TasksPaperbenchmark
Do LLMs Act Like Rational Agents? Measuring Belief Coherence in Probabilistic Decision MakingPaperapplication
Pruning Minimal Reasoning Graphs for Efficient Retrieval-Augmented GenerationPaperbenchmark
Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based AgentsPaperbenchmark
MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement LearningPaperbenchmark
Privasis: Synthesizing the Largest "Public" Private Dataset from ScratchPaperbenchmark
Perfusion Imaging and Single Material Reconstruction in Polychromatic Photon Counting CTPapermedical
RE-MCDF: Closed-Loop Multi-Expert LLM Reasoning for Knowledge-Grounded Clinical DiagnosisPaperbenchmark
MedBeads: An Agent-Native, Immutable Data Substrate for Trustworthy Medical AIPaperframework
ExperienceWeaver: Optimizing Small-sample Experience Learning for LLM-based Clinical Text ImprovementPaperbenchmark
Enhancing Imaging Depth and Sensitivity in Reflectance Mode Near Infrared Optical Imaging with Scatter Reducing AgentsPaperapplication

March

TitlePaper-linkSections
Symphony for Medical Coding: A Next-Generation Agentic System for Scalable and Explainable Medical CodingPaperbenchmark
Knowledge database development by large language models for countermeasures against viruses and marine toxinsPapermedical
Towards a Medical AI ScientistPaperframework
FeDMRA: Federated Incremental Learning with Dynamic Memory Replay AllocationPaperbenchmark
Improving Clinical Diagnosis with Counterfactual Multi-Agent ReasoningPaperbenchmark
MediHive: A Decentralized Agent Collective for Medical ReasoningPaperbenchmark
Autonomous Agent-Orchestrated Digital Twins (AADT): Leveraging the OpenClaw Framework for State Synchronization in Rare Genetic DisordersPaperframework
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AIPaperbenchmark
Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy VideosPaperbenchmark
OMIND: Framework for Knowledge Grounded Finetuning and Multi-Turn Dialogue Benchmark for Mental Health LLMsPaperbenchmark
Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationPaperbenchmark
MedOpenClaw: Auditable Medical Imaging Agents Reasoning over Uncurated Full StudiesPaperbenchmark
Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQAPaperbenchmark
RVLM: Recursive Vision-Language Models with Adaptive DepthPaperframework
CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in HealthcarePaperbenchmark
Dialogue to Question Generation for Evidence-based Medical Guideline Agent DevelopmentPaperbenchmark
3D-LLDM: Label-Guided 3D Latent Diffusion Model for Improving High-Resolution Synthetic MR Imaging in Hepatic Structure SegmentationPaperbenchmark
From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLMPaperframework
Training a Large Language Model for Medical Coding Using Privacy-Preserving Synthetic Clinical DataPapermedical
Privacy-Preserving EHR Data Transformation via Geometric Operators: A Human-AI Co-Design Technical ReportPaperframework
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical DatabasesPaperbenchmark
Cerebra: A Multidisciplinary AI Board for Multimodal Dementia Characterization and Risk AssessmentPaperbenchmark
Agentic Automation of BT-RADS Scoring: End-to-End Multi-Agent System for Standardized Brain Tumor Follow-up AssessmentPapermedical
ARYA: A Physics-Constrained Composable & Deterministic World Model ArchitecturePaperbenchmark
Anatomical Prior-Driven Framework for Autonomous Robotic Cardiac Ultrasound Standard View AcquisitionPaperbenchmark
TuLaBM: Tumor-Biased Latent Bridge Matching for Contrast-Enhanced MRI SynthesisPaperbenchmark
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective IntelligencePaperbenchmark
MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question AnsweringPaperbenchmark
Towards Equitable Robotic Furnishing Agents for Aging-in-Place: ADL-Grounded Design ExplorationPaperother
EviAgent: Evidence-Driven Agent for Radiology Report GenerationPaperbenchmark
OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance DatasetPaperbenchmark
Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in SpacePaperbenchmark
Six Interventions for the Responsible and Ethical Implementation of Medical AI AgentsPaperbenchmark
TheraAgent: Multi-Agent Framework with Self-Evolving Memory and Evidence-Calibrated Reasoning for PET TheranosticsPaperframework
Increasing intelligence in AI agents can worsen collective outcomesPapermedical
A Semi-Decentralized Approach to Multiagent ControlPaperbenchmark
When OpenClaw Meets Hospital: Toward an Agentic Operating System for Dynamic Clinical WorkflowsPaperbenchmark
UAV-MARL: Multi-Agent Reinforcement Learning for Time-Critical and Dynamic Medical Supply DeliveryPaperbenchmark
Human-AI Co-reasoning for Clinical Diagnosis with Evidence-Integrated Language AgentPaperbenchmark
MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent SystemsPaperbenchmark
Meissa: Multi-modal Medical Agentic IntelligencePaperbenchmark
RexDrug: Reliable Multi-Drug Combination Extraction through Reasoning-Enhanced LLMsPaperbenchmark
Empowering Locally Deployable Medical Agent via State Enhanced Logical Skills for FHIR-based Clinical TasksPaperbenchmark
Computational Pathology in the Era of Emerging Foundation and Agentic AI -- International Expert Perspectives on Clinical Integration and Translational ReadinessPaperbenchmark
Shifting Adaptation from Weight Space to Memory Space: A Memory-Augmented Agent for Medical Image SegmentationPaperbenchmark
Evolving Medical Imaging Agents via Experience-driven Self-skill DiscoveryPaperbenchmark
MedCoRAG: Interpretable Hepatology Diagnosis via Hybrid Evidence Retrieval and Multispecialty ConsensusPaperframework
Model Medicine: A Clinical Framework for Understanding, Diagnosing, and Treating AI ModelsPaperframework
Do Mixed-Vendor Multi-Agent LLMs Improve Clinical Diagnosis?Paperframework
A Multi-Agent Framework for Interpreting Multivariate Physiological Time SeriesPapermedical
MIND: Unified Inquiry and Diagnosis RL with Criteria Grounded Clinical Supports for Psychiatric ConsultationPaperframework
From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAGPaperbenchmark
ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical DialoguePaperbenchmark
CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic FrameworkPaperbenchmark
TARSE: Test-Time Adaptation via Retrieval of Skills and Experience for Reasoning AgentsPaperbenchmark
MedCollab: Causal-Driven Multi-Agent Collaboration for Full-Cycle Clinical Diagnosis via IBIS-Structured ArgumentationPaperbenchmark
DUCX: Decomposing Unfairness in Tool-Using Chest X-ray AgentsPaperframework
OPGAgent: An Agent for Auditable Dental Panoramic X-ray InterpretationPaperbenchmark

April

🚀 Year-2025

TitleGitHubSections
MedEyes: Learning Dynamic Visual Focus for Medical Progressive DiagnosisGitHubapplication
Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic RetrievalGitHubapplication
Reinventing Clinical Dialogue: Agentic Paradigms for LLM‑Enabled Healthcare CommunicationGitHubsurvey
Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal ReasoningGitHubevaluation
3mdbench: Medical multimodal multi-agent dialogue benchmarkGitHubevaluation
A co-evolving agentic AI system for medical imaging analysisGitHubother
A multimodal AI agent for clinical decision support in ophthalmologyNot Availableapplication
MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic WorkflowGitHubapplication
A dual-agent collaboration framework based on llms for nursing robots to perform bimanual coordination tasksNot Availablecapability
A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisationNot Availablesafety, evaluation
A hybrid reinforcement learning and knowledge graph framework for financial risk optimization in healthcare systemsNot Availableapplication
A Multi-Agent Approach to Neurological Clinical ReasoningNot Availablecapability, other
A Multimodal Multi-Agent Framework for Radiology Report GenerationNot Availabletask, evaluation, other
A Proposed LLM-Based Supported Treatment Framework for Intracerebral HemorrhageNot Availableintro, capability, application
A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language ModelsNot Availablecapability, application
A two-stage proactive dialogue generator for efficient clinical information collection using large language modelNot Availabletask, application
Actions speak louder than words: Agent decisions reveal implicit biases in language modelsNot Availablesafety
Adagent: Llm agent for alzheimer’s disease analysis with collaborative coordinatorNot Availablecapability, other
Agent-Based Uncertainty Awareness Improves Automated Radiology Report Labeling with an Open-Source Large Language ModelNot Availablecapability, other
Agentic AI for Clinical Decision Support: Real-Time Diagnosis, Triage, and Treatment PlanningNot Availableintro
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical KnowledgeGitHubcapability, task, application
Agentic Workflows in Healthcare: Advancing Clinical Efficiency through AI IntegrationNot Availableintro
Agentic-AI Healthcare: Multilingual, Privacy-First Framework with {MCP} AgentsGitHubcapability, application, safety
AGENTiGraph: A Multi-Agent Knowledge Graph Framework for Interactive, Domain-Specific LLM ChatbotsGitHubtask
Agentmd: Empowering language agents for risk prediction with large-scale clinical tool learningGitHubcapability, task, application
AI Agents in Clinical Medicine: A Systematic ReviewNot Availableevaluation
AI chatbots as professional service agents: developing a professional identityNot Availablecapability, application
AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question AnsweringGitHubcapability, task, other
AMQA: An Adversarial Dataset for Benchmarking Bias of LLMs in Medicine and HealthcareGitHubevaluation
An active inference strategy for prompting reliable responses from large language models in medical practiceNot Availablecapability
An Adaptive Multi-Agent LLM-Based Clinical Decision Support System Integrating Biomedical RAG and Web IntelligenceGitHubapplication
An Agentic Model Context Protocol Framework for Medical Concept StandardizationGitHubtask
An agentic system for rare disease diagnosis with traceable reasoningNot Availabletask, application, other
Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQAGitHubtask, other
ASTRID--An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering SystemsNot Availablesafety
At-cxr: Uncertainty-aware agentic triage for chest x-raysGitHubintro, other
Audited Reasoning Refinement: Fine-Tuning Language Models via LLM-Guided Step-Wise Evaluation and CorrectionNot Availableevaluation
AURA: A Multi-modal Medical Agent for Understanding, Reasoning and AnnotationGitHubother
Autonomous Multi-Modal LLM Agents for Treatment Planning in Focused Ultrasound Ablation SurgeryGitHubcapability, task, application
Autonomous Radiotherapy Treatment Planning Using DOLA: A Privacy-Preserving, LLM-Based Optimization AgentNot Availablesafety, other
Balancing Fairness and Performance in Healthcare {AI}: A Gradient Reconciliation ApproachNot Availablesafety
Benchmarking Automatic Speech Recognition coupled LLM Modules for Medical DiagnosticsNot Availableapplication
Beyond Benchmarks: Dynamic, Automatic and Systematic Red-Teaming Agents for Trustworthy Medical Language ModelsGitHubevaluation
Beyond Benchmarks: Evaluating Generalist Medical Artificial Intelligence With PsychometricsNot Availableevaluation
Bridging Clinical Narratives and ACR Appropriateness Guidelines: A Multi-Agent RAG System for Medical Imaging DecisionsGitHubtask
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False PresuppositionsGitHubother
CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement EstimationGitHubintro, capability, application, other
CARE-AD: a multi-agent large language model framework for Alzheimer’s disease prediction using longitudinal clinical notesNot Availablecapability, task, other
CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery PlanningNot Availablecapability, application, other
Chatbot To Help Patients Understand Their HealthGitHubapplication
ChatMyopia: An AI Agent for Pre-consultation Education in Primary Eye Care SettingsNot Availablecapability, application, other
Cod, towards an interpretable medical agent using chain of diagnosisGitHubcapability, application, safety, evaluation, other
Code Like Humans: A Multi-Agent Solution for Medical CodingGitHubapplication
Conversational health agents: a personalized large language model-powered agent frameworkGitHubsafety
CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question AnsweringNot Availableapplication
Data Overdose? Time for a Quadruple Shot: Knowledge Graph Construction Using Enhanced Triple ExtractionNot Availabletask
Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat AnalysisNot Availablesafety
Developing an Artificial Intelligence Tool for Personalized Breast Cancer Treatment Plans based on the NCCN GuidelinesNot Availableintro, capability, other
Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncologyNot Availableapplication, other
Differential privacy for medical deep learning: methods, tradeoffs, and deployment implicationsNot Availablesafety
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology ReasoningNot Availableevaluation
DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical servicesNot Availabletask, other
Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement LearningGitHubcapability, application
Doctoragent-rl: A multi-agent collaborative reinforcement learning system for multi-turn clinical dialogueGitHubtask
Dr. Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in RomanianNot Availablecapability, task
Drugagent: Multi-agent large language model-based reasoning for drug-target interaction predictionNot Availabletask
EH-Benchmark: Ophthalmic hallucination benchmark and agent-driven top-down traceable reasoning workflowGitHubother
Ehr-mcp: Real-world evaluation of clinical information retrieval by large language models via model context protocolNot Availableevaluation
Emerging cyber attack risks of medical ai agentsNot Availablesafety
Enhancing diagnostic capability with multi-agents conversational large language modelsGitHubcapability, task, application, other
Enhancing Medical Lung X-Ray Diagnosis Through Multi-Agent Vision-Language Model CollaborationNot Availablecapability, application
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in MedicineNot Availableevaluation
Evaluating the accuracy of a state-of-the-art large language model for prediction of admissions from the emergency roomNot Availableevaluation
Evaluating transparency in AI/ML model characteristics for FDA-reviewed medical devicesNot Availablesafety
Explainable AI for medical data: Current methods, limitations, and future directionsNot Availablesafety
Eyecaregpt: Boosting comprehensive ophthalmology understanding with tailored dataset, benchmark and modelGitHubapplication, other
Feat: A multi-agent forensic ai system with domain-adapted large language model for automated cause-of-death analysisGitHubcapability
Fine-tuning vision language models with graph-based knowledge for explainable medical image analysisNot Availablecapability
FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research InsightsGitHubother
GEMA-Score: Granular Explainable Multi-Agent Score for Radiology Report EvaluationGitHubother
Geometry-preserving encoder/decoder in latent generative modelsGitHubevaluation
GMAT: Grounded Multi-agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image ClassificationNot Availabletask
Haibu Mathematical-Medical Intelligent Agent: Enhancing Large Language Model Reliability in Medical Tasks via Verifiable Reasoning ChainsNot Availablecapability
Demo: Healthcare Agent Orchestrator (HAO) for Patient Summarization in Molecular Tumor BoardsNot Availablesections: task, other
Healthcare Agent: Eliciting the Power of Large Language Models for Medical ConsultationNot Availablecapability, application
Image Segmentation Using Only" Better or Worse" Expert FeedbackNot Availabletask
Improving Interactive Diagnostic Ability of a Large Language Model Agent Through Clinical Experience LearningGitHubcapability, application
In-Basket Message Volume in Primary Care: A Cross-sectional Analysis by Gender and SpecialtyNot Availableintro
KERAP: A Knowledge-Enhanced Reasoning Approach for Accurate Zero-shot Diagnosis Prediction Using Multi-agent LLMsGitHubcapability
Large language models in real-world clinical workflows: a systematic review of applications and implementationNot Availableevaluation
Learning to be a doctor: Searching for effective medical agent architecturesNot Availablecapability, application, other
Lessons Learned from Evaluation of LLM based Multi-agents in Safer Therapy RecommendationGitHubtask, other
LINS: A general medical Q&A framework for enhancing the quality and credibility of LLM-generated responsesGitHubcapability
Llms can simulate standardized patients via agent coevolutionGitHubtask, application
M3Builder: A Multi-Agent System for Automated Machine Learning in Medical ImagingGitHubtask
Magnetic Milli-Spinner for Robotic Endovascular SurgeryNot Availabletask
MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized CollaborationGitHubapplication, other
Mdteamgpt: A self-evolving llm-based multi-agent framework for multi-disciplinary team medical consultationGitHubtask, application
Measurement to Meaning: A Validity-Centered Framework for AI EvaluationNot Availableevaluation
Med-TAMARA: Trust-Aware Multi-Agent Risk Assessment in Medical AI DialogueNot Availablecapability, application
Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced AgentsGitHubcapability, application
MedAgent-Pro: Towards Evidence-Based Multi-Modal Medical Diagnosis via Reasoning Agentic WorkflowNot Availablesafety, evaluation, other
MedAgentAudit: Diagnosing and Quantifying Collaborative Failure Modes in Medical Multi-Agent SystemsGitHubcapability, safety, evaluation
MedAgentBench: Dataset for Benchmarking LLMs as AgentsGitHubevaluation
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical TasksGitHubevaluation
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical ReasoningGitHubevaluation
MedAgentSim: Self-evolving Multi-agent Simulations for Realistic Clinical InteractionsGitHubtask, other
MedBrowseComp: Benchmarking Medical Deep Research and Computer UseGitHubevaluation
Medchat: A multi-agent framework for multimodal diagnosis with large language modelsGitHubother
MedCoAct: Confidence-Aware Multi-Agent Collaboration for Complete Clinical DecisionNot Availablecapability, application
Meddxagent: A unified modular agent framework for explainable automatic differential diagnosisGitHubother
MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-Checking of LLM ResponsesGitHubevaluation
Medhallu: A comprehensive benchmark for detecting medical hallucinations in large language modelsGitHubsafety
Mediator-guided multi-agent collaboration among open-source models for medical decision-makingNot Availabletask, other
Medical AI Consensus: A Multi-Agent Framework for Radiology Report Generation and EvaluationNot Availablecapability, task
Medical hallucinations in foundation models and their impact on healthcareGitHubsafety
Position: Medical large language model benchmarks should prioritize construct validityNot Availableevaluation
MedicalOS: An {LLM} Agent based Operating System for Digital HealthcareNot Availablecapability, application
MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge GraphGitHubtask
MedKGEval: A Knowledge Graph-Based Multi-Turn Evaluation Framework for Open-Ended Patient Interactions with Clinical LLMsNot Availabletask
MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language ModelsGitHubapplication
Medmmv: A controllable multimodal multi-agent framework for reliable and verifiable clinical reasoningNot Availabletask, safety, evaluation, other
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible ExtensibilityNot Availabletask, other
MedPAO: A Protocol-Driven Agent for Structuring Medical ReportsGitHubcapability, task, application, other
Medrax: Medical reasoning agent for chest x-rayGitHubcapability, application, other
MedRepBench: A Comprehensive Benchmark for Medical Report InterpretationNot Availablecapability, evaluation
Medresearcher-r1: Expert-level medical deep researcher via a knowledge-informed trajectory synthesis frameworkGitHubevaluation
MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question AnsweringNot Availablecapability, task
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingGitHubevaluation
MEGA-RAG: a retrieval-augmented generation framework with multi-evidence guided answer refinement for mitigating hallucinations of LLMs in public healthNot Availablesafety
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical ReasoningNot Availabletask, other
MoMA: A Mixture-of-Multimodal-Agents Architecture for Enhancing Clinical Prediction ModellingNot Availablecapability, task, other
MRGAgents: A Multi-Agent Framework for Improved Medical Report Generation with Med-LVLMsNot Availabletask, other
Multi agent based medical assistant for edge devicesGitHubsafety
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationNot Availableevaluation
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in {LMIC}sNot Availablecapability, application
Multimodal Models in Healthcare: Methods, Challenges, and Future Directions for Enhanced Clinical Decision SupportNot Availableintro
NurseLLM: The First Specialized Language Model for NursingNot Availablecapability
OAAgent: Multimodal LLM Agent for Predicting Knee Osteoarthritis ProgressionNot Availablecapability
OpenLens AI: Fully Autonomous Research Agent for Health InfomaticsGitHubcapability, application
PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray ReasoningGitHubother
Pathfinder: A multi-modal multi-agent system for medical diagnostic decision-making applied to histopathologyGitHubtask
Patient-Zero: A Unified Framework for Real-Record-Free Patient Agent GenerationNot Availablecapability, application, safety
Performance of Retrieval-Augmented Generation Large Language Models in Guideline-Concordant Prostate-Specific Antigen Testing: Comparative Study With Junior CliniciansNot Availableevaluation
Privacy in action: Towards realistic privacy mitigation and evaluation for llm-powered agentsGitHubsafety
Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language ModelsGitHubcapability
Program Synthesis Dialog Agents for Interactive Decision-MakingNot Availablesafety
Proof-of-TBI--Fine-Tuned Vision Language Model Consortium and OpenAI-o3 Reasoning LLM-Based Medical Diagnosis Support System for Mild Traumatic Brain Injury (TBI) PredictionNot Availablecapability, application
Rapidly benchmarking large language models for diagnosing comorbid patients: comparative study leveraging the LLM-as-a-judge methodNot Availableevaluation
Real-World Evaluation of Large Language Models in Healthcare (RWE-LLM): A New Realm of AI Safety & ValidationNot Availableevaluation
Red-teaming llm multi-agent systems via communication attacksNot Availablesafety
Reducing Hallucinations and Trade-Offs in Responses in Generative AI Chatbots for Cancer Information: Development and Evaluation StudyNot Availablesafety
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical AgentsGitHubcapability, application
Resilient Multi-Agent Negotiation for Medical Supply Chains: Integrating LLMs and Blockchain for Transparent CoordinationNot Availabletask
RxLens: Multi-Agent LLM-powered Scan and Order for PharmacyNot Availablecapability, application, other
SCOPE: Speech-Guided COllaborative PErception Framework for Surgical Scene SegmentationNot Availabletask, other
Self-Assessment of Content, Pedagogy, and Technology Knowledge among Higher Education Academics in BahrainNot Availablecapability
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision MakingNot Availablecapability
SmartState: An Automated Research Protocol Adherence SystemGitHubcapability, application
SOLVE-Med: Specialized Orchestration for Leading Vertical Experts across Medical SpecialtiesGitHubtask
Standard Applicability Judgment and Cross-jurisdictional Reasoning: A RAG-based Framework for Medical Device ComplianceNot Availablecapability
Surgraw: Multi-agent workflow with chain-of-thought reasoning for surgical intelligenceGitHubtask
Survey and improvement strategies for gene prioritization with large language modelsNot Availabletask, other
Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QAGitHubcapability
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM SystemsNot Availablesafety
The evaluation illusion of large language models in medicineNot Availableevaluation
Tiered Agentic Oversight: A Hierarchical Multi-Agent System for AI Safety in HealthcareGitHubcapability, safety
Tool learning with large language models: A surveyGitHubintro
Towards conversational diagnostic artificial intelligenceNot Availabletask
Towards interpretable radiology report generation via concept bottlenecks using a multi-agentic ragGitHubtask
Towards safe ai clinicians: A comprehensive study on large language model jailbreaking in healthcareNot Availablesafety
Transforming healthcare delivery with conversational AI platformsNot Availablesafety
Tree-based RAG-Agent Recommendation System: A Case Study in Medical Test DataNot Availablecapability, task, application
Tree-of-Reasoning: Towards Complex Medical Diagnosis via Multi-Agent Reasoning with Evidence TreeGitHubsafety
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought ProcessesNot Availableevaluation
TxAgent: An AI agent for therapeutic reasoning across a universe of toolsGitHubintro, task, other
Using large language models for enhanced fraud analysis and detection in blockchain based health insurance claimsGitHubcapability, application
Vision-language model for report generation and outcome prediction in CT pulmonary angiogramGitHubcapability, application
Visual-Conversational Interface for Evidence-Based Explanation of Diabetes Risk PredictionGitHubintro
When Avatars Have Personality: Effects on Engagement and Communication in Immersive Medical TrainingNot Availableapplication
World Model for AI Autonomous Navigation in Mechanical ThrombectomyGitHubcapability, task, application
Zero-Shot Large Language Model Agents for Fully Automated Radiotherapy Treatment PlanningNot Availablecapability, application, other
Bias-Aware Agent: Enhancing Fairness in AI-Driven Knowledge RetrievalGitHubsafety
EMR-AGENT: Automating Cohort and Feature Extraction from EMR DatabasesGitHubcapability, task, application
Colacare: Enhancing electronic health record modeling through large language model-driven multi-agent collaborationGitHubcapability, application
A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-makingGitHubcapability, application
MeNTi: Bridging medical calculator and LLM agent with nested tool callingGitHubcapability, application
Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAGGitHubapplication, other
How Can We Diagnose and Treat Bias in Large Language Models for Clinical Decision-Making?Not Availableevaluation

🚀 Year-2024

TitleGitHubSections
A demonstration of adaptive collaboration of large language models for medical decision-makingGitHubcapability, application
A survey on large language model based autonomous agentsNot Availableintro
Achieving health equity through conversational AI: A roadmap for design and implementation of inclusive chatbots in healthcareNot Availablesafety
Adaptive Reasoning and Acting in Medical Language AgentsNot Availablecapability, application
Adversarial attacks on large language models in medicineNot Availablesafety
Agent Hospital: A Simulacrum of Hospital with Evolvable Medical AgentsGitHubcapability, task, application
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environmentsGitHubintro, capability, application, evaluation, other
Agentic llm workflows for generating patient-friendly medical reportsGitHubcapability, task, application, other
Agentigraph: An interactive knowledge graph platform for llm-based chatbots utilizing private dataGitHubcapability, application
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction SimulatorGitHubcapability, task, application, evaluation
Aligning Medical LLMs for Counterfactual FairnessGitHubsafety
ArgMed-Agents: explainable clinical decision reasoning with LLM disscusion via argumentation schemesNot Availablecapability, application
Autohealth: Advanced llm-empowered wearable personalized medical butler for parkinson’s disease managementNot Availableother
Autonomous artificial intelligence agents for clinical decision making in oncologyNot Availablecapability, other
Benchmarking Large Language Models on Communicative Medical Coaching: A Dataset and a Novel SystemGitHubcapability, application
Beyond direct diagnosis: LLM-based multi-specialist agent consultation for automatic diagnosisNot Availablecapability
Chatdev: Communicative agents for software developmentGitHubtask
ClinicalAgent: Clinical Trial Multi-Agent System with Large Language Model-based ReasoningGitHubintro, capability, application, other
ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real WorldGitHubcapability, evaluation
Cxr-agent: Vision-language models for chest x-ray interpretation with uncertainty aware radiology reportingNot Availablecapability, application, other
Development of a Large Language Model-based Multi-Agent Clinical Decision Support System for Korean Triage and Acuity Scale (KTAS)-Based Triage and Treatment Planning in Emergency DepartmentsGitHubcapability, task, other
Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health recordsGitHubcapability, task, application
Enhancing diagnostic accuracy through multi-agent conversations: using large language models to mitigate cognitive biasNot Availableapplication
Enhancing llms for impression generation in radiology reports through a multi-agent systemGitHubcapability, task, application, other
Ethical and regulatory challenges of large language models in medicineNot Availablesafety
Evaluating Large Language Models as Agents in the ClinicNot Availablecapability, evaluation
Exploring llm multi-agents for icd codingNot Availableapplication
Exploring LLM-based Data Annotation Strategies for Medical Dialogue Preference AlignmentNot Availablecapability
Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative AINot Availablesafety
GuidelineGuard: An Agentic Framework for Medical Note Evaluation with Guideline AdherenceNot Availablecapability, application
Imas: A comprehensive agentic approach to rural healthcare deliveryGitHubcapability, application
Improving Clinical Documentation with AI: A Comparative Study of Sporo AI Scribe and GPT-4o miniNot Availablecapability, task, application
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical ReasoningGitHubcapability, application
Integration of multi-source medical data for medical diagnosis question answeringGitHubcapability, application
Iryonlp at mediqa-corr 2024: Tackling the medical error detection & correction task on the shoulders of medical agentsNot Availablecapability, task, application
KG4Diagnosis: A Hierarchical Multi-Agent LLM Framework with Knowledge Graph Enhancement for Medical DiagnosisNot Availablecapability, application
Knowledge-infused llm-powered conversational health agent: A case study for diabetes patientsNot Availablecapability
Large Language Model-Enhanced Interactive Agent for Public Education on Newborn Auricular DeformitiesNot Availablecapability, application
Llm-based framework for administrative task automation in healthcareNot Availablecapability, application
Llm-medqa: Enhancing medical question answering through case studies in large language modelsNot Availablecapability, application
MAGDA: Multi-agent guideline-driven diagnostic assistanceNot Availablecapability, task, application
MALADE: orchestration of LLM-powered agents with retrieval augmented generation for pharmacovigilanceGitHubintro, capability, application, other
Mdagents: An adaptive collaboration of llms for medical decision-makingGitHubcapability, task, application
MedAgents: Large Language Models as Collaborators for Zero-shot Medical ReasoningGitHubcapability, task, other
MedAide: Towards an Omni Medical Aide via Specialized {LLM}-based Multi-Agent CollaborationNot Availablecapability
Medco: Medical education copilots based on a multi-agent frameworkNot Availablecapability, task, application
MedChain: Bridging the Gap Between LLM Agents and Clinical Practice through Interactive Sequential BenchmarkingGitHubcapability, application, evaluation
MedGen: An Explainable Multi-Agent Architecture for Clinical Decision Support through Multisource Knowledge FusionNot Availablecapability, application
Medhalu: Hallucinations in responses to healthcare queries by large language modelsNot Availablesafety
MedQA-CS: Benchmarking Large Language Models’ Clinical Skills Using an AI-SCE FrameworkGitHubevaluation
Mitigating cognitive biases in clinical decision-making through multi-agent conversations using large language models: simulation studyNot Availablecapability, application
Mitigating hallucinations in large language models via self-refinement-enhanced knowledge retrievalNot Availablesafety
Mmedagent: Learning to use medical tools with multi-modal agentGitHubcapability, application, other
MMLU-Pro: A More Robust Benchmark for Multi-Task Language UnderstandingGitHubevaluation
Natural Language Programming in Medicine: Administering Evidence Based Clinical Workflows with Autonomous Agents Powered by Generative Large Language ModelsNot Availablecapability, application
On protecting the data privacy of large language models (llms): A surveyNot Availablesafety
On the resilience of llm-based multi-agent collaboration with faulty agentsNot Availablesafety
Piors: Personalized intelligent outpatient reception based on large language model with multi-agents medical scenario simulationGitHubcapability, task, application
Polaris: A safety-focused llm constellation architecture for healthcareNot Availablecapability, application, evaluation
Privacy-Preserving Large Language Models: MechanismsNot Availablesafety
Advancing healthcare automation: Multi-agent system for medical necessity justificationNot Availablecapability, task, application
RareAgents: Advancing Rare Disease Care through LLM-Empowered Multi-disciplinary TeamNot Availablecapability, application
RareAgents: Autonomous Multi-disciplinary Team for Rare Disease Diagnosis and TreatmentNot Availabletask, other
Regulator-manufacturer AI agents modeling: Mathematical feedback-driven multi-agent LLM frameworkNot availablecapability, task, application
Remoni: An autonomous system integrating wearables and multimodal large language models for enhanced remote health monitoringNot availablecapability, application
Rx strategist: Prescription verification using llm agents systemNot availablecapability, task, application, other
Simulated patient systems are intelligent when powered by large language model-based AI agentsNot availablecapability, application
Smile: Single-turn to multi-turn inclusive language expansion via chatgpt for mental health supportGithubtask
Society of medical simplifiersGitHubcapability, task, application
Surgbox: Agent-driven operating room sandbox with surgery copilotNot availablecapability, task, application
T-agent: A term-aware agent for medical dialogue generationNot availablecapability, application
The role of explainability in AI-supported medical decision-makingNot availablesafety
Towards anatomy education with generative AI-based virtual assistants in immersive virtual reality environmentsNot availableapplication
Towards Automatic Evaluation for {LLM}s' Clinical Capabilities: Metric, Data, and AlgorithmNot availablecapability, application, evaluation
Towards next-generation medical agent: How o1 is reshaping decision-making in medical scenariosNot availablecapability, other
A trustworthy AI reality-check: the lack of transparency of artificial intelligence products in healthcareNot availablesafety
TWIN-GPT: digital twins for clinical trials via large language modelNot availableother
UMass-BioNLP at MEDIQA-M3G 2024: DermPrompt--A Systematic Exploration of Prompt Engineering with GPT-4V for Dermatological DiagnosisGitHubcapability, other
Zodiac: A cardiologist-level llm framework for multi-agent diagnosticsNot availablecapability
OpenAI o1 System CardGitHubcapability
FedAgentBench: Towards Automating Real-World Federated Medical Image Analysis with Server–Client LLM AgentsNot availableevaluation
Evaluating large language models as agents in the clinicNot available

🚀 Year-2023

TitleGitHubSections
A reinforcement learning approach for VQA validation: An application to diabetic macular edema gradingNot availablecapability
Adaptive multi-agent deep reinforcement learning for timely healthcare interventionsNot availablecapability, application
Asynchronous decentralized federated lifelong learning for landmark localization in medical imagingNot availablecapability
Beyond memorization: Violating privacy via inference with large language modelsNot availablesafety
Camel: Communicative agents for" mind" exploration of large language model societyNot availabletask
Clinically-inspired multi-agent transformers for disease trajectory forecasting from multimodal dataGitHubcapability, application
Cognitive architectures for language agentsNot availableintro
Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media ForumNot availableevaluation
Deep Imitation Learning for Automated Drop-In Gamma Probe ManipulationNot availabletask
Diaggpt: An llm-based and multi-agent dialogue system with automatic topic management for flexible task-oriented dialogueGitHubapplication
Dspy: Compiling declarative language model calls into self-improving pipelinesGitHubtask
Federated machine learning, privacy-enhancing technologies, and data protection laws in medical research: scoping reviewNot availablesafety
Generative agents: Interactive simulacra of human behaviorGitHubintro
Interactive medical image segmentation with self-adaptive confidence calibrationGitHubtask
Large language models as agents in the clinicNot availableapplication
MetaGPT: Meta programming for a multi-agent collaborative frameworkGitHubtask
Navigation Through Endoluminal Channels Using Q-LearningNot availablecapability, task
PROFSA: SELF-SUPERVISED POCKET PRETRAINING VIA PROTEIN FRAGMENT-SURROUNDINGS ALIGNGitHubcapability
Reflexion: Language agents with verbal reinforcement learningGitHubcapability
Td-mpc2: Scalable, robust world models for continuous controlGitHubtask
Temporally-extended prompts optimization for sam in interactive medical image segmentationNot availablecapability, task
The NCI Imaging Data Commons as a platform for reproducible research in computational pathologyNot availableintro
Towards Causality-Aware Inferring: A Sequential Discriminative Approach for Medical DiagnosisNot availablecapability, application

🚀 Earlier

TitleGitHubSections
"My Nose is Running." "Are You Also Coughing?": Building a Medical Diagnosis Agent with Interpretable Inquiry LogicsGitHubcapability, application
A Flexible Schema-Guided Dialogue Management Framework: From Friendly Peer to Virtual Standardized Cancer PatientGitHubcapability, application
Building an {ASR} Error Robust Spoken Virtual Patient System in a Highly Class-Imbalanced Scenario Without Speech DataNot availablecapability, application
Constitutional ai: Harmlessness from ai feedbackGitHubcapability
MedDG: An Entity-Centric Medical Consultation Dataset for Entity-Aware Medical Dialogue GenerationGitHubevaluation
MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoningNot availableintro
Multi-agent searching system for medical informationNot availablecapability
React: Synergizing reasoning and acting in language modelsGitHubintro
Scalable Online Disease Diagnosis via Multi-Model-Fused Actor-Critic Reinforcement LearningNot availablecapability, application
MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical Domain Question AnsweringGitHubcapability, evaluation
A grounded well-being conversational agent with multiple interaction modes: Preliminary resultsNot availablecapability, application
Adaptable image quality assessment using meta-reinforcement learning of task amenabilityGitHubcapability
An Edge Based Multi-Agent Model for Improving Hospital Bed ManagementNot availabletask
Cross Modality 3D Navigation Using Reinforcement Learning and Neural Style TransferNot availablecapability
Extracting Training Data from Large Language ModelsGitHubsafety
Human-AI collaboration in healthcare: A review and research agendaNot availabletask
Levels of autonomy and safety assurance for AI-Based clinical decision systemsNot availablesafety
Measuring Massive Multitask Language UnderstandingGitHubevaluation
Autonomous systems and artificial intelligence in healthcare transformation to 5P medicine--ethical challengesNot availableintro
Boundary-aware supervoxel-level iteratively refined interactive 3d image segmentation with multi-agent reinforcement learningNot availabletask
MedDialog: A Large-scale Medical Dialogue DatasetNot availableevaluation
Medical visual question answering via conditional reasoningGitHubtask
PathVQA: 30000+ Questions for Medical Visual Question AnsweringGitHubevaluation
PubMedQA: A Dataset for Biomedical Research Question AnsweringGitHubevaluation
A Dataset of Clinically Generated Visual Questions and Answers About Radiology ImagesNot availableevaluation
Modeling irregularly sampled clinical time seriesGitHubintro
The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive careGitHubtask
Medical robotics—Regulatory, ethical, and legal considerations for increasing levels of autonomyNot availabletask
Tethered to the EHR: primary care physician workload assessment using EHR event log data and time-motion observationsNot availableintro
Allocation of physician time in ambulatory practice: a time and motion study in 4 specialtiesNot availableintro
Assessing electronic note quality using the physician documentation quality instrument (PDQI-9)Not availabletask
Privacy by design: The 7 foundational principlesNot availablesafety
Upper processing stages of the perception--action cycleNot availableintro
What Disease Does This Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical ExamsGitHubevaluation

✨ Papers by Category

🚀 1. Capability

1.1 Planning

TitleGitHubYear
Evaluating Large Language Models as Agents in the ClinicNot Available2024
MedPAO: A Protocol-Driven Agent for Structuring Medical ReportsGitHub2025
Adaptable image quality assessment using meta-reinforcement learning of task amenabilityGitHub2023
World Model for AI Autonomous Navigation in Mechanical ThrombectomyGitHub2025
Polaris: A safety-focused llm constellation architecture for healthcareNot Available2024
MedChain: Bridging the Gap Between LLM Agents and Clinical Practice through Interactive Sequential BenchmarkingGitHub2024
Rx strategist: Prescription verification using llm agents systemNot Available2024
A Flexible Schema-Guided Dialogue Management Framework: From Friendly Peer to Virtual Standardized Cancer PatientGitHub2023
Surgbox: Agent-driven operating room sandbox with surgery copilotNot Available2024
MedicalOS: An {LLM} Agent based Operating System for Digital HealthcareNot Available2025
"My Nose is Running." "Are You Also Coughing?": Building a Medical Diagnosis Agent with Interpretable Inquiry LogicsGitHub2023
Cross Modality 3D Navigation Using Reinforcement Learning and Neural Style TransferNot Available2023
Scalable Online Disease Diagnosis via Multi-Model-Fused Actor-Critic Reinforcement LearningNot Available2023
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction SimulatorGitHub2024
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical ReasoningGitHub2024
Llm-medqa: Enhancing medical question answering through case studies in large language modelsNot Available2024
Medco: Medical education copilots based on a multi-agent frameworkNot Available2024
Mitigating cognitive biases in clinical decision-making through multi-agent conversations using large language models: simulation studyNot Available2024
Regulator-manufacturer AI agents modeling: Mathematical feedback-driven multi-agent LLM frameworkNot Available2024
Society of medical simplifiersGitHub2024
Towards Automatic Evaluation for {LLM}s' Clinical Capabilities: Metric, Data, and AlgorithmNot Available2024
Towards next-generation medical agent: How o1 is reshaping decision-making in medical scenariosNot Available2024
A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-makingGitHub2025
A dual-agent collaboration framework based on llms for nursing robots to perform bimanual coordination tasksNot Available2025
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical KnowledgeGitHub2025
Agentic-AI Healthcare: Multilingual, Privacy-First Framework with {MCP} AgentsGitHub2025
Autonomous Multi-Modal LLM Agents for Treatment Planning in Focused Ultrasound Ablation SurgeryGitHub2025
Colacare: Enhancing electronic health record modeling through large language model-driven multi-agent collaborationGitHub2025
Haibu Mathematical-Medical Intelligent Agent: Enhancing Large Language Model Reliability in Medical Tasks via Verifiable Reasoning ChainsNot Available2025
Improving Interactive Diagnostic Ability of a Large Language Model Agent Through Clinical Experience LearningGitHub2025
Learning to be a doctor: Searching for effective medical agent architecturesNot Available2025
MeNTi: Bridging medical calculator and LLM agent with nested tool callingGitHub2025
OpenLens AI: Fully Autonomous Research Agent for Health InfomaticsGitHub2025
Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QAGitHub2025
Tree-based RAG-Agent Recommendation System: A Case Study in Medical Test DataNot Available2025
AI chatbots as professional service agents: developing a professional identityNot Available2025
A reinforcement learning approach for VQA validation: An application to diabetic macular edema gradingNot Available2023
Building an {ASR} Error Robust Spoken Virtual Patient System in a Highly Class-Imbalanced Scenario Without Speech DataNot Available2023
MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical Domain Question AnsweringGitHub2023
Multi-agent searching system for medical informationNot Available2023
A demonstration of adaptive collaboration of large language models for medical decision-makingGitHub2024
Agentic llm workflows for generating patient-friendly medical reportsGitHub2024
Agentigraph: An interactive knowledge graph platform for llm-based chatbots utilizing private dataGitHub2024
Improving Clinical Documentation with AI: A Comparative Study of Sporo AI Scribe and GPT-4o miniNot Available2024
Piors: Personalized intelligent outpatient reception based on large language model with multi-agents medical scenario simulationGitHub2024
SmartState: An Automated Research Protocol Adherence SystemGitHub2025

1.2 Tool Use

TitleGitHubYear
ClinicalAgent: Clinical Trial Multi-Agent System with Large Language Model-based ReasoningGitHub2024
Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement LearningGitHub2025
KERAP: A Knowledge-Enhanced Reasoning Approach for Accurate Zero-shot Diagnosis Prediction Using Multi-agent LLMsGitHub2025
Iryonlp at mediqa-corr 2024: Tackling the medical error detection & correction task on the shoulders of medical agentsNot Available2024
Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language ModelsGitHub2025
Cxr-agent: Vision-language models for chest x-ray interpretation with uncertainty aware radiology reportingNot Available2024
Enhancing llms for impression generation in radiology reports through a multi-agent systemGitHub2024
GuidelineGuard: An Agentic Framework for Medical Note Evaluation with Guideline AdherenceNot Available2024
Medrax: Medical reasoning agent for chest x-rayGitHub2025
Mmedagent: Learning to use medical tools with multi-modal agentGitHub2024
Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced AgentsGitHub2025
MedAgents: Large Language Models as Collaborators for Zero-shot Medical ReasoningGitHub2024
AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question AnsweringGitHub2025
NurseLLM: The First Specialized Language Model for NursingNot Available2025
A multimodal AI agent for clinical decision support in ophthalmologyNot Available2025
Autonomous artificial intelligence agents for clinical decision making in oncologyNot Available2024
ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real WorldGitHub2024
Llm-based framework for administrative task automation in healthcareNot Available2024
A Multi-Agent Approach to Neurological Clinical ReasoningNot Available2025
Adagent: Llm agent for alzheimer’s disease analysis with collaborative coordinatorNot Available2025
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical AgentsGitHub2025
Development of a Large Language Model-based Multi-Agent Clinical Decision Support System for Korean Triage and Acuity Scale (KTAS)-Based Triage and Treatment Planning in Emergency DepartmentsGitHub2024
Large Language Model-Enhanced Interactive Agent for Public Education on Newborn Auricular DeformitiesNot Available2024
Natural Language Programming in Medicine: Administering Evidence Based Clinical Workflows with Autonomous Agents Powered by Generative Large Language ModelsNot Available2024
Agentmd: Empowering language agents for risk prediction with large-scale clinical tool learningGitHub2025
MoMA: A Mixture-of-Multimodal-Agents Architecture for Enhancing Clinical Prediction ModellingNot Available2025
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in {LMIC}sNot Available2025

1.3 Memory

1.4 Self-Improvement

1.5 Reasoning

TitleGitHubYear
Proof-of-TBI--Fine-Tuned Vision Language Model Consortium and OpenAI-o3 Reasoning LLM-Based Medical Diagnosis Support System for Mild Traumatic Brain Injury (TBI) PredictionNot Available2025
OpenAI o1 System CardGitHub2024
Beyond direct diagnosis: LLM-based multi-specialist agent consultation for automatic diagnosisNot Available2024
MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic WorkflowGitHub2025
MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question AnsweringNot Available2025
KG4Diagnosis: A Hierarchical Multi-Agent LLM Framework with Knowledge Graph Enhancement for Medical DiagnosisNot Available2024
MAGDA: Multi-agent guideline-driven diagnostic assistanceNot Available2024
Zero-Shot Large Language Model Agents for Fully Automated Radiotherapy Treatment PlanningNot Available2025
Cod, towards an interpretable medical agent using chain of diagnosisGitHub2025
Asynchronous decentralized federated lifelong learning for landmark localization in medical imagingNot Available2023
Temporally-extended prompts optimization for sam in interactive medical image segmentationNot Available2023
Integration of multi-source medical data for medical diagnosis question answeringGitHub2024
MedGen: An Explainable Multi-Agent Architecture for Clinical Decision Support through Multisource Knowledge FusionNot Available2024
Zodiac: A cardiologist-level llm framework for multi-agent diagnosticsNot Available2024
An active inference strategy for prompting reliable responses from large language models in medical practiceNot Available2025
CARE-AD: a multi-agent large language model framework for Alzheimer’s disease prediction using longitudinal clinical notesNot Available2025
Feat: A multi-agent forensic ai system with domain-adapted large language model for automated cause-of-death analysisGitHub2025
Fine-tuning vision language models with graph-based knowledge for explainable medical image analysisNot Available2025
Advancing healthcare automation: Multi-agent system for medical necessity justificationNot Available2024
ArgMed-Agents: explainable clinical decision reasoning with LLM disscusion via argumentation schemesNot Available2024
Imas: A comprehensive agentic approach to rural healthcare deliveryGitHub2024
Simulated patient systems are intelligent when powered by large language model-based AI agentsNot Available2024
UMass-BioNLP at MEDIQA-M3G 2024: DermPrompt--A Systematic Exploration of Prompt Engineering with GPT-4V for Dermatological DiagnosisGitHub2024
CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery PlanningNot Available2025
Dr. Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in RomanianNot Available2025
Enhancing Medical Lung X-Ray Diagnosis Through Multi-Agent Vision-Language Model CollaborationNot Available2025
Enhancing diagnostic capability with multi-agents conversational large language modelsGitHub2025
LINS: A general medical Q&A framework for enhancing the quality and credibility of LLM-generated responsesGitHub2025
OAAgent: Multimodal LLM Agent for Predicting Knee Osteoarthritis ProgressionNot Available2025
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision MakingNot Available2025
Standard Applicability Judgment and Cross-jurisdictional Reasoning: A RAG-based Framework for Medical Device ComplianceNot Available2025

1.6 Perception

1.7 Others (continual learning, uncertainty)

🚀 2. Atomic Function

2.1 Basic Technology Empowerment

2.2 Core Diagnostic & Therapeutic Assistance

2.3 Workflow & Documentation Optimization

🚀 3. Application

3.1 Intake & Clinical Dialogue

3.2 Virtual MDT Teams & Multimodal Reasoning

3.3 Treatment Procedures

3.4 Chronic Disease Management & Prescription Safety

3.5 Documentation, Coding & Knowledge Infrastructure

3.6 Simulation & Support Systems

3.7 Regulation, Payer Workflows & Administrative Automation

🚀 4. Safety

4.1 Medical Hallucination

4.2 Privacy & Data-Security

4.3 Explainability & Transparency

4.4 Adversarial Security & Threat Modeling

4.5 AI Governance & Systemic Safety

4.6 Bias, Fairness & Accessibility

🚀 5. Evaluation

5.1 Benchmarks

TitleGitHubYear
Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal ReasoningGitHub2025
A Dataset of Clinically Generated Visual Questions and Answers About Radiology ImagesNot Available2023
Measuring Massive Multitask Language UnderstandingGitHub2023
MedDG: An Entity-Centric Medical Consultation Dataset for Entity-Aware Medical Dialogue GenerationGitHub2023
MedDialog: A Large-scale Medical Dialogue DatasetNot Available2023
PathVQA: 30000+ Questions for Medical Visual Question AnsweringGitHub2023
PubMedQA: A Dataset for Biomedical Research Question AnsweringGitHub2023
FedAgentBench: Towards Automating Real-World Federated Medical Image Analysis with Server–Client LLM AgentsNot Available2024
MMLU-Pro: A More Robust Benchmark for Multi-Task Language UnderstandingGitHub2024
MedQA-CS: Benchmarking Large Language Models’ Clinical Skills Using an AI-SCE FrameworkGitHub2024
3mdbench: Medical multimodal multi-agent dialogue benchmarkGitHub2025
AMQA: An Adversarial Dataset for Benchmarking Bias of LLMs in Medicine and HealthcareGitHub2025
Beyond Benchmarks: Dynamic, Automatic and Systematic Red-Teaming Agents for Trustworthy Medical Language ModelsGitHub2025
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical TasksGitHub2025
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical ReasoningGitHub2025
MedBrowseComp: Benchmarking Medical Deep Research and Computer UseGitHub2025
MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-Checking of LLM ResponsesGitHub2025
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingGitHub2025
Medresearcher-r1: Expert-level medical deep researcher via a knowledge-informed trajectory synthesis frameworkGitHub2025

5.2 Metrics

TitleGitHubYear
Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media ForumNot Available2023
What Disease Does This Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical ExamsGitHub2023
AI Agents in Clinical Medicine: A Systematic ReviewNot Available2025
Audited Reasoning Refinement: Fine-Tuning Language Models via LLM-Guided Step-Wise Evaluation and CorrectionNot Available2025
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology ReasoningNot Available2025
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in MedicineNot Available2025
Evaluating the accuracy of a state-of-the-art large language model for prediction of admissions from the emergency roomNot Available2025
Geometry-preserving encoder/decoder in latent generative modelsGitHub2025
How Can We Diagnose and Treat Bias in Large Language Models for Clinical Decision-Making?Not Available2025
Large language models in real-world clinical workflows: a systematic review of applications and implementationNot Available2025
MedAgentBench: Dataset for Benchmarking LLMs as AgentsGitHub2025
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationNot Available2025
Performance of Retrieval-Augmented Generation Large Language Models in Guideline-Concordant Prostate-Specific Antigen Testing: Comparative Study With Junior CliniciansNot Available2025
Real-World Evaluation of Large Language Models in Healthcare (RWE-LLM): A New Realm of AI Safety & ValidationNot Available2025
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought ProcessesNot Available2025

5.3 Challenge & Discussion

🚀 6. Communication & Collaboration Mechanisms

🚀 7. Others

TitleGitHubYear
Drugagent: Multi-agent large language model-based reasoning for drug-target interaction predictionNot Available2025
An Edge Based Multi-Agent Model for Improving Hospital Bed ManagementNot Available2023
DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical servicesNot Available2025
Lessons Learned from Evaluation of LLM based Multi-agents in Safer Therapy RecommendationGitHub2025
M3Builder: A Multi-Agent System for Automated Machine Learning in Medical ImagingGitHub2025
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical ReasoningNot Available2025
Resilient Multi-Agent Negotiation for Medical Supply Chains: Integrating LLMs and Blockchain for Transparent CoordinationNot Available2025
Survey and improvement strategies for gene prioritization with large language modelsNot Available2025
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible ExtensibilityNot Available2025
Demo: Healthcare Agent Orchestrator (HAO) for Patient Summarization in Molecular Tumor BoardsNot Available2025
Mediator-guided multi-agent collaboration among open-source models for medical decision-makingNot Available2025
A co-evolving agentic AI system for medical imaging analysisGitHub2025
FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research InsightsGitHub2025
Meddxagent: A unified modular agent framework for explainable automatic differential diagnosisGitHub2025
AURA: A Multi-modal Medical Agent for Understanding, Reasoning and AnnotationGitHub2025
GEMA-Score: Granular Explainable Multi-Agent Score for Radiology Report EvaluationGitHub2025
PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray ReasoningGitHub2025
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False PresuppositionsGitHub2025
EH-Benchmark: Ophthalmic hallucination benchmark and agent-driven top-down traceable reasoning workflowGitHub2025
Medchat: A multi-agent framework for multimodal diagnosis with large language modelsGitHub2025
Autohealth: Advanced llm-empowered wearable personalized medical butler for parkinson’s disease managementNot Available2024
TWIN-GPT: digital twins for clinical trials via large language modelNot Available2024
A Proposed LLM-Based Supported Treatment Framework for Intracerebral HemorrhageNot Available2025
Agentic AI for Clinical Decision Support: Real-Time Diagnosis, Triage, and Treatment PlanningNot Available2025
Agentic Workflows in Healthcare: Advancing Clinical Efficiency through AI IntegrationNot Available2025
Developing an Artificial Intelligence Tool for Personalized Breast Cancer Treatment Plans based on the NCCN GuidelinesNot Available2025
TxAgent: An AI agent for therapeutic reasoning across a universe of toolsGitHub2025
Allocation of physician time in ambulatory practice: a time and motion study in 4 specialtiesNot Available2023
Tethered to the EHR: primary care physician workload assessment using EHR event log data and time-motion observationsNot Available2023
The NCI Imaging Data Commons as a platform for reproducible research in computational pathologyNot Available2023
At-cxr: Uncertainty-aware agentic triage for chest x-raysGitHub2025
In-Basket Message Volume in Primary Care: A Cross-sectional Analysis by Gender and SpecialtyNot Available2025
Modeling irregularly sampled clinical time seriesGitHub2023
Multimodal Models in Healthcare: Methods, Challenges, and Future Directions for Enhanced Clinical Decision SupportNot Available2025
Visual-Conversational Interface for Evidence-Based Explanation of Diabetes Risk PredictionGitHub2025
Autonomous systems and artificial intelligence in healthcare transformation to 5P medicine--ethical challengesNot Available2023
Cognitive architectures for language agentsNot Available2023
Generative agents: Interactive simulacra of human behaviorGitHub2023
MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoningNot Available2023
React: Synergizing reasoning and acting in language modelsGitHub2023
Upper processing stages of the perception--action cycleNot Available2023
A survey on large language model based autonomous agentsNot Available2024
Tool learning with large language models: A surveyGitHub2025

Contributors

NUS-Project

183 commits

eolasdfjnas

5 commits

Data-Designer

2 commits

yunhang8658

2 commits

Languages

HTML

100.0%