gudehhh666/Awesome_Scientific_Agent

Awesome Scientific Agent

Python

79

28 commits

updated Aug 25, 2026

See the code

README

Awesome Scientific Agent

A curated map of LLM-based agents for automated scientific research.

Survey Taxonomy Benchmarks Watchlist License

Overview • Taxonomy • Methods • Agentic Repos • Construction • Enhancement • Benchmarks • Contributing • Citation

✨ Overview

Large language model agents are becoming a practical interface for AI for Science (AI4S): reading literature, generating hypotheses, designing experiments, operating tools, analyzing results, and reviewing research outputs. This repository organizes the scientific-agent landscape around the research workflow and the autonomy level of each system.

What this repository mapsHow to read it
Agent levelsAssistant, Partner, and Avatar describe increasing autonomy and responsibility in scientific workflows.
Research stagesLiterature, Hypothesis, Design, Verification, Analysis, and Evaluation show where each method contributes.
Agent componentsReasoning, memory, and collaboration enhancements highlight how systems are built beyond a base LLM.
Evaluation resourcesBenchmarks and datasets help compare scientific-agent capabilities across domains.

Scientific agents across the research lifecycle

📖 Reading Map

SectionUse it for
💡 TaxonomyCompare representative scientific agents by level, domain, backbone, capability, components, and research stage.
🧩 Method GuidesRead focused method notes for construction, enhancement, evaluation, and auto-research systems.
✈️ Scientific Agents ConstructionFind papers and systems about building agent workflows, prompts, tools, context, and domain interfaces.
🚀 Scientific Agents EnhancementExplore reasoning, memory, workflow, and self-improvement techniques for stronger agents.
⚖️ Benchmark For Scientific AgentsLocate benchmarks for scientific reasoning, code generation, data analysis, citation, and domain evaluation.
🧭 Broader Agentic & Auto-Research RepositoriesTrack influential repositories beyond scientific agents, including Karpathy-style training stacks, coding agents, browser agents, and orchestration frameworks.

🖼️ Survey Figure Index

The figures below are synchronized from the TPAMI survey materials so the repository mirrors the paper's structure rather than acting only as a paper list.

FigureWhat it explainsRepository file
Research lifecycleHow scientific agents support literature, hypothesis, design, verification, analysis, and evaluation.overall_short.png
Survey organizationHigh-level organization of the survey and repository map.overall.png
E/M role taxonomyCapability envelope and capability maturity, inducing Assistant, Partner, and Avatar roles.level.png
Extended role viewAdditional role-level view used by the survey materials.level_2.png
Construction overviewAgent construction methodology.construction_overview.png
Knowledge organizationHow scientific agents organize domain knowledge.knowledge_organization.png
OrchestrationCoordination and workflow orchestration in scientific-agent construction.orchestration_flat.png
Enhancement overviewOverview of scientific-agent capability enhancement.enhancement_overview.png
Memory systemsMemory structures for scientific agents.memory.png
Reasoning enhancementReasoning enhancement patterns for scientific agents.reasoning.png
Benchmark overviewScientific-agent benchmark and evaluation metric landscape.benchmark_overview.png

💡 Taxonomy

overall

Taxonomy Dashboard

LevelRole in scientific workCountTypical scope
AssistantHelps with bounded scientific tasks under direct human steering.33Literature synthesis, QA, design assistance, and analysis support.
PartnerCollaborates across multiple workflow steps with stronger tool use or feedback loops.20Ideation, experiment planning, automation, review, and domain reasoning.
AvatarActs as a higher-autonomy research executor in digital or physical environments.15Autonomous labs, discovery loops, and end-to-end research.
Capability / componentMeaning
ECapability envelope: the breadth of scientific workflow coverage.
MCapability maturity: the maturity of autonomy and execution.
R / Mem. / CReasoning enhancement, memory enhancement, and collaboration enhancement.
StagesLiterature, Hypothesis, Design, Verification, Analysis, and Evaluation.
Full taxonomy table — 68 scientific agents across levels, domains, capabilities, components, and research stages

Capability envelope is abbreviated as E, capability maturity as M, reasoning enhancement as R, memory enhancement as Mem., and collaboration enhancement as C.

LevelMethodDomainLLM BackboneEMRMem.CApplication StagesTask Description
AssistantLitLLMGeneralGeneral-purposeE1M1NoNoNoLiterature, AnalysisLiterature review synthesis
Assistantotto-SRMedicalGeneral-purposeE1M1YesNoNoLiterature, AnalysisSystematic review synthesis
AssistantSciMONGeneralGeneral-purposeE1M1YesNoNoLiterature, HypothesisNovel hypothesis generation
AssistantKG-FMMaterialsGeneral-purposeE1M1YesNoNoLiterature, Hypothesis, AnalysisKnowledge-grounded materials QA
AssistantHypoGenGeneralGeneral-purposeE1M1YesNoNoHypothesisResearch hypothesis generation
AssistantLLM-SRPhysicsGeneral-purposeE1M1YesYesNoHypothesis, AnalysisSymbolic equation discovery
AssistantInstructMolChemistryDomain-specializedE1M1NoNoNoVerification, AnalysisMolecular instruction following
AssistantGeneGPTMedicalGeneral-purposeE1M1YesNoNoAnalysisGenomic question answering
AssistantTAISMedicalGeneral-purposeE1M1YesNoYesVerification, AnalysisGene expression analysis
AssistantDrugAgentMedicalGeneral-purposeE1M1YesNoYesAnalysisML-driven drug discovery
AssistantDrugGenMedicalGeneral-purposeE1M1NoNoNoDesign, VerificationTargeted molecular generation
AssistantChemAgentChemistryGeneral-purposeE1M1YesYesNoDesign, Verification, AnalysisMulti-step chemical reasoning
AssistantChatChemTSChemistryDomain-specializedE1M1NoNoNoDesign, VerificationConversational molecule generation
AssistantPaperQAGeneralGeneral-purposeE1M2YesYesNoLiterature, AnalysisScientific document QA
AssistantChatCiteGeneralGeneral-purposeE1M2YesYesNoLiterature, AnalysisEvidence-aware literature synthesis
AssistantCoI-AgentGeneralGeneral-purposeE1M2YesYesNoLiterature, HypothesisChain-structured ideation
AssistantDeep IdeationGeneralGeneral-purposeE1M2YesYesNoHypothesis, AnalysisConcept-network ideation
AssistantIRISGeneralGeneral-purposeE1M2YesNoYesHypothesis, AnalysisInteractive hypothesis search
AssistantLlaSMolChemistryDomain-specializedE1M2YesNoNoVerificationMolecular design assistance
AssistantEther0ChemistryDomain-specializedE1M2YesNoNoDesign, VerificationComplex molecular design
AssistantChemCrowChemistryGeneral-purposeE1M2YesNoNoLiterature, Design, VerificationTool-augmented chemistry assistance
AssistantHoneyCombMaterialsGeneral-purposeE1M2YesNoYesLiterature, Design, VerificationMaterials design assistance
AssistantPaperCoderComputer ScienceGeneral-purposeE1M2YesNoYesLiterature, Design, VerificationPaper-to-code generation
AssistantBioResearcherBiomedicalGeneral-purposeE2M1YesNoNoLiterature, Hypothesis, Design, Verification, AnalysisBiological workflow assistance
AssistantProtAgentsBiologyGeneral-purposeE2M1YesNoYesHypothesis, Design, VerificationProtein design loop
AssistantMOOSE-ChemChemistryGeneral-purposeE2M1NoYesNoHypothesis, VerificationChemical hypothesis generation
AssistantMeta-OpenFoamPhysicsGeneral-purposeE2M1YesNoYesDesign, Verification, AnalysisCFD workflow orchestration
AssistantFoamAgentPhysicsGeneral-purposeE2M1YesNoYesDesign, Verification, AnalysisNatural-language CFD execution
AssistantPiFlowGeneralGeneral-purposeE2M1YesNoYesHypothesis, Design, Verification, AnalysisPrinciple-guided experiment loops
AssistantDrBioRight 2.0BiologyGeneral-purposeE2M1YesNoNoVerification, AnalysisBioinformatics workflow analysis
AssistantOriGeneMedicalGeneral-purposeE2M1YesYesYesHypothesis, Verification, AnalysisTarget discovery and validation
AssistantCellVoyagerBiologyGeneral-purposeE2M1YesNoNoHypothesis, Design, Verification, AnalysisAutonomous scRNA-seq analysis
AssistantVASPilotMaterialsGeneral-purposeE2M1YesNoYesDesign, VerificationAutonomous DFT execution
PartnerDARWIN 1.5Biology/ChemistryGeneral-purposeE2M2NoNoNoLiterature, Verification, AnalysisDomain reasoning and analysis
PartnerCrispr-GPTBiologyGeneral-purposeE2M2YesYesNoHypothesis, Design, EvaluationCRISPR design assistance
PartnerChemmaChemistryGeneral-purposeE2M2YesNoNoLiterature, Hypothesis, DesignProperty-guided synthesis planning
PartnerMRAgentMedicalGeneral-purposeE2M2YesNoNoLiterature, Design, Verification, AnalysisMR-based medical inference
PartnerAviaryHybridGeneral-purposeE2M2YesNoNoLiterature, Hypothesis, Design, Verification, AnalysisGeneral scientific assistance
PartnerVirtual LabGeneralGeneral-purposeE2M2YesNoYesLiterature, Hypothesis, Design, Verification, Analysis, EvaluationPI-guided virtual experimentation
PartnerDeepRareMedicineGeneral-purposeE2M2YesYesYesLiterature, Hypothesis, Analysis, EvaluationRare-disease differential diagnosis
PartnerMatPilotMaterialsGeneral-purposeE2M2YesNoYesLiterature, Hypothesis, Design, Verification, AnalysisLanguage-driven materials design
PartnerSciToolAgentGeneralGeneral-purposeE2M2YesNoYesLiterature, Hypothesis, Design, Verification, AnalysisTool-grounded scientific reasoning
PartnerOrganaChemistryGeneral-purposeE2M2NoNoNoDesign, VerificationHuman-guided robotic chemistry
PartnerFunSearchMathematicsGeneral-purposeE2M2YesYesNoHypothesis, Design, AnalysisEvolutionary mathematical discovery
PartnerStarWhisperAstronomyGeneral-purposeE2M2YesYesYesDesign, VerificationAutonomous telescope operations
PartnerCycleResearcherComputer ScienceGeneral-purposeE2M2YesNoYesLiterature, Hypothesis, EvaluationIterative paper improvement
PartnerBiomniBiologyGeneral-purposeE2M2YesYesYesLiterature, Hypothesis, Design, Verification, AnalysisBroad biological automation
PartnerSciAgentsGeneralGeneral-purposeE2M2YesYesYesLiterature, Hypothesis, AnalysisKG-guided scientific discovery
PartnerAI ScientistComputer ScienceGeneral-purposeE3M1NoNoNoHypothesis, Design, Verification, Analysis, EvaluationEnd-to-end CS research
PartnerAI-ResearcherComputer ScienceGeneral-purposeE3M1YesNoYesLiterature, Hypothesis, Analysis, EvaluationEnd-to-end research assistance
PartnerAgentrxivComputer ScienceGeneral-purposeE3M1YesYesYesLiterature, Hypothesis, Verification, AnalysisPreprint-grounded agent research
PartnerAgent LaboratoryComputer ScienceGeneral-purposeE3M1YesYesYesLiterature, Hypothesis, Design, Analysis, EvaluationEnd-to-end CS experimentation
AvatarA-LabMaterialsGeneral-purposeE2M3NoYesNoLiterature, Hypothesis, Design, Verification, AnalysisAutonomous materials synthesis
AvatarAlphaEvolveGeneralGeneral-purposeE2M3YesYesNoHypothesis, Design, VerificationEvolutionary scientific optimization
AvatarOpenEvidenceMedicineDomain-specializedE2M3YesNoNoLiterature, Analysis, EvaluationPoint-of-care clinical decision
AvatarAILAMaterialsGeneral-purposeE2M3YesNoYesDesign, Verification, AnalysisAutonomous instrument operation
AvatarMOSAIC-chemistryChemistryGeneral-purposeE2M3YesNoYesDesign, Verification, AnalysisCollective synthesis planning
AvatarMARSMaterialsGeneral-purposeE2M3YesNoYesLiterature, Hypothesis, Design, Verification, AnalysisRobotic materials discovery
AvatarScienceOneBiologyDomain-specializedE3M2YesYesYesLiterature, Hypothesis, Design, Verification, Analysis, EvaluationEnd-to-end scientific automation
AvatarAI co-scientistGeneralGeneral-purposeE3M2YesYesYesLiterature, Hypothesis, Design, Verification, AnalysisMulti-agent scientific co-discovery
AvatarAI Scientist-v2Computer ScienceGeneral-purposeE3M2YesNoNoHypothesis, Design, Verification, Analysis, EvaluationWorkshop-level CS research
AvatarCoscientistChemistryGeneral-purposeE3M2YesNoNoLiterature, Hypothesis, Design, Verification, AnalysisAutonomous chemistry execution
AvatarRobinBiologyGeneral-purposeE3M2YesYesYesLiterature, Hypothesis, Design, Verification, Analysis, EvaluationMulti-agent biological discovery
AvatarSparksBiologyGeneral-purposeE3M2YesYesYesHypothesis, Design, Verification, Analysis, EvaluationMulti-agent protein design
AvatarInternAgent-1.5GeneralDomain-specializedE3M2YesYesYesLiterature, Hypothesis, Design, Verification, AnalysisLong-horizon scientific discovery
PartnerSR-ScientistPhysics/MathematicsGeneral-purposeE2M2YesNoNoHypothesis, Design, Verification, AnalysisLong-horizon agentic scientific equation discovery
AvatarEvoScientistComputer ScienceGeneral-purposeE3M2YesYesYesLiterature, Hypothesis, Design, Verification, Analysis, EvaluationSelf-evolving multi-agent end-to-end discovery
AvatarSelf-Evolving Fluid Control AgentPhysics/EngineeringGeneral-purposeE2M3YesNoNoHypothesis, Design, Verification, AnalysisAutonomous physically reasoned controller discovery

Level 1: Agent As Assistant

Level 1: Agent as Assistant - bounded task support and domain-specific assistance
  • AstroLLaMA‑Chat: Scaling AstroLLaMA with Conversational and Diverse Datasets — arXiv, 2024 · paper
  • BioGPT: Generative Pre‑trained Transformer for Biomedical Text Generation and Mining — arXiv, 2022 · paper
  • DARWIN 1.5: Large Language Models as Materials‑Science Foundation Models — arXiv, 2024 · paper
  • ChemBERTa: Large‑Scale Self‑Supervised Pre‑training for Molecular Property Prediction — arXiv, 2020 · paper
  • ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation — arXiv, 2025 · paper
  • ChemDFM: A Large Language Foundation Model for Chemistry — arXiv, 2024 · paper
  • LlaSMol: Advancing Large Language Models for Chemistry with a Large‑Scale, Comprehensive, High‑Quality Instruction Tuning Dataset — arXiv, 2024 · paper
  • InstructMol: Multi‑modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery — arXiv, 2023 · paper
  • ether0: A Scientific Reasoning Model for Chemistry — arXiv, 2025 · paper
  • Leveraging Large Language Models for Predictive Chemistry — Nature, 2024 · paper
  • Multi‑modal Molecule Structure–Text Model for Text‑based Retrieval and Editing — Nature, 2023 · paper
  • ClimateGPT: Towards AI Synthesizing Interdisciplinary Research on Climate Change — arXiv, 2024 · paper
  • Sparks of Science: Hypothesis Generation Using Structured Paper Data — arXiv, 2025 · paper
  • DeepSeek‑Prover‑V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition — arXiv, 2025 · paper
  • BiMediX: Bilingual Medical Mixture of Experts LLM — arXiv, 2024 · paper
  • ChatDoctor: A Medical Chat Model Fine‑tuned on a Large Language Model (LLAMA) Using Medical Domain Knowledge — arXiv, 2023 · paper
  • AgentMD: Empowering Language Agents for Risk Prediction with Large‑Scale Clinical Tool Learning — arXiv, 2024 · paper
  • MedAlpaca: An Open‑Source Collection of Medical Conversational AI Models and Training Data — arXiv, 2023 · paper
  • DrugGen Enhances Drug Discovery with Large Language Models and Reinforcement Learning — Nature, 2025 · paper
  • LLM‑SR: Scientific Equation Discovery via Programming with Large Language Models — arXiv, 2024 · paper
  • LitLLM: A Toolkit for Scientific Literature Review — arXiv, 2024 · paper
  • SciBERT: A Pretrained Language Model for Scientific Text — arXiv, 2019 · paper
  • SciMON: Scientific Inspiration Machines Optimized for Novelty — arXiv, 2023 · paper
  • SCITUNE: Aligning Large Language Models with Scientific Multimodal Instructions — arXiv, 2023 · paper
  • NatureLM: Deciphering the Language of Nature for Scientific Discovery — arXiv, 2025 · paper

Level 2: Agent As Partner

Level 2: Agent as Partner - multi-step collaboration and workflow orchestration
  • StarWhisper Telescope: Agent‑Based Observation Assistant System to Approach an AI Astrophysicist — arXiv, 2024 · paper
  • From Intention to Implementation: Automating Biomedical Research via LLMs — arXiv, 2024 · paper
  • CRISPR‑GPT: An LLM Agent for Automated Design of Gene‑Editing Experiments — arXiv, 2024 · paper
  • Towards an AI Co‑Scientist — arXiv, 2025 · paper
  • DrBioRight 2.0: An LLM‑Powered Bioinformatics Chatbot for Large‑Scale Cancer Functional Proteomics Analysis — Nature, 2025 · paper
  • ProtAgents: Protein Discovery via Large Language Model Multi‑Agent Collaboration — arXiv, 2024 · paper
  • MOOSE‑Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses — arXiv, 2024 · paper
  • Autonomous Chemical Research with Large Language Models — Nature, 2023 · paper
  • ChemCrow: Augmenting Large‑Language Models with Chemistry Tools — arXiv, 2023 · paper
  • ORGANA: A Robotic Assistant for Automated Chemistry Experimentation and Characterization — arXiv, 2024 · paper
  • xChemAgents: Agentic AI for Explainable Quantum Chemistry — arXiv, 2025 · paper
  • ChemAgent: Self‑Updating Memories in Large Language Models Improves Chemical Reasoning — arXiv, 2025 · paper
  • Large Language Models to Accelerate Organic Chemistry Synthesis — arXiv, 2025 · paper
  • MyCrunchGPT: A ChatGPT‑Assisted Framework for Scientific Machine Learning — arXiv, 2023 · paper
  • MetaOpenFoam: An LLM‑Based Multi‑Agent Framework for CFD — arXiv, 2024 · paper
  • FoamAgent: Towards Automated Intelligent CFD Workflows — arXiv, 2025 · paper
  • AI‑Researcher: Autonomous Scientific Innovation — arXiv, 2025 · paper
  • Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs — arXiv, 2025 · paper
  • FlowAgent: Achieving Compliance and Flexibility for Workflow Agents — arXiv, 2025 · paper
  • The AI Scientist: Towards Fully Automated Open‑Ended Scientific Discovery — arXiv, 2024 · paper
  • GeoGPT: Understanding and Processing Geospatial Tasks through an Autonomous GPT — arXiv, 2023 · paper
  • PaperQA: Retrieval‑Augmented Generative Agent for Scientific Research — arXiv, 2023 · paper
  • ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary — arXiv, 2024 · paper
  • PiFlow: Principle‑Aware Scientific Discovery with Multi‑Agent Collaboration — arXiv, 2025 · paper
  • Aviary: Training Language Agents on Challenging Scientific Tasks — arXiv, 2024 · paper
  • An Autonomous Laboratory for the Accelerated Synthesis of Novel Materials — Nature, 2023 · paper
  • Construction of a Knowledge Graph for Framework Material Enabled by Large Language Models and Its Application — npjCM, 2025 · paper
  • MultiCrossModal Automated Agent for Integrating Diverse Materials Science Data — arXiv, 2025 · paper
  • LLMatDesign: Autonomous Materials Discovery with Large Language Models — arXiv, 2024 · paper
  • Honeycomb: A Flexible LLM‑Based Agent System for Materials Science — arXiv, 2024 · paper
  • DrugAgent: Automating AI‑Aided Drug Discovery Programming through LLM Multi‑Agent Collaboration — arXiv, 2024 · paper
  • Toward a Team of AI‑Made Scientists for Scientific Discovery from Gene Expression Data — arXiv, 2024 · paper
  • MedAgents: Large Language Models as Collaborators for Zero‑Shot Medical Reasoning — arXiv, 2023 · paper
  • Automation of Systematic Reviews with Large Language Models — medRxiv, 2025 · paper
  • MRAgent: An LLM‑Based Automated Agent for Causal Knowledge Discovery in Disease via Mendelian Randomization — BriefBioinf, 2025 · website
  • GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information — arXiv, 2023 · paper
  • SR-Scientist: Scientific Equation Discovery With Agentic AI — ICLR, 2026 · paper · code
  • Fantastic Scientific Agents and How to Build Them: AgentBuild for Rietveld Refinement — arXiv, 2026 · paper
  • SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks — arXiv, 2026 · paper

Level 3: Agent As Avatar

Level 3: Agent as Avatar - autonomous research execution and long-horizon discovery
  • CycleResearcher: Improving Automated Research via Automated Review — arXiv, 2024 · paper
  • The AI Scientist‑v2: Workshop‑Level Automated Scientific Discovery via Agentic Tree Search — arXiv, 2025 · paper
  • AgentRxiv: Towards Collaborative Autonomous Research — arXiv, 2025 · paper
  • Agent Laboratory: Using LLM Agents as Research Assistants — arXiv, 2025 · paper
  • BiOMNI: A General‑Purpose Biomedical AI Agent — bioRxiv, 2025 · paper
  • OriGene: A Self‑Evolving Virtual Disease Biologist Automating Therapeutic Target Discovery — bioRxiv, 2025 · paper
  • CellVoyager: AI CompBio Agent Generates New Insights by Autonomously Analyzing Biological Data — bioRxiv, 2025 · paper
  • Sparks: Multi‑Agent Artificial Intelligence Model Discovers Protein Design Principles — arXiv, 2025 · paper
  • AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery — arXiv, 2025 · paper
  • Robin: A Multi‑Agent System for Automating Scientific Discovery — arXiv, 2025 · paper
  • ScienceOne — Project, 2025 · website
  • EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery — arXiv, 2026 · paper · code
  • Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control — arXiv, 2026 · paper

🧩 Method Guides

The root README keeps the curated map compact; the method guides provide deeper explanations for readers who want to understand how scientific agents are built, improved, evaluated, and connected to broader auto-research systems.

GuideWhat it covers
Construction MethodsKnowledge organization, knowledge injection, tool integration, orchestration, and domain interfaces.
Enhancement MethodsReasoning, memory, collaboration, workflow search, and self-review.
Evaluation MethodsBenchmark selection, executable evaluation, citation grounding, and long-horizon validation.
Auto-Research SystemsEnd-to-end research loops, coding/browser agents, training resources, and orchestration frameworks.

🧭 Broader Agentic & Auto-Research Repositories

Beyond scientific-agent papers, the broader agent ecosystem is moving quickly across model training, software engineering, web automation, and multi-agent orchestration. This watchlist keeps a lightweight bridge from the survey taxonomy to practical repositories that shape how autonomous research and agentic systems are built.

RepositoryScopeWhy follow it
karpathy/nanochatMinimal end-to-end LLM training and chat stackTracks Karpathy's compact, hackable path from tokenizer and pretraining to finetuning, evaluation, inference, and chat UI.
karpathy/llm.cLLM training in C/CUDAUseful for understanding low-level training kernels, performance constraints, and reproducible GPT-style training.
SakanaAI/AI-ScientistAutomated idea-to-paper research loopA reference point for autonomous ideation, coding, experimentation, paper writing, and automated review.
SakanaAI/AI-Scientist-v2Agentic tree search for automated discoveryFollows the next iteration of AI Scientist with broader exploration and stronger end-to-end workflow design.
SamuelSchmidgall/AgentLaboratoryHuman-guided autonomous research assistantShows how literature review, experimentation, and report writing can be composed into a full research workflow.
NoviScl/AI-ResearcherResearch ideation and execution studiesProvides agent pipelines and human-study data for comparing LLM-generated ideas with expert research ideas.
OpenHands/OpenHandsAI software development agentsA generalist coding-agent platform for editing repositories, using terminals, browsing, and operating in sandboxed environments.
SWE-agent/SWE-agentGitHub issue fixing and SWE-bench agentsA practical baseline for agentic software engineering, debugging, and repository-level task execution.
huggingface/smolagentsLightweight code-agent frameworkGood for studying minimal abstractions, code-as-action agents, sandboxed execution, and open-model agent workflows.
microsoft/autogenAgentic AI programming frameworkUseful for multi-agent conversations, orchestration patterns, and prototyping collaborative agent systems.
crewAIInc/crewAIMulti-agent orchestrationFocuses on role-based agents, task delegation, crews, flows, and production-style automation workflows.
langchain-ai/langgraphGraph-based agent workflowsUseful for durable, stateful, controllable agent graphs and long-running workflow orchestration.
FoundationAgents/MetaGPTMulti-agent software company metaphorA representative multi-agent framework for decomposing product/software work into role-specialized agents.
browser-use/browser-useBrowser automation for agentsTracks web interaction patterns, browser control, and task automation over ordinary websites.
Significant-Gravitas/AutoGPTEarly autonomous agent platformStill useful as historical context for goal-directed agents, autonomous task decomposition, and agent productization.
EvoScientist/EvoScientistSelf-evolving end-to-end AI scientistTracks persistent research memory, multi-agent experimentation, and human-on-the-loop research workflows.

✈️ Scientific Agents Construction

Agent construction methodology

Knowledge organization in scientific agents Orchestration and coordination in scientific-agent construction

Construction resources - prompts, context, tools, workflows, and domain interfaces
  • PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization — arXiv, 2023 · paper
  • Context Engineering: A Practical Handbook for Context Design, Orchestration, and Optimization — GitHub, 2025 · code
  • Verl: Volcano Engine Reinforcement Learning for Large Language Models — GitHub, 2024 · code
  • GraphRAG: A Modular Graph-Based Retrieval-Augmented Generation (RAG) System — arXiv, 2023 · paper
  • ChemCrow: Augmenting Large‑Language Models with Chemistry Tools — arXiv, 2023 · paper
  • Virtual Lab: AI Agents Design New SARS-CoV-2 Nanobodies with Experimental Validation — arXiv, 2024 · paper
  • OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning — arXiv, 2025 · paper
  • FoamAgent: Towards Automated Intelligent CFD Workflows — arXiv, 2025 · paper
  • MetaOpenFoam: An LLM-Based Multi-Agent Framework for CFD — arXiv, 2024 · paper
  • AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts — EMNLP, 2020 · paper
  • InstructBio: Instruction Tuning for Biomedical LLMs — arXiv, 2024 · paper
  • LangChain: Building Applications with LLMs through Composability — GitHub, 2022 · code
  • LightRAG: Simple and Fast Retrieval‑Augmented Generation — arXiv, 2024 · paper
  • SciTUNE: Aligning Large Language Models with Scientific Multimodal Instructions — arXiv, 2023 · paper
  • ClimateGPT: Towards AI Synthesizing Interdisciplinary Research on Climate Change — arXiv, 2024 · paper
  • SciMON: Scientific Inspiration Machines Optimized for Novelty — arXiv, 2023 · paper
  • TAIS: Gene Expression Agent with LLMs — arXiv, 2025 · paper
  • StarWhisper Telescope: Agent‑Based Observation Assistant System to Approach an AI Astrophysicist — arXiv, 2024 · paper
  • Fantastic Scientific Agents and How to Build Them: AgentBuild for Rietveld Refinement — arXiv, 2026 · paper

🚀 Scientific Agents Enhancement

Scientific-agent ability enhancement

Scientific agents' memory systems Illustration of scientific agent reasoning enhancement

Enhancement resources - reasoning, memory, workflow search, reflection, and collaboration
  • MemGPT: Toward LLMs as Operating Systems — arXiv, 2023 · paper
  • AFlow: Automating Agentic Workflow Generation — arXiv, 2023 · paper
  • ChemAgent: Self‑Updating Memories in Large Language Models Improves Chemical Reasoning — ICLR, 2025 · paper
  • DeepSeek‑Prover‑V2: Advancing Formal Mathematical Reasoning — arXiv, 2025 · paper
  • LLM‑SR: Scientific Equation Discovery via Programming with LLMs — ICLR, 2025 · paper
  • Sparks: Multi‑Agent Artificial Intelligence Model Discovers Protein Design Principles — arXiv, 2025 · paper
  • ether0: A Scientific Reasoning Model for Chemistry — arXiv, 2025 · paper
  • Robin: A Multi‑Agent System for Automating Scientific Discovery — arXiv, 2025 · paper
  • AgentRxiv: Towards Collaborative Autonomous Research — arXiv, 2025 · paper
  • ChatGPT Research Group for Optimizing the Crystallinity of MOFs and COFs — ACS, 2023 · paper
  • AI Co-Scientist: Towards an AI Co-Scientist — arXiv, 2025 · paper
  • ReAct: Synergizing Reasoning and Acting in Language Models — arXiv, 2022 · paper
  • Reflexion: Language Agents with Verbal Reinforcement Learning — NeurIPS, 2023 · paper
  • Self‑Refine: Iterative Self‑Improvement with Self‑Feedback — arXiv, 2023 · paper
  • Self‑Consistency: Reliable Decoding for Complex Reasoning — ICLR, 2023 · paper
  • AquilaChat: Agent with Long-Context Scratchpad Memory — arXiv, 2024 · paper
  • EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery — arXiv, 2026 · paper
  • SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks — arXiv, 2026 · paper
  • Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control — arXiv, 2026 · paper

⚖️ Benchmark For Scientific Agents

Use this section as an evaluation map rather than a flat benchmark list. The resources below cover different failure modes of scientific agents: domain knowledge, executable experiments, citation grounding, data analysis, and long-horizon discovery.

Scientific-agent benchmark and evaluation landscape

Evaluation angleRepresentative focusUseful when you need to test...
Scientific knowledge and reasoningBioMaze, SuperGPQA, Humanity's Last Exam, MR-BenWhether an agent can reason over expert-level scientific concepts.
Citation and literature groundingCiteBench, ALCE, SurveyForgeWhether outputs are traceable, evidence-aware, and literature-faithful.
Code, data, and experiment executionMLAgentBench, DSBench, SciCode, PaperBenchWhether an agent can implement, run, debug, and reproduce research workflows.
Domain and embodied environmentsDiscoveryWorld, AgentClinic, GenoTEX, LLM-SRBenchWhether an agent performs in domain-specific or simulated scientific settings.
Benchmark resources — evaluation suites for scientific reasoning, data analysis, citation, coding, and agentic discovery
  • BioMaze: Benchmarking and Enhancing Large Language Models for Biological Pathway Reasoning — arXiv, 2025 · paper · dataset
  • BioKGBench: A Knowledge Graph Checking Benchmark of AI Agent for Biomedical Science — GitHub, 2025 · code
  • SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-Dimensional Evaluation for Automated Survey Writing — arXiv, 2025 · paper · code
  • MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark — NeurIPS, 2024 · paper · dataset
  • Humanity's Last Exam: A Hard Benchmark at the Frontier of Human Knowledge — arXiv, 2025 · paper · website
  • SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines — arXiv, 2025 · paper · dataset
  • CiteBench: A Benchmark for Scientific Citation Text Generation — arXiv, 2022 · paper · code
  • ALCE: Enabling Large Language Models to Generate Text With Citations — arXiv, 2023 · paper · code
  • Tomato-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses — arXiv, 2024 · paper · code
  • Reviewer2: Optimizing Review Generation Through Prompt Generation — arXiv, 2024 · paper · code
  • Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning — arXiv, 2025 · dataset · paper
  • MR-Ben: A Comprehensive Meta-Reasoning Benchmark for Large Language Models — arXiv, 2024 · paper · website
  • FigureQA: An Annotated Figure Dataset for Visual Reasoning — arXiv, 2017 · paper · dataset
  • SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research — AAAI, 2024 · paper · dataset
  • MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding — arXiv, 2025 · paper · dataset
  • LLM-SRBench: A New Benchmark for Scientific Equation Discovery With Large Language Models — arXiv, 2025 · paper · dataset
  • GenoTEX: A Benchmark for Evaluating LLM-Based Exploration of Gene Expression Data in Alignment With Bioinformaticians — arXiv, 2024 · paper · dataset
  • MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation — arXiv, 2023 · paper · code
  • PaperBench: Evaluating AI's Ability to Replicate AI Research — arXiv, 2025 · paper
  • DSCodeBench: A Realistic Benchmark for Data Science Code Generation — AAAI, 2026 · paper · code
  • DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents — arXiv, 2024 · paper · code
  • SciCode: A Scientist-Curated Benchmark for Scientific Code Generation — arXiv, 2024 · paper · code
  • AgentClinic: A Multimodal Agent Benchmark to Evaluate AI in Simulated Clinical Environments — arXiv, 2024 · code · paper
  • SDRBench: Scientific Data Reduction Benchmark for Lossy Compressors — Website, 2021 · website · website
  • DSBench: How Far Are Data Science Agents from Becoming Data Science Experts? — GitHub, 2025 · code · paper
  • AISB: AI Scientist Benchmark — NLPCC, 2026 · evaluation kit
  • DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? — arXiv, 2026 · paper · code

🤝 Contributing

Catalog tables and lists are generated from structured files in data/. See CONTRIBUTING.md for the evidence and maintenance workflow, and the catalog schema for field definitions. Edit templates/README.md for narrative content, then regenerate this file with python3 scripts/render_readme.py.

🌞 Citation

@article{wang2025hitchhiker,
  title={The Hitchhiker's Guide to Autonomous Research: A Survey of Scientific Agents},
  author={Wang, Xinming and Xu, Jian and Feng, Aslan H and Chen, Yi and Guo, Haiyang and Zhu, Fei and Shao, Yuanqi and Ren, Minsi and Yi, Hongzhu and Lian, Sheng and others},
  year={2025}
}

Contributors

gudehhh666

27 commits

banjiuyufen

1 commits

gudehhh666/Awesome_Scientific_Agent

Awesome Scientific Agent

Python

79

28 commits

updated Aug 25, 2026

See the code

README

Awesome Scientific Agent

A curated map of LLM-based agents for automated scientific research.

Survey Taxonomy Benchmarks Watchlist License

Overview • Taxonomy • Methods • Agentic Repos • Construction • Enhancement • Benchmarks • Contributing • Citation

✨ Overview

Large language model agents are becoming a practical interface for AI for Science (AI4S): reading literature, generating hypotheses, designing experiments, operating tools, analyzing results, and reviewing research outputs. This repository organizes the scientific-agent landscape around the research workflow and the autonomy level of each system.

What this repository mapsHow to read it
Agent levelsAssistant, Partner, and Avatar describe increasing autonomy and responsibility in scientific workflows.
Research stagesLiterature, Hypothesis, Design, Verification, Analysis, and Evaluation show where each method contributes.
Agent componentsReasoning, memory, and collaboration enhancements highlight how systems are built beyond a base LLM.
Evaluation resourcesBenchmarks and datasets help compare scientific-agent capabilities across domains.

Scientific agents across the research lifecycle

📖 Reading Map

SectionUse it for
💡 TaxonomyCompare representative scientific agents by level, domain, backbone, capability, components, and research stage.
🧩 Method GuidesRead focused method notes for construction, enhancement, evaluation, and auto-research systems.
✈️ Scientific Agents ConstructionFind papers and systems about building agent workflows, prompts, tools, context, and domain interfaces.
🚀 Scientific Agents EnhancementExplore reasoning, memory, workflow, and self-improvement techniques for stronger agents.
⚖️ Benchmark For Scientific AgentsLocate benchmarks for scientific reasoning, code generation, data analysis, citation, and domain evaluation.
🧭 Broader Agentic & Auto-Research RepositoriesTrack influential repositories beyond scientific agents, including Karpathy-style training stacks, coding agents, browser agents, and orchestration frameworks.

🖼️ Survey Figure Index

The figures below are synchronized from the TPAMI survey materials so the repository mirrors the paper's structure rather than acting only as a paper list.

FigureWhat it explainsRepository file
Research lifecycleHow scientific agents support literature, hypothesis, design, verification, analysis, and evaluation.overall_short.png
Survey organizationHigh-level organization of the survey and repository map.overall.png
E/M role taxonomyCapability envelope and capability maturity, inducing Assistant, Partner, and Avatar roles.level.png
Extended role viewAdditional role-level view used by the survey materials.level_2.png
Construction overviewAgent construction methodology.construction_overview.png
Knowledge organizationHow scientific agents organize domain knowledge.knowledge_organization.png
OrchestrationCoordination and workflow orchestration in scientific-agent construction.orchestration_flat.png
Enhancement overviewOverview of scientific-agent capability enhancement.enhancement_overview.png
Memory systemsMemory structures for scientific agents.memory.png
Reasoning enhancementReasoning enhancement patterns for scientific agents.reasoning.png
Benchmark overviewScientific-agent benchmark and evaluation metric landscape.benchmark_overview.png

💡 Taxonomy

overall

Taxonomy Dashboard

LevelRole in scientific workCountTypical scope
AssistantHelps with bounded scientific tasks under direct human steering.33Literature synthesis, QA, design assistance, and analysis support.
PartnerCollaborates across multiple workflow steps with stronger tool use or feedback loops.20Ideation, experiment planning, automation, review, and domain reasoning.
AvatarActs as a higher-autonomy research executor in digital or physical environments.15Autonomous labs, discovery loops, and end-to-end research.
Capability / componentMeaning
ECapability envelope: the breadth of scientific workflow coverage.
MCapability maturity: the maturity of autonomy and execution.
R / Mem. / CReasoning enhancement, memory enhancement, and collaboration enhancement.
StagesLiterature, Hypothesis, Design, Verification, Analysis, and Evaluation.
Full taxonomy table — 68 scientific agents across levels, domains, capabilities, components, and research stages

Capability envelope is abbreviated as E, capability maturity as M, reasoning enhancement as R, memory enhancement as Mem., and collaboration enhancement as C.

LevelMethodDomainLLM BackboneEMRMem.CApplication StagesTask Description
AssistantLitLLMGeneralGeneral-purposeE1M1NoNoNoLiterature, AnalysisLiterature review synthesis
Assistantotto-SRMedicalGeneral-purposeE1M1YesNoNoLiterature, AnalysisSystematic review synthesis
AssistantSciMONGeneralGeneral-purposeE1M1YesNoNoLiterature, HypothesisNovel hypothesis generation
AssistantKG-FMMaterialsGeneral-purposeE1M1YesNoNoLiterature, Hypothesis, AnalysisKnowledge-grounded materials QA
AssistantHypoGenGeneralGeneral-purposeE1M1YesNoNoHypothesisResearch hypothesis generation
AssistantLLM-SRPhysicsGeneral-purposeE1M1YesYesNoHypothesis, AnalysisSymbolic equation discovery
AssistantInstructMolChemistryDomain-specializedE1M1NoNoNoVerification, AnalysisMolecular instruction following
AssistantGeneGPTMedicalGeneral-purposeE1M1YesNoNoAnalysisGenomic question answering
AssistantTAISMedicalGeneral-purposeE1M1YesNoYesVerification, AnalysisGene expression analysis
AssistantDrugAgentMedicalGeneral-purposeE1M1YesNoYesAnalysisML-driven drug discovery
AssistantDrugGenMedicalGeneral-purposeE1M1NoNoNoDesign, VerificationTargeted molecular generation
AssistantChemAgentChemistryGeneral-purposeE1M1YesYesNoDesign, Verification, AnalysisMulti-step chemical reasoning
AssistantChatChemTSChemistryDomain-specializedE1M1NoNoNoDesign, VerificationConversational molecule generation
AssistantPaperQAGeneralGeneral-purposeE1M2YesYesNoLiterature, AnalysisScientific document QA
AssistantChatCiteGeneralGeneral-purposeE1M2YesYesNoLiterature, AnalysisEvidence-aware literature synthesis
AssistantCoI-AgentGeneralGeneral-purposeE1M2YesYesNoLiterature, HypothesisChain-structured ideation
AssistantDeep IdeationGeneralGeneral-purposeE1M2YesYesNoHypothesis, AnalysisConcept-network ideation
AssistantIRISGeneralGeneral-purposeE1M2YesNoYesHypothesis, AnalysisInteractive hypothesis search
AssistantLlaSMolChemistryDomain-specializedE1M2YesNoNoVerificationMolecular design assistance
AssistantEther0ChemistryDomain-specializedE1M2YesNoNoDesign, VerificationComplex molecular design
AssistantChemCrowChemistryGeneral-purposeE1M2YesNoNoLiterature, Design, VerificationTool-augmented chemistry assistance
AssistantHoneyCombMaterialsGeneral-purposeE1M2YesNoYesLiterature, Design, VerificationMaterials design assistance
AssistantPaperCoderComputer ScienceGeneral-purposeE1M2YesNoYesLiterature, Design, VerificationPaper-to-code generation
AssistantBioResearcherBiomedicalGeneral-purposeE2M1YesNoNoLiterature, Hypothesis, Design, Verification, AnalysisBiological workflow assistance
AssistantProtAgentsBiologyGeneral-purposeE2M1YesNoYesHypothesis, Design, VerificationProtein design loop
AssistantMOOSE-ChemChemistryGeneral-purposeE2M1NoYesNoHypothesis, VerificationChemical hypothesis generation
AssistantMeta-OpenFoamPhysicsGeneral-purposeE2M1YesNoYesDesign, Verification, AnalysisCFD workflow orchestration
AssistantFoamAgentPhysicsGeneral-purposeE2M1YesNoYesDesign, Verification, AnalysisNatural-language CFD execution
AssistantPiFlowGeneralGeneral-purposeE2M1YesNoYesHypothesis, Design, Verification, AnalysisPrinciple-guided experiment loops
AssistantDrBioRight 2.0BiologyGeneral-purposeE2M1YesNoNoVerification, AnalysisBioinformatics workflow analysis
AssistantOriGeneMedicalGeneral-purposeE2M1YesYesYesHypothesis, Verification, AnalysisTarget discovery and validation
AssistantCellVoyagerBiologyGeneral-purposeE2M1YesNoNoHypothesis, Design, Verification, AnalysisAutonomous scRNA-seq analysis
AssistantVASPilotMaterialsGeneral-purposeE2M1YesNoYesDesign, VerificationAutonomous DFT execution
PartnerDARWIN 1.5Biology/ChemistryGeneral-purposeE2M2NoNoNoLiterature, Verification, AnalysisDomain reasoning and analysis
PartnerCrispr-GPTBiologyGeneral-purposeE2M2YesYesNoHypothesis, Design, EvaluationCRISPR design assistance
PartnerChemmaChemistryGeneral-purposeE2M2YesNoNoLiterature, Hypothesis, DesignProperty-guided synthesis planning
PartnerMRAgentMedicalGeneral-purposeE2M2YesNoNoLiterature, Design, Verification, AnalysisMR-based medical inference
PartnerAviaryHybridGeneral-purposeE2M2YesNoNoLiterature, Hypothesis, Design, Verification, AnalysisGeneral scientific assistance
PartnerVirtual LabGeneralGeneral-purposeE2M2YesNoYesLiterature, Hypothesis, Design, Verification, Analysis, EvaluationPI-guided virtual experimentation
PartnerDeepRareMedicineGeneral-purposeE2M2YesYesYesLiterature, Hypothesis, Analysis, EvaluationRare-disease differential diagnosis
PartnerMatPilotMaterialsGeneral-purposeE2M2YesNoYesLiterature, Hypothesis, Design, Verification, AnalysisLanguage-driven materials design
PartnerSciToolAgentGeneralGeneral-purposeE2M2YesNoYesLiterature, Hypothesis, Design, Verification, AnalysisTool-grounded scientific reasoning
PartnerOrganaChemistryGeneral-purposeE2M2NoNoNoDesign, VerificationHuman-guided robotic chemistry
PartnerFunSearchMathematicsGeneral-purposeE2M2YesYesNoHypothesis, Design, AnalysisEvolutionary mathematical discovery
PartnerStarWhisperAstronomyGeneral-purposeE2M2YesYesYesDesign, VerificationAutonomous telescope operations
PartnerCycleResearcherComputer ScienceGeneral-purposeE2M2YesNoYesLiterature, Hypothesis, EvaluationIterative paper improvement
PartnerBiomniBiologyGeneral-purposeE2M2YesYesYesLiterature, Hypothesis, Design, Verification, AnalysisBroad biological automation
PartnerSciAgentsGeneralGeneral-purposeE2M2YesYesYesLiterature, Hypothesis, AnalysisKG-guided scientific discovery
PartnerAI ScientistComputer ScienceGeneral-purposeE3M1NoNoNoHypothesis, Design, Verification, Analysis, EvaluationEnd-to-end CS research
PartnerAI-ResearcherComputer ScienceGeneral-purposeE3M1YesNoYesLiterature, Hypothesis, Analysis, EvaluationEnd-to-end research assistance
PartnerAgentrxivComputer ScienceGeneral-purposeE3M1YesYesYesLiterature, Hypothesis, Verification, AnalysisPreprint-grounded agent research
PartnerAgent LaboratoryComputer ScienceGeneral-purposeE3M1YesYesYesLiterature, Hypothesis, Design, Analysis, EvaluationEnd-to-end CS experimentation
AvatarA-LabMaterialsGeneral-purposeE2M3NoYesNoLiterature, Hypothesis, Design, Verification, AnalysisAutonomous materials synthesis
AvatarAlphaEvolveGeneralGeneral-purposeE2M3YesYesNoHypothesis, Design, VerificationEvolutionary scientific optimization
AvatarOpenEvidenceMedicineDomain-specializedE2M3YesNoNoLiterature, Analysis, EvaluationPoint-of-care clinical decision
AvatarAILAMaterialsGeneral-purposeE2M3YesNoYesDesign, Verification, AnalysisAutonomous instrument operation
AvatarMOSAIC-chemistryChemistryGeneral-purposeE2M3YesNoYesDesign, Verification, AnalysisCollective synthesis planning
AvatarMARSMaterialsGeneral-purposeE2M3YesNoYesLiterature, Hypothesis, Design, Verification, AnalysisRobotic materials discovery
AvatarScienceOneBiologyDomain-specializedE3M2YesYesYesLiterature, Hypothesis, Design, Verification, Analysis, EvaluationEnd-to-end scientific automation
AvatarAI co-scientistGeneralGeneral-purposeE3M2YesYesYesLiterature, Hypothesis, Design, Verification, AnalysisMulti-agent scientific co-discovery
AvatarAI Scientist-v2Computer ScienceGeneral-purposeE3M2YesNoNoHypothesis, Design, Verification, Analysis, EvaluationWorkshop-level CS research
AvatarCoscientistChemistryGeneral-purposeE3M2YesNoNoLiterature, Hypothesis, Design, Verification, AnalysisAutonomous chemistry execution
AvatarRobinBiologyGeneral-purposeE3M2YesYesYesLiterature, Hypothesis, Design, Verification, Analysis, EvaluationMulti-agent biological discovery
AvatarSparksBiologyGeneral-purposeE3M2YesYesYesHypothesis, Design, Verification, Analysis, EvaluationMulti-agent protein design
AvatarInternAgent-1.5GeneralDomain-specializedE3M2YesYesYesLiterature, Hypothesis, Design, Verification, AnalysisLong-horizon scientific discovery
PartnerSR-ScientistPhysics/MathematicsGeneral-purposeE2M2YesNoNoHypothesis, Design, Verification, AnalysisLong-horizon agentic scientific equation discovery
AvatarEvoScientistComputer ScienceGeneral-purposeE3M2YesYesYesLiterature, Hypothesis, Design, Verification, Analysis, EvaluationSelf-evolving multi-agent end-to-end discovery
AvatarSelf-Evolving Fluid Control AgentPhysics/EngineeringGeneral-purposeE2M3YesNoNoHypothesis, Design, Verification, AnalysisAutonomous physically reasoned controller discovery

Level 1: Agent As Assistant

Level 1: Agent as Assistant - bounded task support and domain-specific assistance
  • AstroLLaMA‑Chat: Scaling AstroLLaMA with Conversational and Diverse Datasets — arXiv, 2024 · paper
  • BioGPT: Generative Pre‑trained Transformer for Biomedical Text Generation and Mining — arXiv, 2022 · paper
  • DARWIN 1.5: Large Language Models as Materials‑Science Foundation Models — arXiv, 2024 · paper
  • ChemBERTa: Large‑Scale Self‑Supervised Pre‑training for Molecular Property Prediction — arXiv, 2020 · paper
  • ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation — arXiv, 2025 · paper
  • ChemDFM: A Large Language Foundation Model for Chemistry — arXiv, 2024 · paper
  • LlaSMol: Advancing Large Language Models for Chemistry with a Large‑Scale, Comprehensive, High‑Quality Instruction Tuning Dataset — arXiv, 2024 · paper
  • InstructMol: Multi‑modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery — arXiv, 2023 · paper
  • ether0: A Scientific Reasoning Model for Chemistry — arXiv, 2025 · paper
  • Leveraging Large Language Models for Predictive Chemistry — Nature, 2024 · paper
  • Multi‑modal Molecule Structure–Text Model for Text‑based Retrieval and Editing — Nature, 2023 · paper
  • ClimateGPT: Towards AI Synthesizing Interdisciplinary Research on Climate Change — arXiv, 2024 · paper
  • Sparks of Science: Hypothesis Generation Using Structured Paper Data — arXiv, 2025 · paper
  • DeepSeek‑Prover‑V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition — arXiv, 2025 · paper
  • BiMediX: Bilingual Medical Mixture of Experts LLM — arXiv, 2024 · paper
  • ChatDoctor: A Medical Chat Model Fine‑tuned on a Large Language Model (LLAMA) Using Medical Domain Knowledge — arXiv, 2023 · paper
  • AgentMD: Empowering Language Agents for Risk Prediction with Large‑Scale Clinical Tool Learning — arXiv, 2024 · paper
  • MedAlpaca: An Open‑Source Collection of Medical Conversational AI Models and Training Data — arXiv, 2023 · paper
  • DrugGen Enhances Drug Discovery with Large Language Models and Reinforcement Learning — Nature, 2025 · paper
  • LLM‑SR: Scientific Equation Discovery via Programming with Large Language Models — arXiv, 2024 · paper
  • LitLLM: A Toolkit for Scientific Literature Review — arXiv, 2024 · paper
  • SciBERT: A Pretrained Language Model for Scientific Text — arXiv, 2019 · paper
  • SciMON: Scientific Inspiration Machines Optimized for Novelty — arXiv, 2023 · paper
  • SCITUNE: Aligning Large Language Models with Scientific Multimodal Instructions — arXiv, 2023 · paper
  • NatureLM: Deciphering the Language of Nature for Scientific Discovery — arXiv, 2025 · paper

Level 2: Agent As Partner

Level 2: Agent as Partner - multi-step collaboration and workflow orchestration
  • StarWhisper Telescope: Agent‑Based Observation Assistant System to Approach an AI Astrophysicist — arXiv, 2024 · paper
  • From Intention to Implementation: Automating Biomedical Research via LLMs — arXiv, 2024 · paper
  • CRISPR‑GPT: An LLM Agent for Automated Design of Gene‑Editing Experiments — arXiv, 2024 · paper
  • Towards an AI Co‑Scientist — arXiv, 2025 · paper
  • DrBioRight 2.0: An LLM‑Powered Bioinformatics Chatbot for Large‑Scale Cancer Functional Proteomics Analysis — Nature, 2025 · paper
  • ProtAgents: Protein Discovery via Large Language Model Multi‑Agent Collaboration — arXiv, 2024 · paper
  • MOOSE‑Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses — arXiv, 2024 · paper
  • Autonomous Chemical Research with Large Language Models — Nature, 2023 · paper
  • ChemCrow: Augmenting Large‑Language Models with Chemistry Tools — arXiv, 2023 · paper
  • ORGANA: A Robotic Assistant for Automated Chemistry Experimentation and Characterization — arXiv, 2024 · paper
  • xChemAgents: Agentic AI for Explainable Quantum Chemistry — arXiv, 2025 · paper
  • ChemAgent: Self‑Updating Memories in Large Language Models Improves Chemical Reasoning — arXiv, 2025 · paper
  • Large Language Models to Accelerate Organic Chemistry Synthesis — arXiv, 2025 · paper
  • MyCrunchGPT: A ChatGPT‑Assisted Framework for Scientific Machine Learning — arXiv, 2023 · paper
  • MetaOpenFoam: An LLM‑Based Multi‑Agent Framework for CFD — arXiv, 2024 · paper
  • FoamAgent: Towards Automated Intelligent CFD Workflows — arXiv, 2025 · paper
  • AI‑Researcher: Autonomous Scientific Innovation — arXiv, 2025 · paper
  • Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs — arXiv, 2025 · paper
  • FlowAgent: Achieving Compliance and Flexibility for Workflow Agents — arXiv, 2025 · paper
  • The AI Scientist: Towards Fully Automated Open‑Ended Scientific Discovery — arXiv, 2024 · paper
  • GeoGPT: Understanding and Processing Geospatial Tasks through an Autonomous GPT — arXiv, 2023 · paper
  • PaperQA: Retrieval‑Augmented Generative Agent for Scientific Research — arXiv, 2023 · paper
  • ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary — arXiv, 2024 · paper
  • PiFlow: Principle‑Aware Scientific Discovery with Multi‑Agent Collaboration — arXiv, 2025 · paper
  • Aviary: Training Language Agents on Challenging Scientific Tasks — arXiv, 2024 · paper
  • An Autonomous Laboratory for the Accelerated Synthesis of Novel Materials — Nature, 2023 · paper
  • Construction of a Knowledge Graph for Framework Material Enabled by Large Language Models and Its Application — npjCM, 2025 · paper
  • MultiCrossModal Automated Agent for Integrating Diverse Materials Science Data — arXiv, 2025 · paper
  • LLMatDesign: Autonomous Materials Discovery with Large Language Models — arXiv, 2024 · paper
  • Honeycomb: A Flexible LLM‑Based Agent System for Materials Science — arXiv, 2024 · paper
  • DrugAgent: Automating AI‑Aided Drug Discovery Programming through LLM Multi‑Agent Collaboration — arXiv, 2024 · paper
  • Toward a Team of AI‑Made Scientists for Scientific Discovery from Gene Expression Data — arXiv, 2024 · paper
  • MedAgents: Large Language Models as Collaborators for Zero‑Shot Medical Reasoning — arXiv, 2023 · paper
  • Automation of Systematic Reviews with Large Language Models — medRxiv, 2025 · paper
  • MRAgent: An LLM‑Based Automated Agent for Causal Knowledge Discovery in Disease via Mendelian Randomization — BriefBioinf, 2025 · website
  • GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information — arXiv, 2023 · paper
  • SR-Scientist: Scientific Equation Discovery With Agentic AI — ICLR, 2026 · paper · code
  • Fantastic Scientific Agents and How to Build Them: AgentBuild for Rietveld Refinement — arXiv, 2026 · paper
  • SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks — arXiv, 2026 · paper

Level 3: Agent As Avatar

Level 3: Agent as Avatar - autonomous research execution and long-horizon discovery
  • CycleResearcher: Improving Automated Research via Automated Review — arXiv, 2024 · paper
  • The AI Scientist‑v2: Workshop‑Level Automated Scientific Discovery via Agentic Tree Search — arXiv, 2025 · paper
  • AgentRxiv: Towards Collaborative Autonomous Research — arXiv, 2025 · paper
  • Agent Laboratory: Using LLM Agents as Research Assistants — arXiv, 2025 · paper
  • BiOMNI: A General‑Purpose Biomedical AI Agent — bioRxiv, 2025 · paper
  • OriGene: A Self‑Evolving Virtual Disease Biologist Automating Therapeutic Target Discovery — bioRxiv, 2025 · paper
  • CellVoyager: AI CompBio Agent Generates New Insights by Autonomously Analyzing Biological Data — bioRxiv, 2025 · paper
  • Sparks: Multi‑Agent Artificial Intelligence Model Discovers Protein Design Principles — arXiv, 2025 · paper
  • AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery — arXiv, 2025 · paper
  • Robin: A Multi‑Agent System for Automating Scientific Discovery — arXiv, 2025 · paper
  • ScienceOne — Project, 2025 · website
  • EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery — arXiv, 2026 · paper · code
  • Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control — arXiv, 2026 · paper

🧩 Method Guides

The root README keeps the curated map compact; the method guides provide deeper explanations for readers who want to understand how scientific agents are built, improved, evaluated, and connected to broader auto-research systems.

GuideWhat it covers
Construction MethodsKnowledge organization, knowledge injection, tool integration, orchestration, and domain interfaces.
Enhancement MethodsReasoning, memory, collaboration, workflow search, and self-review.
Evaluation MethodsBenchmark selection, executable evaluation, citation grounding, and long-horizon validation.
Auto-Research SystemsEnd-to-end research loops, coding/browser agents, training resources, and orchestration frameworks.

🧭 Broader Agentic & Auto-Research Repositories

Beyond scientific-agent papers, the broader agent ecosystem is moving quickly across model training, software engineering, web automation, and multi-agent orchestration. This watchlist keeps a lightweight bridge from the survey taxonomy to practical repositories that shape how autonomous research and agentic systems are built.

RepositoryScopeWhy follow it
karpathy/nanochatMinimal end-to-end LLM training and chat stackTracks Karpathy's compact, hackable path from tokenizer and pretraining to finetuning, evaluation, inference, and chat UI.
karpathy/llm.cLLM training in C/CUDAUseful for understanding low-level training kernels, performance constraints, and reproducible GPT-style training.
SakanaAI/AI-ScientistAutomated idea-to-paper research loopA reference point for autonomous ideation, coding, experimentation, paper writing, and automated review.
SakanaAI/AI-Scientist-v2Agentic tree search for automated discoveryFollows the next iteration of AI Scientist with broader exploration and stronger end-to-end workflow design.
SamuelSchmidgall/AgentLaboratoryHuman-guided autonomous research assistantShows how literature review, experimentation, and report writing can be composed into a full research workflow.
NoviScl/AI-ResearcherResearch ideation and execution studiesProvides agent pipelines and human-study data for comparing LLM-generated ideas with expert research ideas.
OpenHands/OpenHandsAI software development agentsA generalist coding-agent platform for editing repositories, using terminals, browsing, and operating in sandboxed environments.
SWE-agent/SWE-agentGitHub issue fixing and SWE-bench agentsA practical baseline for agentic software engineering, debugging, and repository-level task execution.
huggingface/smolagentsLightweight code-agent frameworkGood for studying minimal abstractions, code-as-action agents, sandboxed execution, and open-model agent workflows.
microsoft/autogenAgentic AI programming frameworkUseful for multi-agent conversations, orchestration patterns, and prototyping collaborative agent systems.
crewAIInc/crewAIMulti-agent orchestrationFocuses on role-based agents, task delegation, crews, flows, and production-style automation workflows.
langchain-ai/langgraphGraph-based agent workflowsUseful for durable, stateful, controllable agent graphs and long-running workflow orchestration.
FoundationAgents/MetaGPTMulti-agent software company metaphorA representative multi-agent framework for decomposing product/software work into role-specialized agents.
browser-use/browser-useBrowser automation for agentsTracks web interaction patterns, browser control, and task automation over ordinary websites.
Significant-Gravitas/AutoGPTEarly autonomous agent platformStill useful as historical context for goal-directed agents, autonomous task decomposition, and agent productization.
EvoScientist/EvoScientistSelf-evolving end-to-end AI scientistTracks persistent research memory, multi-agent experimentation, and human-on-the-loop research workflows.

✈️ Scientific Agents Construction

Agent construction methodology

Knowledge organization in scientific agents Orchestration and coordination in scientific-agent construction

Construction resources - prompts, context, tools, workflows, and domain interfaces
  • PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization — arXiv, 2023 · paper
  • Context Engineering: A Practical Handbook for Context Design, Orchestration, and Optimization — GitHub, 2025 · code
  • Verl: Volcano Engine Reinforcement Learning for Large Language Models — GitHub, 2024 · code
  • GraphRAG: A Modular Graph-Based Retrieval-Augmented Generation (RAG) System — arXiv, 2023 · paper
  • ChemCrow: Augmenting Large‑Language Models with Chemistry Tools — arXiv, 2023 · paper
  • Virtual Lab: AI Agents Design New SARS-CoV-2 Nanobodies with Experimental Validation — arXiv, 2024 · paper
  • OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning — arXiv, 2025 · paper
  • FoamAgent: Towards Automated Intelligent CFD Workflows — arXiv, 2025 · paper
  • MetaOpenFoam: An LLM-Based Multi-Agent Framework for CFD — arXiv, 2024 · paper
  • AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts — EMNLP, 2020 · paper
  • InstructBio: Instruction Tuning for Biomedical LLMs — arXiv, 2024 · paper
  • LangChain: Building Applications with LLMs through Composability — GitHub, 2022 · code
  • LightRAG: Simple and Fast Retrieval‑Augmented Generation — arXiv, 2024 · paper
  • SciTUNE: Aligning Large Language Models with Scientific Multimodal Instructions — arXiv, 2023 · paper
  • ClimateGPT: Towards AI Synthesizing Interdisciplinary Research on Climate Change — arXiv, 2024 · paper
  • SciMON: Scientific Inspiration Machines Optimized for Novelty — arXiv, 2023 · paper
  • TAIS: Gene Expression Agent with LLMs — arXiv, 2025 · paper
  • StarWhisper Telescope: Agent‑Based Observation Assistant System to Approach an AI Astrophysicist — arXiv, 2024 · paper
  • Fantastic Scientific Agents and How to Build Them: AgentBuild for Rietveld Refinement — arXiv, 2026 · paper

🚀 Scientific Agents Enhancement

Scientific-agent ability enhancement

Scientific agents' memory systems Illustration of scientific agent reasoning enhancement

Enhancement resources - reasoning, memory, workflow search, reflection, and collaboration
  • MemGPT: Toward LLMs as Operating Systems — arXiv, 2023 · paper
  • AFlow: Automating Agentic Workflow Generation — arXiv, 2023 · paper
  • ChemAgent: Self‑Updating Memories in Large Language Models Improves Chemical Reasoning — ICLR, 2025 · paper
  • DeepSeek‑Prover‑V2: Advancing Formal Mathematical Reasoning — arXiv, 2025 · paper
  • LLM‑SR: Scientific Equation Discovery via Programming with LLMs — ICLR, 2025 · paper
  • Sparks: Multi‑Agent Artificial Intelligence Model Discovers Protein Design Principles — arXiv, 2025 · paper
  • ether0: A Scientific Reasoning Model for Chemistry — arXiv, 2025 · paper
  • Robin: A Multi‑Agent System for Automating Scientific Discovery — arXiv, 2025 · paper
  • AgentRxiv: Towards Collaborative Autonomous Research — arXiv, 2025 · paper
  • ChatGPT Research Group for Optimizing the Crystallinity of MOFs and COFs — ACS, 2023 · paper
  • AI Co-Scientist: Towards an AI Co-Scientist — arXiv, 2025 · paper
  • ReAct: Synergizing Reasoning and Acting in Language Models — arXiv, 2022 · paper
  • Reflexion: Language Agents with Verbal Reinforcement Learning — NeurIPS, 2023 · paper
  • Self‑Refine: Iterative Self‑Improvement with Self‑Feedback — arXiv, 2023 · paper
  • Self‑Consistency: Reliable Decoding for Complex Reasoning — ICLR, 2023 · paper
  • AquilaChat: Agent with Long-Context Scratchpad Memory — arXiv, 2024 · paper
  • EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery — arXiv, 2026 · paper
  • SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks — arXiv, 2026 · paper
  • Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control — arXiv, 2026 · paper

⚖️ Benchmark For Scientific Agents

Use this section as an evaluation map rather than a flat benchmark list. The resources below cover different failure modes of scientific agents: domain knowledge, executable experiments, citation grounding, data analysis, and long-horizon discovery.

Scientific-agent benchmark and evaluation landscape

Evaluation angleRepresentative focusUseful when you need to test...
Scientific knowledge and reasoningBioMaze, SuperGPQA, Humanity's Last Exam, MR-BenWhether an agent can reason over expert-level scientific concepts.
Citation and literature groundingCiteBench, ALCE, SurveyForgeWhether outputs are traceable, evidence-aware, and literature-faithful.
Code, data, and experiment executionMLAgentBench, DSBench, SciCode, PaperBenchWhether an agent can implement, run, debug, and reproduce research workflows.
Domain and embodied environmentsDiscoveryWorld, AgentClinic, GenoTEX, LLM-SRBenchWhether an agent performs in domain-specific or simulated scientific settings.
Benchmark resources — evaluation suites for scientific reasoning, data analysis, citation, coding, and agentic discovery
  • BioMaze: Benchmarking and Enhancing Large Language Models for Biological Pathway Reasoning — arXiv, 2025 · paper · dataset
  • BioKGBench: A Knowledge Graph Checking Benchmark of AI Agent for Biomedical Science — GitHub, 2025 · code
  • SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-Dimensional Evaluation for Automated Survey Writing — arXiv, 2025 · paper · code
  • MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark — NeurIPS, 2024 · paper · dataset
  • Humanity's Last Exam: A Hard Benchmark at the Frontier of Human Knowledge — arXiv, 2025 · paper · website
  • SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines — arXiv, 2025 · paper · dataset
  • CiteBench: A Benchmark for Scientific Citation Text Generation — arXiv, 2022 · paper · code
  • ALCE: Enabling Large Language Models to Generate Text With Citations — arXiv, 2023 · paper · code
  • Tomato-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses — arXiv, 2024 · paper · code
  • Reviewer2: Optimizing Review Generation Through Prompt Generation — arXiv, 2024 · paper · code
  • Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning — arXiv, 2025 · dataset · paper
  • MR-Ben: A Comprehensive Meta-Reasoning Benchmark for Large Language Models — arXiv, 2024 · paper · website
  • FigureQA: An Annotated Figure Dataset for Visual Reasoning — arXiv, 2017 · paper · dataset
  • SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research — AAAI, 2024 · paper · dataset
  • MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding — arXiv, 2025 · paper · dataset
  • LLM-SRBench: A New Benchmark for Scientific Equation Discovery With Large Language Models — arXiv, 2025 · paper · dataset
  • GenoTEX: A Benchmark for Evaluating LLM-Based Exploration of Gene Expression Data in Alignment With Bioinformaticians — arXiv, 2024 · paper · dataset
  • MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation — arXiv, 2023 · paper · code
  • PaperBench: Evaluating AI's Ability to Replicate AI Research — arXiv, 2025 · paper
  • DSCodeBench: A Realistic Benchmark for Data Science Code Generation — AAAI, 2026 · paper · code
  • DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents — arXiv, 2024 · paper · code
  • SciCode: A Scientist-Curated Benchmark for Scientific Code Generation — arXiv, 2024 · paper · code
  • AgentClinic: A Multimodal Agent Benchmark to Evaluate AI in Simulated Clinical Environments — arXiv, 2024 · code · paper
  • SDRBench: Scientific Data Reduction Benchmark for Lossy Compressors — Website, 2021 · website · website
  • DSBench: How Far Are Data Science Agents from Becoming Data Science Experts? — GitHub, 2025 · code · paper
  • AISB: AI Scientist Benchmark — NLPCC, 2026 · evaluation kit
  • DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? — arXiv, 2026 · paper · code

🤝 Contributing

Catalog tables and lists are generated from structured files in data/. See CONTRIBUTING.md for the evidence and maintenance workflow, and the catalog schema for field definitions. Edit templates/README.md for narrative content, then regenerate this file with python3 scripts/render_readme.py.

🌞 Citation

@article{wang2025hitchhiker,
  title={The Hitchhiker's Guide to Autonomous Research: A Survey of Scientific Agents},
  author={Wang, Xinming and Xu, Jian and Feng, Aslan H and Chen, Yi and Guo, Haiyang and Zhu, Fei and Shao, Yuanqi and Ren, Minsi and Yi, Hongzhu and Lian, Sheng and others},
  year={2025}
}

Contributors

gudehhh666

27 commits

banjiuyufen

1 commits

Languages

Python

100.0%