Omni-Scientist/Awesome-AI-Scientist

🧪 Awesome list of AI Scientist papers, systems, benchmarks, datasets and open-source platforms.

Python

100

19 commits

updated Sep 9, 2026

See the code

README

Awesome AI Scientist: end-to-end AI scientists, co-scientists and runnable workbenches, across the research loop of ideate, discover, experiment, write, verify and knowledge, plus models, data, tools, benchmarks and leaderboards

Awesome AI Scientist Awesome

Papers, systems, benchmarks, datasets and leaderboards for AI that does science.

Website

Entries Papers Systems Workbenches Benchmarks Datasets Daily Papers

🔬 Systems · 🧰 Workbenches you can run · 📊 Benchmarks · 📦 Data & environments · 🎓 Learning resources


Contents


🔥 News

🧰 2026-09 · Open-Source Workbenches section. Platforms you can actually clone and run now have their own home.

🤗 2026-09 · Hugging Face Daily Papers links. Every entry whose paper has a 🤗 Daily Papers page now carries a direct badge, so you can jump straight to the community discussion and upvotes.

🚀 2026-08 · Repository launch. First release covering AI Scientist systems, co-scientists, benchmarks, datasets, and scientific environments.

💡 Ongoing · PRs welcome. Missing something strong? Open a pull request. One excellent entry beats five weak ones.


🤖 End-to-End AI Scientists

Systems that connect several research stages into one loop. Human supervision is still normal, and "autonomous" is not treated as a binary claim.

  • OmniScientist, "An Omni-Modal Omni-Discipline AI Scientist". arXiv Code Daily Papers Website
  • The AI Scientist, "Towards Fully Automated Open-Ended Scientific Discovery". Nature arXiv Code Daily Papers Website
  • The AI Scientist-v2, "Workshop-Level Automated Scientific Discovery via Agentic Tree Search". arXiv Code Daily Papers Website
  • Robin, "A multi-agent system for automating scientific discovery". Nature Code Website
  • data-to-paper, "Autonomous LLM-driven research from data to human-verifiable research papers". NEJM AI arXiv Code
  • Agent Laboratory, "Using LLM Agents as Research Assistants". arXiv Code Daily Papers Website
  • AutoResearchClaw, "Self-Reinforcing Autonomous Research with Human-AI Collaboration". arXiv Code Daily Papers
  • EvoScientist, "Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery". arXiv Code Daily Papers Website
  • DeepScientist, "Advancing Frontier-Pushing Scientific Findings Progressively". arXiv Code Daily Papers Website
  • Kosmos, "An AI Scientist for Autonomous Discovery". arXiv Code Daily Papers Website
  • Denario, "The Denario project: Deep knowledge AI agents for scientific discovery". arXiv Code Daily Papers Website
  • CycleResearcher, "Improving Automated Research via Automated Review". ICLR 2025 arXiv Code Daily Papers Website
  • Dolphin, "Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback". ACL 2025 Main arXiv Code Daily Papers Website
  • CORAL, "Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery". COLM 2026 arXiv Code Daily Papers Website
  • ASI-Evolve, "AI Accelerates AI". arXiv Code Daily Papers
  • Darwin Godel Machine, "Open-Ended Evolution of Self-Improving Agents". ICLR 2026 (Poster) arXiv Code Daily Papers Website
  • AlphaEvolve, "A coding agent for scientific and algorithmic discovery". arXiv Code Daily Papers Website
  • FunSearch, "Mathematical discoveries from program search with large language models". Nature 2024 Paper Code
  • ShinkaEvolve, "Towards Open-Ended And Sample-Efficient Program Evolution". ICLR 2026 arXiv Code Daily Papers Website
  • MLE-STAR, "Machine Learning Engineering Agent via Search and Targeted Refinement". NeurIPS 2025 arXiv Daily Papers Website
  • ML-Master 2.0, "Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering". arXiv Code Daily Papers Website
  • Kolb-Based Experiential Learning (Agent K), "Kolb-Based Experiential Learning for Generalist Agents with Human-Level Kaggle Data Science Performance". arXiv Daily Papers
  • AutoKaggle, "A Multi-Agent Framework for Autonomous Data Science Competitions". arXiv Code Daily Papers
  • DS-Agent, "Automated Data Science by Empowering Large Language Models with Case-Based Reasoning". ICML 2024 arXiv Code
  • AutoML-Agent, "A Multi-Agent LLM Framework for Full-Pipeline AutoML". ICML 2025 arXiv Code Daily Papers Website
  • MLR-Copilot, "Autonomous Machine Learning Research based on Large Language Models Agents". arXiv Code Daily Papers Website
  • AI-Newton, "A Concept-Driven Physical Law Discovery System without Prior Physical Knowledge". arXiv Code
  • AtomAgents, "Alloy design and discovery through physics-aware multi-modal multi-agent artificial intelligence". PNAS 2025 arXiv Code
  • MatAgent, "Accelerated Inorganic Materials Design with Generative AI Agents". arXiv Code
  • El Agente, "An Autonomous Agent for Quantum Chemistry". Matter 2025 arXiv
  • LabOS, "The AI-XR Co-Scientist That Sees and Works With Humans". arXiv Code Website
  • ORGANA, "A Robotic Assistant for Automated Chemistry Experimentation and Characterization". Matter 2025 arXiv Code Website
  • ProtAgents, "Protein discovery via large language model multi-agent collaborations combining physics and machine learning". Digital Discovery 2024 arXiv Code Daily Papers
  • STELLA, "Self-Evolving LLM Agent for Biomedical Research". arXiv Code Daily Papers
  • CRISPR-GPT, "CRISPR-GPT for Agentic Automation of Gene-editing Experiments". Nature Biomedical Engineering 2025 arXiv Code Daily Papers Website

🔬 Co-Scientists & Research Agents

Highly relevant to AI Scientist research, but focused on part of the loop or explicitly keeping the human scientist as decision maker.

  • Co-Scientist, "Accelerating scientific discovery with Co-Scientist". Nature arXiv Daily Papers Website
  • PaperQA2, "Language agents achieve superhuman synthesis of scientific knowledge". arXiv Code Daily Papers Website
  • OpenScholar, "Synthesizing Scientific Literature with Retrieval-augmented LMs". Nature arXiv Code Daily Papers Website
  • Biomni, "A General-Purpose Biomedical AI Agent". bioRxiv Code Website
  • Coscientist, "Autonomous chemical research with large language models". Nature Code
  • ChemCrow, "Augmenting large-language models with chemistry tools". Nat Mach Intell arXiv Code Daily Papers
  • ResearchAgent, "Iterative Research Idea Generation over Scientific Literature with Large Language Models". arXiv Daily Papers
  • Aviary, "training language agents on challenging scientific tasks". arXiv Code Agent Daily Papers
  • SciAgents, "Automating scientific discovery through multi-agent intelligent graph reasoning". Advanced Materials 2024 arXiv Code Daily Papers
  • ToolUniverse, "An open platform for democratizing AI scientists". arXiv Code Daily Papers Website
  • CoI-Agent (Chain of Ideas), "Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents". arXiv Code Daily Papers
  • SciMaster / X-Master, "SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?". arXiv Code Daily Papers
  • BioDiscoveryAgent, "An AI Agent for Designing Genetic Perturbation Experiments". arXiv Code Daily Papers
  • GeneAgent, "Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases". Nature Methods 2025 arXiv Code Daily Papers
  • TAIS, "Toward a Team of AI-made Scientists for Scientific Discovery from Gene Expression Data". arXiv Code Daily Papers Website
  • GenoMAS, "A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis". arXiv Code Daily Papers
  • dZiner, "Rational Inverse Design of Materials with AI Agents". arXiv Code Website
  • ChemAgent, "Self-updating Library in Large Language Models Improves Chemical Reasoning". ICLR 2025 arXiv Code Daily Papers
  • LLMatDesign, "Autonomous Materials Discovery with Large Language Models". arXiv Code
  • BioResearcher, Multi-agent pipeline searches literature and datasets, then drafts and reviews dry and wet-lab biomedical experimental protocols. Science China Information Sciences 2025 Code
  • AstroAgents, "A Multi-Agent AI for Hypothesis Generation from Mass Spectrometry Data". ICLR 2025 Workshop arXiv Code Website
  • The AI Cosmologist, "The AI Cosmologist I: An Agentic System for Automated Data Analysis". arXiv Code Daily Papers
  • Nova, "An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas". arXiv Daily Papers
  • SciSciGPT, "Advancing Human-AI Collaboration in the Science of Science". arXiv Code
  • SciAgent, "Tool-augmented Language Models for Scientific Reasoning". arXiv Daily Papers
  • Elicit, Screens and extracts structured data from 125M papers, automating systematic-review screening into reusable tables. Website
  • Undermind, Iteratively reads and scores hundreds of papers and follows citation trails until a search is exhaustive. whitepaper Paper Website

🧰 Open-Source Workbenches

Platforms you can clone, install, and drive today.

🖥️ End-to-end research platforms

  • OmniScientist, Omni-modal, omni-discipline AI scientist you can run locally across heterogeneous scientific evidence. arXiv Code Daily Papers Website
  • Open Science Desktop, Local-first desktop workbench wiring agents, notebooks, runs, figures and review into one auditable provenance trail. Code Website
  • InternAgent, "When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification". Code arXiv Daily Papers Website
  • RD-Agent, "R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science". Code arXiv Daily Papers Website
  • freephdlabor, "Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation". Code arXiv Daily Papers Website
  • The Virtual Lab, "The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies". Code Nature
  • Curie, "Toward Rigorous and Automated Scientific Experimentation with AI Agents". Code arXiv Daily Papers Website
  • TxAgent, "An AI Agent for Therapeutic Reasoning Across a Universe of Tools". Code arXiv Daily Papers Website
  • Galaxy, "Galaxy for accessible, reproducible, and collaborative data analyses: 2026 update". Code Nucleic_Acids_Res Website
  • MLE-Agent, Plans and implements ML engineering work with arXiv integration and code retrieval. Code

📖 Literature and deep-research platforms

  • STORM and Co-STORM, "Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models". Code arXiv Daily Papers Website
  • Asta, AI2's science agent family, reproducible and benchmarkable against a rigorous multi-task research suite. Agent Bench arXiv Website
  • GPT Researcher, Autonomous deep research over web and local documents with any provider, emitting a cited report. Code Website
  • DeerFlow, Long-horizon agent harness that researches, writes code, and produces artifacts. Code Website
  • Local Deep Research, Fully local, encrypted research agent over arXiv, PubMed, and your own private document collection. Code
  • OpenResearcher, "Unleashing AI for Accelerated Scientific Research". Code arXiv Daily Papers
  • DeepResearchAgent, Hierarchical planner plus specialist agents for deep research and general task execution. Code
  • DeepLiterature, Open research assistant combining search, code execution, link resolution and information expansion. Code
  • Deep Research from Scratch, LangGraph reference implementation of a deep-research agent, the maintained successor to Open Deep Research. Code

🔬 Lab automation and self-driving labs

  • Opentrons, Write Python protocols and execute them on physical Flex and OT-2 liquid-handling robots. Code Website
  • PyLabRobot, "An open-source, hardware-agnostic interface for liquid-handling robots and accessories". Code Device Website
  • Bluesky, Orchestrates beamline and laboratory experiments plus data acquisition, in production at NSLS-II. Code Website
  • Atlas, "a brain for self-driving laboratories". Code Digital_Discovery Website
  • Olympus, "a benchmarking framework for noisy optimization and experiment planning". Code arXiv Website
  • Self-Driving Lab Demo, Build and run a real low-cost autonomous experimentation rig from dimmable LEDs and a spectrophotometer. Code Website

🧩 Libraries, skill packs and components

  • Scientific Agent Skills, 165 validated science skills plus database connectors, installable into Claude Code, Cursor or Codex. Code Website
  • AI4S Skills, Agent skills for topic exploration, literature survey, experiments, paper writing and integrity audit. Code
  • AIDE, "AI-Driven Exploration in the Space of Code". Code arXiv Daily Papers Website
  • TinyScientist, "An Interactive, Extensible, and Controllable Framework for Building Research Agents". Code arXiv Daily Papers
  • ResearchTown, "Simulator of Human Research Community". Code arXiv Daily Papers
  • Virtual Scientists, Multi-agent simulation of science-of-science dynamics over real publication data. Code
  • Intern-S1, "A Scientific Multimodal Foundation Model". Code arXiv Daily Papers
  • Data Formulator, "Data Formulator 2: Iterative Creation of Data Visualizations, with AI Transforming Data Along the Way". Code CHI_2025 Website
  • OpenChemIE, "An Information Extraction Toolkit for Chemistry Literature". Code JCIM
  • MolScribe, "Robust Molecular Structure Recognition with Image-to-Graph Generation". Code JCIM

🧱 General agent frameworks these are built on

  • MetaGPT and Data Interpreter, "Data Interpreter: An LLM Agent For Data Science". Code arXiv Daily Papers
  • CAMEL, "Communicative Agents for "Mind" Exploration of Large Language Model Society". Code arXiv Daily Papers
  • OWL, "Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation". Code arXiv Daily Papers
  • ChatDev, "Communicative Agents for Software Development". Code arXiv Daily Papers
  • smolagents, Minimal library for code-writing agents, shipping the reference Open Deep Research implementation. Code Website

🔧 More runnable systems

  • autoresearch (karpathy), Gives an agent a real single-GPU nanochat training setup and lets it modify code, train, evaluate, keep or discard. Code
  • ARIS (Auto-Research-In-Sleep), Markdown-only skill pack for autonomous ML research providing cross-model review loops, idea discovery and experiment automation. Code
  • Paper2Agent, "Reimagining Research Papers As Interactive and Reliable AI Agents". arXiv Code Daily Papers
  • Paper2Code, "Automating Code Generation from Scientific Papers in Machine Learning". ICLR 2026 arXiv Code Daily Papers
  • claude-scholar, Semi-automated research assistant spanning ideation, coding, experiments, writing and publication across Claude Code, Codex, Kimi and OpenCode. Code
  • Dr. Claw, "An AI Scientist Workspace for Vibe Research". arXiv Code Website
  • ScienceClaw, Self-evolving research colleague with 285 skills across 28 disciplines and persistent memory over literature and databases. Code Website
  • MLGym, "A New Framework and Benchmark for Advancing AI Research Agents". arXiv Code Daily Papers
  • AIRA / aira-dojo, "AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench". NeurIPS 2025 arXiv Code Daily Papers
  • DeepReview / DeepReviewer, "DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process". ACL 2025 Long arXiv Code Daily Papers Website
  • CMBAgent, "Open Source Planning & Control System with Language Agents for Autonomous Scientific Discovery". ICML 2025 Workshop arXiv Code Daily Papers Website
  • LLM-SR, "Scientific Equation Discovery via Programming with Large Language Models". ICLR 2025 (Oral) arXiv Code Daily Papers
  • AI Hilbert, "Evolving scientific discovery by unifying data and background knowledge with AI Hilbert". Nature Communications 2024 Paper Code Website
  • PaSa, "An LLM Agent for Comprehensive Academic Paper Search". arXiv Code Daily Papers
  • AutoSurvey, "Large Language Models Can Automatically Write Surveys". NeurIPS 2024 arXiv Code Daily Papers
  • AutoSota, CLI and leaderboard that autonomously optimizes existing research codebases, publishing results only when an internal ledger confirms improvement. Code Website
  • ChemGraph, "An Agentic Framework for Computational Chemistry Workflows". Communications Chemistry 2026 arXiv Code Website
  • MDCrow, "Automating Molecular Dynamics Workflows with Large Language Models". MLST 2026 arXiv Code
  • LLaMP, "Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation". EMNLP 2025 Main arXiv Code Daily Papers
  • HoneyComb, "A Flexible LLM-Based Agent System for Materials Science". Findings of EMNLP 2024 arXiv Code
  • CellAgent, "An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis". arXiv Code Daily Papers
  • ResearchCodeAgent, "An LLM Multi-Agent System for Automated Codification of Research Methodologies". AI4Research 2025 arXiv Daily Papers
  • EvoMaster, Foundational evolving-agent framework reimplementing the SciMaster line including ML-Master, X-Master and Browse-Master. Code
  • LabClaw, Skills-only operating layer for LabOS, with no engine of its own. Code Website

⚠️ License traps worth knowing

AI-Researcher (HKUDS/AI-Researcher, 5.7k stars) ships substantial code but has no license file at all, so all rights are reserved. Coscientist (gomesgroup/coscientist) carries a Commons Clause rider forbidding sale, which is not OSI open source. The AI Scientist and v2 use a custom AI Scientist Source Code License derived from the Responsible AI license, carrying use restrictions. UniScientist (UniPat-AI/UniScientist) has no license file.

🧭 Surveys & Position Papers


🔭 Research Stages

Work that targets one stage of the research loop rather than the whole thing.

💡 Ideation & hypothesis generation

📖 Literature review & retrieval

🧪 Experimentation & code execution

  • AutoMind, "Adaptive Knowledgeable Agent for Automated Data Science". arXiv Code Daily Papers
  • A Self-Improving Coding Agent arXiv Code Daily Papers
  • CodeScientist, "End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation". arXiv Code
  • MLZero, "A Multi-Agent System for End-to-end Machine Learning Automation". arXiv Code
  • AutoReproduce, "Automatic AI Experiment Reproduction with Paper Lineage". arXiv Code Daily Papers
  • I-MCTS, "Enhancing Agentic AutoML via Introspective Monte Carlo Tree Search". EACL 2026 Findings arXiv Code Daily Papers
  • MLAgentBench, "Evaluating Language Agents on Machine Learning Experimentation". ICML 2024 arXiv Code Daily Papers
  • EXP-Bench, "Can AI Conduct AI Research Experiments?". arXiv Code Daily Papers

✍️ Scientific writing

🧾 Peer review & verification

🧠 Knowledge, tools & environments



🌍 Domains

Systems built for one science rather than for research in general.

⚗️ Chemistry & materials

  • A-Lab / AlabOS, "AlabOS: A Python-based Reconfigurable Workflow Management Framework for Autonomous Laboratories". Nature 2023 arXiv Code Website
  • GNoME, "Scaling deep learning for materials discovery". Nature 2023 Paper Code Website
  • MatterGen, "a generative model for inorganic materials design". Nature 2025 arXiv Code Daily Papers Model Website
  • MatterSim, "A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures". arXiv Code Daily Papers
  • ChemOS, "An orchestration software to democratize autonomous discovery". PLOS ONE 2020 Paper Code
  • ChemOS 2.0, "An orchestration architecture for chemical self-driving laboratories". Matter 2024 Paper Code
  • Chemist-X, "Large Language Model-empowered Agent for Reaction Condition Recommendation in Chemical Synthesis". arXiv Code
  • ChemReasoner, "Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback". ICML 2024 arXiv Code
  • CACTUS, "Chemistry Agent Connecting Tool-Usage to Science". ACS Omega 2024 arXiv Code Daily Papers
  • SynAsk, "Unleashing the Power of Large Language Models in Organic Synthesis". Chemical Science 2025 arXiv
  • AlphaFlow, "autonomous discovery and optimization of multi-step chemistry using a self-driven fluidic lab guided by reinforcement learning". Nature Communications 2023 Paper
  • RoboChem, "Automated self-optimization, intensification, and scale-up of photocatalysis in flow". Science 2024 Paper
  • LLM-RDF, "An automatic end-to-end chemical synthesis development platform powered by large language models". Nature Communications 2024 Paper
  • Mobile Robotic Chemist, "A mobile robotic chemist". Nature 2020 Paper
  • Autonomous Mobile Robots for Exploratory Synthetic Chemistry Nature 2024 Paper
  • Multiagent-Driven Robotic AI Chemist, "A Multiagent-Driven Robotic AI Chemist Enabling Autonomous Chemical Research On Demand". JACS 2025 Paper
  • The World Avatar, "A dynamic knowledge graph approach to distributed self-driving laboratories". Nature Communications 2024 Paper

🧬 Biology & medicine

  • AMIE, "Towards Conversational Diagnostic AI". Nature 2025 arXiv Daily Papers
  • PathChat, "A multimodal generative AI copilot for human pathology". Nature 2024 Paper
  • SpatialAgent, "An autonomous AI agent for spatial biology". bioRxiv 2025 Paper Code
  • DrugAgent, "Automating AI-aided Drug Discovery Programming through LLM Multi-Agent Collaboration". arXiv Code Daily Papers
  • BioPlanner, "Automatic Evaluation of LLMs on Protocol Planning in Biology". EMNLP 2023 arXiv Code
  • BioMARS, "A Multi-Agent Robotic System for Autonomous Biological Experiments". arXiv Daily Papers
  • MedRAX, "Medical Reasoning Agent for Chest X-ray". ICML 2025 arXiv Code Daily Papers

🔭 Physics, astronomy & earth

  • AI Feynman, "a Physics-Inspired Method for Symbolic Regression". Science Advances 2020 arXiv Code
  • PySR, "Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl". arXiv Code Daily Papers Website
  • AlphaTensor, "Discovering faster matrix multiplication algorithms with reinforcement learning". Nature 2022 Paper Code
  • DeepMind Tokamak Plasma Control, "Magnetic control of tokamak plasmas through deep reinforcement learning". Nature 2022 Paper
  • Scientific Generative Agent, "LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery". ICML 2024 arXiv Daily Papers
  • KAN, "Kolmogorov-Arnold Networks". arXiv Code Daily Papers
  • ClimateGPT, "Towards AI Synthesizing Interdisciplinary Research on Climate Change". arXiv Code Daily Papers
  • GeoGalactica, "A Scientific Large Language Model in Geoscience". arXiv Code Daily Papers

➗ Mathematics

  • AlphaProof, "Olympiad-level formal mathematical reasoning with reinforcement learning". Nature 2025 Paper
  • AlphaGeometry, "Solving olympiad geometry without human demonstrations". Nature 2024 Paper Code
  • AlphaGeometry2, "Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2". arXiv Code Daily Papers
  • LeanDojo, "Theorem Proving with Retrieval-Augmented Language Models". NeurIPS 2023 Datasets and Benchmarks arXiv Code Daily Papers Website
  • Lean Copilot, "Large Language Models as Copilots for Theorem Proving in Lean". arXiv Code Daily Papers Model Website
  • LeanAgent, "Lifelong Learning for Formal Theorem Proving". arXiv Code Daily Papers
  • DeepSeek-Prover, "Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data". arXiv Daily Papers
  • DeepSeek-Prover-V1.5, "Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search". ICLR 2025 arXiv Code Daily Papers
  • DeepSeek-Prover-V2, "Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition". arXiv Code Daily Papers Model
  • Kimina-Prover, "Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning". arXiv Code Daily Papers
  • Seed-Prover, "Solving Formal Math Problems by Decomposition and Iterative Reflection". arXiv Code Daily Papers
  • Goedel-Prover, "A Frontier Model for Open-Source Automated Theorem Proving". arXiv Code Daily Papers Model
  • Goedel-Prover-V2, Scaffolded data synthesis plus verifier-guided self-correction, shipping open 8B and 32B theorem-proving weights. Code Model
  • NuminaMath / AIMO Progress Prize, Winning AI Mathematical Olympiad pipeline plus the 860k chain-of-thought corpus that became the standard open math dataset. Code Dataset
  • PutnamBench, "Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition". NeurIPS 2024 Datasets and Benchmarks arXiv Code Daily Papers
  • miniF2F, "a cross-system benchmark for formal Olympiad-level mathematics". ICLR 2022 arXiv Code Daily Papers
  • InternLM-Math, "Open Math Large Language Models Toward Verifiable Reasoning". arXiv Code Daily Papers Model
  • Equational Theories Project, Tao-led crowdsourced Lean formalization settling all 22 million implications between 4694 magma equational laws. Code Website
  • Formal Conjectures, "An Open and Evolving Benchmark for Verified Discovery in Mathematics". arXiv Code Daily Papers Website
  • Aristotle, "IMO-level Automated Theorem Proving". arXiv Daily Papers Website

🏛️ Social science & humanities


🧠 Scientific Foundation Models

Open or landmark models that autonomous discovery systems call as tools.

  • AlphaFold 3, "Accurate structure prediction of biomolecular interactions with AlphaFold 3". Nature 2024 Paper Code Website
  • ESM3, "Simulating 500 million years of evolution with a language model". Science 2025 Paper Code Model Website
  • Evo 2, "Genome modeling and design across all domains of life with Evo 2". bioRxiv 2025 Paper Code Model Website
  • Boltz-2, "Towards Accurate and Efficient Binding Affinity Prediction". bioRxiv 2025 Paper Code Model Website
  • Chai-1, "Decoding the molecular interactions of life". bioRxiv 2024 Paper Code Website
  • Aurora, "A Foundation Model for the Earth System". Nature 2025 arXiv Code Daily Papers Model Website
  • AstroLLaMA, "Towards Specialized Foundation Models in Astronomy". IJCNLP-AACL 2023 workshop arXiv Daily Papers Model
  • Galactica, "A Large Language Model for Science". arXiv Code Daily Papers Model
  • SciBERT, "A Pretrained Language Model for Scientific Text". EMNLP 2019 arXiv Code Daily Papers Model
  • ChemDFM, "Developing ChemDFM as a large language foundation model for chemistry". Cell Rep Phys Sci 2025 arXiv Code Daily Papers Model
  • nach0, "Multimodal Natural and Chemical Languages Foundation Model". Chemical Science 2024 arXiv Code Daily Papers Model
  • DPLM, "Diffusion Language Models Are Versatile Protein Learners". ICML 2024 arXiv Code Daily Papers Model
  • GenomeOcean, Efficient 4B genome foundation model trained on 600 Gbp of metagenomic assemblies for de novo sequence generation. Code Model
  • scGPT, "toward building a foundation model for single-cell multi-omics using generative AI". Nature Methods 2024 Paper Code

📊 Benchmarks & Evaluation

Evaluation is the bottleneck. These test scientific knowledge, literature grounding, code and experiment execution, whole research workflows, or interaction with scientific environments.

  • AstaBench, "Rigorous Benchmarking of AI Agents with a Scientific Research Suite". arXiv Code Daily Papers
  • ScienceBoard, "Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows". arXiv Code Daily Papers Website
  • ScienceAgentBench, "Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery". arXiv Code Daily Papers Website
  • PaperBench, "Evaluating AI's Ability to Replicate AI Research". arXiv Code Daily Papers Website
  • CORE-Bench, "Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark". arXiv Code Daily Papers Leaderboard
  • SciCode, "A Research Coding Benchmark Curated by Scientists". arXiv Code Daily Papers Website
  • MLE-bench, "Evaluating Machine Learning Agents on Machine Learning Engineering". arXiv Code Daily Papers
  • LAB-Bench, "Measuring Capabilities of Language Models for Biology Research". arXiv Code Dataset Daily Papers
  • AIRS-Bench, "a Suite of Tasks for Frontier AI Research Science Agents". arXiv Code Dataset Daily Papers

🤖 Research automation

  • RE-Bench, "Evaluating frontier AI R&D capabilities of language model agents against human experts". arXiv Code Daily Papers Website
  • MLRC-Bench, "Can Language Agents Solve Machine Learning Research Challenges?". NeurIPS 2025 Datasets and Benchmarks arXiv Code Daily Papers Dataset
  • MLR-Bench, "Evaluating AI Agents on Open-Ended Machine Learning Research". NeurIPS 2025 Datasets and Benchmarks arXiv Code Daily Papers Dataset Website
  • InnovatorBench, "Evaluating Agents' Ability to Conduct Innovative LLM Research". arXiv Code Dataset
  • MLE-Dojo, "Interactive Environments for Empowering LLM Agents in Machine Learning Engineering". arXiv Code Daily Papers
  • AutoSDT, "Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists". EMNLP 2025 arXiv Code Daily Papers Dataset
  • The Automated LLM Speedrunning Benchmark, "Reproducing NanoGPT Improvements". arXiv Code Daily Papers
  • ResearchGym, "Evaluating Language Model Agents on Real-World AI Research". ICLR 2026 Workshop arXiv Code Daily Papers
  • MLS-Bench, "A Holistic and Rigorous Assessment of AI Systems on Building Better AI". arXiv Code Daily Papers Dataset Website
  • FML-bench, "Benchmarking Machine Learning Agents for Scientific Research". arXiv Code Daily Papers
  • NatureBench, "Can Coding Agents Match the Published SOTA of Nature-Family Papers?". arXiv Code Daily Papers Dataset Website
  • FIRE-Bench, "Evaluating AI Agents on the Rediscovery of Scientific Insights". ICML 2026 arXiv Daily Papers Website
  • SGI-Bench, "Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows". arXiv Code Daily Papers Website
  • DeltaML-Bench, "Evaluating Machine Learning Agents on Real-World Research Repositories". arXiv Code
  • AARRI-Bench, "Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle". arXiv Code
  • SciAgentArena, "Benchmarking AI Agents for Addressing Scientific Challenges Across Scales". arXiv Code Daily Papers Dataset Website

💻 Scientific coding

  • SUPER, "Evaluating Agents on Setting Up and Executing Tasks from Research Repositories". arXiv Code Daily Papers
  • DSBench, "How Far Are Data Science Agents from Becoming Data Science Experts?". ICLR 2025 arXiv Code Daily Papers Dataset Website
  • DABStep, "Data Agent Benchmark for Multi-step Reasoning". arXiv Daily Papers Dataset Website
  • BixBench, "a Comprehensive Benchmark for LLM-b

Truncated — view the full README on GitHub.

agentic-ai
ai4science
ai-for-science
ai-scientist
automated-science
autonomous-research
awesome
awesome-list
benchmark
co-scientist
deep-research
large-language-models
llm
llm-agents
machine-learning
multi-agent-systems
paper-list
research-agent
scientific-discovery
self-driving-lab

Contributors

unikcc

18 commits

reacher-z

1 commits

Omni-Scientist/Awesome-AI-Scientist

🧪 Awesome list of AI Scientist papers, systems, benchmarks, datasets and open-source platforms.

Python

100

19 commits

updated Sep 9, 2026

See the code

README

Awesome AI Scientist: end-to-end AI scientists, co-scientists and runnable workbenches, across the research loop of ideate, discover, experiment, write, verify and knowledge, plus models, data, tools, benchmarks and leaderboards

Awesome AI Scientist Awesome

Papers, systems, benchmarks, datasets and leaderboards for AI that does science.

Website

Entries Papers Systems Workbenches Benchmarks Datasets Daily Papers

🔬 Systems · 🧰 Workbenches you can run · 📊 Benchmarks · 📦 Data & environments · 🎓 Learning resources


Contents


🔥 News

🧰 2026-09 · Open-Source Workbenches section. Platforms you can actually clone and run now have their own home.

🤗 2026-09 · Hugging Face Daily Papers links. Every entry whose paper has a 🤗 Daily Papers page now carries a direct badge, so you can jump straight to the community discussion and upvotes.

🚀 2026-08 · Repository launch. First release covering AI Scientist systems, co-scientists, benchmarks, datasets, and scientific environments.

💡 Ongoing · PRs welcome. Missing something strong? Open a pull request. One excellent entry beats five weak ones.


🤖 End-to-End AI Scientists

Systems that connect several research stages into one loop. Human supervision is still normal, and "autonomous" is not treated as a binary claim.

  • OmniScientist, "An Omni-Modal Omni-Discipline AI Scientist". arXiv Code Daily Papers Website
  • The AI Scientist, "Towards Fully Automated Open-Ended Scientific Discovery". Nature arXiv Code Daily Papers Website
  • The AI Scientist-v2, "Workshop-Level Automated Scientific Discovery via Agentic Tree Search". arXiv Code Daily Papers Website
  • Robin, "A multi-agent system for automating scientific discovery". Nature Code Website
  • data-to-paper, "Autonomous LLM-driven research from data to human-verifiable research papers". NEJM AI arXiv Code
  • Agent Laboratory, "Using LLM Agents as Research Assistants". arXiv Code Daily Papers Website
  • AutoResearchClaw, "Self-Reinforcing Autonomous Research with Human-AI Collaboration". arXiv Code Daily Papers
  • EvoScientist, "Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery". arXiv Code Daily Papers Website
  • DeepScientist, "Advancing Frontier-Pushing Scientific Findings Progressively". arXiv Code Daily Papers Website
  • Kosmos, "An AI Scientist for Autonomous Discovery". arXiv Code Daily Papers Website
  • Denario, "The Denario project: Deep knowledge AI agents for scientific discovery". arXiv Code Daily Papers Website
  • CycleResearcher, "Improving Automated Research via Automated Review". ICLR 2025 arXiv Code Daily Papers Website
  • Dolphin, "Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback". ACL 2025 Main arXiv Code Daily Papers Website
  • CORAL, "Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery". COLM 2026 arXiv Code Daily Papers Website
  • ASI-Evolve, "AI Accelerates AI". arXiv Code Daily Papers
  • Darwin Godel Machine, "Open-Ended Evolution of Self-Improving Agents". ICLR 2026 (Poster) arXiv Code Daily Papers Website
  • AlphaEvolve, "A coding agent for scientific and algorithmic discovery". arXiv Code Daily Papers Website
  • FunSearch, "Mathematical discoveries from program search with large language models". Nature 2024 Paper Code
  • ShinkaEvolve, "Towards Open-Ended And Sample-Efficient Program Evolution". ICLR 2026 arXiv Code Daily Papers Website
  • MLE-STAR, "Machine Learning Engineering Agent via Search and Targeted Refinement". NeurIPS 2025 arXiv Daily Papers Website
  • ML-Master 2.0, "Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering". arXiv Code Daily Papers Website
  • Kolb-Based Experiential Learning (Agent K), "Kolb-Based Experiential Learning for Generalist Agents with Human-Level Kaggle Data Science Performance". arXiv Daily Papers
  • AutoKaggle, "A Multi-Agent Framework for Autonomous Data Science Competitions". arXiv Code Daily Papers
  • DS-Agent, "Automated Data Science by Empowering Large Language Models with Case-Based Reasoning". ICML 2024 arXiv Code
  • AutoML-Agent, "A Multi-Agent LLM Framework for Full-Pipeline AutoML". ICML 2025 arXiv Code Daily Papers Website
  • MLR-Copilot, "Autonomous Machine Learning Research based on Large Language Models Agents". arXiv Code Daily Papers Website
  • AI-Newton, "A Concept-Driven Physical Law Discovery System without Prior Physical Knowledge". arXiv Code
  • AtomAgents, "Alloy design and discovery through physics-aware multi-modal multi-agent artificial intelligence". PNAS 2025 arXiv Code
  • MatAgent, "Accelerated Inorganic Materials Design with Generative AI Agents". arXiv Code
  • El Agente, "An Autonomous Agent for Quantum Chemistry". Matter 2025 arXiv
  • LabOS, "The AI-XR Co-Scientist That Sees and Works With Humans". arXiv Code Website
  • ORGANA, "A Robotic Assistant for Automated Chemistry Experimentation and Characterization". Matter 2025 arXiv Code Website
  • ProtAgents, "Protein discovery via large language model multi-agent collaborations combining physics and machine learning". Digital Discovery 2024 arXiv Code Daily Papers
  • STELLA, "Self-Evolving LLM Agent for Biomedical Research". arXiv Code Daily Papers
  • CRISPR-GPT, "CRISPR-GPT for Agentic Automation of Gene-editing Experiments". Nature Biomedical Engineering 2025 arXiv Code Daily Papers Website

🔬 Co-Scientists & Research Agents

Highly relevant to AI Scientist research, but focused on part of the loop or explicitly keeping the human scientist as decision maker.

  • Co-Scientist, "Accelerating scientific discovery with Co-Scientist". Nature arXiv Daily Papers Website
  • PaperQA2, "Language agents achieve superhuman synthesis of scientific knowledge". arXiv Code Daily Papers Website
  • OpenScholar, "Synthesizing Scientific Literature with Retrieval-augmented LMs". Nature arXiv Code Daily Papers Website
  • Biomni, "A General-Purpose Biomedical AI Agent". bioRxiv Code Website
  • Coscientist, "Autonomous chemical research with large language models". Nature Code
  • ChemCrow, "Augmenting large-language models with chemistry tools". Nat Mach Intell arXiv Code Daily Papers
  • ResearchAgent, "Iterative Research Idea Generation over Scientific Literature with Large Language Models". arXiv Daily Papers
  • Aviary, "training language agents on challenging scientific tasks". arXiv Code Agent Daily Papers
  • SciAgents, "Automating scientific discovery through multi-agent intelligent graph reasoning". Advanced Materials 2024 arXiv Code Daily Papers
  • ToolUniverse, "An open platform for democratizing AI scientists". arXiv Code Daily Papers Website
  • CoI-Agent (Chain of Ideas), "Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents". arXiv Code Daily Papers
  • SciMaster / X-Master, "SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?". arXiv Code Daily Papers
  • BioDiscoveryAgent, "An AI Agent for Designing Genetic Perturbation Experiments". arXiv Code Daily Papers
  • GeneAgent, "Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases". Nature Methods 2025 arXiv Code Daily Papers
  • TAIS, "Toward a Team of AI-made Scientists for Scientific Discovery from Gene Expression Data". arXiv Code Daily Papers Website
  • GenoMAS, "A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis". arXiv Code Daily Papers
  • dZiner, "Rational Inverse Design of Materials with AI Agents". arXiv Code Website
  • ChemAgent, "Self-updating Library in Large Language Models Improves Chemical Reasoning". ICLR 2025 arXiv Code Daily Papers
  • LLMatDesign, "Autonomous Materials Discovery with Large Language Models". arXiv Code
  • BioResearcher, Multi-agent pipeline searches literature and datasets, then drafts and reviews dry and wet-lab biomedical experimental protocols. Science China Information Sciences 2025 Code
  • AstroAgents, "A Multi-Agent AI for Hypothesis Generation from Mass Spectrometry Data". ICLR 2025 Workshop arXiv Code Website
  • The AI Cosmologist, "The AI Cosmologist I: An Agentic System for Automated Data Analysis". arXiv Code Daily Papers
  • Nova, "An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas". arXiv Daily Papers
  • SciSciGPT, "Advancing Human-AI Collaboration in the Science of Science". arXiv Code
  • SciAgent, "Tool-augmented Language Models for Scientific Reasoning". arXiv Daily Papers
  • Elicit, Screens and extracts structured data from 125M papers, automating systematic-review screening into reusable tables. Website
  • Undermind, Iteratively reads and scores hundreds of papers and follows citation trails until a search is exhaustive. whitepaper Paper Website

🧰 Open-Source Workbenches

Platforms you can clone, install, and drive today.

🖥️ End-to-end research platforms

  • OmniScientist, Omni-modal, omni-discipline AI scientist you can run locally across heterogeneous scientific evidence. arXiv Code Daily Papers Website
  • Open Science Desktop, Local-first desktop workbench wiring agents, notebooks, runs, figures and review into one auditable provenance trail. Code Website
  • InternAgent, "When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification". Code arXiv Daily Papers Website
  • RD-Agent, "R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science". Code arXiv Daily Papers Website
  • freephdlabor, "Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation". Code arXiv Daily Papers Website
  • The Virtual Lab, "The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies". Code Nature
  • Curie, "Toward Rigorous and Automated Scientific Experimentation with AI Agents". Code arXiv Daily Papers Website
  • TxAgent, "An AI Agent for Therapeutic Reasoning Across a Universe of Tools". Code arXiv Daily Papers Website
  • Galaxy, "Galaxy for accessible, reproducible, and collaborative data analyses: 2026 update". Code Nucleic_Acids_Res Website
  • MLE-Agent, Plans and implements ML engineering work with arXiv integration and code retrieval. Code

📖 Literature and deep-research platforms

  • STORM and Co-STORM, "Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models". Code arXiv Daily Papers Website
  • Asta, AI2's science agent family, reproducible and benchmarkable against a rigorous multi-task research suite. Agent Bench arXiv Website
  • GPT Researcher, Autonomous deep research over web and local documents with any provider, emitting a cited report. Code Website
  • DeerFlow, Long-horizon agent harness that researches, writes code, and produces artifacts. Code Website
  • Local Deep Research, Fully local, encrypted research agent over arXiv, PubMed, and your own private document collection. Code
  • OpenResearcher, "Unleashing AI for Accelerated Scientific Research". Code arXiv Daily Papers
  • DeepResearchAgent, Hierarchical planner plus specialist agents for deep research and general task execution. Code
  • DeepLiterature, Open research assistant combining search, code execution, link resolution and information expansion. Code
  • Deep Research from Scratch, LangGraph reference implementation of a deep-research agent, the maintained successor to Open Deep Research. Code

🔬 Lab automation and self-driving labs

  • Opentrons, Write Python protocols and execute them on physical Flex and OT-2 liquid-handling robots. Code Website
  • PyLabRobot, "An open-source, hardware-agnostic interface for liquid-handling robots and accessories". Code Device Website
  • Bluesky, Orchestrates beamline and laboratory experiments plus data acquisition, in production at NSLS-II. Code Website
  • Atlas, "a brain for self-driving laboratories". Code Digital_Discovery Website
  • Olympus, "a benchmarking framework for noisy optimization and experiment planning". Code arXiv Website
  • Self-Driving Lab Demo, Build and run a real low-cost autonomous experimentation rig from dimmable LEDs and a spectrophotometer. Code Website

🧩 Libraries, skill packs and components

  • Scientific Agent Skills, 165 validated science skills plus database connectors, installable into Claude Code, Cursor or Codex. Code Website
  • AI4S Skills, Agent skills for topic exploration, literature survey, experiments, paper writing and integrity audit. Code
  • AIDE, "AI-Driven Exploration in the Space of Code". Code arXiv Daily Papers Website
  • TinyScientist, "An Interactive, Extensible, and Controllable Framework for Building Research Agents". Code arXiv Daily Papers
  • ResearchTown, "Simulator of Human Research Community". Code arXiv Daily Papers
  • Virtual Scientists, Multi-agent simulation of science-of-science dynamics over real publication data. Code
  • Intern-S1, "A Scientific Multimodal Foundation Model". Code arXiv Daily Papers
  • Data Formulator, "Data Formulator 2: Iterative Creation of Data Visualizations, with AI Transforming Data Along the Way". Code CHI_2025 Website
  • OpenChemIE, "An Information Extraction Toolkit for Chemistry Literature". Code JCIM
  • MolScribe, "Robust Molecular Structure Recognition with Image-to-Graph Generation". Code JCIM

🧱 General agent frameworks these are built on

  • MetaGPT and Data Interpreter, "Data Interpreter: An LLM Agent For Data Science". Code arXiv Daily Papers
  • CAMEL, "Communicative Agents for "Mind" Exploration of Large Language Model Society". Code arXiv Daily Papers
  • OWL, "Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation". Code arXiv Daily Papers
  • ChatDev, "Communicative Agents for Software Development". Code arXiv Daily Papers
  • smolagents, Minimal library for code-writing agents, shipping the reference Open Deep Research implementation. Code Website

🔧 More runnable systems

  • autoresearch (karpathy), Gives an agent a real single-GPU nanochat training setup and lets it modify code, train, evaluate, keep or discard. Code
  • ARIS (Auto-Research-In-Sleep), Markdown-only skill pack for autonomous ML research providing cross-model review loops, idea discovery and experiment automation. Code
  • Paper2Agent, "Reimagining Research Papers As Interactive and Reliable AI Agents". arXiv Code Daily Papers
  • Paper2Code, "Automating Code Generation from Scientific Papers in Machine Learning". ICLR 2026 arXiv Code Daily Papers
  • claude-scholar, Semi-automated research assistant spanning ideation, coding, experiments, writing and publication across Claude Code, Codex, Kimi and OpenCode. Code
  • Dr. Claw, "An AI Scientist Workspace for Vibe Research". arXiv Code Website
  • ScienceClaw, Self-evolving research colleague with 285 skills across 28 disciplines and persistent memory over literature and databases. Code Website
  • MLGym, "A New Framework and Benchmark for Advancing AI Research Agents". arXiv Code Daily Papers
  • AIRA / aira-dojo, "AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench". NeurIPS 2025 arXiv Code Daily Papers
  • DeepReview / DeepReviewer, "DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process". ACL 2025 Long arXiv Code Daily Papers Website
  • CMBAgent, "Open Source Planning & Control System with Language Agents for Autonomous Scientific Discovery". ICML 2025 Workshop arXiv Code Daily Papers Website
  • LLM-SR, "Scientific Equation Discovery via Programming with Large Language Models". ICLR 2025 (Oral) arXiv Code Daily Papers
  • AI Hilbert, "Evolving scientific discovery by unifying data and background knowledge with AI Hilbert". Nature Communications 2024 Paper Code Website
  • PaSa, "An LLM Agent for Comprehensive Academic Paper Search". arXiv Code Daily Papers
  • AutoSurvey, "Large Language Models Can Automatically Write Surveys". NeurIPS 2024 arXiv Code Daily Papers
  • AutoSota, CLI and leaderboard that autonomously optimizes existing research codebases, publishing results only when an internal ledger confirms improvement. Code Website
  • ChemGraph, "An Agentic Framework for Computational Chemistry Workflows". Communications Chemistry 2026 arXiv Code Website
  • MDCrow, "Automating Molecular Dynamics Workflows with Large Language Models". MLST 2026 arXiv Code
  • LLaMP, "Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation". EMNLP 2025 Main arXiv Code Daily Papers
  • HoneyComb, "A Flexible LLM-Based Agent System for Materials Science". Findings of EMNLP 2024 arXiv Code
  • CellAgent, "An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis". arXiv Code Daily Papers
  • ResearchCodeAgent, "An LLM Multi-Agent System for Automated Codification of Research Methodologies". AI4Research 2025 arXiv Daily Papers
  • EvoMaster, Foundational evolving-agent framework reimplementing the SciMaster line including ML-Master, X-Master and Browse-Master. Code
  • LabClaw, Skills-only operating layer for LabOS, with no engine of its own. Code Website

⚠️ License traps worth knowing

AI-Researcher (HKUDS/AI-Researcher, 5.7k stars) ships substantial code but has no license file at all, so all rights are reserved. Coscientist (gomesgroup/coscientist) carries a Commons Clause rider forbidding sale, which is not OSI open source. The AI Scientist and v2 use a custom AI Scientist Source Code License derived from the Responsible AI license, carrying use restrictions. UniScientist (UniPat-AI/UniScientist) has no license file.

🧭 Surveys & Position Papers


🔭 Research Stages

Work that targets one stage of the research loop rather than the whole thing.

💡 Ideation & hypothesis generation

📖 Literature review & retrieval

🧪 Experimentation & code execution

  • AutoMind, "Adaptive Knowledgeable Agent for Automated Data Science". arXiv Code Daily Papers
  • A Self-Improving Coding Agent arXiv Code Daily Papers
  • CodeScientist, "End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation". arXiv Code
  • MLZero, "A Multi-Agent System for End-to-end Machine Learning Automation". arXiv Code
  • AutoReproduce, "Automatic AI Experiment Reproduction with Paper Lineage". arXiv Code Daily Papers
  • I-MCTS, "Enhancing Agentic AutoML via Introspective Monte Carlo Tree Search". EACL 2026 Findings arXiv Code Daily Papers
  • MLAgentBench, "Evaluating Language Agents on Machine Learning Experimentation". ICML 2024 arXiv Code Daily Papers
  • EXP-Bench, "Can AI Conduct AI Research Experiments?". arXiv Code Daily Papers

✍️ Scientific writing

🧾 Peer review & verification

🧠 Knowledge, tools & environments



🌍 Domains

Systems built for one science rather than for research in general.

⚗️ Chemistry & materials

  • A-Lab / AlabOS, "AlabOS: A Python-based Reconfigurable Workflow Management Framework for Autonomous Laboratories". Nature 2023 arXiv Code Website
  • GNoME, "Scaling deep learning for materials discovery". Nature 2023 Paper Code Website
  • MatterGen, "a generative model for inorganic materials design". Nature 2025 arXiv Code Daily Papers Model Website
  • MatterSim, "A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures". arXiv Code Daily Papers
  • ChemOS, "An orchestration software to democratize autonomous discovery". PLOS ONE 2020 Paper Code
  • ChemOS 2.0, "An orchestration architecture for chemical self-driving laboratories". Matter 2024 Paper Code
  • Chemist-X, "Large Language Model-empowered Agent for Reaction Condition Recommendation in Chemical Synthesis". arXiv Code
  • ChemReasoner, "Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback". ICML 2024 arXiv Code
  • CACTUS, "Chemistry Agent Connecting Tool-Usage to Science". ACS Omega 2024 arXiv Code Daily Papers
  • SynAsk, "Unleashing the Power of Large Language Models in Organic Synthesis". Chemical Science 2025 arXiv
  • AlphaFlow, "autonomous discovery and optimization of multi-step chemistry using a self-driven fluidic lab guided by reinforcement learning". Nature Communications 2023 Paper
  • RoboChem, "Automated self-optimization, intensification, and scale-up of photocatalysis in flow". Science 2024 Paper
  • LLM-RDF, "An automatic end-to-end chemical synthesis development platform powered by large language models". Nature Communications 2024 Paper
  • Mobile Robotic Chemist, "A mobile robotic chemist". Nature 2020 Paper
  • Autonomous Mobile Robots for Exploratory Synthetic Chemistry Nature 2024 Paper
  • Multiagent-Driven Robotic AI Chemist, "A Multiagent-Driven Robotic AI Chemist Enabling Autonomous Chemical Research On Demand". JACS 2025 Paper
  • The World Avatar, "A dynamic knowledge graph approach to distributed self-driving laboratories". Nature Communications 2024 Paper

🧬 Biology & medicine

  • AMIE, "Towards Conversational Diagnostic AI". Nature 2025 arXiv Daily Papers
  • PathChat, "A multimodal generative AI copilot for human pathology". Nature 2024 Paper
  • SpatialAgent, "An autonomous AI agent for spatial biology". bioRxiv 2025 Paper Code
  • DrugAgent, "Automating AI-aided Drug Discovery Programming through LLM Multi-Agent Collaboration". arXiv Code Daily Papers
  • BioPlanner, "Automatic Evaluation of LLMs on Protocol Planning in Biology". EMNLP 2023 arXiv Code
  • BioMARS, "A Multi-Agent Robotic System for Autonomous Biological Experiments". arXiv Daily Papers
  • MedRAX, "Medical Reasoning Agent for Chest X-ray". ICML 2025 arXiv Code Daily Papers

🔭 Physics, astronomy & earth

  • AI Feynman, "a Physics-Inspired Method for Symbolic Regression". Science Advances 2020 arXiv Code
  • PySR, "Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl". arXiv Code Daily Papers Website
  • AlphaTensor, "Discovering faster matrix multiplication algorithms with reinforcement learning". Nature 2022 Paper Code
  • DeepMind Tokamak Plasma Control, "Magnetic control of tokamak plasmas through deep reinforcement learning". Nature 2022 Paper
  • Scientific Generative Agent, "LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery". ICML 2024 arXiv Daily Papers
  • KAN, "Kolmogorov-Arnold Networks". arXiv Code Daily Papers
  • ClimateGPT, "Towards AI Synthesizing Interdisciplinary Research on Climate Change". arXiv Code Daily Papers
  • GeoGalactica, "A Scientific Large Language Model in Geoscience". arXiv Code Daily Papers

➗ Mathematics

  • AlphaProof, "Olympiad-level formal mathematical reasoning with reinforcement learning". Nature 2025 Paper
  • AlphaGeometry, "Solving olympiad geometry without human demonstrations". Nature 2024 Paper Code
  • AlphaGeometry2, "Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2". arXiv Code Daily Papers
  • LeanDojo, "Theorem Proving with Retrieval-Augmented Language Models". NeurIPS 2023 Datasets and Benchmarks arXiv Code Daily Papers Website
  • Lean Copilot, "Large Language Models as Copilots for Theorem Proving in Lean". arXiv Code Daily Papers Model Website
  • LeanAgent, "Lifelong Learning for Formal Theorem Proving". arXiv Code Daily Papers
  • DeepSeek-Prover, "Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data". arXiv Daily Papers
  • DeepSeek-Prover-V1.5, "Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search". ICLR 2025 arXiv Code Daily Papers
  • DeepSeek-Prover-V2, "Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition". arXiv Code Daily Papers Model
  • Kimina-Prover, "Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning". arXiv Code Daily Papers
  • Seed-Prover, "Solving Formal Math Problems by Decomposition and Iterative Reflection". arXiv Code Daily Papers
  • Goedel-Prover, "A Frontier Model for Open-Source Automated Theorem Proving". arXiv Code Daily Papers Model
  • Goedel-Prover-V2, Scaffolded data synthesis plus verifier-guided self-correction, shipping open 8B and 32B theorem-proving weights. Code Model
  • NuminaMath / AIMO Progress Prize, Winning AI Mathematical Olympiad pipeline plus the 860k chain-of-thought corpus that became the standard open math dataset. Code Dataset
  • PutnamBench, "Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition". NeurIPS 2024 Datasets and Benchmarks arXiv Code Daily Papers
  • miniF2F, "a cross-system benchmark for formal Olympiad-level mathematics". ICLR 2022 arXiv Code Daily Papers
  • InternLM-Math, "Open Math Large Language Models Toward Verifiable Reasoning". arXiv Code Daily Papers Model
  • Equational Theories Project, Tao-led crowdsourced Lean formalization settling all 22 million implications between 4694 magma equational laws. Code Website
  • Formal Conjectures, "An Open and Evolving Benchmark for Verified Discovery in Mathematics". arXiv Code Daily Papers Website
  • Aristotle, "IMO-level Automated Theorem Proving". arXiv Daily Papers Website

🏛️ Social science & humanities


🧠 Scientific Foundation Models

Open or landmark models that autonomous discovery systems call as tools.

  • AlphaFold 3, "Accurate structure prediction of biomolecular interactions with AlphaFold 3". Nature 2024 Paper Code Website
  • ESM3, "Simulating 500 million years of evolution with a language model". Science 2025 Paper Code Model Website
  • Evo 2, "Genome modeling and design across all domains of life with Evo 2". bioRxiv 2025 Paper Code Model Website
  • Boltz-2, "Towards Accurate and Efficient Binding Affinity Prediction". bioRxiv 2025 Paper Code Model Website
  • Chai-1, "Decoding the molecular interactions of life". bioRxiv 2024 Paper Code Website
  • Aurora, "A Foundation Model for the Earth System". Nature 2025 arXiv Code Daily Papers Model Website
  • AstroLLaMA, "Towards Specialized Foundation Models in Astronomy". IJCNLP-AACL 2023 workshop arXiv Daily Papers Model
  • Galactica, "A Large Language Model for Science". arXiv Code Daily Papers Model
  • SciBERT, "A Pretrained Language Model for Scientific Text". EMNLP 2019 arXiv Code Daily Papers Model
  • ChemDFM, "Developing ChemDFM as a large language foundation model for chemistry". Cell Rep Phys Sci 2025 arXiv Code Daily Papers Model
  • nach0, "Multimodal Natural and Chemical Languages Foundation Model". Chemical Science 2024 arXiv Code Daily Papers Model
  • DPLM, "Diffusion Language Models Are Versatile Protein Learners". ICML 2024 arXiv Code Daily Papers Model
  • GenomeOcean, Efficient 4B genome foundation model trained on 600 Gbp of metagenomic assemblies for de novo sequence generation. Code Model
  • scGPT, "toward building a foundation model for single-cell multi-omics using generative AI". Nature Methods 2024 Paper Code

📊 Benchmarks & Evaluation

Evaluation is the bottleneck. These test scientific knowledge, literature grounding, code and experiment execution, whole research workflows, or interaction with scientific environments.

  • AstaBench, "Rigorous Benchmarking of AI Agents with a Scientific Research Suite". arXiv Code Daily Papers
  • ScienceBoard, "Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows". arXiv Code Daily Papers Website
  • ScienceAgentBench, "Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery". arXiv Code Daily Papers Website
  • PaperBench, "Evaluating AI's Ability to Replicate AI Research". arXiv Code Daily Papers Website
  • CORE-Bench, "Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark". arXiv Code Daily Papers Leaderboard
  • SciCode, "A Research Coding Benchmark Curated by Scientists". arXiv Code Daily Papers Website
  • MLE-bench, "Evaluating Machine Learning Agents on Machine Learning Engineering". arXiv Code Daily Papers
  • LAB-Bench, "Measuring Capabilities of Language Models for Biology Research". arXiv Code Dataset Daily Papers
  • AIRS-Bench, "a Suite of Tasks for Frontier AI Research Science Agents". arXiv Code Dataset Daily Papers

🤖 Research automation

  • RE-Bench, "Evaluating frontier AI R&D capabilities of language model agents against human experts". arXiv Code Daily Papers Website
  • MLRC-Bench, "Can Language Agents Solve Machine Learning Research Challenges?". NeurIPS 2025 Datasets and Benchmarks arXiv Code Daily Papers Dataset
  • MLR-Bench, "Evaluating AI Agents on Open-Ended Machine Learning Research". NeurIPS 2025 Datasets and Benchmarks arXiv Code Daily Papers Dataset Website
  • InnovatorBench, "Evaluating Agents' Ability to Conduct Innovative LLM Research". arXiv Code Dataset
  • MLE-Dojo, "Interactive Environments for Empowering LLM Agents in Machine Learning Engineering". arXiv Code Daily Papers
  • AutoSDT, "Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists". EMNLP 2025 arXiv Code Daily Papers Dataset
  • The Automated LLM Speedrunning Benchmark, "Reproducing NanoGPT Improvements". arXiv Code Daily Papers
  • ResearchGym, "Evaluating Language Model Agents on Real-World AI Research". ICLR 2026 Workshop arXiv Code Daily Papers
  • MLS-Bench, "A Holistic and Rigorous Assessment of AI Systems on Building Better AI". arXiv Code Daily Papers Dataset Website
  • FML-bench, "Benchmarking Machine Learning Agents for Scientific Research". arXiv Code Daily Papers
  • NatureBench, "Can Coding Agents Match the Published SOTA of Nature-Family Papers?". arXiv Code Daily Papers Dataset Website
  • FIRE-Bench, "Evaluating AI Agents on the Rediscovery of Scientific Insights". ICML 2026 arXiv Daily Papers Website
  • SGI-Bench, "Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows". arXiv Code Daily Papers Website
  • DeltaML-Bench, "Evaluating Machine Learning Agents on Real-World Research Repositories". arXiv Code
  • AARRI-Bench, "Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle". arXiv Code
  • SciAgentArena, "Benchmarking AI Agents for Addressing Scientific Challenges Across Scales". arXiv Code Daily Papers Dataset Website

💻 Scientific coding

  • SUPER, "Evaluating Agents on Setting Up and Executing Tasks from Research Repositories". arXiv Code Daily Papers
  • DSBench, "How Far Are Data Science Agents from Becoming Data Science Experts?". ICLR 2025 arXiv Code Daily Papers Dataset Website
  • DABStep, "Data Agent Benchmark for Multi-step Reasoning". arXiv Daily Papers Dataset Website
  • BixBench, "a Comprehensive Benchmark for LLM-b

Truncated — view the full README on GitHub.

agentic-ai
ai4science
ai-for-science
ai-scientist
automated-science
autonomous-research
awesome
awesome-list
benchmark
co-scientist
deep-research
large-language-models
llm
llm-agents
machine-learning
multi-agent-systems
paper-list
research-agent
scientific-discovery
self-driving-lab

Contributors

unikcc

18 commits

reacher-z

1 commits

Languages

Python

98.3%

JavaScript

1.7%