Contents
🔥 News
🧰 2026-09 · Open-Source Workbenches section. Platforms you can actually clone and run now have their own home.
🤗 2026-09 · Hugging Face Daily Papers links. Every entry whose paper has a 🤗 Daily Papers page now carries a direct badge, so you can jump straight to the community discussion and upvotes.
🚀 2026-08 · Repository launch. First release covering AI Scientist systems, co-scientists, benchmarks, datasets, and scientific environments.
💡 Ongoing · PRs welcome. Missing something strong? Open a pull request. One excellent entry beats five weak ones.
🤖 End-to-End AI Scientists
Systems that connect several research stages into one loop. Human supervision is still normal, and "autonomous" is not treated as a binary claim.
- OmniScientist, "An Omni-Modal Omni-Discipline AI Scientist".

- The AI Scientist, "Towards Fully Automated Open-Ended Scientific Discovery".

- The AI Scientist-v2, "Workshop-Level Automated Scientific Discovery via Agentic Tree Search".

- Robin, "A multi-agent system for automating scientific discovery".

- data-to-paper, "Autonomous LLM-driven research from data to human-verifiable research papers".

- Agent Laboratory, "Using LLM Agents as Research Assistants".

- AutoResearchClaw, "Self-Reinforcing Autonomous Research with Human-AI Collaboration".

- EvoScientist, "Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery".

- DeepScientist, "Advancing Frontier-Pushing Scientific Findings Progressively".

- Kosmos, "An AI Scientist for Autonomous Discovery".

- Denario, "The Denario project: Deep knowledge AI agents for scientific discovery".

- CycleResearcher, "Improving Automated Research via Automated Review".

- Dolphin, "Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback".

- CORAL, "Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery".

- ASI-Evolve, "AI Accelerates AI".

- Darwin Godel Machine, "Open-Ended Evolution of Self-Improving Agents".

- AlphaEvolve, "A coding agent for scientific and algorithmic discovery".

- FunSearch, "Mathematical discoveries from program search with large language models".

- ShinkaEvolve, "Towards Open-Ended And Sample-Efficient Program Evolution".

- MLE-STAR, "Machine Learning Engineering Agent via Search and Targeted Refinement".

- ML-Master 2.0, "Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering".

- Kolb-Based Experiential Learning (Agent K), "Kolb-Based Experiential Learning for Generalist Agents with Human-Level Kaggle Data Science Performance".

- AutoKaggle, "A Multi-Agent Framework for Autonomous Data Science Competitions".

- DS-Agent, "Automated Data Science by Empowering Large Language Models with Case-Based Reasoning".

- AutoML-Agent, "A Multi-Agent LLM Framework for Full-Pipeline AutoML".

- MLR-Copilot, "Autonomous Machine Learning Research based on Large Language Models Agents".

- AI-Newton, "A Concept-Driven Physical Law Discovery System without Prior Physical Knowledge".

- AtomAgents, "Alloy design and discovery through physics-aware multi-modal multi-agent artificial intelligence".

- MatAgent, "Accelerated Inorganic Materials Design with Generative AI Agents".

- El Agente, "An Autonomous Agent for Quantum Chemistry".

- LabOS, "The AI-XR Co-Scientist That Sees and Works With Humans".

- ORGANA, "A Robotic Assistant for Automated Chemistry Experimentation and Characterization".

- ProtAgents, "Protein discovery via large language model multi-agent collaborations combining physics and machine learning".

- STELLA, "Self-Evolving LLM Agent for Biomedical Research".

- CRISPR-GPT, "CRISPR-GPT for Agentic Automation of Gene-editing Experiments".

🔬 Co-Scientists & Research Agents
Highly relevant to AI Scientist research, but focused on part of the loop or explicitly keeping the human scientist as decision maker.
- Co-Scientist, "Accelerating scientific discovery with Co-Scientist".

- PaperQA2, "Language agents achieve superhuman synthesis of scientific knowledge".

- OpenScholar, "Synthesizing Scientific Literature with Retrieval-augmented LMs".

- Biomni, "A General-Purpose Biomedical AI Agent".

- Coscientist, "Autonomous chemical research with large language models".

- ChemCrow, "Augmenting large-language models with chemistry tools".

- ResearchAgent, "Iterative Research Idea Generation over Scientific Literature with Large Language Models".

- Aviary, "training language agents on challenging scientific tasks".

- SciAgents, "Automating scientific discovery through multi-agent intelligent graph reasoning".

- ToolUniverse, "An open platform for democratizing AI scientists".

- CoI-Agent (Chain of Ideas), "Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents".

- SciMaster / X-Master, "SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?".

- BioDiscoveryAgent, "An AI Agent for Designing Genetic Perturbation Experiments".

- GeneAgent, "Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases".

- TAIS, "Toward a Team of AI-made Scientists for Scientific Discovery from Gene Expression Data".

- GenoMAS, "A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis".

- dZiner, "Rational Inverse Design of Materials with AI Agents".

- ChemAgent, "Self-updating Library in Large Language Models Improves Chemical Reasoning".

- LLMatDesign, "Autonomous Materials Discovery with Large Language Models".

- BioResearcher, Multi-agent pipeline searches literature and datasets, then drafts and reviews dry and wet-lab biomedical experimental protocols.

- AstroAgents, "A Multi-Agent AI for Hypothesis Generation from Mass Spectrometry Data".

- The AI Cosmologist, "The AI Cosmologist I: An Agentic System for Automated Data Analysis".

- Nova, "An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas".

- SciSciGPT, "Advancing Human-AI Collaboration in the Science of Science".

- SciAgent, "Tool-augmented Language Models for Scientific Reasoning".

- Elicit, Screens and extracts structured data from 125M papers, automating systematic-review screening into reusable tables.

- Undermind, Iteratively reads and scores hundreds of papers and follows citation trails until a search is exhaustive.

🧰 Open-Source Workbenches
Platforms you can clone, install, and drive today.
- OmniScientist, Omni-modal, omni-discipline AI scientist you can run locally across heterogeneous scientific evidence.

- Open Science Desktop, Local-first desktop workbench wiring agents, notebooks, runs, figures and review into one auditable provenance trail.

- InternAgent, "When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification".

- RD-Agent, "R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science".

- freephdlabor, "Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation".

- The Virtual Lab, "The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies".

- Curie, "Toward Rigorous and Automated Scientific Experimentation with AI Agents".

- TxAgent, "An AI Agent for Therapeutic Reasoning Across a Universe of Tools".

- Galaxy, "Galaxy for accessible, reproducible, and collaborative data analyses: 2026 update".

- MLE-Agent, Plans and implements ML engineering work with arXiv integration and code retrieval.

- STORM and Co-STORM, "Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models".

- Asta, AI2's science agent family, reproducible and benchmarkable against a rigorous multi-task research suite.

- GPT Researcher, Autonomous deep research over web and local documents with any provider, emitting a cited report.

- DeerFlow, Long-horizon agent harness that researches, writes code, and produces artifacts.

- Local Deep Research, Fully local, encrypted research agent over arXiv, PubMed, and your own private document collection.

- OpenResearcher, "Unleashing AI for Accelerated Scientific Research".

- DeepResearchAgent, Hierarchical planner plus specialist agents for deep research and general task execution.

- DeepLiterature, Open research assistant combining search, code execution, link resolution and information expansion.

- Deep Research from Scratch, LangGraph reference implementation of a deep-research agent, the maintained successor to Open Deep Research.

🔬 Lab automation and self-driving labs
- Opentrons, Write Python protocols and execute them on physical Flex and OT-2 liquid-handling robots.

- PyLabRobot, "An open-source, hardware-agnostic interface for liquid-handling robots and accessories".

- Bluesky, Orchestrates beamline and laboratory experiments plus data acquisition, in production at NSLS-II.

- Atlas, "a brain for self-driving laboratories".

- Olympus, "a benchmarking framework for noisy optimization and experiment planning".

- Self-Driving Lab Demo, Build and run a real low-cost autonomous experimentation rig from dimmable LEDs and a spectrophotometer.

🧩 Libraries, skill packs and components
- Scientific Agent Skills, 165 validated science skills plus database connectors, installable into Claude Code, Cursor or Codex.

- AI4S Skills, Agent skills for topic exploration, literature survey, experiments, paper writing and integrity audit.

- AIDE, "AI-Driven Exploration in the Space of Code".

- TinyScientist, "An Interactive, Extensible, and Controllable Framework for Building Research Agents".

- ResearchTown, "Simulator of Human Research Community".

- Virtual Scientists, Multi-agent simulation of science-of-science dynamics over real publication data.

- Intern-S1, "A Scientific Multimodal Foundation Model".

- Data Formulator, "Data Formulator 2: Iterative Creation of Data Visualizations, with AI Transforming Data Along the Way".

- OpenChemIE, "An Information Extraction Toolkit for Chemistry Literature".

- MolScribe, "Robust Molecular Structure Recognition with Image-to-Graph Generation".

🧱 General agent frameworks these are built on
- MetaGPT and Data Interpreter, "Data Interpreter: An LLM Agent For Data Science".

- CAMEL, "Communicative Agents for "Mind" Exploration of Large Language Model Society".

- OWL, "Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation".

- ChatDev, "Communicative Agents for Software Development".

- smolagents, Minimal library for code-writing agents, shipping the reference Open Deep Research implementation.

🔧 More runnable systems
- autoresearch (karpathy), Gives an agent a real single-GPU nanochat training setup and lets it modify code, train, evaluate, keep or discard.

- ARIS (Auto-Research-In-Sleep), Markdown-only skill pack for autonomous ML research providing cross-model review loops, idea discovery and experiment automation.

- Paper2Agent, "Reimagining Research Papers As Interactive and Reliable AI Agents".

- Paper2Code, "Automating Code Generation from Scientific Papers in Machine Learning".

- claude-scholar, Semi-automated research assistant spanning ideation, coding, experiments, writing and publication across Claude Code, Codex, Kimi and OpenCode.

- Dr. Claw, "An AI Scientist Workspace for Vibe Research".

- ScienceClaw, Self-evolving research colleague with 285 skills across 28 disciplines and persistent memory over literature and databases.

- MLGym, "A New Framework and Benchmark for Advancing AI Research Agents".

- AIRA / aira-dojo, "AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench".

- DeepReview / DeepReviewer, "DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process".

- CMBAgent, "Open Source Planning & Control System with Language Agents for Autonomous Scientific Discovery".

- LLM-SR, "Scientific Equation Discovery via Programming with Large Language Models".

- AI Hilbert, "Evolving scientific discovery by unifying data and background knowledge with AI Hilbert".

- PaSa, "An LLM Agent for Comprehensive Academic Paper Search".

- AutoSurvey, "Large Language Models Can Automatically Write Surveys".

- AutoSota, CLI and leaderboard that autonomously optimizes existing research codebases, publishing results only when an internal ledger confirms improvement.

- ChemGraph, "An Agentic Framework for Computational Chemistry Workflows".

- MDCrow, "Automating Molecular Dynamics Workflows with Large Language Models".

- LLaMP, "Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation".

- HoneyComb, "A Flexible LLM-Based Agent System for Materials Science".

- CellAgent, "An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis".

- ResearchCodeAgent, "An LLM Multi-Agent System for Automated Codification of Research Methodologies".

- EvoMaster, Foundational evolving-agent framework reimplementing the SciMaster line including ML-Master, X-Master and Browse-Master.

- LabClaw, Skills-only operating layer for LabOS, with no engine of its own.

⚠️ License traps worth knowing
AI-Researcher (HKUDS/AI-Researcher, 5.7k stars) ships substantial code but has no license file at all, so all rights are reserved.
Coscientist (gomesgroup/coscientist) carries a Commons Clause rider forbidding sale, which is not OSI open source.
The AI Scientist and v2 use a custom AI Scientist Source Code License derived from the Responsible AI license, carrying use restrictions.
UniScientist (UniPat-AI/UniScientist) has no license file.
🧭 Surveys & Position Papers
- From AI for Science to Agentic Science, "A Survey on Autonomous Scientific Discovery".

- Agentic AI for Scientific Discovery, "A Survey of Progress, Challenges, and Future Directions".

- From Automation to Autonomy, "A Survey on Large Language Models in Scientific Discovery".

- Agent Systems for Academic Research Automation, Survey of research agents across literature, ideation, experimentation, and paper production.

- Exploring the role of LLMs in the scientific method, "Exploring the role of large language models in the scientific method: from hypothesis to discovery".

- Automated Scientific Discovery, "From Equation Discovery to Autonomous Discovery Systems".

- AI4Research, "A Survey of Artificial Intelligence for Scientific Research".

- LLM4SR, "A Survey on Large Language Models for Scientific Research".

- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery

- A Survey of Scientific Large Language Models, "From Data Foundations to Agent Frontiers".

- Towards Scientific Intelligence, "A Survey of LLM-based Scientific Agents".

- Transforming Science with Large Language Models, "A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation".

- From Hypothesis to Publication, "A Comprehensive Survey of AI-Driven Research Support Systems".

- A Survey of AI Scientists

- Autonomous Research Agents, "A Survey of AI Scientists and the Verification Gap".

- Deep Research, "A Survey of Autonomous Research Agents".

- A Comprehensive Survey of Deep Research, "Systems, Methodologies, and Applications".

- A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models

- Large Language Models for Scientific Idea Generation, "A Creativity-Centered Survey".

- A review of large language models and autonomous agents in chemistry

- A Survey on Large Language Model-based Agents for Statistics and Data Science

- Measuring Data Science Automation, "A Survey of Evaluation Tools for AI Assistants and Agents".

- LLM-Based Data Science Agents, "A Survey of Capabilities, Challenges, and Future Directions".

- The rise of self-driving labs in chemical and materials sciences

- Empowering biomedical discovery with AI agents

- Nobel Turing Challenge, "creating the engine for scientific discovery".

- Exploring the use of AI authors and reviewers at Agents4Science

- Evolving Roles of LLMs in Scientific Innovation, "Assistant, Collaborator, Scientist, and Evaluator".

- Position, "AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems".

- Scaling Laws in Scientific Discovery with AI and Robot Scientists

- The Ideation-Execution Gap, "Execution Outcomes of LLM-Generated versus Human Research Ideas".

- Can AI make scientific discoveries?

- The More You Automate, the Less You See, "Hidden Pitfalls of AI Scientist Systems".

- Why LLMs Aren't Scientists Yet, "Lessons from Four Autonomous Research Attempts".

- AI Research Agents Narrow Scientific Exploration

- Can AI agents conduct open-ended AI research? Early evidence from two case studies

- How Far Are AI Scientists from Changing the World?

- Fully Autonomous AI Agents Should Not be Developed

🔭 Research Stages
Work that targets one stage of the research loop rather than the whole thing.
💡 Ideation & hypothesis generation
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

- SciMON, "Scientific Inspiration Machines Optimized for Novelty".

- SciPIP, "An LLM-based Scientific Paper Idea Proposer".

- MOOSE-Chem, "Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses".

- Large Language Models for Automated Open-domain Scientific Hypotheses Discovery

- MOOSE-Chem3, "Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback".

- Many Heads Are Better Than One, "Improved Scientific Idea Generation by A LLM-Based Multi-Agent System".

- IRIS, "Interactive Research Ideation System for Accelerating Scientific Discovery".

- HypoBench, "Towards Systematic and Principled Benchmarking for Hypothesis Generation".

- Literature Meets Data, "A Synergistic Approach to Hypothesis Generation".

- AI Idea Bench 2025, "AI Research Idea Generation Benchmark".

- IdeaBench, "Benchmarking Large Language Models for Research Idea Generation".

- Can Large Language Models Unlock Novel Scientific Research Ideas?

- Sparks of Science, "Hypothesis Generation Using Structured Paper Data".

- Large Language Models as Biomedical Hypothesis Generators, "A Comprehensive Evaluation".

- Spark, "A System for Scientifically Creative Idea Generation".

- LDC: Learning to Generate Research Idea with Dynamic Control

- TrustResearcher, "Automating Knowledge-Grounded and Transparent Research Ideation with Multi-Agent Collaboration".

- Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas

📖 Literature review & retrieval
- SurveyForge, "On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing".

- SurveyX, "Academic Survey Automation via Large Language Models".

- LitSearch, "A Retrieval Benchmark for Scientific Literature Search".

- ArxivDIGESTables, "Synthesizing Scientific Literature into Tables using Language Models".

- LitLLM, "A Toolkit for Scientific Literature Review".

- SPAR, "Scholar Paper Retrieval with LLM-based Agents for Enhanced Academic Search".

- Citegeist, "Automated Generation of Related Work Analysis on the arXiv Corpus".

- SurveyBench, "Can LLM(-Agents) Write Academic Surveys that Align with Reader Needs?".

- SurveyGen, "Quality-Aware Scientific Survey Generation with Large Language Models".

- ScholarGym, "Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research".

- Patience is all you need! An agentic system for performing scientific literature review

- Multi-Turn Agentic Scientific Literature Search via Workflow Induction

🧪 Experimentation & code execution
- AutoMind, "Adaptive Knowledgeable Agent for Automated Data Science".

- A Self-Improving Coding Agent

- CodeScientist, "End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation".

- MLZero, "A Multi-Agent System for End-to-end Machine Learning Automation".

- AutoReproduce, "Automatic AI Experiment Reproduction with Paper Lineage".

- I-MCTS, "Enhancing Agentic AutoML via Introspective Monte Carlo Tree Search".

- MLAgentBench, "Evaluating Language Agents on Machine Learning Experimentation".

- EXP-Bench, "Can AI Conduct AI Research Experiments?".

✍️ Scientific writing
- AutoP2C, "An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers".

- Paper2Poster, "Towards Multimodal Poster Automation from Scientific Papers".

- P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark

- AutoPresent, "Designing Structured Visuals from Scratch".

- Paper2Video, "Automatic Video Generation from Scientific Papers".

- ScholarCopilot, "Training Large Language Models for Academic Writing with Accurate Citations".

- SciLitLLM, "How to Adapt LLMs for Scientific Literature Understanding".

- ScholaWrite, "A Dataset of End-to-End Scholarly Writing Process".

- PaperOrchestra, "A Multi-Agent Framework for Automated AI Research Paper Writing".

- DeTikZify, "Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ".

- AutomaTikZ, "Text-Guided Synthesis of Scientific Vector Graphics with TikZ".

- SciDoc2Diagrammer-MAF, "Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement".

- Attribution in Scientific Literature, "New Benchmark and Methods".

- Do Language Models Know When They're Hallucinating References?

- From Sparse to Dense, "GPT-4 Summarization with Chain of Density Prompting".

🧾 Peer review & verification
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis

- OpenReviewer, "A Specialized Large Language Model for Generating Critical Scientific Paper Reviews".

- AgentReview, "Exploring Peer Review Dynamics with LLM Agents".

- Reviewer2, "Optimizing Review Generation Through Prompt Generation".

- LLMs Assist NLP Researchers, "Critique Paper (Meta-)Reviewing".

- Monitoring AI-Modified Content at Scale, "A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews".

- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review

- BadScientist, "Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?".

- When AI Co-Scientists Fail, "SPOT-a Benchmark for Automated Verification of Scientific Research".

- AI-Assisted Peer Review at Scale, "The AAAI-26 AI Review Pilot".

- Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community

- Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers

- Fact or Fiction, "Verifying Scientific Claims".

- MultiVerS, "Improving scientific claim verification with weak supervision and full-document context".

- SciClaimHunt, "A Large Dataset for Evidence-based Scientific Claim Verification".

🌍 Domains
Systems built for one science rather than for research in general.
⚗️ Chemistry & materials
- A-Lab / AlabOS, "AlabOS: A Python-based Reconfigurable Workflow Management Framework for Autonomous Laboratories".

- GNoME, "Scaling deep learning for materials discovery".

- MatterGen, "a generative model for inorganic materials design".

- MatterSim, "A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures".

- ChemOS, "An orchestration software to democratize autonomous discovery".

- ChemOS 2.0, "An orchestration architecture for chemical self-driving laboratories".

- Chemist-X, "Large Language Model-empowered Agent for Reaction Condition Recommendation in Chemical Synthesis".

- ChemReasoner, "Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback".

- CACTUS, "Chemistry Agent Connecting Tool-Usage to Science".

- SynAsk, "Unleashing the Power of Large Language Models in Organic Synthesis".

- AlphaFlow, "autonomous discovery and optimization of multi-step chemistry using a self-driven fluidic lab guided by reinforcement learning".

- RoboChem, "Automated self-optimization, intensification, and scale-up of photocatalysis in flow".

- LLM-RDF, "An automatic end-to-end chemical synthesis development platform powered by large language models".

- Mobile Robotic Chemist, "A mobile robotic chemist".

- Autonomous Mobile Robots for Exploratory Synthetic Chemistry

- Multiagent-Driven Robotic AI Chemist, "A Multiagent-Driven Robotic AI Chemist Enabling Autonomous Chemical Research On Demand".

- The World Avatar, "A dynamic knowledge graph approach to distributed self-driving laboratories".

🧬 Biology & medicine
- AMIE, "Towards Conversational Diagnostic AI".

- PathChat, "A multimodal generative AI copilot for human pathology".

- SpatialAgent, "An autonomous AI agent for spatial biology".

- DrugAgent, "Automating AI-aided Drug Discovery Programming through LLM Multi-Agent Collaboration".

- BioPlanner, "Automatic Evaluation of LLMs on Protocol Planning in Biology".

- BioMARS, "A Multi-Agent Robotic System for Autonomous Biological Experiments".

- MedRAX, "Medical Reasoning Agent for Chest X-ray".

🔭 Physics, astronomy & earth
- AI Feynman, "a Physics-Inspired Method for Symbolic Regression".

- PySR, "Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl".

- AlphaTensor, "Discovering faster matrix multiplication algorithms with reinforcement learning".

- DeepMind Tokamak Plasma Control, "Magnetic control of tokamak plasmas through deep reinforcement learning".

- Scientific Generative Agent, "LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery".

- KAN, "Kolmogorov-Arnold Networks".

- ClimateGPT, "Towards AI Synthesizing Interdisciplinary Research on Climate Change".

- GeoGalactica, "A Scientific Large Language Model in Geoscience".

➗ Mathematics
- AlphaProof, "Olympiad-level formal mathematical reasoning with reinforcement learning".

- AlphaGeometry, "Solving olympiad geometry without human demonstrations".

- AlphaGeometry2, "Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2".

- LeanDojo, "Theorem Proving with Retrieval-Augmented Language Models".

- Lean Copilot, "Large Language Models as Copilots for Theorem Proving in Lean".

- LeanAgent, "Lifelong Learning for Formal Theorem Proving".

- DeepSeek-Prover, "Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data".

- DeepSeek-Prover-V1.5, "Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search".

- DeepSeek-Prover-V2, "Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition".

- Kimina-Prover, "Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning".

- Seed-Prover, "Solving Formal Math Problems by Decomposition and Iterative Reflection".

- Goedel-Prover, "A Frontier Model for Open-Source Automated Theorem Proving".

- Goedel-Prover-V2, Scaffolded data synthesis plus verifier-guided self-correction, shipping open 8B and 32B theorem-proving weights.

- NuminaMath / AIMO Progress Prize, Winning AI Mathematical Olympiad pipeline plus the 860k chain-of-thought corpus that became the standard open math dataset.

- PutnamBench, "Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition".

- miniF2F, "a cross-system benchmark for formal Olympiad-level mathematics".

- InternLM-Math, "Open Math Large Language Models Toward Verifiable Reasoning".

- Equational Theories Project, Tao-led crowdsourced Lean formalization settling all 22 million implications between 4694 magma equational laws.

- Formal Conjectures, "An Open and Evolving Benchmark for Verified Discovery in Mathematics".

- Aristotle, "IMO-level Automated Theorem Proving".

🏛️ Social science & humanities
🧠 Scientific Foundation Models
Open or landmark models that autonomous discovery systems call as tools.
- AlphaFold 3, "Accurate structure prediction of biomolecular interactions with AlphaFold 3".

- ESM3, "Simulating 500 million years of evolution with a language model".

- Evo 2, "Genome modeling and design across all domains of life with Evo 2".

- Boltz-2, "Towards Accurate and Efficient Binding Affinity Prediction".

- Chai-1, "Decoding the molecular interactions of life".

- Aurora, "A Foundation Model for the Earth System".

- AstroLLaMA, "Towards Specialized Foundation Models in Astronomy".

- Galactica, "A Large Language Model for Science".

- SciBERT, "A Pretrained Language Model for Scientific Text".

- ChemDFM, "Developing ChemDFM as a large language foundation model for chemistry".

- nach0, "Multimodal Natural and Chemical Languages Foundation Model".

- DPLM, "Diffusion Language Models Are Versatile Protein Learners".

- GenomeOcean, Efficient 4B genome foundation model trained on 600 Gbp of metagenomic assemblies for de novo sequence generation.

- scGPT, "toward building a foundation model for single-cell multi-omics using generative AI".

📊 Benchmarks & Evaluation
Evaluation is the bottleneck. These test scientific knowledge, literature grounding, code and experiment execution, whole research workflows, or interaction with scientific environments.
- AstaBench, "Rigorous Benchmarking of AI Agents with a Scientific Research Suite".

- ScienceBoard, "Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows".

- ScienceAgentBench, "Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery".

- PaperBench, "Evaluating AI's Ability to Replicate AI Research".

- CORE-Bench, "Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark".

- SciCode, "A Research Coding Benchmark Curated by Scientists".

- MLE-bench, "Evaluating Machine Learning Agents on Machine Learning Engineering".

- LAB-Bench, "Measuring Capabilities of Language Models for Biology Research".

- AIRS-Bench, "a Suite of Tasks for Frontier AI Research Science Agents".

🤖 Research automation
- RE-Bench, "Evaluating frontier AI R&D capabilities of language model agents against human experts".

- MLRC-Bench, "Can Language Agents Solve Machine Learning Research Challenges?".

- MLR-Bench, "Evaluating AI Agents on Open-Ended Machine Learning Research".

- InnovatorBench, "Evaluating Agents' Ability to Conduct Innovative LLM Research".

- MLE-Dojo, "Interactive Environments for Empowering LLM Agents in Machine Learning Engineering".

- AutoSDT, "Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists".

- The Automated LLM Speedrunning Benchmark, "Reproducing NanoGPT Improvements".

- ResearchGym, "Evaluating Language Model Agents on Real-World AI Research".

- MLS-Bench, "A Holistic and Rigorous Assessment of AI Systems on Building Better AI".

- FML-bench, "Benchmarking Machine Learning Agents for Scientific Research".

- NatureBench, "Can Coding Agents Match the Published SOTA of Nature-Family Papers?".

- FIRE-Bench, "Evaluating AI Agents on the Rediscovery of Scientific Insights".

- SGI-Bench, "Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows".

- DeltaML-Bench, "Evaluating Machine Learning Agents on Real-World Research Repositories".

- AARRI-Bench, "Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle".

- SciAgentArena, "Benchmarking AI Agents for Addressing Scientific Challenges Across Scales".

💻 Scientific coding
- SUPER, "Evaluating Agents on Setting Up and Executing Tasks from Research Repositories".

- DSBench, "How Far Are Data Science Agents from Becoming Data Science Experts?".

- DABStep, "Data Agent Benchmark for Multi-step Reasoning".

- BixBench, "a Comprehensive Benchmark for LLM-b