handsome-rich/Awesome-Auto-Research-Tools

A curated collection of automated research tools, covering literature search, paper reading, experiment management, and code generation to help researchers accelerate their workflow.

Python

1,216

32 commits

updated Oct 5, 2026

See the code

README

πŸ”¬ Awesome Auto Research Awesome

English | δΈ­ζ–‡

Awesome Auto Research

πŸ€– A curated list of open-source projects that automate scientific research β€” from literature review to idea generation, experiment execution, paper writing, and peer review.

πŸ“… Star counts last verified: 2026-07-25


πŸ“‘ Table of Contents


πŸ§ͺ Autonomous Research Systems

Multi-stage systems that autonomously handle several parts of the research loop, such as hypothesis generation, experimentation, analysis, and manuscript preparation.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
autoresearchCustom (PyTorch, nanochat)External coding agents such as Anthropic Claude Code and OpenAI CodexBy Andrej Karpathy. A minimal single-GPU research harness where an external coding agent repeatedly edits train.py and runs fixed five-minute nanochat experiments under instructions in program.md.
AI-ScientistCustom (templates, LaTeX pipeline)OpenAI, Anthropic Claude, DeepSeek, Gemini, OpenRouter, open-weight modelsThe first comprehensive system for fully automated open-ended scientific discovery. Automates idea generation, coding, experiments, and manuscript writing.
RD-AgentCustom + LiteLLM, Docker, Streamlit, QlibOpenAI (GPT-4o/o1/o3), Azure OpenAI, DeepSeek; any LiteLLM providerMicrosoft. Automates R&D processes β€” factor/model evolution for quant, Kaggle automation, paper-to-code implementation. Top MLE-bench agent.
AutoResearchClawOpenClaw + Docker, LaTeX (NeurIPS/ICML/ICLR), OpenAlex, Semantic ScholarOpenAI (GPT-4o), OpenRouter, DeepSeek, MiniMax; Claude/Gemini/Kimi via ACPAutonomous or human-in-the-loop research: idea β†’ literature retrieval β†’ sandbox experiments β†’ multi-agent peer review β†’ LaTeX paper output, with six configurable intervention modes.
ARISClaude Code + MCP servers (Codex, llm-chat, Zotero, Obsidian)Anthropic Claude, OpenAI GPT, GLM-5, MiniMax, Kimi, Qwen, DeepSeek, LongCat; any OpenAI-compatibleClaude Code skills for autonomous ML research: cross-model review loops, idea discovery, experiment automation, and paper writing.
AI-Scientist-v2Custom (BFTS agentic tree search, AIDE)OpenAI (o1/o3/GPT-4o), Anthropic (Bedrock), GeminiUpgraded version using agentic tree search. Generated the first AI-written workshop paper accepted through peer review.
Agent LaboratoryCustom multi-agent (arXiv, HuggingFace, LaTeX)OpenAI (o1/o3/GPT-4o), DeepSeekEnd-to-end autonomous research workflow with specialized agents for literature review, experimentation, and report writing.
AI-ResearcherCustom + LiteLLM, Docker, GradioAnthropic, OpenAI, Gemini, DeepSeek, OpenRouter, GitHub AI (via LiteLLM)NeurIPS 2025 Spotlight. Fully autonomous system covering literature review, hypothesis generation, algorithm implementation, and manuscript preparation.
claude-scholarClaude Code / Codex CLI / OpenCode, Zotero MCP, Obsidian, LaTeXAnthropic Claude, OpenAI (via Codex)Semi-automated academic research assistant covering ideation β†’ coding β†’ experiments β†’ writing β†’ publication.
EvoScientistLangChain + DeepAgents, Docker (Python 3.11 + Node.js 24)Anthropic Claude, OpenAI, Google Gemini, MiniMax, NVIDIA NIMSelf-evolving AI Scientists. Six-agent team with persistent memory autonomously explores and iteratively improves. Built-in messaging channels (Slack/Discord/Telegram/Feishu/WeChat).
BiomniCustom biomedical agent + code execution, datalake, know-how libraryAnthropic Claude, OpenAI, Azure OpenAI, Gemini, Groq, AWS Bedrock, custom OpenAI-compatible APIsStanford. General-purpose biomedical AI agent that autonomously executes research tasks across biology and medicine, combining LLM reasoning, retrieval, and tool/code use.
DeepScientistCustom (Bayesian optimization, Findings Memory, Research Map), Git worktrees, LaTeXOpenAI (Codex CLI), Anthropic Claude, Moonshot Kimi, OpenCode; local backendsLocal-first autonomous research studio. Findings Memory + Bayesian optimization orchestrate baseline reproduction β†’ branched experiments β†’ LaTeX paper drafts.
DATAGENLangChain + LangGraph, MCP servers, FirecrawlOpenAI, Anthropic Claude, Gemini, Ollama, GroqAI-driven multi-agent research assistant automating hypothesis generation, data analysis, visualization, and report writing.
AutoSciMemory-centric agent framework, persistent knowledge graph, web dashboardClaude Code; preview support for Codex and OpenCodeFull-lifecycle research system with persistent memory across literature review, ideation, experiments, analysis, and paper writing.
NanoResearchNine-stage pipeline, local/SLURM execution, LaTeXOpenAI-compatible APIs; Claude Code and CodexEnd-to-end research pipeline from idea to paper, including real experiment execution on local machines or SLURM clusters.
InternAgentCustom (Aider for codegen, persistent memory), Conda; Google Search, Semantic ScholarOpenAI (incl. OpenAI-compatible), Anthropic ClaudeShanghai AI Lab. Unified agentic framework for long-horizon autonomous discovery across physics, biology, earth, and life sciences β€” reaction yield, molecular dynamics, protein engineering, climate diagnostics.
Idea2PaperAgentAlpha Framework (Multi-Agent), Vector DB, Knowledge Graph (KG)DeepSeek V3/R1, Claude 3.5, GPT-4o; Semantic Scholar, ArXiv APIAdvanced Research Idea Exploration Engine: Orchestrates multi-agent workflows for deep literature mining and KG alignment; Refines raw ideas into novel, structured research proposals.
K-Dense BYOKLocal-first desktop workspace, scientific skills, specialist agents, MCP, OllamaOpenRouter, OpenAI Codex, Anthropic Claude, GitHub Copilot, xAI, OllamaLocal AI co-scientist that searches literature, analyzes real datasets, runs code, creates figures and reports, and records an inspectable living lab notebook.
data-to-paperMulti-agent pipeline, code execution, LaTeX, Semantic ScholarOpenAI API; optional DeepInfraTurns research datasets into transparent, traceable, and verifiable manuscripts, with agents for analysis, interpretation, literature search, and writing.
RobinMulti-agent scientific discovery system, LiteLLM, Edison platformLiteLLM-supported models; Edison access requiredFutureHouse's multi-agent system for scientific discovery, coordinating literature research, data analysis, and experimental planning.

πŸ“š Deep Research & Literature Synthesis

Projects focused on automated information gathering, literature review, and report generation.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
DeerFlowLangChain + LangGraph, InfoQuestAny OpenAI-compatible API (GPT-4, Gemini via OpenRouter, etc.)ByteDance. Open-source SuperAgent harness. Orchestrates sub-agents, memory, and sandboxes for deep research, code generation, and report writing.
STORMDSPy + LiteLLM, StreamlitAll LiteLLM models (OpenAI, Azure, etc.); Search: You.com, Bing, Google, Brave, Tavily, SearXNGStanford. LLM-powered knowledge curation system that generates full-length Wikipedia-like articles with citations. Features Co-STORM.
GPT ResearcherLangGraph, MCP, FastAPI, NextJSOpenAI, Anthropic Claude, Gemini; any OpenAI-compatible APIAutonomous agent for deep web & local research. Generates 5-6 page factual reports with citations in PDF/Docx/Markdown.
Tongyi DeepResearchCustom (ReAct, IterResearch, GRPO RL); Serper, Jina, SandboxFusionOpenAI-compatible, OpenRouter; Tongyi-30B-A3B, Dashscope/BailianAlibaba. Agentic LLM (30.5B params, 3.3B activated) for long-horizon deep information-seeking. SOTA on multiple benchmarks.
Open Deep ResearchLangChain + LangGraph, MCP, LangSmithOpenAI (GPT-5/4.1), Anthropic (Sonnet 4), OpenRouter, Ollama (local)LangChain. Open-source deep research framework with configurable MCP tools and search APIs.
PaperQA2Custom + LiteLLM, Pydantic, tantivyOpenAI, Anthropic, Gemini, Ollama, llama.cpp; any LiteLLM providerHigh-accuracy RAG for scientific documents. Dynamically retrieves full-text papers and iterates on answers. Published at ICLR.
local-deep-researchLangChain + LangGraph, FastAPI, FAISS, SQLCipher, SearXNGOllama, LM Studio, llama.cpp (local); OpenAI, Anthropic, Gemini, OpenRouterLocal-first deep research agent reaching ~95% on SimpleQA with local LLMs. Integrates arXiv/PubMed/Semantic Scholar/Wikipedia and 10+ other sources with encrypted storage.
DeepResearchAgentCustom (Autogenesis self-evolution), MMEngine configsOpenRouter (multi-model access)Skywork. Hierarchical multi-agent system with top-level planning agent coordinating specialized lower-level agents.
Auto-Deep-ResearchAutoAgent Framework + LiteLLM, DockerAnthropic, OpenAI, Gemini, Mistral, Groq, OpenRouter, DeepSeek; any OpenAI-compatibleOpen-source, cost-efficient alternative to OpenAI's Deep Research. Universal LLM compatibility, zero-config launch. Strong GAIA Benchmark results.
OpenScholarCustom RAG (PyTorch, HuggingFace, Contriever)OpenAI (GPT-4o), Llama 3.1 8B (self-hosted); Semantic Scholar API, You.comRetrieval-augmented LM searching 45M open-access papers. Published in Nature. Outperforms PaperQA2 and Perplexity Pro.
OpenResearcherMegatron-LM (training), vLLM (serving), HuggingFace, Tevatron, BM25 + Qwen3-Embedding, SerperOpenResearcher-30B-A3B (open-weight release); OpenAI API (scoring)Fully open training + inference pipeline for long-horizon deep research. Releases 30B-A3B model, surpassing GPT-4.1 and Claude Opus 4 on BrowseComp-Plus.

βš™οΈ Research Implementation & Experimentation

Coding, orchestration, and experiment-optimization systems that implement research ideas and run iterative empirical loops.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
AutoGPTCustom (Agent Builder, workflow blocks), DockerOpenAI, Anthropic, Groq, Llama, AI/ML API (300+ models)One of the earliest autonomous AI agent frameworks. Includes Forge for agent creation, benchmarking suite, and user-friendly UI.
OpenHandsCustom agentic framework, composable Python libAnthropic Claude, OpenAI GPT, MiniMax; any LLMAI-driven software development platform. Autonomous coding agents that edit files, run commands, browse web. 72% on SWE-Bench Verified.
AiderCustom (AI pair-programming CLI), Git integrationAnthropic Claude, OpenAI, DeepSeek, OpenRouter, Ollama; nearly any LLMAI pair programming in your terminal. Supports multi-file edits, git integration. Widely used as the coding backbone in research pipelines.
SWE-agentCustom (YAML-config-driven), purpose-built for researchOpenAI (GPT-4o), Anthropic (Sonnet 4, Claude 3.7); configurablePrinceton. Turns LLMs into software engineering agents that fix real GitHub issues. Pioneered the SWE-Bench benchmark.
DeepCodeMulti-agent code generation, MCP, sandbox executionOpenAI, Anthropic, Gemini, OpenRouter; OpenAI-compatible APIsAgentic coding platform for research and software tasks, including a Paper2Code workflow that turns technical papers into executable implementations.
ClawTeamMulti-agent swarm, Git worktrees, shared task state, isolated sandboxesClaude Code, Codex, OpenClaw, Cursor, Gemini CLI, and other coding agentsSwarm orchestration infrastructure for parallel research experiments; its autoresearch workflow coordinates many isolated agents and GPU runs.
Paper2CodeMulti-agent planning, analysis, and code generation; vLLMOpenAI (o3-mini); open-weight models through vLLMICLR 2026. Converts machine-learning papers into runnable codebases through planning, analysis, and generation agents.
Paper2AgentClaude Code, MCP, paper/codebase analysisAnthropic ClaudeConverts research papers and their codebases into interactive MCP agents that expose methods, data, and workflows as callable tools.
MLE-agentPython, Kaggle integration, arXiv, Papers with CodeOpenAI, Anthropic Claude, Ollama (Llama3), MistralIntelligent companion for ML engineering and research. Integrates with arXiv and Papers with Code for better code/research plans. Auto-debugging.
AIDEPython, Streamlit, DockerOpenAI (GPT-4-turbo/4o), Anthropic Claude, Gemini, Ollama (local)AI-Driven Exploration in the Space of Code. LLM agent that writes, evaluates, and improves ML code via agentic tree search. [paper] 4x more Kaggle medals than best linear agent. Hosted platform: Weco AI.
CORALMulti-agent worktrees, graders, shared state, optional LiteLLMClaude Code, Codex, Cursor; LiteLLM-supported modelsInfrastructure for self-evolving research agents: parallel worktrees, automated grading, shared discoveries, and autoresearch-style experiment loops.

✍️ Academic Writing & Communication

Tools focused on reading, polishing, illustrating, presenting, and responding to reviews for scientific work.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
ChatPaperPyMuPDF, arxiv.py, Flask, DockerOpenAI (GPT-3.5/4)Uses ChatGPT to summarize arXiv papers, provide professional translation, polish manuscripts, analyze reviews, and draft reviewer responses.
PaperBananaStreamlit, OpenRouterOpenAI, Anthropic, Gemini (via OpenRouter)Reference-driven multi-agent framework for automated academic illustration. Five specialized agents produce publication-quality diagrams.
Paper2PosterMulti-agent pipeline, editable PowerPoint output, vLLMGPT-4o; open-weight models through vLLMNeurIPS 2025. Converts a paper PDF into an editable academic poster (.pptx), with both a full pipeline and lightweight coding-agent skill.
ChatReviewerPython, tiktoken, Docker, HuggingFace SpacesOpenAI (GPT-3.5/4)Uses ChatGPT to analyze paper strengths and weaknesses, suggest improvements, and draft reviewer responses. Companion to ChatPaper.

πŸ”§ Research Skills & Plugin Collections

Reusable skill sets and plugin ecosystems that integrate with coding agents (Claude Code, Codex, Gemini CLI, etc.) to enable research workflows.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
scientific-agent-skillsPyTorch Lightning, scikit-learn, BioPython, RDKit, DeepChem, Scanpy, OpenMMAgent-agnostic (Claude Code, Cursor, Codex, Gemini CLI)150 ready-to-use scientific skills across bioinformatics, drug discovery, clinical research, medical imaging, and materials science.
AI-Research-SKILLsDeepSpeed, vLLM, LangChain, W&B, MLflow, and 80+ frameworksAgent-agnostic (Claude Code, Codex, Gemini CLI, Qwen Code)98 skills across 23 categories covering the full AI research lifecycle: literature review, idea generation, experimentation, and paper authoring.
OpenClaw-Medical-SkillsBioPython, GATK, Scanpy, RDKit, DeepChem, OpenMM, AlphaFold, pysam, MDAnalysisClaude-based agents via OpenClaw / NanoClaw frameworks869 medical AI skills spanning clinical reports, genomics, drug discovery, bioinformatics, structural biology, and biomedical databases.

πŸ“‹ Awesome Lists & Surveys

Curated collections and survey papers on the auto-research landscape.

ProjectStarsDescription
awesome-autoresearchCurated index of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch. 50+ entries.
awesome-ai-for-scienceCurated list of AI tools, libraries, papers, datasets, and frameworks for scientific discovery across physics, chemistry, biology, and materials.
Autonomous-AgentsDaily-updated curated collection of research papers on autonomous LLM agents. Covers multi-agent systems, scientific computing, robotics, and more.
Awesome-Deep-ResearchCurated collection of deep research agents β€” industry products, open-source implementations, 70+ recent papers, benchmarks, and 2026 agent systems.

πŸ’‘ How This Differs from General AI Agent Lists

This list focuses specifically on automating the scientific research process β€” not general-purpose AI agents. We include projects that target one or more stages of the research lifecycle:

πŸ“– Literature Review β†’ πŸ’‘ Idea Generation β†’ πŸ” Novelty Check β†’ πŸ“ Experiment Design β†’
πŸ’» Code Implementation β†’ πŸš€ Experiment Execution β†’ πŸ“Š Result Analysis β†’ ✍️ Paper Writing β†’ πŸ“ Peer Review

General-purpose coding agents (OpenHands, Aider, SWE-agent) are included because they serve as critical infrastructure for the experiment execution stage.


🀝 Contributing

PRs welcome! Please ensure the project:

  • Has 500+ GitHub stars
  • Is directly related to automating scientific research
  • Is open-source with an active repository

Please keep entries sorted by star count (descending) within each category.


πŸ“ˆ Star History

Star History Chart


πŸ“„ License

CC0 1.0 Universal

Significant stargazers

Xi Zhang

48 followers Β· starred May 2026

Ziyang Guo

13 followers Β· starred May 2026

Leon Zhang

174 followers Β· starred May 2026

Hao Ren

40 followers Β· starred Jul 2026

handsome-rich/Awesome-Auto-Research-Tools

A curated collection of automated research tools, covering literature search, paper reading, experiment management, and code generation to help researchers accelerate their workflow.

Python

1,216

32 commits

updated Oct 5, 2026

See the code

README

πŸ”¬ Awesome Auto Research Awesome

English | δΈ­ζ–‡

Awesome Auto Research

πŸ€– A curated list of open-source projects that automate scientific research β€” from literature review to idea generation, experiment execution, paper writing, and peer review.

πŸ“… Star counts last verified: 2026-07-25


πŸ“‘ Table of Contents


πŸ§ͺ Autonomous Research Systems

Multi-stage systems that autonomously handle several parts of the research loop, such as hypothesis generation, experimentation, analysis, and manuscript preparation.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
autoresearchCustom (PyTorch, nanochat)External coding agents such as Anthropic Claude Code and OpenAI CodexBy Andrej Karpathy. A minimal single-GPU research harness where an external coding agent repeatedly edits train.py and runs fixed five-minute nanochat experiments under instructions in program.md.
AI-ScientistCustom (templates, LaTeX pipeline)OpenAI, Anthropic Claude, DeepSeek, Gemini, OpenRouter, open-weight modelsThe first comprehensive system for fully automated open-ended scientific discovery. Automates idea generation, coding, experiments, and manuscript writing.
RD-AgentCustom + LiteLLM, Docker, Streamlit, QlibOpenAI (GPT-4o/o1/o3), Azure OpenAI, DeepSeek; any LiteLLM providerMicrosoft. Automates R&D processes β€” factor/model evolution for quant, Kaggle automation, paper-to-code implementation. Top MLE-bench agent.
AutoResearchClawOpenClaw + Docker, LaTeX (NeurIPS/ICML/ICLR), OpenAlex, Semantic ScholarOpenAI (GPT-4o), OpenRouter, DeepSeek, MiniMax; Claude/Gemini/Kimi via ACPAutonomous or human-in-the-loop research: idea β†’ literature retrieval β†’ sandbox experiments β†’ multi-agent peer review β†’ LaTeX paper output, with six configurable intervention modes.
ARISClaude Code + MCP servers (Codex, llm-chat, Zotero, Obsidian)Anthropic Claude, OpenAI GPT, GLM-5, MiniMax, Kimi, Qwen, DeepSeek, LongCat; any OpenAI-compatibleClaude Code skills for autonomous ML research: cross-model review loops, idea discovery, experiment automation, and paper writing.
AI-Scientist-v2Custom (BFTS agentic tree search, AIDE)OpenAI (o1/o3/GPT-4o), Anthropic (Bedrock), GeminiUpgraded version using agentic tree search. Generated the first AI-written workshop paper accepted through peer review.
Agent LaboratoryCustom multi-agent (arXiv, HuggingFace, LaTeX)OpenAI (o1/o3/GPT-4o), DeepSeekEnd-to-end autonomous research workflow with specialized agents for literature review, experimentation, and report writing.
AI-ResearcherCustom + LiteLLM, Docker, GradioAnthropic, OpenAI, Gemini, DeepSeek, OpenRouter, GitHub AI (via LiteLLM)NeurIPS 2025 Spotlight. Fully autonomous system covering literature review, hypothesis generation, algorithm implementation, and manuscript preparation.
claude-scholarClaude Code / Codex CLI / OpenCode, Zotero MCP, Obsidian, LaTeXAnthropic Claude, OpenAI (via Codex)Semi-automated academic research assistant covering ideation β†’ coding β†’ experiments β†’ writing β†’ publication.
EvoScientistLangChain + DeepAgents, Docker (Python 3.11 + Node.js 24)Anthropic Claude, OpenAI, Google Gemini, MiniMax, NVIDIA NIMSelf-evolving AI Scientists. Six-agent team with persistent memory autonomously explores and iteratively improves. Built-in messaging channels (Slack/Discord/Telegram/Feishu/WeChat).
BiomniCustom biomedical agent + code execution, datalake, know-how libraryAnthropic Claude, OpenAI, Azure OpenAI, Gemini, Groq, AWS Bedrock, custom OpenAI-compatible APIsStanford. General-purpose biomedical AI agent that autonomously executes research tasks across biology and medicine, combining LLM reasoning, retrieval, and tool/code use.
DeepScientistCustom (Bayesian optimization, Findings Memory, Research Map), Git worktrees, LaTeXOpenAI (Codex CLI), Anthropic Claude, Moonshot Kimi, OpenCode; local backendsLocal-first autonomous research studio. Findings Memory + Bayesian optimization orchestrate baseline reproduction β†’ branched experiments β†’ LaTeX paper drafts.
DATAGENLangChain + LangGraph, MCP servers, FirecrawlOpenAI, Anthropic Claude, Gemini, Ollama, GroqAI-driven multi-agent research assistant automating hypothesis generation, data analysis, visualization, and report writing.
AutoSciMemory-centric agent framework, persistent knowledge graph, web dashboardClaude Code; preview support for Codex and OpenCodeFull-lifecycle research system with persistent memory across literature review, ideation, experiments, analysis, and paper writing.
NanoResearchNine-stage pipeline, local/SLURM execution, LaTeXOpenAI-compatible APIs; Claude Code and CodexEnd-to-end research pipeline from idea to paper, including real experiment execution on local machines or SLURM clusters.
InternAgentCustom (Aider for codegen, persistent memory), Conda; Google Search, Semantic ScholarOpenAI (incl. OpenAI-compatible), Anthropic ClaudeShanghai AI Lab. Unified agentic framework for long-horizon autonomous discovery across physics, biology, earth, and life sciences β€” reaction yield, molecular dynamics, protein engineering, climate diagnostics.
Idea2PaperAgentAlpha Framework (Multi-Agent), Vector DB, Knowledge Graph (KG)DeepSeek V3/R1, Claude 3.5, GPT-4o; Semantic Scholar, ArXiv APIAdvanced Research Idea Exploration Engine: Orchestrates multi-agent workflows for deep literature mining and KG alignment; Refines raw ideas into novel, structured research proposals.
K-Dense BYOKLocal-first desktop workspace, scientific skills, specialist agents, MCP, OllamaOpenRouter, OpenAI Codex, Anthropic Claude, GitHub Copilot, xAI, OllamaLocal AI co-scientist that searches literature, analyzes real datasets, runs code, creates figures and reports, and records an inspectable living lab notebook.
data-to-paperMulti-agent pipeline, code execution, LaTeX, Semantic ScholarOpenAI API; optional DeepInfraTurns research datasets into transparent, traceable, and verifiable manuscripts, with agents for analysis, interpretation, literature search, and writing.
RobinMulti-agent scientific discovery system, LiteLLM, Edison platformLiteLLM-supported models; Edison access requiredFutureHouse's multi-agent system for scientific discovery, coordinating literature research, data analysis, and experimental planning.

πŸ“š Deep Research & Literature Synthesis

Projects focused on automated information gathering, literature review, and report generation.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
DeerFlowLangChain + LangGraph, InfoQuestAny OpenAI-compatible API (GPT-4, Gemini via OpenRouter, etc.)ByteDance. Open-source SuperAgent harness. Orchestrates sub-agents, memory, and sandboxes for deep research, code generation, and report writing.
STORMDSPy + LiteLLM, StreamlitAll LiteLLM models (OpenAI, Azure, etc.); Search: You.com, Bing, Google, Brave, Tavily, SearXNGStanford. LLM-powered knowledge curation system that generates full-length Wikipedia-like articles with citations. Features Co-STORM.
GPT ResearcherLangGraph, MCP, FastAPI, NextJSOpenAI, Anthropic Claude, Gemini; any OpenAI-compatible APIAutonomous agent for deep web & local research. Generates 5-6 page factual reports with citations in PDF/Docx/Markdown.
Tongyi DeepResearchCustom (ReAct, IterResearch, GRPO RL); Serper, Jina, SandboxFusionOpenAI-compatible, OpenRouter; Tongyi-30B-A3B, Dashscope/BailianAlibaba. Agentic LLM (30.5B params, 3.3B activated) for long-horizon deep information-seeking. SOTA on multiple benchmarks.
Open Deep ResearchLangChain + LangGraph, MCP, LangSmithOpenAI (GPT-5/4.1), Anthropic (Sonnet 4), OpenRouter, Ollama (local)LangChain. Open-source deep research framework with configurable MCP tools and search APIs.
PaperQA2Custom + LiteLLM, Pydantic, tantivyOpenAI, Anthropic, Gemini, Ollama, llama.cpp; any LiteLLM providerHigh-accuracy RAG for scientific documents. Dynamically retrieves full-text papers and iterates on answers. Published at ICLR.
local-deep-researchLangChain + LangGraph, FastAPI, FAISS, SQLCipher, SearXNGOllama, LM Studio, llama.cpp (local); OpenAI, Anthropic, Gemini, OpenRouterLocal-first deep research agent reaching ~95% on SimpleQA with local LLMs. Integrates arXiv/PubMed/Semantic Scholar/Wikipedia and 10+ other sources with encrypted storage.
DeepResearchAgentCustom (Autogenesis self-evolution), MMEngine configsOpenRouter (multi-model access)Skywork. Hierarchical multi-agent system with top-level planning agent coordinating specialized lower-level agents.
Auto-Deep-ResearchAutoAgent Framework + LiteLLM, DockerAnthropic, OpenAI, Gemini, Mistral, Groq, OpenRouter, DeepSeek; any OpenAI-compatibleOpen-source, cost-efficient alternative to OpenAI's Deep Research. Universal LLM compatibility, zero-config launch. Strong GAIA Benchmark results.
OpenScholarCustom RAG (PyTorch, HuggingFace, Contriever)OpenAI (GPT-4o), Llama 3.1 8B (self-hosted); Semantic Scholar API, You.comRetrieval-augmented LM searching 45M open-access papers. Published in Nature. Outperforms PaperQA2 and Perplexity Pro.
OpenResearcherMegatron-LM (training), vLLM (serving), HuggingFace, Tevatron, BM25 + Qwen3-Embedding, SerperOpenResearcher-30B-A3B (open-weight release); OpenAI API (scoring)Fully open training + inference pipeline for long-horizon deep research. Releases 30B-A3B model, surpassing GPT-4.1 and Claude Opus 4 on BrowseComp-Plus.

βš™οΈ Research Implementation & Experimentation

Coding, orchestration, and experiment-optimization systems that implement research ideas and run iterative empirical loops.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
AutoGPTCustom (Agent Builder, workflow blocks), DockerOpenAI, Anthropic, Groq, Llama, AI/ML API (300+ models)One of the earliest autonomous AI agent frameworks. Includes Forge for agent creation, benchmarking suite, and user-friendly UI.
OpenHandsCustom agentic framework, composable Python libAnthropic Claude, OpenAI GPT, MiniMax; any LLMAI-driven software development platform. Autonomous coding agents that edit files, run commands, browse web. 72% on SWE-Bench Verified.
AiderCustom (AI pair-programming CLI), Git integrationAnthropic Claude, OpenAI, DeepSeek, OpenRouter, Ollama; nearly any LLMAI pair programming in your terminal. Supports multi-file edits, git integration. Widely used as the coding backbone in research pipelines.
SWE-agentCustom (YAML-config-driven), purpose-built for researchOpenAI (GPT-4o), Anthropic (Sonnet 4, Claude 3.7); configurablePrinceton. Turns LLMs into software engineering agents that fix real GitHub issues. Pioneered the SWE-Bench benchmark.
DeepCodeMulti-agent code generation, MCP, sandbox executionOpenAI, Anthropic, Gemini, OpenRouter; OpenAI-compatible APIsAgentic coding platform for research and software tasks, including a Paper2Code workflow that turns technical papers into executable implementations.
ClawTeamMulti-agent swarm, Git worktrees, shared task state, isolated sandboxesClaude Code, Codex, OpenClaw, Cursor, Gemini CLI, and other coding agentsSwarm orchestration infrastructure for parallel research experiments; its autoresearch workflow coordinates many isolated agents and GPU runs.
Paper2CodeMulti-agent planning, analysis, and code generation; vLLMOpenAI (o3-mini); open-weight models through vLLMICLR 2026. Converts machine-learning papers into runnable codebases through planning, analysis, and generation agents.
Paper2AgentClaude Code, MCP, paper/codebase analysisAnthropic ClaudeConverts research papers and their codebases into interactive MCP agents that expose methods, data, and workflows as callable tools.
MLE-agentPython, Kaggle integration, arXiv, Papers with CodeOpenAI, Anthropic Claude, Ollama (Llama3), MistralIntelligent companion for ML engineering and research. Integrates with arXiv and Papers with Code for better code/research plans. Auto-debugging.
AIDEPython, Streamlit, DockerOpenAI (GPT-4-turbo/4o), Anthropic Claude, Gemini, Ollama (local)AI-Driven Exploration in the Space of Code. LLM agent that writes, evaluates, and improves ML code via agentic tree search. [paper] 4x more Kaggle medals than best linear agent. Hosted platform: Weco AI.
CORALMulti-agent worktrees, graders, shared state, optional LiteLLMClaude Code, Codex, Cursor; LiteLLM-supported modelsInfrastructure for self-evolving research agents: parallel worktrees, automated grading, shared discoveries, and autoresearch-style experiment loops.

✍️ Academic Writing & Communication

Tools focused on reading, polishing, illustrating, presenting, and responding to reviews for scientific work.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
ChatPaperPyMuPDF, arxiv.py, Flask, DockerOpenAI (GPT-3.5/4)Uses ChatGPT to summarize arXiv papers, provide professional translation, polish manuscripts, analyze reviews, and draft reviewer responses.
PaperBananaStreamlit, OpenRouterOpenAI, Anthropic, Gemini (via OpenRouter)Reference-driven multi-agent framework for automated academic illustration. Five specialized agents produce publication-quality diagrams.
Paper2PosterMulti-agent pipeline, editable PowerPoint output, vLLMGPT-4o; open-weight models through vLLMNeurIPS 2025. Converts a paper PDF into an editable academic poster (.pptx), with both a full pipeline and lightweight coding-agent skill.
ChatReviewerPython, tiktoken, Docker, HuggingFace SpacesOpenAI (GPT-3.5/4)Uses ChatGPT to analyze paper strengths and weaknesses, suggest improvements, and draft reviewer responses. Companion to ChatPaper.

πŸ”§ Research Skills & Plugin Collections

Reusable skill sets and plugin ecosystems that integrate with coding agents (Claude Code, Codex, Gemini CLI, etc.) to enable research workflows.

ProjectStarsFramework / ToolsSupported LLM APIsDescription
scientific-agent-skillsPyTorch Lightning, scikit-learn, BioPython, RDKit, DeepChem, Scanpy, OpenMMAgent-agnostic (Claude Code, Cursor, Codex, Gemini CLI)150 ready-to-use scientific skills across bioinformatics, drug discovery, clinical research, medical imaging, and materials science.
AI-Research-SKILLsDeepSpeed, vLLM, LangChain, W&B, MLflow, and 80+ frameworksAgent-agnostic (Claude Code, Codex, Gemini CLI, Qwen Code)98 skills across 23 categories covering the full AI research lifecycle: literature review, idea generation, experimentation, and paper authoring.
OpenClaw-Medical-SkillsBioPython, GATK, Scanpy, RDKit, DeepChem, OpenMM, AlphaFold, pysam, MDAnalysisClaude-based agents via OpenClaw / NanoClaw frameworks869 medical AI skills spanning clinical reports, genomics, drug discovery, bioinformatics, structural biology, and biomedical databases.

πŸ“‹ Awesome Lists & Surveys

Curated collections and survey papers on the auto-research landscape.

ProjectStarsDescription
awesome-autoresearchCurated index of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch. 50+ entries.
awesome-ai-for-scienceCurated list of AI tools, libraries, papers, datasets, and frameworks for scientific discovery across physics, chemistry, biology, and materials.
Autonomous-AgentsDaily-updated curated collection of research papers on autonomous LLM agents. Covers multi-agent systems, scientific computing, robotics, and more.
Awesome-Deep-ResearchCurated collection of deep research agents β€” industry products, open-source implementations, 70+ recent papers, benchmarks, and 2026 agent systems.

πŸ’‘ How This Differs from General AI Agent Lists

This list focuses specifically on automating the scientific research process β€” not general-purpose AI agents. We include projects that target one or more stages of the research lifecycle:

πŸ“– Literature Review β†’ πŸ’‘ Idea Generation β†’ πŸ” Novelty Check β†’ πŸ“ Experiment Design β†’
πŸ’» Code Implementation β†’ πŸš€ Experiment Execution β†’ πŸ“Š Result Analysis β†’ ✍️ Paper Writing β†’ πŸ“ Peer Review

General-purpose coding agents (OpenHands, Aider, SWE-agent) are included because they serve as critical infrastructure for the experiment execution stage.


🀝 Contributing

PRs welcome! Please ensure the project:

  • Has 500+ GitHub stars
  • Is directly related to automating scientific research
  • Is open-source with an active repository

Please keep entries sorted by star count (descending) within each category.


πŸ“ˆ Star History

Star History Chart


πŸ“„ License

CC0 1.0 Universal

Significant stargazers

Xi Zhang

48 followers Β· starred May 2026

Ziyang Guo

13 followers Β· starred May 2026

Leon Zhang

174 followers Β· starred May 2026

Hao Ren

40 followers Β· starred Jul 2026