SpectrAI-Initiative/Vibe-Research-Guide

A curated guide for LLM-agent-driven scientific research automation — from getting started to the frontier.

JavaScript

36

27 commits

updated Apr 21, 2026

See the code

README

Vibe Research Guide

A curated guide for LLM-agent-driven scientific research automation

Stars Last Commit Issues MIT

🌐 View Interactive Multilingual README →
🇨🇳 中文 · 🇺🇸 English · 🇰🇷 한국어 · 🇯🇵 日本語 · 🇩🇪 Deutsch · 🇫🇷 Français · 🇪🇸 Español · 🇮🇹 Italiano · 🇵🇹 Português · 🇸🇦 العربية · 🇹🇭 ไทย · 🇻🇳 Tiếng Việt · 🇷🇺 Русский

Vibe Research Guide Overview


Vibe Research At The Center, With The Broader Agent-Native Stack Around It

Automate the research loop with LLM agents: literature review → idea generation → experiment execution → paper writing → peer review.

This repo is a research-first landing page for the field. The center is still Vibe Research, but the guide now gives Auto Research / AI Scientist its own top-level map and then places Claw, coding agents, connectors, and adjacent assistant ecosystems around that core.

Start here: Getting Started · Auto Research · Tools & Platforms · Claw Park

Vibe Research: AI assistant workflow (idea → literature → experiment → code → result → paper)

At A Glance

Core Question

How far can AI move from research assistant to research operator?

Focus: literature, ideation, experiment, writing, and evaluation.
What Changed In 2026

Research copilots got stronger, learning layers became real, autonomous research systems got more credible, and Vibe Coding became the execution layer.
How To Use This Repo

Treat the README as a map. Treat the topic pages as the actual guide.

2026 Landscape Snapshot

Five shifts now define the field:

  1. Research copilots are stronger and easier to trust: Deep Research, NotebookLM-style source-grounded reading, and scientific workspaces such as Prism are making synthesis and report-writing meaningfully faster.
  2. Auto Research is becoming a recognizable system category: The AI Scientist, The AI Scientist-v2, Agent Laboratory, AI-Researcher, RD-Agent, Auto-Deep-Research, and EvoScientist now form a real system family.
  3. Benchmarks are getting closer to scientific reality: ScienceAgentBench, FIRE-Bench, ResearchClawBench, SGI-Bench, and RE-Bench shift evaluation away from vague "agent capability" claims.
  4. Execution substrates now determine whether research agents actually work: Claude Code, Codex, OpenHands, SWE-agent, OpenClaw, and chat bridges such as cc-connect increasingly matter because experiment loops fail on tooling long before they fail on prose.
  5. A wider agent-native ecosystem is forming around research: personal assistants, registries, skills, self-evolving stacks, and companion UX are increasingly part of the same operational environment.

2026 Auto Research Signals

Several current signals make the field feel less like a loose collection of demos and more like an emerging research stack:

  1. The field has moved from "can an agent summarize papers?" to "can an agent operate a research loop?"
  2. System families are diverging clearly: end-to-end scientists, human-in-the-loop research copilots, R&D execution agents, deep-research assistants, and self-evolving scientist stacks now look meaningfully different.
  3. Benchmarking is improving fast: scientist-aligned workflows, rediscovery tasks, expert comparison, and checklist-based evaluation are becoming standard expectations instead of afterthoughts.
  4. Frameworks matter as much as flagship demos: AI-Researcher, RD-Agent, and Auto-Deep-Research show that orchestration and harness design are now first-class concerns.
  5. Platformization is visible: FutureHouse Platform, Robin, BixBench, and Edison Scientific / Kosmos show how AI-scientist ideas are moving from one-off papers to persistent public or commercial surfaces.

Choose a Path

🟢 New to Vibe Research

Start: Getting Started
Then: Auto Research · Tools & Platforms
🔵 Developer / Builder

Start: Auto Research
Then: Tools & Platforms · Vibe Coding · Systems
🔴 Researcher

Start: Surveys
Then: Auto Research · Benchmarks · Ideation
🟣 Creator / Operator

Start: Tools & Platforms
Then: Vibe Coding · Vibe Anything

Only have 5 minutes? Install InnoClaw and try it out.


Auto Research Stack

This is the shortest useful way to read the field in 2026:

LayerRepresentative resourcesWhy it matters
Research copilotsDeep Research · NotebookLM · PrismBest entry point for synthesis, reading, and report generation
Auto Research systemsThe AI Scientist · The AI Scientist-v2 · Agent Laboratory · EvoScientistDefines what end-to-end or near-end-to-end research automation looks like
Orchestration frameworksAI-Researcher · RD-Agent · Auto-Deep-ResearchShows the framework layer growing around research loops, not only one-off papers
Benchmarks & scientist-aligned evalScienceAgentBench · FIRE-Bench · ResearchClawBench · SGI-Bench · RE-BenchKeeps the field grounded in rediscovery, expert comparison, and workflow realism
Execution substrateClaude Code · Codex · OpenHands · SWE-agent · OpenClawMost research-agent failures now happen here: tool use, code execution, environment control, and iteration
Platform signalsFutureHouse Platform · Robin · Edison ScientificShows the move from repos and papers toward durable product surfaces

For a dedicated overview, start here: → Auto Research


Representative Auto Research Systems

FamilyRepresentative resourcesWhat it optimizes for
End-to-end AI scientistThe AI Scientist · The AI Scientist-v2Idea generation, experiment execution, and paper/report production
Human-in-the-loop research copilotAgent LaboratoryCollaboration and controllable research assistance
Research orchestration frameworkAI-Researcher · Auto-Deep-ResearchSearch, synthesis, and multi-stage research workflow control
Autonomous R&D / data scienceRD-AgentReal implementation loops, evaluation, and applied experimentation
Self-evolving scientistEvoScientistMemory, iteration, and continual improvement of the research process itself

Auto Research Benchmarks

BenchmarkWhat it measuresWhy it matters
ScienceAgentBenchScientific discovery tasks with grounded evaluationOne of the clearest early attempts to benchmark real research-agent capability
FIRE-BenchRediscovery of known scientific insightsMakes "can the system rediscover something real?" a first-class metric
ResearchClawBenchAutonomous research from rediscovery to new-discoveryStrong recent signal that agentic research-workspace evaluation is becoming more realistic
SGI-BenchScientist-aligned workflows and scientific general intelligenceSeparates scientist-like process quality from generic language-model fluency
RE-BenchFrontier AI R&D against human expertsUseful reality check for how far autonomous R&D actually is from human performance

More detail: → Benchmarks


Agent-Native Landscape Beyond Research

Vibe Research remains the center of this repo. But in practice, research agents now live inside a larger operating environment: personal assistants, software-surface layers, self-improving agents, and companion UX around coding agents.

Personal Agent Assistants

These are not all "research agents", but they increasingly shape the environment research agents live in.

ProjectWhat it isWhy it matters here
OpenClawGateway-native assistant runtime with chat control, plugins, bundles, and deployment surfacesShows how personal assistants, plugins, and research flows can share one substrate instead of staying as separate demos
Hermes AgentGeneral-purpose personal agent stack with gateway, CLI, plugins, skills, and long-running plansImportant signal that "personal assistant" is becoming a real open-source platform layer, not only a chatbot wrapper
GooseOpen-source extensible agent that can install, execute, edit, and test with any LLMRepresents the dev-native branch of personal assistants that sits close to repo work and engineering execution
KhojSelf-hostable AI second brain and autonomous personal AI with deep-research and automation hooksShows the knowledge-native assistant pattern where long-term memory, personal docs, and web retrieval become part of the assistant surface
AnythingLLMPrivacy-first workspace-style AI productivity layerUseful reference for the workspace-native assistant pattern where teams want one local-first surface for models, docs, and agents

Agent-Native Software / CLI-Anything

This layer matters because the frontier is no longer only "which agent is best", but also "which software surfaces are now agent-operable".

ProjectWhat it isWhy it matters here
CLI-AnythingTurns software into agent-native CLI surfaces and ships many agent-harness adaptersStrong signal that existing tools are being retrofitted into agent-operable interfaces instead of being rebuilt from scratch
cc-connectChat control plane for Claude Code, Codex, Cursor, Gemini CLI, and related agentsShows how chat surfaces are becoming remote-control shells for coding and research agents
Official MCP RegistryCanonical discovery layer for MCP servers and toolsMakes tool discovery and installation a real infrastructure concern instead of ad hoc glue
anthropics/skillsPublic reusable skill substrate for Claude CodeShows how skills are becoming portable capability units rather than only local prompts
ClawHub · OpenClaw Plugin BundlesRegistry, bundle compatibility, and install layer around OpenClawGood reference for how agent-native ecosystems package and distribute capabilities

Learning, RL & Self-Evolving Agents

This is still one of the most important structural shifts in the field: not just tool-using agents, but agents that can be trained, optimized, or improved over time.

LayerRepresentative resourcesWhy it matters
Agent training / optimizationAgent LightningBrings RL, automatic prompt optimization, and SFT to arbitrary agent systems with near-zero code changes
Zero-data self-evolutionAgent0 · AgentEvolverShows how agents can generate tasks, feedback, and training signals without human-curated data pipelines
Evolving workflowsEvoAgentX · EvoScientist · MetaClawShifts the focus from optimizing one prompt to evolving whole workflows, skill graphs, or scientist loops
Skill & memory substrateAcontext · anthropics/skillsMakes skills, context, and reusable experience part of the learning layer
Landscape mapAwesome-Self-Evolving-AgentsBest current GitHub-native overview of the optimization, evolution, and lifelong-agent literature

Vibe Coding Apps & Companion UX

This is a newer branch, but it is growing quickly enough that it deserves to be called out separately.

ProjectWhat it isWhy it matters here
CrushOpen-source agentic coding UX from CharmGood signal that the UI layer around coding agents is now its own product surface, not just a wrapper around a model API
Vibe IslandCommercial macOS Dynamic Island companion for many coding agentsShows the new monitor / approve / ask / jump-back interaction layer around long-running coding-agent sessions
xislandFree macOS notch companion for Claude Code, Codex, Gemini CLI, and OpenCodeSimilar signal from an indie distribution angle: approval and session-monitoring UX is becoming a category of its own

Claw Stack At A Glance

The Claw family now spans more than research agents. The cleaner reading is:

LayerRepresentative projectsWhy it matters
Gateway / foundationOpenClawCore runtime, control surface, and the base layer other Claw projects increasingly build on
Registry / discoveryClawHub · awesome-openclaw-skillsSkill and plugin discovery are now a separate ecosystem layer
Compatibility / bundlesOpenClaw Plugin BundlesMakes it easier to import Codex, Claude, and Cursor ecosystem formats into OpenClaw-native features
Packaging / deploymentnix-openclawGives the ecosystem a reproducible deployment path for serious self-hosting and ops
Research surfacesInnoClaw · ResearchClaw · ResearchClaw Desktop AppCovers grounded workspaces, daily research copilots, and lighter-weight reading surfaces
Scientific specialist / evolutionScienceClaw · MetaClawPushes deeper into scientific specialization, persistent memory, and online learning
Autonomous pipelineAutoResearchClawRepresents the maximum-autonomy idea-to-paper direction

Full map: → Claw Park


Ecosystem Snapshot

CLAW Ecosystem - Vibe Research tools and platforms

LayerRepresentative projectsWhy it matters
Research copilotsOpenAI Deep Research · Gemini Deep Research · NotebookLM · PrismFast literature synthesis, source-grounded reading, and scientific writing assistance
Auto Research systemsThe AI Scientist · The AI Scientist-v2 · Agent Laboratory · AI-Researcher · RD-Agent · Auto-Deep-Research · EvoScientistThe current system family for automating or orchestrating meaningful parts of the research loop
Scientist-aligned benchmarksScienceAgentBench · FIRE-Bench · ResearchClawBench · SGI-Bench · RE-BenchKeeps the field grounded in rediscovery, expert comparison, and research-workflow realism
AI scientist platformsFutureHouse Platform · Robin · Edison Scientific · KosmosShows the field moving from paper demos to persistent web/API platforms and validated scientific workflows
Research workspacesInnoClaw · ResearchClaw · FARS · OpenClawWorkspaces and assistant shells that keep papers, files, notes, code, and local execution in one loop
Learning / self-evolving layerAgent Lightning · Agent0 · AgentEvolver · EvoAgentX · AcontextTurns agent training, self-generated data, evolving workflows, and persistent skill/context memory into a real stack layer
Agent-native software / harnessesCLI-Anything · cc-connect · Official MCP Registry · anthropics/skillsShows how software surfaces, registries, and skills are becoming agent-operable infrastructure
Claw ecosystemOpenClaw · ClawHub · OpenClaw Plugin Bundles · nix-openclaw · InnoClaw · ResearchClaw · ScienceClaw · MetaClaw · AutoResearchClawGateway, registry, compatibility, deployment, research workspaces, scientific specialization, online learning, and autonomous pipelines
Execution layerClaude Code · Codex · Cursor Background Agents · GitHub Copilot Coding AgentThe coding and repo workflow layer that increasingly powers research execution
Personal agent assistantsHermes Agent · Goose · Khoj · AnythingLLMShows the wider assistant layer surrounding research: messaging-native, dev-native, knowledge-native, and workspace-native agents
Companion apps / coding UXCrush · Vibe Island · xislandShows the monitoring, approval, and session-jump UX layer forming around long-running coding agents
Adjacent prompt-native toolsv0 · Lovable · Replit AgentUseful for prototyping, but not the core of Vibe Research

Plugins, Bridges & Research Connectors

A new layer is forming between "agent" and "workflow": plugin surfaces, MCP registries, skill catalogs, and chat bridges that make research agents easier to extend, discover, and operate.

LayerRepresentative resourcesWhy it matters
Bridge & control surfacescc-connectRuns Claude Code, Cursor, Gemini CLI, Codex, and similar agents from chat surfaces such as Feishu/Lark, Slack, Telegram, and WeCom
Plugin / customization layerCLI-Anything · ClawHub · OpenClaw Plugin Bundles · awesome-claude-code-pluginsShows how agent ecosystems are moving toward software harnesses, skill registries, plugin marketplaces, bundle compatibility, and installable capability packs
Learning / memory substrateAcontext · anthropics/skillsShows how context, memory, and reusable skills are turning into persistent substrates for agent improvement
Claude Code workflow layerwshobson/agents · SuperClaude Framework · claude-task-masterShows how commands, agent teams, skills, and task systems are turning Claude Code into a fuller development environment
Routing / agent-ops layerclaude-code-router · Claude Squad · RepomixHighlights provider routing, multi-agent session management, and codebase packaging as new operational layers around coding agents
Registry / discovery layerOfficial MCP Registry · awesome-mcp-servers · awesome-openclaw-skillsMakes it easier to find, compare, and install the rapidly growing tool and skill ecosystem
Research connectorsOpenAlex Research MCP · Academia MCP · PapersWithCode MCPConnects agents directly to literature graphs, code artifacts, datasets, and benchmark metadata

More detailed map: → Tools & Platforms


Topic Map

Core Guides

TopicDescriptionLink
🚀 Getting Started5-min demo → 30-min agent deployment → full automation→ Getting Started
🔬 Auto ResearchMap the AI Scientist stack across papers, systems, benchmarks, frameworks, and platform signals→ Auto Research
🧰 Tools & PlatformsCore research platforms plus assistant, connector, and software-surface layers→ Tools & Platforms
🦞 Claw ParkEcosystem map for what each Claw project is building and where it fits→ Claw Park
💻 Vibe CodingTerminal agents, coding agents, companion apps, and repo guardrails→ Vibe Coding
🎨 Vibe AnythingAdjacent prompt-native workflows for apps, design, writing, slides, and ops→ Vibe Anything

Research Topics (39+ papers)

TopicCore QuestionPapersLink
📄 SurveysLandscape, scientific-method framing, and evolution of the field6→ Surveys
⚙️ SystemsHow to design end-to-end research systems7→ Systems
💡 IdeationCan LLMs generate novel ideas6→ Ideation
📚 SynthesisHow to synthesize literature at scale5→ Synthesis
🧪 ExperimentHow agents automate experiments4→ Experiment
✍️ Writing & ReviewLLM-assisted writing & peer review4→ Writing & Review
📊 BenchmarksHow to evaluate research agents7→ Benchmarks

Reading Modes

Read The Field

Auto Research · Surveys · Systems · Benchmarks
Build The Stack

Tools & Platforms · Claw Park · Vibe Coding
Prototype Beyond Research

Vibe Anything

Useful Resources

Introductions: AI for Science (Nature) · LLM Agents (Lilian Weng) · Agentic Patterns (Andrew Ng)

Awesome Lists: LLM Agent Survey · AI Agents · Scientific Idea Generation

Research Core: Semantic Scholar · Elicit · Consensus · Connected Papers

Auto Research / AI Scientist: The AI Scientist · The AI Scientist-v2 · Agent Laboratory · AI-Researcher · RD-Agent · Auto-Deep-Research · EvoScientist

Benchmarks & Evaluation: ScienceAgentBench · FIRE-Bench · ResearchClawBench · SGI-Bench · RE-Bench · MLE-bench

Scientific-Method Framing: From Automation to Autonomy · LLM Agents as AI Scientists: A Survey · A Survey of LLM-based Scientific Agents · Exploring the role of large language models in the scientific method

AI Scientist Platforms: FutureHouse Platform · Robin · BixBench · Edison Scientific · Kosmos

Personal Agent Assistants: OpenClaw · Hermes Agent · Goose · Khoj · AnythingLLM

Agent-Native Interfaces & Harnesses: CLI-Anything · cc-connect · Official MCP Registry · anthropics/skills · ClawHub

Claw Ecosystem: OpenClaw · ClawHub · Plugin Bundles · nix-openclaw · Claw Park

Learning / Self-Evolving Agents: Agent Lightning · Agent0 · AgentEvolver · EvoAgentX · Acontext · Awesome-Self-Evolving-Agents

Execution / Vibe Coding: Claude Code · Codex · Cursor Background Agents · GitHub Copilot Coding Agent · Gemini CLI · Crush

Claude Code Ecosystem: anthropics/skills · wshobson/agents · SuperClaude Framework · claude-code-router · Claude Squad · claude-task-master · Repomix

Companion Apps: Vibe Island · xisland

Prototyping: v0 · Lovable · Replit Agent · Figma AI · Canva AI

Conferences: NeurIPS · ICML · ICLR · ACL · AAAI · EMNLP


Contribute

Submit resources via Resource Suggestion · Contribute via PR · Follow the curation guidelines


Citation
@misc{viberesearch2026,
  title = {Vibe Research Guide},
  author = {Aaron Wang and Contributors},
  year = {2026},
  url = {https://github.com/SpectrAI-Initiative/Vibe-Research-Guide},
}
Changelog
  • 2026-04-21: Added a dedicated Auto Research core guide, reweighted the README back toward Vibe Research, refreshed Systems / Surveys / Benchmarks / Tools, and synced the landing page to the stronger auto-research framing
  • 2026-04-15: Reframed the README as a research-first but broader agent-native map, adding personal assistants, CLI-Anything / harness layers, and companion coding apps while keeping Vibe Research as the center
  • 2026-W17: Expanded Claw coverage from a short project list into a fuller family map, including ClawHub, Plugin Bundles, nix-openclaw, ResearchClaw Desktop App, and a clearer stack-layer taxonomy
  • 2026-W16: Added a dedicated learning / RL / self-evolving layer to the guide, including Agent Lightning, Agent0, AgentEvolver, EvoAgentX, Acontext, and Awesome-Self-Evolving-Agents
  • 2026-W14: Added 2026 Q1 signals for OpenClaw platformization, FutureHouse / Robin / BixBench, and Edison Scientific / Kosmos; refreshed ecosystem framing across the guide
  • 2026-W14: Added recent Claude Code ecosystem signals, including anthropics/skills, wshobson/agents, SuperClaude, claude-code-router, Claude Squad, claude-task-master, and Repomix
  • 2026-W13: Added a new plugin / bridge / registry layer to the guide, including cc-connect, OpenAlex Research MCP, Academia MCP, PapersWithCode MCP, and more Claw ecosystem positioning
  • 2026-W13: Added core tools & platforms (InnoClaw, ResearchClaw, FARS, Orchestra, OpenClaw, EvoScientist); added Deep Research tools, OpenAI Prism, MCP Servers; switched all content to English; expanded to 35+ papers across 9 topic files
  • 2026-W12: Redesigned README into a stronger landing page with cleaner hierarchy, card-style path selection, and a more visual ecosystem map
  • 2026-W12: Hub-and-spoke architecture reorganization
  • 2026-W12: Initial public release

Full history: CHANGELOG.md

MIT License · Star History Chart

ai-science
literature-review
llm-agents
scientific-discovery
vibe-research

Significant stargazers

Seungwoo hong

98 followers · starred May 2026

SpectrAI-Initiative/Vibe-Research-Guide

A curated guide for LLM-agent-driven scientific research automation — from getting started to the frontier.

JavaScript

36

27 commits

updated Apr 21, 2026

See the code

README

Vibe Research Guide

A curated guide for LLM-agent-driven scientific research automation

Stars Last Commit Issues MIT

🌐 View Interactive Multilingual README →
🇨🇳 中文 · 🇺🇸 English · 🇰🇷 한국어 · 🇯🇵 日本語 · 🇩🇪 Deutsch · 🇫🇷 Français · 🇪🇸 Español · 🇮🇹 Italiano · 🇵🇹 Português · 🇸🇦 العربية · 🇹🇭 ไทย · 🇻🇳 Tiếng Việt · 🇷🇺 Русский

Vibe Research Guide Overview


Vibe Research At The Center, With The Broader Agent-Native Stack Around It

Automate the research loop with LLM agents: literature review → idea generation → experiment execution → paper writing → peer review.

This repo is a research-first landing page for the field. The center is still Vibe Research, but the guide now gives Auto Research / AI Scientist its own top-level map and then places Claw, coding agents, connectors, and adjacent assistant ecosystems around that core.

Start here: Getting Started · Auto Research · Tools & Platforms · Claw Park

Vibe Research: AI assistant workflow (idea → literature → experiment → code → result → paper)

At A Glance

Core Question

How far can AI move from research assistant to research operator?

Focus: literature, ideation, experiment, writing, and evaluation.
What Changed In 2026

Research copilots got stronger, learning layers became real, autonomous research systems got more credible, and Vibe Coding became the execution layer.
How To Use This Repo

Treat the README as a map. Treat the topic pages as the actual guide.

2026 Landscape Snapshot

Five shifts now define the field:

  1. Research copilots are stronger and easier to trust: Deep Research, NotebookLM-style source-grounded reading, and scientific workspaces such as Prism are making synthesis and report-writing meaningfully faster.
  2. Auto Research is becoming a recognizable system category: The AI Scientist, The AI Scientist-v2, Agent Laboratory, AI-Researcher, RD-Agent, Auto-Deep-Research, and EvoScientist now form a real system family.
  3. Benchmarks are getting closer to scientific reality: ScienceAgentBench, FIRE-Bench, ResearchClawBench, SGI-Bench, and RE-Bench shift evaluation away from vague "agent capability" claims.
  4. Execution substrates now determine whether research agents actually work: Claude Code, Codex, OpenHands, SWE-agent, OpenClaw, and chat bridges such as cc-connect increasingly matter because experiment loops fail on tooling long before they fail on prose.
  5. A wider agent-native ecosystem is forming around research: personal assistants, registries, skills, self-evolving stacks, and companion UX are increasingly part of the same operational environment.

2026 Auto Research Signals

Several current signals make the field feel less like a loose collection of demos and more like an emerging research stack:

  1. The field has moved from "can an agent summarize papers?" to "can an agent operate a research loop?"
  2. System families are diverging clearly: end-to-end scientists, human-in-the-loop research copilots, R&D execution agents, deep-research assistants, and self-evolving scientist stacks now look meaningfully different.
  3. Benchmarking is improving fast: scientist-aligned workflows, rediscovery tasks, expert comparison, and checklist-based evaluation are becoming standard expectations instead of afterthoughts.
  4. Frameworks matter as much as flagship demos: AI-Researcher, RD-Agent, and Auto-Deep-Research show that orchestration and harness design are now first-class concerns.
  5. Platformization is visible: FutureHouse Platform, Robin, BixBench, and Edison Scientific / Kosmos show how AI-scientist ideas are moving from one-off papers to persistent public or commercial surfaces.

Choose a Path

🟢 New to Vibe Research

Start: Getting Started
Then: Auto Research · Tools & Platforms
🔵 Developer / Builder

Start: Auto Research
Then: Tools & Platforms · Vibe Coding · Systems
🔴 Researcher

Start: Surveys
Then: Auto Research · Benchmarks · Ideation
🟣 Creator / Operator

Start: Tools & Platforms
Then: Vibe Coding · Vibe Anything

Only have 5 minutes? Install InnoClaw and try it out.


Auto Research Stack

This is the shortest useful way to read the field in 2026:

LayerRepresentative resourcesWhy it matters
Research copilotsDeep Research · NotebookLM · PrismBest entry point for synthesis, reading, and report generation
Auto Research systemsThe AI Scientist · The AI Scientist-v2 · Agent Laboratory · EvoScientistDefines what end-to-end or near-end-to-end research automation looks like
Orchestration frameworksAI-Researcher · RD-Agent · Auto-Deep-ResearchShows the framework layer growing around research loops, not only one-off papers
Benchmarks & scientist-aligned evalScienceAgentBench · FIRE-Bench · ResearchClawBench · SGI-Bench · RE-BenchKeeps the field grounded in rediscovery, expert comparison, and workflow realism
Execution substrateClaude Code · Codex · OpenHands · SWE-agent · OpenClawMost research-agent failures now happen here: tool use, code execution, environment control, and iteration
Platform signalsFutureHouse Platform · Robin · Edison ScientificShows the move from repos and papers toward durable product surfaces

For a dedicated overview, start here: → Auto Research


Representative Auto Research Systems

FamilyRepresentative resourcesWhat it optimizes for
End-to-end AI scientistThe AI Scientist · The AI Scientist-v2Idea generation, experiment execution, and paper/report production
Human-in-the-loop research copilotAgent LaboratoryCollaboration and controllable research assistance
Research orchestration frameworkAI-Researcher · Auto-Deep-ResearchSearch, synthesis, and multi-stage research workflow control
Autonomous R&D / data scienceRD-AgentReal implementation loops, evaluation, and applied experimentation
Self-evolving scientistEvoScientistMemory, iteration, and continual improvement of the research process itself

Auto Research Benchmarks

BenchmarkWhat it measuresWhy it matters
ScienceAgentBenchScientific discovery tasks with grounded evaluationOne of the clearest early attempts to benchmark real research-agent capability
FIRE-BenchRediscovery of known scientific insightsMakes "can the system rediscover something real?" a first-class metric
ResearchClawBenchAutonomous research from rediscovery to new-discoveryStrong recent signal that agentic research-workspace evaluation is becoming more realistic
SGI-BenchScientist-aligned workflows and scientific general intelligenceSeparates scientist-like process quality from generic language-model fluency
RE-BenchFrontier AI R&D against human expertsUseful reality check for how far autonomous R&D actually is from human performance

More detail: → Benchmarks


Agent-Native Landscape Beyond Research

Vibe Research remains the center of this repo. But in practice, research agents now live inside a larger operating environment: personal assistants, software-surface layers, self-improving agents, and companion UX around coding agents.

Personal Agent Assistants

These are not all "research agents", but they increasingly shape the environment research agents live in.

ProjectWhat it isWhy it matters here
OpenClawGateway-native assistant runtime with chat control, plugins, bundles, and deployment surfacesShows how personal assistants, plugins, and research flows can share one substrate instead of staying as separate demos
Hermes AgentGeneral-purpose personal agent stack with gateway, CLI, plugins, skills, and long-running plansImportant signal that "personal assistant" is becoming a real open-source platform layer, not only a chatbot wrapper
GooseOpen-source extensible agent that can install, execute, edit, and test with any LLMRepresents the dev-native branch of personal assistants that sits close to repo work and engineering execution
KhojSelf-hostable AI second brain and autonomous personal AI with deep-research and automation hooksShows the knowledge-native assistant pattern where long-term memory, personal docs, and web retrieval become part of the assistant surface
AnythingLLMPrivacy-first workspace-style AI productivity layerUseful reference for the workspace-native assistant pattern where teams want one local-first surface for models, docs, and agents

Agent-Native Software / CLI-Anything

This layer matters because the frontier is no longer only "which agent is best", but also "which software surfaces are now agent-operable".

ProjectWhat it isWhy it matters here
CLI-AnythingTurns software into agent-native CLI surfaces and ships many agent-harness adaptersStrong signal that existing tools are being retrofitted into agent-operable interfaces instead of being rebuilt from scratch
cc-connectChat control plane for Claude Code, Codex, Cursor, Gemini CLI, and related agentsShows how chat surfaces are becoming remote-control shells for coding and research agents
Official MCP RegistryCanonical discovery layer for MCP servers and toolsMakes tool discovery and installation a real infrastructure concern instead of ad hoc glue
anthropics/skillsPublic reusable skill substrate for Claude CodeShows how skills are becoming portable capability units rather than only local prompts
ClawHub · OpenClaw Plugin BundlesRegistry, bundle compatibility, and install layer around OpenClawGood reference for how agent-native ecosystems package and distribute capabilities

Learning, RL & Self-Evolving Agents

This is still one of the most important structural shifts in the field: not just tool-using agents, but agents that can be trained, optimized, or improved over time.

LayerRepresentative resourcesWhy it matters
Agent training / optimizationAgent LightningBrings RL, automatic prompt optimization, and SFT to arbitrary agent systems with near-zero code changes
Zero-data self-evolutionAgent0 · AgentEvolverShows how agents can generate tasks, feedback, and training signals without human-curated data pipelines
Evolving workflowsEvoAgentX · EvoScientist · MetaClawShifts the focus from optimizing one prompt to evolving whole workflows, skill graphs, or scientist loops
Skill & memory substrateAcontext · anthropics/skillsMakes skills, context, and reusable experience part of the learning layer
Landscape mapAwesome-Self-Evolving-AgentsBest current GitHub-native overview of the optimization, evolution, and lifelong-agent literature

Vibe Coding Apps & Companion UX

This is a newer branch, but it is growing quickly enough that it deserves to be called out separately.

ProjectWhat it isWhy it matters here
CrushOpen-source agentic coding UX from CharmGood signal that the UI layer around coding agents is now its own product surface, not just a wrapper around a model API
Vibe IslandCommercial macOS Dynamic Island companion for many coding agentsShows the new monitor / approve / ask / jump-back interaction layer around long-running coding-agent sessions
xislandFree macOS notch companion for Claude Code, Codex, Gemini CLI, and OpenCodeSimilar signal from an indie distribution angle: approval and session-monitoring UX is becoming a category of its own

Claw Stack At A Glance

The Claw family now spans more than research agents. The cleaner reading is:

LayerRepresentative projectsWhy it matters
Gateway / foundationOpenClawCore runtime, control surface, and the base layer other Claw projects increasingly build on
Registry / discoveryClawHub · awesome-openclaw-skillsSkill and plugin discovery are now a separate ecosystem layer
Compatibility / bundlesOpenClaw Plugin BundlesMakes it easier to import Codex, Claude, and Cursor ecosystem formats into OpenClaw-native features
Packaging / deploymentnix-openclawGives the ecosystem a reproducible deployment path for serious self-hosting and ops
Research surfacesInnoClaw · ResearchClaw · ResearchClaw Desktop AppCovers grounded workspaces, daily research copilots, and lighter-weight reading surfaces
Scientific specialist / evolutionScienceClaw · MetaClawPushes deeper into scientific specialization, persistent memory, and online learning
Autonomous pipelineAutoResearchClawRepresents the maximum-autonomy idea-to-paper direction

Full map: → Claw Park


Ecosystem Snapshot

CLAW Ecosystem - Vibe Research tools and platforms

LayerRepresentative projectsWhy it matters
Research copilotsOpenAI Deep Research · Gemini Deep Research · NotebookLM · PrismFast literature synthesis, source-grounded reading, and scientific writing assistance
Auto Research systemsThe AI Scientist · The AI Scientist-v2 · Agent Laboratory · AI-Researcher · RD-Agent · Auto-Deep-Research · EvoScientistThe current system family for automating or orchestrating meaningful parts of the research loop
Scientist-aligned benchmarksScienceAgentBench · FIRE-Bench · ResearchClawBench · SGI-Bench · RE-BenchKeeps the field grounded in rediscovery, expert comparison, and research-workflow realism
AI scientist platformsFutureHouse Platform · Robin · Edison Scientific · KosmosShows the field moving from paper demos to persistent web/API platforms and validated scientific workflows
Research workspacesInnoClaw · ResearchClaw · FARS · OpenClawWorkspaces and assistant shells that keep papers, files, notes, code, and local execution in one loop
Learning / self-evolving layerAgent Lightning · Agent0 · AgentEvolver · EvoAgentX · AcontextTurns agent training, self-generated data, evolving workflows, and persistent skill/context memory into a real stack layer
Agent-native software / harnessesCLI-Anything · cc-connect · Official MCP Registry · anthropics/skillsShows how software surfaces, registries, and skills are becoming agent-operable infrastructure
Claw ecosystemOpenClaw · ClawHub · OpenClaw Plugin Bundles · nix-openclaw · InnoClaw · ResearchClaw · ScienceClaw · MetaClaw · AutoResearchClawGateway, registry, compatibility, deployment, research workspaces, scientific specialization, online learning, and autonomous pipelines
Execution layerClaude Code · Codex · Cursor Background Agents · GitHub Copilot Coding AgentThe coding and repo workflow layer that increasingly powers research execution
Personal agent assistantsHermes Agent · Goose · Khoj · AnythingLLMShows the wider assistant layer surrounding research: messaging-native, dev-native, knowledge-native, and workspace-native agents
Companion apps / coding UXCrush · Vibe Island · xislandShows the monitoring, approval, and session-jump UX layer forming around long-running coding agents
Adjacent prompt-native toolsv0 · Lovable · Replit AgentUseful for prototyping, but not the core of Vibe Research

Plugins, Bridges & Research Connectors

A new layer is forming between "agent" and "workflow": plugin surfaces, MCP registries, skill catalogs, and chat bridges that make research agents easier to extend, discover, and operate.

LayerRepresentative resourcesWhy it matters
Bridge & control surfacescc-connectRuns Claude Code, Cursor, Gemini CLI, Codex, and similar agents from chat surfaces such as Feishu/Lark, Slack, Telegram, and WeCom
Plugin / customization layerCLI-Anything · ClawHub · OpenClaw Plugin Bundles · awesome-claude-code-pluginsShows how agent ecosystems are moving toward software harnesses, skill registries, plugin marketplaces, bundle compatibility, and installable capability packs
Learning / memory substrateAcontext · anthropics/skillsShows how context, memory, and reusable skills are turning into persistent substrates for agent improvement
Claude Code workflow layerwshobson/agents · SuperClaude Framework · claude-task-masterShows how commands, agent teams, skills, and task systems are turning Claude Code into a fuller development environment
Routing / agent-ops layerclaude-code-router · Claude Squad · RepomixHighlights provider routing, multi-agent session management, and codebase packaging as new operational layers around coding agents
Registry / discovery layerOfficial MCP Registry · awesome-mcp-servers · awesome-openclaw-skillsMakes it easier to find, compare, and install the rapidly growing tool and skill ecosystem
Research connectorsOpenAlex Research MCP · Academia MCP · PapersWithCode MCPConnects agents directly to literature graphs, code artifacts, datasets, and benchmark metadata

More detailed map: → Tools & Platforms


Topic Map

Core Guides

TopicDescriptionLink
🚀 Getting Started5-min demo → 30-min agent deployment → full automation→ Getting Started
🔬 Auto ResearchMap the AI Scientist stack across papers, systems, benchmarks, frameworks, and platform signals→ Auto Research
🧰 Tools & PlatformsCore research platforms plus assistant, connector, and software-surface layers→ Tools & Platforms
🦞 Claw ParkEcosystem map for what each Claw project is building and where it fits→ Claw Park
💻 Vibe CodingTerminal agents, coding agents, companion apps, and repo guardrails→ Vibe Coding
🎨 Vibe AnythingAdjacent prompt-native workflows for apps, design, writing, slides, and ops→ Vibe Anything

Research Topics (39+ papers)

TopicCore QuestionPapersLink
📄 SurveysLandscape, scientific-method framing, and evolution of the field6→ Surveys
⚙️ SystemsHow to design end-to-end research systems7→ Systems
💡 IdeationCan LLMs generate novel ideas6→ Ideation
📚 SynthesisHow to synthesize literature at scale5→ Synthesis
🧪 ExperimentHow agents automate experiments4→ Experiment
✍️ Writing & ReviewLLM-assisted writing & peer review4→ Writing & Review
📊 BenchmarksHow to evaluate research agents7→ Benchmarks

Reading Modes

Read The Field

Auto Research · Surveys · Systems · Benchmarks
Build The Stack

Tools & Platforms · Claw Park · Vibe Coding
Prototype Beyond Research

Vibe Anything

Useful Resources

Introductions: AI for Science (Nature) · LLM Agents (Lilian Weng) · Agentic Patterns (Andrew Ng)

Awesome Lists: LLM Agent Survey · AI Agents · Scientific Idea Generation

Research Core: Semantic Scholar · Elicit · Consensus · Connected Papers

Auto Research / AI Scientist: The AI Scientist · The AI Scientist-v2 · Agent Laboratory · AI-Researcher · RD-Agent · Auto-Deep-Research · EvoScientist

Benchmarks & Evaluation: ScienceAgentBench · FIRE-Bench · ResearchClawBench · SGI-Bench · RE-Bench · MLE-bench

Scientific-Method Framing: From Automation to Autonomy · LLM Agents as AI Scientists: A Survey · A Survey of LLM-based Scientific Agents · Exploring the role of large language models in the scientific method

AI Scientist Platforms: FutureHouse Platform · Robin · BixBench · Edison Scientific · Kosmos

Personal Agent Assistants: OpenClaw · Hermes Agent · Goose · Khoj · AnythingLLM

Agent-Native Interfaces & Harnesses: CLI-Anything · cc-connect · Official MCP Registry · anthropics/skills · ClawHub

Claw Ecosystem: OpenClaw · ClawHub · Plugin Bundles · nix-openclaw · Claw Park

Learning / Self-Evolving Agents: Agent Lightning · Agent0 · AgentEvolver · EvoAgentX · Acontext · Awesome-Self-Evolving-Agents

Execution / Vibe Coding: Claude Code · Codex · Cursor Background Agents · GitHub Copilot Coding Agent · Gemini CLI · Crush

Claude Code Ecosystem: anthropics/skills · wshobson/agents · SuperClaude Framework · claude-code-router · Claude Squad · claude-task-master · Repomix

Companion Apps: Vibe Island · xisland

Prototyping: v0 · Lovable · Replit Agent · Figma AI · Canva AI

Conferences: NeurIPS · ICML · ICLR · ACL · AAAI · EMNLP


Contribute

Submit resources via Resource Suggestion · Contribute via PR · Follow the curation guidelines


Citation
@misc{viberesearch2026,
  title = {Vibe Research Guide},
  author = {Aaron Wang and Contributors},
  year = {2026},
  url = {https://github.com/SpectrAI-Initiative/Vibe-Research-Guide},
}
Changelog
  • 2026-04-21: Added a dedicated Auto Research core guide, reweighted the README back toward Vibe Research, refreshed Systems / Surveys / Benchmarks / Tools, and synced the landing page to the stronger auto-research framing
  • 2026-04-15: Reframed the README as a research-first but broader agent-native map, adding personal assistants, CLI-Anything / harness layers, and companion coding apps while keeping Vibe Research as the center
  • 2026-W17: Expanded Claw coverage from a short project list into a fuller family map, including ClawHub, Plugin Bundles, nix-openclaw, ResearchClaw Desktop App, and a clearer stack-layer taxonomy
  • 2026-W16: Added a dedicated learning / RL / self-evolving layer to the guide, including Agent Lightning, Agent0, AgentEvolver, EvoAgentX, Acontext, and Awesome-Self-Evolving-Agents
  • 2026-W14: Added 2026 Q1 signals for OpenClaw platformization, FutureHouse / Robin / BixBench, and Edison Scientific / Kosmos; refreshed ecosystem framing across the guide
  • 2026-W14: Added recent Claude Code ecosystem signals, including anthropics/skills, wshobson/agents, SuperClaude, claude-code-router, Claude Squad, claude-task-master, and Repomix
  • 2026-W13: Added a new plugin / bridge / registry layer to the guide, including cc-connect, OpenAlex Research MCP, Academia MCP, PapersWithCode MCP, and more Claw ecosystem positioning
  • 2026-W13: Added core tools & platforms (InnoClaw, ResearchClaw, FARS, Orchestra, OpenClaw, EvoScientist); added Deep Research tools, OpenAI Prism, MCP Servers; switched all content to English; expanded to 35+ papers across 9 topic files
  • 2026-W12: Redesigned README into a stronger landing page with cleaner hierarchy, card-style path selection, and a more visual ecosystem map
  • 2026-W12: Hub-and-spoke architecture reorganization
  • 2026-W12: Initial public release

Full history: CHANGELOG.md

MIT License · Star History Chart

ai-science
literature-review
llm-agents
scientific-discovery
vibe-research

Significant stargazers

Seungwoo hong

98 followers · starred May 2026