Wholiver/metis

Metis is a coding agent that boosts AI/LLM coding performance by 50%

138

stars

132

commits

TypeScript

primary language

Sep 10, 2026

updated

agentic-workflow
agent-orchestration
agent-skills
ai-agents
ai-coding
automated-testing
autonomous-agents
claude-code
cli
code-review
codex
coding-agents
developer-tools
github-copilot
multi-agent-systems
opencode
prompt-engineering
subagents
test-driven-development
workflow-automation

README

Metis app icon

English · 简体中文

TypeScript npm version latest GitHub release Node.js 22.19.0 or newer MIT License Powered by OrcaRouter

A coding agent that searches, remembers, executes, and verifies across terminal and desktop.

Quick start · Benchmark & Comparison · Key features · Documentation

Quick start

Desktop

Standalone application with built-in Metis CLI and Server runtime (no Node.js required):

CLI installation

Requires Node.js >=22.19.0.

npm install -g @wholiver_hu/metis
metis

Run in any repository or directory:

metis "Explain this repository"
metis @src/main.ts "Review this file"
git diff | metis -p "Review this diff"

Use /login for subscription providers or configure an API key. See Quickstart for the complete guide.

Benchmark & Comparison

Terminal-Bench 2.1 Benchmark Results

In a controlled benchmark run using the same model (DeepSeek V4 Flash), same 89 real-world tasks, identical budget, and environment:

Agent FrameworkModelBenchmarkSolved (Accuracy)Architecture & Harness Advantage
🏆 MetisDeepSeek V4 FlashTerminal-Bench 2.1 (89 tasks)73 / 89 (82.02%)✅ Recursive 5-role agents + SQLite memory + Plan/Build
OpenCodeDeepSeek V4 FlashTerminal-Bench 2.1 (89 tasks)60 / 89 (67.42%)⚠️ Single-thread flat tool execution
📈 ImprovementSame Model & BudgetSame Environment+14.6% (+13 tasks)🚀 Harness, memory, and verification gates alone

Feature Comparison Matrix

CapabilityMetisClaude CodeOpenCodeCursor / Cline
License & PricingMIT ($0 Free)❌ Proprietary✅ MIT ($0 Free)⚠️ Commercial
Model FreedomAny Model / OrcaRouter❌ Anthropic Only✅ Multi-Provider⚠️ Limited / BYOK
User InterfacesTUI + React Desktop⚠️ Terminal Only⚠️ Terminal Only⚠️ IDE Only
Workflow ModePlan ↔ Build Dual-Mode⚠️ Single Flow⚠️ Single Flow⚠️ Chat / Inline
Multi-Agent SystemRecursive L0→L4 (5 Roles)⚠️ Flat Subagents⚠️ Basic❌ None
Durable MemorySQLite + Vector Search❌ Ephemeral❌ Ephemeral⚠️ Code Embeddings
Verification GatesTest Gates + Video Evidence⚠️ Manual Bash⚠️ Manual Bash⚠️ Basic Linter
Headless BenchmarkPython Adapter + Trace❌ None⚠️ Partial❌ None

Key Features

  • Plan & Build Dual Workflows — Safely investigate in read-only Plan mode, then execute approved plans with a live-updating checklist in Build mode.
  • Dual Interface for Terminal & Desktop — Work directly in your terminal via the rich interactive TUI, or use the dedicated React/Vite Desktop workspace on macOS and Windows.
  • Recursive Multi-Agent System — Native named agents (coordinator, planner, implementer, reviewer, verifier) with L0→L4 recursive delegation and Git Worktree isolation.
  • Durable Memory & Resumable Sessions — Project knowledge and decisions persist in SQLite across restarts, context compactions, and session forks.
  • Extensible & Model-Agnostic — Use any LLM provider (OpenAI, Anthropic, DeepSeek, OrcaRouter, Gemini, Groq, Ollama, vLLM) and extend with TypeScript plugins, Agent Skills, and MCP.
  • Benchmark-Grade Reliability — Automated verification gates, video evidence inspection, and full Terminal-Bench & Harbor evaluation readiness.

Documentation

TopicGuide
Install, authenticate, and startQuickstart
Commands and terminal UIUsing Metis · TUI
Providers and custom modelsProviders · Custom models · Custom providers
Multi-Agent SystemNamed Agents & Delegation
Benchmark & EvaluationTerminalBench & Harbor
Sessions and compactionSessions · Compaction
Extensions, skills, and packagesExtensions · Skills · Packages
Prompts and interface customizationPrompt templates · Themes · Keybindings
Programmatic integrationSDK · RPC · JSON
Video inspectionVideo tool
Security and configurationSecurity · Settings
Platforms and isolationWindows · Termux · tmux · Containers

See the documentation index for every guide.

Developer information
npm run build                 # Compile TypeScript and copy runtime assets
npm test                      # Run the Vitest suite
npm run clean                 # Remove compiled output
npm run build:binary          # Build the standalone binary
npm --prefix desktop run dev  # Start the React/Vite Desktop app in development
npm --prefix desktop run build # Build the renderer and Electron artifact

The package exports the Node.js SDK from @wholiver_hu/metis and the RPC entry point from @wholiver_hu/metis/rpc-entry.

Contributing

Contributions are welcome. See CONTRIBUTING.md for development, Extension and Package integration, testing, and AI-assisted contribution guidance.

License

Distributed under the MIT License.

Contributors

Wholiver

122 commits

ccw-dy

9 commits

Wholiver/metis

Metis is a coding agent that boosts AI/LLM coding performance by 50%

138

stars

132

commits

TypeScript

primary language

Sep 10, 2026

updated

agentic-workflow
agent-orchestration
agent-skills
ai-agents
ai-coding
automated-testing
autonomous-agents
claude-code
cli
code-review
codex
coding-agents
developer-tools
github-copilot
multi-agent-systems
opencode
prompt-engineering
subagents
test-driven-development
workflow-automation

README

Metis app icon

English · 简体中文

TypeScript npm version latest GitHub release Node.js 22.19.0 or newer MIT License Powered by OrcaRouter

A coding agent that searches, remembers, executes, and verifies across terminal and desktop.

Quick start · Benchmark & Comparison · Key features · Documentation

Quick start

Desktop

Standalone application with built-in Metis CLI and Server runtime (no Node.js required):

CLI installation

Requires Node.js >=22.19.0.

npm install -g @wholiver_hu/metis
metis

Run in any repository or directory:

metis "Explain this repository"
metis @src/main.ts "Review this file"
git diff | metis -p "Review this diff"

Use /login for subscription providers or configure an API key. See Quickstart for the complete guide.

Benchmark & Comparison

Terminal-Bench 2.1 Benchmark Results

In a controlled benchmark run using the same model (DeepSeek V4 Flash), same 89 real-world tasks, identical budget, and environment:

Agent FrameworkModelBenchmarkSolved (Accuracy)Architecture & Harness Advantage
🏆 MetisDeepSeek V4 FlashTerminal-Bench 2.1 (89 tasks)73 / 89 (82.02%)✅ Recursive 5-role agents + SQLite memory + Plan/Build
OpenCodeDeepSeek V4 FlashTerminal-Bench 2.1 (89 tasks)60 / 89 (67.42%)⚠️ Single-thread flat tool execution
📈 ImprovementSame Model & BudgetSame Environment+14.6% (+13 tasks)🚀 Harness, memory, and verification gates alone

Feature Comparison Matrix

CapabilityMetisClaude CodeOpenCodeCursor / Cline
License & PricingMIT ($0 Free)❌ Proprietary✅ MIT ($0 Free)⚠️ Commercial
Model FreedomAny Model / OrcaRouter❌ Anthropic Only✅ Multi-Provider⚠️ Limited / BYOK
User InterfacesTUI + React Desktop⚠️ Terminal Only⚠️ Terminal Only⚠️ IDE Only
Workflow ModePlan ↔ Build Dual-Mode⚠️ Single Flow⚠️ Single Flow⚠️ Chat / Inline
Multi-Agent SystemRecursive L0→L4 (5 Roles)⚠️ Flat Subagents⚠️ Basic❌ None
Durable MemorySQLite + Vector Search❌ Ephemeral❌ Ephemeral⚠️ Code Embeddings
Verification GatesTest Gates + Video Evidence⚠️ Manual Bash⚠️ Manual Bash⚠️ Basic Linter
Headless BenchmarkPython Adapter + Trace❌ None⚠️ Partial❌ None

Key Features

  • Plan & Build Dual Workflows — Safely investigate in read-only Plan mode, then execute approved plans with a live-updating checklist in Build mode.
  • Dual Interface for Terminal & Desktop — Work directly in your terminal via the rich interactive TUI, or use the dedicated React/Vite Desktop workspace on macOS and Windows.
  • Recursive Multi-Agent System — Native named agents (coordinator, planner, implementer, reviewer, verifier) with L0→L4 recursive delegation and Git Worktree isolation.
  • Durable Memory & Resumable Sessions — Project knowledge and decisions persist in SQLite across restarts, context compactions, and session forks.
  • Extensible & Model-Agnostic — Use any LLM provider (OpenAI, Anthropic, DeepSeek, OrcaRouter, Gemini, Groq, Ollama, vLLM) and extend with TypeScript plugins, Agent Skills, and MCP.
  • Benchmark-Grade Reliability — Automated verification gates, video evidence inspection, and full Terminal-Bench & Harbor evaluation readiness.

Documentation

TopicGuide
Install, authenticate, and startQuickstart
Commands and terminal UIUsing Metis · TUI
Providers and custom modelsProviders · Custom models · Custom providers
Multi-Agent SystemNamed Agents & Delegation
Benchmark & EvaluationTerminalBench & Harbor
Sessions and compactionSessions · Compaction
Extensions, skills, and packagesExtensions · Skills · Packages
Prompts and interface customizationPrompt templates · Themes · Keybindings
Programmatic integrationSDK · RPC · JSON
Video inspectionVideo tool
Security and configurationSecurity · Settings
Platforms and isolationWindows · Termux · tmux · Containers

See the documentation index for every guide.

Developer information
npm run build                 # Compile TypeScript and copy runtime assets
npm test                      # Run the Vitest suite
npm run clean                 # Remove compiled output
npm run build:binary          # Build the standalone binary
npm --prefix desktop run dev  # Start the React/Vite Desktop app in development
npm --prefix desktop run build # Build the renderer and Electron artifact

The package exports the Node.js SDK from @wholiver_hu/metis and the RPC entry point from @wholiver_hu/metis/rpc-entry.

Contributing

Contributions are welcome. See CONTRIBUTING.md for development, Extension and Package integration, testing, and AI-assisted contribution guidance.

License

Distributed under the MIT License.

Contributors

Wholiver

122 commits

ccw-dy

9 commits

Languages

TypeScript

67.3%

JavaScript

31.8%