Codebase intelligence for AI and humans: code health scores, auto-generated docs, git analytics, dead code detection, and architectural decisions via MCP.
See the codeRepowise indexes your code, call graph, git history, tests, docs and design
decisions once, on your machine. Then you and your coding agent ask it things:
what calls this, what breaks if I change it, what is dead, what to fix first, and why it was built this way.
−31.6% | 2.3× | 58,924 |
| less agent output 3.8 vs 7.2 tool calls 43 questions · p<0.0001 | more defects surfaced than CodeScene same 20% review budget 2,770 files · p=0.003 | files indexed in one run dotnet/runtime, one laptop 120 min · 11.7 GiB peak |
Call graph checked against the compiler: 0.976 to 0.995 precision on Go and TypeScript,
and no tool that finds as much of the graph gets more of it right, in 7 of 7 comparisons.
5 tools · 37,853 compiler edges · every losing row published
Graph, risk, health, tests and dead code make zero LLM calls · no API key needed · free and self-hosted · AGPL-3.0 or commercial
What it does · Quickstart · Agents · Changes · Code health · Workspaces · Evidence · Enterprise · Docs
Repowise is an ambitious project. We want every engineer, and every agent working beside them, to understand a codebase the way the person who has maintained it for five years does: what calls what, what tends to break, what nobody uses anymore, and why it was built this way. Cutting tokens was never the goal. It happens anyway, because an agent that can ask the index stops searching, opening and re-reading files to find out. Measured against the other context tools on the same agent tasks, it is also the largest saving.
Open Claude Code, Codex, Cursor or any MCP-capable agent in your repository and paste:
Read https://docs.repowise.dev/setup.md and set up Repowise in this repository.
The agent installs Repowise, indexes the repo with no API key, wires itself to the index, and asks you before anything costs money. Prefer to do it yourself:
uv tool install repowise # or: pipx install repowise / pip install repowise
cd /path/to/your/repo
repowise init --yes --no-prose # graph, git, health, dead code, docs. No key, no spend.
repowise serve # local dashboard + MCP server
init wires Claude Code automatically. Then ask your agent "Use Repowise
get_overview to summarize this repository" or "What breaks if I change
src/auth.py?"
Full setup, every agent, optional model-written docs →
One index, three ways to use it. Find the question you came with; each one links to the page that answers it.
| You ask | Repowise gives you |
|---|---|
| How does checkout work in this repo? | A cited answer built from the call graph and the generated docs, in one call. Search and answers |
| What calls this function, and what does it call? | A call graph across 26 parsed languages, every edge stamped with how it was resolved and how far to trust it, plus traced execution flows from each entry point. The graph |
| Can I get docs for this codebase? | A wiki for every module and file, rendered from the code's structure with no key, or written by a model when you choose. It updates incrementally after each commit. Docs |
| Which of our docs are wrong? | Markdown checked against the tree: every reference to a file or symbol the code no longer has, with the line to edit. Doc drift |
| Why is it built this way? | Decisions mined from ADRs, # WHY: comments, commit and PR history and your agent sessions, each tied to the code it governs and flagged when it goes stale. Decisions |
| Who knows this code? | Owners, bus factor, knowledge-loss risk when the main author goes quiet, and suggested reviewers. Ownership |
| Can I see the architecture? | An explorable dependency map, C4 views, and a Structurizr export, no model involved. Dashboard |
| You ask | Repowise gives you |
|---|---|
| What breaks if I change this? | Symbol-level blast radius: the callers of what you changed, the files that historically change with it but are missing from your diff, and the tests that reach it. Change risk |
| How risky is this PR? | Where the change ranks against your repository's own recent commits, with the reasons, as a directive your agent can act on. Change risk |
| Which tests should run? | The tests a diff actually exercises, from a coverage report if you have one and from the call graph if you do not. Test intelligence |
| Did this PR add untested lines? | Patch coverage, branch coverage on changed lines and path-scoped gates in CI, on GitHub, GitLab or any runner. CI gates |
| Will this API change break another repo? | HTTP, gRPC, topic and OpenAPI contracts matched across repositories, with a breaking-change guard and the consumer files it affects. Workspaces |
| Is anyone else editing these files? | Other open branches touching the same files or their co-change partners, each with the reason it is listed. Branch overlap |
| Did we just commit a secret? | Keys, tokens and risky calls found in the working tree and in full git history, with a pre-commit check and a CI gate. Security signals |
| You ask | Repowise gives you |
|---|---|
| What should we fix first? | A ranked queue weighing impact against effort, using churn, fan-in, coverage and bug history. Fix first |
| Where is the debt? | A 1 to 10 score for every file from 53 deterministic detectors, split into defect risk, maintainability and performance, validated against real bug history. Code health |
| Why is this slow? | N+1 queries, I/O in loops, blocking calls inside async code and quadratic loops, traced across function and file boundaries. Performance |
| How do I break this up safely? | Concrete refactoring plans: Extract Method, Extract Class, Move Method, Split File, Break Cycle, with the exact symbols that move and what moves with them. Ready to hand to an agent. Refactoring |
| What can we delete? | Unreachable files, unused exports and unused packages, each with a confidence tier and the evidence behind it. Dead code |
| Where do bugs keep landing? | Bug-fix commits traced to files and symbols, and a warning when your agent edits a bug magnet. Bug history |
| Are our tests testing anything? | Tests with no assertions, tests that only check their own mocks, and untested hotspots. Test-quality smells |
CLAUDE.md and AGENTS.md from the real index and keeps them
current, so even an agent with no MCP support starts informed.repowise distill pytest
keeps every failure and drops the noise, and nothing is lost: an expand command
restores any cut.
Recorded on this repository and its workspace: an agent asking the index, then the dashboard and CLI on the same local index. No API key and nothing uploaded.
| If you care about... | Start here |
|---|---|
| A coding agent that knows the repository | Task-shaped context in fewer calls, with decisions and risk delivered before the agent asks. For agents ↓ |
| Safer pull requests and faster CI | Change risk, symbol-level callers, missing co-changes and the tests a diff needs. Change intelligence ↓ |
| Paying down the code most likely to hurt you | A defect-validated health score, then the concrete fix. Code health ↓ |
| An estate of many repositories | Contracts matched across repos, breaking-change guards, architecture rules in CI, one MCP endpoint for everything. Workspaces ↓ |
| Rolling it out across a company | Self-hosted with nothing leaving your network, per-language accuracy, sizing, compliance status and licensing in one place. Teams and enterprise ↓ |
Every question your agent asks about a repository has an answer that could have been computed ahead of time. Who calls this function? What breaks if I change it? Why is it written this way? Which files are actually dangerous? Without an index, the agent rediscovers that answer on every task: grep, read, re-read, forget.
Repowise gives Claude Code, Codex, Cursor, VS Code and any other MCP host ten task-shaped MCP tools backed by one index of graph, git, docs and decisions. Most code tools are built around data entities, one file or one symbol at a time, which pushes agents into long chains of sequential calls. These are built around tasks: pass several targets in one call and get the whole picture back. The tool list ↓
About tokens. Every tool in this category promises to cut your token bill, and
there are a lot of tools in this category. We think tokens are a symptom. An agent
burns them because it does not know the codebase, so it searches, opens files, opens
more files, and searches again. Give it an index that already knows, and the savings
show up on their own. They also happen to be the best we have measured: in a paired
agent loop over 43 questions on django/django, Repowise cut the agent's own output
by 31.6% (p<0.0001) and got there in 3.8 tool calls instead of 7.2, ahead
of every other context tool in the same run. On 42 sealed retrieval tasks it found
0.876 of the files a fix needed, against 0.610 for the next tool.
Method and every row we lose →
Context arrives before the agent asks. Optional hooks push it into the session when it matters: the governing decision when your agent edits a file that decision covers, a warning when it touches a file with a run of recent bug fixes, a short briefing at session start, and a correction when it reaches for a path that does not exist.
It learns from how you work. Switch on transcript capture
(repowise decision source set session --on) and Repowise reads your own agent
transcripts for the corrections you keep making, turning the durable ones into
tracked decisions it delivers back later. Transcripts never leave your machine; one
batched model call per update turns the candidates that clear the deterministic
gates into records, and --no-llm keeps the gates and drops that call.
Five layers, one index:
| Layer | What it contributes |
|---|---|
| 1. Graph | File and symbol dependencies across 26 AST-parsed languages, confidence-stamped call resolution, communities, centrality, cycles and execution flows |
| 2. Git history | Hotspots, ownership, co-change, bus factor and bug-fix history: behavioural signals static analysis cannot see |
| 3. Docs | A wiki for every module and file, hybrid search, and your own markdown checked against the tree for claims the code no longer supports |
| 4. Decisions | Architectural rationale from ADRs, inline markers, commits, PRs and agent sessions, each claim traced to evidence |
| 5. Health and change | 53 deterministic detectors across defect risk, maintainability and performance, change risk, test impact, dead code and concrete refactoring plans |
The structural wiki needs no model. Model-written prose is an optional upgrade, one page or directory at a time.
Most of what an agent reads back from a shell command is noise: 300 lines of passing
tests wrapped around 4 failures, full commit bodies when it asked what changed
recently. repowise distill <cmd> compresses command output before the agent reads
it, errors first, exit code preserved.
repowise distill pytest # 61% fewer tokens, all 11 failure lines kept
repowise distill git log -50 # 89% fewer tokens
repowise saved # what distillation saved you, in tokens and dollars
Every omission leaves an inline [repowise#<ref>] marker that repowise expand <ref>
reverses in full, so the agent can pull the detail back without re-running the
command. Small outputs pass through untouched. An opt-in hook rewrites noisy commands
for the agent automatically.
The Costs dashboard tallies both savings surfaces. Every event is priced from the model that produced it, and where the evidence is ambiguous it declines to claim a saving. Example from a week of heavy local use.
Full guide: docs/agent/DISTILL.md →
Four deterministic signals, all computed from the graph and git history, no LLM:
base..HEAD range 0-10 from the shape of
the diff, ranked against your repository's own recent commits. PR mode returns
directives an agent can act on: may_break, missing_cochanges, missing_tests,
tests_to_run. One command: repowise risk main..HEAD.
(reference →)repowise overlap and repowise risk.
(reference →)Ingest LCOV, Cobertura, Clover, JaCoCo or a Go coverprofile and you get the measured answer. Most repositories never produce one, so the graph answers instead: a test file that imports a source file reaches it, which is a recorded edge, where most tools fall back to matching file names.
repowise impacted-tests main..HEAD # only the tests this diff exercises
repowise health # untested hotspots, graph-aware
Checked against a real coverage run --contexts=test on this repository:
95.7% precision on what reaches a file and 97.5% on the run
list. Every row is stamped basis: "measured" or
"inferred", measured wins where both can answer, and an empty answer means
unknown, never "no tests".
Test intelligence →
Patch coverage, doc drift, security and change risk run as gates in your own pipeline through the GitHub Action, a GitLab template or plain CLI commands anywhere else, with annotations, SARIF and GitLab Code Quality output. The gates need no API key. Repowise in CI →
On GitHub you can also install the free Repowise PR Bot, a hosted GitHub App that puts the same analysis on every pull request. One comment, edited in place on every push, and a green PR gets no comment at all. It shows symbol-level blast radius (the contracts the PR changed and every caller outside the PR), the tests and co-change partners missing from the change, change risk against the repository's own history, and a public analysis page per PR. Zero LLM calls, so the same diff always gets the same review.
A real comment on a real PR: repowise-dev/repowise#1204 · its analysis page → · install the PR bot →
A score that says "this file is risky" is where most tools stop. Repowise scores every file, finds where the risk concentrates, and names the specific fix.
Every file is scored 1-10 by 53 deterministic detectors (McCabe complexity, brain methods, LCOM4 cohesion, god classes, clone detection, untested hotspots, change entropy, prior-defect history and more), read through three lenses: defect risk, maintainability and performance. Performance findings such as N+1 queries and I/O in loops are traced across functions and files through the call graph, which is where file-local linters lose them. Only 25 of the 53 detectors move the defect score, because that is the number carrying published accuracy claims.
Zero LLM calls, zero cloud. Detector weights are calibrated against a real defect corpus, not hand-tuned: every file scored at a commit before the bug window so nothing leaks backward, with file size as an explicit control, so a detector only earns weight for defect lift beyond a file being big.
It checks itself on your repository. After every index, Repowise compares its own flags with your git history and tells you what it found: "16 of the 20 lowest-health files had a bug fix in the last 6 months, 3.3x the 24% baseline." If that number is bad on your codebase, you will see it.
Then it names the fix. Extract Class, Extract Helper, Move Method, Break Cycle, Split File or Extract Method, with the exact methods, edges and symbols that move, the callers and co-changing files that move with them, and a ranking that puts a fix on a central hub above the same fix on a leaf. Extract Method runs a dataflow pass over the function to lift the exact span and infer a behaviour-preserving signature.
repowise next # what to fix first, ranked by impact and effort
repowise health # KPIs and lowest-scoring files
repowise health --refactoring-targets # ranked, concrete plans
repowise health --trend # snapshots plus declining-health alerts
repowise dead-code # what can go, by confidence tier
The dashboard renders each plan as a card with a copy-to-agent button. An optional model step, never in the indexing path, expands a plan into generated code and a unified diff.
Validated on 21 open-source repositories across 9 languages (2,826 files scored at a fixed point and checked against the following 6 months of bug fixes): ROC AUC 0.737 [0.683, 0.787]. Against CodeScene on the 2,770 files both tools scored, ranking by Repowise surfaces 2.3x the defects under a fixed review budget (paired, p = 0.003). CodeScene keeps a shorter, slightly more precise list. Full head-to-head and its limits →
Guides: code health · refactoring · dead code
repowise serve starts the full web dashboard next to the MCP server. No separate
setup, all local.
![]() Architecture · the dependency graph, laid out and explorable, with per-node context and change coupling | ![]() Code Health · every file as a bubble, hover any one to inspect its score, size, coverage and findings |
![]() Chat · ask the codebase a question, answers cite the files and pages they came from | ![]() Docs · generated wiki pages for the whole codebase, with confidence and freshness badges |
Also in there: Architecture and C4 views, the Knowledge Graph and a zoomable map, Risk, Hotspots, Coupling and Blast radius, Contributors and Ownership, Decisions with an evidence drawer and timeline, Symbols, Security, Dead code, Costs and Workspace. Every view and what it answers: docs/start/DASHBOARD.md →
Real systems are not one repository, and the expensive failures live in the gaps between them. Change a backend contract and Repowise can name the frontend calls that consume it, the services downstream, the companion files missing from the change, and the architecture rule the new dependency violates, before it ships.
| Workspace intelligence | What it answers |
|---|---|
| Contract map | Which services provide and consume each HTTP, gRPC, event, socket and data contract? Links keep their exact or candidate confidence and the source evidence. |
| Cross-repo blast radius | If this provider changes, which downstream services are in structural reach, and which may drift through historical co-change? |
| Breaking-change guard | Was an endpoint removed or an OpenAPI, proto or signature shape changed incompatibly, and which consumer files are linked to that contract? |
| Test impact | Which tests in the consumer repos should run for this provider change, measured from coverage or inferred from the call graph? |
| Architecture as code | Does the live system graph violate declared dependency rules or contain cycles? repowise workspace check gates CI. |
| Architecture health | How coupled is the estate? Propagation cost, the cyclic core, service roles and a deterministic 1-10 architecture score. |
| Federated context | One dashboard and one MCP server answer across every repository while keeping repo-level evidence. |
The system map models services, not just repository boxes, and never conflates a real contract with "these files often changed together". HTTP field-level comparison covers a bounded OpenAPI 3.x subset; matched consumers prove endpoint exposure, not field use or runtime failure. Workspace guide and exact support matrix →
Worktrees and updates stay light: a linked worktree seeds its index from the base checkout, and post-commit hooks, file watching, webhooks or polling keep each repository and the cross-repo graph current. Keeping the index fresh →
26 languages parsed to an AST, 40 on a five-rung ladder, framework-aware where an ecosystem handler exists. Every language ships in the open-source distribution.
Full
Good
· Partial
| Rung | Languages | What you get |
|---|---|---|
| Full (13) | Python · TypeScript · JavaScript · Svelte · Vue · Java · Kotlin · Go · Rust · C++ · C# · Scala · Ruby | The whole pipeline: AST symbols, import resolution, a resolved call graph, heritage, docstrings, framework edges and code-health markers |
| Good (11) | C · Swift · PHP · Dart · Object Pascal · COBOL · GDScript · VB.NET · Elixir · F# · Objective-C | All of the above except the full health suite, within the language-specific ceilings in the full matrix |
| Partial (2) | Luau / Roblox · Razor / Blazor | Luau: AST symbols and require() resolution, Rojo and .luaurc aware. Razor: component symbols, @code and component-tag call edges, C# health markers; no import resolution yet |
| Lightweight (6) | Clojure · Haskell · Lean 4 · Erlang · HTML · QML | A real file-to-file import graph, and no symbol-level claims |
| Structural (8) | R · Zig · Julia · Elm · OCaml · Crystal · Nim · D | Git history: blame, hotspots, co-change, ownership, bug history |
SQL and dbt projects get ref() / source() lineage, shell scripts get function-level
symbols, HTML pages contribute their <script src> and <link href> dependencies, and
OpenAPI, Protobuf, GraphQL, Dockerfile, Terraform and similar formats get dedicated
handlers. Anything else is still tracked through git history.
Every call edge is stamped with how it was resolved and how much to trust it, from
same_file at 0.95 down to a repo-wide name match at 0.50, labelled as the guess it
is. Accuracy per language, graded against each language's own compiler:
docs/BENCHMARKS.md#accuracy-by-language.
Full matrix: docs/layers/LANGUAGE_SUPPORT.md → · adding a language takes five small steps and no changes to the parser core: docs/architecture/language-support.md → · languages moving up the ladder: roadmap →
Six agents wired end to end · two at the Full tier · every other MCP host one paste away.
Full tier
Good tier
Full is every surface Repowise has: MCP tools, skills, slash commands, a managed
instructions file, hook-level interception of tool calls, and transcript mining after
the session. Good is MCP tools and the config to reach them, without hooks or
transcript mining. Anything else that speaks MCP is one snippet away:
repowise agents print-config claude-code prints a server entry for Cline, Windsurf,
Zed, Gemini CLI or any host that reads mcpServers.
Integration matrix →
In VS Code, the Repowise extension shows what your change breaks before you push (riskiest files, what is downstream, forgotten companion files, missing tests, suggested reviewers), health in the gutter and status bar, callers and ownership on hover, and refactoring plans as CodeLens. One install also registers the MCP server, so the same index serves you and your agent. Install from the Marketplace or Open VSX and run Repowise: Set Up This Repository. VS Code guide →
In Claude Code, Lens ships inside the Repowise plugin and shows you what the index
knows while Claude works: the file it is on and how many files depend on it, a change
review after a turn that edits files (code health, tests to run, other branches on the
same files), and /lens, a pane with Flow (what to check before accepting each turn),
a map of the repo lit by what Claude searched, opened and edited and what its edits
reach, Ask, and a session recap. Nothing reaches Claude unless you press a button, and
Lens makes no model calls of its own. It needs Claude Code 2.1.287 or later, and the
map needs repowise serve --no-ui running. Lens guide →
Every response carries a _meta envelope with the indexed commit, the index age and a
stale warning when the index has fallen behind your checkout, so your agent always
knows how much to trust what it just read.
| Tool | What only this tool answers |
|---|---|
get_overview() | Architecture summary, module map, entry points, git health. The first call on an unfamiliar codebase. |
get_answer(question) | Hybrid retrieval (full-text plus vector), graph expansion and one cited answer with a calibrated retrieval_quality. Search, read and reason in a single round-trip. |
get_context(targets, include?) | Triage card for files, modules or symbols: summary, signatures, hotspot flag, governing decisions, symbol ids. include opens callers, callees, ownership and metrics. Batch many targets in one call. |
get_symbol("file.py::Name") | One indexed symbol's source with exact line bounds. |
search_codebase(query) | Hybrid search over code and docs, by symbol, path or concept. |
get_risk(targets?, changed_files?) | Hotspots, dependents, co-change partners, ownership, test gaps, bug history. Pass changed_files for PR mode and get a directive back. |
get_change_risk(revspec) | What a commit, range or uncommitted change made worse across defect risk, maintainability and performance, the tests that touch it, and how the diff ranks against recent commits. |
get_why(query?, targets?) | Architectural decisions with their verbatim evidence. Falls back to git archaeology when no decision exists. |
get_dead_code(...) | Unreachable code by confidence tier, with cross-repo consumers in workspace mode. |
get_health(targets?, include?) | Health scores and findings across all three lenses, Fix first, coverage, trends, doc drift and refactoring plans. |
Ten is a deliberate ceiling: a small, task-shaped surface is easier for an agent to choose from than a large one. Seven more tools (dependency paths, execution flows, refactoring code generation, finding triage, and three workspace architecture tools) are opt-in. Parameters, examples and when to use which: docs/agent/MCP_TOOLS.md →
Open-source agent-context tools, the same repositories, the same pinned commits, the same questions, each tool given its full advertised surface. The full page carries the rows we lose beside the rows we win.
The full results, the methodology, and the rows we lose → · The research it is built on →
No single product competes with all of this, so there is no single table. Rows marked measured are head-to-head numbers that link to docs/BENCHMARKS.md, where the sample sizes, the tests and the rows we lose live. Unmarked rows are capability presence, not measurements.
| repowise | CodeGraph | Serena | DeepWiki | |
|---|---|---|---|---|
| Self-hostable, open source | ✅ AGPL-3.0 | ✅ | ✅ | ❌ cloud only |
| Private repo, no cloud | ✅ | ✅ | ✅ | ❌ OSS forks only |
| MCP tools served | 10 default + 7 opt-in | 1 | 29 | 3 |
| Finds the gold files (measured, n=42 sealed) | ✅ 0.876 | 0.610 | not in this run | not measured |
| Output tokens vs a bare agent (measured, n=43) | ✅ -31.6% | -24.4% | -14.8% | not measured |
| Memory to build the graph (measured, 5 tools, 35 repos) | ✅ 75 MB, lowest on 35 of 35 | 757 MB | not measured | n/a, cloud |
| Time to build the graph (measured, same run) | 2.77s, fastest on 14 of 35 | 3.65s, fastest on 16 | not measured | n/a, cloud |
| Time to build the full index, django (measured) | ⚠️ 366.8s, slowest here | ✅ 16.4s | not measured | n/a, cloud |
| Call-edge precision, hand-graded (measured, 560 rows, 9 languages) | ✅ 85.7% | 58.6% | not measured | not measured |
| Call-edge precision, judged by a compiler (measured, 5 tools, 7 cells) | ✅ nothing that finds as much gets more of it right, 7 of 7 | lower precision in 7 | not measured | not measured |
| Generated documentation | ✅ | ❌ | ❌ | ✅ |
| Proactive agent hooks | ✅ Claude + Codex | ❌ | ❌ | ❌ |
Generated CLAUDE.md / AGENTS.md | ✅ | ❌ | ❌ | ❌ |
| Command-output distillation | ✅ reversible | ❌ | ❌ | ❌ |
| Architectural decision records | ✅ | ❌ | ❌ | ❌ |
| Multi-repo workspace intelligence | ✅ contracts, co-change, federated MCP | ❌ | ❌ | ❌ |
The two cost rows answer different questions. Building the call graph, we are the lightest tool measured, about ten times lighter than the next, and roughly as fast as the fastest. Building the whole index, CodeGraph is 22x faster than we are, because by then we have also built the git-history layer, the wiki, the decisions and the health pass. If a call graph is all you need, that is the right trade and you should take it. With model-written prose on, it is 135x.
The hand-graded precision row cuts both ways. About fourteen percent of the call
edges we draw are wrong, and on seastar CodeGraph grades better than we do. The
compiler row exists because we graded the hand-read one ourselves: on Go and
TypeScript the answer key is the Go team's own call graph and the tsc checker's own
resolution, which we neither wrote nor can tune. Precision alone is easy to win by
drawing almost nothing, and recall alone by drawing everything, so the claim is the
pair.
Competitors measured at CodeGraph 1.5.0, Graphify 0.9.31, Serena 1.6.2.dev0 and
code-review-graph 2.3.7 in August 2026. Repowise compiler-graded figures are from main
on 2026-10-03; the other Repowise rows are from 081a59fa, August 2026.
| repowise | CodeScene | |
|---|---|---|
| Self-hostable, open source | ✅ AGPL-3.0 | ⚠️ on-prem Docker, proprietary |
| Code health score (1-10) | ✅ 53 detectors, 25 scoring | ✅ 25-30 |
| Brain Method / LCOM4 / god class | ✅ | ✅ |
| Defects found at a 20% review budget (measured, 2,770 files) | ✅ 0.173 | 0.074 |
| Effort-aware ranking, Popt (measured, p=0.003) | ✅ 0.607 | 0.462 |
| Precision at that budget (measured) | 0.580 | ✅ 0.636, a shorter list |
| Discrimination, ROC AUC (measured, paired) | 0.731 | 0.705, p=0.054, not significant |
| Business impact (resolution time) | ❌ we could not replicate this on open data | ✅ Code Red study |
| Git intelligence (hotspots, ownership, co-change) | ✅ | ✅ |
| Pre-merge change-risk scoring | ✅ 0-10 + directives | ✅ |
| Concrete cross-file refactoring plans | ✅ graph-aware + blast radius | ⚠️ within-function only |
| Test-coverage intelligence | ✅ LCOV/Cobertura/Clover/JaCoCo/Go | ❌ |
| Dead code detection | ✅ | ❌ |
| Serves it to an AI agent over MCP | ✅ | ✅ |
CodeScene is the only other vendor in this category with a published empirical defect study, which is why it is the one we ran head to head. It flags about 27 files where we flag 132, so if you want a short list to act on, its threshold is the better fit.
DeepWiki, Google Code Wiki and Swimm generate documentation from a repository, which overlaps one of our layers. We have not measured against them, so there is no table here.
| Repowise PR Bot | CodeRabbit | Greptile | |
|---|---|---|---|
| LLM calls per PR | ✅ zero | ❌ every review | ❌ every review |
| Same diff, same review | ✅ deterministic | ❌ sampled output | ❌ sampled output |
| Your code sent to a model provider | ✅ never | ❌ yes | ❌ yes |
| Symbol-level blast radius | ✅ call graph | ❌ | ⚠️ prose, from context |
| Co-change partners missing from the PR | ✅ git history | ❌ | ❌ |
| Change risk vs the repo's own history | ✅ 0-10 + percentile | ❌ | ❌ |
| Silent on a clean PR | ✅ by default | ⚠️ configurable | ⚠️ configurable |
An LLM reviewer is a different product: it can read intent, and it can be wrong in a new way on every run. This one does set arithmetic over a call graph and a git history, so pushing the same diff twice gives the same review twice. Full side-by-side comparisons: repowise.dev/compare →
AI makes producing a change cheaper. It does not make understanding its consequences cheaper. In a large estate that answer crosses repositories, ownership boundaries, service contracts, test suites and years of architectural history. Repowise gives developers, agents, reviewers and platform teams the same evidence about what exists, what depends on it, what is risky and what will break, from one index, in place of a separate health tool, dead-code tool, code-search layer for agents, docs generator and test-impact service.
| Deployment | Where source is read | Where the index lives | What leaves your network |
|---|---|---|---|
Self-hosted, open source (pip install repowise) | your machine | .repowise/ next to the repo | anonymous CLI telemetry you can switch off, and your own LLM provider only if you turn prose on |
| Self-hosted, commercial | your VPC or an air-gapped network | Postgres plus LanceDB or pgvector, inside your network | the same, plus integrations you configure. Nothing at all in air-gapped mode |
| Hosted (repowise.dev) | the hosted indexer | infrastructure we operate | your code goes to the platform |
Graph, git, health, change risk, tests, dead code and PR review make zero LLM calls. Raw source is parsed in memory and never persisted. Prose is optional and runs on your own provider contract or fully offline through Ollama, chosen per repository. The threat model and data flows are in the security review pack. The published research each layer is built on, and how we checked it, is in LINEAGE.md.
Call-graph accuracy is graded against each language's own compiler toolchain, and we publish precision and recall per language, with the misses, in the accuracy table. On current main, precision runs 0.95 to 0.99 for Go, TypeScript, C and most Python, 0.85 to 0.90 for Rust, and 0.73 to 0.92 for Java and C#, where overloaded methods are the main gap.
The largest repository indexed so far is dotnet/runtime: 58,924 files in 120 minutes at 11.7 GiB peak memory, on one 31 GB laptop with no API key. Time and memory by repository size: scale.
Any git host works, because indexing reads a local checkout. On top of that:
repowise init at a plain directory, an export, or a Perforce
or SVN workspace. The graph, docs, decisions and health layers build normally; the
history layer needs a commit log, and Perforce and SVN adapters are
on the roadmap.| Status | Capability |
|---|---|
| Shipping in open source | Every deterministic layer, ten MCP tools, multi-repo workspaces, contract extraction and cross-repo blast radius, test intelligence, architecture conformance, local dashboard, auto-sync, full-history secret scanning. |
| GA commercially | Hosted graph-aware security, CVE prioritization, CycloneDX SBOM and VEX, PCI-DSS and SOC 2 control-coverage reports, audit export and webhook stream, Jira and Confluence, reference HA topology on your infrastructure, custom language extensions, SLA support, IP indemnification. |
| Rolling out | Managed GitHub Enterprise, Azure DevOps, GitLab and Bitbucket integrations. SAML/OIDC SSO and SCIM. Engineering-leader dashboards. |
| Planned | RBAC and multi-tenancy, a packaged air-gap install bundle, the Helm chart. |
Repowise holds no SOC 2, ISO 27001 or other audited certification today. The SOC 2 and PCI-DSS reports above are control-coverage signals from your own findings, and every export says so. Every item is tracked with its status in the capability matrix.
Commercial licences are priced per indexed repository, with unlimited seats inside the licensed set. More and more of the code in a repository is written and read by agents and CI, so seat counts stop tracking the value. Details →
Do we have to open-source our code because of the AGPL? No. Using Repowise inside your company, including on internal servers, creates no obligation to publish your code. The obligations apply if you modify Repowise and offer it to others over a network, or ship it inside your own product. A commercial licence removes them.
Commercial support comes with a named contact, a response-time SLA and a quarterly architecture review. Security fixes land on the latest minor release; reports go to security@repowise.dev under the disclosure policy.
Commercial detail · Security review pack · Roadmap · hello@repowise.dev
repowise.dev runs the same engine fully managed. We run it on our own codebase in the open: live snapshot · explore public repos.
Already indexed locally? repowise publish puts the same repo on repowise.dev in one
command: it asks repowise.dev to index the repo's GitHub remote, so nothing is uploaded
from your machine and only what you have pushed is published. A hosted index usually
takes about 10 minutes. Public repos publish on a free account (up to 2 repos, no card);
private repos and more repos need Pro, free for 10 days with a card.
What hosted adds →
--no-prose, code-derived content stays on
your infrastructure. The CLI reports anonymous, opt-out usage telemetry (command
names and coarse environment only); turn it off with repowise telemetry disable,
DO_NOT_TRACK=1, or by running fully offline.
What's collected →Doing a security review? docs/business/SECURITY_COMPLIANCE.md →
repowise init [PATH] # index a codebase (asks; --yes --no-prose needs no key)
repowise generate [PATH] # write wiki pages with a model, on demand
repowise serve [PATH] # MCP server + local dashboard
repowise update [PATH] # incremental update (--workspace for every repo)
repowise watch # re-index on file change
repowise search "<q>" # hybrid search (fulltext / semantic / symbol / path)
repowise ask "<q>" # a synthesized answer with citations
repowise context <files> # triage card: layer, hotspot, fix history, freshness
repowise symbol <id> # one symbol's body, with verified line bounds
repowise why <q|path> # decisions, rationale, git archaeology
repowise next # what to fix first
repowise health # code-health KPIs and lowest-scoring files
repowise risk main..HEAD # score a branch or PR range
repowise overlap # other branches editing the same files
repowise impacted-tests # only the tests a diff exercises
repowise dead-code # unreachable code by confidence tier
repowise doc-drift # documentation the code no longer supports
repowise security # secrets and risky patterns, tree or full history
repowise decision list # architectural decisions
repowise export --format structurizr # the architecture as Structurizr DSL
repowise distill pytest # compact, errors-first, reversible command output
repowise saved # tokens and dollars saved by distillation
repowise workspace add # multi-repo workspace management
repowise doctor # check setup, API keys, index drift
repowise publish # put this repo on repowise.dev (indexed from GitHub; public repos free)
repowise uninstall # remove what repowise wrote, and say what it left
Every command and flag: docs/reference/CLI_REFERENCE.md · config: docs/reference/CONFIG.md · something not working: docs/start/TROUBLESHOOTING.md · all docs: docs/README.md
git clone https://github.com/repowise-dev/repowise
cd repowise
uv sync --all-packages
uv run repowise --version
uv run pytest tests/unit/
New here? You do not have to read 3,000 files to start. We keep a public index of this repository built by Repowise itself, re-indexed on every push: explore repowise with repowise → (architecture, hotspots, ownership, decisions, and a ranked refactoring backlog you are welcome to pick from).
Full guide, including how to add languages and LLM providers: CONTRIBUTING.md · architecture: docs/architecture/
AGPL-3.0. Free for individuals, teams and companies using Repowise internally.
For commercial licensing (the enterprise security and compliance layer, SSO/SCIM, RBAC, workflow integrations, priority support and SLA, or embedding Repowise in a product without AGPL obligations), see docs/business/COMMERCIAL.md or contact hello@repowise.dev.
Built for engineers who got tired of watching their AI agent cat the same file for the fourth time.
⭐ If Repowise earns a place in your workflow, give it a star. It costs you nothing, and it's the signal that keeps a small team building this in the open.
repowise.dev · Explore → · Discord · X · hello@repowise.dev
(top 24 of 59)
602 followers · starred Aug 2026
822 followers · starred May 2026
569 followers · starred Apr 2026
336 followers · starred Apr 2026
Codebase intelligence for AI and humans: code health scores, auto-generated docs, git analytics, dead code detection, and architectural decisions via MCP.
See the codeRepowise indexes your code, call graph, git history, tests, docs and design
decisions once, on your machine. Then you and your coding agent ask it things:
what calls this, what breaks if I change it, what is dead, what to fix first, and why it was built this way.
−31.6% | 2.3× | 58,924 |
| less agent output 3.8 vs 7.2 tool calls 43 questions · p<0.0001 | more defects surfaced than CodeScene same 20% review budget 2,770 files · p=0.003 | files indexed in one run dotnet/runtime, one laptop 120 min · 11.7 GiB peak |
Call graph checked against the compiler: 0.976 to 0.995 precision on Go and TypeScript,
and no tool that finds as much of the graph gets more of it right, in 7 of 7 comparisons.
5 tools · 37,853 compiler edges · every losing row published
Graph, risk, health, tests and dead code make zero LLM calls · no API key needed · free and self-hosted · AGPL-3.0 or commercial
What it does · Quickstart · Agents · Changes · Code health · Workspaces · Evidence · Enterprise · Docs
Repowise is an ambitious project. We want every engineer, and every agent working beside them, to understand a codebase the way the person who has maintained it for five years does: what calls what, what tends to break, what nobody uses anymore, and why it was built this way. Cutting tokens was never the goal. It happens anyway, because an agent that can ask the index stops searching, opening and re-reading files to find out. Measured against the other context tools on the same agent tasks, it is also the largest saving.
Open Claude Code, Codex, Cursor or any MCP-capable agent in your repository and paste:
Read https://docs.repowise.dev/setup.md and set up Repowise in this repository.
The agent installs Repowise, indexes the repo with no API key, wires itself to the index, and asks you before anything costs money. Prefer to do it yourself:
uv tool install repowise # or: pipx install repowise / pip install repowise
cd /path/to/your/repo
repowise init --yes --no-prose # graph, git, health, dead code, docs. No key, no spend.
repowise serve # local dashboard + MCP server
init wires Claude Code automatically. Then ask your agent "Use Repowise
get_overview to summarize this repository" or "What breaks if I change
src/auth.py?"
Full setup, every agent, optional model-written docs →
One index, three ways to use it. Find the question you came with; each one links to the page that answers it.
| You ask | Repowise gives you |
|---|---|
| How does checkout work in this repo? | A cited answer built from the call graph and the generated docs, in one call. Search and answers |
| What calls this function, and what does it call? | A call graph across 26 parsed languages, every edge stamped with how it was resolved and how far to trust it, plus traced execution flows from each entry point. The graph |
| Can I get docs for this codebase? | A wiki for every module and file, rendered from the code's structure with no key, or written by a model when you choose. It updates incrementally after each commit. Docs |
| Which of our docs are wrong? | Markdown checked against the tree: every reference to a file or symbol the code no longer has, with the line to edit. Doc drift |
| Why is it built this way? | Decisions mined from ADRs, # WHY: comments, commit and PR history and your agent sessions, each tied to the code it governs and flagged when it goes stale. Decisions |
| Who knows this code? | Owners, bus factor, knowledge-loss risk when the main author goes quiet, and suggested reviewers. Ownership |
| Can I see the architecture? | An explorable dependency map, C4 views, and a Structurizr export, no model involved. Dashboard |
| You ask | Repowise gives you |
|---|---|
| What breaks if I change this? | Symbol-level blast radius: the callers of what you changed, the files that historically change with it but are missing from your diff, and the tests that reach it. Change risk |
| How risky is this PR? | Where the change ranks against your repository's own recent commits, with the reasons, as a directive your agent can act on. Change risk |
| Which tests should run? | The tests a diff actually exercises, from a coverage report if you have one and from the call graph if you do not. Test intelligence |
| Did this PR add untested lines? | Patch coverage, branch coverage on changed lines and path-scoped gates in CI, on GitHub, GitLab or any runner. CI gates |
| Will this API change break another repo? | HTTP, gRPC, topic and OpenAPI contracts matched across repositories, with a breaking-change guard and the consumer files it affects. Workspaces |
| Is anyone else editing these files? | Other open branches touching the same files or their co-change partners, each with the reason it is listed. Branch overlap |
| Did we just commit a secret? | Keys, tokens and risky calls found in the working tree and in full git history, with a pre-commit check and a CI gate. Security signals |
| You ask | Repowise gives you |
|---|---|
| What should we fix first? | A ranked queue weighing impact against effort, using churn, fan-in, coverage and bug history. Fix first |
| Where is the debt? | A 1 to 10 score for every file from 53 deterministic detectors, split into defect risk, maintainability and performance, validated against real bug history. Code health |
| Why is this slow? | N+1 queries, I/O in loops, blocking calls inside async code and quadratic loops, traced across function and file boundaries. Performance |
| How do I break this up safely? | Concrete refactoring plans: Extract Method, Extract Class, Move Method, Split File, Break Cycle, with the exact symbols that move and what moves with them. Ready to hand to an agent. Refactoring |
| What can we delete? | Unreachable files, unused exports and unused packages, each with a confidence tier and the evidence behind it. Dead code |
| Where do bugs keep landing? | Bug-fix commits traced to files and symbols, and a warning when your agent edits a bug magnet. Bug history |
| Are our tests testing anything? | Tests with no assertions, tests that only check their own mocks, and untested hotspots. Test-quality smells |
CLAUDE.md and AGENTS.md from the real index and keeps them
current, so even an agent with no MCP support starts informed.repowise distill pytest
keeps every failure and drops the noise, and nothing is lost: an expand command
restores any cut.
Recorded on this repository and its workspace: an agent asking the index, then the dashboard and CLI on the same local index. No API key and nothing uploaded.
| If you care about... | Start here |
|---|---|
| A coding agent that knows the repository | Task-shaped context in fewer calls, with decisions and risk delivered before the agent asks. For agents ↓ |
| Safer pull requests and faster CI | Change risk, symbol-level callers, missing co-changes and the tests a diff needs. Change intelligence ↓ |
| Paying down the code most likely to hurt you | A defect-validated health score, then the concrete fix. Code health ↓ |
| An estate of many repositories | Contracts matched across repos, breaking-change guards, architecture rules in CI, one MCP endpoint for everything. Workspaces ↓ |
| Rolling it out across a company | Self-hosted with nothing leaving your network, per-language accuracy, sizing, compliance status and licensing in one place. Teams and enterprise ↓ |
Every question your agent asks about a repository has an answer that could have been computed ahead of time. Who calls this function? What breaks if I change it? Why is it written this way? Which files are actually dangerous? Without an index, the agent rediscovers that answer on every task: grep, read, re-read, forget.
Repowise gives Claude Code, Codex, Cursor, VS Code and any other MCP host ten task-shaped MCP tools backed by one index of graph, git, docs and decisions. Most code tools are built around data entities, one file or one symbol at a time, which pushes agents into long chains of sequential calls. These are built around tasks: pass several targets in one call and get the whole picture back. The tool list ↓
About tokens. Every tool in this category promises to cut your token bill, and
there are a lot of tools in this category. We think tokens are a symptom. An agent
burns them because it does not know the codebase, so it searches, opens files, opens
more files, and searches again. Give it an index that already knows, and the savings
show up on their own. They also happen to be the best we have measured: in a paired
agent loop over 43 questions on django/django, Repowise cut the agent's own output
by 31.6% (p<0.0001) and got there in 3.8 tool calls instead of 7.2, ahead
of every other context tool in the same run. On 42 sealed retrieval tasks it found
0.876 of the files a fix needed, against 0.610 for the next tool.
Method and every row we lose →
Context arrives before the agent asks. Optional hooks push it into the session when it matters: the governing decision when your agent edits a file that decision covers, a warning when it touches a file with a run of recent bug fixes, a short briefing at session start, and a correction when it reaches for a path that does not exist.
It learns from how you work. Switch on transcript capture
(repowise decision source set session --on) and Repowise reads your own agent
transcripts for the corrections you keep making, turning the durable ones into
tracked decisions it delivers back later. Transcripts never leave your machine; one
batched model call per update turns the candidates that clear the deterministic
gates into records, and --no-llm keeps the gates and drops that call.
Five layers, one index:
| Layer | What it contributes |
|---|---|
| 1. Graph | File and symbol dependencies across 26 AST-parsed languages, confidence-stamped call resolution, communities, centrality, cycles and execution flows |
| 2. Git history | Hotspots, ownership, co-change, bus factor and bug-fix history: behavioural signals static analysis cannot see |
| 3. Docs | A wiki for every module and file, hybrid search, and your own markdown checked against the tree for claims the code no longer supports |
| 4. Decisions | Architectural rationale from ADRs, inline markers, commits, PRs and agent sessions, each claim traced to evidence |
| 5. Health and change | 53 deterministic detectors across defect risk, maintainability and performance, change risk, test impact, dead code and concrete refactoring plans |
The structural wiki needs no model. Model-written prose is an optional upgrade, one page or directory at a time.
Most of what an agent reads back from a shell command is noise: 300 lines of passing
tests wrapped around 4 failures, full commit bodies when it asked what changed
recently. repowise distill <cmd> compresses command output before the agent reads
it, errors first, exit code preserved.
repowise distill pytest # 61% fewer tokens, all 11 failure lines kept
repowise distill git log -50 # 89% fewer tokens
repowise saved # what distillation saved you, in tokens and dollars
Every omission leaves an inline [repowise#<ref>] marker that repowise expand <ref>
reverses in full, so the agent can pull the detail back without re-running the
command. Small outputs pass through untouched. An opt-in hook rewrites noisy commands
for the agent automatically.
The Costs dashboard tallies both savings surfaces. Every event is priced from the model that produced it, and where the evidence is ambiguous it declines to claim a saving. Example from a week of heavy local use.
Full guide: docs/agent/DISTILL.md →
Four deterministic signals, all computed from the graph and git history, no LLM:
base..HEAD range 0-10 from the shape of
the diff, ranked against your repository's own recent commits. PR mode returns
directives an agent can act on: may_break, missing_cochanges, missing_tests,
tests_to_run. One command: repowise risk main..HEAD.
(reference →)repowise overlap and repowise risk.
(reference →)Ingest LCOV, Cobertura, Clover, JaCoCo or a Go coverprofile and you get the measured answer. Most repositories never produce one, so the graph answers instead: a test file that imports a source file reaches it, which is a recorded edge, where most tools fall back to matching file names.
repowise impacted-tests main..HEAD # only the tests this diff exercises
repowise health # untested hotspots, graph-aware
Checked against a real coverage run --contexts=test on this repository:
95.7% precision on what reaches a file and 97.5% on the run
list. Every row is stamped basis: "measured" or
"inferred", measured wins where both can answer, and an empty answer means
unknown, never "no tests".
Test intelligence →
Patch coverage, doc drift, security and change risk run as gates in your own pipeline through the GitHub Action, a GitLab template or plain CLI commands anywhere else, with annotations, SARIF and GitLab Code Quality output. The gates need no API key. Repowise in CI →
On GitHub you can also install the free Repowise PR Bot, a hosted GitHub App that puts the same analysis on every pull request. One comment, edited in place on every push, and a green PR gets no comment at all. It shows symbol-level blast radius (the contracts the PR changed and every caller outside the PR), the tests and co-change partners missing from the change, change risk against the repository's own history, and a public analysis page per PR. Zero LLM calls, so the same diff always gets the same review.
A real comment on a real PR: repowise-dev/repowise#1204 · its analysis page → · install the PR bot →
A score that says "this file is risky" is where most tools stop. Repowise scores every file, finds where the risk concentrates, and names the specific fix.
Every file is scored 1-10 by 53 deterministic detectors (McCabe complexity, brain methods, LCOM4 cohesion, god classes, clone detection, untested hotspots, change entropy, prior-defect history and more), read through three lenses: defect risk, maintainability and performance. Performance findings such as N+1 queries and I/O in loops are traced across functions and files through the call graph, which is where file-local linters lose them. Only 25 of the 53 detectors move the defect score, because that is the number carrying published accuracy claims.
Zero LLM calls, zero cloud. Detector weights are calibrated against a real defect corpus, not hand-tuned: every file scored at a commit before the bug window so nothing leaks backward, with file size as an explicit control, so a detector only earns weight for defect lift beyond a file being big.
It checks itself on your repository. After every index, Repowise compares its own flags with your git history and tells you what it found: "16 of the 20 lowest-health files had a bug fix in the last 6 months, 3.3x the 24% baseline." If that number is bad on your codebase, you will see it.
Then it names the fix. Extract Class, Extract Helper, Move Method, Break Cycle, Split File or Extract Method, with the exact methods, edges and symbols that move, the callers and co-changing files that move with them, and a ranking that puts a fix on a central hub above the same fix on a leaf. Extract Method runs a dataflow pass over the function to lift the exact span and infer a behaviour-preserving signature.
repowise next # what to fix first, ranked by impact and effort
repowise health # KPIs and lowest-scoring files
repowise health --refactoring-targets # ranked, concrete plans
repowise health --trend # snapshots plus declining-health alerts
repowise dead-code # what can go, by confidence tier
The dashboard renders each plan as a card with a copy-to-agent button. An optional model step, never in the indexing path, expands a plan into generated code and a unified diff.
Validated on 21 open-source repositories across 9 languages (2,826 files scored at a fixed point and checked against the following 6 months of bug fixes): ROC AUC 0.737 [0.683, 0.787]. Against CodeScene on the 2,770 files both tools scored, ranking by Repowise surfaces 2.3x the defects under a fixed review budget (paired, p = 0.003). CodeScene keeps a shorter, slightly more precise list. Full head-to-head and its limits →
Guides: code health · refactoring · dead code
repowise serve starts the full web dashboard next to the MCP server. No separate
setup, all local.
![]() Architecture · the dependency graph, laid out and explorable, with per-node context and change coupling | ![]() Code Health · every file as a bubble, hover any one to inspect its score, size, coverage and findings |
![]() Chat · ask the codebase a question, answers cite the files and pages they came from | ![]() Docs · generated wiki pages for the whole codebase, with confidence and freshness badges |
Also in there: Architecture and C4 views, the Knowledge Graph and a zoomable map, Risk, Hotspots, Coupling and Blast radius, Contributors and Ownership, Decisions with an evidence drawer and timeline, Symbols, Security, Dead code, Costs and Workspace. Every view and what it answers: docs/start/DASHBOARD.md →
Real systems are not one repository, and the expensive failures live in the gaps between them. Change a backend contract and Repowise can name the frontend calls that consume it, the services downstream, the companion files missing from the change, and the architecture rule the new dependency violates, before it ships.
| Workspace intelligence | What it answers |
|---|---|
| Contract map | Which services provide and consume each HTTP, gRPC, event, socket and data contract? Links keep their exact or candidate confidence and the source evidence. |
| Cross-repo blast radius | If this provider changes, which downstream services are in structural reach, and which may drift through historical co-change? |
| Breaking-change guard | Was an endpoint removed or an OpenAPI, proto or signature shape changed incompatibly, and which consumer files are linked to that contract? |
| Test impact | Which tests in the consumer repos should run for this provider change, measured from coverage or inferred from the call graph? |
| Architecture as code | Does the live system graph violate declared dependency rules or contain cycles? repowise workspace check gates CI. |
| Architecture health | How coupled is the estate? Propagation cost, the cyclic core, service roles and a deterministic 1-10 architecture score. |
| Federated context | One dashboard and one MCP server answer across every repository while keeping repo-level evidence. |
The system map models services, not just repository boxes, and never conflates a real contract with "these files often changed together". HTTP field-level comparison covers a bounded OpenAPI 3.x subset; matched consumers prove endpoint exposure, not field use or runtime failure. Workspace guide and exact support matrix →
Worktrees and updates stay light: a linked worktree seeds its index from the base checkout, and post-commit hooks, file watching, webhooks or polling keep each repository and the cross-repo graph current. Keeping the index fresh →
26 languages parsed to an AST, 40 on a five-rung ladder, framework-aware where an ecosystem handler exists. Every language ships in the open-source distribution.
Full
Good
· Partial
| Rung | Languages | What you get |
|---|---|---|
| Full (13) | Python · TypeScript · JavaScript · Svelte · Vue · Java · Kotlin · Go · Rust · C++ · C# · Scala · Ruby | The whole pipeline: AST symbols, import resolution, a resolved call graph, heritage, docstrings, framework edges and code-health markers |
| Good (11) | C · Swift · PHP · Dart · Object Pascal · COBOL · GDScript · VB.NET · Elixir · F# · Objective-C | All of the above except the full health suite, within the language-specific ceilings in the full matrix |
| Partial (2) | Luau / Roblox · Razor / Blazor | Luau: AST symbols and require() resolution, Rojo and .luaurc aware. Razor: component symbols, @code and component-tag call edges, C# health markers; no import resolution yet |
| Lightweight (6) | Clojure · Haskell · Lean 4 · Erlang · HTML · QML | A real file-to-file import graph, and no symbol-level claims |
| Structural (8) | R · Zig · Julia · Elm · OCaml · Crystal · Nim · D | Git history: blame, hotspots, co-change, ownership, bug history |
SQL and dbt projects get ref() / source() lineage, shell scripts get function-level
symbols, HTML pages contribute their <script src> and <link href> dependencies, and
OpenAPI, Protobuf, GraphQL, Dockerfile, Terraform and similar formats get dedicated
handlers. Anything else is still tracked through git history.
Every call edge is stamped with how it was resolved and how much to trust it, from
same_file at 0.95 down to a repo-wide name match at 0.50, labelled as the guess it
is. Accuracy per language, graded against each language's own compiler:
docs/BENCHMARKS.md#accuracy-by-language.
Full matrix: docs/layers/LANGUAGE_SUPPORT.md → · adding a language takes five small steps and no changes to the parser core: docs/architecture/language-support.md → · languages moving up the ladder: roadmap →
Six agents wired end to end · two at the Full tier · every other MCP host one paste away.
Full tier
Good tier
Full is every surface Repowise has: MCP tools, skills, slash commands, a managed
instructions file, hook-level interception of tool calls, and transcript mining after
the session. Good is MCP tools and the config to reach them, without hooks or
transcript mining. Anything else that speaks MCP is one snippet away:
repowise agents print-config claude-code prints a server entry for Cline, Windsurf,
Zed, Gemini CLI or any host that reads mcpServers.
Integration matrix →
In VS Code, the Repowise extension shows what your change breaks before you push (riskiest files, what is downstream, forgotten companion files, missing tests, suggested reviewers), health in the gutter and status bar, callers and ownership on hover, and refactoring plans as CodeLens. One install also registers the MCP server, so the same index serves you and your agent. Install from the Marketplace or Open VSX and run Repowise: Set Up This Repository. VS Code guide →
In Claude Code, Lens ships inside the Repowise plugin and shows you what the index
knows while Claude works: the file it is on and how many files depend on it, a change
review after a turn that edits files (code health, tests to run, other branches on the
same files), and /lens, a pane with Flow (what to check before accepting each turn),
a map of the repo lit by what Claude searched, opened and edited and what its edits
reach, Ask, and a session recap. Nothing reaches Claude unless you press a button, and
Lens makes no model calls of its own. It needs Claude Code 2.1.287 or later, and the
map needs repowise serve --no-ui running. Lens guide →
Every response carries a _meta envelope with the indexed commit, the index age and a
stale warning when the index has fallen behind your checkout, so your agent always
knows how much to trust what it just read.
| Tool | What only this tool answers |
|---|---|
get_overview() | Architecture summary, module map, entry points, git health. The first call on an unfamiliar codebase. |
get_answer(question) | Hybrid retrieval (full-text plus vector), graph expansion and one cited answer with a calibrated retrieval_quality. Search, read and reason in a single round-trip. |
get_context(targets, include?) | Triage card for files, modules or symbols: summary, signatures, hotspot flag, governing decisions, symbol ids. include opens callers, callees, ownership and metrics. Batch many targets in one call. |
get_symbol("file.py::Name") | One indexed symbol's source with exact line bounds. |
search_codebase(query) | Hybrid search over code and docs, by symbol, path or concept. |
get_risk(targets?, changed_files?) | Hotspots, dependents, co-change partners, ownership, test gaps, bug history. Pass changed_files for PR mode and get a directive back. |
get_change_risk(revspec) | What a commit, range or uncommitted change made worse across defect risk, maintainability and performance, the tests that touch it, and how the diff ranks against recent commits. |
get_why(query?, targets?) | Architectural decisions with their verbatim evidence. Falls back to git archaeology when no decision exists. |
get_dead_code(...) | Unreachable code by confidence tier, with cross-repo consumers in workspace mode. |
get_health(targets?, include?) | Health scores and findings across all three lenses, Fix first, coverage, trends, doc drift and refactoring plans. |
Ten is a deliberate ceiling: a small, task-shaped surface is easier for an agent to choose from than a large one. Seven more tools (dependency paths, execution flows, refactoring code generation, finding triage, and three workspace architecture tools) are opt-in. Parameters, examples and when to use which: docs/agent/MCP_TOOLS.md →
Open-source agent-context tools, the same repositories, the same pinned commits, the same questions, each tool given its full advertised surface. The full page carries the rows we lose beside the rows we win.
The full results, the methodology, and the rows we lose → · The research it is built on →
No single product competes with all of this, so there is no single table. Rows marked measured are head-to-head numbers that link to docs/BENCHMARKS.md, where the sample sizes, the tests and the rows we lose live. Unmarked rows are capability presence, not measurements.
| repowise | CodeGraph | Serena | DeepWiki | |
|---|---|---|---|---|
| Self-hostable, open source | ✅ AGPL-3.0 | ✅ | ✅ | ❌ cloud only |
| Private repo, no cloud | ✅ | ✅ | ✅ | ❌ OSS forks only |
| MCP tools served | 10 default + 7 opt-in | 1 | 29 | 3 |
| Finds the gold files (measured, n=42 sealed) | ✅ 0.876 | 0.610 | not in this run | not measured |
| Output tokens vs a bare agent (measured, n=43) | ✅ -31.6% | -24.4% | -14.8% | not measured |
| Memory to build the graph (measured, 5 tools, 35 repos) | ✅ 75 MB, lowest on 35 of 35 | 757 MB | not measured | n/a, cloud |
| Time to build the graph (measured, same run) | 2.77s, fastest on 14 of 35 | 3.65s, fastest on 16 | not measured | n/a, cloud |
| Time to build the full index, django (measured) | ⚠️ 366.8s, slowest here | ✅ 16.4s | not measured | n/a, cloud |
| Call-edge precision, hand-graded (measured, 560 rows, 9 languages) | ✅ 85.7% | 58.6% | not measured | not measured |
| Call-edge precision, judged by a compiler (measured, 5 tools, 7 cells) | ✅ nothing that finds as much gets more of it right, 7 of 7 | lower precision in 7 | not measured | not measured |
| Generated documentation | ✅ | ❌ | ❌ | ✅ |
| Proactive agent hooks | ✅ Claude + Codex | ❌ | ❌ | ❌ |
Generated CLAUDE.md / AGENTS.md | ✅ | ❌ | ❌ | ❌ |
| Command-output distillation | ✅ reversible | ❌ | ❌ | ❌ |
| Architectural decision records | ✅ | ❌ | ❌ | ❌ |
| Multi-repo workspace intelligence | ✅ contracts, co-change, federated MCP | ❌ | ❌ | ❌ |
The two cost rows answer different questions. Building the call graph, we are the lightest tool measured, about ten times lighter than the next, and roughly as fast as the fastest. Building the whole index, CodeGraph is 22x faster than we are, because by then we have also built the git-history layer, the wiki, the decisions and the health pass. If a call graph is all you need, that is the right trade and you should take it. With model-written prose on, it is 135x.
The hand-graded precision row cuts both ways. About fourteen percent of the call
edges we draw are wrong, and on seastar CodeGraph grades better than we do. The
compiler row exists because we graded the hand-read one ourselves: on Go and
TypeScript the answer key is the Go team's own call graph and the tsc checker's own
resolution, which we neither wrote nor can tune. Precision alone is easy to win by
drawing almost nothing, and recall alone by drawing everything, so the claim is the
pair.
Competitors measured at CodeGraph 1.5.0, Graphify 0.9.31, Serena 1.6.2.dev0 and
code-review-graph 2.3.7 in August 2026. Repowise compiler-graded figures are from main
on 2026-10-03; the other Repowise rows are from 081a59fa, August 2026.
| repowise | CodeScene | |
|---|---|---|
| Self-hostable, open source | ✅ AGPL-3.0 | ⚠️ on-prem Docker, proprietary |
| Code health score (1-10) | ✅ 53 detectors, 25 scoring | ✅ 25-30 |
| Brain Method / LCOM4 / god class | ✅ | ✅ |
| Defects found at a 20% review budget (measured, 2,770 files) | ✅ 0.173 | 0.074 |
| Effort-aware ranking, Popt (measured, p=0.003) | ✅ 0.607 | 0.462 |
| Precision at that budget (measured) | 0.580 | ✅ 0.636, a shorter list |
| Discrimination, ROC AUC (measured, paired) | 0.731 | 0.705, p=0.054, not significant |
| Business impact (resolution time) | ❌ we could not replicate this on open data | ✅ Code Red study |
| Git intelligence (hotspots, ownership, co-change) | ✅ | ✅ |
| Pre-merge change-risk scoring | ✅ 0-10 + directives | ✅ |
| Concrete cross-file refactoring plans | ✅ graph-aware + blast radius | ⚠️ within-function only |
| Test-coverage intelligence | ✅ LCOV/Cobertura/Clover/JaCoCo/Go | ❌ |
| Dead code detection | ✅ | ❌ |
| Serves it to an AI agent over MCP | ✅ | ✅ |
CodeScene is the only other vendor in this category with a published empirical defect study, which is why it is the one we ran head to head. It flags about 27 files where we flag 132, so if you want a short list to act on, its threshold is the better fit.
DeepWiki, Google Code Wiki and Swimm generate documentation from a repository, which overlaps one of our layers. We have not measured against them, so there is no table here.
| Repowise PR Bot | CodeRabbit | Greptile | |
|---|---|---|---|
| LLM calls per PR | ✅ zero | ❌ every review | ❌ every review |
| Same diff, same review | ✅ deterministic | ❌ sampled output | ❌ sampled output |
| Your code sent to a model provider | ✅ never | ❌ yes | ❌ yes |
| Symbol-level blast radius | ✅ call graph | ❌ | ⚠️ prose, from context |
| Co-change partners missing from the PR | ✅ git history | ❌ | ❌ |
| Change risk vs the repo's own history | ✅ 0-10 + percentile | ❌ | ❌ |
| Silent on a clean PR | ✅ by default | ⚠️ configurable | ⚠️ configurable |
An LLM reviewer is a different product: it can read intent, and it can be wrong in a new way on every run. This one does set arithmetic over a call graph and a git history, so pushing the same diff twice gives the same review twice. Full side-by-side comparisons: repowise.dev/compare →
AI makes producing a change cheaper. It does not make understanding its consequences cheaper. In a large estate that answer crosses repositories, ownership boundaries, service contracts, test suites and years of architectural history. Repowise gives developers, agents, reviewers and platform teams the same evidence about what exists, what depends on it, what is risky and what will break, from one index, in place of a separate health tool, dead-code tool, code-search layer for agents, docs generator and test-impact service.
| Deployment | Where source is read | Where the index lives | What leaves your network |
|---|---|---|---|
Self-hosted, open source (pip install repowise) | your machine | .repowise/ next to the repo | anonymous CLI telemetry you can switch off, and your own LLM provider only if you turn prose on |
| Self-hosted, commercial | your VPC or an air-gapped network | Postgres plus LanceDB or pgvector, inside your network | the same, plus integrations you configure. Nothing at all in air-gapped mode |
| Hosted (repowise.dev) | the hosted indexer | infrastructure we operate | your code goes to the platform |
Graph, git, health, change risk, tests, dead code and PR review make zero LLM calls. Raw source is parsed in memory and never persisted. Prose is optional and runs on your own provider contract or fully offline through Ollama, chosen per repository. The threat model and data flows are in the security review pack. The published research each layer is built on, and how we checked it, is in LINEAGE.md.
Call-graph accuracy is graded against each language's own compiler toolchain, and we publish precision and recall per language, with the misses, in the accuracy table. On current main, precision runs 0.95 to 0.99 for Go, TypeScript, C and most Python, 0.85 to 0.90 for Rust, and 0.73 to 0.92 for Java and C#, where overloaded methods are the main gap.
The largest repository indexed so far is dotnet/runtime: 58,924 files in 120 minutes at 11.7 GiB peak memory, on one 31 GB laptop with no API key. Time and memory by repository size: scale.
Any git host works, because indexing reads a local checkout. On top of that:
repowise init at a plain directory, an export, or a Perforce
or SVN workspace. The graph, docs, decisions and health layers build normally; the
history layer needs a commit log, and Perforce and SVN adapters are
on the roadmap.| Status | Capability |
|---|---|
| Shipping in open source | Every deterministic layer, ten MCP tools, multi-repo workspaces, contract extraction and cross-repo blast radius, test intelligence, architecture conformance, local dashboard, auto-sync, full-history secret scanning. |
| GA commercially | Hosted graph-aware security, CVE prioritization, CycloneDX SBOM and VEX, PCI-DSS and SOC 2 control-coverage reports, audit export and webhook stream, Jira and Confluence, reference HA topology on your infrastructure, custom language extensions, SLA support, IP indemnification. |
| Rolling out | Managed GitHub Enterprise, Azure DevOps, GitLab and Bitbucket integrations. SAML/OIDC SSO and SCIM. Engineering-leader dashboards. |
| Planned | RBAC and multi-tenancy, a packaged air-gap install bundle, the Helm chart. |
Repowise holds no SOC 2, ISO 27001 or other audited certification today. The SOC 2 and PCI-DSS reports above are control-coverage signals from your own findings, and every export says so. Every item is tracked with its status in the capability matrix.
Commercial licences are priced per indexed repository, with unlimited seats inside the licensed set. More and more of the code in a repository is written and read by agents and CI, so seat counts stop tracking the value. Details →
Do we have to open-source our code because of the AGPL? No. Using Repowise inside your company, including on internal servers, creates no obligation to publish your code. The obligations apply if you modify Repowise and offer it to others over a network, or ship it inside your own product. A commercial licence removes them.
Commercial support comes with a named contact, a response-time SLA and a quarterly architecture review. Security fixes land on the latest minor release; reports go to security@repowise.dev under the disclosure policy.
Commercial detail · Security review pack · Roadmap · hello@repowise.dev
repowise.dev runs the same engine fully managed. We run it on our own codebase in the open: live snapshot · explore public repos.
Already indexed locally? repowise publish puts the same repo on repowise.dev in one
command: it asks repowise.dev to index the repo's GitHub remote, so nothing is uploaded
from your machine and only what you have pushed is published. A hosted index usually
takes about 10 minutes. Public repos publish on a free account (up to 2 repos, no card);
private repos and more repos need Pro, free for 10 days with a card.
What hosted adds →
--no-prose, code-derived content stays on
your infrastructure. The CLI reports anonymous, opt-out usage telemetry (command
names and coarse environment only); turn it off with repowise telemetry disable,
DO_NOT_TRACK=1, or by running fully offline.
What's collected →Doing a security review? docs/business/SECURITY_COMPLIANCE.md →
repowise init [PATH] # index a codebase (asks; --yes --no-prose needs no key)
repowise generate [PATH] # write wiki pages with a model, on demand
repowise serve [PATH] # MCP server + local dashboard
repowise update [PATH] # incremental update (--workspace for every repo)
repowise watch # re-index on file change
repowise search "<q>" # hybrid search (fulltext / semantic / symbol / path)
repowise ask "<q>" # a synthesized answer with citations
repowise context <files> # triage card: layer, hotspot, fix history, freshness
repowise symbol <id> # one symbol's body, with verified line bounds
repowise why <q|path> # decisions, rationale, git archaeology
repowise next # what to fix first
repowise health # code-health KPIs and lowest-scoring files
repowise risk main..HEAD # score a branch or PR range
repowise overlap # other branches editing the same files
repowise impacted-tests # only the tests a diff exercises
repowise dead-code # unreachable code by confidence tier
repowise doc-drift # documentation the code no longer supports
repowise security # secrets and risky patterns, tree or full history
repowise decision list # architectural decisions
repowise export --format structurizr # the architecture as Structurizr DSL
repowise distill pytest # compact, errors-first, reversible command output
repowise saved # tokens and dollars saved by distillation
repowise workspace add # multi-repo workspace management
repowise doctor # check setup, API keys, index drift
repowise publish # put this repo on repowise.dev (indexed from GitHub; public repos free)
repowise uninstall # remove what repowise wrote, and say what it left
Every command and flag: docs/reference/CLI_REFERENCE.md · config: docs/reference/CONFIG.md · something not working: docs/start/TROUBLESHOOTING.md · all docs: docs/README.md
git clone https://github.com/repowise-dev/repowise
cd repowise
uv sync --all-packages
uv run repowise --version
uv run pytest tests/unit/
New here? You do not have to read 3,000 files to start. We keep a public index of this repository built by Repowise itself, re-indexed on every push: explore repowise with repowise → (architecture, hotspots, ownership, decisions, and a ranked refactoring backlog you are welcome to pick from).
Full guide, including how to add languages and LLM providers: CONTRIBUTING.md · architecture: docs/architecture/
AGPL-3.0. Free for individuals, teams and companies using Repowise internally.
For commercial licensing (the enterprise security and compliance layer, SSO/SCIM, RBAC, workflow integrations, priority support and SLA, or embedding Repowise in a product without AGPL obligations), see docs/business/COMMERCIAL.md or contact hello@repowise.dev.
Built for engineers who got tired of watching their AI agent cat the same file for the fourth time.
⭐ If Repowise earns a place in your workflow, give it a star. It costs you nothing, and it's the signal that keeps a small team building this in the open.
repowise.dev · Explore → · Discord · X · hello@repowise.dev
(top 24 of 59)
602 followers · starred Aug 2026
822 followers · starred May 2026
569 followers · starred Apr 2026
336 followers · starred Apr 2026