Benchmark coding agents and models on your own repository with private Harbor-formatted evals built from its merged pull requests.
See the code
Find the best models for your repo.
Private coding-agent benchmarks built from your repository's own merged pull requests.
Public benchmarks tell you how a model does on someone else's code. SelfBench tells you how it does on yours: it turns your merged PRs into tasks with hidden tests, runs agents and models on them, and plots accuracy against cost.
Every open-source repository released on selfbench.dev gets a live leaderboard: each model, harness, and reasoning setting placed by accuracy and cost per task, with the Pareto frontier drawn through the settings nothing else beats on both.
Browse the leaderboards: vercel/next.js · supabase/supabase · earendil-works/pi · getsentry/sentry · PostHog/posthog · pingdotgg/t3code · vercel/vercel · all repositories →
For each merged PR, SelfBench rebuilds the task from the commit before the change: the PR's own request becomes the instruction, and an authoring agent writes hidden tests and a reference solution. A task is accepted only if the tests fail without a solution, pass with the original implementation, pass again on a rerun, and survive an independent review. Every accepted task is a native Harbor task.
Everything happens in the web app at app.selfbench.dev:
Models and sandboxes run on your organization's own keys under Credentials. Read released results through the Public Results API, or automate your workspace with an API key.
SelfBench's reference deployment runs on GCP: Cloud Run serves the API, while GKE Autopilot runs the Temporal workflow worker and KEDA-scaled Harbor jobs, with Cloud SQL and GCS. A Cloud Run worker pool remains available as an alternative.
dev and prod environments to deploy through Actions.For prerequisites, exact Terraform commands, runtime configuration, GitHub Actions setup, and GKE worker setup, see the self-hosting and infrastructure guide.
Requires Bun 1.3.14+ and Docker with Compose.
bun install --frozen-lockfile
bun run validate
Run the whole stack (API serving the app, worker, Temporal, Postgres, local Docker sandboxes) from a checkout:
cp .env.example .env # GitHub OAuth app, session secret, credential key, managed keys
SELFBENCH_PUBLIC_URL=https://your-tunnel.example docker compose --profile sandbox up -d --build
docker compose port api 8080
Compose names the project after the checkout directory and publishes ephemeral host ports, so worktrees run side by side. Set SELFBENCH_PUBLIC_URL to the origin the browser opens (a tunnel or reverse proxy) before up, and register <origin>/auth/github/callback on the OAuth app.
For frontend work, bun run dev:site runs the API and Vite with hot reload (secrets in .env.site; see scripts/dev-site.ts).
cd docs && npm ci && npm run dev)MIT © 2026 Mupt AI.
234 followers · starred Aug 2026
511 followers · starred Aug 2026
12 followers · starred Aug 2026
Benchmark coding agents and models on your own repository with private Harbor-formatted evals built from its merged pull requests.
See the code
Find the best models for your repo.
Private coding-agent benchmarks built from your repository's own merged pull requests.
Public benchmarks tell you how a model does on someone else's code. SelfBench tells you how it does on yours: it turns your merged PRs into tasks with hidden tests, runs agents and models on them, and plots accuracy against cost.
Every open-source repository released on selfbench.dev gets a live leaderboard: each model, harness, and reasoning setting placed by accuracy and cost per task, with the Pareto frontier drawn through the settings nothing else beats on both.
Browse the leaderboards: vercel/next.js · supabase/supabase · earendil-works/pi · getsentry/sentry · PostHog/posthog · pingdotgg/t3code · vercel/vercel · all repositories →
For each merged PR, SelfBench rebuilds the task from the commit before the change: the PR's own request becomes the instruction, and an authoring agent writes hidden tests and a reference solution. A task is accepted only if the tests fail without a solution, pass with the original implementation, pass again on a rerun, and survive an independent review. Every accepted task is a native Harbor task.
Everything happens in the web app at app.selfbench.dev:
Models and sandboxes run on your organization's own keys under Credentials. Read released results through the Public Results API, or automate your workspace with an API key.
SelfBench's reference deployment runs on GCP: Cloud Run serves the API, while GKE Autopilot runs the Temporal workflow worker and KEDA-scaled Harbor jobs, with Cloud SQL and GCS. A Cloud Run worker pool remains available as an alternative.
dev and prod environments to deploy through Actions.For prerequisites, exact Terraform commands, runtime configuration, GitHub Actions setup, and GKE worker setup, see the self-hosting and infrastructure guide.
Requires Bun 1.3.14+ and Docker with Compose.
bun install --frozen-lockfile
bun run validate
Run the whole stack (API serving the app, worker, Temporal, Postgres, local Docker sandboxes) from a checkout:
cp .env.example .env # GitHub OAuth app, session secret, credential key, managed keys
SELFBENCH_PUBLIC_URL=https://your-tunnel.example docker compose --profile sandbox up -d --build
docker compose port api 8080
Compose names the project after the checkout directory and publishes ephemeral host ports, so worktrees run side by side. Set SELFBENCH_PUBLIC_URL to the origin the browser opens (a tunnel or reverse proxy) before up, and register <origin>/auth/github/callback on the OAuth app.
For frontend work, bun run dev:site runs the API and Vite with hot reload (secrets in .env.site; see scripts/dev-site.ts).
cd docs && npm ci && npm run dev)MIT © 2026 Mupt AI.
234 followers · starred Aug 2026
511 followers · starred Aug 2026
12 followers · starred Aug 2026