mupt-ai/self-bench

Benchmark coding agents and models on your own repository with private Harbor-formatted evals built from its merged pull requests.

TypeScript

28

792 commits

updated Oct 5, 2026

See the code

See what people are saying

README

SelfBench

SelfBench

Find the best models for your repo.
Private coding-agent benchmarks built from your repository's own merged pull requests.

Leaderboards at selfbench.dev Benchmark your repo at app.selfbench.dev

CI License

Browsing selfbench.dev: searching for a repository, opening its accuracy vs cost chart, and reading every model setting's score

Public benchmarks tell you how a model does on someone else's code. SelfBench tells you how it does on yours: it turns your merged PRs into tasks with hidden tests, runs agents and models on them, and plots accuracy against cost.

Results

Every open-source repository released on selfbench.dev gets a live leaderboard: each model, harness, and reasoning setting placed by accuracy and cost per task, with the Pareto frontier drawn through the settings nothing else beats on both.

Browse the leaderboards: vercel/next.js · supabase/supabase · earendil-works/pi · getsentry/sentry · PostHog/posthog · pingdotgg/t3code · vercel/vercel · all repositories →

How it works

How SelfBench works: a merged PR is rebuilt from the commit before the change into an instruction, hidden tests, and a reference solution; Harbor's smoke, nop, oracle, and determinism gates and an independent review accept it; agents and models attempt every task; and accuracy is plotted against cost with the Pareto frontier

For each merged PR, SelfBench rebuilds the task from the commit before the change: the PR's own request becomes the instruction, and an authoring agent writes hidden tests and a reference solution. A task is accepted only if the tests fail without a solution, pass with the original implementation, pass again on a rerun, and survive an independent review. Every accepted task is a native Harbor task.

Using SelfBench

Everything happens in the web app at app.selfbench.dev:

  1. Sign in with GitHub and connect a repository.
  2. Batch Generation: choose how many easy, medium, and hard tasks to build, or use Add PRs on the Dataset page to build one task from each pull request you pick. A batch takes hours; it keeps running after you close the page.
  3. Dataset: inspect each task (instruction, environment, hidden tests, reference patch, pipeline artifacts) and approve or reject it.
  4. Run: pick models, harnesses, and a sandbox, and run them on the approved tasks.
  5. Results: compare accuracy against cost, and open any trial's transcript and scores.
  6. Releases: publish a public repository's results to selfbench.dev.

Models and sandboxes run on your organization's own keys under Credentials. Read released results through the Public Results API, or automate your workspace with an API key.

Self-hosting

SelfBench's reference deployment runs on GCP: Cloud Run serves the API, while GKE Autopilot runs the Temporal workflow worker and KEDA-scaled Harbor jobs, with Cloud SQL and GCS. A Cloud Run worker pool remains available as an alternative.

  1. Provision a GCP project, billing, Terraform state bucket, and GitHub Actions Workload Identity Federation.
  2. Configure Terraform inputs and store each runtime secret value in its own Secret Manager secret.
  3. Apply the environment with Terraform, or configure the protected GitHub dev and prod environments to deploy through Actions.
  4. Point your domain at the provisioned load balancer and configure GitHub OAuth for the app URL.

For prerequisites, exact Terraform commands, runtime configuration, GitHub Actions setup, and GKE worker setup, see the self-hosting and infrastructure guide.

Development

Requires Bun 1.3.14+ and Docker with Compose.

bun install --frozen-lockfile
bun run validate

Run the whole stack (API serving the app, worker, Temporal, Postgres, local Docker sandboxes) from a checkout:

cp .env.example .env   # GitHub OAuth app, session secret, credential key, managed keys
SELFBENCH_PUBLIC_URL=https://your-tunnel.example docker compose --profile sandbox up -d --build
docker compose port api 8080

Compose names the project after the checkout directory and publishes ephemeral host ports, so worktrees run side by side. Set SELFBENCH_PUBLIC_URL to the origin the browser opens (a tunnel or reverse proxy) before up, and register <origin>/auth/github/callback on the OAuth app.

For frontend work, bun run dev:site runs the API and Vite with hot reload (secrets in .env.site; see scripts/dev-site.ts).

Documentation

License

MIT © 2026 Mupt AI.

agent-evaluation
ai-agents
benchmark
coding-agents
evals
harbor
llm
llm-evaluation
swe-bench

Significant stargazers

Eugene Klimov

234 followers · starred Aug 2026

Andrew Qu

511 followers · starred Aug 2026

Aleksandr Filippov

12 followers · starred Aug 2026

mupt-ai/self-bench

Benchmark coding agents and models on your own repository with private Harbor-formatted evals built from its merged pull requests.

TypeScript

28

792 commits

updated Oct 5, 2026

See the code

See what people are saying

README

SelfBench

SelfBench

Find the best models for your repo.
Private coding-agent benchmarks built from your repository's own merged pull requests.

Leaderboards at selfbench.dev Benchmark your repo at app.selfbench.dev

CI License

Browsing selfbench.dev: searching for a repository, opening its accuracy vs cost chart, and reading every model setting's score

Public benchmarks tell you how a model does on someone else's code. SelfBench tells you how it does on yours: it turns your merged PRs into tasks with hidden tests, runs agents and models on them, and plots accuracy against cost.

Results

Every open-source repository released on selfbench.dev gets a live leaderboard: each model, harness, and reasoning setting placed by accuracy and cost per task, with the Pareto frontier drawn through the settings nothing else beats on both.

Browse the leaderboards: vercel/next.js · supabase/supabase · earendil-works/pi · getsentry/sentry · PostHog/posthog · pingdotgg/t3code · vercel/vercel · all repositories →

How it works

How SelfBench works: a merged PR is rebuilt from the commit before the change into an instruction, hidden tests, and a reference solution; Harbor's smoke, nop, oracle, and determinism gates and an independent review accept it; agents and models attempt every task; and accuracy is plotted against cost with the Pareto frontier

For each merged PR, SelfBench rebuilds the task from the commit before the change: the PR's own request becomes the instruction, and an authoring agent writes hidden tests and a reference solution. A task is accepted only if the tests fail without a solution, pass with the original implementation, pass again on a rerun, and survive an independent review. Every accepted task is a native Harbor task.

Using SelfBench

Everything happens in the web app at app.selfbench.dev:

  1. Sign in with GitHub and connect a repository.
  2. Batch Generation: choose how many easy, medium, and hard tasks to build, or use Add PRs on the Dataset page to build one task from each pull request you pick. A batch takes hours; it keeps running after you close the page.
  3. Dataset: inspect each task (instruction, environment, hidden tests, reference patch, pipeline artifacts) and approve or reject it.
  4. Run: pick models, harnesses, and a sandbox, and run them on the approved tasks.
  5. Results: compare accuracy against cost, and open any trial's transcript and scores.
  6. Releases: publish a public repository's results to selfbench.dev.

Models and sandboxes run on your organization's own keys under Credentials. Read released results through the Public Results API, or automate your workspace with an API key.

Self-hosting

SelfBench's reference deployment runs on GCP: Cloud Run serves the API, while GKE Autopilot runs the Temporal workflow worker and KEDA-scaled Harbor jobs, with Cloud SQL and GCS. A Cloud Run worker pool remains available as an alternative.

  1. Provision a GCP project, billing, Terraform state bucket, and GitHub Actions Workload Identity Federation.
  2. Configure Terraform inputs and store each runtime secret value in its own Secret Manager secret.
  3. Apply the environment with Terraform, or configure the protected GitHub dev and prod environments to deploy through Actions.
  4. Point your domain at the provisioned load balancer and configure GitHub OAuth for the app URL.

For prerequisites, exact Terraform commands, runtime configuration, GitHub Actions setup, and GKE worker setup, see the self-hosting and infrastructure guide.

Development

Requires Bun 1.3.14+ and Docker with Compose.

bun install --frozen-lockfile
bun run validate

Run the whole stack (API serving the app, worker, Temporal, Postgres, local Docker sandboxes) from a checkout:

cp .env.example .env   # GitHub OAuth app, session secret, credential key, managed keys
SELFBENCH_PUBLIC_URL=https://your-tunnel.example docker compose --profile sandbox up -d --build
docker compose port api 8080

Compose names the project after the checkout directory and publishes ephemeral host ports, so worktrees run side by side. Set SELFBENCH_PUBLIC_URL to the origin the browser opens (a tunnel or reverse proxy) before up, and register <origin>/auth/github/callback on the OAuth app.

For frontend work, bun run dev:site runs the API and Vite with hot reload (secrets in .env.site; see scripts/dev-site.ts).

Documentation

License

MIT © 2026 Mupt AI.

agent-evaluation
ai-agents
benchmark
coding-agents
evals
harbor
llm
llm-evaluation
swe-bench

Significant stargazers

Eugene Klimov

234 followers · starred Aug 2026

Andrew Qu

511 followers · starred Aug 2026

Aleksandr Filippov

12 followers · starred Aug 2026