Coding agent CLI for local and hosted models, designed around small local models first
See the code
A coding agent CLI for local and hosted models,
designed around small local models first.
Website · Quick start · Docs · Tested models · 简体中文

A real local run of Qwen3.6 35B A3B (Unsloth UD-IQ2_M) via llama.cpp, sped up 3×. The model's first test fails and it fixes it from the error.
Reika is a coding agent CLI for local and hosted models. It was designed around small local models first (8B–35B, often at Q2–Q4, on a 16–32k context window), so it is careful with context, fails more gracefully, and says when it's stuck. It doesn't make a small model smarter. It makes working with one less frustrating: less of the window wasted, fewer loops on the same read, and less chance of the task quietly getting lost.
Day to day it runs on large hosted models, and small local ones are where it gets tested. At the small end, expect a better experience rather than frontier results. In testing, 14–35B models handled multi-file tasks most reliably, 8–9B models managed simple, well-scoped ones, and below 7B rarely got far.
Reika is pre-1.0, so commands, settings and behavior can still change between releases.
The harness's design choices were measured rather than assumed, and the record is published. It measures what a harness can and can't do about the problems small models run into, not how good the models become. The short version, with the runs and the numbers behind it in docs/findings.md:
/issue and /review ship with Reika and appear wherever gh and a GitHub remote are.REIKA_VISION_MODEL to read pasted images with a vision model instead). Windows may work under WSL2.Install it with any of these (npm and pnpm need Node ≥ 22; Homebrew pulls Node in):
npm i -g @alexwkleung/reika
pnpm add -g @alexwkleung/reika
brew install alexwkleung/tap/reika
Serve a model. Example below with llama.cpp:
llama-server -m <model.gguf> -c 24576 --jinja <other-launch-args>
Then point Reika at it:
export REIKA_MODEL=model # or put it in ~/.config/reika/.env
cd your-project && reika
Using a hosted API instead? Skip the server and set REIKA_BASE_URL, REIKA_API_KEY and REIKA_MODEL to the provider's endpoint, your key and its model id; Install and first run has the details.
REIKA_BASE_URL defaults to http://localhost:8080/v1, llama-server's default. The context window is read from the server when it reports one, and for hosted endpoints from the models.dev catalog; set REIKA_CONTEXT_WINDOW when neither has it. Reika sends no sampling parameters of its own, so your server's flags are what apply.
To run from a checkout instead — git clone https://github.com/alexwkleung/reika.git && cd reika, then pnpm install (npm users: npm i -g pnpm) and pnpm run install:global to build and install the reika binary, or cp .env.example .env and pnpm run dev to run without installing. pnpm run uninstall:global removes the global binary.
Set these in your shell, a project .env, or ~/.config/reika/.env (in that order of precedence). The full list, experimental flags, and multi-model profiles are in docs/configuration.md.
| Key | Default | What |
|---|---|---|
REIKA_MODEL | required | Model name, or a comma-separated list served by the same endpoint |
REIKA_BASE_URL | http://localhost:8080/v1 | OpenAI-compatible endpoint |
REIKA_API_KEY | no-key | API key (any non-empty value for local servers) |
REIKA_CONTEXT_WINDOW | probed from the server | Context window in tokens; drives compaction, the context gauge, and the repo map size |
REIKA_MIN_GEN_TOKENS | learned (from 2048) | Room reserved for the reply; learned from the model's rounds, set it to pin |
REIKA_AUTO_APPROVE | safe | off confirms every edit and command; bypass confirms nothing |
Shift+Tab cycles modes, or switch with a slash command. The mode and model you end a session in are the ones the next session opens in.
| Mode | What it does |
|---|---|
agent | The default: the model reads, edits, and runs commands |
plan | Read-only exploration that ends in a written plan; /implement runs it |
vibe | Plans first, then implements the plan, on every prompt |
minimal | Shell only, with no repo map or project context loaded upfront |
grind | Agent turns run a fixed, careful procedure: define done, test, review |
chat | Plain conversation with a separate history; web tools only |
shell | Your input runs as a shell command, and the output joins the agent's context |
reika -p "<prompt>" runs a single turn headless and prints the reply, for scripts and other harnesses. See docs/usage.md for modes in detail, headless flags, and every slash command.
safe: ordinary edits and commands run on their own, but anything matching a dangerous pattern (rm -rf, force pushes, package installs, curl, …) and any write outside the project still asks you first. bypass can only be set at launch, never mid-session.git/gh. A command you explicitly approve runs unsandboxed. REIKA_SANDBOX=0 turns it off.Found a way around one of these? Report it privately; see SECURITY.md.
The same pages are on the website, reikacode.com.
| Doc | Covers |
|---|---|
| Install and first run | Requirements, installing, serving a model, and the first run |
| Configuration | Every .env key, experimental flags, named profiles, last-session state |
| Modes, headless, commands | Each mode in depth, reika -p, slash commands, /save transcripts |
| Tools | The model's tools, approval prompts, web search setup, MCP servers, .gitignore |
| Instructions and skills | AGENTS.md, skills as slash commands, plain-English routing, pasted URLs |
| Tested models | Local quants and APIs Reika has been run against |
| Findings | The measured record: what broke, what held, and how each number was obtained |
| Platforms | Requirements, running on a weak machine, what differs on Linux and Windows |
| Architecture and caveats | How the harness works, and known limitations |
| Contributing | Scripts, design philosophy, a pointer to AGENTS.md, and external contributor guidelines |
Claude Code, Codex, Crush, OpenCode, Pi, Aider, DeepSeek Harness, Qwen Code, Kimi Code CLI, Gemini CLI, Junie, and DS4.
Apache License 2.0. See LICENSE and NOTICE.
If Reika is useful to you, consider supporting it via GitHub Sponsors, Ko-fi, or Buy Me a Coffee. See Support for what it funds and other ways to help.
Coding agent CLI for local and hosted models, designed around small local models first
See the code
A coding agent CLI for local and hosted models,
designed around small local models first.
Website · Quick start · Docs · Tested models · 简体中文

A real local run of Qwen3.6 35B A3B (Unsloth UD-IQ2_M) via llama.cpp, sped up 3×. The model's first test fails and it fixes it from the error.
Reika is a coding agent CLI for local and hosted models. It was designed around small local models first (8B–35B, often at Q2–Q4, on a 16–32k context window), so it is careful with context, fails more gracefully, and says when it's stuck. It doesn't make a small model smarter. It makes working with one less frustrating: less of the window wasted, fewer loops on the same read, and less chance of the task quietly getting lost.
Day to day it runs on large hosted models, and small local ones are where it gets tested. At the small end, expect a better experience rather than frontier results. In testing, 14–35B models handled multi-file tasks most reliably, 8–9B models managed simple, well-scoped ones, and below 7B rarely got far.
Reika is pre-1.0, so commands, settings and behavior can still change between releases.
The harness's design choices were measured rather than assumed, and the record is published. It measures what a harness can and can't do about the problems small models run into, not how good the models become. The short version, with the runs and the numbers behind it in docs/findings.md:
/issue and /review ship with Reika and appear wherever gh and a GitHub remote are.REIKA_VISION_MODEL to read pasted images with a vision model instead). Windows may work under WSL2.Install it with any of these (npm and pnpm need Node ≥ 22; Homebrew pulls Node in):
npm i -g @alexwkleung/reika
pnpm add -g @alexwkleung/reika
brew install alexwkleung/tap/reika
Serve a model. Example below with llama.cpp:
llama-server -m <model.gguf> -c 24576 --jinja <other-launch-args>
Then point Reika at it:
export REIKA_MODEL=model # or put it in ~/.config/reika/.env
cd your-project && reika
Using a hosted API instead? Skip the server and set REIKA_BASE_URL, REIKA_API_KEY and REIKA_MODEL to the provider's endpoint, your key and its model id; Install and first run has the details.
REIKA_BASE_URL defaults to http://localhost:8080/v1, llama-server's default. The context window is read from the server when it reports one, and for hosted endpoints from the models.dev catalog; set REIKA_CONTEXT_WINDOW when neither has it. Reika sends no sampling parameters of its own, so your server's flags are what apply.
To run from a checkout instead — git clone https://github.com/alexwkleung/reika.git && cd reika, then pnpm install (npm users: npm i -g pnpm) and pnpm run install:global to build and install the reika binary, or cp .env.example .env and pnpm run dev to run without installing. pnpm run uninstall:global removes the global binary.
Set these in your shell, a project .env, or ~/.config/reika/.env (in that order of precedence). The full list, experimental flags, and multi-model profiles are in docs/configuration.md.
| Key | Default | What |
|---|---|---|
REIKA_MODEL | required | Model name, or a comma-separated list served by the same endpoint |
REIKA_BASE_URL | http://localhost:8080/v1 | OpenAI-compatible endpoint |
REIKA_API_KEY | no-key | API key (any non-empty value for local servers) |
REIKA_CONTEXT_WINDOW | probed from the server | Context window in tokens; drives compaction, the context gauge, and the repo map size |
REIKA_MIN_GEN_TOKENS | learned (from 2048) | Room reserved for the reply; learned from the model's rounds, set it to pin |
REIKA_AUTO_APPROVE | safe | off confirms every edit and command; bypass confirms nothing |
Shift+Tab cycles modes, or switch with a slash command. The mode and model you end a session in are the ones the next session opens in.
| Mode | What it does |
|---|---|
agent | The default: the model reads, edits, and runs commands |
plan | Read-only exploration that ends in a written plan; /implement runs it |
vibe | Plans first, then implements the plan, on every prompt |
minimal | Shell only, with no repo map or project context loaded upfront |
grind | Agent turns run a fixed, careful procedure: define done, test, review |
chat | Plain conversation with a separate history; web tools only |
shell | Your input runs as a shell command, and the output joins the agent's context |
reika -p "<prompt>" runs a single turn headless and prints the reply, for scripts and other harnesses. See docs/usage.md for modes in detail, headless flags, and every slash command.
safe: ordinary edits and commands run on their own, but anything matching a dangerous pattern (rm -rf, force pushes, package installs, curl, …) and any write outside the project still asks you first. bypass can only be set at launch, never mid-session.git/gh. A command you explicitly approve runs unsandboxed. REIKA_SANDBOX=0 turns it off.Found a way around one of these? Report it privately; see SECURITY.md.
The same pages are on the website, reikacode.com.
| Doc | Covers |
|---|---|
| Install and first run | Requirements, installing, serving a model, and the first run |
| Configuration | Every .env key, experimental flags, named profiles, last-session state |
| Modes, headless, commands | Each mode in depth, reika -p, slash commands, /save transcripts |
| Tools | The model's tools, approval prompts, web search setup, MCP servers, .gitignore |
| Instructions and skills | AGENTS.md, skills as slash commands, plain-English routing, pasted URLs |
| Tested models | Local quants and APIs Reika has been run against |
| Findings | The measured record: what broke, what held, and how each number was obtained |
| Platforms | Requirements, running on a weak machine, what differs on Linux and Windows |
| Architecture and caveats | How the harness works, and known limitations |
| Contributing | Scripts, design philosophy, a pointer to AGENTS.md, and external contributor guidelines |
Claude Code, Codex, Crush, OpenCode, Pi, Aider, DeepSeek Harness, Qwen Code, Kimi Code CLI, Gemini CLI, Junie, and DS4.
Apache License 2.0. See LICENSE and NOTICE.
If Reika is useful to you, consider supporting it via GitHub Sponsors, Ko-fi, or Buy Me a Coffee. See Support for what it funds and other ways to help.