KikeVen/zerikai_memory

A standalone local-only Python MCP server that gives any IDE persistent, workspace-isolated memory. works with any IDE supporting MCP servers

Python

33

96 commits

updated Sep 26, 2026

See the code

See what people are saying

README

zerikai_memory 🧠

⭐ Bookmark the project: If you use this tool, drop a star to save it to your GitHub profile and track new performance updates.

Never lose your AI context again.
zerikai_memory provides persistent, workspace-isolated memory for every IDE that is local-first, cost-aware, and instant. It uses deterministic Tree-Sitter code parsing indexing to capture entities and deep code descriptions like functions, classes, and docstrings into a local ChromaDB vector store. Accessed via a local MCP interface to slash token costs while maintaining high-resolution codebase mapping, it retrieves hyper-relevant context on query through L2 and Lexical re-indexing with strict source verification (Entity, File, Line Number, and L2). Designed to pair perfectly with low-cost DeepSeek APIs, it injects structured, highly precise local context instead of dumping raw, massive files, maximizing KV cache hits to radically reduce your active token costs.

Python 3.11+ ChromaDB Ollama DeepSeek MCP MIT License
Platform Support

GitHub Release GitHub Actions Workflow Status GitHub commit activity

πŸ’‘Status: Active & Self-Contained.
This project is used daily and actively maintained by the author. Pull Requests and Issues are closed to keep maintenance overhead low. It is provided fully functional and ready for production use.

⚠️ Platform Support Notice:

Note: This project is developed and actively maintained on Windows. GitHub Actions verifies that basic installation compiles across Windows, macOS, and Linux, but the maintainer cannot troubleshoot platform-specific runtime errors on Mac or Linux. Community pull requests fixing Mac/Linux bugs are highly welcome!


πŸ†• New β€” Jev Semantic Judgment Layer

zerikai_memory now integrates Jev, TypeSafe AI's System One model β€” a fast, structured judgment engine that decides which retrieved passages actually answer your question and attaches a plain-text evidence report to every answer.

  • What it is: a second, narrow AI layer beside your LLM. Your LLM writes the answer; Jev judges the evidence it is built from. It returns calibrated probabilities β€” not prose.
  • What it does: per-passage relevance / evidence / contradiction / prompt-injection scoring, better passage ordering (replaces keyword rerank), and a plain-text Assessment / Evidence / Guidance report the agent can act on. Injection attempts are dropped.
  • Off by default: with ENABLE_JEV=false behavior is byte-identical to before, and it is fully fail-open β€” if Jev is unavailable, the normal pipeline runs.
  • Activate it: set TYPESAFE_API_KEY and ENABLE_JEV=true in .env, then restart the server.
  • Get the API key: TypeSafe early access β†’ console.typesafe.ai.

πŸ“– Full details β€” how the judgment layer works and every parameter: documentation/10-jev-judgment-layer.md


The Problem

Every new chat session starts completely cold. When you switch contexts or open a new window:

  • Your AI Agent forgets every architectural decision, convention, and stack choice made over hours
  • You waste critical tokens and 10–15 minutes re-explaining the codebase setup in every single chat
  • Large raw file dumps inflate your token costs and shrink your available context window instantly
  • Switching IDEs (e.g., VS Code to Cursor) forces you to restart your conversation history from scratch

How Zerikai Memory Solves It

Zerikai Memory runs as a local STDIO MCP server between your IDE and your LLM. It parses your codebase using tree-sitter, indexes code entities into a local ChromaDB vector store, and injects highly relevant context snippets dynamically through natural language.

Your Codebase  β†’  tree-sitter (local parse)  β†’  ChromaDB (.brain/)
                                                      β”‚
Your IDE       β†’  MCP Server (:stdio)        β†’        β–Ό
                                             Ollama / DeepSeek
                    β”‚                      (auto-routed synthesis)
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚   4-Stage Pipeline     β”‚
          β”‚   L1  Vector Search    β”‚  ChromaDB L2 distance matching
          β”‚   L2  Lexical Re-rank  β”‚  Keyword overlap boost on names
          β”‚   L3  Auto-Routing     β”‚  Ollama (free) vs. DeepSeek Cloud
          β”‚   L4  LLM Synthesis    β”‚  Answer + inline #file:line citations
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Architecture

Cost Model β€” The Context Tax Mitigation

What gets taxedWithout ZerikaiWith Zerikai
πŸ”΄ Monthly quotaRe-explaining stack, decisions, and conventions every sessionIndexed once. Retrieved as compact snippets per query.
🟑 Context windowRaw file dumps shrink the window available for code generation1,000–1,200 token brief prefix*. Window stays wide open.
βšͺ IDE switchingFull re-explanation required in every new toolShared zerikai_memory workspace .brain/ directory.

Tip: The project brief acts as a stable prefix. After your first query, DeepSeek caches it β€” making subsequent repeated queries 50–100Γ— cheaper (rate depends on whether you are in peak or off-peak hours). See the DeepSeek Pricing section for current rates.


Quick Start

I have added a YouTube video walkthrough of the installation and setup process in the first step below. If you prefer text instructions, just follow along with the code snippets.

1. Install

Watch the installation video

Click the image below to watch a step-by-step walkthrough of the installation and setup process:

Zerikai Memory installation video

git clone https://github.com/your-username/zerikai_memory.git
cd zeriakai_memory

# Create and activate a virtual environment (Python 3.11+)
python -m venv venv
source venv/bin/activate  # Windows: .\venv\Scripts\activate

pip install -r requirements.txt

2. Configure Environment

Remove the .example from the .env.example file in the root directory and rename it to .env:

Expand to view .env
DEEPSEEK_API_KEY=your_deepseek_key_here

# Memory Mode controls which LLM is used for operations:
# - "cloud": Use DeepSeek for all operations (scan, brief, queries) - highest quality, tracked usage
# - "hybrid": Use Ollama for file scanning, DeepSeek for briefs and escalated queries
# - "local": Use Ollama for everything (free, but lower quality briefs)
MEMORY_MODE=cloud

# Enable token tracking and cost reporting (SQLite database at .brain/token_usage.db)
# Set to "false" to disable tracking
ENABLE_TOKEN_TRACKING=true

# Enable deepseek-v4-pro for complex architectural queries (design, architecture, tradeoffs)
# v4-pro is ~4Γ— more expensive than deepseek-flash. See DeepSeek Pricing section below
# for current peak/off-peak rates.
# Recommended: keep this "false" unless you need maximum reasoning capability
ENABLE_DEEPSEEK_PRO=false

# DeepSeek thinking mode (chain-of-thought). Docs:
# https://api-docs.deepseek.com/guides/thinking_mode/
# Thinking is ON by default at effort "high" when no parameter is sent.
# Options per path: enabled | disabled. "disabled" is right for extractive work.
# NOTE: briefs and query synthesis are separate pipelines with separate toggles.
DEEPSEEK_THINKING_BRIEF=disabled   # 9-section project brief generation
DEEPSEEK_THINKING_SCAN=disabled    # per-file indexing summaries
DEEPSEEK_THINKING_QUERY=disabled   # query_memory answer synthesis

# Reasoning effort per path when that path is "enabled": low | high | max.
# Ignored when the matching path is "disabled". "low" is the cheapest enabled tier.
DEEPSEEK_REASONING_EFFORT_BRIEF=low
DEEPSEEK_REASONING_EFFORT_SCAN=low
DEEPSEEK_REASONING_EFFORT_QUERY=low

# Semantic search relevance cutoff for query_memory (L2 distance).
# Lower = stricter. Watch "best dist=X.XX" in server.log to calibrate.
# Typical: <0.8 strong match, 0.8-1.5 related, >1.5 noise.
QUERY_DISTANCE_THRESHOLD=1.0

# File extensions to skip during scanning when tree-sitter produces zero
# entities (no functions, classes, headings, semantic HTML elements, etc.).
# Saves API calls on bare config files, trivial templates, empty CSS, etc.
# Format: ['.py', '.html', '.md', '.css']
# Default: [] (empty β€” no extensions skipped, all fall through to LLM).
SKIP_BARE_FILES=['.py', '.html', '.md', '.css']

# Enable lexical re-ranking in query_memory.
# When true, results passing the distance threshold are reordered by a
# weighted combination of semantic distance and keyword overlap in entity
# name and docstring text. Nothing is dropped β€” pure reorder.
# Default: false (existing pure-semantic behaviour preserved).
ENABLE_LEXICAL_RERANK=true

# Weight applied per keyword hit during lexical re-ranking.
# The 1/dist spread across the valid-hit band (0.85–0.98) is ~0.156.
# Keep this value below that spread to avoid keyword hits overriding
# a genuinely closer semantic result.
# Recommended starting point: 0.05 (one hit = +0.05, two hits = +0.10).
LEXICAL_RERANK_WEIGHT=0.05

# Candidate pool per section for project-brief synthesis. Each section queries
# ChromaDB, re-ranks locally, then trims to a per-section cap (20/25/30).
# Decoupled from FETCH_CAP (query-only) so a tight query pool doesn't starve
# the brief. Default: 20.
BRIEF_FETCH_CAP=20

2a. Verify

python -c "from main import scan_workspace, query_memory; print('OK')"

You should see the startup banner followed by OK.

3. IDE Rule Enforcement

To stop your AI agent from ignoring the memory protocol, copy these directives into your IDE's agent rules profile (e.g., .cursorrules or system prompt guidelines):

IDE Rules in: agent_rules/ide_agent_rules.md

  • Universal-Brain First: The agent must query universal-brain before attempting raw file searches.
  • Source Discipline: Every answer must surface actual file.py:line citations with zero fabrication.

4. IDE Registration

  1. Press Ctrl+Shift+P β†’ MCP: Add Local Server
  2. Choose STDIO
  3. Set command: C:\path\to\zerikai_memory\venv\Scripts\python.exe C:\path\to\zerikai_memory\main.py

Add to your claude_desktop_config.json profile:

{
  "mcpServers": {
    "universal-brain": {
      "command": "C:\\path\\to\\zerikai_memory\\venv\\Scripts\\python.exe",
      "args": ["C:\\path\\to\\zerikai_memory\\main.py"]
    }
  }
}

5. Setup the .memignore file

Works like .gitignore: one pattern per line. scan_workspace reads this file and skips matching paths.

Each project should have its own .memignore in its root directory. Forgetting to configure it before the first scan is the most common reason to use drop_memory.py and start fresh:

Examples of what to ignore: Expand to view

Sample .memignore
# Directories (trailing slash required)
.git/
node_modules/
venv/
__pycache__/
.brain/
dist/
build/

# File/Folder patterns
**/test/
**/tests/
.env
*.log
*.lock
*.pyc

6. Embedding-Docstring Skill

Before running your first index scan, optimize your codebase's docstrings for vector search. Ask your AI Agent:

  • To install the embedding-docstring globally in your IDE and run it against your codebase to rewrite docstrings into a more embedding-friendly format.

"Audit and optimize docstrings across this project using the embedding-docstring skill, respecting .memignore."

RequirementWhy It MattersTarget Impact
Explicit Tech NamesUse "Uses Redis" instead of "key-value store"Embeddings match precise tokens, not abstract concepts.
Routing / BranchesDocument specific route paths and logical pivot optionsEnsures structural code matches are surfaceable.
Guarantees & EffectsExplicitly state code idempotency, atomicity, or mutation side-effectsPrevents agent generation from breaking runtime boundaries.

7. Chat with Memory

Simply instruct your IDE's active AI agent using natural language commands:

prefix queries with "universal-brain: <command>" to ensure they route through the MCP server and leverage your indexed memory:

  • Scan the workspace for the first time: "Set up memory for this project"
  • Ask a question: "What are the main architectural components of this project?"

Frequently used follow-ups:

  • After a code change: "Rescan the workspace and force a refresh of the project brief."
  • Save part of a chat: "Save the following context to memory: [your custom notes or constraints here]"
  • Ask how much have you used: "Get me a cost report for my memory usage so far."

See below for a full reference of available commands and their descriptions.


Safe Upgrade Process

Upgrading zerikai_memory while keeping your already indexed workspaces is safe as long as the workspace path never changes.

Your workspace identity is a stable UUID derived from its absolute filesystem path (_derive_workspace_id in main.py). Because .brain/ is gitignored, a git pull never touches your indexed memory, project briefs, or workspace registry.

Upgrade in place β€” never copy, rename, clone, or move the project directory.

The whole point is to preserve your already-indexed .brain/ data β€” back it up, upgrade the code, restore it, and you get your old memory back exactly as it was. No re-indexing needed.

# 1. Stop the MCP server (close the IDE / kill main.py) β€” releases the ChromaDB lock

# 2. Back up your existing indexed data so you don't lose it
#    (do NOT rename the whole project β€” only back up .brain/)
#    Windows (PowerShell):  Copy-Item -Recurse .brain .brain.bak
#    macOS / Linux:         cp -r .brain .brain.bak

# 3. Pull the latest code
git pull

# 4. Reinstall dependencies only if requirements.txt changed
#    Windows:  .\venv\Scripts\python.exe -m pip install -r requirements.txt
#    macOS/Linux:  venv/bin/python -m pip install -r requirements.txt

# 5. Restore .brain/ (only needed if the upgrade replaced it), restart, and verify
#    the workspace still resolves to the same UUID β€” your old memory is back

⚠️ Do NOT rename the folder (e.g. zerikai_memory_old), clone into a differently-named folder, or move the project to a new parent directory. Any of these changes the path β†’ generates a new UUID β†’ orphans your old collection and brief. If you already did this, recover with merge_workspaces.

Full step-by-step guide, backup/restore, and recovery instructions: documentation/09-upgrading.md


MCP Tools Reference

You never run these commands directly; your active AI agent executes them on your behalf.

Workspace Management

ToolDescription
init_workspaceRegisters a project folder, assigns a UUID, and creates a pending brief file. Idempotent; safe to run multiple times.
list_workspacesLists all known workspaces that have a brief or stored memories.
resolve_workspaceResolves a workspace identifier (UUID, short-UUID, or display name) to its filesystem path.
merge_workspacesConsolidates duplicate workspace IDs into one. Irreversible.
debug_workspace_idDiagnostic tool; shows what workspace ID would be generated from a given path.

Memory & Briefs

ToolDescription
scan_workspaceStarts a background scan. Returns immediately; use scan_status to track progress. Walks the directory, respects .memignore, saves all readable text files to persistent memory. Idempotent and self-cleaning. Concurrent (4 workers, batch writes).
scan_statusReturns progress of a running or recently completed background scan: files scanned, entities indexed, errors, elapsed time, brief status.
save_to_memoryManually saves an architectural decision, fact, or technical note with an optional category tag.
list_memoryLists stored memories for a workspace, optionally filtered by category.
query_memoryRetrieves relevant context via vector search and synthesises an answer via Ollama or DeepSeek (auto-routed). Returns the answer as plain text plus a trailing Sources: block of file:line citations with relevance scores (L2 distance or rerank).
get_briefRetrieves the current project brief from .brain/contexts/.
update_briefManually updates the markdown content of a project brief.

Usage & Diagnostics

ToolDescription
get_token_usageReturns DeepSeek API token usage and cost statistics.
get_cost_reportGenerates a cost breakdown by operation type. Prepends a live PEAK / OFF-PEAK banner showing currently active rates.
get_cache_statsShows cache hit/miss rates by operation type.
purge_usage_dataDeletes historical token tracking records.

Project Brief Matrix

When a workspace is scanned, Zerikai compiles a dense 1,000–1,200 token project brief across 9 locked components:

SectionWhat It Captures
1. OverviewProject domain, primary type, and functional scope.
2. Technical StackBackend engines, databases, integrations, and core libraries.
3. Core ArchitectureInteractivity between frontend, backend, and processing layers.
4. Primary ConventionsLocal code styling, custom error handling, and validation schema rules.
5. PurposeBusiness logic problems solved and key underlying objectives.
6. Key FilesDefinitive app entry points, central routers, and specific domain tasks.
7. Dev & TestingEnvironment installation setups, execution triggers, and testing runs.
8. Data FlowComplete systemic request lifecycle tracing from gateway to database layer.
9. Future RoadmapPlanned engineering steps and dangling TODO items parsed directly from code.

Memory Modes

Adjust your operation profile via the MEMORY_MODE environment toggle to balance privacy, speed, and API costs.

πŸ’‘ Deterministic First: All high-resolution code parsing (functions, classes, methods) is performed locally and deterministically using Tree-Sitter for $0 cost. The engines below are only used for text-file fallbacks and generating the architectural Project Brief.

ModeAnalysis & BriefsQuery EngineTotal CostIdeal Use Case
🟒 cloudDeepSeekDeepSeekLowRecommended. High-fidelity architectural briefs.
🟑 hybridOllamaOllama + DeepSeekLowestLocal privacy with cloud reasoning escalation.
πŸ”΄ localOllamaOllama$0.00100% air-gapped hardware-local tracking.

Configuration Reference

KeyDefaultDescription
DEEPSEEK_API_KEYRequiredActive API authorization key from platform.deepseek.com.
MEMORY_MODEcloudSets target engines: choices include cloud, hybrid, or local.
ENABLE_TOKEN_TRACKINGtrueCalculates continuous usage and outputs summaries to SQLite.
QUERY_DISTANCE_THRESHOLD1.5Sets L2 vector distance cutoff limits. Lower inputs restrict matches.
ENABLE_LEXICAL_RERANKfalseActivates secondary hybrid reordering layer via keyword matching.
SKIP_BARE_FILES[]Extension list to bypass when tree-sitter finds zero valid code entities.

DeepSeek Pricing

DeepSeek uses peak / off-peak pricing across all tiers. Off-peak rates are exactly half of peak rates.

⏰ Peak hours (UTC): 01:00–04:00 and 06:00–10:00, Monday–Friday only. Weekends are always off-peak.

Rate Table (USD per 1M tokens)

Tierdeepseek-flash inputdeepseek-flash outputdeepseek-flash cachedv4-pro inputv4-pro outputv4-pro cached
Peak$0.30$1.20$0.006$1.32$3.96$0.044
Off-peak$0.15$0.60$0.003$0.66$1.98$0.022

The tool automatically resolves the correct tier at call time β€” no manual configuration needed.

Off-Peak Windows by Region

⚠️ Peak hours apply Monday–Friday only. Weekends are always off-peak regardless of time. Offsets shown for summer / daylight saving time (DST). In winter, US timezones shift 1 hour later; European zones shift 1 hour earlier β€” meaning off-peak windows shift accordingly.

RegionUTC offset (summer)Peak local time (Mon–Fri)βœ… Off-peak local time
EST (New York, Miami)UTCβˆ’58pm–11pm & 1am–5am5am–8pm and 11pm–1am (+ all weekend)
CST (Chicago, Dallas)UTCβˆ’67pm–10pm & midnight–4am4am–7pm and 10pm–midnight (+ all weekend)
PST (Los Angeles, Seattle)UTCβˆ’85pm–8pm & 10pm–2am2am–5pm and 8pm–10pm (+ all weekend)
Ireland (Dublin)UTC+12am–5am & 7am–11am11am–2am and 5am–7am (+ all weekend)
Spain (Madrid)UTC+23am–6am & 8am–noonnoon–3am and 6am–8am (+ all weekend)
Germany (Berlin)UTC+23am–6am & 8am–noonnoon–3am and 6am–8am (+ all weekend)
Norway (Oslo)UTC+23am–6am & 8am–noonnoon–3am and 6am–8am (+ all weekend)

Key insight: For US users working standard hours on weekdays, most of the working day is already off-peak (cheaper). European users in GMT+2 zones benefit from off-peak pricing through most of the afternoon and evening β€” and all weekend queries cost even less. Saturday and Sunday queries are always billed at off-peak rates.


Auxiliary Scripts & Troubleshooting

Workspace Reset

If you accidentally execute a workspace crawl before setting up your .memignore configurations, run the auxiliary wipe script to delete stale workspace data:

# Windows
.\venv\Scripts\python.exe drop_memory.py "Workspace Name"

# macOS / Linux
venv/bin/python drop_memory.py "Workspace Name"

Live Log Diagnostics

Monitor server activity, runtime operations, and auto-routing logs inside .brain/server.log:

# Live stream logs (macOS/Linux)
tail -f .brain/server.log

# Live stream logs (Windows PowerShell)
Get-Content .brain\server.log -Wait -Tail 30


Security & Data Privacy

  • All active vector spaces, tracking registries, and context details reside directly on your local machine.
  • Add .env and .brain/ explicitly to your global or project .gitignore patterns to prevent API keys and secure indexes from leaking to version control platforms.

To read more about the underlying design principles, architecture decisions, and future roadmap for Zerikai Memory, check out the insight article.

License

MIT License Β© Zerikai

πŸ› οΈ Support: This project is provided as-is for personal use.

ai
ai-tools
chromadb
deepseek-v4-pro
jev
lexical
mcp
mcp-memory
mcp-server
mcp-tools
memory
mistral-7b
ollama
ornith
rag-memory
re-ranking
tree-sitter-parser
typesafe-ai
typesafe-jev
zerikai

KikeVen/zerikai_memory

A standalone local-only Python MCP server that gives any IDE persistent, workspace-isolated memory. works with any IDE supporting MCP servers

Python

33

96 commits

updated Sep 26, 2026

See the code

See what people are saying

README

zerikai_memory 🧠

⭐ Bookmark the project: If you use this tool, drop a star to save it to your GitHub profile and track new performance updates.

Never lose your AI context again.
zerikai_memory provides persistent, workspace-isolated memory for every IDE that is local-first, cost-aware, and instant. It uses deterministic Tree-Sitter code parsing indexing to capture entities and deep code descriptions like functions, classes, and docstrings into a local ChromaDB vector store. Accessed via a local MCP interface to slash token costs while maintaining high-resolution codebase mapping, it retrieves hyper-relevant context on query through L2 and Lexical re-indexing with strict source verification (Entity, File, Line Number, and L2). Designed to pair perfectly with low-cost DeepSeek APIs, it injects structured, highly precise local context instead of dumping raw, massive files, maximizing KV cache hits to radically reduce your active token costs.

Python 3.11+ ChromaDB Ollama DeepSeek MCP MIT License
Platform Support

GitHub Release GitHub Actions Workflow Status GitHub commit activity

πŸ’‘Status: Active & Self-Contained.
This project is used daily and actively maintained by the author. Pull Requests and Issues are closed to keep maintenance overhead low. It is provided fully functional and ready for production use.

⚠️ Platform Support Notice:

Note: This project is developed and actively maintained on Windows. GitHub Actions verifies that basic installation compiles across Windows, macOS, and Linux, but the maintainer cannot troubleshoot platform-specific runtime errors on Mac or Linux. Community pull requests fixing Mac/Linux bugs are highly welcome!


πŸ†• New β€” Jev Semantic Judgment Layer

zerikai_memory now integrates Jev, TypeSafe AI's System One model β€” a fast, structured judgment engine that decides which retrieved passages actually answer your question and attaches a plain-text evidence report to every answer.

  • What it is: a second, narrow AI layer beside your LLM. Your LLM writes the answer; Jev judges the evidence it is built from. It returns calibrated probabilities β€” not prose.
  • What it does: per-passage relevance / evidence / contradiction / prompt-injection scoring, better passage ordering (replaces keyword rerank), and a plain-text Assessment / Evidence / Guidance report the agent can act on. Injection attempts are dropped.
  • Off by default: with ENABLE_JEV=false behavior is byte-identical to before, and it is fully fail-open β€” if Jev is unavailable, the normal pipeline runs.
  • Activate it: set TYPESAFE_API_KEY and ENABLE_JEV=true in .env, then restart the server.
  • Get the API key: TypeSafe early access β†’ console.typesafe.ai.

πŸ“– Full details β€” how the judgment layer works and every parameter: documentation/10-jev-judgment-layer.md


The Problem

Every new chat session starts completely cold. When you switch contexts or open a new window:

  • Your AI Agent forgets every architectural decision, convention, and stack choice made over hours
  • You waste critical tokens and 10–15 minutes re-explaining the codebase setup in every single chat
  • Large raw file dumps inflate your token costs and shrink your available context window instantly
  • Switching IDEs (e.g., VS Code to Cursor) forces you to restart your conversation history from scratch

How Zerikai Memory Solves It

Zerikai Memory runs as a local STDIO MCP server between your IDE and your LLM. It parses your codebase using tree-sitter, indexes code entities into a local ChromaDB vector store, and injects highly relevant context snippets dynamically through natural language.

Your Codebase  β†’  tree-sitter (local parse)  β†’  ChromaDB (.brain/)
                                                      β”‚
Your IDE       β†’  MCP Server (:stdio)        β†’        β–Ό
                                             Ollama / DeepSeek
                    β”‚                      (auto-routed synthesis)
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚   4-Stage Pipeline     β”‚
          β”‚   L1  Vector Search    β”‚  ChromaDB L2 distance matching
          β”‚   L2  Lexical Re-rank  β”‚  Keyword overlap boost on names
          β”‚   L3  Auto-Routing     β”‚  Ollama (free) vs. DeepSeek Cloud
          β”‚   L4  LLM Synthesis    β”‚  Answer + inline #file:line citations
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Architecture

Cost Model β€” The Context Tax Mitigation

What gets taxedWithout ZerikaiWith Zerikai
πŸ”΄ Monthly quotaRe-explaining stack, decisions, and conventions every sessionIndexed once. Retrieved as compact snippets per query.
🟑 Context windowRaw file dumps shrink the window available for code generation1,000–1,200 token brief prefix*. Window stays wide open.
βšͺ IDE switchingFull re-explanation required in every new toolShared zerikai_memory workspace .brain/ directory.

Tip: The project brief acts as a stable prefix. After your first query, DeepSeek caches it β€” making subsequent repeated queries 50–100Γ— cheaper (rate depends on whether you are in peak or off-peak hours). See the DeepSeek Pricing section for current rates.


Quick Start

I have added a YouTube video walkthrough of the installation and setup process in the first step below. If you prefer text instructions, just follow along with the code snippets.

1. Install

Watch the installation video

Click the image below to watch a step-by-step walkthrough of the installation and setup process:

Zerikai Memory installation video

git clone https://github.com/your-username/zerikai_memory.git
cd zeriakai_memory

# Create and activate a virtual environment (Python 3.11+)
python -m venv venv
source venv/bin/activate  # Windows: .\venv\Scripts\activate

pip install -r requirements.txt

2. Configure Environment

Remove the .example from the .env.example file in the root directory and rename it to .env:

Expand to view .env
DEEPSEEK_API_KEY=your_deepseek_key_here

# Memory Mode controls which LLM is used for operations:
# - "cloud": Use DeepSeek for all operations (scan, brief, queries) - highest quality, tracked usage
# - "hybrid": Use Ollama for file scanning, DeepSeek for briefs and escalated queries
# - "local": Use Ollama for everything (free, but lower quality briefs)
MEMORY_MODE=cloud

# Enable token tracking and cost reporting (SQLite database at .brain/token_usage.db)
# Set to "false" to disable tracking
ENABLE_TOKEN_TRACKING=true

# Enable deepseek-v4-pro for complex architectural queries (design, architecture, tradeoffs)
# v4-pro is ~4Γ— more expensive than deepseek-flash. See DeepSeek Pricing section below
# for current peak/off-peak rates.
# Recommended: keep this "false" unless you need maximum reasoning capability
ENABLE_DEEPSEEK_PRO=false

# DeepSeek thinking mode (chain-of-thought). Docs:
# https://api-docs.deepseek.com/guides/thinking_mode/
# Thinking is ON by default at effort "high" when no parameter is sent.
# Options per path: enabled | disabled. "disabled" is right for extractive work.
# NOTE: briefs and query synthesis are separate pipelines with separate toggles.
DEEPSEEK_THINKING_BRIEF=disabled   # 9-section project brief generation
DEEPSEEK_THINKING_SCAN=disabled    # per-file indexing summaries
DEEPSEEK_THINKING_QUERY=disabled   # query_memory answer synthesis

# Reasoning effort per path when that path is "enabled": low | high | max.
# Ignored when the matching path is "disabled". "low" is the cheapest enabled tier.
DEEPSEEK_REASONING_EFFORT_BRIEF=low
DEEPSEEK_REASONING_EFFORT_SCAN=low
DEEPSEEK_REASONING_EFFORT_QUERY=low

# Semantic search relevance cutoff for query_memory (L2 distance).
# Lower = stricter. Watch "best dist=X.XX" in server.log to calibrate.
# Typical: <0.8 strong match, 0.8-1.5 related, >1.5 noise.
QUERY_DISTANCE_THRESHOLD=1.0

# File extensions to skip during scanning when tree-sitter produces zero
# entities (no functions, classes, headings, semantic HTML elements, etc.).
# Saves API calls on bare config files, trivial templates, empty CSS, etc.
# Format: ['.py', '.html', '.md', '.css']
# Default: [] (empty β€” no extensions skipped, all fall through to LLM).
SKIP_BARE_FILES=['.py', '.html', '.md', '.css']

# Enable lexical re-ranking in query_memory.
# When true, results passing the distance threshold are reordered by a
# weighted combination of semantic distance and keyword overlap in entity
# name and docstring text. Nothing is dropped β€” pure reorder.
# Default: false (existing pure-semantic behaviour preserved).
ENABLE_LEXICAL_RERANK=true

# Weight applied per keyword hit during lexical re-ranking.
# The 1/dist spread across the valid-hit band (0.85–0.98) is ~0.156.
# Keep this value below that spread to avoid keyword hits overriding
# a genuinely closer semantic result.
# Recommended starting point: 0.05 (one hit = +0.05, two hits = +0.10).
LEXICAL_RERANK_WEIGHT=0.05

# Candidate pool per section for project-brief synthesis. Each section queries
# ChromaDB, re-ranks locally, then trims to a per-section cap (20/25/30).
# Decoupled from FETCH_CAP (query-only) so a tight query pool doesn't starve
# the brief. Default: 20.
BRIEF_FETCH_CAP=20

2a. Verify

python -c "from main import scan_workspace, query_memory; print('OK')"

You should see the startup banner followed by OK.

3. IDE Rule Enforcement

To stop your AI agent from ignoring the memory protocol, copy these directives into your IDE's agent rules profile (e.g., .cursorrules or system prompt guidelines):

IDE Rules in: agent_rules/ide_agent_rules.md

  • Universal-Brain First: The agent must query universal-brain before attempting raw file searches.
  • Source Discipline: Every answer must surface actual file.py:line citations with zero fabrication.

4. IDE Registration

  1. Press Ctrl+Shift+P β†’ MCP: Add Local Server
  2. Choose STDIO
  3. Set command: C:\path\to\zerikai_memory\venv\Scripts\python.exe C:\path\to\zerikai_memory\main.py

Add to your claude_desktop_config.json profile:

{
  "mcpServers": {
    "universal-brain": {
      "command": "C:\\path\\to\\zerikai_memory\\venv\\Scripts\\python.exe",
      "args": ["C:\\path\\to\\zerikai_memory\\main.py"]
    }
  }
}

5. Setup the .memignore file

Works like .gitignore: one pattern per line. scan_workspace reads this file and skips matching paths.

Each project should have its own .memignore in its root directory. Forgetting to configure it before the first scan is the most common reason to use drop_memory.py and start fresh:

Examples of what to ignore: Expand to view

Sample .memignore
# Directories (trailing slash required)
.git/
node_modules/
venv/
__pycache__/
.brain/
dist/
build/

# File/Folder patterns
**/test/
**/tests/
.env
*.log
*.lock
*.pyc

6. Embedding-Docstring Skill

Before running your first index scan, optimize your codebase's docstrings for vector search. Ask your AI Agent:

  • To install the embedding-docstring globally in your IDE and run it against your codebase to rewrite docstrings into a more embedding-friendly format.

"Audit and optimize docstrings across this project using the embedding-docstring skill, respecting .memignore."

RequirementWhy It MattersTarget Impact
Explicit Tech NamesUse "Uses Redis" instead of "key-value store"Embeddings match precise tokens, not abstract concepts.
Routing / BranchesDocument specific route paths and logical pivot optionsEnsures structural code matches are surfaceable.
Guarantees & EffectsExplicitly state code idempotency, atomicity, or mutation side-effectsPrevents agent generation from breaking runtime boundaries.

7. Chat with Memory

Simply instruct your IDE's active AI agent using natural language commands:

prefix queries with "universal-brain: <command>" to ensure they route through the MCP server and leverage your indexed memory:

  • Scan the workspace for the first time: "Set up memory for this project"
  • Ask a question: "What are the main architectural components of this project?"

Frequently used follow-ups:

  • After a code change: "Rescan the workspace and force a refresh of the project brief."
  • Save part of a chat: "Save the following context to memory: [your custom notes or constraints here]"
  • Ask how much have you used: "Get me a cost report for my memory usage so far."

See below for a full reference of available commands and their descriptions.


Safe Upgrade Process

Upgrading zerikai_memory while keeping your already indexed workspaces is safe as long as the workspace path never changes.

Your workspace identity is a stable UUID derived from its absolute filesystem path (_derive_workspace_id in main.py). Because .brain/ is gitignored, a git pull never touches your indexed memory, project briefs, or workspace registry.

Upgrade in place β€” never copy, rename, clone, or move the project directory.

The whole point is to preserve your already-indexed .brain/ data β€” back it up, upgrade the code, restore it, and you get your old memory back exactly as it was. No re-indexing needed.

# 1. Stop the MCP server (close the IDE / kill main.py) β€” releases the ChromaDB lock

# 2. Back up your existing indexed data so you don't lose it
#    (do NOT rename the whole project β€” only back up .brain/)
#    Windows (PowerShell):  Copy-Item -Recurse .brain .brain.bak
#    macOS / Linux:         cp -r .brain .brain.bak

# 3. Pull the latest code
git pull

# 4. Reinstall dependencies only if requirements.txt changed
#    Windows:  .\venv\Scripts\python.exe -m pip install -r requirements.txt
#    macOS/Linux:  venv/bin/python -m pip install -r requirements.txt

# 5. Restore .brain/ (only needed if the upgrade replaced it), restart, and verify
#    the workspace still resolves to the same UUID β€” your old memory is back

⚠️ Do NOT rename the folder (e.g. zerikai_memory_old), clone into a differently-named folder, or move the project to a new parent directory. Any of these changes the path β†’ generates a new UUID β†’ orphans your old collection and brief. If you already did this, recover with merge_workspaces.

Full step-by-step guide, backup/restore, and recovery instructions: documentation/09-upgrading.md


MCP Tools Reference

You never run these commands directly; your active AI agent executes them on your behalf.

Workspace Management

ToolDescription
init_workspaceRegisters a project folder, assigns a UUID, and creates a pending brief file. Idempotent; safe to run multiple times.
list_workspacesLists all known workspaces that have a brief or stored memories.
resolve_workspaceResolves a workspace identifier (UUID, short-UUID, or display name) to its filesystem path.
merge_workspacesConsolidates duplicate workspace IDs into one. Irreversible.
debug_workspace_idDiagnostic tool; shows what workspace ID would be generated from a given path.

Memory & Briefs

ToolDescription
scan_workspaceStarts a background scan. Returns immediately; use scan_status to track progress. Walks the directory, respects .memignore, saves all readable text files to persistent memory. Idempotent and self-cleaning. Concurrent (4 workers, batch writes).
scan_statusReturns progress of a running or recently completed background scan: files scanned, entities indexed, errors, elapsed time, brief status.
save_to_memoryManually saves an architectural decision, fact, or technical note with an optional category tag.
list_memoryLists stored memories for a workspace, optionally filtered by category.
query_memoryRetrieves relevant context via vector search and synthesises an answer via Ollama or DeepSeek (auto-routed). Returns the answer as plain text plus a trailing Sources: block of file:line citations with relevance scores (L2 distance or rerank).
get_briefRetrieves the current project brief from .brain/contexts/.
update_briefManually updates the markdown content of a project brief.

Usage & Diagnostics

ToolDescription
get_token_usageReturns DeepSeek API token usage and cost statistics.
get_cost_reportGenerates a cost breakdown by operation type. Prepends a live PEAK / OFF-PEAK banner showing currently active rates.
get_cache_statsShows cache hit/miss rates by operation type.
purge_usage_dataDeletes historical token tracking records.

Project Brief Matrix

When a workspace is scanned, Zerikai compiles a dense 1,000–1,200 token project brief across 9 locked components:

SectionWhat It Captures
1. OverviewProject domain, primary type, and functional scope.
2. Technical StackBackend engines, databases, integrations, and core libraries.
3. Core ArchitectureInteractivity between frontend, backend, and processing layers.
4. Primary ConventionsLocal code styling, custom error handling, and validation schema rules.
5. PurposeBusiness logic problems solved and key underlying objectives.
6. Key FilesDefinitive app entry points, central routers, and specific domain tasks.
7. Dev & TestingEnvironment installation setups, execution triggers, and testing runs.
8. Data FlowComplete systemic request lifecycle tracing from gateway to database layer.
9. Future RoadmapPlanned engineering steps and dangling TODO items parsed directly from code.

Memory Modes

Adjust your operation profile via the MEMORY_MODE environment toggle to balance privacy, speed, and API costs.

πŸ’‘ Deterministic First: All high-resolution code parsing (functions, classes, methods) is performed locally and deterministically using Tree-Sitter for $0 cost. The engines below are only used for text-file fallbacks and generating the architectural Project Brief.

ModeAnalysis & BriefsQuery EngineTotal CostIdeal Use Case
🟒 cloudDeepSeekDeepSeekLowRecommended. High-fidelity architectural briefs.
🟑 hybridOllamaOllama + DeepSeekLowestLocal privacy with cloud reasoning escalation.
πŸ”΄ localOllamaOllama$0.00100% air-gapped hardware-local tracking.

Configuration Reference

KeyDefaultDescription
DEEPSEEK_API_KEYRequiredActive API authorization key from platform.deepseek.com.
MEMORY_MODEcloudSets target engines: choices include cloud, hybrid, or local.
ENABLE_TOKEN_TRACKINGtrueCalculates continuous usage and outputs summaries to SQLite.
QUERY_DISTANCE_THRESHOLD1.5Sets L2 vector distance cutoff limits. Lower inputs restrict matches.
ENABLE_LEXICAL_RERANKfalseActivates secondary hybrid reordering layer via keyword matching.
SKIP_BARE_FILES[]Extension list to bypass when tree-sitter finds zero valid code entities.

DeepSeek Pricing

DeepSeek uses peak / off-peak pricing across all tiers. Off-peak rates are exactly half of peak rates.

⏰ Peak hours (UTC): 01:00–04:00 and 06:00–10:00, Monday–Friday only. Weekends are always off-peak.

Rate Table (USD per 1M tokens)

Tierdeepseek-flash inputdeepseek-flash outputdeepseek-flash cachedv4-pro inputv4-pro outputv4-pro cached
Peak$0.30$1.20$0.006$1.32$3.96$0.044
Off-peak$0.15$0.60$0.003$0.66$1.98$0.022

The tool automatically resolves the correct tier at call time β€” no manual configuration needed.

Off-Peak Windows by Region

⚠️ Peak hours apply Monday–Friday only. Weekends are always off-peak regardless of time. Offsets shown for summer / daylight saving time (DST). In winter, US timezones shift 1 hour later; European zones shift 1 hour earlier β€” meaning off-peak windows shift accordingly.

RegionUTC offset (summer)Peak local time (Mon–Fri)βœ… Off-peak local time
EST (New York, Miami)UTCβˆ’58pm–11pm & 1am–5am5am–8pm and 11pm–1am (+ all weekend)
CST (Chicago, Dallas)UTCβˆ’67pm–10pm & midnight–4am4am–7pm and 10pm–midnight (+ all weekend)
PST (Los Angeles, Seattle)UTCβˆ’85pm–8pm & 10pm–2am2am–5pm and 8pm–10pm (+ all weekend)
Ireland (Dublin)UTC+12am–5am & 7am–11am11am–2am and 5am–7am (+ all weekend)
Spain (Madrid)UTC+23am–6am & 8am–noonnoon–3am and 6am–8am (+ all weekend)
Germany (Berlin)UTC+23am–6am & 8am–noonnoon–3am and 6am–8am (+ all weekend)
Norway (Oslo)UTC+23am–6am & 8am–noonnoon–3am and 6am–8am (+ all weekend)

Key insight: For US users working standard hours on weekdays, most of the working day is already off-peak (cheaper). European users in GMT+2 zones benefit from off-peak pricing through most of the afternoon and evening β€” and all weekend queries cost even less. Saturday and Sunday queries are always billed at off-peak rates.


Auxiliary Scripts & Troubleshooting

Workspace Reset

If you accidentally execute a workspace crawl before setting up your .memignore configurations, run the auxiliary wipe script to delete stale workspace data:

# Windows
.\venv\Scripts\python.exe drop_memory.py "Workspace Name"

# macOS / Linux
venv/bin/python drop_memory.py "Workspace Name"

Live Log Diagnostics

Monitor server activity, runtime operations, and auto-routing logs inside .brain/server.log:

# Live stream logs (macOS/Linux)
tail -f .brain/server.log

# Live stream logs (Windows PowerShell)
Get-Content .brain\server.log -Wait -Tail 30


Security & Data Privacy

  • All active vector spaces, tracking registries, and context details reside directly on your local machine.
  • Add .env and .brain/ explicitly to your global or project .gitignore patterns to prevent API keys and secure indexes from leaking to version control platforms.

To read more about the underlying design principles, architecture decisions, and future roadmap for Zerikai Memory, check out the insight article.

License

MIT License Β© Zerikai

πŸ› οΈ Support: This project is provided as-is for personal use.

ai
ai-tools
chromadb
deepseek-v4-pro
jev
lexical
mcp
mcp-memory
mcp-server
mcp-tools
memory
mistral-7b
ollama
ornith
rag-memory
re-ranking
tree-sitter-parser
typesafe-ai
typesafe-jev
zerikai

Languages

Python

100.0%