JackalopeTechnologies/SaddleRAG

A documentation RAG system — scrape, index, and search documentation via MCP tools for AI assistants

C#

2

826 commits

updated Aug 21, 2026

See the code

README

SaddleRAG

Documentation Retrieval-Augmented Generation for AI coding assistants.

SaddleRAG scrapes documentation websites, classifies and chunks the content with a local LLM, generates vector embeddings, and stores everything in MongoDB. It exposes the indexed documentation through MCP (Model Context Protocol) tools so that AI assistants like Claude Code, GitHub Copilot, and others can search your documentation library in real time.

Why SaddleRAG?

AI coding assistants are limited by their training cutoff and context window. When you're working with a niche library, a new release, or internal documentation, the assistant doesn't know about it. SaddleRAG bridges that gap:

  • Scrape any documentation site into a searchable vector database
  • Index local PDF, DOCX, Markdown, text, and HTML document libraries with citations and subject classification
  • Auto-index project dependencies from NuGet, npm, and pip
  • Serve documentation to your AI assistant via MCP tools during coding sessions
  • Track multiple versions of the same library and diff changes between them
  • Share a company-wide documentation database across your team

Architecture

Documentation Sites          SaddleRAG Pipeline                    AI Assistants

docs.example.com  --+
                    |      +-------------+
github.com/repo   --+-->   |  Playwright  |  (headless browser)
                    |      |   Crawler    |
learn.microsoft   --+      +------+------+
                                  |
                           +------v------+
                           |   Ollama    |  (local LLM)
                           |  Classifier |  phi4-mini:3.8b
                           +------+------+
                                  |
                           +------v------+
                           |  Category-  |
                           |   Aware     |
                           |  Chunker    |
                           +------+------+
                                  |
                           +------v------+
                           |   Ollama    |  nomic-embed-text
                           |  Embedder   |  (768 dimensions)
                           +------+------+
                                  |
                           +------v------+     +--------------+
                           |   MongoDB   |<--->|  MCP Server  |--> Claude Code
                           |  (storage)  |     |   (HTTP)     |--> Copilot
                           +-------------+     +--------------+--> Any MCP client

Quick Start (Windows Installer)

The fastest way to get SaddleRAG running is the MSI installer from GitHub Releases. It installs SaddleRAG as a Windows service, configures connections to MongoDB and Ollama, and starts automatically.

Website indexing requires two free, open-source prerequisites: MongoDB and Ollama. Local PDF and DOCX libraries additionally use an optional, user-operated Docling Serve endpoint. Tesseract is an optional, separately user-installed OCR engine; SaddleRAG does not select it for Docling.

Step 1: Install MongoDB Community Edition (free)

MongoDB stores all scraped documentation, chunks, and vector embeddings.

  1. Download the Community Edition from mongodb.com/try/download/community
  2. Run the installer, choose Complete setup type
  3. Keep the default settings: port 27017, Run as a Service checked
  4. After install, verify it's running: open a terminal and run mongosh -- you should see a connection prompt

Using Docker or a remote server? No problem. The SaddleRAG installer lets you enter any MongoDB connection string (e.g. mongodb://your-server:27017). You can also run MongoDB in Docker: docker run -d -p 27017:27017 --name saddlerag-mongo mongo:latest

Step 2: Install Ollama (free)

Ollama runs AI models locally for document classification and embedding generation. No API keys or cloud accounts needed.

  1. Download from ollama.com
  2. Run the installer -- Ollama runs as a background service on port 11434
  3. After install, verify it's running: open a terminal and run ollama list

SaddleRAG automatically pulls the required models on first use:

  • nomic-embed-text -- generates vector embeddings (768 dimensions)
  • phi4-mini:3.8b -- classifies documentation pages and optional re-ranking

Running Ollama elsewhere? The SaddleRAG installer lets you point to any Ollama endpoint (e.g. http://your-gpu-server:11434).

Step 3: Enable local PDF and DOCX libraries (optional)

SaddleRAG does not install, license, configure, or upgrade Docling Serve or Tesseract. You install and operate them only if you want PDF/DOCX ingestion. If you register a Docling start command in the SaddleRAG tray, the tray can run it as you, at your request; the SaddleRAG MCP service never starts, stops, or restarts anything. Markdown, text, HTML, and website ingestion do not require either tool.

  1. Review the current official Docling Serve deployment guide and latest Docling Serve release.

  2. Install Docling Serve in a user-owned Python environment. A simple Windows setup is:

    py -3.12 -m venv .venv
    .\.venv\Scripts\python -m pip install --upgrade "docling-serve[ui]"
    .\.venv\Scripts\docling-serve run
    

    Keep that process running when SaddleRAG scans PDF or DOCX files. A local server normally listens at http://localhost:5001.

  3. Tesseract is optional. SaddleRAG asks Docling to perform OCR and sends an OCR-engine selection only when you set DocumentIngestion:Docling:OcrEngine; while that setting is empty the request omits the field entirely and Docling uses its own configured/default OCR behavior, so installing Tesseract alone does not change how documents are converted. If you deliberately configure your user-owned Docling environment to use Tesseract, follow the official Tesseract installation guide; on Windows, that guide directs users to the current UB Mannheim installers. Include the required language data, add the Tesseract program directory to PATH if necessary, and set TESSDATA_PREFIX to the installed tessdata directory with a trailing path separator (for example, C:\Program Files\Tesseract-OCR\tessdata\). Restart your user-owned Docling Serve process after changing its OCR environment.

  4. In the SaddleRAG installer, leave the endpoint at http://localhost:5001 or enter your own Docling Serve URL. Click Test Docling. For an endpoint that does not require an API key, the test checks health, model readiness, and a bounded asynchronous conversion of a SaddleRAG-owned PDF before reporting DOCLING_READY.

If conversions are slower than you expect, see Making Docling faster. Docling installs a CPU-only PyTorch by default, so an NVIDIA GPU sits idle until you replace it with a CUDA build; that guide also covers thread and batch settings, and the fact that Tesseract is CPU-only while RapidOCR's torch backend is the GPU-capable OCR path.

The installer test is deliberately unauthenticated and never asks for, collects, or stores secrets. For an API-key-protected endpoint, configure DocumentIngestion:Docling:ApiKey through an access-restricted runtime configuration source, then use SaddleRAG's runtime status/conversion probe to verify it. Neither probe installs either prerequisite, tests private documents, or proves that Tesseract is the OCR engine selected inside Docling. Use a harmless image-only scanned PDF and the user-owned Docling logs to verify a specific OCR-engine configuration.

Step 4: Install SaddleRAG

  1. Download SaddleRAG.Mcp-*.msi from the latest release
  2. Run the installer
  3. MongoDB Configuration -- the installer defaults to mongodb://localhost:27017 with database SaddleRAG. Use the Test Connection button to verify MongoDB is reachable. If your MongoDB is on a different host, enter the connection string. Reset to Local Defaults reverts to the standard local settings.
  4. Ollama Configuration -- defaults to http://localhost:11434. Use Test Connection to verify. Change only if Ollama is running on another machine.
  5. Optional Document Ingestion -- enter a user-managed, unauthenticated Docling Serve endpoint and click Test Docling. The page links to current Docling and Tesseract setup guidance but never installs or controls either product and never collects secrets. You can continue without Docling when PDF/DOCX ingestion is not needed.
  6. Click Install -- files are copied to Program Files\SaddleRAG\SaddleRAG.Mcp, your connection settings are written to appsettings.json, and the SaddleRAGMcp Windows service starts automatically.

Don't have the prerequisites yet? The installer opens the official MongoDB, Ollama, Docling, and Tesseract guidance in your browser. You remain in control of every external installation. Return to the installer and use its connection tests when ready.

Step 5: Connect Your AI Assistant

The MSI installer wires SaddleRAG into all your installed AI tools automatically:

  • Claude Code — adds a user-level MCP server entry and refreshes the SaddleRAG skills in ~/.claude/skills
  • Claude Desktop — adds a server entry to claude_desktop_config.json
  • VSCode (GitHub Copilot Chat MCP) — installs a local SaddleRAG agent plugin under your VS Code user profile and updates %APPDATA%\Code\User\settings.json so Copilot can load the plugin and your per-user skill folders in all workspaces
  • GitHub Copilot CLI — updates ~/.copilot/mcp-config.json and refreshes the SaddleRAG skills in ~/.copilot/skills

No manual .mcp.json editing required. If a tool is not installed, its registration is skipped silently. Re-running the installer or SaddleRAG.Cli register-clients updates existing entries in place, refreshes the local VS Code SaddleRAG plugin, and replaces same-name SaddleRAG skill files with the current packaged content.

Step 6: Verify

Open your AI assistant and ask it to list libraries:

"Use the list_libraries tool to show what documentation is indexed."

If SaddleRAG is running, you'll get an empty list (nothing indexed yet). Then try:

"Scrape the documentation at https://docs.example.com for me."

The assistant will use the scrape_docs tool to index the site.

Verify the Service

  • Health check: visit http://localhost:6100/health in a browser
  • Service status: run Get-Service SaddleRAGMcp in PowerShell
  • Logs: check %ProgramData%\SaddleRAG\logs\ or use the get_server_logs MCP tool

Quick Start (Developer / Build from Source)

If you want to build and run from source instead of the MSI:

Prerequisites

DependencyVersionPurpose
.NET SDK10.0+Build and run
MongoDB6.0+Document storage (port 27017)
OllamaLatestLocal LLM for embeddings (port 11434)
Docling ServeOptional; latest release must pass SaddleRAG's probePDF and DOCX extraction (port 5001 by default)
TesseractOptionalUser-configured OCR engine; installing it alone does not select it for SaddleRAG or Docling

Build and Run

git clone https://github.com/JackalopeTechnologies/saddlerag.git
cd SaddleRAG
dotnet build SaddleRAG.slnx
dotnet run --project SaddleRAG.Mcp

The server starts on http://localhost:6100 by default. Configuration is in SaddleRAG.Mcp/appsettings.Development.json.

Running Tests with Coverage

The repo has a pinned dotnet-reportgenerator-globaltool in .config/dotnet-tools.json and two helper scripts that wrap dotnet test --collect:"XPlat Code Coverage" and produce an HTML report.

# Windows
scripts/coverage.ps1
# Linux / macOS
scripts/coverage.sh

Both scripts:

  1. Restore the local tool manifest (reportgenerator)
  2. Run the test project with coverage collection into ./coverage-results
  3. Generate an HTML report in ./coverage-results/html/index.html, print the text summary, and (by default) open it in your browser
  4. Exit 0 on a successful test run — no coverage gate is enforced

Pass --no-open (bash) or -NoOpen (PowerShell) to skip the browser launch; pass --filter <expr> / -Filter <expr> to override the default Category!=Integration xUnit filter.

CI collects coverage from two jobs — build-linux (unit) and integration-test-linux (Mongo / Playwright / ONNX integration) — and merges them in a coverage-report job. The merged summary is rendered on the workflow run page (via $GITHUB_STEP_SUMMARY) and posted as a sticky comment on each PR; the full cobertura XML and HTML drill-down are uploaded as a workflow artifact for download.

Connect Your AI Assistant

For the recommended per-user setup across all projects, run:

SaddleRAG.Cli register-clients

This writes the per-user client config for the supported tools, refreshes the SaddleRAG skill folders, and for VS Code also installs or updates the local SaddleRAG Copilot plugin plus the user settings that enable plugin and skill discovery.

On Linux, run SaddleRAG.Cli register-clients as the desktop user who runs VS Code or Copilot, not as root. The install.sh server install is system-wide, but the VS Code/Copilot integration is per-user.

If you only want a workspace-local MCP entry for the current repo, add this to .mcp.json in your project root:

{
  "mcpServers": {
    "saddlerag": {
      "type": "http",
      "url": "http://localhost:6100/mcp",
      "timeout": 60
    }
  }
}

The workspace-local .mcp.json path is useful for repo-specific setup, but it does not install the per-user VS Code plugin, does not refresh per-user skills, and does not configure Copilot to discover the SaddleRAG integration automatically.

Linux / Docker

Requires Docker with Compose. Tested on Ubuntu 22.04+.

Start the stack:

docker compose up -d

Download models (one-time, ~3 GB):

./warmup.sh

On first run, warmup.sh downloads ONNX embedding and reranker models from HuggingFace and the Ollama classification model (phi4-mini:3.8b). Models are stored in named Docker volumes and are not re-downloaded on restart.

The optional recon model (phi4:14b, ~8 GB) can be pulled separately:

docker compose exec ollama ollama pull phi4:14b

Access: http://localhost:6100

Logs: docker compose logs -f saddlerag

Stop: docker compose down (data preserved). docker compose down -v deletes all volumes including downloaded models — use with caution.

Bare-metal (Ubuntu/Debian or Rocky/RHEL)

curl -fsSL https://github.com/JackalopeTechnologies/SaddleRAG/releases/latest/download/install.sh | sudo bash

The script installs .NET ASP.NET Core Runtime 10, MongoDB 8, Ollama, and SaddleRAG. It registers a systemd service and downloads models during install (same prewarm step as the Windows MSI).

Uninstall: sudo /opt/saddlerag/uninstall.sh

Model warmup behaviour

SaddleRAG downloads models on first run, not at image build time. This applies to both Docker and bare-metal:

PlatformWhen models download
Windows (MSI)During MSI install (prewarm custom action)
DockerWhen you run ./warmup.sh after docker compose up -d
Bare-metal LinuxDuring install.sh (automatic)

Warmup sequence (logged to stdout / journalctl):

[Warmup] MongoDB profiles discovered
[Warmup] Ollama bootstrap — pulls phi4-mini:3.8b if absent
[Warmup] ONNX models ready — downloads nomic-embed-text-v1.5 + mxbai-rerank-base-v1 from HuggingFace
[Warmup] Vector indices loaded
[Warmup] Full pipeline warm

Managing AI Client Registrations

The SaddleRAG.Cli tool manages which AI tools SaddleRAG is wired into.

Check registration status

SaddleRAG.Cli clients-status

Reports whether SaddleRAG is registered in each supported tool, with config file paths and any errors.

Re-register after adding a new tool

SaddleRAG.Cli register-clients

Registers SaddleRAG in all installed AI tools' per-user config. Safe to run multiple times — existing entries are updated in place, same-name SaddleRAG skills are replaced with the current packaged versions, and the VS Code registration refreshes Copilot autostart plus skill-discovery settings.

For VS Code specifically, register-clients installs or updates a local SaddleRAG agent plugin and updates %APPDATA%\Code\User\settings.json so the plugin is enabled across workspaces without per-project setup.

Disable SaddleRAG for specific tools

SaddleRAG.Cli unregister-clients --claude-desktop=true --claude-code=false

Each flag controls one tool. Omitted flags default to true (remove from that tool). The example above removes SaddleRAG from Claude Desktop but leaves Claude Code wired.

MCP Tools Reference

SaddleRAG exposes 33 tools through the MCP protocol. Six load eagerly into every session; the rest are deferred and pulled in by ToolSearch when needed.

Entry-point tools (eager — in every session)

ToolDescription
get_dashboard_indexStart here in any fresh session. Returns a single-call status overview: library/version counts, recent scrape jobs, server health
list_librariesList all indexed libraries with current version and all ingested versions
search_docsNatural language search across all libraries or filtered by library, version, and category
get_class_referenceLook up API reference for a class or type by name — exact match, then fuzzy
get_library_overviewGet Overview-category chunks for a library: concepts, architecture, getting-started guides
list_symbolsList documented symbols for a library, optionally filtered by kind (class, enum, function, parameter)

Ingestion

ToolDescription
start_ingestSingle ingestion entry point — inspects (library, version) state and returns the next recommended action
scrape_docsScrape a documentation URL with auto-derived crawl settings. Cache-aware: skips already-indexed libraries unless force=true. Use for first-time ingest or URL/pattern overrides
rescrape_libraryRe-scrape an already-indexed library from its source. Takes library + version only — reuses the prior scrape's config and seeds the crawler from stored page URLs so dead/changed/new pages are all picked up
dryrun_scrapeTest a scrape configuration without writing to the database. Reports page counts, depth distribution, and GitHub repos that would be cloned
index_project_dependenciesScan a project's NuGet/npm/pip dependencies and auto-index their documentation

Job management

ToolDescription
get_scrape_statusPoll a scrape job's progress by job ID
list_scrape_jobsList recent scrape jobs with status, most recent first
cancel_jobCancel a running scrape, dryrun, rechunk, reembed, or reextract job

Library & pages

ToolDescription
list_pagesList the URLs of every page indexed for a (library, version) — useful for auditing scrape completeness
add_pageFetch a single URL and add it to an existing (library, version) index without re-crawling

Version management

ToolDescription
get_version_changesDiff two versions of a library — added, removed, and changed pages with summaries

Health

ToolDescription
get_library_healthPer-version diagnostic snapshot: chunk count, hostname distribution, language mix, boundary-issue rate, suspect markers

Library administration

ToolDescription
rename_libraryRename a library across every collection. Defaults to dryRun=true — preview before committing
delete_versionHard-delete one (library, version) with all its chunks, pages, indexes, and profile. Defaults to dryRun=true
delete_libraryHard-delete an entire library across every collection. Defaults to dryRun=true

Index maintenance

ToolDescription
rechunk_libraryRe-run the chunker over stored pages, replace all chunks, and re-embed. Requires reextract_library as a follow-up
reembed_libraryRe-embed every stored chunk via the current embedding provider; updates the version's provider/model/dimensions. Use after swapping embedding provider or model
reextract_libraryRe-run the symbol extractor and classifier over existing chunks without re-crawling or re-embedding
recon_libraryGet the instructions and JSON schema needed to characterize a docs site before scraping (LLM-assisted reconnaissance)
submit_library_profileSubmit the reconnaissance JSON produced by recon_library to persist it as the LibraryProfile

Symbol management

ToolDescription
list_excluded_symbolsList symbols on the extraction stoplist for a library
add_to_likely_symbolsAdd a symbol to the high-confidence list (overrides heuristic rejection)
add_to_stoplistAdd a symbol to the stoplist so it is excluded from future extraction passes

URL correction

ToolDescription
submit_url_correctionSubmit a corrected canonical URL for a page that was indexed under a redirect or wrong URL

Configuration

ToolDescription
list_profilesList all configured MongoDB database profiles
reload_profileReload the in-memory vector index from MongoDB (useful after manual data changes)

Settings

ToolDescription
set_rerank_strategySet the reranker strategy at runtime: Off, Llm, or CrossEncoder
toggle_loggingToggle verbose request logging without restarting the server

Diagnostics

ToolDescription
get_server_logsRetrieve recent server log lines, with optional text filter

CLI Tool

The CLI provides direct access to ingestion and management without the MCP server.

dotnet build SaddleRAG.Cli/SaddleRAG.Cli.csproj

Commands

Ingest a documentation library:

saddlerag ingest \
  --root-url https://docs.example.com/ \
  --library-id example-lib \
  --version 2.0 \
  --hint "Example library for building widgets" \
  --allowed "docs.example.com" \
  --max-pages 500 \
  --delay 1000

Dry-run a scrape (no database writes):

saddlerag dryrun \
  --root-url https://docs.example.com/ \
  --allowed "docs.example.com" \
  --max-pages 200

Inspect a page's link/sidebar structure (useful for tuning URL patterns):

saddlerag inspect --url https://docs.example.com/getting-started

List indexed libraries:

saddlerag list

Show ingestion status:

saddlerag status --library-id example-lib

Re-classify pages with the LLM (fix unclassified pages):

saddlerag reclassify --library-id example-lib
saddlerag reclassify --all  # Reclassify everything, even already-classified pages

Scan project dependencies and auto-index:

saddlerag scan --path ./MyProject.sln
saddlerag scan --path ./package.json --profile company

Manage database profiles:

saddlerag profile list

Configuration

MongoDB Profiles

SaddleRAG supports multiple MongoDB databases via named profiles. Configure them in appsettings.json:

{
  "MongoDB": {
    "ActiveProfile": "local",
    "Profiles": {
      "local": {
        "ConnectionString": "mongodb://localhost:27017",
        "DatabaseName": "SaddleRAG",
        "Description": "Local development database"
      },
      "company": {
        "ConnectionString": "mongodb://saddlerag.internal.company.com:27017",
        "DatabaseName": "SaddleRAG",
        "Description": "Shared company documentation database"
      }
    }
  }
}

Every MCP tool accepts an optional profile parameter to target a specific database. This enables scenarios like:

  • Personal local index for experiments
  • Shared team database with pre-indexed company libraries
  • CI/CD pipeline that indexes docs on release

Ollama Settings

{
  "Ollama": {
    "Endpoint": "http://localhost:11434",
    "EmbeddingModel": "nomic-embed-text",
    "EmbeddingDimensions": 768,
    "ClassificationModel": "phi4-mini:3.8b",
    "ReRankingModel": "phi4-mini:3.8b",
    "ModelPullTimeoutSeconds": 600
  }
}

Environment Variables

All settings can be overridden via environment variables prefixed with SADDLERAG_:

SADDLERAG_MONGODB_PROFILE=company          # Override active profile
ASPNETCORE_ENVIRONMENT=Development      # Enable dev settings (disables re-ranking)

Troubleshooting

SaddleRAG isn't visible in Claude Code / Claude Desktop / VSCode / Copilot

Run the diagnostics command:

SaddleRAG.Cli clients-status

Then re-register:

SaddleRAG.Cli register-clients

If registration succeeds but the tool still doesn't appear, restart the AI tool so it picks up the new config. In VS Code, trust and enabled-state are stored separately from mcp.json; you may need to trust or enable the server once the first time, but after that the user-level autostart setting should start it automatically in future Copilot sessions.

When using the local SaddleRAG plugin path in VS Code, there should not be a separate MCP trust prompt for SaddleRAG because plugin-provided MCP servers are trusted as part of the plugin install. If it still does not appear, check that agent plugins are enabled and that the local plugin is enabled in the Agent Plugins view.

I want to disable SaddleRAG for one specific tool

Use unregister-clients with explicit flags. Each flag controls one tool; unspecified tools are also unregistered by default, so be explicit:

# Remove from Claude Desktop only, leave everything else registered
SaddleRAG.Cli unregister-clients --claude-desktop=true --claude-code=false --vscode=false --copilot-cli=false

To re-enable a tool later, run register-clients (it re-wires all installed tools).

The MCP server health check fails

Visit http://localhost:6100/health. If it returns an error, check the service:

Get-Service SaddleRAGMcp
Start-Service SaddleRAGMcp

Logs are in %ProgramData%\SaddleRAG\logs\.

Releasing

To create a new release with an MSI installer:

git tag v1.0.0
git push origin v1.0.0

The tag workflow builds and tests Windows and Linux, packages the MSI, desktop extension, and Linux archive, pushes the release container images, and uploads the assets to a draft GitHub Release. Verify the workflow and assets, then publish that draft.

Project Structure

SaddleRAG.slnx                    # Solution file
SaddleRAG.Core/                   # Domain models, interfaces, enums
SaddleRAG.Database/               # MongoDB repositories and context factory
SaddleRAG.Ingestion/              # Scraping, classification, chunking, embedding pipeline
  Crawling/                    #   Playwright web crawler + GitHub repo scraper
  Classification/              #   Ollama LLM page classifier
  Chunking/                    #   Category-aware semantic chunker
  Embedding/                   #   Ollama embedding provider
  Symbols/                     #   Symbol extraction and stoplist management
  Recon/                       #   LLM-assisted library profiling (recon/reextract)
  Scanning/                    #   Project dependency scanner
  Ecosystems/                  #   NuGet, npm, pip registry clients
SaddleRAG.Mcp/                    # ASP.NET Core MCP server (HTTP transport)
  Tools/                       #   33 MCP tool definitions across 18 files
SaddleRAG.Cli/                    # Command-line interface (ingest, status, register-clients, ...)
SaddleRAG.Installer/              # WiX MSI installer definition
SaddleRAG.Tests/                  # Integration and unit tests

Support

SaddleRAG is free and MIT-licensed. If it's saving you time and you'd like to chip in, GitHub Sponsors is the easiest way to do that — any amount helps and there are no perks gated behind it. Totally optional; stars and bug reports are appreciated just as much.

License

SaddleRAG is released under the MIT License. You may use, copy, modify, merge, publish, distribute, sublicense, and sell the software, including for commercial purposes, subject only to the requirement that the copyright notice and the MIT permission notice be included in copies or substantial portions of the software.

Contributions are welcome — see CONTRIBUTING.md. No contributor agreement is required.

JackalopeTechnologies/SaddleRAG

A documentation RAG system — scrape, index, and search documentation via MCP tools for AI assistants

C#

2

826 commits

updated Aug 21, 2026

See the code

README

SaddleRAG

Documentation Retrieval-Augmented Generation for AI coding assistants.

SaddleRAG scrapes documentation websites, classifies and chunks the content with a local LLM, generates vector embeddings, and stores everything in MongoDB. It exposes the indexed documentation through MCP (Model Context Protocol) tools so that AI assistants like Claude Code, GitHub Copilot, and others can search your documentation library in real time.

Why SaddleRAG?

AI coding assistants are limited by their training cutoff and context window. When you're working with a niche library, a new release, or internal documentation, the assistant doesn't know about it. SaddleRAG bridges that gap:

  • Scrape any documentation site into a searchable vector database
  • Index local PDF, DOCX, Markdown, text, and HTML document libraries with citations and subject classification
  • Auto-index project dependencies from NuGet, npm, and pip
  • Serve documentation to your AI assistant via MCP tools during coding sessions
  • Track multiple versions of the same library and diff changes between them
  • Share a company-wide documentation database across your team

Architecture

Documentation Sites          SaddleRAG Pipeline                    AI Assistants

docs.example.com  --+
                    |      +-------------+
github.com/repo   --+-->   |  Playwright  |  (headless browser)
                    |      |   Crawler    |
learn.microsoft   --+      +------+------+
                                  |
                           +------v------+
                           |   Ollama    |  (local LLM)
                           |  Classifier |  phi4-mini:3.8b
                           +------+------+
                                  |
                           +------v------+
                           |  Category-  |
                           |   Aware     |
                           |  Chunker    |
                           +------+------+
                                  |
                           +------v------+
                           |   Ollama    |  nomic-embed-text
                           |  Embedder   |  (768 dimensions)
                           +------+------+
                                  |
                           +------v------+     +--------------+
                           |   MongoDB   |<--->|  MCP Server  |--> Claude Code
                           |  (storage)  |     |   (HTTP)     |--> Copilot
                           +-------------+     +--------------+--> Any MCP client

Quick Start (Windows Installer)

The fastest way to get SaddleRAG running is the MSI installer from GitHub Releases. It installs SaddleRAG as a Windows service, configures connections to MongoDB and Ollama, and starts automatically.

Website indexing requires two free, open-source prerequisites: MongoDB and Ollama. Local PDF and DOCX libraries additionally use an optional, user-operated Docling Serve endpoint. Tesseract is an optional, separately user-installed OCR engine; SaddleRAG does not select it for Docling.

Step 1: Install MongoDB Community Edition (free)

MongoDB stores all scraped documentation, chunks, and vector embeddings.

  1. Download the Community Edition from mongodb.com/try/download/community
  2. Run the installer, choose Complete setup type
  3. Keep the default settings: port 27017, Run as a Service checked
  4. After install, verify it's running: open a terminal and run mongosh -- you should see a connection prompt

Using Docker or a remote server? No problem. The SaddleRAG installer lets you enter any MongoDB connection string (e.g. mongodb://your-server:27017). You can also run MongoDB in Docker: docker run -d -p 27017:27017 --name saddlerag-mongo mongo:latest

Step 2: Install Ollama (free)

Ollama runs AI models locally for document classification and embedding generation. No API keys or cloud accounts needed.

  1. Download from ollama.com
  2. Run the installer -- Ollama runs as a background service on port 11434
  3. After install, verify it's running: open a terminal and run ollama list

SaddleRAG automatically pulls the required models on first use:

  • nomic-embed-text -- generates vector embeddings (768 dimensions)
  • phi4-mini:3.8b -- classifies documentation pages and optional re-ranking

Running Ollama elsewhere? The SaddleRAG installer lets you point to any Ollama endpoint (e.g. http://your-gpu-server:11434).

Step 3: Enable local PDF and DOCX libraries (optional)

SaddleRAG does not install, license, configure, or upgrade Docling Serve or Tesseract. You install and operate them only if you want PDF/DOCX ingestion. If you register a Docling start command in the SaddleRAG tray, the tray can run it as you, at your request; the SaddleRAG MCP service never starts, stops, or restarts anything. Markdown, text, HTML, and website ingestion do not require either tool.

  1. Review the current official Docling Serve deployment guide and latest Docling Serve release.

  2. Install Docling Serve in a user-owned Python environment. A simple Windows setup is:

    py -3.12 -m venv .venv
    .\.venv\Scripts\python -m pip install --upgrade "docling-serve[ui]"
    .\.venv\Scripts\docling-serve run
    

    Keep that process running when SaddleRAG scans PDF or DOCX files. A local server normally listens at http://localhost:5001.

  3. Tesseract is optional. SaddleRAG asks Docling to perform OCR and sends an OCR-engine selection only when you set DocumentIngestion:Docling:OcrEngine; while that setting is empty the request omits the field entirely and Docling uses its own configured/default OCR behavior, so installing Tesseract alone does not change how documents are converted. If you deliberately configure your user-owned Docling environment to use Tesseract, follow the official Tesseract installation guide; on Windows, that guide directs users to the current UB Mannheim installers. Include the required language data, add the Tesseract program directory to PATH if necessary, and set TESSDATA_PREFIX to the installed tessdata directory with a trailing path separator (for example, C:\Program Files\Tesseract-OCR\tessdata\). Restart your user-owned Docling Serve process after changing its OCR environment.

  4. In the SaddleRAG installer, leave the endpoint at http://localhost:5001 or enter your own Docling Serve URL. Click Test Docling. For an endpoint that does not require an API key, the test checks health, model readiness, and a bounded asynchronous conversion of a SaddleRAG-owned PDF before reporting DOCLING_READY.

If conversions are slower than you expect, see Making Docling faster. Docling installs a CPU-only PyTorch by default, so an NVIDIA GPU sits idle until you replace it with a CUDA build; that guide also covers thread and batch settings, and the fact that Tesseract is CPU-only while RapidOCR's torch backend is the GPU-capable OCR path.

The installer test is deliberately unauthenticated and never asks for, collects, or stores secrets. For an API-key-protected endpoint, configure DocumentIngestion:Docling:ApiKey through an access-restricted runtime configuration source, then use SaddleRAG's runtime status/conversion probe to verify it. Neither probe installs either prerequisite, tests private documents, or proves that Tesseract is the OCR engine selected inside Docling. Use a harmless image-only scanned PDF and the user-owned Docling logs to verify a specific OCR-engine configuration.

Step 4: Install SaddleRAG

  1. Download SaddleRAG.Mcp-*.msi from the latest release
  2. Run the installer
  3. MongoDB Configuration -- the installer defaults to mongodb://localhost:27017 with database SaddleRAG. Use the Test Connection button to verify MongoDB is reachable. If your MongoDB is on a different host, enter the connection string. Reset to Local Defaults reverts to the standard local settings.
  4. Ollama Configuration -- defaults to http://localhost:11434. Use Test Connection to verify. Change only if Ollama is running on another machine.
  5. Optional Document Ingestion -- enter a user-managed, unauthenticated Docling Serve endpoint and click Test Docling. The page links to current Docling and Tesseract setup guidance but never installs or controls either product and never collects secrets. You can continue without Docling when PDF/DOCX ingestion is not needed.
  6. Click Install -- files are copied to Program Files\SaddleRAG\SaddleRAG.Mcp, your connection settings are written to appsettings.json, and the SaddleRAGMcp Windows service starts automatically.

Don't have the prerequisites yet? The installer opens the official MongoDB, Ollama, Docling, and Tesseract guidance in your browser. You remain in control of every external installation. Return to the installer and use its connection tests when ready.

Step 5: Connect Your AI Assistant

The MSI installer wires SaddleRAG into all your installed AI tools automatically:

  • Claude Code — adds a user-level MCP server entry and refreshes the SaddleRAG skills in ~/.claude/skills
  • Claude Desktop — adds a server entry to claude_desktop_config.json
  • VSCode (GitHub Copilot Chat MCP) — installs a local SaddleRAG agent plugin under your VS Code user profile and updates %APPDATA%\Code\User\settings.json so Copilot can load the plugin and your per-user skill folders in all workspaces
  • GitHub Copilot CLI — updates ~/.copilot/mcp-config.json and refreshes the SaddleRAG skills in ~/.copilot/skills

No manual .mcp.json editing required. If a tool is not installed, its registration is skipped silently. Re-running the installer or SaddleRAG.Cli register-clients updates existing entries in place, refreshes the local VS Code SaddleRAG plugin, and replaces same-name SaddleRAG skill files with the current packaged content.

Step 6: Verify

Open your AI assistant and ask it to list libraries:

"Use the list_libraries tool to show what documentation is indexed."

If SaddleRAG is running, you'll get an empty list (nothing indexed yet). Then try:

"Scrape the documentation at https://docs.example.com for me."

The assistant will use the scrape_docs tool to index the site.

Verify the Service

  • Health check: visit http://localhost:6100/health in a browser
  • Service status: run Get-Service SaddleRAGMcp in PowerShell
  • Logs: check %ProgramData%\SaddleRAG\logs\ or use the get_server_logs MCP tool

Quick Start (Developer / Build from Source)

If you want to build and run from source instead of the MSI:

Prerequisites

DependencyVersionPurpose
.NET SDK10.0+Build and run
MongoDB6.0+Document storage (port 27017)
OllamaLatestLocal LLM for embeddings (port 11434)
Docling ServeOptional; latest release must pass SaddleRAG's probePDF and DOCX extraction (port 5001 by default)
TesseractOptionalUser-configured OCR engine; installing it alone does not select it for SaddleRAG or Docling

Build and Run

git clone https://github.com/JackalopeTechnologies/saddlerag.git
cd SaddleRAG
dotnet build SaddleRAG.slnx
dotnet run --project SaddleRAG.Mcp

The server starts on http://localhost:6100 by default. Configuration is in SaddleRAG.Mcp/appsettings.Development.json.

Running Tests with Coverage

The repo has a pinned dotnet-reportgenerator-globaltool in .config/dotnet-tools.json and two helper scripts that wrap dotnet test --collect:"XPlat Code Coverage" and produce an HTML report.

# Windows
scripts/coverage.ps1
# Linux / macOS
scripts/coverage.sh

Both scripts:

  1. Restore the local tool manifest (reportgenerator)
  2. Run the test project with coverage collection into ./coverage-results
  3. Generate an HTML report in ./coverage-results/html/index.html, print the text summary, and (by default) open it in your browser
  4. Exit 0 on a successful test run — no coverage gate is enforced

Pass --no-open (bash) or -NoOpen (PowerShell) to skip the browser launch; pass --filter <expr> / -Filter <expr> to override the default Category!=Integration xUnit filter.

CI collects coverage from two jobs — build-linux (unit) and integration-test-linux (Mongo / Playwright / ONNX integration) — and merges them in a coverage-report job. The merged summary is rendered on the workflow run page (via $GITHUB_STEP_SUMMARY) and posted as a sticky comment on each PR; the full cobertura XML and HTML drill-down are uploaded as a workflow artifact for download.

Connect Your AI Assistant

For the recommended per-user setup across all projects, run:

SaddleRAG.Cli register-clients

This writes the per-user client config for the supported tools, refreshes the SaddleRAG skill folders, and for VS Code also installs or updates the local SaddleRAG Copilot plugin plus the user settings that enable plugin and skill discovery.

On Linux, run SaddleRAG.Cli register-clients as the desktop user who runs VS Code or Copilot, not as root. The install.sh server install is system-wide, but the VS Code/Copilot integration is per-user.

If you only want a workspace-local MCP entry for the current repo, add this to .mcp.json in your project root:

{
  "mcpServers": {
    "saddlerag": {
      "type": "http",
      "url": "http://localhost:6100/mcp",
      "timeout": 60
    }
  }
}

The workspace-local .mcp.json path is useful for repo-specific setup, but it does not install the per-user VS Code plugin, does not refresh per-user skills, and does not configure Copilot to discover the SaddleRAG integration automatically.

Linux / Docker

Requires Docker with Compose. Tested on Ubuntu 22.04+.

Start the stack:

docker compose up -d

Download models (one-time, ~3 GB):

./warmup.sh

On first run, warmup.sh downloads ONNX embedding and reranker models from HuggingFace and the Ollama classification model (phi4-mini:3.8b). Models are stored in named Docker volumes and are not re-downloaded on restart.

The optional recon model (phi4:14b, ~8 GB) can be pulled separately:

docker compose exec ollama ollama pull phi4:14b

Access: http://localhost:6100

Logs: docker compose logs -f saddlerag

Stop: docker compose down (data preserved). docker compose down -v deletes all volumes including downloaded models — use with caution.

Bare-metal (Ubuntu/Debian or Rocky/RHEL)

curl -fsSL https://github.com/JackalopeTechnologies/SaddleRAG/releases/latest/download/install.sh | sudo bash

The script installs .NET ASP.NET Core Runtime 10, MongoDB 8, Ollama, and SaddleRAG. It registers a systemd service and downloads models during install (same prewarm step as the Windows MSI).

Uninstall: sudo /opt/saddlerag/uninstall.sh

Model warmup behaviour

SaddleRAG downloads models on first run, not at image build time. This applies to both Docker and bare-metal:

PlatformWhen models download
Windows (MSI)During MSI install (prewarm custom action)
DockerWhen you run ./warmup.sh after docker compose up -d
Bare-metal LinuxDuring install.sh (automatic)

Warmup sequence (logged to stdout / journalctl):

[Warmup] MongoDB profiles discovered
[Warmup] Ollama bootstrap — pulls phi4-mini:3.8b if absent
[Warmup] ONNX models ready — downloads nomic-embed-text-v1.5 + mxbai-rerank-base-v1 from HuggingFace
[Warmup] Vector indices loaded
[Warmup] Full pipeline warm

Managing AI Client Registrations

The SaddleRAG.Cli tool manages which AI tools SaddleRAG is wired into.

Check registration status

SaddleRAG.Cli clients-status

Reports whether SaddleRAG is registered in each supported tool, with config file paths and any errors.

Re-register after adding a new tool

SaddleRAG.Cli register-clients

Registers SaddleRAG in all installed AI tools' per-user config. Safe to run multiple times — existing entries are updated in place, same-name SaddleRAG skills are replaced with the current packaged versions, and the VS Code registration refreshes Copilot autostart plus skill-discovery settings.

For VS Code specifically, register-clients installs or updates a local SaddleRAG agent plugin and updates %APPDATA%\Code\User\settings.json so the plugin is enabled across workspaces without per-project setup.

Disable SaddleRAG for specific tools

SaddleRAG.Cli unregister-clients --claude-desktop=true --claude-code=false

Each flag controls one tool. Omitted flags default to true (remove from that tool). The example above removes SaddleRAG from Claude Desktop but leaves Claude Code wired.

MCP Tools Reference

SaddleRAG exposes 33 tools through the MCP protocol. Six load eagerly into every session; the rest are deferred and pulled in by ToolSearch when needed.

Entry-point tools (eager — in every session)

ToolDescription
get_dashboard_indexStart here in any fresh session. Returns a single-call status overview: library/version counts, recent scrape jobs, server health
list_librariesList all indexed libraries with current version and all ingested versions
search_docsNatural language search across all libraries or filtered by library, version, and category
get_class_referenceLook up API reference for a class or type by name — exact match, then fuzzy
get_library_overviewGet Overview-category chunks for a library: concepts, architecture, getting-started guides
list_symbolsList documented symbols for a library, optionally filtered by kind (class, enum, function, parameter)

Ingestion

ToolDescription
start_ingestSingle ingestion entry point — inspects (library, version) state and returns the next recommended action
scrape_docsScrape a documentation URL with auto-derived crawl settings. Cache-aware: skips already-indexed libraries unless force=true. Use for first-time ingest or URL/pattern overrides
rescrape_libraryRe-scrape an already-indexed library from its source. Takes library + version only — reuses the prior scrape's config and seeds the crawler from stored page URLs so dead/changed/new pages are all picked up
dryrun_scrapeTest a scrape configuration without writing to the database. Reports page counts, depth distribution, and GitHub repos that would be cloned
index_project_dependenciesScan a project's NuGet/npm/pip dependencies and auto-index their documentation

Job management

ToolDescription
get_scrape_statusPoll a scrape job's progress by job ID
list_scrape_jobsList recent scrape jobs with status, most recent first
cancel_jobCancel a running scrape, dryrun, rechunk, reembed, or reextract job

Library & pages

ToolDescription
list_pagesList the URLs of every page indexed for a (library, version) — useful for auditing scrape completeness
add_pageFetch a single URL and add it to an existing (library, version) index without re-crawling

Version management

ToolDescription
get_version_changesDiff two versions of a library — added, removed, and changed pages with summaries

Health

ToolDescription
get_library_healthPer-version diagnostic snapshot: chunk count, hostname distribution, language mix, boundary-issue rate, suspect markers

Library administration

ToolDescription
rename_libraryRename a library across every collection. Defaults to dryRun=true — preview before committing
delete_versionHard-delete one (library, version) with all its chunks, pages, indexes, and profile. Defaults to dryRun=true
delete_libraryHard-delete an entire library across every collection. Defaults to dryRun=true

Index maintenance

ToolDescription
rechunk_libraryRe-run the chunker over stored pages, replace all chunks, and re-embed. Requires reextract_library as a follow-up
reembed_libraryRe-embed every stored chunk via the current embedding provider; updates the version's provider/model/dimensions. Use after swapping embedding provider or model
reextract_libraryRe-run the symbol extractor and classifier over existing chunks without re-crawling or re-embedding
recon_libraryGet the instructions and JSON schema needed to characterize a docs site before scraping (LLM-assisted reconnaissance)
submit_library_profileSubmit the reconnaissance JSON produced by recon_library to persist it as the LibraryProfile

Symbol management

ToolDescription
list_excluded_symbolsList symbols on the extraction stoplist for a library
add_to_likely_symbolsAdd a symbol to the high-confidence list (overrides heuristic rejection)
add_to_stoplistAdd a symbol to the stoplist so it is excluded from future extraction passes

URL correction

ToolDescription
submit_url_correctionSubmit a corrected canonical URL for a page that was indexed under a redirect or wrong URL

Configuration

ToolDescription
list_profilesList all configured MongoDB database profiles
reload_profileReload the in-memory vector index from MongoDB (useful after manual data changes)

Settings

ToolDescription
set_rerank_strategySet the reranker strategy at runtime: Off, Llm, or CrossEncoder
toggle_loggingToggle verbose request logging without restarting the server

Diagnostics

ToolDescription
get_server_logsRetrieve recent server log lines, with optional text filter

CLI Tool

The CLI provides direct access to ingestion and management without the MCP server.

dotnet build SaddleRAG.Cli/SaddleRAG.Cli.csproj

Commands

Ingest a documentation library:

saddlerag ingest \
  --root-url https://docs.example.com/ \
  --library-id example-lib \
  --version 2.0 \
  --hint "Example library for building widgets" \
  --allowed "docs.example.com" \
  --max-pages 500 \
  --delay 1000

Dry-run a scrape (no database writes):

saddlerag dryrun \
  --root-url https://docs.example.com/ \
  --allowed "docs.example.com" \
  --max-pages 200

Inspect a page's link/sidebar structure (useful for tuning URL patterns):

saddlerag inspect --url https://docs.example.com/getting-started

List indexed libraries:

saddlerag list

Show ingestion status:

saddlerag status --library-id example-lib

Re-classify pages with the LLM (fix unclassified pages):

saddlerag reclassify --library-id example-lib
saddlerag reclassify --all  # Reclassify everything, even already-classified pages

Scan project dependencies and auto-index:

saddlerag scan --path ./MyProject.sln
saddlerag scan --path ./package.json --profile company

Manage database profiles:

saddlerag profile list

Configuration

MongoDB Profiles

SaddleRAG supports multiple MongoDB databases via named profiles. Configure them in appsettings.json:

{
  "MongoDB": {
    "ActiveProfile": "local",
    "Profiles": {
      "local": {
        "ConnectionString": "mongodb://localhost:27017",
        "DatabaseName": "SaddleRAG",
        "Description": "Local development database"
      },
      "company": {
        "ConnectionString": "mongodb://saddlerag.internal.company.com:27017",
        "DatabaseName": "SaddleRAG",
        "Description": "Shared company documentation database"
      }
    }
  }
}

Every MCP tool accepts an optional profile parameter to target a specific database. This enables scenarios like:

  • Personal local index for experiments
  • Shared team database with pre-indexed company libraries
  • CI/CD pipeline that indexes docs on release

Ollama Settings

{
  "Ollama": {
    "Endpoint": "http://localhost:11434",
    "EmbeddingModel": "nomic-embed-text",
    "EmbeddingDimensions": 768,
    "ClassificationModel": "phi4-mini:3.8b",
    "ReRankingModel": "phi4-mini:3.8b",
    "ModelPullTimeoutSeconds": 600
  }
}

Environment Variables

All settings can be overridden via environment variables prefixed with SADDLERAG_:

SADDLERAG_MONGODB_PROFILE=company          # Override active profile
ASPNETCORE_ENVIRONMENT=Development      # Enable dev settings (disables re-ranking)

Troubleshooting

SaddleRAG isn't visible in Claude Code / Claude Desktop / VSCode / Copilot

Run the diagnostics command:

SaddleRAG.Cli clients-status

Then re-register:

SaddleRAG.Cli register-clients

If registration succeeds but the tool still doesn't appear, restart the AI tool so it picks up the new config. In VS Code, trust and enabled-state are stored separately from mcp.json; you may need to trust or enable the server once the first time, but after that the user-level autostart setting should start it automatically in future Copilot sessions.

When using the local SaddleRAG plugin path in VS Code, there should not be a separate MCP trust prompt for SaddleRAG because plugin-provided MCP servers are trusted as part of the plugin install. If it still does not appear, check that agent plugins are enabled and that the local plugin is enabled in the Agent Plugins view.

I want to disable SaddleRAG for one specific tool

Use unregister-clients with explicit flags. Each flag controls one tool; unspecified tools are also unregistered by default, so be explicit:

# Remove from Claude Desktop only, leave everything else registered
SaddleRAG.Cli unregister-clients --claude-desktop=true --claude-code=false --vscode=false --copilot-cli=false

To re-enable a tool later, run register-clients (it re-wires all installed tools).

The MCP server health check fails

Visit http://localhost:6100/health. If it returns an error, check the service:

Get-Service SaddleRAGMcp
Start-Service SaddleRAGMcp

Logs are in %ProgramData%\SaddleRAG\logs\.

Releasing

To create a new release with an MSI installer:

git tag v1.0.0
git push origin v1.0.0

The tag workflow builds and tests Windows and Linux, packages the MSI, desktop extension, and Linux archive, pushes the release container images, and uploads the assets to a draft GitHub Release. Verify the workflow and assets, then publish that draft.

Project Structure

SaddleRAG.slnx                    # Solution file
SaddleRAG.Core/                   # Domain models, interfaces, enums
SaddleRAG.Database/               # MongoDB repositories and context factory
SaddleRAG.Ingestion/              # Scraping, classification, chunking, embedding pipeline
  Crawling/                    #   Playwright web crawler + GitHub repo scraper
  Classification/              #   Ollama LLM page classifier
  Chunking/                    #   Category-aware semantic chunker
  Embedding/                   #   Ollama embedding provider
  Symbols/                     #   Symbol extraction and stoplist management
  Recon/                       #   LLM-assisted library profiling (recon/reextract)
  Scanning/                    #   Project dependency scanner
  Ecosystems/                  #   NuGet, npm, pip registry clients
SaddleRAG.Mcp/                    # ASP.NET Core MCP server (HTTP transport)
  Tools/                       #   33 MCP tool definitions across 18 files
SaddleRAG.Cli/                    # Command-line interface (ingest, status, register-clients, ...)
SaddleRAG.Installer/              # WiX MSI installer definition
SaddleRAG.Tests/                  # Integration and unit tests

Support

SaddleRAG is free and MIT-licensed. If it's saving you time and you'd like to chip in, GitHub Sponsors is the easiest way to do that — any amount helps and there are no perks gated behind it. Totally optional; stars and bug reports are appreciated just as much.

License

SaddleRAG is released under the MIT License. You may use, copy, modify, merge, publish, distribute, sublicense, and sell the software, including for commercial purposes, subject only to the requirement that the copyright notice and the MIT permission notice be included in copies or substantial portions of the software.

Contributions are welcome — see CONTRIBUTING.md. No contributor agreement is required.

Languages

C#

97.0%

HTML

1.2%