laban254/insight-orchestra

Self-hostable AI data analyst: a multi-agent LLM pipeline with natural-language querying and sandboxed Python execution.

Python

9

174 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

Built a self-hostable AI data analyst — upload a file or connect a DB, 4 agents analyze it, works with Ollama, OpenAI, Anthropic, or DeepSeek (r/ollama)

Every time I needed to explore a dataset quickly, the same wall appeared — cleaning boilerplate, encoding errors, deciding which chart actually makes sense. The mechanical parts aren't interesting, they're just in the way. So I built something to handle them. You connect a file (CSV, Excel,…

1

Oct 1, 2026

README

Insight Orchestra

Your data, analyzed by a team of AI agents.

Connect a data file or a database and watch specialized agents clean it, form hypotheses,
debate them, and visualize what matters — then ask follow-ups in plain English.

Website · Docs · Report a bug

CI CodeQL Latest release License Stars Open in GitHub Codespaces

Insight Orchestra — four agents cleaning, hypothesising, debating and visualising a dataset

A real run on the bundled Sales dataset — unedited.


What is Insight Orchestra?

Insight Orchestra is an open-source AI data analyst you can self-host — think Julius AI or ChatGPT's data analysis, but running on your own hardware, with your choice of LLM, where your data never leaves your machine. Upload a data file — CSV, TSV, Excel, JSON, or Parquet — or connect a PostgreSQL, MySQL, SQLite, or DuckDB database, and a 4-agent pipeline cleans the data, generates evidence-backed hypotheses, scores them in an LLM-refereed debate, and builds interactive Plotly charts. Then keep asking questions in plain English: an NLQ agent writes pandas code and executes it in a locked-down sandbox — or, for a connected database, writes and runs read-only SQL directly, joining across tables as needed.

It works with your choice of LLM — OpenAI, Anthropic, or DeepSeek in the cloud, or fully local and private with Ollama.

Quick Start

Prerequisites: Docker & Docker Compose v2 · Git · 4 GB RAM (8 GB recommended for local LLMs)

One line, clones into ./insight-orchestra and runs the setup wizard:

curl -fsSL https://raw.githubusercontent.com/laban254/insight-orchestra/main/install.sh | bash

Or clone it yourself first:

git clone https://github.com/laban254/insight-orchestra.git
cd insight-orchestra
./setup.sh

The script asks which LLM provider to use (Ollama by default — local, private, no API key needed), writes backend/.env, starts the containers, and pulls the Ollama model automatically. Run it again any time; it won't clobber an existing backend/.env without asking.

Images are pulled prebuilt from GitHub Container Registry, so there's no local build to sit through. To pin a specific release instead of tracking latest:

IO_IMAGE_TAG=v1.0.0 ./setup.sh

To build from source instead — for development, or on a platform we don't publish images for — use ./setup.sh --build. See Contributing for the development workflow.

Fully non-interactive (works with either path above — pass the flags after bash -s -- for the curl one-liner):

./setup.sh --provider ollama -y                        # local, no API key
./setup.sh --provider openai --api-key sk-... -y       # or anthropic / deepseek

# equivalent, without cloning first:
curl -fsSL https://raw.githubusercontent.com/laban254/insight-orchestra/main/install.sh | bash -s -- --provider ollama -y

Something not working? ./setup.sh doctor checks Docker, ports, config, and running services. Prefer to configure backend/.env by hand instead? See the Setup Guide.

Open the app

Once ./setup.sh finishes:

Pick one of the five bundled demo datasets, upload your own file (CSV, TSV, Excel, JSON, or Parquet), or connect a PostgreSQL/MySQL/SQLite/DuckDB database — the pipeline runs automatically either way.

How It Works

The pipeline runs the moment you upload or select a dataset. Results — narrative, ranked insights, charts, suggested follow-ups — appear in the chat as the first message.

StageAgentFunction
1Data JanitorRemoves duplicates; imputes missing values (median for numeric, mode for categorical); flags bias (>30% missing); detects outliers via IQR
2Hypothesis BotBuilds descriptive statistics + correlations, then asks the LLM to generate 5–8 specific, directional, evidence-backed insights referencing actual column names and numbers
3Debate ManagerLLM scores each hypothesis on confidence and business_value (0–1) using the real data stats as evidence; sorts by combined score; selects consensus winner
4Viz WhizAsks the LLM which columns best illustrate the top insight; falls back to regex extraction then structured heuristics; generates up to 6 Plotly charts
5Insight SummarizerLLM writes a 3–5 sentence narrative summarising all findings; generates 4–5 specific follow-up questions using actual column names

Each stage streams real-time progress to the UI via SSE. See the Agent Pipeline Guide for the full breakdown.

Features

  • Natural Language Queries — the NLQ agent generates pandas code, executes it in the RestrictedPython sandbox, and returns results + optional Plotly charts
  • Four LLM Providers — OpenAI, Anthropic, DeepSeek, or Ollama (any locally-hosted model); switch provider/model at runtime, no restart needed
  • Multiple File Formats — upload CSV, TSV, Excel (.xlsx), JSON, or Parquet; each is sniffed for encoding, delimiter, and date columns on the way in
  • Multi-Database Support — PostgreSQL, MySQL, SQLite, and DuckDB, all read-only, connected through the UI (BigQuery has an experimental endpoint; see the API Reference). Includes a JOIN-capable natural-language SQL agent that answers questions directly against a connected database, across every table in scope, without materializing a table first
  • Sandboxed Code Execution — no file I/O, no network access, no dangerous imports; configurable timeout
  • Real-Time Agent Progress — SSE streaming shows each agent's status, output, and duration
  • Workspace, Share & Export — pin and compare charts, workspace history saved server-side (reopen past runs from any browser), one-click read-only share links (72 h TTL), export as an interactive HTML report, PDF, Markdown summary, or Q&A CSV
  • Optional Auth & Access Control — off by default for local/single-user use; turn on AUTH_ENABLED for login, role-based access (admin/member/viewer), OIDC SSO, self-service API keys, and an audit log — see API Reference
  • 5 Demo Datasets — try it without bringing your own data

Documentation

DocumentPurpose
Setup GuideDocker and local development setup, troubleshooting
ArchitectureSystem design, component breakdown, data flow
Agent PipelineDeep dive into all 4 agents + NLQ agent
API ReferenceAll REST endpoints with request/response examples

Roadmap

Have a request or want to influence priorities? Open an issue.

Contributing

Contributions are welcome — the Contributing Guide covers the development workflow, code style, and how to add a new agent to the pipeline.

If Insight Orchestra is useful to you, consider starring the repo — it helps others find the project.

License

Apache 2.0 — see LICENSE.

Author: @laban254

ai
ai-agents
data-analysis
data-science
data-visualization
duckdb
fastapi
llm
multi-agent
natural-language-query
nextjs
ollama
openai
pandas
plotly
python
self-hosted

laban254/insight-orchestra

Self-hostable AI data analyst: a multi-agent LLM pipeline with natural-language querying and sandboxed Python execution.

Python

9

174 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

Built a self-hostable AI data analyst — upload a file or connect a DB, 4 agents analyze it, works with Ollama, OpenAI, Anthropic, or DeepSeek (r/ollama)

Every time I needed to explore a dataset quickly, the same wall appeared — cleaning boilerplate, encoding errors, deciding which chart actually makes sense. The mechanical parts aren't interesting, they're just in the way. So I built something to handle them. You connect a file (CSV, Excel,…

1

Oct 1, 2026

README

Insight Orchestra

Your data, analyzed by a team of AI agents.

Connect a data file or a database and watch specialized agents clean it, form hypotheses,
debate them, and visualize what matters — then ask follow-ups in plain English.

Website · Docs · Report a bug

CI CodeQL Latest release License Stars Open in GitHub Codespaces

Insight Orchestra — four agents cleaning, hypothesising, debating and visualising a dataset

A real run on the bundled Sales dataset — unedited.


What is Insight Orchestra?

Insight Orchestra is an open-source AI data analyst you can self-host — think Julius AI or ChatGPT's data analysis, but running on your own hardware, with your choice of LLM, where your data never leaves your machine. Upload a data file — CSV, TSV, Excel, JSON, or Parquet — or connect a PostgreSQL, MySQL, SQLite, or DuckDB database, and a 4-agent pipeline cleans the data, generates evidence-backed hypotheses, scores them in an LLM-refereed debate, and builds interactive Plotly charts. Then keep asking questions in plain English: an NLQ agent writes pandas code and executes it in a locked-down sandbox — or, for a connected database, writes and runs read-only SQL directly, joining across tables as needed.

It works with your choice of LLM — OpenAI, Anthropic, or DeepSeek in the cloud, or fully local and private with Ollama.

Quick Start

Prerequisites: Docker & Docker Compose v2 · Git · 4 GB RAM (8 GB recommended for local LLMs)

One line, clones into ./insight-orchestra and runs the setup wizard:

curl -fsSL https://raw.githubusercontent.com/laban254/insight-orchestra/main/install.sh | bash

Or clone it yourself first:

git clone https://github.com/laban254/insight-orchestra.git
cd insight-orchestra
./setup.sh

The script asks which LLM provider to use (Ollama by default — local, private, no API key needed), writes backend/.env, starts the containers, and pulls the Ollama model automatically. Run it again any time; it won't clobber an existing backend/.env without asking.

Images are pulled prebuilt from GitHub Container Registry, so there's no local build to sit through. To pin a specific release instead of tracking latest:

IO_IMAGE_TAG=v1.0.0 ./setup.sh

To build from source instead — for development, or on a platform we don't publish images for — use ./setup.sh --build. See Contributing for the development workflow.

Fully non-interactive (works with either path above — pass the flags after bash -s -- for the curl one-liner):

./setup.sh --provider ollama -y                        # local, no API key
./setup.sh --provider openai --api-key sk-... -y       # or anthropic / deepseek

# equivalent, without cloning first:
curl -fsSL https://raw.githubusercontent.com/laban254/insight-orchestra/main/install.sh | bash -s -- --provider ollama -y

Something not working? ./setup.sh doctor checks Docker, ports, config, and running services. Prefer to configure backend/.env by hand instead? See the Setup Guide.

Open the app

Once ./setup.sh finishes:

Pick one of the five bundled demo datasets, upload your own file (CSV, TSV, Excel, JSON, or Parquet), or connect a PostgreSQL/MySQL/SQLite/DuckDB database — the pipeline runs automatically either way.

How It Works

The pipeline runs the moment you upload or select a dataset. Results — narrative, ranked insights, charts, suggested follow-ups — appear in the chat as the first message.

StageAgentFunction
1Data JanitorRemoves duplicates; imputes missing values (median for numeric, mode for categorical); flags bias (>30% missing); detects outliers via IQR
2Hypothesis BotBuilds descriptive statistics + correlations, then asks the LLM to generate 5–8 specific, directional, evidence-backed insights referencing actual column names and numbers
3Debate ManagerLLM scores each hypothesis on confidence and business_value (0–1) using the real data stats as evidence; sorts by combined score; selects consensus winner
4Viz WhizAsks the LLM which columns best illustrate the top insight; falls back to regex extraction then structured heuristics; generates up to 6 Plotly charts
5Insight SummarizerLLM writes a 3–5 sentence narrative summarising all findings; generates 4–5 specific follow-up questions using actual column names

Each stage streams real-time progress to the UI via SSE. See the Agent Pipeline Guide for the full breakdown.

Features

  • Natural Language Queries — the NLQ agent generates pandas code, executes it in the RestrictedPython sandbox, and returns results + optional Plotly charts
  • Four LLM Providers — OpenAI, Anthropic, DeepSeek, or Ollama (any locally-hosted model); switch provider/model at runtime, no restart needed
  • Multiple File Formats — upload CSV, TSV, Excel (.xlsx), JSON, or Parquet; each is sniffed for encoding, delimiter, and date columns on the way in
  • Multi-Database Support — PostgreSQL, MySQL, SQLite, and DuckDB, all read-only, connected through the UI (BigQuery has an experimental endpoint; see the API Reference). Includes a JOIN-capable natural-language SQL agent that answers questions directly against a connected database, across every table in scope, without materializing a table first
  • Sandboxed Code Execution — no file I/O, no network access, no dangerous imports; configurable timeout
  • Real-Time Agent Progress — SSE streaming shows each agent's status, output, and duration
  • Workspace, Share & Export — pin and compare charts, workspace history saved server-side (reopen past runs from any browser), one-click read-only share links (72 h TTL), export as an interactive HTML report, PDF, Markdown summary, or Q&A CSV
  • Optional Auth & Access Control — off by default for local/single-user use; turn on AUTH_ENABLED for login, role-based access (admin/member/viewer), OIDC SSO, self-service API keys, and an audit log — see API Reference
  • 5 Demo Datasets — try it without bringing your own data

Documentation

DocumentPurpose
Setup GuideDocker and local development setup, troubleshooting
ArchitectureSystem design, component breakdown, data flow
Agent PipelineDeep dive into all 4 agents + NLQ agent
API ReferenceAll REST endpoints with request/response examples

Roadmap

Have a request or want to influence priorities? Open an issue.

Contributing

Contributions are welcome — the Contributing Guide covers the development workflow, code style, and how to add a new agent to the pipeline.

If Insight Orchestra is useful to you, consider starring the repo — it helps others find the project.

License

Apache 2.0 — see LICENSE.

Author: @laban254

ai
ai-agents
data-analysis
data-science
data-visualization
duckdb
fastapi
llm
multi-agent
natural-language-query
nextjs
ollama
openai
pandas
plotly
python
self-hosted

Languages

Python

67.7%

TypeScript

29.4%

Shell

1.7%