Self-hosted, source-available AI data platform for batch pipelines, CDC, scheduled models, and lineage.
See the codeSelf-hosted, source-available AI data platform for batch pipelines, CDC, scheduled models, and lineage. Describe a pipeline in plain English, approve the plan, and see exactly what ran, failed, or became stale.

You describe the pipeline in plain English. It resolves the plan, then stops for your approval — batch, CDC, or changes-only — before a single row moves.
rsync.ai moves data between databases, warehouses, object stores and APIs. You describe the job in a sentence; an agent turns it into an explicit, staged plan, pauses for you when something is ambiguous, and executes it on Temporal so a long sync survives restarts. Batch and change-data-capture are both first-class. Twenty-one connectors ship in the box.
It is unrelated to rsync(1), the file-synchronisation
tool — this moves rows between systems, not files between hosts.
It is source-available under the Elastic License 2.0: run it, modify it, and use it internally for free — you just cannot resell it as a hosted service. The full summary is below.
More guides: all solutions.

Ask in plain English, review the SQL it wrote, run it against a connected source. Here: how many swipes went left versus right.

A finished run, stage by stage: what each one did, the rows it moved, and how stale the destination has become since. Nothing here was typed in by hand.

Every pipeline in a workspace on one screen: batch or CDC, source to destination, and whether it is running right now.

Lineage across pipelines and scheduled SQL models: which table a model writes, which model reads it, and which one runs after which.
Managed ELT tools move data well but hand off at the warehouse door. Orchestrators and automation tools are general-purpose and leave the data semantics to you. rsync.ai aims at the middle: get the data moving and keep it modelled, on hardware you control.
| Instead of | What it does well | What rsync.ai does differently |
|---|---|---|
| Fivetran | Managed and reliable, hundreds of connectors, someone else is on call | Runs on your infrastructure with your keys. A connector you need is a container you can write, not a support ticket. |
| Airbyte | Large connector ecosystem, self-hostable, mature ELT | You describe the pipeline in a sentence and approve a plan instead of configuring each sync by hand, and batch and CDC are the same product rather than separate paths. |
| dbt | The standard for SQL transformation, with deep testing and a large package ecosystem | Scheduled, dependency-aware SQL models are built in, so moving and modelling data is one tool instead of two. dbt's testing and packages are considerably deeper. |
| Debezium on its own | Best-in-class change data capture | rsync.ai runs Debezium and adds the provisioning, sinks, retries and UI around it, so you are not assembling Kafka Connect by hand. |
| Airflow / n8n | General orchestration and automation, enormously flexible | A pipeline is a first-class object with row counts, lineage and CDC built in, rather than something you assemble from operators or nodes. |
Where it is honestly weaker. There is no managed option — every install is yours to run. The catalogue is 21 connectors, not hundreds. Data-quality assertions are not built yet. And the Kubernetes path is younger than the Docker one (see Project status). If you want someone else carrying the pager, use a managed tool.
http://localhost:3000 and click Start with sample data. The stack bundles a
sample-data source and a throwaway demo-warehouse PostgreSQL, so this needs no
credential of your own./chat, ask for "sync customers and orders from sample data to the demo warehouse",
pick the tables, and confirm.That path is a batch pipeline. CDC, Shopify and your own databases need a source of your own — see the quickstart and the self-hosting guide.
flowchart LR
U["You, in plain English"] --> FE["Frontend<br/>Next.js"]
FE --> GW["API Gateway<br/>Go"]
GW --> ORCH["Orchestrator<br/>Go workers"]
ORCH --> TMP["Temporal<br/>durable workflows"]
TMP --> CON["MCP connectors<br/>versioned containers"]
CON --> DATA[("Your sources and<br/>destinations")]
For CDC, Debezium on Kafka Connect and a sink worker carry the change stream; they start with the rest of the default install. ARCHITECTURE.md explains why each piece was chosen, and docs/architecture/overview.md has the component and data-flow diagrams.
curl -sSL https://raw.githubusercontent.com/rsync-ai/rsync/main/install.sh | bash
Requires Docker and nothing else. The installer asks which LLM you want — your own
OpenAI key, the Ollama it bundles, or none for now — generates every other secret itself,
and starts the full stack. Choose Ollama and there is no key to find and no model to pull
by hand: the stack ships an Ollama container and a one-shot job that downloads the model
before anything that would ask for one starts. Choose none and pipelines, raw SQL and the
shipped connectors still work; the LLM features say Set up an LLM first until you add one
(which LLM is used). Open
http://localhost:3000 when it finishes. If
the stack does not come up, the installer says so and exits non-zero — it does not print a
success banner over a dead stack.
Which code you get.
v0.1.7, the current release. Both halves of the install come from that one tag: the compose file is fetched fromRSYNC_REFand the images are pulled at a tag derived from it, so the file and the containers it starts are the same commit. Every image the default compose starts is published at that tag and pullable anonymously — a test pins that, so a release cannot ship half-built.What it starts. Everything needed for both sync modes, change data capture included — Kafka Connect, Debezium and the sink worker come up with the rest. They are not an add-on: pick a streaming sync without them and the run fails a pre-flight two minutes in rather than falling back to batch. On a machine that will only ever run batch syncs,
curl -sSL … | RSYNC_PROFILES= bashleaves the JVM out and drops the memory floor back to 6 GB.Settings go on the
bashside of the pipe. AVAR=xwritten beforecurlsets it forcurl, which never reads it, and the installer runs with the default — no error, just the setting silently ignored. That is true of every variable here.Pass
RSYNC_REF=main(ascurl -sSL … | RSYNC_REF=main bash) to track the branch instead. That install is not reproducible: the compose file comes from the branch tip and changes with every commit, whilemainimages track the last publish rather than the newest commit, so the two halves move at different rates.A mirror, or your own build.
RSYNC_IMAGE_REGISTRY=registry.example.com/rsync-aipulls every first-party image from there instead ofghcr.io/rsync-ai— the same variableinstall-k8s.shreads — and is kept in.env, so a re-run keeps it. To install a checkout instead of a release (a fork, or a commit no tag carries yet), runRSYNC_COMPOSE_DIR=<checkout> RSYNC_VERSION=<tag> bash <checkout>/install.sh, where<tag>is the tag you pushed the checkout's images under. The compose files are copied from the checkout instead of downloaded. It refuses to run withoutRSYNC_VERSION, because the default would pair the checkout's compose file with the last release's images.
Point kubectl at any cluster and run:
curl -sSL https://raw.githubusercontent.com/rsync-ai/rsync/main/install-k8s.sh | bash
That is the whole install. It generates every secret, installs the platform, the demo
warehouse and a working set of connectors, waits for the release, and prints (or, on a
terminal, opens) the two port-forwards that put the UI at http://localhost:3000. Edit
~/rsync-ai-k8s/.env and run it again to change anything — that is also the upgrade path.
Back that file up: it holds ENCRYPTION_KEY. The default install asks for about 8.8 GiB
of memory and 3.7 CPU in requests, and it measures what the cluster has left before it
starts: on a smaller cluster it trims to a set that fits — first the connectors and demo
you did not choose, then the spare api-gateway and frontend replicas — instead of leaving
pods Pending. CPU is what runs out first: a single 4-vCPU node is not enough for the
default (GKE's own DaemonSets leave ~3.4 of it), and what fits there is the trimmed set
at ~3.0 CPU. Everything it accepts is listed in the
Kubernetes guide.
Prefer to run helm yourself? A bare helm install also needs connectors.fleet set, or no
connector pod starts and no pipeline can reach a source
(why):
git clone https://github.com/rsync-ai/rsync.git && cd rsync
helm install rsync ./deploy/helm/rsync-ai \
--namespace rsync --create-namespace \
--set secrets.jwtSecret="$(openssl rand -base64 32)" \
--set secrets.encryptionKey="$(openssl rand -base64 32)" \
--set secrets.internalServiceSecret="$(openssl rand -hex 24)" \
--set secrets.postgresPassword="$(openssl rand -hex 24)" \
--set secrets.minioAccessKey="$(openssl rand -hex 16)" \
--set secrets.minioSecretKey="$(openssl rand -base64 32)" \
--set frontend.publicUrl=https://app.example.com \
--set frontend.apiUrl=https://api.example.com \
-f my-values.yaml # at least connectors.fleet
That is the evaluation footprint — in-chart Postgres, Redis, Kafka, MinIO and Temporal, one replica each, no backups. The chart runs the same images as the compose stack and can point at managed Postgres, Redis, Kafka and object storage instead; per-provider value files ship for EKS, GKE and AKS. See the Kubernetes guide for a production install.
[!IMPORTANT] Save
secrets.encryptionKey. It encrypts every stored connection credential. Read it back withkubectl -n rsync get secret rsync-secrets -o jsonpath='{.data.ENCRYPTION_KEY}' | base64 -dand keep it somewhere you will still have it after the cluster is gone — reinstalling with a different key makes every saved connection permanently undecryptable.
[!TIP] The chart is also published to the registry, so you can install without cloning:
helm install rsync oci://ghcr.io/rsync-ai/charts/rsync-ai --version 0.1.7 \ --namespace rsync --create-namespace \ --set secrets.jwtSecret="$(openssl rand -base64 32)" \ --set secrets.encryptionKey="$(openssl rand -base64 32)" \ --set secrets.internalServiceSecret="$(openssl rand -hex 24)" \ --set secrets.postgresPassword="$(openssl rand -hex 24)" \ --set secrets.minioAccessKey="$(openssl rand -hex 16)" \ --set secrets.minioSecretKey="$(openssl rand -base64 32)" \ --set frontend.publicUrl=https://app.example.com \ --set frontend.apiUrl=https://api.example.comThe two
frontend.*flags are not optional on either path — the chart refuses to render without them, because the browser calls the API directly and NextAuth builds its callback URLs frompublicUrl. Point them at the hostnames your ingress will serve. MinIO withdrew anonymous pulls fromdocker.io/minio/*and then fromquay.io/minio/*. Chart 0.1.6 onward and a checkout'svalues.yamlname Chainguard's build instead, so neither path needs a MinIO override. Chart 0.1.5 and older still name the withdrawn images; to install one of those, add--set objectStorage.minio.image=cgr.dev/chainguard/minio@sha256:bd014394a80898e68c149f2311fdf8d5a2c2f3bb2c33b9327ae6d02b4b065ae1and the same value forobjectStorage.minio.mcImage. Both paths pull rsync's own images at.Chart.AppVersion(0.1.7), and everyghcr.io/rsync-aiimage the chart names is published at that tag for bothamd64andarm64(0.1.2 and older areamd64only, so they will not start on Apple Silicon, Graviton, Axion or Ampere nodes).
| Pipelines from a sentence | Type "sync MySQL orders to S3 every hour". An agent resolves it into named stages you can read before anything runs. |
| Batch and CDC, both first-class | Batch loads for anything, plus Debezium-backed change data capture on five databases — PostgreSQL, MySQL, SQL Server, Oracle and MongoDB. |
| It asks instead of guessing | When the source is ambiguous — which tables, which schema, which key — the run pauses on a human-in-the-loop gate rather than picking for you. |
| Durable execution | Stages run as Temporal workflows, so a multi-hour sync survives a restart, a redeploy, or a crashed worker. |
| You can answer "why did it do that?" | Every run emits domain events carrying stage state, row counts and a trace id, and the UI shows them stage by stage. |
| A SQL and NL query surface | The Data Explorer queries the systems you connected — no second BI tool to stand up first. |
| Your infrastructure, your keys | One Docker command or one Helm chart. Credentials are encrypted at rest with a key you hold; point the LLM at OpenAI or at the Ollama the installer bundles, or run without one. |
21 connectors ship in the box — every one is a source, 17 are also destinations, and five support change data capture. Each runs as its own versioned container, so you can upgrade or pin one without touching the rest.
| Category | Connectors | CDC |
|---|---|---|
| Relational | PostgreSQL, MySQL, SQL Server, Oracle, ClickHouse, Amazon Redshift | PostgreSQL, MySQL, SQL Server, Oracle |
| Data warehouse | Snowflake, Google BigQuery, Databricks | — |
| Document | MongoDB | MongoDB |
| Object storage | AWS S3, Google Cloud Storage, Azure Blob Storage | — |
| APIs | Stripe, Shopify, GitHub, Notion, Google Sheets | — |
| Demo and reference | Sample Data (credential-free demo source), Petstore (OpenAPI example), Widgets-GraphQL (GraphQL example) | — |
The connector reference is generated from the connector tree itself and lists exact ids, versions and per-connector source/destination support — CI fails if it drifts, and a second guard fails if the table above stops matching it. To add your own, start with the connector developer guide.
Once data has landed somewhere, you can query it without leaving rsync. Ask a question in English and get SQL back, or write the SQL yourself; browse the schema; then keep the useful ones — as a saved query with versions and diffs, or as a model: a table that rebuilds itself on a cron, an interval, or after a given pipeline finishes. Results export to CSV, TSV and JSON. See the Data Explorer guide and the deep dive on saved queries, models and schedules.
/chat. An
agent reads it and drafts a staged plan.Set up an LLM first until then
(which LLM is used)| Quick start | Local dev setup and first pipeline |
| Solutions | PostgreSQL CDC, PostgreSQL to MySQL, Shopify to PostgreSQL, scheduled SQL models, lineage |
| Self-hosting | Production deployment with TLS |
| Kubernetes | Helm chart install on EKS, GKE, AKS, or any cluster |
| Oracle Cloud (free) | Free 4 OCPU / 24 GB VM |
| Connector reference | Every shipped source and destination |
| Connector developer guide | Build a new connector |
| Data Explorer | SQL, natural-language queries, saved models and schedules |
| Architecture | System design and data flows |
| API reference | REST + WebSocket endpoints |
| Environment variables | Full configuration reference |
| Errors | What each error code means and what to do about it |
| All docs | Full documentation index |
git clone https://github.com/rsync-ai/rsync.git
cd rsync
cp .env.example .env # add your OPENAI_API_KEY, if you have one
cp llm-service/.env.example llm-service/.env # or set LLM_PROVIDER=none here
docker compose -p rsync-ai up -d
open http://localhost:3000
See CONTRIBUTING.md for building individual services, running the test suites, and the PR process.
rsync.ai is young and self-hosted. It runs, it has been driven end to end, and the connector and deployment claims on this page are checked by tests rather than asserted — but you are early. The rough edge today is Kubernetes: a managed-cluster install (EKS, GKE or AKS against real RDS, MSK and S3) has not been run end to end, so the cloud value files are reviewed starting points rather than verified recipes — the Kubernetes guide says so where you meet it. There is no hosted offering: every install is yours.
What that means in practice: pin a tag rather than tracking main if you want
reproducibility, keep ENCRYPTION_KEY somewhere durable before you store a credential,
and read CHANGELOG.md before upgrading. Bugs and gaps are tracked as
GitHub issues — that list is the register.
rsync.ai is source-available under the Elastic License 2.0 (ELv2) — not an OSI "open source" license.
The LICENSE file is the binding text; the following is a plain-English summary (not
legal advice):
You can:
You cannot:
The rsync.ai name and logo are trademarks — see TRADEMARK.md. Licenses of bundled third-party dependencies are listed in THIRD_PARTY_NOTICES.md.
Python
39.2%
Go
38.0%
TypeScript
18.8%
Shell
2.7%
Self-hosted, source-available AI data platform for batch pipelines, CDC, scheduled models, and lineage.
See the codeSelf-hosted, source-available AI data platform for batch pipelines, CDC, scheduled models, and lineage. Describe a pipeline in plain English, approve the plan, and see exactly what ran, failed, or became stale.

You describe the pipeline in plain English. It resolves the plan, then stops for your approval — batch, CDC, or changes-only — before a single row moves.
rsync.ai moves data between databases, warehouses, object stores and APIs. You describe the job in a sentence; an agent turns it into an explicit, staged plan, pauses for you when something is ambiguous, and executes it on Temporal so a long sync survives restarts. Batch and change-data-capture are both first-class. Twenty-one connectors ship in the box.
It is unrelated to rsync(1), the file-synchronisation
tool — this moves rows between systems, not files between hosts.
It is source-available under the Elastic License 2.0: run it, modify it, and use it internally for free — you just cannot resell it as a hosted service. The full summary is below.
More guides: all solutions.

Ask in plain English, review the SQL it wrote, run it against a connected source. Here: how many swipes went left versus right.

A finished run, stage by stage: what each one did, the rows it moved, and how stale the destination has become since. Nothing here was typed in by hand.

Every pipeline in a workspace on one screen: batch or CDC, source to destination, and whether it is running right now.

Lineage across pipelines and scheduled SQL models: which table a model writes, which model reads it, and which one runs after which.
Managed ELT tools move data well but hand off at the warehouse door. Orchestrators and automation tools are general-purpose and leave the data semantics to you. rsync.ai aims at the middle: get the data moving and keep it modelled, on hardware you control.
| Instead of | What it does well | What rsync.ai does differently |
|---|---|---|
| Fivetran | Managed and reliable, hundreds of connectors, someone else is on call | Runs on your infrastructure with your keys. A connector you need is a container you can write, not a support ticket. |
| Airbyte | Large connector ecosystem, self-hostable, mature ELT | You describe the pipeline in a sentence and approve a plan instead of configuring each sync by hand, and batch and CDC are the same product rather than separate paths. |
| dbt | The standard for SQL transformation, with deep testing and a large package ecosystem | Scheduled, dependency-aware SQL models are built in, so moving and modelling data is one tool instead of two. dbt's testing and packages are considerably deeper. |
| Debezium on its own | Best-in-class change data capture | rsync.ai runs Debezium and adds the provisioning, sinks, retries and UI around it, so you are not assembling Kafka Connect by hand. |
| Airflow / n8n | General orchestration and automation, enormously flexible | A pipeline is a first-class object with row counts, lineage and CDC built in, rather than something you assemble from operators or nodes. |
Where it is honestly weaker. There is no managed option — every install is yours to run. The catalogue is 21 connectors, not hundreds. Data-quality assertions are not built yet. And the Kubernetes path is younger than the Docker one (see Project status). If you want someone else carrying the pager, use a managed tool.
http://localhost:3000 and click Start with sample data. The stack bundles a
sample-data source and a throwaway demo-warehouse PostgreSQL, so this needs no
credential of your own./chat, ask for "sync customers and orders from sample data to the demo warehouse",
pick the tables, and confirm.That path is a batch pipeline. CDC, Shopify and your own databases need a source of your own — see the quickstart and the self-hosting guide.
flowchart LR
U["You, in plain English"] --> FE["Frontend<br/>Next.js"]
FE --> GW["API Gateway<br/>Go"]
GW --> ORCH["Orchestrator<br/>Go workers"]
ORCH --> TMP["Temporal<br/>durable workflows"]
TMP --> CON["MCP connectors<br/>versioned containers"]
CON --> DATA[("Your sources and<br/>destinations")]
For CDC, Debezium on Kafka Connect and a sink worker carry the change stream; they start with the rest of the default install. ARCHITECTURE.md explains why each piece was chosen, and docs/architecture/overview.md has the component and data-flow diagrams.
curl -sSL https://raw.githubusercontent.com/rsync-ai/rsync/main/install.sh | bash
Requires Docker and nothing else. The installer asks which LLM you want — your own
OpenAI key, the Ollama it bundles, or none for now — generates every other secret itself,
and starts the full stack. Choose Ollama and there is no key to find and no model to pull
by hand: the stack ships an Ollama container and a one-shot job that downloads the model
before anything that would ask for one starts. Choose none and pipelines, raw SQL and the
shipped connectors still work; the LLM features say Set up an LLM first until you add one
(which LLM is used). Open
http://localhost:3000 when it finishes. If
the stack does not come up, the installer says so and exits non-zero — it does not print a
success banner over a dead stack.
Which code you get.
v0.1.7, the current release. Both halves of the install come from that one tag: the compose file is fetched fromRSYNC_REFand the images are pulled at a tag derived from it, so the file and the containers it starts are the same commit. Every image the default compose starts is published at that tag and pullable anonymously — a test pins that, so a release cannot ship half-built.What it starts. Everything needed for both sync modes, change data capture included — Kafka Connect, Debezium and the sink worker come up with the rest. They are not an add-on: pick a streaming sync without them and the run fails a pre-flight two minutes in rather than falling back to batch. On a machine that will only ever run batch syncs,
curl -sSL … | RSYNC_PROFILES= bashleaves the JVM out and drops the memory floor back to 6 GB.Settings go on the
bashside of the pipe. AVAR=xwritten beforecurlsets it forcurl, which never reads it, and the installer runs with the default — no error, just the setting silently ignored. That is true of every variable here.Pass
RSYNC_REF=main(ascurl -sSL … | RSYNC_REF=main bash) to track the branch instead. That install is not reproducible: the compose file comes from the branch tip and changes with every commit, whilemainimages track the last publish rather than the newest commit, so the two halves move at different rates.A mirror, or your own build.
RSYNC_IMAGE_REGISTRY=registry.example.com/rsync-aipulls every first-party image from there instead ofghcr.io/rsync-ai— the same variableinstall-k8s.shreads — and is kept in.env, so a re-run keeps it. To install a checkout instead of a release (a fork, or a commit no tag carries yet), runRSYNC_COMPOSE_DIR=<checkout> RSYNC_VERSION=<tag> bash <checkout>/install.sh, where<tag>is the tag you pushed the checkout's images under. The compose files are copied from the checkout instead of downloaded. It refuses to run withoutRSYNC_VERSION, because the default would pair the checkout's compose file with the last release's images.
Point kubectl at any cluster and run:
curl -sSL https://raw.githubusercontent.com/rsync-ai/rsync/main/install-k8s.sh | bash
That is the whole install. It generates every secret, installs the platform, the demo
warehouse and a working set of connectors, waits for the release, and prints (or, on a
terminal, opens) the two port-forwards that put the UI at http://localhost:3000. Edit
~/rsync-ai-k8s/.env and run it again to change anything — that is also the upgrade path.
Back that file up: it holds ENCRYPTION_KEY. The default install asks for about 8.8 GiB
of memory and 3.7 CPU in requests, and it measures what the cluster has left before it
starts: on a smaller cluster it trims to a set that fits — first the connectors and demo
you did not choose, then the spare api-gateway and frontend replicas — instead of leaving
pods Pending. CPU is what runs out first: a single 4-vCPU node is not enough for the
default (GKE's own DaemonSets leave ~3.4 of it), and what fits there is the trimmed set
at ~3.0 CPU. Everything it accepts is listed in the
Kubernetes guide.
Prefer to run helm yourself? A bare helm install also needs connectors.fleet set, or no
connector pod starts and no pipeline can reach a source
(why):
git clone https://github.com/rsync-ai/rsync.git && cd rsync
helm install rsync ./deploy/helm/rsync-ai \
--namespace rsync --create-namespace \
--set secrets.jwtSecret="$(openssl rand -base64 32)" \
--set secrets.encryptionKey="$(openssl rand -base64 32)" \
--set secrets.internalServiceSecret="$(openssl rand -hex 24)" \
--set secrets.postgresPassword="$(openssl rand -hex 24)" \
--set secrets.minioAccessKey="$(openssl rand -hex 16)" \
--set secrets.minioSecretKey="$(openssl rand -base64 32)" \
--set frontend.publicUrl=https://app.example.com \
--set frontend.apiUrl=https://api.example.com \
-f my-values.yaml # at least connectors.fleet
That is the evaluation footprint — in-chart Postgres, Redis, Kafka, MinIO and Temporal, one replica each, no backups. The chart runs the same images as the compose stack and can point at managed Postgres, Redis, Kafka and object storage instead; per-provider value files ship for EKS, GKE and AKS. See the Kubernetes guide for a production install.
[!IMPORTANT] Save
secrets.encryptionKey. It encrypts every stored connection credential. Read it back withkubectl -n rsync get secret rsync-secrets -o jsonpath='{.data.ENCRYPTION_KEY}' | base64 -dand keep it somewhere you will still have it after the cluster is gone — reinstalling with a different key makes every saved connection permanently undecryptable.
[!TIP] The chart is also published to the registry, so you can install without cloning:
helm install rsync oci://ghcr.io/rsync-ai/charts/rsync-ai --version 0.1.7 \ --namespace rsync --create-namespace \ --set secrets.jwtSecret="$(openssl rand -base64 32)" \ --set secrets.encryptionKey="$(openssl rand -base64 32)" \ --set secrets.internalServiceSecret="$(openssl rand -hex 24)" \ --set secrets.postgresPassword="$(openssl rand -hex 24)" \ --set secrets.minioAccessKey="$(openssl rand -hex 16)" \ --set secrets.minioSecretKey="$(openssl rand -base64 32)" \ --set frontend.publicUrl=https://app.example.com \ --set frontend.apiUrl=https://api.example.comThe two
frontend.*flags are not optional on either path — the chart refuses to render without them, because the browser calls the API directly and NextAuth builds its callback URLs frompublicUrl. Point them at the hostnames your ingress will serve. MinIO withdrew anonymous pulls fromdocker.io/minio/*and then fromquay.io/minio/*. Chart 0.1.6 onward and a checkout'svalues.yamlname Chainguard's build instead, so neither path needs a MinIO override. Chart 0.1.5 and older still name the withdrawn images; to install one of those, add--set objectStorage.minio.image=cgr.dev/chainguard/minio@sha256:bd014394a80898e68c149f2311fdf8d5a2c2f3bb2c33b9327ae6d02b4b065ae1and the same value forobjectStorage.minio.mcImage. Both paths pull rsync's own images at.Chart.AppVersion(0.1.7), and everyghcr.io/rsync-aiimage the chart names is published at that tag for bothamd64andarm64(0.1.2 and older areamd64only, so they will not start on Apple Silicon, Graviton, Axion or Ampere nodes).
| Pipelines from a sentence | Type "sync MySQL orders to S3 every hour". An agent resolves it into named stages you can read before anything runs. |
| Batch and CDC, both first-class | Batch loads for anything, plus Debezium-backed change data capture on five databases — PostgreSQL, MySQL, SQL Server, Oracle and MongoDB. |
| It asks instead of guessing | When the source is ambiguous — which tables, which schema, which key — the run pauses on a human-in-the-loop gate rather than picking for you. |
| Durable execution | Stages run as Temporal workflows, so a multi-hour sync survives a restart, a redeploy, or a crashed worker. |
| You can answer "why did it do that?" | Every run emits domain events carrying stage state, row counts and a trace id, and the UI shows them stage by stage. |
| A SQL and NL query surface | The Data Explorer queries the systems you connected — no second BI tool to stand up first. |
| Your infrastructure, your keys | One Docker command or one Helm chart. Credentials are encrypted at rest with a key you hold; point the LLM at OpenAI or at the Ollama the installer bundles, or run without one. |
21 connectors ship in the box — every one is a source, 17 are also destinations, and five support change data capture. Each runs as its own versioned container, so you can upgrade or pin one without touching the rest.
| Category | Connectors | CDC |
|---|---|---|
| Relational | PostgreSQL, MySQL, SQL Server, Oracle, ClickHouse, Amazon Redshift | PostgreSQL, MySQL, SQL Server, Oracle |
| Data warehouse | Snowflake, Google BigQuery, Databricks | — |
| Document | MongoDB | MongoDB |
| Object storage | AWS S3, Google Cloud Storage, Azure Blob Storage | — |
| APIs | Stripe, Shopify, GitHub, Notion, Google Sheets | — |
| Demo and reference | Sample Data (credential-free demo source), Petstore (OpenAPI example), Widgets-GraphQL (GraphQL example) | — |
The connector reference is generated from the connector tree itself and lists exact ids, versions and per-connector source/destination support — CI fails if it drifts, and a second guard fails if the table above stops matching it. To add your own, start with the connector developer guide.
Once data has landed somewhere, you can query it without leaving rsync. Ask a question in English and get SQL back, or write the SQL yourself; browse the schema; then keep the useful ones — as a saved query with versions and diffs, or as a model: a table that rebuilds itself on a cron, an interval, or after a given pipeline finishes. Results export to CSV, TSV and JSON. See the Data Explorer guide and the deep dive on saved queries, models and schedules.
/chat. An
agent reads it and drafts a staged plan.Set up an LLM first until then
(which LLM is used)| Quick start | Local dev setup and first pipeline |
| Solutions | PostgreSQL CDC, PostgreSQL to MySQL, Shopify to PostgreSQL, scheduled SQL models, lineage |
| Self-hosting | Production deployment with TLS |
| Kubernetes | Helm chart install on EKS, GKE, AKS, or any cluster |
| Oracle Cloud (free) | Free 4 OCPU / 24 GB VM |
| Connector reference | Every shipped source and destination |
| Connector developer guide | Build a new connector |
| Data Explorer | SQL, natural-language queries, saved models and schedules |
| Architecture | System design and data flows |
| API reference | REST + WebSocket endpoints |
| Environment variables | Full configuration reference |
| Errors | What each error code means and what to do about it |
| All docs | Full documentation index |
git clone https://github.com/rsync-ai/rsync.git
cd rsync
cp .env.example .env # add your OPENAI_API_KEY, if you have one
cp llm-service/.env.example llm-service/.env # or set LLM_PROVIDER=none here
docker compose -p rsync-ai up -d
open http://localhost:3000
See CONTRIBUTING.md for building individual services, running the test suites, and the PR process.
rsync.ai is young and self-hosted. It runs, it has been driven end to end, and the connector and deployment claims on this page are checked by tests rather than asserted — but you are early. The rough edge today is Kubernetes: a managed-cluster install (EKS, GKE or AKS against real RDS, MSK and S3) has not been run end to end, so the cloud value files are reviewed starting points rather than verified recipes — the Kubernetes guide says so where you meet it. There is no hosted offering: every install is yours.
What that means in practice: pin a tag rather than tracking main if you want
reproducibility, keep ENCRYPTION_KEY somewhere durable before you store a credential,
and read CHANGELOG.md before upgrading. Bugs and gaps are tracked as
GitHub issues — that list is the register.
rsync.ai is source-available under the Elastic License 2.0 (ELv2) — not an OSI "open source" license.
The LICENSE file is the binding text; the following is a plain-English summary (not
legal advice):
You can:
You cannot:
The rsync.ai name and logo are trademarks — see TRADEMARK.md. Licenses of bundled third-party dependencies are listed in THIRD_PARTY_NOTICES.md.
Python
39.2%
Go
38.0%
TypeScript
18.8%
Shell
2.7%