ubunye-ai-ecosystems/ubunye_engine

Config-driven Spark framework for data and ML pipelines. Describe a pipeline once as config plus Python, then run that exact folder on a laptop, Docker, Kubernetes, cloud clusters or Databricks.

Python

17

327 commits

updated Sep 30, 2026

See the code

See what people are saying

README

Ubunye Engine

Ubunye (oo-BOON-yeh) — isiZulu for "unity"

One framework. Every pipeline. Any environment.

Docs • Docs map • Quickstart • Why Ubunye • Community


Hey there 👋

A data pipeline is a program that moves data from one place to another — a database to a file, a REST API to a data warehouse — and usually reshapes the data along the way. Building one from scratch is mostly plumbing: wire up the connection, juggle credentials, learn a framework's quirks, write the same "read → transform → write" scaffold for the tenth time this year. It's a lot of glue code standing between you and the three lines that actually matter.

Ubunye Engine writes that plumbing for you. You describe the pipeline in a short YAML file and put your transformation in a normal Python class. Ubunye takes care of connections, the compute engine (Apache Spark), and the read/write loop.

Same pipeline runs on your laptop today and on a production cluster tomorrow, with no code changes. That sentence is tested, not hoped: a build job runs one pipeline on six environments and fails if the outputs differ by a single byte.


What the engine gives you today

  • Proven portability. One task, unchanged, gives the same rows on pandas (no Java), local Spark, Kubernetes, AWS Glue, GCP Dataproc, Azure Container Apps and Databricks: the same data hash in all seven, from run records compared by ubunye prove (see Run Anywhere).
  • Models saved anywhere. The model registry writes to a local folder, a Databricks volume, S3 or GCS, chosen purely by the path. New storage kinds are one class and one entry point.
  • A truly open plugin system. Connectors, storage backends and registries are all added from the outside, with no engine edits and no inheritance required. A test proves it with a connector the engine has never seen.
  • Types that ship. The package carries its type information, and a type checker guards every merge.
  • Errors that help. Failures say what went wrong, show the context, and suggest the fix.
  • Nothing here is decorative. Eleven worked examples have been run for real. Claims that could not be executed were removed from these docs rather than left to mislead.

Quickstart

On a laptop, with no Java and no cloud account:

pip install "ubunye-engine[pandas]"
ubunye init -d pipelines -u demo -p starter -t filter_adults
ubunye plan -d pipelines -u demo -p starter -t filter_adults --backend pandas
ubunye run -d pipelines -u demo -p starter -t filter_adults --backend pandas --lineage
ubunye lineage list -d pipelines -u demo -p starter -t filter_adults

These four commands are run by the test suite exactly as written. init makes a folder, and the folder is the whole task:

pipelines/demo/starter/filter_adults/
  config.yaml            what to read and write
  transformations.py     your code
  data/people.csv        a small sample to start from

plan checks it without moving any data, run reads the CSV and writes Parquet, and lineage list shows the record the run left, including a hash of every row written. The code is one line:

class FilterAdults(Task):
    def transform(self, sources):
        people = sources["people"]
        return {"adults": people[people["age"] >= 18]}

That line means the same thing in pandas and in Spark. With Java installed (pip install "ubunye-engine[spark]"), leave out --backend pandas and the same folder runs on Spark, leaving the same data hash. On Databricks, call it from a notebook and the notebook's session is used:

import ubunye
outputs = ubunye.run_task("pipelines/demo/starter/filter_adults")

Step by step: the Quickstart. Realistic end to end examples live in ubunye-examples.


Docs

The full docs live at ubunye-ai-ecosystems.github.io/ubunye_engine. These are the pages most people need:

You want toRead
Install itInstallation
Run a first taskQuickstart
Write a config.yamlConfig overview, then Inputs and outputs
Check the data a task writesExpectations
Read what a run did (the run record)Lineage commands and the run record
Run on pandas or SparkExecution backends
Look up a commandCLI reference
Call it from Python or a notebookPython API
Add a language model stepLanguage model steps
Copy a working exampleExamples
Read from or write to a systemConnectors
Deploy itDeployment
Understand an errorErrors
See what changedChangelog

Why Ubunye

We've all been there. You join a new team, open the repo, and find five Spark projects — each structured differently, each with its own way of handling configs, credentials, and deployment. One uses a JSON file, another has everything hardcoded, a third has a 300-line bash script that "Dave wrote and it just works."

Ubunye says: let's agree on how pipelines look. One folder structure. One config format. One CLI. Whether you're building an ETL job, a feature pipeline, or an ML training run.

Without UbunyeWith Ubunye
Every project looks differentOne standard: use_case / pipeline / task
Spark setup scattered everywhereEngine handles it from YAML config
Credentials hardcoded or inconsistent{{ env.DB_PASSWORD }} everywhere
"Works on my machine"Same config runs local, YARN, K8s, Databricks
New teammate needs a week to onboardubunye init and they're running in minutes

How It Works

Three simple ideas:

Config over code. Your pipeline is a YAML file. Inputs, outputs, Spark settings, scheduling — all declared, not coded.

Plugins for everything. The format field in your config picks which connector to use. A connector is a small Python class that knows how to read from or write to one specific place (a database, a REST API, a cloud bucket). Built-ins include hive, jdbc, delta, s3, unity, and rest_api. Need a new data source? Write one and register it — Ubunye discovers plugins automatically.

Folders as architecture. Pipelines are organized as project / use_case / pipeline / task. The CLI uses this structure for scaffolding, execution, and discovery:

pipelines/
  fraud_detection/
    ingestion/
      claim_etl/
      policy_etl/
    feature_engineering/
      claim_features/
    risk_scoring/
      train_model/
      score_claims/

What Can You Build With It

ETL pipelines — move data between Hive, JDBC databases, Delta Lake, S3, REST APIs. Config-driven, scheduled, reproducible.

ML training and inference — define your model behind a simple contract, let the engine handle versioning, storage, and deployment.

RAG document pipelines — ingest documents, extract text, chunk, compute embeddings, load into a vector store. All from YAML.

Feature engineering — compute features once, write to a shared table, reuse across use cases.

Data drift detection — monitor feature distributions between runs, flag when things shift.

Check out the Examples and the RAG document pipeline guide in our docs.


Examples

All worked examples live in ubunye-examples. They install the engine from PyPI, so what you run there is what you get from pip install.

They are grouped by where they run. The same task folder is used in every environment. Only environment variables change.

Databricks (free workspace is enough)

ExampleWhat it shows
Tables and SQL pushdownread a table, push a join down as SQL, write with merge
REST API ingestionthe rest_api connector against a public weather API
Unstructured filesbinary file reading and text chunking
The ML lifecyclefeatures, train, quality gate, registry, promote, score
RAGembeddings and a chat model, retrieval, grounded answers
Fine-tune an open LLMa large model labels data, DistilBERT learns from it
Data qualitya contract with severities, bad rows quarantined
Model monitoringdrift, decay, and a rollback that is a decision, not a reflex

Local machine, Docker, and Kubernetes (no cloud account needed)

ExampleWhat it shows
Run anywhereone task, identical output hash on local Spark, Docker, Kubernetes, object storage and Databricks
JDBCa partitioned parallel read from a real public PostgreSQL database
RAG and fine-tuning on open modelsthe same pipelines with local open source models instead of hosted endpoints
The three-framework racethe same model in scikit-learn, PyTorch and TensorFlow behind identical configs

AWS, GCP and Azure

ubunye deploy glue, dataproc and container-apps run a task on AWS Glue, GCP Dataproc Serverless and Azure Container Apps; each has run the proving workload with the same result as a laptop (see Run Anywhere). EMR Serverless is built but not yet run.

Connectors

FormatReadWriteDescription
hive✓✓Apache Hive tables
jdbc✓✓PostgreSQL, MySQL, Teradata, and more
delta✓✓Delta Lake (standalone or Unity Catalog)
s3✓✓S3, HDFS, or local filesystem
unity✓✓Databricks Unity Catalog
binary✓Binary files (images, PDFs)
rest_api✓✓REST APIs with pagination and auth

Want to add one? See the plugin guide.


Run Anywhere

The same task folder, unchanged, runs in every environment below and writes the same rows. This is measured, not claimed: the proving ground (ubunye prove) compares each run's record (every input and output row hashed, the schema, the row counts, the task code) with a local Spark run, and an environment without evidence is reported NOT RUN.

Workload C01 (a portable join, filter, null group keys, money and timestamps cut to a day), 2026-09-28, generated report:

EnvironmentHow it was launchedResult
Your laptop, pandas (no Java)ubunye run ... --backend pandasPASS, digest bb08a7d7a9fd
Your laptop, Sparkubunye run ... --backend sparkPASS, bb08a7d7a9fd
Kubernetes (kind)ubunye deploy k8sPASS, bb08a7d7a9fd
AWS Glue 5.0ubunye deploy gluePASS, bb08a7d7a9fd
GCP Dataproc Serverless 2.2ubunye deploy dataprocPASS, bb08a7d7a9fd
Azure Container Appsubunye deploy container-appsPASS, bb08a7d7a9fd
Databricks serverlessubunye.run_task() in a notebook jobPASS, bb08a7d7a9fd

Try the first two rows yourself in ten minutes: Tutorial 1.

Built, not yet run: AWS EMR Serverless (ubunye deploy emr-serverless; the AWS free plan blocks EMR). Not tested: Hadoop/YARN. Neither is claimed until a run says so.

One rule makes this work: the task never chooses its own cluster. Which Spark master to use belongs to whoever launches the job, so leave spark.master out of your config. On a laptop the runner sets local[*]. On a cloud the platform sets it, and the engine refuses a config that tries to override it, because a silent single-node run on paid compute is worse than an error. Time is cut in UTC on every backend unless the task says otherwise (ADR 007).


Jinja Templating

Anywhere a string appears in your YAML, you can plug in a variable using {{ … }} syntax (this is called Jinja templating). That's how you keep secrets out of your config, change paths per environment, and inject the run date from the CLI:

# Environment variables
password: "{{ env.DB_PASSWORD }}"

# CLI variables (--var ds=2025-01-01)
path: "s3a://bucket/{{ ds }}/"

# Defaults
path: "s3a://bucket/{{ ds | default('2025-01-01') }}/"

CLI

ubunye init     -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # scaffold
ubunye validate -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # check config
ubunye plan     -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # preview plan
ubunye run      -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # execute
ubunye test run -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # test mode
ubunye lineage list -d ./pipelines -u <use_case> -p <pipeline> -t <task>  # run history
ubunye models list -u <use_case> -m <model> -s <store>                 # model versions

Python API

import ubunye

# Run from Databricks or any Python environment
outputs = ubunye.run_task(task_dir="./pipelines/...", mode="DEV", dt="2024-06-01")

# Multiple tasks
results = ubunye.run_pipeline(
    usecase_dir="./pipelines", usecase="fraud", package="etl",
    tasks=["claim_etl", "features"], mode="DEV",
)

What Ubunye Is Not

It's not an agent framework — use LangChain or CrewAI for that. It's not an orchestrator — use Airflow, Prefect, or Dagster. It's not a compute engine — it runs on Spark.

Ubunye is the standardization layer between your data sources and your applications. It makes the plumbing boring so you can focus on what matters.


Roadmap

  • Config-driven ETL pipelines
  • Multi-environment profiles
  • Jinja templating
  • Plugin-based connectors
  • CLI scaffolding and execution
  • Pydantic config validation
  • ML model contract
  • Model registry with versioning
  • Lineage tracking
  • Python API for Databricks
  • Databricks Asset Bundles deployment
  • Dev notebook scaffolding
  • Data drift detection
  • ubunye deploy CLI command

Get Involved

We'd love your help. Whether it's a new connector, a bug fix, a typo, or just telling us what you're building — all contributions matter.


License

MIT License


Built with 🇿🇦 by Ubunye AI Ecosystems

apache-spark
databricks
data-engineering
data-pipelines
etl
kubernetes
machine-learning
mlops
pipeline-framework
pyspark
python
reproducibility

Significant stargazers

Humbulani

9 followers · starred Sep 2026

ubunye-ai-ecosystems/ubunye_engine

Config-driven Spark framework for data and ML pipelines. Describe a pipeline once as config plus Python, then run that exact folder on a laptop, Docker, Kubernetes, cloud clusters or Databricks.

Python

17

327 commits

updated Sep 30, 2026

See the code

See what people are saying

README

Ubunye Engine

Ubunye (oo-BOON-yeh) — isiZulu for "unity"

One framework. Every pipeline. Any environment.

Docs • Docs map • Quickstart • Why Ubunye • Community


Hey there 👋

A data pipeline is a program that moves data from one place to another — a database to a file, a REST API to a data warehouse — and usually reshapes the data along the way. Building one from scratch is mostly plumbing: wire up the connection, juggle credentials, learn a framework's quirks, write the same "read → transform → write" scaffold for the tenth time this year. It's a lot of glue code standing between you and the three lines that actually matter.

Ubunye Engine writes that plumbing for you. You describe the pipeline in a short YAML file and put your transformation in a normal Python class. Ubunye takes care of connections, the compute engine (Apache Spark), and the read/write loop.

Same pipeline runs on your laptop today and on a production cluster tomorrow, with no code changes. That sentence is tested, not hoped: a build job runs one pipeline on six environments and fails if the outputs differ by a single byte.


What the engine gives you today

  • Proven portability. One task, unchanged, gives the same rows on pandas (no Java), local Spark, Kubernetes, AWS Glue, GCP Dataproc, Azure Container Apps and Databricks: the same data hash in all seven, from run records compared by ubunye prove (see Run Anywhere).
  • Models saved anywhere. The model registry writes to a local folder, a Databricks volume, S3 or GCS, chosen purely by the path. New storage kinds are one class and one entry point.
  • A truly open plugin system. Connectors, storage backends and registries are all added from the outside, with no engine edits and no inheritance required. A test proves it with a connector the engine has never seen.
  • Types that ship. The package carries its type information, and a type checker guards every merge.
  • Errors that help. Failures say what went wrong, show the context, and suggest the fix.
  • Nothing here is decorative. Eleven worked examples have been run for real. Claims that could not be executed were removed from these docs rather than left to mislead.

Quickstart

On a laptop, with no Java and no cloud account:

pip install "ubunye-engine[pandas]"
ubunye init -d pipelines -u demo -p starter -t filter_adults
ubunye plan -d pipelines -u demo -p starter -t filter_adults --backend pandas
ubunye run -d pipelines -u demo -p starter -t filter_adults --backend pandas --lineage
ubunye lineage list -d pipelines -u demo -p starter -t filter_adults

These four commands are run by the test suite exactly as written. init makes a folder, and the folder is the whole task:

pipelines/demo/starter/filter_adults/
  config.yaml            what to read and write
  transformations.py     your code
  data/people.csv        a small sample to start from

plan checks it without moving any data, run reads the CSV and writes Parquet, and lineage list shows the record the run left, including a hash of every row written. The code is one line:

class FilterAdults(Task):
    def transform(self, sources):
        people = sources["people"]
        return {"adults": people[people["age"] >= 18]}

That line means the same thing in pandas and in Spark. With Java installed (pip install "ubunye-engine[spark]"), leave out --backend pandas and the same folder runs on Spark, leaving the same data hash. On Databricks, call it from a notebook and the notebook's session is used:

import ubunye
outputs = ubunye.run_task("pipelines/demo/starter/filter_adults")

Step by step: the Quickstart. Realistic end to end examples live in ubunye-examples.


Docs

The full docs live at ubunye-ai-ecosystems.github.io/ubunye_engine. These are the pages most people need:

You want toRead
Install itInstallation
Run a first taskQuickstart
Write a config.yamlConfig overview, then Inputs and outputs
Check the data a task writesExpectations
Read what a run did (the run record)Lineage commands and the run record
Run on pandas or SparkExecution backends
Look up a commandCLI reference
Call it from Python or a notebookPython API
Add a language model stepLanguage model steps
Copy a working exampleExamples
Read from or write to a systemConnectors
Deploy itDeployment
Understand an errorErrors
See what changedChangelog

Why Ubunye

We've all been there. You join a new team, open the repo, and find five Spark projects — each structured differently, each with its own way of handling configs, credentials, and deployment. One uses a JSON file, another has everything hardcoded, a third has a 300-line bash script that "Dave wrote and it just works."

Ubunye says: let's agree on how pipelines look. One folder structure. One config format. One CLI. Whether you're building an ETL job, a feature pipeline, or an ML training run.

Without UbunyeWith Ubunye
Every project looks differentOne standard: use_case / pipeline / task
Spark setup scattered everywhereEngine handles it from YAML config
Credentials hardcoded or inconsistent{{ env.DB_PASSWORD }} everywhere
"Works on my machine"Same config runs local, YARN, K8s, Databricks
New teammate needs a week to onboardubunye init and they're running in minutes

How It Works

Three simple ideas:

Config over code. Your pipeline is a YAML file. Inputs, outputs, Spark settings, scheduling — all declared, not coded.

Plugins for everything. The format field in your config picks which connector to use. A connector is a small Python class that knows how to read from or write to one specific place (a database, a REST API, a cloud bucket). Built-ins include hive, jdbc, delta, s3, unity, and rest_api. Need a new data source? Write one and register it — Ubunye discovers plugins automatically.

Folders as architecture. Pipelines are organized as project / use_case / pipeline / task. The CLI uses this structure for scaffolding, execution, and discovery:

pipelines/
  fraud_detection/
    ingestion/
      claim_etl/
      policy_etl/
    feature_engineering/
      claim_features/
    risk_scoring/
      train_model/
      score_claims/

What Can You Build With It

ETL pipelines — move data between Hive, JDBC databases, Delta Lake, S3, REST APIs. Config-driven, scheduled, reproducible.

ML training and inference — define your model behind a simple contract, let the engine handle versioning, storage, and deployment.

RAG document pipelines — ingest documents, extract text, chunk, compute embeddings, load into a vector store. All from YAML.

Feature engineering — compute features once, write to a shared table, reuse across use cases.

Data drift detection — monitor feature distributions between runs, flag when things shift.

Check out the Examples and the RAG document pipeline guide in our docs.


Examples

All worked examples live in ubunye-examples. They install the engine from PyPI, so what you run there is what you get from pip install.

They are grouped by where they run. The same task folder is used in every environment. Only environment variables change.

Databricks (free workspace is enough)

ExampleWhat it shows
Tables and SQL pushdownread a table, push a join down as SQL, write with merge
REST API ingestionthe rest_api connector against a public weather API
Unstructured filesbinary file reading and text chunking
The ML lifecyclefeatures, train, quality gate, registry, promote, score
RAGembeddings and a chat model, retrieval, grounded answers
Fine-tune an open LLMa large model labels data, DistilBERT learns from it
Data qualitya contract with severities, bad rows quarantined
Model monitoringdrift, decay, and a rollback that is a decision, not a reflex

Local machine, Docker, and Kubernetes (no cloud account needed)

ExampleWhat it shows
Run anywhereone task, identical output hash on local Spark, Docker, Kubernetes, object storage and Databricks
JDBCa partitioned parallel read from a real public PostgreSQL database
RAG and fine-tuning on open modelsthe same pipelines with local open source models instead of hosted endpoints
The three-framework racethe same model in scikit-learn, PyTorch and TensorFlow behind identical configs

AWS, GCP and Azure

ubunye deploy glue, dataproc and container-apps run a task on AWS Glue, GCP Dataproc Serverless and Azure Container Apps; each has run the proving workload with the same result as a laptop (see Run Anywhere). EMR Serverless is built but not yet run.

Connectors

FormatReadWriteDescription
hive✓✓Apache Hive tables
jdbc✓✓PostgreSQL, MySQL, Teradata, and more
delta✓✓Delta Lake (standalone or Unity Catalog)
s3✓✓S3, HDFS, or local filesystem
unity✓✓Databricks Unity Catalog
binary✓Binary files (images, PDFs)
rest_api✓✓REST APIs with pagination and auth

Want to add one? See the plugin guide.


Run Anywhere

The same task folder, unchanged, runs in every environment below and writes the same rows. This is measured, not claimed: the proving ground (ubunye prove) compares each run's record (every input and output row hashed, the schema, the row counts, the task code) with a local Spark run, and an environment without evidence is reported NOT RUN.

Workload C01 (a portable join, filter, null group keys, money and timestamps cut to a day), 2026-09-28, generated report:

EnvironmentHow it was launchedResult
Your laptop, pandas (no Java)ubunye run ... --backend pandasPASS, digest bb08a7d7a9fd
Your laptop, Sparkubunye run ... --backend sparkPASS, bb08a7d7a9fd
Kubernetes (kind)ubunye deploy k8sPASS, bb08a7d7a9fd
AWS Glue 5.0ubunye deploy gluePASS, bb08a7d7a9fd
GCP Dataproc Serverless 2.2ubunye deploy dataprocPASS, bb08a7d7a9fd
Azure Container Appsubunye deploy container-appsPASS, bb08a7d7a9fd
Databricks serverlessubunye.run_task() in a notebook jobPASS, bb08a7d7a9fd

Try the first two rows yourself in ten minutes: Tutorial 1.

Built, not yet run: AWS EMR Serverless (ubunye deploy emr-serverless; the AWS free plan blocks EMR). Not tested: Hadoop/YARN. Neither is claimed until a run says so.

One rule makes this work: the task never chooses its own cluster. Which Spark master to use belongs to whoever launches the job, so leave spark.master out of your config. On a laptop the runner sets local[*]. On a cloud the platform sets it, and the engine refuses a config that tries to override it, because a silent single-node run on paid compute is worse than an error. Time is cut in UTC on every backend unless the task says otherwise (ADR 007).


Jinja Templating

Anywhere a string appears in your YAML, you can plug in a variable using {{ … }} syntax (this is called Jinja templating). That's how you keep secrets out of your config, change paths per environment, and inject the run date from the CLI:

# Environment variables
password: "{{ env.DB_PASSWORD }}"

# CLI variables (--var ds=2025-01-01)
path: "s3a://bucket/{{ ds }}/"

# Defaults
path: "s3a://bucket/{{ ds | default('2025-01-01') }}/"

CLI

ubunye init     -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # scaffold
ubunye validate -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # check config
ubunye plan     -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # preview plan
ubunye run      -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # execute
ubunye test run -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # test mode
ubunye lineage list -d ./pipelines -u <use_case> -p <pipeline> -t <task>  # run history
ubunye models list -u <use_case> -m <model> -s <store>                 # model versions

Python API

import ubunye

# Run from Databricks or any Python environment
outputs = ubunye.run_task(task_dir="./pipelines/...", mode="DEV", dt="2024-06-01")

# Multiple tasks
results = ubunye.run_pipeline(
    usecase_dir="./pipelines", usecase="fraud", package="etl",
    tasks=["claim_etl", "features"], mode="DEV",
)

What Ubunye Is Not

It's not an agent framework — use LangChain or CrewAI for that. It's not an orchestrator — use Airflow, Prefect, or Dagster. It's not a compute engine — it runs on Spark.

Ubunye is the standardization layer between your data sources and your applications. It makes the plumbing boring so you can focus on what matters.


Roadmap

  • Config-driven ETL pipelines
  • Multi-environment profiles
  • Jinja templating
  • Plugin-based connectors
  • CLI scaffolding and execution
  • Pydantic config validation
  • ML model contract
  • Model registry with versioning
  • Lineage tracking
  • Python API for Databricks
  • Databricks Asset Bundles deployment
  • Dev notebook scaffolding
  • Data drift detection
  • ubunye deploy CLI command

Get Involved

We'd love your help. Whether it's a new connector, a bug fix, a typo, or just telling us what you're building — all contributions matter.


License

MIT License


Built with 🇿🇦 by Ubunye AI Ecosystems

apache-spark
databricks
data-engineering
data-pipelines
etl
kubernetes
machine-learning
mlops
pipeline-framework
pyspark
python
reproducibility

Significant stargazers

Humbulani

9 followers · starred Sep 2026