Desbordante/desbordante-core

Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.

C++

510

2,094 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

Show HN: Desbordante 2.5.0 Released

1

Sep 30, 2026

README

Downloads Downloads License: AGPL v3

Desbordante logo

About

Desbordante is a high-performance data profiler that is capable of discovering and validating many different patterns in data using various algorithms.

Tool Positioning

Why Desbordante? Why not just use an LLM or other black-box machine learning approaches?

Unlike these, Desbordante takes a symbolic, explainable approach to data analysis, producing explicit rules or rule-shaped artifacts that are verifiable, such as data dependencies, association rules, or other patterns. This is a distinct paradigm for data analysis that complements black-box learning. In contrast to free-form model outputs, discovered patterns have precisely defined semantics and can be algorithmically validated against the underlying data.

These patterns add data-grounded and rule-based structure, which provides useful properties for systems built around machine learning and large language models: consistency, provenance, explainability, and auditability. Further, the patterns can serve as inputs to downstream data-quality checks, validation procedures, integrity constraints, training-data analysis, and other rule-based safeguards. They can also be used to choose a direction of an ablation study, to generate examples that violate rules for use in machine learning model training and testing, to assist fact completion, and to facilitate acceptance checks and guardrails for pipelines involving retrieval-augmented generation. This is roughly the motivation used for the work in [1], but it applies more broadly to the patterns Desbordante works with.

[1] François Amat. Mining Rules on Tabular Data. Institut Polytechnique de Paris, 2025.

From other angles, the discovered patterns can be used in many ways:

  • For scientific data, especially those obtained experimentally, an interesting pattern allows to formulate a hypothesis that could lead to a scientific discovery. In some cases it even allows one to draw conclusions immediately, if there is enough data. At the very least, the found pattern can be used to provide a direction for further study.

  • For business data, it is also possible to obtain a hypothesis based on found patterns. However, there are more down-to-earth and more in-demand applications in this case: clearing errors in data, finding and removing inexact duplicates, performing schema matching, and many more.

  • For database data, found patterns can help with recovering/defining primary and foreign keys, checking/setting up all kinds of integrity constraints.

General

Desbordante supports three types of tasks.

The Discovery task is designed to identify all instances of a specified pattern type of a given dataset.

The Validation task is different: it is designed to check whether a specified pattern instance is present in a given dataset. This task not only returns True or False, but it also explains why the instance does not hold (e.g. it can list table rows with conflicting values).

For some patterns Desbordante supports a dynamic task variant. The distinguishing feature of dynamic algorithms compared to classic (static) algorithms is that after a result is obtained, the table can be changed and a dynamic algorithm will update the result based just on those changes instead of processing the whole table again. As a result, they can be up to several orders of magnitude faster than classic (static) ones in some situations.

Desbordante support four data types:

  • Tabular
  • Graph
  • Transactional
  • Event sequences

Tabular data patterns

  • Exact functional dependencies (discovery and validation)
  • Approximate functional dependencies, with
  • Probabilistic functional dependencies, with PerTuple and PerValue metrics (discovery and validation)
  • Classic soft functional dependencies (with correlations), with $\rho$ metric (discovery and validation)
  • Dynamic validation of exact and approximate ($g_1$) functional dependencies
  • Relaxed functional dependencies (discovery)
  • Numerical dependencies (validation)
  • Conditional functional dependencies (discovery and validation)
  • Inclusion dependencies
  • Conditional inclusion dependencies (discovery and validation)
  • Order dependencies:
    • set-based axiomatization (discovery and validation)
    • list-based axiomatization (discovery)
  • Approximate order dependencies:
    • set-based axiomatization (discovery and validation)
  • Sequential dependencies (validation)
  • Metric functional dependencies (validation)
  • Fuzzy algebraic constraints (discovery)
  • Domain Probabilistic and Approximate Constraints (validation)
  • Differential Dependencies (discovery and validation)
  • Unique column combinations:
    • Exact unique column combination (discovery and validation)
    • Approximate unique column combination, with $g_1$ metric (discovery and validation)
  • Numerical association rules (discovery)
  • Matching dependencies (discovery and validation)
  • Denial constraints

Graph data patterns

  • Graph functional dependencies (discovery and validation)
  • Graph differential dependencies (validation)
  • Frequent subgraphs (discovery)

Transactional data patterns

Event sequence data patterns

  • Frequent episode mining episode (discovery)
  • Maximal frequent episode (discovery)
  • Top-k frequent episode (discovery)

Desbordante can be used via two interfaces:

  • Console application. This is a classic command-line interface that aims to provide basic profiling functionality, i.e. discovery and validation of patterns. A user can specify pattern type, task type, algorithm, input file(s) and output results to the screen or into a file.
  • Python bindings. Desbordante functionality can be accessed from within Python programs by employing the Desbordante Python library. This interface offers everything that is currently provided by the console version and allows advanced use, such as building interactive applications and designing scenarios for solving a particular real-life task. Relational data processing algorithms accept pandas DataFrames as input, allowing the user to conveniently preprocess the data before mining patterns.

A brief introduction to the tool and its use cases can be found here (in English) and here (in Russian). Also, an extensive list of tutorial examples that cover each supported pattern is available here.

Table of Contents

Console

For information about the console interface check the repository.

Python bindings

Desbordante features can be accessed from within Python programs by employing the Desbordante Python library. The library is implemented in the form of Python bindings to the interface of the Desbordante C++ core library, using pybind11. Apart from discovery and validation of patterns, this interface is capable of providing valuable additional information which can, for example, describe why a given pattern does not hold.

We want to demonstrate the power of Desbordante through examples where some patterns are extracted from tabular data, providing non-trivial insights. The patterns are quite complex and require detailed explanations, as well as a significant amount of code. This takes up quite a bit of space. Therefore, we do not include the actual code here; instead, we provide a clear (albeit simplified) definition and a link to a Colab notebook with interactive examples. The examples themselves are very detailed and allow users to understand the pattern and how to extract it using Desbordante.

  1. Differential Dependencies (DD). DD is a statement of the form X -> Y, where X and Y are sets of attributes. It indicates that for any two rows, $t$ and $s$, if the attributes in $X$ are similar, then the attributes in $Y$ will also be similar. The similarity for each attribute is defined as: $diff(t[X_i], s[X_i]) \in [val_1, val_2]$, where $t[X_i]$ is the value of attribute $X_i$ in row $t$, $val$ is a constant, and $diff$ is a function that typically calculates the difference, often through simple subtraction. A live Python example that provides insight into the definition and demonstrates how to use this pattern in Desbordante is available here.
  2. Numeric Association Rules (NAR). NAR is a statement of the form X -> Y, where X and Y are conditions, specified on disjoint sets of attributes. Each condition takes a form of $A_1 \wedge A_2 \wedge \ldots \wedge A_n$, where $A_i$ is either $Attribute_i \in$ $[constant_{i}^{1}; constant_{i}^{2}]$ or $Attribute_i$ = $constant_i^3$. Furthermore, the statement includes the support (sup) and confidence (conf) values, which lie in $[0; 1]$. The rule can be interpreted as follows: 1) the supp share of rows in the dataset satisfies both the X and Y conditions, and 2) the conf share of rows that satisfy the X also satisfies Y. A live Python example that provides insight into the definition and demonstrates how to use this pattern in Desbordante is available here.
  3. Matching Dependencies (MD). MD is a statement of the form X -> Y, where X and Y are sets of so-called column matches. Each column match includes: 1) a metric (e.g., Levenshtein distance, Jaccard similarity, etc.), 2) a left column, and 3) a right column. Note that this pattern may involve two tables in its column matches. Finally, each match has its own threshold, which is applied to the corresponding metric and lies in the $[0; 1]$ range. The dependency can be interpreted as follows: any two records that satisfy X will also satisfy Y. A live Python example that provides insight into the definition and demonstrates how to use this pattern in Desbordante is available here.
  4. Denial Constraints (DC). A denial constraint is a statement that says: "For all pairs of rows in a table, it should never happen that some condition is true". Formally, DC $\varphi$ is a conjunction of predicates of the following form: $\forall s,t \in R, s \neq t: \lnot (p_1 \wedge \ldots \wedge p_m)$. Each $p_k$ has the form $column_i$ $op$ $column_j$, where $op \in {>, <, \leq, \geq, =, \neq}$. A live Python example that provides insight into the definition and demonstrates how to use this pattern in Desbordante is available here

Desbordante offers examples for each supported pattern, sometimes several if the pattern is complex or needs to highlight its unique characteristics compared to others in the same family. We have mentioned only a small portion here, which is available in Colab. The rest can be found in our example folder.

Finally, Desbordante allows end users to solve various data quality problems by constructing ad-hoc Python programs, incorporating different Python libraries, and utilizing the search and validation of various patterns. To demonstrate the power of this approach, we have implemented several demo scenarios:

  1. Typo detection
  2. Data deduplication
  3. Anomaly detection

There is also an interactive demo for all of them , and all of these python scripts are here. The ideas behind them are briefly discussed in this preprint (Section 3).

I still don't understand how to use Desbordante and patterns :(

No worries! Desbordante offers a novel type of data profiling, which may require that you first familiarize yourself with its concepts and usage. The most challenging part of Desbordante are the primitives: their definitions and applications in practice. To help you get started, here’s a step-by-step guide:

  1. First of all, we provide several examples for each supported pattern. These examples illustrate both the pattern itself and how to use it in Python. You can check them out here.
  2. Each of our patterns was introduced in a research paper. These papers typically provide a formal definition of the pattern, examples of use, and its application scope. We recommend at least skimming through them. Don't be discouraged by the complexity of the papers! To effectively use the patterns, you only need to read the more accessible parts, such as the introduction and the example sections.
  3. Finally, do not hesitate to ask questions in the mailing list (link below) or create an issue.

Papers about patterns

Here is a list of papers about patterns, organized in the recommended reading order in each item.

Tabular data patterns

Graph data patterns

Transactional data patterns

Event sequence data patterns

Installation (this is what you probably want if you are not a project maintainer)

Desbordante is available at the Python Package Index (PyPI). Dependencies:

  • Python >=3.10

To install Desbordante type:

$ pip install desbordante

However, as Desbordante core uses C++, additional requirements on the machine are imposed. Therefore this installation option may not work for everyone. Currently, only manylinux_2_28 (Ubuntu 21.04+) and macOS 15.0+ (arm64, x86_64) is supported. If the above does not work for you consider building from sources.

Build instructions

Ubuntu and macOS

The following instructions are expected to work for Ubuntu 21.04+ LTS and macOS Sequoia 15.0+ (Apple Silicon).

Dependencies

Prior to cloning the repository and attempting to build the project, ensure that you have the following software:

  • GNU GCC, version 14+, LLVM Clang, version 16+, or Apple Clang, version 16+
  • CMake, version 3.25+
  • Boost library built with compiler you're going to use (GCC or Clang), version 1.85-1.86, 1.88+

Instructions below are given for GCC (on Linux) and Apple Clang (on macOS). Instructions for other supported compilers can be found in Desbordante wiki.

Ubuntu dependencies installation (GCC)

For Ubuntu versions earlier than 24.04, you need to add the Kitware APT repository to your system by following their official guide to install the latest version of CMake.

Then run the following commands:

sudo apt update && sudo apt upgrade
sudo apt install g++ cmake ninja-build libboost-all-dev python3 python3-venv
export CXX=g++

The last line sets g++ as CMake compiler in your terminal session. You can also set it by default in all sessions: echo 'export CXX=g++' >> ~/.profile

For Ubuntu 24.04 and above, you can skip to the build steps. For older versions the Ubuntu APT repository might not have a compatible version of Boost, so you'll need to install it manually:

wget https://archives.boost.io/release/1.89.0/source/boost_1_89_0.tar.gz
tar xzvf boost_1_89_0.tar.gz
cd boost_1_89_0 && ./bootstrap.sh
sudo ./b2 install --prefix=/usr/

macOS dependencies installation (Apple Clang)

Install Xcode Command Line Tools if you don't have them. Run:

xcode-select --install

Follow the prompts to continue.

To install the build dependencies on macOS we recommend to use Homebrew package manager. With Homebrew installed, run the following commands:

brew install cmake boost

After installation, check cmake --version. If command is not found, then you need to add to environment path to homebrew installed packages. To do this open ~/.zprofile (for Zsh) or ~/.bash_profile (for Bash) and add to the end of the file the output of brew shellenv. After that, restart the terminal and check the version of CMake again, now it should be displayed.

Run the following commands:

export CXX=clang++
export BOOST_ROOT=$(brew --prefix boost)

These commands set Apple Clang and Homebrew Boost as default in CMake in your terminal session. You can also add them to the end of ~/.profile to set this by default in all sessions.

Building the project

Building the Python module using pip

Clone the repository, change the current directory to the project directory and run the following commands:

python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install .

Now it is possible to import desbordante as a module from within the created virtual environment.

Building tests & the Python module manually

Build the tests themselves:

./build.sh

The Python module can be built by providing the --pybind switch:

./build.sh --pybind 

See ./build.sh --help for more available options.

The ./build.sh script generates the following file structure in /path/to/desbordante-core/build/target:

├───input_data
│   └───some-sample-csv\'s.csv
├───desbordante.cpython-*.so

The input_data directory contains several .csv files that are used by unit tests. You can run tests with CTest from any directory in the Desbordante tree:

ctest --test-dir build --exclude-regex ".*HeavyDatasets.*" -j $JOBS

where $JOBS is the desired number of concurrent jobs.

desbordante.cpython-*.so is a Python module, packaging Python bindings for the Desbordante core library. In order to use it, simply import it:

cd build/target
python3
>>> import desbordante

The core library uses spdlog. All log messages are automatically bridged to Python's standard logging module.

Log level control:

import logging
import desbordante

# Get the logger provided by the library
log = logging.getLogger("desbordante")

# Set the desired log level using standard logging constants
log.setLevel(logging.INFO) # Or logging.DEBUG, logging.TRACE, etc.

Troubleshooting

No type hints in IDE

If type hints don't work for you in Visual Studio Code, for example, then install stubs using the command:

pip install desbordante-stubs

NOTE: Stubs may not fully support current version of desbordante package, as they are updated independently.

Cite

If you use this software for research, please cite our core paper:

@inproceedings{10.1145/3703323.3703725,
   author = {Chernishev, George and Polyntsov, Michael and Chizhov, Anton and Stupakov, Kirill and Shchuckin, Ilya and Smirnov, Alexander and Strutovsky, Maxim and Shlyonskikh, Alexey and Firsov, Mikhail and Manannikov, Stepan and Bobrov, Nikita and Goncharov, Daniil and Barutkin, Ilia and Yakshigulov, Vadim and Shalnev, Vladislav and Muraviev, Kirill and Rakhmukova, Anna and Shcheka, Dmitriy and Chernikov, Anton and Kuzin, Yakov and Sinelnikov, Michael and Abrosimov, Grigorii and Popov, Dmitriy and Demchenko, Artem and Belokonny, Sergey and Soloveva, Liana-Iuliia and Kurbatov, Yaroslav and Vyrodov, Mikhail and Saliou, Arthur and Gaisin, Eduard and Smirnov, Kirill},
   title = {Desbordante: from benchmarking suite to high-performance science-intensive data profiler},
   year = {2025},
   isbn = {9798400711244},
   publisher = {Association for Computing Machinery},
   address = {New York, NY, USA},
   url = {https://doi.org/10.1145/3703323.3703725},
   doi = {10.1145/3703323.3703725},
   booktitle = {Proceedings of the 8th International Conference on Data Science and Management of Data (12th ACM IKDD CODS and 30th COMAD)},
   pages = {234--243},
   numpages = {10},
   keywords = {Data Mining, Data Profiling, Pattern Extraction, Data Analysis, Knowledge Discovery, Data Exploration, Anomaly Detection, Data Wrangling},
   location = {},
   series = {CODS-COMAD '24}
}

or cite one of our papers, if you use a particular part:

  1. George Chernishev, et al. Solving Data Quality Problems with Desbordante: a Demo. CoRR abs/2307.14935 (2023).
  2. M. Strutovskiy, N. Bobrov, K. Smirnov and G. Chernishev, "Desbordante: a Framework for Exploring Limits of Dependency Discovery Algorithms," 2021 29th Conference of Open Innovations Association (FRUCT), 2021, pp. 344-354.
  3. A. Smirnov, A. Chizhov, I. Shchuckin, N. Bobrov and G. Chernishev, "Fast Discovery of Inclusion Dependencies with Desbordante," 2023 33rd Conference of Open Innovations Association (FRUCT), Zilina, Slovakia, 2023, pp. 264-275.
  4. A. Chernikov, Y. Litvinov, K. Smirnov, and G. Chernishev, "FastGFDs: Efficient Validation of Graph Functional Dependencies with Desbordante," 2023 33rd Conference of Open Innovations Association (FRUCT), Zilina, Slovakia, 2023, Issue 2 (Works in Progress), pp. 346-352.
  5. Y. Kuzin, D. Shcheka, M. Polyntsov, K. Stupakov, M. Firsov and G. Chernishev, "Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms," 2024 35th Conference of Open Innovations Association (FRUCT), Tampere, Finland, 2024, pp. 413-424.
  6. I. Barutkin, M. Fofanov, S. Belokonny, V. Makeev and G. Chernishev, "Extending Desbordante with Probabilistic Functional Dependency Discovery Support," 2024 35th Conference of Open Innovations Association (FRUCT), Tampere, Finland, 2024, pp. 158-169.
  7. A. Shlyonskikh, M. Sinelnikov, D. Nikolaev, Y. Litvinov and G. Chernishev, "Lightning Fast Matching Dependency Discovery with Desbordante," 2024 36th Conference of Open Innovations Association (FRUCT), Lappeenranta, Finland, 2024, pp. 729-740.
  8. M. Ivanov, M. Smirnov, A. Strazdina and G. Chernishev, "Scalable Maximal Frequent Episode Mining with Desbordante," 2026 39th Conference of Open Innovations Association (FRUCT), Helsinki, Finland, 2026, pp. 102-113.
  9. I. Kozhukov et al., "Efficient Discovery of Conditional Dependencies with Desbordante," 2026 39th Conference of Open Innovations Association (FRUCT), Helsinki, Finland, 2026, pp. 130-141.

Contacts and Q&A

If you have any questions regarding the tool usage you can ask it in our google group. To contact dev team email George Chernishev, Alexey Shlyonskikh or Michael Polyntsov.

anomaly-detection
correlations
data-analytics
data-cleaning
data-cleansing
data-engineering
data-exploration
data-mining
data-mining-algorithms
data-preprocessing
data-profiling
data-science
data-wrangling
exploratory-data-analysis
feature-engineering
feature-extraction
feature-selection
knowledge-discovery
spreadsheets
tabular-data

Significant stargazers

Vincent Koc

2,112 followers · starred May 2024

Daria Fomina

10 followers · starred Sep 2026

Maksim Dergousov

14 followers · starred Feb 2025

Desbordante/desbordante-core

Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.

C++

510

2,094 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

Show HN: Desbordante 2.5.0 Released

1

Sep 30, 2026

README

Downloads Downloads License: AGPL v3

Desbordante logo

About

Desbordante is a high-performance data profiler that is capable of discovering and validating many different patterns in data using various algorithms.

Tool Positioning

Why Desbordante? Why not just use an LLM or other black-box machine learning approaches?

Unlike these, Desbordante takes a symbolic, explainable approach to data analysis, producing explicit rules or rule-shaped artifacts that are verifiable, such as data dependencies, association rules, or other patterns. This is a distinct paradigm for data analysis that complements black-box learning. In contrast to free-form model outputs, discovered patterns have precisely defined semantics and can be algorithmically validated against the underlying data.

These patterns add data-grounded and rule-based structure, which provides useful properties for systems built around machine learning and large language models: consistency, provenance, explainability, and auditability. Further, the patterns can serve as inputs to downstream data-quality checks, validation procedures, integrity constraints, training-data analysis, and other rule-based safeguards. They can also be used to choose a direction of an ablation study, to generate examples that violate rules for use in machine learning model training and testing, to assist fact completion, and to facilitate acceptance checks and guardrails for pipelines involving retrieval-augmented generation. This is roughly the motivation used for the work in [1], but it applies more broadly to the patterns Desbordante works with.

[1] François Amat. Mining Rules on Tabular Data. Institut Polytechnique de Paris, 2025.

From other angles, the discovered patterns can be used in many ways:

  • For scientific data, especially those obtained experimentally, an interesting pattern allows to formulate a hypothesis that could lead to a scientific discovery. In some cases it even allows one to draw conclusions immediately, if there is enough data. At the very least, the found pattern can be used to provide a direction for further study.

  • For business data, it is also possible to obtain a hypothesis based on found patterns. However, there are more down-to-earth and more in-demand applications in this case: clearing errors in data, finding and removing inexact duplicates, performing schema matching, and many more.

  • For database data, found patterns can help with recovering/defining primary and foreign keys, checking/setting up all kinds of integrity constraints.

General

Desbordante supports three types of tasks.

The Discovery task is designed to identify all instances of a specified pattern type of a given dataset.

The Validation task is different: it is designed to check whether a specified pattern instance is present in a given dataset. This task not only returns True or False, but it also explains why the instance does not hold (e.g. it can list table rows with conflicting values).

For some patterns Desbordante supports a dynamic task variant. The distinguishing feature of dynamic algorithms compared to classic (static) algorithms is that after a result is obtained, the table can be changed and a dynamic algorithm will update the result based just on those changes instead of processing the whole table again. As a result, they can be up to several orders of magnitude faster than classic (static) ones in some situations.

Desbordante support four data types:

  • Tabular
  • Graph
  • Transactional
  • Event sequences

Tabular data patterns

  • Exact functional dependencies (discovery and validation)
  • Approximate functional dependencies, with
  • Probabilistic functional dependencies, with PerTuple and PerValue metrics (discovery and validation)
  • Classic soft functional dependencies (with correlations), with $\rho$ metric (discovery and validation)
  • Dynamic validation of exact and approximate ($g_1$) functional dependencies
  • Relaxed functional dependencies (discovery)
  • Numerical dependencies (validation)
  • Conditional functional dependencies (discovery and validation)
  • Inclusion dependencies
  • Conditional inclusion dependencies (discovery and validation)
  • Order dependencies:
    • set-based axiomatization (discovery and validation)
    • list-based axiomatization (discovery)
  • Approximate order dependencies:
    • set-based axiomatization (discovery and validation)
  • Sequential dependencies (validation)
  • Metric functional dependencies (validation)
  • Fuzzy algebraic constraints (discovery)
  • Domain Probabilistic and Approximate Constraints (validation)
  • Differential Dependencies (discovery and validation)
  • Unique column combinations:
    • Exact unique column combination (discovery and validation)
    • Approximate unique column combination, with $g_1$ metric (discovery and validation)
  • Numerical association rules (discovery)
  • Matching dependencies (discovery and validation)
  • Denial constraints

Graph data patterns

  • Graph functional dependencies (discovery and validation)
  • Graph differential dependencies (validation)
  • Frequent subgraphs (discovery)

Transactional data patterns

Event sequence data patterns

  • Frequent episode mining episode (discovery)
  • Maximal frequent episode (discovery)
  • Top-k frequent episode (discovery)

Desbordante can be used via two interfaces:

  • Console application. This is a classic command-line interface that aims to provide basic profiling functionality, i.e. discovery and validation of patterns. A user can specify pattern type, task type, algorithm, input file(s) and output results to the screen or into a file.
  • Python bindings. Desbordante functionality can be accessed from within Python programs by employing the Desbordante Python library. This interface offers everything that is currently provided by the console version and allows advanced use, such as building interactive applications and designing scenarios for solving a particular real-life task. Relational data processing algorithms accept pandas DataFrames as input, allowing the user to conveniently preprocess the data before mining patterns.

A brief introduction to the tool and its use cases can be found here (in English) and here (in Russian). Also, an extensive list of tutorial examples that cover each supported pattern is available here.

Table of Contents

Console

For information about the console interface check the repository.

Python bindings

Desbordante features can be accessed from within Python programs by employing the Desbordante Python library. The library is implemented in the form of Python bindings to the interface of the Desbordante C++ core library, using pybind11. Apart from discovery and validation of patterns, this interface is capable of providing valuable additional information which can, for example, describe why a given pattern does not hold.

We want to demonstrate the power of Desbordante through examples where some patterns are extracted from tabular data, providing non-trivial insights. The patterns are quite complex and require detailed explanations, as well as a significant amount of code. This takes up quite a bit of space. Therefore, we do not include the actual code here; instead, we provide a clear (albeit simplified) definition and a link to a Colab notebook with interactive examples. The examples themselves are very detailed and allow users to understand the pattern and how to extract it using Desbordante.

  1. Differential Dependencies (DD). DD is a statement of the form X -> Y, where X and Y are sets of attributes. It indicates that for any two rows, $t$ and $s$, if the attributes in $X$ are similar, then the attributes in $Y$ will also be similar. The similarity for each attribute is defined as: $diff(t[X_i], s[X_i]) \in [val_1, val_2]$, where $t[X_i]$ is the value of attribute $X_i$ in row $t$, $val$ is a constant, and $diff$ is a function that typically calculates the difference, often through simple subtraction. A live Python example that provides insight into the definition and demonstrates how to use this pattern in Desbordante is available here.
  2. Numeric Association Rules (NAR). NAR is a statement of the form X -> Y, where X and Y are conditions, specified on disjoint sets of attributes. Each condition takes a form of $A_1 \wedge A_2 \wedge \ldots \wedge A_n$, where $A_i$ is either $Attribute_i \in$ $[constant_{i}^{1}; constant_{i}^{2}]$ or $Attribute_i$ = $constant_i^3$. Furthermore, the statement includes the support (sup) and confidence (conf) values, which lie in $[0; 1]$. The rule can be interpreted as follows: 1) the supp share of rows in the dataset satisfies both the X and Y conditions, and 2) the conf share of rows that satisfy the X also satisfies Y. A live Python example that provides insight into the definition and demonstrates how to use this pattern in Desbordante is available here.
  3. Matching Dependencies (MD). MD is a statement of the form X -> Y, where X and Y are sets of so-called column matches. Each column match includes: 1) a metric (e.g., Levenshtein distance, Jaccard similarity, etc.), 2) a left column, and 3) a right column. Note that this pattern may involve two tables in its column matches. Finally, each match has its own threshold, which is applied to the corresponding metric and lies in the $[0; 1]$ range. The dependency can be interpreted as follows: any two records that satisfy X will also satisfy Y. A live Python example that provides insight into the definition and demonstrates how to use this pattern in Desbordante is available here.
  4. Denial Constraints (DC). A denial constraint is a statement that says: "For all pairs of rows in a table, it should never happen that some condition is true". Formally, DC $\varphi$ is a conjunction of predicates of the following form: $\forall s,t \in R, s \neq t: \lnot (p_1 \wedge \ldots \wedge p_m)$. Each $p_k$ has the form $column_i$ $op$ $column_j$, where $op \in {>, <, \leq, \geq, =, \neq}$. A live Python example that provides insight into the definition and demonstrates how to use this pattern in Desbordante is available here

Desbordante offers examples for each supported pattern, sometimes several if the pattern is complex or needs to highlight its unique characteristics compared to others in the same family. We have mentioned only a small portion here, which is available in Colab. The rest can be found in our example folder.

Finally, Desbordante allows end users to solve various data quality problems by constructing ad-hoc Python programs, incorporating different Python libraries, and utilizing the search and validation of various patterns. To demonstrate the power of this approach, we have implemented several demo scenarios:

  1. Typo detection
  2. Data deduplication
  3. Anomaly detection

There is also an interactive demo for all of them , and all of these python scripts are here. The ideas behind them are briefly discussed in this preprint (Section 3).

I still don't understand how to use Desbordante and patterns :(

No worries! Desbordante offers a novel type of data profiling, which may require that you first familiarize yourself with its concepts and usage. The most challenging part of Desbordante are the primitives: their definitions and applications in practice. To help you get started, here’s a step-by-step guide:

  1. First of all, we provide several examples for each supported pattern. These examples illustrate both the pattern itself and how to use it in Python. You can check them out here.
  2. Each of our patterns was introduced in a research paper. These papers typically provide a formal definition of the pattern, examples of use, and its application scope. We recommend at least skimming through them. Don't be discouraged by the complexity of the papers! To effectively use the patterns, you only need to read the more accessible parts, such as the introduction and the example sections.
  3. Finally, do not hesitate to ask questions in the mailing list (link below) or create an issue.

Papers about patterns

Here is a list of papers about patterns, organized in the recommended reading order in each item.

Tabular data patterns

Graph data patterns

Transactional data patterns

Event sequence data patterns

Installation (this is what you probably want if you are not a project maintainer)

Desbordante is available at the Python Package Index (PyPI). Dependencies:

  • Python >=3.10

To install Desbordante type:

$ pip install desbordante

However, as Desbordante core uses C++, additional requirements on the machine are imposed. Therefore this installation option may not work for everyone. Currently, only manylinux_2_28 (Ubuntu 21.04+) and macOS 15.0+ (arm64, x86_64) is supported. If the above does not work for you consider building from sources.

Build instructions

Ubuntu and macOS

The following instructions are expected to work for Ubuntu 21.04+ LTS and macOS Sequoia 15.0+ (Apple Silicon).

Dependencies

Prior to cloning the repository and attempting to build the project, ensure that you have the following software:

  • GNU GCC, version 14+, LLVM Clang, version 16+, or Apple Clang, version 16+
  • CMake, version 3.25+
  • Boost library built with compiler you're going to use (GCC or Clang), version 1.85-1.86, 1.88+

Instructions below are given for GCC (on Linux) and Apple Clang (on macOS). Instructions for other supported compilers can be found in Desbordante wiki.

Ubuntu dependencies installation (GCC)

For Ubuntu versions earlier than 24.04, you need to add the Kitware APT repository to your system by following their official guide to install the latest version of CMake.

Then run the following commands:

sudo apt update && sudo apt upgrade
sudo apt install g++ cmake ninja-build libboost-all-dev python3 python3-venv
export CXX=g++

The last line sets g++ as CMake compiler in your terminal session. You can also set it by default in all sessions: echo 'export CXX=g++' >> ~/.profile

For Ubuntu 24.04 and above, you can skip to the build steps. For older versions the Ubuntu APT repository might not have a compatible version of Boost, so you'll need to install it manually:

wget https://archives.boost.io/release/1.89.0/source/boost_1_89_0.tar.gz
tar xzvf boost_1_89_0.tar.gz
cd boost_1_89_0 && ./bootstrap.sh
sudo ./b2 install --prefix=/usr/

macOS dependencies installation (Apple Clang)

Install Xcode Command Line Tools if you don't have them. Run:

xcode-select --install

Follow the prompts to continue.

To install the build dependencies on macOS we recommend to use Homebrew package manager. With Homebrew installed, run the following commands:

brew install cmake boost

After installation, check cmake --version. If command is not found, then you need to add to environment path to homebrew installed packages. To do this open ~/.zprofile (for Zsh) or ~/.bash_profile (for Bash) and add to the end of the file the output of brew shellenv. After that, restart the terminal and check the version of CMake again, now it should be displayed.

Run the following commands:

export CXX=clang++
export BOOST_ROOT=$(brew --prefix boost)

These commands set Apple Clang and Homebrew Boost as default in CMake in your terminal session. You can also add them to the end of ~/.profile to set this by default in all sessions.

Building the project

Building the Python module using pip

Clone the repository, change the current directory to the project directory and run the following commands:

python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install .

Now it is possible to import desbordante as a module from within the created virtual environment.

Building tests & the Python module manually

Build the tests themselves:

./build.sh

The Python module can be built by providing the --pybind switch:

./build.sh --pybind 

See ./build.sh --help for more available options.

The ./build.sh script generates the following file structure in /path/to/desbordante-core/build/target:

├───input_data
│   └───some-sample-csv\'s.csv
├───desbordante.cpython-*.so

The input_data directory contains several .csv files that are used by unit tests. You can run tests with CTest from any directory in the Desbordante tree:

ctest --test-dir build --exclude-regex ".*HeavyDatasets.*" -j $JOBS

where $JOBS is the desired number of concurrent jobs.

desbordante.cpython-*.so is a Python module, packaging Python bindings for the Desbordante core library. In order to use it, simply import it:

cd build/target
python3
>>> import desbordante

The core library uses spdlog. All log messages are automatically bridged to Python's standard logging module.

Log level control:

import logging
import desbordante

# Get the logger provided by the library
log = logging.getLogger("desbordante")

# Set the desired log level using standard logging constants
log.setLevel(logging.INFO) # Or logging.DEBUG, logging.TRACE, etc.

Troubleshooting

No type hints in IDE

If type hints don't work for you in Visual Studio Code, for example, then install stubs using the command:

pip install desbordante-stubs

NOTE: Stubs may not fully support current version of desbordante package, as they are updated independently.

Cite

If you use this software for research, please cite our core paper:

@inproceedings{10.1145/3703323.3703725,
   author = {Chernishev, George and Polyntsov, Michael and Chizhov, Anton and Stupakov, Kirill and Shchuckin, Ilya and Smirnov, Alexander and Strutovsky, Maxim and Shlyonskikh, Alexey and Firsov, Mikhail and Manannikov, Stepan and Bobrov, Nikita and Goncharov, Daniil and Barutkin, Ilia and Yakshigulov, Vadim and Shalnev, Vladislav and Muraviev, Kirill and Rakhmukova, Anna and Shcheka, Dmitriy and Chernikov, Anton and Kuzin, Yakov and Sinelnikov, Michael and Abrosimov, Grigorii and Popov, Dmitriy and Demchenko, Artem and Belokonny, Sergey and Soloveva, Liana-Iuliia and Kurbatov, Yaroslav and Vyrodov, Mikhail and Saliou, Arthur and Gaisin, Eduard and Smirnov, Kirill},
   title = {Desbordante: from benchmarking suite to high-performance science-intensive data profiler},
   year = {2025},
   isbn = {9798400711244},
   publisher = {Association for Computing Machinery},
   address = {New York, NY, USA},
   url = {https://doi.org/10.1145/3703323.3703725},
   doi = {10.1145/3703323.3703725},
   booktitle = {Proceedings of the 8th International Conference on Data Science and Management of Data (12th ACM IKDD CODS and 30th COMAD)},
   pages = {234--243},
   numpages = {10},
   keywords = {Data Mining, Data Profiling, Pattern Extraction, Data Analysis, Knowledge Discovery, Data Exploration, Anomaly Detection, Data Wrangling},
   location = {},
   series = {CODS-COMAD '24}
}

or cite one of our papers, if you use a particular part:

  1. George Chernishev, et al. Solving Data Quality Problems with Desbordante: a Demo. CoRR abs/2307.14935 (2023).
  2. M. Strutovskiy, N. Bobrov, K. Smirnov and G. Chernishev, "Desbordante: a Framework for Exploring Limits of Dependency Discovery Algorithms," 2021 29th Conference of Open Innovations Association (FRUCT), 2021, pp. 344-354.
  3. A. Smirnov, A. Chizhov, I. Shchuckin, N. Bobrov and G. Chernishev, "Fast Discovery of Inclusion Dependencies with Desbordante," 2023 33rd Conference of Open Innovations Association (FRUCT), Zilina, Slovakia, 2023, pp. 264-275.
  4. A. Chernikov, Y. Litvinov, K. Smirnov, and G. Chernishev, "FastGFDs: Efficient Validation of Graph Functional Dependencies with Desbordante," 2023 33rd Conference of Open Innovations Association (FRUCT), Zilina, Slovakia, 2023, Issue 2 (Works in Progress), pp. 346-352.
  5. Y. Kuzin, D. Shcheka, M. Polyntsov, K. Stupakov, M. Firsov and G. Chernishev, "Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms," 2024 35th Conference of Open Innovations Association (FRUCT), Tampere, Finland, 2024, pp. 413-424.
  6. I. Barutkin, M. Fofanov, S. Belokonny, V. Makeev and G. Chernishev, "Extending Desbordante with Probabilistic Functional Dependency Discovery Support," 2024 35th Conference of Open Innovations Association (FRUCT), Tampere, Finland, 2024, pp. 158-169.
  7. A. Shlyonskikh, M. Sinelnikov, D. Nikolaev, Y. Litvinov and G. Chernishev, "Lightning Fast Matching Dependency Discovery with Desbordante," 2024 36th Conference of Open Innovations Association (FRUCT), Lappeenranta, Finland, 2024, pp. 729-740.
  8. M. Ivanov, M. Smirnov, A. Strazdina and G. Chernishev, "Scalable Maximal Frequent Episode Mining with Desbordante," 2026 39th Conference of Open Innovations Association (FRUCT), Helsinki, Finland, 2026, pp. 102-113.
  9. I. Kozhukov et al., "Efficient Discovery of Conditional Dependencies with Desbordante," 2026 39th Conference of Open Innovations Association (FRUCT), Helsinki, Finland, 2026, pp. 130-141.

Contacts and Q&A

If you have any questions regarding the tool usage you can ask it in our google group. To contact dev team email George Chernishev, Alexey Shlyonskikh or Michael Polyntsov.

anomaly-detection
correlations
data-analytics
data-cleaning
data-cleansing
data-engineering
data-exploration
data-mining
data-mining-algorithms
data-preprocessing
data-profiling
data-science
data-wrangling
exploratory-data-analysis
feature-engineering
feature-extraction
feature-selection
knowledge-discovery
spreadsheets
tabular-data

Significant stargazers

Vincent Koc

2,112 followers · starred May 2024

Daria Fomina

10 followers · starred Sep 2026

Maksim Dergousov

14 followers · starred Feb 2025

Languages

C++

96.5%

CMake

2.4%