cs.DBJul 26, 2026

Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms

Authors: Yakov KuzinDmitriy ShchekaMichael PolyntsovKirill StupakovMikhail FirsovGeorge Chernishev

Organizations: Saint-Petersburg University Saint-Petersburg, Russia

Abstract

Science-intensive data profiling focuses on discovery and validation of various patterns in datasets. This study considers discovery of one such pattern - order dependency (OD). Simply put, OD states that some list of columns is ordered according to another one. It is of use for database query optimization, data cleaning and deduplication, anomaly detection, and much more. Existing discovery methods have approached this problem solely from the algorithmic standpoint, without focusing on the implementation side. At the same time, this problem is very computationally intensive, and therefore this part should not be ignored, as it brings ODs closer to industrial use. In this paper, we study two algorithms for OD discovery which target different OD axiomatizations - FASTOD and ORDER. We start by reimplementing these algorithms in C++ in order to speed them up and lower their memory consumption. We then analyze their bottlenecks and propose several techniques which improve their performance even further. To perform evaluation, we have implemented these algorithms inside Desbordante - a science-intensive, high-performance, and open-source data profiling tool developed in C++. Experiments have demonstrated a performance improvement of up to 3x obtained by reimplemented versions, and, with the application of our techniques, up to 10x. Memory consumption has been lowered by up to 2.9x.

Explore similar work

CardsList
  1. Fast Discovery of Inclusion Dependencies with Desbordante

    Aug 3, 2026Alexander Smirnov, Anton Chizhov, Ilya Shchuckin +2ProfilingData Mining

  2. Lightning Fast Matching Dependency Discovery with Desbordante

    Jul 12, 2026Alexey Shlyonskikh, Michael Sinelnikov, Daniil Nikolaev +2Data MiningSemantic Deduplication

  3. Efficient Discovery of Conditional Dependencies with Desbordante

    Jul 4, 2026Ivan Kozhukov, Dmitry Fedoseev, Maksim Emelyanov +4Data MiningFalse Discovery Rate