Unit Testing Data with Deequ
Summary: Deequ is a Spark-based library that automates data quality verification at scale with a declarative constraints API and custom validation. Open-source, production-ready at Amazon; scales to billions of records and supports incremental validation. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sebastian Schelter (Amazon)
- 2. Felix Biessmann (Amazon)
- 3. Dustin Lange (Amazon)
- 4. Tammo Rukat (Amazon)
- 5. Philipp Schmidt (Amazon)
- 6. Stephan Seufert (Amazon)
- 7. Pierre Brunelle (Amazon)
- 8. Andrey Taptunov (Amazon)
BibTeX Citation
@inproceedings{schelter_sigmod19,
title = {{Unit Testing Data with Deequ}},
author = {Schelter, Sebastian and Biessmann, Felix and Lange, Dustin and Rukat, Tammo and Schmidt, Philipp and Seufert, Stephan and Brunelle, Pierre and Taptunov, Andrey},
series = {{SIGMOD} '19},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3299869.3320210},
url = {https://dl.acm.org/doi/10.1145/3299869.3320210},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,574 | Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications | 2023 | SIGMOD | 6.0773771e-05 |
| 7,742 | Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes | 2021 | SIGMOD | 5.4599258e-05 |
| 9,239 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD | 5.2032182e-05 |
| 11,417 | Demonstrating Matelda for Multi-Table Error Detection | 2025 | VLDB | 4.9769913e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 23 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00055384955 |
| 1,153 | Data Management Challenges in Production Machine Learning | 2017 | SIGMOD | 0.00011793347 |
| 1,308 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.00011073863 |
| 5,308 | Probabilistic Demand Forecasting at Scale | 2017 | VLDB | 6.1844285e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,706 | TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines | 2020 | SIGMOD |
| 2 | 1,123 | Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark | 2018 | SIGMOD |
| 3 | 8,513 | DataProf: Semantic Profiling for Iterative Data Cleansing and Business Rule Acquisition | 2018 | SIGMOD |
| 4 | 11,803 | To UDFs and Beyond: Demonstration of a Fully Decomposed Data Processor for General Data Wrangling Tasks | 2023 | VLDB |
| 5 | 9,835 | Transforming ML Predictive Pipelines into SQL with MASQ | 2021 | SIGMOD |
| 6 | 4,451 | Data Debugging and Exploration with Vizier | 2019 | SIGMOD |
| 7 | 5,074 | Data Collection and Quality Challenges for Deep Learning | 2020 | VLDB |
| 8 | 9,253 | TsQuality: Measuring Time Series Data Quality in Apache IoTDB | 2023 | VLDB |
| 9 | 9,256 | DQDF: Data-Quality-Aware Dataframes | 2022 | VLDB |
| 10 | 1,308 | Automating Large-Scale Data Quality Verification | 2018 | VLDB |