Unit Testing Data with Deequ
Summary: Deequ is a Spark-based library that automates data quality verification at scale with a declarative constraints API and custom validation. Open-source, production-ready at Amazon; scales to billions of records and supports incremental validation. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sebastian Schelter (Amazon)
- 2. Felix Biessmann (Amazon)
- 3. Dustin Lange (Amazon)
- 4. Tammo Rukat (Amazon)
- 5. Philipp Schmidt (Amazon)
- 6. Stephan Seufert (Amazon)
- 7. Pierre Brunelle (Amazon)
- 8. Andrey Taptunov (Amazon)
BibTeX Citation
@inproceedings{schelter_sigmod19,
title = {{Unit Testing Data with Deequ}},
author = {Schelter, Sebastian and Biessmann, Felix and Lange, Dustin and Rukat, Tammo and Schmidt, Philipp and Seufert, Stephan and Brunelle, Pierre and Taptunov, Andrey},
series = {{SIGMOD} '19},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3299869.3320210},
url = {https://dl.acm.org/doi/10.1145/3299869.3320210},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,232 | Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications | 2023 | SIGMOD | 5.6659017e-05 |
| 7,690 | Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes | 2021 | SIGMOD | 5.5669307e-05 |
| 10,785 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD | 5.093636e-05 |
| 11,048 | Demonstrating Matelda for Multi-Table Error Detection | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 1,147 | Data Management Challenges in Production Machine Learning | 2017 | SIGMOD | 0.00011974846 |
| 1,350 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.00011065626 |
| 5,182 | Probabilistic Demand Forecasting at Scale | 2017 | VLDB | 6.3287692e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,630 | TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines | 2020 | SIGMOD |
| 2 | 1,190 | Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark | 2018 | SIGMOD |
| 3 | 8,399 | DataProf: Semantic Profiling for Iterative Data Cleansing and Business Rule Acquisition | 2018 | SIGMOD |
| 4 | 11,487 | To UDFs and Beyond: Demonstration of a Fully Decomposed Data Processor for General Data Wrangling Tasks | 2023 | VLDB |
| 5 | 9,649 | Transforming ML Predictive Pipelines into SQL with MASQ | 2021 | SIGMOD |
| 6 | 4,373 | Data Debugging and Exploration with Vizier | 2019 | SIGMOD |
| 7 | 6,315 | Data Collection and Quality Challenges for Deep Learning | 2020 | VLDB |
| 8 | 9,066 | TsQuality: Measuring Time Series Data Quality in Apache IoTDB | 2023 | VLDB |
| 9 | 9,069 | DQDF: Data-Quality-Aware Dataframes | 2022 | VLDB |
| 10 | 1,350 | Automating Large-Scale Data Quality Verification | 2018 | VLDB |