DBScholar

Back to papers

Reptile: Aggregation-level Explanations for Hierarchical Data

Summary: Iterative, human-in-the-loop system that explains and cleans hierarchical data by learning group-level statistics and guiding drill-downs to fix distributive aggregation errors. Introduces factorised learning for aggregation-join queries with hierarchical optimisations, delivering >6× speedups and real-world deployments on Covid-19 and African farmer surveys used for policy-relevant data cleaning. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6368
Venue
SIGMOD
Year
2022
Pagerank
5.1814573e-05
Overall Rank
10,001 | 31.39%
DOI
10.1145/3514221.3517854

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{huang_sigmod22,
        title = {{Reptile: Aggregation-level Explanations for Hierarchical Data}},
        author = {Huang, Zezhou and Wu, Eugene},
        series = {{SIGMOD} '22},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3514221.3517854},
        url = {https://dl.acm.org/doi/10.1145/3514221.3517854},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
11,098 SDEcho: Efficient Explanation of Aggregated Sequence Difference 2025 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 27 of 27 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
112 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00032801121
191 Scorpion: Explaining Away Outliers in Aggregate Queries 2013 VLDB 0.00026096009
376 Discovering Denial Constraints 2013 VLDB 0.00019677674
536 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.0001693369
549 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016692839
661 Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes 2013 SIGMOD 0.0001519162
663 A Formal Approach to Finding Explanations for Database Queries 2014 SIGMOD 0.00015174751
700 Explaining differences in multidimensional aggregates 1999 VLDB 0.00014871032
714 Guided Data Repair 2011 VLDB 0.00014662041
946 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013054126
1,121 ARDA: Automatic Relational Data Augmentation for Machine Learning 2020 VLDB 0.00012093059
1,323 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00011152602
1,476 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010659277
1,792 MacroBase: Prioritizing Attention in Fast Data 2017 SIGMOD 9.7436856e-05
2,161 DIFF: A Relational Interface for Large-Scale Data Explanation 2019 VLDB 9.0606664e-05
2,297 Raha: A Configuration-Free Error Detection System 2019 SIGMOD 8.7897221e-05
3,206 Incremental View Maintenance with Triple Lock Factorization Benefits 2018 SIGMOD 7.6367549e-05
3,361 Cleaning Denial Constraint Violations through Relaxation 2020 SIGMOD 7.4872032e-05
3,380 UGuide – User-Guided Discovery of FD-Detectable Errors 2017 SIGMOD 7.457539e-05
4,392 Multi-Structural Databases 2005 PODS 6.7288249e-05
4,638 Going Beyond Provenance: Explaining Query Answers with Pattern-based Counterbalances 2019 SIGMOD 6.5919801e-05
4,890 Descriptive and Prescriptive Data Cleaning 2014 SIGMOD 6.4590072e-05
5,397 KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing 2015 VLDB 6.232136e-05
5,508 LMFAO: An Engine for Batches of Group-By Aggregates 2020 VLDB 6.1922654e-05
6,763 Smart Drill-Down: A New Data Exploration Operator 2015 VLDB 5.7810308e-05
6,819 Estimating the Impact of Unknown Unknowns on Aggregate Query Results 2016 SIGMOD 5.7635226e-05
8,153 The Cascading Analysts Algorithm 2018 SIGMOD 5.4769543e-05
Previous Page 1 / 1 Next

Semantically Similar Papers