Complaint-Driven Training Data Debugging at Interactive Speeds
Summary: Rain++ enables complaint-driven debugging of training data for inference queries by ranking offending examples from complaints. Precomputation decouples cost from model size, enabling interactive ~1 ms latency for multi-million-parameter models and supporting standing/streaming queries. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Lampros Flokas (Columbia University)
- 2. Weiyuan Wu (Simon Fraser University)
- 3. Yejia Liu (Simon Fraser University)
- 4. Jiannan Wang (Simon Fraser University)
- 5. Nakul Verma (Columbia University)
- 6. Eugene Wu (Columbia University)
BibTeX Citation
@inproceedings{flokas_sigmod22,
title = {{Complaint-Driven Training Data Debugging at Interactive Speeds}},
author = {Flokas, Lampros and Wu, Weiyuan and Liu, Yejia and Wang, Jiannan and Verma, Nakul and Wu, Eugene},
series = {{SIGMOD} '22},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3514221.3517849},
url = {https://dl.acm.org/doi/10.1145/3514221.3517849},
year = {2022}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,982 | XInsight: eXplainable Data Analysis Through The Lens of Causality | 2023 | SIGMOD | 6.328859e-05 |
| 7,538 | Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines | 2023 | SIGMOD | 5.4995874e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 17 | Provenance Semirings | 2007 | PODS | 0.00059752575 |
| 105 | The MADlib Analytics Library or MAD Skills, the SQL | 2012 | VLDB | 0.00033638251 |
| 189 | Scorpion: Explaining Away Outliers in Aggregate Queries | 2013 | VLDB | 0.00025840026 |
| 819 | Provenance for Aggregate Queries | 2011 | PODS | 0.00013666629 |
| 1,051 | Tiresias: The Database Oracle for How-To Queries | 2012 | SIGMOD | 0.00012279821 |
| 1,308 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.0001107886 |
| 2,160 | DIFF: A Relational Interface for Large-Scale Data Explanation | 2019 | VLDB | 8.9364035e-05 |
| 2,598 | Complaint-driven Training Data Debugging for Query 2.0 | 2020 | SIGMOD | 8.2385793e-05 |
| 2,686 | Data X-Ray: A Diagnostic Tool for Data Errors | 2015 | SIGMOD | 8.1308928e-05 |
| 4,195 | PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models | 2020 | SIGMOD | 6.7422305e-05 |
| 4,739 | Going Beyond Provenance: Explaining Query Answers with Pattern-based Counterbalances | 2019 | SIGMOD | 6.4441962e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,363 | PerfGuard: Deploying ML-for-Systems without Performance Regressions, Almost! | 2021 | VLDB |
| 2 | 6,844 | Serving Deep Learning Models with Deduplication from Relational Databases | 2022 | VLDB |
| 3 | 5,272 | InferDB: In-Database Machine Learning Inference Using Indexes | 2024 | VLDB |
| 4 | 560 | Plan-Structured Deep Neural Network Models for Query Performance Prediction | 2019 | VLDB |
| 5 | 318 | DeepDB: Learn from Data, not from Queries! | 2020 | VLDB |
| 6 | 5,158 | Facilitating SQL Query Composition and Analysis | 2020 | SIGMOD |
| 7 | 281 | Accelerating Machine Learning Inference with Probabilistic Predicates | 2018 | SIGMOD |
| 8 | 7,236 | RCRank: Multimodal Ranking of Root Causes of Slow Queries in Cloud Database Systems | 2025 | VLDB |
| 9 | 5,247 | Enabling SQL-based Training Data Debugging for Federated Learning | 2022 | VLDB |
| 10 | 2,598 | Complaint-driven Training Data Debugging for Query 2.0 | 2020 | SIGMOD |