DBScholar

Back to papers

ICARUS: Minimizing Human Effort in Iterative Data Completion

Summary: ICARUS reduces expert labor by presenting small, high-impact matrix subsets for edits to the matrix. Schema-informed hierarchies amplify edits into rules; heuristic subset selection yields ~50% improvement, with users filling 68% in an hour. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
11925
Venue
VLDB
Year
2018
Pagerank
5.4754476e-05
Overall Rank
8,160 | 44.02%
DOI
10.14778/3275366.3275374

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{rahman_vldb18,
        title = {{ICARUS: Minimizing Human Effort in Iterative Data Completion}},
        author = {Rahman, Protiva and Hebert, Courtney and Nandi, Arnab},
        journal = {PVLDB},
        series = {{VLDB} '18},
        volume = {11},
        number = {13},
        pages = {2263--2276},
        doi = {10.14778/3275366.3275374},
        url = {https://doi.org/10.14778/3275366.3275374},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
5,722 Adaptive Rule Discovery for Labeling Text Data 2021 SIGMOD 6.1089867e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
112 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00032801121
159 CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies 2004 SIGMOD 0.00028129426
533 Improving Data Quality: Consistency and Accuracy 2007 VLDB 0.0001705859
549 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016692839
647 Discovering Data Quality Rules 2008 VLDB 0.00015334666
714 Guided Data Repair 2011 VLDB 0.00014662041
725 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014617251
1,033 On Generating Near-Optimal Tableaux for Conditional Functional Dependencies 2008 VLDB 0.00012529852
1,101 KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing 2015 SIGMOD 0.00012168934
1,736 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.8984415e-05
2,570 Query-Oriented Data Cleaning with Oracles 2015 SIGMOD 8.4061164e-05
2,636 CrowdFill: Collecting Structured Data from the Crowd 2014 SIGMOD 8.31868e-05
3,147 Auto-Detect: Data-Driven Error Detection in Tables 2018 SIGMOD 7.7077175e-05
3,380 UGuide – User-Guided Discovery of FD-Detectable Errors 2017 SIGMOD 7.457539e-05
3,547 Learning Semantic String Transformations from Examples 2012 VLDB 7.3225973e-05
5,253 Rudolf: Interactive Rule Refinement System for Fraud Detection 2016 VLDB 6.2989009e-05
6,059 Evaluating Interactive Data Systems: Workloads, Metrics, and Guidelines 2018 SIGMOD 5.9915079e-05
8,399 DataProf: Semantic Profiling for Iterative Data Cleansing and Business Rule Acquisition 2018 SIGMOD 5.4333402e-05
Previous Page 1 / 1 Next

Semantically Similar Papers