Back to papers
Parallel Rule Discovery from Large Datasets by Sampling
Summary: Parallel rule discovery for REEs across tables via multi-round sampling with alpha precision and beta recall guarantees. Deep Q-learning selects predicates for multi-variable rules; tableau boosts recall; parallelization yields 12.2x speedups at 10% sample.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 6476
- Venue
- SIGMOD
- Year
- 2022
- Pagerank
- 4.2254157e-05
- Overall Rank
- 9,962 | 30.77%
- DOI
-
10.1145/3514221.3526165
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 11 of 11 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 9,354 |
GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models |
2024 |
SIGMOD |
4.3484715e-05 |
| 9,362 |
Discovering Top-k Rules using Subjective and Objective Criteria |
2023 |
SIGMOD |
4.3472627e-05 |
| 9,439 |
Rock: Cleaning Data by Embedding ML in Logic Rules |
2024 |
SIGMOD |
4.3389137e-05 |
| 9,846 |
HyperBlocker: Accelerating Rule-based Blocking in Entity Resolution using GPUs |
2025 |
VLDB |
4.2680295e-05 |
| 9,847 |
Discovering Top-k Relevant and Diversified Rules |
2024 |
SIGMOD |
4.2680295e-05 |
| 10,029 |
Outliers: The Good, the Bad and the Ugly |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,499 |
Incremental Rule Discovery in Response to Parameter Updates |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,984 |
Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and Quality |
2024 |
SIGMOD |
4.1905499e-05 |
| 11,004 |
Capturing More Associations by Referencing External Graphs |
2024 |
VLDB |
4.1905499e-05 |
| 11,114 |
Rock: Cleaning Data with both ML and Logic Rules |
2024 |
VLDB |
4.1905499e-05 |
| 11,225 |
Splitting Tuples of Mismatched Entities |
2023 |
SIGMOD |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 49 |
Consistent Query Answers in Inconsistent Databases |
1999 |
PODS |
0.00067607389 |
| 192 |
HoloClean: Holistic Data Repairs with Probabilistic Inference |
2017 |
VLDB |
0.00035692958 |
| 219 |
Deep Entity Matching with Pre-Trained Language Models |
2021 |
VLDB |
0.00033354456 |
| 318 |
Evaluation of entity resolution approaches on real-world match problems |
2010 |
VLDB |
0.00027850417 |
| 473 |
Sampling Large Databases for Association Rules |
1996 |
VLDB |
0.00022304724 |
| 556 |
Discovering Denial Constraints |
2013 |
VLDB |
0.00020214701 |
| 890 |
A Hybrid Approach to Functional Dependency Discovery |
2016 |
SIGMOD |
0.00015542177 |
| 1,187 |
On Generating Near-Optimal Tableaux for Conditional Functional Dependencies |
2008 |
VLDB |
0.00013432048 |
| 1,340 |
HoloDetect: Few-Shot Learning for Error Detection |
2019 |
SIGMOD |
0.00012492795 |
| 1,821 |
Synthesizing Entity Matching Rules by Examples |
2018 |
VLDB |
0.00010406856 |
| 2,082 |
Efficient Discovery of Approximate Dependencies |
2018 |
VLDB |
9.5875364e-05 |
| 2,258 |
Efficient Denial Constraint Discovery with Hydra |
2018 |
VLDB |
9.1804145e-05 |
| 2,484 |
Discovery of Approximate (and Exact) Denial Constraints |
2020 |
VLDB |
8.6737275e-05 |
| 3,448 |
Approximate Denial Constraints |
2020 |
VLDB |
7.081095e-05 |
| 4,129 |
A Statistical Perspective on Discovering Functional Dependencies in Noisy Data |
2020 |
SIGMOD |
6.4208557e-05 |
| 5,203 |
Pattern Functional Dependencies for Data Cleaning |
2020 |
VLDB |
5.628087e-05 |
| 5,258 |
Error-bounded Sampling for Analytics on Big Sparse Data |
2014 |
VLDB |
5.5973455e-05 |
| 5,622 |
Distributed implementations of dependency discovery algorithms |
2019 |
VLDB |
5.4050344e-05 |
| 6,047 |
MDedup: Duplicate Detection with Matching Dependencies |
2020 |
VLDB |
5.2355891e-05 |
| 7,283 |
Discovering Association Rules from Big Graphs |
2022 |
VLDB |
4.7716465e-05 |
Semantically Similar Papers