Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise
Summary: Studies selective data cleaning for fact-checking, contrasting uncertainty minimization with maximizing counterargument probability—and showing their divergence can induce bias. Develops efficient algorithms for complex nonlinear objectives, generalizable beyond fact-checking. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Stavros Sintos (Duke University)
- 2. Pankaj K. Agarwal (Duke University)
- 3. Jun Yang (Duke University)
BibTeX Citation
@article{sintos_vldb19,
title = {{Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise}},
author = {Sintos, Stavros and Agarwal, Pankaj K. and Yang, Jun},
journal = {PVLDB},
series = {{VLDB} '19},
volume = {12},
number = {13},
pages = {2408--2421},
doi = {10.14778/3358701.3358708},
url = {https://doi.org/10.14778/3358701.3358708},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,663 | To Intervene or Not To Intervene: Cost based Intervention for Combating Fake News | 2021 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 112 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00032801121 |
| 557 | The Theory of Probabilistic Databases | 1987 | VLDB | 0.00016541752 |
| 582 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB | 0.00016148948 |
| 725 | NADEEF: A Commodity Data Cleaning System | 2013 | SIGMOD | 0.00014617251 |
| 890 | Integrating Conflicting Data: The Role of Source Dependence | 2009 | VLDB | 0.00013382697 |
| 1,736 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD | 9.8984415e-05 |
| 3,127 | Toward Computational Fact-Checking | 2014 | VLDB | 7.7308958e-05 |
| 3,281 | Cleaning Crowdsourced Labels Using Oracles for Statistical Classification | 2019 | VLDB | 7.5706653e-05 |
| 5,318 | Cleaning Uncertain Data with Quality Guarantees | 2008 | VLDB | 6.2672687e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,628 | User Guidance for Efficient Fact Checking | 2019 | VLDB |
| 2 | 10,979 | Finding Convincing Views to Endorse a Claim | 2025 | VLDB |
| 3 | 1,736 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD |
| 4 | 1,914 | Statistical Distortion: Consequences of Data Cleaning | 2012 | VLDB |
| 5 | 5,548 | Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond | 2020 | VLDB |
| 6 | 5,318 | Cleaning Uncertain Data with Quality Guarantees | 2008 | VLDB |
| 7 | 9,193 | Query-Guided Resolution in Uncertain Databases | 2023 | SIGMOD |
| 8 | 1,323 | Data Cleaning: Overview and Emerging Challenges | 2016 | SIGMOD |
| 9 | 7,125 | Computational Fact Checking: A Content Management Perspective | 2018 | VLDB |
| 10 | 3,127 | Toward Computational Fact-Checking | 2014 | VLDB |