Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise
Summary: Studies selective data cleaning for fact-checking, contrasting uncertainty minimization with maximizing counterargument probability—and showing their divergence can induce bias. Develops efficient algorithms for complex nonlinear objectives, generalizable beyond fact-checking. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Stavros Sintos (Duke University)
- 2. Pankaj K. Agarwal (Duke University)
- 3. Jun Yang (Duke University)
BibTeX Citation
@article{sintos_vldb19,
title = {{Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise}},
author = {Sintos, Stavros and Agarwal, Pankaj K. and Yang, Jun},
journal = {PVLDB},
series = {{VLDB} '19},
volume = {12},
number = {13},
pages = {2408--2421},
doi = {10.14778/3358701.3358708},
url = {https://doi.org/10.14778/3358701.3358708},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,970 | To Intervene or Not To Intervene: Cost based Intervention for Combating Fake News | 2021 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 104 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00033690989 |
| 483 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB | 0.00017590977 |
| 571 | The Theory of Probabilistic Databases | 1987 | VLDB | 0.00016210743 |
| 697 | NADEEF: A Commodity Data Cleaning System | 2013 | SIGMOD | 0.00014694048 |
| 877 | Integrating Conflicting Data: The Role of Source Dependence | 2009 | VLDB | 0.00013306717 |
| 1,720 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD | 9.7965659e-05 |
| 3,178 | Toward Computational Fact-Checking | 2014 | VLDB | 7.5655186e-05 |
| 3,307 | Cleaning Crowdsourced Labels Using Oracles for Statistical Classification | 2019 | VLDB | 7.4414303e-05 |
| 5,175 | Cleaning Uncertain Data with Quality Guarantees | 2008 | VLDB | 6.2429707e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,778 | User Guidance for Efficient Fact Checking | 2019 | VLDB |
| 2 | 11,362 | Finding Convincing Views to Endorse a Claim | 2025 | VLDB |
| 3 | 1,720 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD |
| 4 | 1,914 | Statistical Distortion: Consequences of Data Cleaning | 2012 | VLDB |
| 5 | 5,679 | Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond | 2020 | VLDB |
| 6 | 5,175 | Cleaning Uncertain Data with Quality Guarantees | 2008 | VLDB |
| 7 | 9,373 | Query-Guided Resolution in Uncertain Databases | 2023 | SIGMOD |
| 8 | 1,043 | Data Cleaning: Overview and Emerging Challenges | 2016 | SIGMOD |
| 9 | 7,274 | Computational Fact Checking: A Content Management Perspective | 2018 | VLDB |
| 10 | 3,178 | Toward Computational Fact-Checking | 2014 | VLDB |