DBScholar

Back to papers

Human-in-the-loop Data Integration

Summary: Hybrid human–machine data integration for entity matching: learned similarity/knowledge rules and distributed in-memory candidate generation (DIMA) precede crowd verification. CDB selects, infers via transitivity/partial orders, and refines uncertain judgments through SQL-like crowd queries. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
11697
Venue
VLDB
Year
2017
Pagerank
5.6046905e-05
Overall Rank
7,493 | 48.60%
DOI
10.14778/3137765.3137833

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{li_vldb17,
        title = {{Human-in-the-loop Data Integration}},
        author = {Li, Guoliang},
        journal = {PVLDB},
        series = {{VLDB} '17},
        volume = {10},
        number = {12},
        pages = {2006--2009},
        doi = {10.14778/3137765.3137833},
        url = {https://doi.org/10.14778/3137765.3137833},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 9 of 9 citing papers.

Rank Citing Paper Year Venue Pagerank
498 QTune: A Query-Aware Database Tuning System with Deep Reinforcement Learning 2019 VLDB 0.00017440583
2,216 Open Data Integration 2018 VLDB 8.9374127e-05
2,888 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.9941489e-05
6,908 Cost-Effective Data Annotation using Game-Based Crowdsourcing 2019 VLDB 5.7415834e-05
7,743 Entity Resolution On-Demand 2022 VLDB 5.5545622e-05
9,381 Deduplicated Sampling On-Demand 2025 VLDB 5.2755515e-05
10,049 Towards Interpretable and Learnable Risk Analysis for Entity Resolution 2020 SIGMOD 5.1685424e-05
10,505 The Case For Language Model Approximated LIKE Predicate 2026 SIGMOD 5.093636e-05
11,912 A Rating-Ranking Method for Crowdsourced Top-k Computation 2018 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 43 of 43 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
90 CrowdDB: Answering Queries with Crowdsourcing 2011 SIGMOD 0.00034951786
169 Efficient Exact Set-Similarity Joins 2006 VLDB 0.0002743469
196 CrowdER: Crowdsourcing Entity Resolution 2012 VLDB 0.00025780596
200 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025597287
251 Crowdsourced Databases: Query Processing with People 2011 CIDR 0.00023261113
265 Human-powered Sorts and Joins 2012 VLDB 0.00022935368
356 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020303289
439 Corleone: Hands-Off Crowdsourcing for Entity Matching 2014 SIGMOD 0.00018464913
529 Magellan: Toward Building Entity Matching Management Systems 2016 VLDB 0.00017096361
743 So Who Won? Dynamic Max Discovery with the Crowd 2012 SIGMOD 0.00014421358
852 Leveraging Transitive Relations for Crowdsourced Joins 2013 SIGMOD 0.00013604253
911 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013283031
933 Question Selection for Crowd Entity Resolution 2013 VLDB 0.00013111293
975 Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search 2012 SIGMOD 0.00012870645
991 Bayesian Locality Sensitive Hashing for Fast Similarity Search 2012 VLDB 0.00012793339
997 CrowdScreen: Algorithms for Filtering Data with Humans 2012 SIGMOD 0.00012755983
1,248 Entity Matching: How Similar Is Similar 2011 VLDB 0.00011498301
1,293 Entity Resolution with Iterative Blocking 2009 SIGMOD 0.00011292804
1,415 V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors 2012 VLDB 0.00010840141
1,643 Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services 2017 SIGMOD 0.00010134956
1,886 Pass-Join: A Partition-based Method for Similarity Joins 2012 VLDB 9.5358137e-05
1,910 Counting with the Crowd 2013 VLDB 9.4972788e-05
1,935 ATLAS: A Probabilistic Algorithm for High Dimensional Similarity Search 2011 SIGMOD 9.4560124e-05
2,186 String Similarity Joins: An Experimental Evaluation 2014 VLDB 9.0001436e-05
2,307 Resolving Conflicts in Heterogeneous Data by Truth Discovery and Source Reliability Estimation 2014 SIGMOD 8.7750499e-05
2,409 Deco: A System for Declarative Crowdsourcing 2012 VLDB 8.6145078e-05
2,636 CrowdFill: Collecting Structured Data from the Crowd 2014 SIGMOD 8.31868e-05
3,231 Truth Inference in Crowdsourcing: Is the Problem Solved? 2017 VLDB 7.6191767e-05
3,235 Large-Scale Collective Entity Matching 2011 VLDB 7.6139368e-05
3,262 iCrowd: An Adaptive Crowdsourcing Framework 2015 SIGMOD 7.5846052e-05
3,474 An Efficient Partition Based Method for Exact Set Similarity Joins 2016 VLDB 7.3859271e-05
3,542 A Confidence-Aware Approach for Truth Discovery on Long-Tail Data 2015 VLDB 7.3279455e-05
3,665 BLAST: a Loosely Schema-aware Meta-blocking Approach for Entity Resolution 2016 VLDB 7.2142217e-05
3,904 QASCA: A Quality-Aware Task Assignment System for Crowdsourcing Applications 2015 SIGMOD 7.0304212e-05
3,983 Crowd-Based Deduplication: An Adaptive Approach 2015 SIGMOD 6.9739851e-05
4,538 Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach 2016 SIGMOD 6.6404776e-05
4,539 Reducing Uncertainty of Schema Matching via Crowdsourcing 2013 VLDB 6.6403822e-05
6,252 Putting Context into Schema Matching 2006 VLDB 5.941182e-05
6,290 Dima: A Distributed In-Memory Similarity-Based Query Processing System 2017 VLDB 5.9253163e-05
6,906 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.7418509e-05
7,575 Scalable Column Concept Determination for Web Tables Using Large Knowledge Bases 2013 VLDB 5.5937684e-05
9,401 CDB: Optimizing Queries with Crowd-Based Selections and Joins 2017 SIGMOD 5.2755515e-05
9,705 META: An Efficient Matching-Based Method for Error-Tolerant Autocompletion 2016 VLDB 5.2351259e-05
Previous Page 1 / 1 Next

Semantically Similar Papers