DBScholar

Back to papers

Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction

Summary: Faerie provides a unified framework for approximate dictionary-based entity extraction, supporting diverse similarity/dissimilarity measures. It uses overlap-aware filtering and pruning to share work across substrings, achieving top performance. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h08bd159a71b60486
Venue
SIGMOD
Year
2011
Pagerank
6.088104e-05
Overall Rank
5,549 | 62.70%
DOI
10.1145/1989323.1989379

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{li_sigmod11,
        title = {{Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction}},
        author = {Li, Guoliang and Deng, Dong and Feng, Jianhua},
        series = {{SIGMOD} '11},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/1989323.1989379},
        url = {https://dl.acm.org/doi/10.1145/1989323.1989379},
        year = {2011}
}

Incoming Citations (Sorted by Pagerank)

Showing 11 of 11 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
108 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.0003305531
161 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00027718195
168 Efficient Exact Set-Similarity Joins 2006 VLDB 0.00027163517
201 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025331535
929 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013020115
1,061 VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams 2007 VLDB 0.00012213729
1,643 Relaxing Join and Selection Queries 2006 VLDB 0.00010006399
2,313 n-Gram/2L: A Space and Time Efficient Two-Level n-Gram Inverted Index Structure 2005 VLDB 8.6550779e-05
2,354 Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries 2008 VLDB 8.5896515e-05
2,967 Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance 2007 VLDB 7.8029216e-05
3,501 An Efficient Filter for Approximate Membership Checking 2008 SIGMOD 7.2531528e-05
3,512 Efficient Approximate Entity Extraction with Edit Distance Constraints 2009 SIGMOD 7.2439455e-05
4,439 Incremental Maintenance of Length Normalized Indexes for Approximate String Matching 2009 SIGMOD 6.5998035e-05
4,536 Power-Law Based Estimation of Set Similarity Join Size 2009 VLDB 6.5557137e-05
4,736 Scalable Ad-hoc Entity Extraction from Text Collections 2008 VLDB 6.4451961e-05
5,403 Mining Document Collections to Facilitate Accurate Approximate Entity Matching 2009 VLDB 6.1459021e-05
Previous Page 1 / 1 Next

Semantically Similar Papers