DBScholar

Back to papers

Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction

Summary: Faerie provides a unified framework for approximate dictionary-based entity extraction, supporting diverse similarity/dissimilarity measures. It uses overlap-aware filtering and pruning to share work across substrings, achieving top performance. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
4472
Venue
SIGMOD
Year
2011
Pagerank
6.2272563e-05
Overall Rank
5,412 | 62.87%
DOI
10.1145/1989323.1989379

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{li_sigmod11,
        title = {{Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction}},
        author = {Li, Guoliang and Deng, Dong and Feng, Jianhua},
        series = {{SIGMOD} '11},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/1989323.1989379},
        url = {https://dl.acm.org/doi/10.1145/1989323.1989379},
        year = {2011}
}

Incoming Citations (Sorted by Pagerank)

Showing 11 of 11 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
107 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.00033511706
158 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00028199923
169 Efficient Exact Set-Similarity Joins 2006 VLDB 0.0002743469
200 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025597287
911 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013283031
1,040 VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams 2007 VLDB 0.00012466499
1,636 Relaxing Join and Selection Queries 2006 VLDB 0.00010156479
2,262 n-Gram/2L: A Space and Time Efficient Two-Level n-Gram Inverted Index Structure 2005 VLDB 8.8440146e-05
2,308 Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries 2008 VLDB 8.7738996e-05
2,899 Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance 2007 VLDB 7.97814e-05
3,442 An Efficient Filter for Approximate Membership Checking 2008 SIGMOD 7.4122197e-05
3,446 Efficient Approximate Entity Extraction with Edit Distance Constraints 2009 SIGMOD 7.4087786e-05
4,359 Incremental Maintenance of Length Normalized Indexes for Approximate String Matching 2009 SIGMOD 6.7451753e-05
4,450 Power-Law Based Estimation of Set Similarity Join Size 2009 VLDB 6.6972929e-05
4,992 Scalable Ad-hoc Entity Extraction from Text Collections 2008 VLDB 6.4100657e-05
5,279 Mining Document Collections to Facilitate Accurate Approximate Entity Matching 2009 VLDB 6.2849137e-05
Previous Page 1 / 1 Next

Semantically Similar Papers