Back to papers
Fainder: A Fast and Accurate Index for Distribution-Aware Dataset Search
Summary: Fainder introduces a distribution-aware index for percentile predicates over heterogeneous histogram summaries, enabling dataset discovery based on distributional properties rather than keywords. It uses binary search plus multi-step pruning on summary bounds to prune candidates and yields order-of-magnitude speedups.
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 13541
- Venue
- VLDB
- Year
- 2024
- Pagerank
- 4.2470891e-05
- Overall Rank
- 9,928 | 31.01%
- DOI
-
10.14778/3681954.3681999
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 18 of 18 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 326 |
Optimal Histograms with Quality Guarantees |
1998 |
VLDB |
0.0002737538 |
| 609 |
Goods: Organizing Google's Datasets |
2016 |
SIGMOD |
0.00019223217 |
| 1,403 |
Detecting Data Errors: Where are we and what needs to be done? |
2016 |
VLDB |
0.00012180046 |
| 1,643 |
Finding Related Tables in Data Lakes for Interactive Data Science |
2020 |
SIGMOD |
0.00011031534 |
| 1,742 |
Auctus: A Dataset Search Engine for Data Discovery and Augmentation |
2021 |
VLDB |
0.00010695388 |
| 3,360 |
Organizing Data Lakes for Navigation |
2020 |
SIGMOD |
7.1719486e-05 |
| 3,520 |
GitTables: A Large-Scale Corpus of Relational Tables |
2023 |
SIGMOD |
7.0136102e-05 |
| 5,022 |
Towards Distribution-aware Query Answering in Data Markets |
2022 |
VLDB |
5.7479778e-05 |
| 5,386 |
Selective Data Acquisition in the Wild for Model Charging |
2022 |
VLDB |
5.5346315e-05 |
| 5,805 |
Discovering Related Data At Scale |
2021 |
VLDB |
5.3197639e-05 |
| 6,268 |
MATE: Multi-Attribute Table Extraction |
2022 |
VLDB |
5.1288179e-05 |
| 6,432 |
RONIN: Data Lake Exploration |
2021 |
VLDB |
5.0571585e-05 |
| 6,462 |
Tailoring Data Source Distributions for Fairness-aware Data Integration |
2021 |
VLDB |
5.0479645e-05 |
| 6,948 |
DataPrism: Exposing Disconnect between Data and Systems |
2022 |
SIGMOD |
4.8865863e-05 |
| 7,276 |
DICE: Data Discovery by Example |
2021 |
VLDB |
4.773228e-05 |
| 7,853 |
Consistent Range Approximation for Fair Predictive Modeling |
2023 |
VLDB |
4.6308623e-05 |
| 7,869 |
Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach |
2023 |
SIGMOD |
4.6275089e-05 |
| 8,616 |
Nexus: Correlation Discovery over Collections of Spatio-Temporal Tabular Data |
2024 |
SIGMOD |
4.4795277e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 1,805 |
Top-k Query Evaluation with Probabilistic Guarantees |
2004 |
VLDB |
0.00010479371 |
| 11,602 |
IDAR: Fast Supergraph Search Using DAG Integration |
2020 |
VLDB |
4.1905499e-05 |
| 7,916 |
HINT: A Hierarchical Index for Intervals in Main Memory |
2022 |
SIGMOD |
4.6133471e-05 |
| 3,136 |
FINEdex: A Fine-grained Learned Index Scheme for Scalable and Concurrent Memory Systems |
2022 |
VLDB |
7.4926368e-05 |
| 9,180 |
RDFind: Scalable Conditional Inclusion Dependency Discovery in RDF Datasets |
2016 |
SIGMOD |
4.3793465e-05 |
| 10,963 |
FairHash: A Fair and Memory/Time-efficient Hashmap |
2024 |
SIGMOD |
4.1905499e-05 |
| 11,253 |
Fast Search-By-Classification for Large-Scale Databases Using Index-Aware Decision Trees and Random Forests |
2023 |
VLDB |
4.1905499e-05 |
| 11,381 |
Fast Dataset Search with Earth Mover’s Distance |
2022 |
VLDB |
4.1905499e-05 |
| 10,449 |
Finding What You’re Looking For: A Distribution-Aware Dataset Search Engine in Action |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,353 |
A Theoretical Framework for Distribution-Aware Dataset Search |
2025 |
PODS |
4.1905499e-05 |