DBScholar

Back to papers

Building Wavelet Histograms on Large Data in MapReduce

Summary: Proposes exact and approximate wavelet histogram algorithms for MapReduce, cutting communication and runtime vs naive approaches. Implemented in Hadoop and evaluated on a 16-node cluster with real and synthetic data, showing large improvements. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
10536
Venue
VLDB
Year
2012
Pagerank
6.1871697e-05
Overall Rank
5,522 | 62.12%
DOI
10.14778/2078324.2078327

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{jestes_vldb12,
        title = {{Building Wavelet Histograms on Large Data in MapReduce}},
        author = {Jestes, Jeffrey and Yi, Ke and Li, Feifei},
        journal = {PVLDB},
        series = {{VLDB} '12},
        volume = {5},
        number = {2},
        pages = {109--120},
        doi = {10.14778/2078324.2078327},
        url = {https://doi.org/10.14778/2078324.2078327},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 7 of 7 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 22 of 22 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
5 Optimal Aggregation Algorithms for Middleware [Extended Abstract] 2001 PODS 0.0010828372
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010686205
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00051174276
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
35 Improved Histograms for Selectivity Estimation of Range Predicates 1996 SIGMOD 0.00048481081
44 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00046055057
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.00031680027
155 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028713176
168 Wavelet-Based Histograms for Selectivity Estimation 1998 SIGMOD 0.00027541029
267 Optimal Histograms with Quality Guarantees 1998 VLDB 0.00022798161
307 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021792475
311 Surfing Wavelets on Streams: One-Pass Summaries for Approximate Aggregate Queries 2001 VLDB 0.00021760621
356 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020303289
642 Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience 2009 VLDB 0.00015395331
660 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00015198804
885 Finding Frequent Items in Data Streams 2008 VLDB 0.00013419017
977 Dynamic Maintenance of Wavelet-Based Histograms 2000 VLDB 0.00012864017
1,156 Wavelet Synopses with Error Guarantees 2002 SIGMOD 0.00011929041
1,436 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010797443
2,312 Online Aggregation and Continuous Query support in MapReduce 2010 SIGMOD 8.7642158e-05
2,813 KLEE: A Framework for Distributed Top-k Query Algorithms 2005 VLDB 8.0975254e-05
6,060 Finding Global Icebergs over Distributed Data Sets 2006 PODS 5.9914746e-05
Previous Page 1 / 1 Next

Semantically Similar Papers