DBScholar

Back to papers

Building Wavelet Histograms on Large Data in MapReduce

Summary: Proposes exact and approximate wavelet histogram algorithms for MapReduce, cutting communication and runtime vs naive approaches. Implemented in Hadoop and evaluated on a 16-node cluster with real and synthetic data, showing large improvements. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h32436c103450cb9a
Venue
VLDB
Year
2012
Pagerank
6.0515502e-05
Overall Rank
5,652 | 62.00%
DOI
10.14778/2078324.2078327

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{jestes_vldb12,
        title = {{Building Wavelet Histograms on Large Data in MapReduce}},
        author = {Jestes, Jeffrey and Yi, Ke and Li, Feifei},
        journal = {PVLDB},
        series = {{VLDB} '12},
        volume = {5},
        number = {2},
        pages = {109--120},
        doi = {10.14778/2078324.2078327},
        url = {https://doi.org/10.14778/2078324.2078327},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 7 of 7 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 22 of 22 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
5 Optimal Aggregation Algorithms for Middleware [Extended Abstract] 2001 PODS 0.0010679641
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.001052036
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
37 Improved Histograms for Selectivity Estimation of Range Predicates 1996 SIGMOD 0.00047731453
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.000311132
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028579704
169 Wavelet-Based Histograms for Selectivity Estimation 1998 SIGMOD 0.00027134723
275 Optimal Histograms with Quality Guarantees 1998 VLDB 0.00022413521
309 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021384073
312 Surfing Wavelets on Streams: One-Pass Summaries for Approximate Aggregate Queries 2001 VLDB 0.00021311793
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020009936
651 Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience 2009 VLDB 0.00015130782
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
909 Finding Frequent Items in Data Streams 2008 VLDB 0.00013125647
1,004 Dynamic Maintenance of Wavelet-Based Histograms 2000 VLDB 0.00012594522
1,173 Wavelet Synopses with Error Guarantees 2002 SIGMOD 0.00011686985
1,466 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010569837
2,361 Online Aggregation and Continuous Query support in MapReduce 2010 SIGMOD 8.5761274e-05
2,874 KLEE: A Framework for Distributed Top-k Query Algorithms 2005 VLDB 7.9221111e-05
6,188 Finding Global Icebergs over Distributed Data Sets 2006 PODS 5.8571256e-05
Previous Page 1 / 1 Next

Semantically Similar Papers