DBScholar

Back to papers

Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams

Summary: Hydra enables real-time, general subpopulation analytics on multidimensional streams with a 'sketch of sketches' and universal sketching to bound errors across combinatorial subpopulations. Spark plugin implementation minimizes overhead and memory, delivering interactive estimates with order-of-magnitude gains versus Spark/Druid. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hff8d722904f38426
Venue
VLDB
Year
2022
Pagerank
5.4723863e-05
Overall Rank
7,707 | 48.21%
DOI
10.14778/3551793.3551867
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{manousis_vldb22,
        title = {{Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams}},
        author = {Manousis, Antonis and Cheng, Zhuo and Basat, Ran Ben and Liu, Zaoxing and Sekar, Vyas},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {11},
        pages = {3249--3262},
        doi = {10.14778/3551793.3551867},
        url = {https://doi.org/10.14778/3551793.3551867},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 40 of 40 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010515896
9 Online Aggregation 1997 SIGMOD 0.00076265429
11 Implementing Data Cubes Efficiently 1996 SIGMOD 0.00071056708
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055384955
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050475202
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049821554
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004314366
83 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00035962466
127 The Design of the Borealis Stream Processing Engine 2005 CIDR 0.00030414379
148 Gorilla: A Fast, Scalable, In-Memory Time Series Database 2015 VLDB 0.00029001141
218 MillWheel: Fault-Tolerant Stream Processing at Internet Scale 2013 VLDB 0.00024379041
222 Approximate Computation of Multidimensional Aggregates of Sparse Data Using Wavelets 1999 SIGMOD 0.00024210103
240 Gigascope: A Stream Database for Network Applications 2003 SIGMOD 0.00023425462
327 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.00020942751
335 The Aqua Approximate Query Answering System 1999 SIGMOD 0.000206533
395 SeeDB: Efficient Data-Driven Visualization Recommendations to Support Visual Analytics 2015 VLDB 0.00019156481
456 Mergeable Summaries 2012 PODS 0.00017904764
564 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016297598
596 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00015782051
772 VerdictDB: Universalizing Approximate Query Processing 2018 SIGMOD 0.0001409096
1,021 Online Aggregation for Large MapReduce Jobs 2011 VLDB 0.00012437619
1,061 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012208639
1,238 Druid: A Real-time Analytical Data Store 2014 SIGMOD 0.00011397949
1,381 Scuba: Diving into Data at Facebook 2013 VLDB 0.00010856731
1,608 SciBORQ: Scientific data management with Bounds On Runtime and Quality 2011 CIDR 0.00010085907
1,828 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 9.547768e-05
1,829 Fault-Tolerance in the Borealis Distributed Stream Processing System 2005 SIGMOD 9.5470848e-05
1,833 MacroBase: Prioritizing Attention in Fast Data 2017 SIGMOD 9.5376219e-05
1,875 Efficient Computation of Iceberg Cubes with Complex Measures 2001 SIGMOD 9.4535173e-05
2,799 Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries 2018 VLDB 7.9895512e-05
3,082 Persistent Data Sketching 2015 SIGMOD 7.6626082e-05
3,124 Analytics in Motion: High Performance Event-Processing AND Real-Time Analytics in the Same Database 2015 SIGMOD 7.6222513e-05
3,460 Quotient Cube: How to Summarize the Semantics of a Data Cube 2002 VLDB 7.2811286e-05
3,553 Interactive Analysis of Web-Scale Data 2009 CIDR 7.20663e-05
3,819 Spatial Online Sampling and Aggregation 2016 VLDB 7.0027383e-05
3,842 High-Dimensional OLAP: A Minimal Cubing Approach 2004 VLDB 6.9871413e-05
5,653 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.048773e-05
6,124 Approximate Distinct Counts for Billions of Datasets 2019 SIGMOD 5.8769926e-05
6,515 Geospatial Stream Query Processing using Microsoft SQL Server StreamInsight 2010 VLDB 5.7569075e-05
8,783 CoopStore: Optimizing Precomputed Summaries for Aggregation 2020 VLDB 5.2776273e-05
Previous Page 1 / 1 Next

Semantically Similar Papers