DBScholar

Back to papers

Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams

Summary: Hydra enables real-time, general subpopulation analytics on multidimensional streams with a 'sketch of sketches' and universal sketching to bound errors across combinatorial subpopulations. Spark plugin implementation minimizes overhead and memory, delivering interactive estimates with order-of-magnitude gains versus Spark/Druid. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hff8d722904f38426
Venue
VLDB
Year
2022
Pagerank
5.474978e-05
Overall Rank
7,701 | 48.23%
DOI
10.14778/3551793.3551867

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{manousis_vldb22,
        title = {{Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams}},
        author = {Manousis, Antonis and Cheng, Zhuo and Basat, Ran Ben and Liu, Zaoxing and Sekar, Vyas},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {11},
        pages = {3249--3262},
        doi = {10.14778/3551793.3551867},
        url = {https://doi.org/10.14778/3551793.3551867},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 40 of 40 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.001052036
9 Online Aggregation 1997 SIGMOD 0.00076195956
11 Implementing Data Cubes Efficiently 1996 SIGMOD 0.00071084324
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055406774
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.00043160717
83 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00035978046
127 The Design of the Borealis Stream Processing Engine 2005 CIDR 0.00030427614
148 Gorilla: A Fast, Scalable, In-Memory Time Series Database 2015 VLDB 0.0002900671
218 MillWheel: Fault-Tolerant Stream Processing at Internet Scale 2013 VLDB 0.00024390324
222 Approximate Computation of Multidimensional Aggregates of Sparse Data Using Wavelets 1999 SIGMOD 0.00024218831
240 Gigascope: A Stream Database for Network Applications 2003 SIGMOD 0.00023435436
327 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.0002095191
336 The Aqua Approximate Query Answering System 1999 SIGMOD 0.00020657819
395 SeeDB: Efficient Data-Driven Visualization Recommendations to Support Visual Analytics 2015 VLDB 0.00019165452
456 Mergeable Summaries 2012 PODS 0.0001791284
564 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016296665
596 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00015785583
784 VerdictDB: Universalizing Approximate Query Processing 2018 SIGMOD 0.00014012614
1,022 Online Aggregation for Large MapReduce Jobs 2011 VLDB 0.00012438826
1,082 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012122749
1,236 Druid: A Real-time Analytical Data Store 2014 SIGMOD 0.00011402848
1,381 Scuba: Diving into Data at Facebook 2013 VLDB 0.00010860462
1,607 SciBORQ: Scientific data management with Bounds On Runtime and Quality 2011 CIDR 0.0001008742
1,828 Fault-Tolerance in the Borealis Distributed Stream Processing System 2005 SIGMOD 9.5516021e-05
1,829 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 9.5510333e-05
1,833 MacroBase: Prioritizing Attention in Fast Data 2017 SIGMOD 9.5405247e-05
1,874 Efficient Computation of Iceberg Cubes with Complex Measures 2001 SIGMOD 9.4579757e-05
2,798 Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries 2018 VLDB 7.9933329e-05
3,080 Persistent Data Sketching 2015 SIGMOD 7.6662346e-05
3,122 Analytics in Motion: High Performance Event-Processing AND Real-Time Analytics in the Same Database 2015 SIGMOD 7.6236292e-05
3,460 Quotient Cube: How to Summarize the Semantics of a Data Cube 2002 VLDB 7.2845441e-05
3,560 Interactive Analysis of Web-Scale Data 2009 CIDR 7.2072209e-05
3,818 Spatial Online Sampling and Aggregation 2016 VLDB 7.0060535e-05
3,841 High-Dimensional OLAP: A Minimal Cubing Approach 2004 VLDB 6.9904505e-05
5,652 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.0515502e-05
6,122 Approximate Distinct Counts for Billions of Datasets 2019 SIGMOD 5.879545e-05
6,513 Geospatial Stream Query Processing using Microsoft SQL Server StreamInsight 2010 VLDB 5.7596327e-05
8,776 CoopStore: Optimizing Precomputed Summaries for Aggregation 2020 VLDB 5.2800094e-05
Previous Page 1 / 1 Next

Semantically Similar Papers