DBScholar

Back to papers

Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams

Summary: Hydra enables real-time, general subpopulation analytics on multidimensional streams with a 'sketch of sketches' and universal sketching to bound errors across combinatorial subpopulations. Spark plugin implementation minimizes overhead and memory, delivering interactive estimates with order-of-magnitude gains versus Spark/Druid. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12993
Venue
VLDB
Year
2022
Pagerank
5.6006414e-05
Overall Rank
7,551 | 48.20%
DOI
10.14778/3551793.3551867

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{manousis_vldb22,
        title = {{Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams}},
        author = {Manousis, Antonis and Cheng, Zhuo and Basat, Ran Ben and Liu, Zaoxing and Sekar, Vyas},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {11},
        pages = {3249--3262},
        doi = {10.14778/3551793.3551867},
        url = {https://doi.org/10.14778/3551793.3551867},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 40 of 40 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010686205
9 Online Aggregation 1997 SIGMOD 0.00077458002
11 Implementing Data Cubes Efficiently 1996 SIGMOD 0.00071822821
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00051174276
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
51 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004291425
82 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00036378991
127 The Design of the Borealis Stream Processing Engine 2005 CIDR 0.00030738755
148 Gorilla: A Fast, Scalable, In-Memory Time Series Database 2015 VLDB 0.00029250767
213 Approximate Computation of Multidimensional Aggregates of Sparse Data Using Wavelets 1999 SIGMOD 0.00024723025
224 MillWheel: Fault-Tolerant Stream Processing at Internet Scale 2013 VLDB 0.00024130894
230 Gigascope: A Stream Database for Network Applications 2003 SIGMOD 0.00023891474
327 The Aqua Approximate Query Answering System 1999 SIGMOD 0.00021091539
330 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.0002104801
410 SeeDB: Efficient Data-Driven Visualization Recommendations to Support Visual Analytics 2015 VLDB 0.0001890421
451 Mergeable Summaries 2012 PODS 0.00018151445
553 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016590619
593 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00016027871
772 VerdictDB: Universalizing Approximate Query Processing 2018 SIGMOD 0.00014147905
1,009 Online Aggregation for Large MapReduce Jobs 2011 VLDB 0.00012684342
1,108 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012145154
1,242 Druid: A Real-time Analytical Data Store 2014 SIGMOD 0.00011516162
1,352 Scuba: Diving into Data at Facebook 2013 VLDB 0.00011064595
1,582 SciBORQ: Scientific data management with Bounds On Runtime and Quality 2011 CIDR 0.00010295367
1,789 Fault-Tolerance in the Borealis Distributed Stream Processing System 2005 SIGMOD 9.7502324e-05
1,792 MacroBase: Prioritizing Attention in Fast Data 2017 SIGMOD 9.7436856e-05
1,799 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 9.7326398e-05
1,825 Efficient Computation of Iceberg Cubes with Complex Measures 2001 SIGMOD 9.6726613e-05
2,747 Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries 2018 VLDB 8.1711208e-05
3,023 Persistent Data Sketching 2015 SIGMOD 7.8398269e-05
3,082 Analytics in Motion: High Performance Event-Processing AND Real-Time Analytics in the Same Database 2015 SIGMOD 7.7725834e-05
3,407 Quotient Cube: How to Summarize the Semantics of a Data Cube 2002 VLDB 7.4391131e-05
3,497 Interactive Analysis of Web-Scale Data 2009 CIDR 7.3634378e-05
3,741 Spatial Online Sampling and Aggregation 2016 VLDB 7.1586403e-05
3,753 High-Dimensional OLAP: A Minimal Cubing Approach 2004 VLDB 7.1505817e-05
5,522 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.1871697e-05
5,998 Approximate Distinct Counts for Billions of Datasets 2019 SIGMOD 6.0142553e-05
6,384 Geospatial Stream Query Processing using Microsoft SQL Server StreamInsight 2010 VLDB 5.8909379e-05
8,616 CoopStore: Optimizing Precomputed Summaries for Aggregation 2020 VLDB 5.4004741e-05
Previous Page 1 / 1 Next

Semantically Similar Papers