Database Paper Browser

Back to papers

Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams

Summary: Hydra enables real-time, general subpopulation analytics on multidimensional streams with a 'sketch of sketches' and universal sketching to bound errors across combinatorial subpopulations. Spark plugin implementation minimizes overhead and memory, delivering interactive estimates with order-of-magnitude gains versus Spark/Druid. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12806
Venue
VLDB
Year
2022
Pagerank
4.7134753e-05
Overall Rank
7,533 | 47.65%
DOI
10.14778/3551793.3551867

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 40 of 40 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
3 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0024217964
11 Implementing Data Cubes Efficiently 1996 SIGMOD 0.0011695087
14 Online Aggregation 1997 SIGMOD 0.0010813443
22 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00084679526
66 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00061707583
70 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00059744625
109 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.00048217028
126 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00044753012
191 The Design of the Borealis Stream Processing Engine 2005 CIDR 0.00035714897
211 Gorilla: A Fast, Scalable, In-Memory Time Series Database 2015 VLDB 0.0003401421
275 Approximate Computation of Multidimensional Aggregates of Sparse Data Using Wavelets 1999 SIGMOD 0.00029381206
314 MillWheel: Fault-Tolerant Stream Processing at Internet Scale 2013 VLDB 0.00028059664
324 Gigascope: A Stream Database for Network Applications 2003 SIGMOD 0.00027465124
398 Mergeable Summaries 2012 PODS 0.00024383201
431 The Aqua Approximate Query Answering System 1999 SIGMOD 0.00023397171
461 SeeDB: Efficient Data-Driven Visualization Recommendations to Support Visual Analytics 2015 VLDB 0.00022615628
476 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.00022216002
736 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00017414831
941 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00015147831
1,161 VerdictDB: Universalizing Approximate Query Processing 2018 SIGMOD 0.00013579831
1,451 Online Aggregation for Large MapReduce Jobs 2011 VLDB 0.00011925842
1,488 Scuba: Diving into Data at Facebook 2013 VLDB 0.00011690191
1,574 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00011289028
1,588 Druid: A Real-time Analytical Data Store 2014 SIGMOD 0.00011232949
1,735 SciBORQ: Scientific data management with Bounds On Runtime and Quality 2011 CIDR 0.00010713691
1,967 Efficient Computation of Iceberg Cubes with Complex Measures 2001 SIGMOD 9.9108189e-05
1,995 Fault-Tolerance in the Borealis Distributed Stream Processing System 2005 SIGMOD 9.83817e-05
2,129 MacroBase: Prioritizing Attention in Fast Data 2017 SIGMOD 9.4799835e-05
2,494 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 8.6457436e-05
2,954 Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries 2018 VLDB 7.8218804e-05
3,162 High-Dimensional OLAP: A Minimal Cubing Approach 2004 VLDB 7.4622284e-05
3,387 Analytics in Motion: High Performance Event-Processing AND Real-Time Analytics in the Same Database 2015 SIGMOD 7.1505623e-05
3,585 Quotient Cube: How to Summarize the Semantics of a Data Cube 2002 VLDB 6.9408892e-05
3,618 Persistent Data Sketching 2015 SIGMOD 6.9080647e-05
4,032 Spatial Online Sampling and Aggregation 2016 VLDB 6.5131946e-05
4,037 Interactive Analysis of Web-Scale Data 2009 CIDR 6.5076279e-05
5,908 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 5.2731311e-05
6,243 Approximate Distinct Counts for Billions of Datasets 2019 SIGMOD 5.1348218e-05
6,652 Geospatial Stream Query Processing using Microsoft SQL Server StreamInsight 2010 VLDB 4.9710292e-05
8,669 CoopStore: Optimizing Precomputed Summaries for Aggregation 2020 VLDB 4.4667395e-05
Previous Page 1 / 1 Next

Semantically Similar Papers