DBScholar

Back to papers

Data Canopy: Accelerating Exploratory Statistical Analysis

Summary: Data Canopy provides an in-memory library of basic aggregates to reuse statistics across overlapping data parts, cutting recomputation in exploratory analysis. It decomposes stats into reusable aggregates, with storage/maintenance and hardware-aware tuning, yielding ~10x speedup after 100 queries vs. state-of-the-art. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
ha9d86a8dbe7fa5af
Venue
SIGMOD
Year
2017
Pagerank
6.8741188e-05
Overall Rank
3,984 | 73.22%
DOI
10.1145/3035918.3064051

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{wasay_sigmod17,
        title = {{Data Canopy: Accelerating Exploratory Statistical Analysis}},
        author = {Wasay, Abdul and Wei, Xinding and Dayan, Niv and Idreos, Stratos},
        series = {{SIGMOD} '17},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3035918.3064051},
        url = {https://dl.acm.org/doi/10.1145/3035918.3064051},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 9 of 9 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 41 of 41 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
27 Database Architecture Optimized for the New Bottleneck: Memory Access 1999 VLDB 0.0005158963
39 Efficiently Updating Materialized Views 1986 SIGMOD 0.00046602544
216 Small Materialized Aggregates: A Light Weight Index Structure for Data Warehousing 1998 VLDB 0.00024485024
242 Fast Incremental Maintenance of Approximate Histograms 1997 VLDB 0.00023363722
323 An Array-Based Algorithm for Simultaneous Multidimensional Aggregates 1997 SIGMOD 0.0002100085
378 AutoAdmin "What-if" Index Analysis Utility 1998 SIGMOD 0.00019549382
388 Incremental Maintenance of Views with Duplicates 1995 SIGMOD 0.00019350381
403 Bottom-Up Computation of Sparse and Iceberg CUBEs 1999 SIGMOD 0.0001910396
510 TelegraphCQ: Continuous Dataflow Processing 2003 SIGMOD 0.00017065714
519 Random Sampling for Histogram Construction: How much is enough? 1998 SIGMOD 0.00016942879
521 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.00016929744
678 StatStream: Statistical Monitoring of Thousands of Data Streams in Real Time 2002 VLDB 0.00014838454
685 Trill: A High-Performance Incremental Query Processor for Diverse Analytics 2015 VLDB 0.00014782777
720 TIMBER: A Sophisticated Relation Browser 1982 VLDB 0.00014528145
821 Query Execution Techniques for Caching Expensive Methods 1996 SIGMOD 0.00013660347
901 Maintenance of Data Cubes and Summary Tables in a Warehouse 1997 SIGMOD 0.00013185553
1,215 Overview of Data Exploration Techniques 2015 SIGMOD 0.00011492648
1,272 Discovering Queries based on Example Tuples 2014 SIGMOD 0.00011248674
1,607 SciBORQ: Scientific data management with Bounds On Runtime and Quality 2011 CIDR 0.0001008742
1,608 Dynamic Prefetching of Data Tiles for Interactive Visualization 2016 SIGMOD 0.00010084292
1,764 Caching Multidimensional Queries Using Chunks 1998 SIGMOD 9.6965217e-05
1,792 Fast Approximate Correlation for Massive Time-series Data 2010 SIGMOD 9.6238255e-05
1,826 Effective Use of Block-Level Sampling in Statistics Estimation 2004 SIGMOD 9.5575424e-05
1,948 VizDeck: Self-Organizing Dashboards for Visual Analytics 2012 SIGMOD 9.3212031e-05
2,247 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.7585767e-05
2,402 i3: Intelligent, Interactive Investigation of OLAP data cubes 2000 SIGMOD 8.5189151e-05
2,520 VizQL: A Language for Query, Analysis and Visualization 2006 SIGMOD 8.3479795e-05
2,544 Star-Cubing: Computing Iceberg Cubes by Top-Down and Bottom-Up Integration 2003 VLDB 8.3179863e-05
2,550 SQLShare: Results from a Multi-Year SQL-as-a-Service Experiment 2016 SIGMOD 8.3121033e-05
2,648 Explore-by-Example: An Automatic Query Steering Framework for Interactive Data Exploration 2014 SIGMOD 8.1761271e-05
2,718 Answering Queries from Statistics and Probabilistic Views 2005 VLDB 8.0978764e-05
3,087 Continuous Sampling for Online Aggregation Over Multiple Queries 2010 SIGMOD 7.6624333e-05
3,196 Assessing and Ranking Structural Correlations in Graphs 2011 SIGMOD 7.5460261e-05
4,328 Playful Query Specification with DataPlay 2012 VLDB 6.6614241e-05
5,168 Querying Without Keyboards 2013 CIDR 6.2460238e-05
5,180 The Case for Data Visualization Management Systems 2014 VLDB 6.240722e-05
5,783 Sampling Cube: A Framework for Statistical OLAP Over Sampling Data 2008 SIGMOD 5.9963988e-05
8,142 Information Retrieval from an Incomplete Data Cube 1996 VLDB 5.3913918e-05
8,814 ARCube: Supporting Ranking Aggregate Queries in Partially Materialized Data Cubes 2008 SIGMOD 5.2710039e-05
9,686 Tracking Set Correlations at Large Scale 2014 SIGMOD 5.1414953e-05
10,201 Structures, Semantics and Statistics 2004 VLDB 5.0611832e-05
Previous Page 1 / 1 Next

Semantically Similar Papers