DBScholar

Back to papers

Dremel: Interactive Analysis of Web-Scale Datasets

Summary: Dremel: scalable interactive query system for nested data, using multilevel execution trees and columnar layout for aggregations. Nested-columnar storage and tree-based execution enable petabyte-scale analysis on thousands of CPUs, complementing MapReduce. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
10279
Venue
VLDB
Year
2010
Pagerank
0.0004291425
Overall Rank
51 | 99.66%
DOI
10.14778/1920841.1920886

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{melnik_vldb10,
        title = {{Dremel: Interactive Analysis of Web-Scale Datasets}},
        author = {Melnik, Sergey and Gubarev, Andrey and Long, Jing Jing and Romer, Geoffrey and Shivakumar, Shiva and Tolton, Matt and Vassilakis, Theo},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {330--339},
        doi = {10.14778/1920841.1920886},
        url = {https://doi.org/10.14778/1920841.1920886},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 128 citing papers.

Rank Citing Paper Year Venue Pagerank
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
66 The Snowflake Elastic Data Warehouse 2016 SIGMOD 0.00038561587
330 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.0002104801
425 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018704491
445 Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources 2018 SIGMOD 0.00018336751
514 Goods: Organizing Google's Datasets 2016 SIGMOD 0.00017178673
819 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013815639
845 Spanner: Becoming a SQL System 2017 SIGMOD 0.00013660379
872 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013486409
954 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012997301
1,044 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.0001244236
1,054 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012390673
1,108 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012145154
1,207 Communication Steps for Parallel Query Processing 2013 PODS 0.00011663155
1,225 Processing a Trillion Cells per Mouse Click 2012 VLDB 0.00011590013
1,242 Druid: A Real-time Analytical Data Store 2014 SIGMOD 0.00011516162
1,320 From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System 2015 SIGMOD 0.00011166426
1,352 Scuba: Diving into Data at Facebook 2013 VLDB 0.00011064595
1,453 Dremel: A Decade of Interactive SQL Analysis at Web Scale 2020 VLDB 0.00010742227
1,494 Procella: Unifying serving and analytical data at YouTube 2019 VLDB 0.00010577585
1,589 Manu: A Cloud Native Vector Database Management System 2022 VLDB 0.00010264469
1,606 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing 2014 VLDB 0.00010228576
1,621 Orca: A Modular Query Optimizer Architecture for Big Data 2014 SIGMOD 0.00010203114
1,738 Sinew: A SQL System for Multi-Structured Data 2014 SIGMOD 9.8823494e-05
1,829 Axiomatic Foundations and Algorithms for Deciding Semantic Equivalences of SQL Queries 2018 VLDB 9.6671311e-05
1,889 SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures 2014 VLDB 9.5335988e-05
1,925 POLARIS: The Distributed SQL Engine in Azure Synapse 2020 VLDB 9.4764398e-05
2,161 DIFF: A Relational Interface for Large-Scale Data Explanation 2019 VLDB 9.0606664e-05
2,391 Mison: A Fast JSON Parser for Data Analytics 2017 VLDB 8.6413407e-05
2,539 Minimal MapReduce Algorithms 2013 SIGMOD 8.4526595e-05
2,563 AnalyticDB: Real-time OLAP Database System at Alibaba Cloud 2019 VLDB 8.412445e-05
2,706 Major Technical Advancements in Apache Hive 2014 SIGMOD 8.2287564e-05
2,788 BtrBlocks: Efficient Columnar Compression for Data Lakes 2023 SIGMOD 8.1205155e-05
2,849 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 8.053191e-05
2,850 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 8.0499212e-05
2,889 F1 Query: Declarative Querying at Scale 2018 VLDB 7.9935046e-05
2,942 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.9358593e-05
2,972 F1 Lightning: HTAP as a Service 2020 VLDB 7.9114157e-05
3,062 Adaptive Query Processing on RAW Data 2014 VLDB 7.8037446e-05
3,074 Speculative Distributed CSV Data Parsing for Big Data Analytics 2019 SIGMOD 7.7844208e-05
3,106 Skipping-oriented Partitioning for Columnar Layouts 2017 VLDB 7.7515666e-05
3,366 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.4748604e-05
3,413 Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing 2017 VLDB 7.4326381e-05
3,460 DIAMetrics: Benchmarking Query Engines at Scale 2020 VLDB 7.3939262e-05
3,466 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.3909785e-05
3,552 Rethinking Data-Intensive Science Using Scalable Analytics Systems 2015 SIGMOD 7.3188008e-05
3,555 Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing 2019 SIGMOD 7.3115321e-05
3,570 Davos: A System for Interactive Data-Driven Decision Making 2021 VLDB 7.3008782e-05
3,632 FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMS 2021 VLDB 7.2368817e-05
3,638 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2338361e-05
Previous Page 1 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers