DBScholar

Back to papers

Dremel: Interactive Analysis of Web-Scale Datasets

Summary: Dremel: scalable interactive query system for nested data, using multilevel execution trees and columnar layout for aggregations. Nested-columnar storage and tree-based execution enable petabyte-scale analysis on thousands of CPUs, complementing MapReduce. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
ha2f19a0707fe9579
Venue
VLDB
Year
2010
Pagerank
0.00043160717
Overall Rank
49 | 99.68%
DOI
10.14778/1920841.1920886

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{melnik_vldb10,
        title = {{Dremel: Interactive Analysis of Web-Scale Datasets}},
        author = {Melnik, Sergey and Gubarev, Andrey and Long, Jing Jing and Romer, Geoffrey and Shivakumar, Shiva and Tolton, Matt and Vassilakis, Theo},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {330--339},
        doi = {10.14778/1920841.1920886},
        url = {https://doi.org/10.14778/1920841.1920886},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 132 citing papers.

Rank Citing Paper Year Venue Pagerank
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055406774
52 The Snowflake Elastic Data Warehouse 2016 SIGMOD 0.00041219077
327 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.0002095191
379 Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources 2018 SIGMOD 0.00019514689
432 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018339357
509 Goods: Organizing Google's Datasets 2016 SIGMOD 0.00017071087
746 Spanner: Becoming a SQL System 2017 SIGMOD 0.00014286107
840 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.0001354605
876 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013309176
977 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012731794
1,036 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.00012377471
1,072 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012168947
1,082 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012122749
1,135 Dremel: A Decade of Interactive SQL Analysis at Web Scale 2020 VLDB 0.00011891907
1,210 Processing a Trillion Cells per Mouse Click 2012 VLDB 0.00011527605
1,231 Communication Steps for Parallel Query Processing 2013 PODS 0.00011410373
1,236 Druid: A Real-time Analytical Data Store 2014 SIGMOD 0.00011402848
1,276 Orca: A Modular Query Optimizer Architecture for Big Data 2014 SIGMOD 0.00011239266
1,292 From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System 2015 SIGMOD 0.00011152286
1,381 Scuba: Diving into Data at Facebook 2013 VLDB 0.00010860462
1,423 Procella: Unifying serving and analytical data at YouTube 2019 VLDB 0.00010719956
1,462 Manu: A Cloud Native Vector Database Management System 2022 VLDB 0.0001058099
1,622 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing 2014 VLDB 0.00010051462
1,680 Axiomatic Foundations and Algorithms for Deciding Semantic Equivalences of SQL Queries 2018 VLDB 9.9031411e-05
1,743 Sinew: A SQL System for Multi-Structured Data 2014 SIGMOD 9.7368807e-05
1,790 POLARIS: The Distributed SQL Engine in Azure Synapse 2020 VLDB 9.6272006e-05
1,884 SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures 2014 VLDB 9.4388097e-05
2,160 DIFF: A Relational Interface for Large-Scale Data Explanation 2019 VLDB 8.9364035e-05
2,314 BtrBlocks: Efficient Columnar Compression for Data Lakes 2023 SIGMOD 8.6533171e-05
2,436 Mison: A Fast JSON Parser for Data Analytics 2017 VLDB 8.4708399e-05
2,484 AnalyticDB: Real-time OLAP Database System at Alibaba Cloud 2019 VLDB 8.3993882e-05
2,573 Minimal MapReduce Algorithms 2013 SIGMOD 8.2821647e-05
2,689 Major Technical Advancements in Apache Hive 2014 SIGMOD 8.1264718e-05
2,735 Speculative Distributed CSV Data Parsing for Big Data Analytics 2019 SIGMOD 8.0761736e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8853204e-05
2,903 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 7.8810111e-05
2,915 F1 Query: Declarative Querying at Scale 2018 VLDB 7.8616593e-05
2,997 F1 Lightning: HTAP as a Service 2020 VLDB 7.7689996e-05
3,007 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.7607173e-05
3,069 Adaptive Query Processing on RAW Data 2014 VLDB 7.6822697e-05
3,081 Skipping-oriented Partitioning for Columnar Layouts 2017 VLDB 7.6653727e-05
3,424 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.3117029e-05
3,440 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.2986853e-05
3,445 Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing 2017 VLDB 7.2943981e-05
3,446 Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing 2019 SIGMOD 7.2942885e-05
3,465 DIAMetrics: Benchmarking Query Engines at Scale 2020 VLDB 7.2808932e-05
3,505 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2508161e-05
3,520 FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMS 2021 VLDB 7.2360022e-05
3,532 Rethinking Data-Intensive Science Using Scalable Analytics Systems 2015 SIGMOD 7.2261031e-05
3,648 Davos: A System for Interactive Data-Driven Decision Making 2021 VLDB 7.1370661e-05
Previous Page 1 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers