Database Paper Browser

Back to papers

The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward

Summary: Cosmos' exabyte-scale evolution at Microsoft spans reliability, scale, efficiency, and usability, with next steps toward security, compliance, and heterogeneous analytics. The paper links Cosmos workload evolution to broad big-data trends, offering platform-driven design insights for researchers. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12520
Venue
VLDB
Year
2021
Pagerank
5.1241654e-05
Overall Rank
6,278 | 56.37%
DOI
10.14778/3476311.3476390

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 10 of 10 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 32 of 32 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
183 Automatic Database Management System Tuning Through Large-scale Machine Learning 2017 SIGMOD 0.00036859633
332 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00027173479
739 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores 2020 VLDB 0.00017365933
796 SageDB: A Learned Database System 2019 CIDR 0.00016541749
1,048 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00014442178
1,097 Trill: A High-Performance Incremental Query Processor for Diverse Analytics 2015 VLDB 0.00014083973
1,320 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00012606067
1,356 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012409986
1,921 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 0.00010085899
2,080 Towards a Learning Optimizer for Shared Clouds 2019 VLDB 9.5954034e-05
2,410 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 8.8643562e-05
2,658 Data Warehousing and Analytics Infrastructure at Facebook 2010 SIGMOD 8.3634079e-05
2,955 Magpie: Python at Speed and Scale using Cloud Backends 2021 CIDR 7.8188583e-05
3,044 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 7.6624689e-05
3,139 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 7.4915127e-05
3,623 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings 2020 SIGMOD 6.9017341e-05
3,866 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 6.6795497e-05
3,923 Pushing Data-Induced Predicates Through Joins in Big-Data Clusters 2020 VLDB 6.6232068e-05
4,068 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 6.4748133e-05
4,171 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 6.3800823e-05
4,226 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.3382156e-05
4,563 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.0752535e-05
4,859 The "Big Data" Ecosystem at LinkedIn 2013 SIGMOD 5.8680444e-05
5,258 Error-bounded Sampling for Analytics on Big Sparse Data 2014 VLDB 5.5973455e-05
5,371 Efficient Estimation of Inclusion Coefficient using HyperLogLog Sketches 2018 VLDB 5.5424564e-05
6,241 Helios: Hyperscale Indexing for the Cloud & Edge 2020 VLDB 5.1357476e-05
6,671 Incorporating Super-Operators in Big-Data Query Optimizers 2020 VLDB 4.9625353e-05
7,099 KEA: Tuning an Exabyte-Scale Data Infrastructure 2021 SIGMOD 4.8263529e-05
7,386 Bubble Execution: Resource-aware Reliable Analytics at Cloud Scale 2018 VLDB 4.7398757e-05
7,685 AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft 2020 VLDB 4.6753414e-05
8,235 Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters 2019 VLDB 4.5481384e-05
9,530 Winds from Seattle: Database Research Directions 2020 VLDB 4.3251179e-05
Previous Page 1 / 1 Next

Semantically Similar Papers