DBScholar

Back to papers

The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward

Summary: Cosmos' exabyte-scale evolution at Microsoft spans reliability, scale, efficiency, and usability, with next steps toward security, compliance, and heterogeneous analytics. The paper links Cosmos workload evolution to broad big-data trends, offering platform-driven design insights for researchers. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12707
Venue
VLDB
Year
2021
Pagerank
5.9688569e-05
Overall Rank
6,121 | 58.01%
DOI
10.14778/3476311.3476390

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{power_vldb21,
        title = {{The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward}},
        author = {Power, Conor and Patel, Hiren and Jindal, Alekh and Leeka, Jyoti and Jenkins, Bob and Rys, Michael and Triou, Ed and Zhu, Dexin and Katahanas, Lucky and Talapady, Chakrapani Bhat and Rowe, Joshua and Zhang, Fan and Draves, Rich and Friedman, Marc and Filho, Ivan Santa Maria and Kumar, Amrish},
        journal = {PVLDB},
        series = {{VLDB} '21},
        volume = {14},
        number = {12},
        pages = {3148--3161},
        doi = {10.14778/3476311.3476390},
        url = {https://doi.org/10.14778/3476311.3476390},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 10 of 10 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 32 of 32 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
86 Automatic Database Management System Tuning Through Large-scale Machine Learning 2017 SIGMOD 0.00035316107
295 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022238183
520 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores 2020 VLDB 0.00017136828
568 SageDB: A Learned Database System 2019 CIDR 0.0001641553
710 Trill: A High-Performance Incremental Query Processor for Diverse Analytics 2015 VLDB 0.00014715033
819 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013815639
923 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00013189886
1,138 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012023643
1,468 Towards a Learning Optimizer for Shared Clouds 2019 VLDB 0.00010686496
1,765 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.8079546e-05
1,852 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.6134443e-05
2,191 Data Warehousing and Analytics Infrastructure at Facebook 2010 SIGMOD 8.9780154e-05
2,477 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 8.5239378e-05
2,567 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.4098241e-05
2,651 Magpie: Python at Speed and Scale using Cloud Backends 2021 CIDR 8.2918086e-05
2,822 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings 2020 SIGMOD 8.0898536e-05
3,137 Pushing Data-Induced Predicates Through Joins in Big-Data Clusters 2020 VLDB 7.7204167e-05
3,466 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.3909785e-05
3,605 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2640711e-05
3,614 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 7.2568185e-05
3,964 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9855158e-05
4,169 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.8561463e-05
4,595 The "Big Data" Ecosystem at LinkedIn 2013 SIGMOD 6.6137084e-05
4,696 Error-bounded Sampling for Analytics on Big Sparse Data 2014 VLDB 6.557612e-05
5,257 Efficient Estimation of Inclusion Coefficient using HyperLogLog Sketches 2018 VLDB 6.2971456e-05
6,045 Helios: Hyperscale Indexing for the Cloud & Edge 2020 VLDB 5.9957544e-05
6,194 Incorporating Super-Operators in Big-Data Query Optimizers 2020 VLDB 5.9470844e-05
6,946 KEA: Tuning an Exabyte-Scale Data Infrastructure 2021 SIGMOD 5.7309848e-05
7,182 Bubble Execution: Resource-aware Reliable Analytics at Cloud Scale 2018 VLDB 5.6793679e-05
7,619 AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft 2020 VLDB 5.5810604e-05
8,108 Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters 2019 VLDB 5.4850569e-05
9,681 Winds from Seattle: Database Research Directions 2020 VLDB 5.2357516e-05
Previous Page 1 / 1 Next

Semantically Similar Papers