DBScholar

Back to papers

The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward

Summary: Cosmos' exabyte-scale evolution at Microsoft spans reliability, scale, efficiency, and usability, with next steps toward security, compliance, and heterogeneous analytics. The paper links Cosmos workload evolution to broad big-data trends, offering platform-driven design insights for researchers. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hdf562d8aa555c147
Venue
VLDB
Year
2021
Pagerank
5.837329e-05
Overall Rank
6,245 | 58.02%
DOI
10.14778/3476311.3476390

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{power_vldb21,
        title = {{The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward}},
        author = {Power, Conor and Patel, Hiren and Jindal, Alekh and Leeka, Jyoti and Jenkins, Bob and Rys, Michael and Triou, Ed and Zhu, Dexin and Katahanas, Lucky and Talapady, Chakrapani Bhat and Rowe, Joshua and Zhang, Fan and Draves, Rich and Friedman, Marc and Filho, Ivan Santa Maria and Kumar, Amrish},
        journal = {PVLDB},
        series = {{VLDB} '21},
        volume = {14},
        number = {12},
        pages = {3148--3161},
        doi = {10.14778/3476311.3476390},
        url = {https://doi.org/10.14778/3476311.3476390},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 10 of 10 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 32 of 32 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
78 Automatic Database Management System Tuning Through Large-scale Machine Learning 2017 SIGMOD 0.00036684414
281 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022295232
459 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores 2020 VLDB 0.00017856221
555 SageDB: A Learned Database System 2019 CIDR 0.00016506678
685 Trill: A High-Performance Incremental Query Processor for Diverse Analytics 2015 VLDB 0.00014782777
840 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.0001354605
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012964445
950 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012895553
1,433 Towards a Learning Optimizer for Shared Clouds 2019 VLDB 0.00010677711
1,745 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.7343818e-05
1,883 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.4391795e-05
2,207 Magpie: Python at Speed and Scale using Cloud Backends 2021 CIDR 8.8487039e-05
2,230 Data Warehousing and Analytics Infrastructure at Facebook 2010 SIGMOD 8.799877e-05
2,446 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 8.4547121e-05
2,577 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.2738285e-05
2,834 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings 2020 SIGMOD 7.9560627e-05
3,073 Pushing Data-Induced Predicates Through Joins in Big-Data Clusters 2020 VLDB 7.6777283e-05
3,440 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.2986853e-05
3,545 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2134803e-05
3,683 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 7.1006425e-05
3,895 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9393157e-05
4,249 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.7037533e-05
4,691 The "Big Data" Ecosystem at LinkedIn 2013 SIGMOD 6.4685355e-05
4,792 Error-bounded Sampling for Analytics on Big Sparse Data 2014 VLDB 6.4130671e-05
5,108 Efficient Estimation of Inclusion Coefficient using HyperLogLog Sketches 2018 VLDB 6.2696201e-05
6,174 Helios: Hyperscale Indexing for the Cloud & Edge 2020 VLDB 5.8612829e-05
6,290 Incorporating Super-Operators in Big-Data Query Optimizers 2020 VLDB 5.8223076e-05
7,084 KEA: Tuning an Exabyte-Scale Data Infrastructure 2021 SIGMOD 5.6029455e-05
7,312 Bubble Execution: Resource-aware Reliable Analytics at Cloud Scale 2018 VLDB 5.5555613e-05
7,759 AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft 2020 VLDB 5.4575614e-05
8,283 Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters 2019 VLDB 5.3627138e-05
9,859 Winds from Seattle: Database Research Directions 2020 VLDB 5.1183325e-05
Previous Page 1 / 1 Next

Semantically Similar Papers