DBScholar

Back to papers

The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward

Summary: Cosmos' exabyte-scale evolution at Microsoft spans reliability, scale, efficiency, and usability, with next steps toward security, compliance, and heterogeneous analytics. The paper links Cosmos workload evolution to broad big-data trends, offering platform-driven design insights for researchers. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hdf562d8aa555c147
Venue
VLDB
Year
2021
Pagerank
5.8345657e-05
Overall Rank
6,248 | 58.01%
DOI
10.14778/3476311.3476390
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{power_vldb21,
        title = {{The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward}},
        author = {Power, Conor and Patel, Hiren and Jindal, Alekh and Leeka, Jyoti and Jenkins, Bob and Rys, Michael and Triou, Ed and Zhu, Dexin and Katahanas, Lucky and Talapady, Chakrapani Bhat and Rowe, Joshua and Zhang, Fan and Draves, Rich and Friedman, Marc and Filho, Ivan Santa Maria and Kumar, Amrish},
        journal = {PVLDB},
        series = {{VLDB} '21},
        volume = {14},
        number = {12},
        pages = {3148--3161},
        doi = {10.14778/3476311.3476390},
        url = {https://doi.org/10.14778/3476311.3476390},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 10 of 10 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 32 of 32 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
78 Automatic Database Management System Tuning Through Large-scale Machine Learning 2017 SIGMOD 0.00036675568
282 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022302793
460 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores 2020 VLDB 0.00017850215
555 SageDB: A Learned Database System 2019 CIDR 0.0001650754
685 Trill: A High-Performance Incremental Query Processor for Diverse Analytics 2015 VLDB 0.00014778299
841 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013543
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012959992
950 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012889569
1,432 Towards a Learning Optimizer for Shared Clouds 2019 VLDB 0.00010676754
1,747 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.7303647e-05
1,884 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.4349703e-05
2,208 Magpie: Python at Speed and Scale using Cloud Backends 2021 CIDR 8.8445332e-05
2,231 Data Warehousing and Analytics Infrastructure at Facebook 2010 SIGMOD 8.7958881e-05
2,448 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 8.4508839e-05
2,577 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.2699584e-05
2,833 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings 2020 SIGMOD 7.9539771e-05
3,075 Pushing Data-Induced Predicates Through Joins in Big-Data Clusters 2020 VLDB 7.6742518e-05
3,440 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.2954357e-05
3,544 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2108612e-05
3,685 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 7.0972826e-05
3,895 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9360458e-05
4,250 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.7005978e-05
4,693 The "Big Data" Ecosystem at LinkedIn 2013 SIGMOD 6.4654777e-05
4,795 Error-bounded Sampling for Analytics on Big Sparse Data 2014 VLDB 6.4101044e-05
5,112 Efficient Estimation of Inclusion Coefficient using HyperLogLog Sketches 2018 VLDB 6.2668086e-05
6,176 Helios: Hyperscale Indexing for the Cloud & Edge 2020 VLDB 5.8585082e-05
6,293 Incorporating Super-Operators in Big-Data Query Optimizers 2020 VLDB 5.8195888e-05
7,086 KEA: Tuning an Exabyte-Scale Data Infrastructure 2021 SIGMOD 5.6002958e-05
7,315 Bubble Execution: Resource-aware Reliable Analytics at Cloud Scale 2018 VLDB 5.5529695e-05
7,767 AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft 2020 VLDB 5.4550466e-05
8,289 Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters 2019 VLDB 5.360349e-05
9,866 Winds from Seattle: Database Research Directions 2020 VLDB 5.1159095e-05
Previous Page 1 / 1 Next

Semantically Similar Papers