Back to papers
Magpie: Python at Speed and Scale using Cloud Backends
Summary: Magpie exposes the Pandas API but lazily pushes dataframe work into cloud query engines (SQL DW, Spark, SCOPE) through a common data layer, avoiding cross-engine transfer and leveraging DB-grade features. It auto-selects optimal backends to deliver database-scale performance to Python analytics; production traces show ~25% of internal computations could benefit.
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 407
- Venue
- CIDR
- Year
- 2021
- Pagerank
- 7.8188583e-05
- Overall Rank
- 2,955 | 79.47%
- DOI
-
-
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 16 of 16 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 3,393 |
Lux: Always-on Visualization Recommendations for Exploratory Dataframe Workflows |
2022 |
VLDB |
7.1415762e-05 |
| 3,409 |
End-to-end Optimization of Machine Learning Prediction Queries |
2022 |
SIGMOD |
7.1240791e-05 |
| 3,772 |
Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System |
2022 |
VLDB |
6.7736479e-05 |
| 4,776 |
PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes |
2021 |
VLDB |
5.9263134e-05 |
| 5,741 |
Babelfish: Efficient Execution of Polyglot Queries |
2022 |
VLDB |
5.3450701e-05 |
| 6,278 |
The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward |
2021 |
VLDB |
5.1241654e-05 |
| 6,539 |
ConnectorX: Accelerating Data Loading From Databases to Dataframes |
2022 |
VLDB |
5.0168759e-05 |
| 6,703 |
YeSQL: “You extend SQL” with Rich and Highly Performant User-Defined Functions in Relational Databases |
2022 |
VLDB |
4.9514593e-05 |
| 6,898 |
Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine |
2025 |
VLDB |
4.8878659e-05 |
| 8,580 |
Efficient Execution of User-Defined Functions in SQL Queries |
2023 |
VLDB |
4.4876382e-05 |
| 8,625 |
Predicate Pushdown for Data Science Pipelines |
2023 |
SIGMOD |
4.4784651e-05 |
| 9,348 |
The Key to Effective UDF Optimization: Before Inlining, First Perform Outlining |
2025 |
VLDB |
4.3504473e-05 |
| 9,764 |
QURE: AI-Assisted and Automatically Verified UDF Inlining |
2025 |
SIGMOD |
4.2815042e-05 |
| 9,910 |
Dias: Dynamic Rewriting of Pandas Code |
2024 |
SIGMOD |
4.2524493e-05 |
| 10,934 |
Proactive Resume and Pause of Resources for Microsoft Azure SQL Database Serverless |
2024 |
SIGMOD |
4.1905499e-05 |
| 11,027 |
SplitDF: Splitting Dataframes for Memory-Efficient Data Analysis |
2024 |
VLDB |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 1,107 |
Froid: Optimization of Imperative Programs in a Relational Database |
2018 |
VLDB |
0.0001397627 |
| 1,427 |
Towards Scalable Dataframe Systems |
2020 |
VLDB |
0.00012033503 |
| 1,457 |
Rewriting Procedures for Batched Bindings |
2008 |
VLDB |
0.00011891025 |
| 1,632 |
Garlic: A New Flavor of Federated Query Processing for DB2 |
2002 |
SIGMOD |
0.00011070537 |
| 1,856 |
AI Meets AI: Leveraging Query Executions to Improve Index Recommendations |
2019 |
SIGMOD |
0.00010319105 |
| 2,939 |
AIDA - Abstraction for Advanced In-Database Analytics |
2018 |
VLDB |
7.8514162e-05 |
| 3,044 |
Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics |
2017 |
SIGMOD |
7.6624689e-05 |
| 3,298 |
Extracting Equivalent SQL from Imperative Code in Database Applications |
2016 |
SIGMOD |
7.2527707e-05 |
| 3,312 |
Automatic Partitioning of Database Applications |
2012 |
VLDB |
7.2356116e-05 |
| 3,623 |
Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings |
2020 |
SIGMOD |
6.9017341e-05 |
| 3,983 |
The Myria Big Data Management and Analytics System and Cloud Service |
2017 |
CIDR |
6.5588011e-05 |
| 4,164 |
Sloth: Being Lazy is a Virtue (When Issuing Database Queries) |
2014 |
SIGMOD |
6.3859482e-05 |
| 7,044 |
Seagull: An Infrastructure for Load Prediction and Optimized Resource Allocation |
2021 |
VLDB |
4.8475963e-05 |
| 7,447 |
DBridge: Translating Imperative Code to SQL |
2017 |
SIGMOD |
4.7227748e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 2,072 |
HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics |
2016 |
VLDB |
9.6300019e-05 |
| 6,605 |
MotherDuck: DuckDB in the cloud and in the client |
2024 |
CIDR |
4.9923144e-05 |
| 4,776 |
PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes |
2021 |
VLDB |
5.9263134e-05 |
| 1,884 |
Tuplex: Data Science in Python at Native Code Speed |
2021 |
SIGMOD |
0.00010206514 |
| 4,681 |
Cloud Analytics Benchmark |
2023 |
VLDB |
5.9953512e-05 |
| 6,191 |
Accelerating Python UDFs in Vectorized Query Execution |
2022 |
CIDR |
5.1598046e-05 |
| 9,422 |
When sweet and cute isn't enough anymore: Solving scalability issues in Python Pandas with Grizzly |
2020 |
CIDR |
4.3399748e-05 |
| 4,813 |
Putting Pandas in a Box |
2021 |
CIDR |
5.8993009e-05 |
| 1,427 |
Towards Scalable Dataframe Systems |
2020 |
VLDB |
0.00012033503 |
| 1,324 |
Starling: A Scalable Query Engine on Cloud Functions |
2020 |
SIGMOD |
0.00012585081 |