PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes
Summary: Retargets AFrame from AsterixDB to a backend-agnostic, query-based layer for scalable DataFrame analytics. Introduces PolyFrame, preserving Pandas API while incrementally shaping queries for diverse composable languages across DBMS backends. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Phanwadee Sinthong (University of California Irvine)
- 2. Michael J. Carey (University of California Irvine)
BibTeX Citation
@article{sinthong_vldb21,
title = {{PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes}},
author = {Sinthong, Phanwadee and Carey, Michael J.},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {11},
pages = {2296--2304},
doi = {10.14778/3476249.3476281},
url = {https://doi.org/10.14778/3476249.3476281},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,460 | Query Processing on Tensor Computation Runtimes | 2022 | VLDB | 8.4308406e-05 |
| 7,272 | Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine | 2025 | VLDB | 5.5668569e-05 |
| 10,284 | Dias: Dynamic Rewriting of Pandas Code | 2024 | SIGMOD | 5.0431349e-05 |
| 11,576 | SplitDF: Splitting Dataframes for Memory-Efficient Data Analysis | 2024 | VLDB | 4.9769913e-05 |
| 11,959 | Wisconsin Benchmark Data Generator: To JSON and Beyond | 2021 | SIGMOD | 4.9769913e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 23 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00055384955 |
| 921 | AsterixDB: A Scalable, Open Source BDMS | 2014 | VLDB | 0.00013064043 |
| 1,370 | Towards Scalable Dataframe Systems | 2020 | VLDB | 0.00010895207 |
| 2,208 | Magpie: Python at Speed and Scale using Cloud Backends | 2021 | CIDR | 8.8445332e-05 |
| 3,385 | Putting Pandas in a Box | 2021 | CIDR | 7.3515238e-05 |
| 3,455 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB | 7.2850317e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,385 | Putting Pandas in a Box | 2021 | CIDR |
| 2 | 2,208 | Magpie: Python at Speed and Scale using Cloud Backends | 2021 | CIDR |
| 3 | 7,612 | On Scale Independence for Querying Big Data | 2014 | PODS |
| 4 | 11,004 | Painless and Efficient Scaling of Pandas Programs: Demo | 2026 | VLDB |
| 5 | 6,117 | ConnectorX: Accelerating Data Loading From Databases to Dataframes | 2022 | VLDB |
| 6 | 3,931 | Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System | 2022 | VLDB |
| 7 | 9,256 | DQDF: Data-Quality-Aware Dataframes | 2022 | VLDB |
| 8 | 9,683 | An IDEA: An Ingestion Framework for Data Enrichment in AsterixDB | 2019 | VLDB |
| 9 | 8,252 | Large-scale Complex Analytics on Semi-structured Datasets using AsterixDB and Spark | 2016 | VLDB |
| 10 | 1,370 | Towards Scalable Dataframe Systems | 2020 | VLDB |