PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes
Summary: Retargets AFrame from AsterixDB to a backend-agnostic, query-based layer for scalable DataFrame analytics. Introduces PolyFrame, preserving Pandas API while incrementally shaping queries for diverse composable languages across DBMS backends. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Phanwadee Sinthong (University of California Irvine)
- 2. Michael J. Carey (University of California Irvine)
BibTeX Citation
@article{sinthong_vldb21,
title = {{PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes}},
author = {Sinthong, Phanwadee and Carey, Michael J.},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {11},
pages = {2296--2304},
doi = {10.14778/3476249.3476281},
url = {https://doi.org/10.14778/3476249.3476281},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,460 | Query Processing on Tensor Computation Runtimes | 2022 | VLDB | 8.4348335e-05 |
| 7,269 | Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine | 2025 | VLDB | 5.5694935e-05 |
| 10,278 | Dias: Dynamic Rewriting of Pandas Code | 2024 | SIGMOD | 5.0455234e-05 |
| 11,570 | SplitDF: Splitting Dataframes for Memory-Efficient Data Analysis | 2024 | VLDB | 4.9793485e-05 |
| 11,953 | Wisconsin Benchmark Data Generator: To JSON and Beyond | 2021 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 23 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00055406774 |
| 922 | AsterixDB: A Scalable, Open Source BDMS | 2014 | VLDB | 0.00013068048 |
| 1,369 | Towards Scalable Dataframe Systems | 2020 | VLDB | 0.00010899832 |
| 2,207 | Magpie: Python at Speed and Scale using Cloud Backends | 2021 | CIDR | 8.8487039e-05 |
| 3,384 | Putting Pandas in a Box | 2021 | CIDR | 7.3550055e-05 |
| 3,455 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB | 7.2884813e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,384 | Putting Pandas in a Box | 2021 | CIDR |
| 2 | 2,207 | Magpie: Python at Speed and Scale using Cloud Backends | 2021 | CIDR |
| 3 | 7,610 | On Scale Independence for Querying Big Data | 2014 | PODS |
| 4 | 10,995 | Painless and Efficient Scaling of Pandas Programs: Demo | 2026 | VLDB |
| 5 | 6,116 | ConnectorX: Accelerating Data Loading From Databases to Dataframes | 2022 | VLDB |
| 6 | 3,930 | Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System | 2022 | VLDB |
| 7 | 9,246 | DQDF: Data-Quality-Aware Dataframes | 2022 | VLDB |
| 8 | 9,676 | An IDEA: An Ingestion Framework for Data Enrichment in AsterixDB | 2019 | VLDB |
| 9 | 8,246 | Large-scale Complex Analytics on Semi-structured Datasets using AsterixDB and Spark | 2016 | VLDB |
| 10 | 1,369 | Towards Scalable Dataframe Systems | 2020 | VLDB |