PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes
Summary: Retargets AFrame from AsterixDB to a backend-agnostic, query-based layer for scalable DataFrame analytics. Introduces PolyFrame, preserving Pandas API while incrementally shaping queries for diverse composable languages across DBMS backends. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Phanwadee Sinthong (University of California Irvine)
- 2. Michael J. Carey (University of California Irvine)
BibTeX Citation
@article{sinthong_vldb21,
title = {{PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes}},
author = {Sinthong, Phanwadee and Carey, Michael J.},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {11},
pages = {2296--2304},
doi = {10.14778/3476249.3476281},
url = {https://doi.org/10.14778/3476249.3476281},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,823 | Query Processing on Tensor Computation Runtimes | 2022 | VLDB | 8.0893814e-05 |
| 7,502 | Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine | 2025 | VLDB | 5.6029996e-05 |
| 10,074 | Dias: Dynamic Rewriting of Pandas Code | 2024 | SIGMOD | 5.1613298e-05 |
| 11,235 | SplitDF: Splitting Dataframes for Memory-Efficient Data Analysis | 2024 | VLDB | 5.093636e-05 |
| 11,646 | Wisconsin Benchmark Data Generator: To JSON and Beyond | 2021 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 1,015 | AsterixDB: A Scalable, Open Source BDMS | 2014 | VLDB | 0.00012647763 |
| 1,431 | Towards Scalable Dataframe Systems | 2020 | VLDB | 0.00010807221 |
| 2,651 | Magpie: Python at Speed and Scale using Cloud Backends | 2021 | CIDR | 8.2918086e-05 |
| 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB | 7.436229e-05 |
| 4,292 | Putting Pandas in a Box | 2021 | CIDR | 6.7784325e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,567 | When sweet and cute isn't enough anymore: Solving scalability issues in Python Pandas with Grizzly | 2020 | CIDR |
| 2 | 4,292 | Putting Pandas in a Box | 2021 | CIDR |
| 3 | 2,651 | Magpie: Python at Speed and Scale using Cloud Backends | 2021 | CIDR |
| 4 | 7,484 | On Scale Independence for Querying Big Data | 2014 | PODS |
| 5 | 6,628 | ConnectorX: Accelerating Data Loading From Databases to Dataframes | 2022 | VLDB |
| 6 | 4,079 | Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System | 2022 | VLDB |
| 7 | 9,069 | DQDF: Data-Quality-Aware Dataframes | 2022 | VLDB |
| 8 | 9,508 | An IDEA: An Ingestion Framework for Data Enrichment in AsterixDB | 2019 | VLDB |
| 9 | 8,129 | Large-scale Complex Analytics on Semi-structured Datasets using AsterixDB and Spark | 2016 | VLDB |
| 10 | 1,431 | Towards Scalable Dataframe Systems | 2020 | VLDB |