DBScholar

Back to papers

ConnectorX: Accelerating Data Loading From Databases to Dataframes

Summary: ConnectorX speeds loading from DBMSs to dataframes by reducing client-side overhead—the dominant cost—rather than query execution or data transfer. Server-side result partitioning and a modular design enable easy extension to multiple databases and dataframes, yielding speedups over prior libraries. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hc3d1a67a0411a671
Venue
VLDB
Year
2022
Pagerank
5.8815309e-05
Overall Rank
6,116 | 58.89%
DOI
10.14778/3551793.3551847

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{wang_vldb22,
        title = {{ConnectorX: Accelerating Data Loading From Databases to Dataframes}},
        author = {Wang, Xiaoying and Wu, Weiyuan and Wu, Jinze and Chen, Yizhou and Zrymiak, Nick and Qu, Changbo and Flokas, Lampros and Chow, George and Wang, Jiannan and Wang, Tianzheng and Wu, Eugene and Zhou, Qingqing},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {11},
        pages = {2994--3003},
        doi = {10.14778/3551793.3551847},
        url = {https://doi.org/10.14778/3551793.3551847},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055406774
52 The Snowflake Elastic Data Warehouse 2016 SIGMOD 0.00041219077
950 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012895553
1,135 Dremel: A Decade of Interactive SQL Analysis at Web Scale 2020 VLDB 0.00011891907
1,256 Towards Linear Algebra over Normalized Data 2017 VLDB 0.00011314687
1,369 Towards Scalable Dataframe Systems 2020 VLDB 0.00010899832
1,371 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.00010899041
1,825 Data Management for Data Science: Towards Embedded Analytics 2020 CIDR 9.5603293e-05
2,182 Extending Relational Query Processing with ML Inference 2020 CIDR 8.8982998e-05
2,207 Magpie: Python at Speed and Scale using Cloud Backends 2021 CIDR 8.8487039e-05
2,778 AIDA - Abstraction for Advanced In-Database Analytics 2018 VLDB 8.0299506e-05
2,830 DB4ML – An In-Memory Database Kernel with Machine Learning Support 2020 SIGMOD 7.9619979e-05
2,975 In-Database Learning with Sparse Tensors 2018 PODS 7.7907759e-05
3,384 Putting Pandas in a Box 2021 CIDR 7.3550055e-05
3,524 MLog: Towards Declarative In-Database Machine Learning 2017 VLDB 7.2337006e-05
4,112 Don’t Hold My Data Hostage – A Case For Client Protocol Redesign 2017 VLDB 6.7994301e-05
5,944 Bridging Two Worlds with RICE: Integrating R into the SAP In-Memory Computing Engine 2011 VLDB 5.9387238e-05
6,297 Mainlining Databases: Supporting Fast Transactional Workloads on Universal Columnar Data File Formats 2021 VLDB 5.81916e-05
Previous Page 1 / 1 Next

Semantically Similar Papers