DBScholar

Back to papers

DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines

Summary: DAPHNE: extensible infra unifying DM, HPC and ML pipelines via shared language abstractions, compiler/runtime integration and multi-level scheduling. Key novelty: vectorized engine for computational storage and accelerators to cut data-movement/format overheads for local and distributed ops; prelim results show notable speedups. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
450
Venue
CIDR
Year
2022
Pagerank
5.6855887e-05
Overall Rank
7,160 | 50.88%
DOI
-

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{damme_cidr22,
        address = {Amsterdam, Netherlands},
        series = {{CIDR} '22},
        title = {{DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines}},
        booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
        author = {Damme, Patrick and Birkenbach, Marius and Bitsakos, Constantinos and Boehm, Matthias and Bonnet, Philippe and Ciorba, Florina and Dokter, Mark and Dowgiallo, Pawel and Eleliemy, Ahmed and Faerber, Christian and Goumas, Georgios and Habich, Dirk and Hedam, Niclas and Hofer, Marlies and Huang, Wenjun and Innerebner, Kevin and Karakostas, Vasileios and Kern, Roman and Kosar, Tomaž and Krause, Alexander and Krems, Daniel and Laber, Andreas and Lehner, Wolfgang and Mier, Eric and Paradies, Marcus and Peischl, Bernhard and Poerwawinata, Gabrielle and Psomadakis, Stratos and Rabl, Tilmann and Ratuszniak, Piotr and Silva, Pedro and Skuppin, Nikolai and Starzacher, Andreas and Steinwender, Benjamin and Tolovski, Ilin and Tözün, Pınar and Ulatowski, Wojciech and Wang, Yuanyuan and Wrosz, Izajasz and Zamuda, Aleš and Zhang, Ce and Zhu, Xiao Xiang},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 5 of 5 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 37 of 37 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
3 Pregel: A System for Large-Scale Graph Processing 2010 SIGMOD 0.0012250108
14 MonetDB/X100: Hyper-Pipelining Query Execution 2005 CIDR 0.0006312782
66 The Snowflake Elastic Data Warehouse 2016 SIGMOD 0.00038561587
241 Morsel-Driven Parallelism: A NUMA-Aware Query Evaluation Framework for the Many-Core Age 2014 SIGMOD 0.00023654664
518 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017167492
532 MLbase: A Distributed Machine-learning System 2013 CIDR 0.00017072641
1,002 Designing And Mining Multi-Terabyte Astronomy Archives: The Sloan Digital Sky Survey 2000 SIGMOD 0.00012706131
1,079 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012258469
1,250 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011485301
1,333 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.0001112858
1,675 Garlic: A New Flavor of Federated Query Processing for DB2 2002 SIGMOD 0.00010035333
1,756 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.8172465e-05
1,768 Tuplex: Data Science in Python at Native Code Speed 2021 SIGMOD 9.8041636e-05
2,000 Data Management for Data Science: Towards Embedded Analytics 2020 CIDR 9.3336258e-05
2,164 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 9.0521951e-05
2,208 Query Optimization for Dynamic Imputation 2017 VLDB 8.9512455e-05
2,273 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.8230899e-05
2,566 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.4116562e-05
2,997 Adapting to Source Properties in Processing Data Integration Queries 2004 SIGMOD 7.8745158e-05
3,205 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6386536e-05
3,874 Tensors: An abstraction for general data processing 2021 VLDB 7.0561161e-05
4,090 The bionic DBMS is coming, but what will it look like? 2013 CIDR 6.9096664e-05
4,223 Accelerating Queries with Group-By and Join by Groupjoin 2011 VLDB 6.8224393e-05
4,240 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.809685e-05
4,346 Estimating Compilation Time of a Query Optimizer 2003 SIGMOD 6.7510183e-05
4,557 Not your Grandpa's SSD: The Era of Co-Designed Storage Devices 2021 SIGMOD 6.6333658e-05
4,726 Computational Storage: Where Are We Today? 2021 CIDR 6.5364137e-05
5,204 QuERy: A Framework for Integrating Entity Resolution with Query Processing 2016 VLDB 6.3187983e-05
5,855 Bridging Two Worlds with RICE: Integrating R into the SAP In-Memory Computing Engine 2011 VLDB 6.064726e-05
5,876 BAGUA: Scaling up Distributed Learning with System Relaxations 2022 VLDB 6.0557672e-05
6,572 A Morsel-Driven Query Execution Engine for Heterogeneous Multi-Cores 2019 VLDB 5.8373744e-05
6,835 GPU-accelerated data management under the test of time 2020 CIDR 5.7584924e-05
7,424 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.6214566e-05
7,666 Hardware-Oblivious SIMD Parallelism for In-Memory Column-Stores 2020 CIDR 5.57179e-05
7,687 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.5671645e-05
8,394 Topology-aware Parallel Data Processing: Models, Algorithms and Systems at Scale 2020 CIDR 5.4349805e-05
8,958 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.3449654e-05
Previous Page 1 / 1 Next

Semantically Similar Papers