DBScholar

Back to papers

DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines

Summary: DAPHNE: extensible infra unifying DM, HPC and ML pipelines via shared language abstractions, compiler/runtime integration and multi-level scheduling. Key novelty: vectorized engine for computational storage and accelerators to cut data-movement/format overheads for local and distributed ops; prelim results show notable speedups. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
he2e1ccb5f2480d87
Venue
CIDR
Year
2022
Pagerank
5.6552024e-05
Overall Rank
6,883 | 53.74%
DOI
-
PDF
Download (CC BY 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{damme_cidr22,
        address = {Amsterdam, Netherlands},
        series = {{CIDR} '22},
        title = {{DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines}},
        booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
        author = {Damme, Patrick and Birkenbach, Marius and Bitsakos, Constantinos and Boehm, Matthias and Bonnet, Philippe and Ciorba, Florina and Dokter, Mark and Dowgiallo, Pawel and Eleliemy, Ahmed and Faerber, Christian and Goumas, Georgios and Habich, Dirk and Hedam, Niclas and Hofer, Marlies and Huang, Wenjun and Innerebner, Kevin and Karakostas, Vasileios and Kern, Roman and Kosar, Tomaž and Krause, Alexander and Krems, Daniel and Laber, Andreas and Lehner, Wolfgang and Mier, Eric and Paradies, Marcus and Peischl, Bernhard and Poerwawinata, Gabrielle and Psomadakis, Stratos and Rabl, Tilmann and Ratuszniak, Piotr and Silva, Pedro and Skuppin, Nikolai and Starzacher, Andreas and Steinwender, Benjamin and Tolovski, Ilin and Tözün, Pınar and Ulatowski, Wojciech and Wang, Yuanyuan and Wrosz, Izajasz and Zamuda, Aleš and Zhang, Ce and Zhu, Xiao Xiang},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 37 of 37 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
3 Pregel: A System for Large-Scale Graph Processing 2010 SIGMOD 0.0012087459
14 MonetDB/X100: Hyper-Pipelining Query Execution 2005 CIDR 0.00064013679
52 The Snowflake Elastic Data Warehouse 2016 SIGMOD 0.00041210636
215 Morsel-Driven Parallelism: A NUMA-Aware Query Evaluation Framework for the Many-Core Age 2014 SIGMOD 0.00024589307
503 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017195428
537 MLbase: A Distributed Machine-learning System 2013 CIDR 0.00016762761
1,018 Designing And Mining Multi-Terabyte Astronomy Archives: The Sloan Digital Sky Survey 2000 SIGMOD 0.00012460728
1,082 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012118261
1,223 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011468426
1,371 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.00010894415
1,669 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9324573e-05
1,686 Garlic: A New Flavor of Federated Query Processing for DB2 2002 SIGMOD 9.8662635e-05
1,803 Tuplex: Data Science in Python at Native Code Speed 2021 SIGMOD 9.602292e-05
1,825 Data Management for Data Science: Towards Embedded Analytics 2020 CIDR 9.5558093e-05
2,197 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 8.8740089e-05
2,241 Query Optimization for Dynamic Imputation 2017 VLDB 8.7704255e-05
2,287 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.691301e-05
2,329 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.6268411e-05
2,989 Adapting to Source Properties in Processing Data Integration Queries 2004 SIGMOD 7.7736772e-05
3,103 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6456038e-05
3,955 Tensors: An abstraction for general data processing 2021 VLDB 6.8994712e-05
4,123 The bionic DBMS is coming, but what will it look like? 2013 CIDR 6.7901576e-05
4,209 Accelerating Queries with Group-By and Join by Groupjoin 2011 VLDB 6.7307652e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6537801e-05
4,417 Estimating Compilation Time of a Query Optimizer 2003 SIGMOD 6.6064397e-05
4,610 Not your Grandpa's SSD: The Era of Co-Designed Storage Devices 2021 SIGMOD 6.5005576e-05
4,816 Computational Storage: Where Are We Today? 2021 CIDR 6.39818e-05
5,199 QuERy: A Framework for Integrating Entity Resolution with Query Processing 2016 VLDB 6.2309669e-05
5,944 Bridging Two Worlds with RICE: Integrating R into the SAP In-Memory Computing Engine 2011 VLDB 5.9360073e-05
5,999 BAGUA: Scaling up Distributed Learning with System Relaxations 2022 VLDB 5.9170896e-05
6,518 GPU-accelerated data management under the test of time 2020 CIDR 5.7561332e-05
6,662 A Morsel-Driven Query Execution Engine for Heterogeneous Multi-Cores 2019 VLDB 5.7155423e-05
7,418 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.5306468e-05
7,735 Hardware-Oblivious SIMD Parallelism for In-Memory Column-Stores 2020 CIDR 5.4632401e-05
7,843 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4406331e-05
8,557 Topology-aware Parallel Data Processing: Models, Algorithms and Systems at Scale 2020 CIDR 5.314366e-05
9,042 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.2310219e-05
Previous Page 1 / 1 Next

Semantically Similar Papers