APEX-DAG: Library and Language independent Pipeline EXtraction
Summary: APEX-DAG statically extracts dataflow, transformations, and dependencies from notebooks/scripts, avoiding execution or code changes. A graph-attention model enables pipeline discovery across languages and libraries, including previously unseen ones after training. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Sebastian Eggers (Berlin Institute for the Foundations of Learning and Data; Technical University of Berlin)
- 2. Nina Żukowska (Berlin Institute for the Foundations of Learning and Data; Technical University of Berlin)
- 3. Ziawasch Abedjan (Berlin Institute for the Foundations of Learning and Data; Technical University of Berlin)
BibTeX Citation
@article{eggers_vldb25,
title = {{APEX-DAG: Library and Language independent Pipeline EXtraction}},
author = {Eggers, Sebastian and Żukowska, Nina and Abedjan, Ziawasch},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {12},
pages = {5375--5378},
doi = {10.14778/3750601.3750675},
url = {https://doi.org/10.14778/3750601.3750675},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,303 | Finding Related Tables in Data Lakes for Interactive Data Science | 2020 | SIGMOD | 0.0001123653 |
| 2,431 | noWorkflow: a Tool for Collecting, Analyzing, and Managing Provenance from Python Scripts | 2017 | VLDB | 8.5875984e-05 |
| 2,657 | Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities | 2021 | SIGMOD | 8.2887895e-05 |
| 4,518 | MLINSPECT: A Data Distribution Debugger for Machine Learning Pipelines | 2021 | SIGMOD | 6.6474737e-05 |
| 7,065 | Dataset Relationship Management | 2019 | CIDR | 5.7123431e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,245 | Towards Observability for Production Machine Learning Pipelines | 2022 | VLDB |
| 2 | 8,914 | CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning | 2024 | SIGMOD |
| 3 | 7,395 | Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines | 2023 | SIGMOD |
| 4 | 2,657 | Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities | 2021 | SIGMOD |
| 5 | 11,512 | Towards Observability for Machine Learning Pipelines | 2022 | CIDR |
| 6 | 11,509 | Screening Native ML Pipelines with “ArgusEyes” | 2022 | CIDR |
| 7 | 5,901 | A Scalable AutoML Approach Based on Graph Neural Networks | 2022 | VLDB |
| 8 | 4,875 | Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search | 2021 | VLDB |
| 9 | 4,518 | MLINSPECT: A Data Distribution Debugger for Machine Learning Pipelines | 2021 | SIGMOD |
| 10 | 6,250 | Lightweight Inspection of Data Preprocessing in Native Machine Learning Pipelines | 2021 | CIDR |