DBScholar

Back to papers

ExDRa: Exploratory Data Science on Federated Raw Data

Summary: ExDRa enables exploratory data science on federated raw data: ad-hoc integration, intermediates reuse, and lifecycle optimization for partially accessible data. It adds a federated SystemDS backend for linear algebra, PS, and data prep to enable enterprise federated ML and privacy-aware data management. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h91e6339f1b24baee
Venue
SIGMOD
Year
2021
Pagerank
5.4406331e-05
Overall Rank
7,843 | 47.29%
DOI
10.1145/3448016.3457549

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{baunsgaard_sigmod21,
        title = {{ExDRa: Exploratory Data Science on Federated Raw Data}},
        author = {Baunsgaard, Sebastian and Boehm, Matthias and Chaudhary, Ankit and Derakhshan, Behrouz and Geißelsöder, Stefan and Grulich, Philipp M. and Hildebrand, Michael and Innerebner, Kevin and Markl, Volker and Neubauer, Claus and Osterburg, Sarah and Ovcharenko, Olga and Redyuk, Sergey and Rieger, Tobias and Mahdiraji, Alireza Rezaei and Wrede, Sebastian Benjamin and Zeuch, Steffen},
        series = {{SIGMOD} '21},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3448016.3457549},
        url = {https://dl.acm.org/doi/10.1145/3448016.3457549},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 7 of 7 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 37 of 37 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055384955
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033676943
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028568843
252 Database Cracking 2007 CIDR 0.00023101361
464 An Architecture for Parallel Topic Models 2010 VLDB 0.00017783871
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016083582
654 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.00015096817
884 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013263269
972 The Data Civilizer System 2017 CIDR 0.00012757732
976 Democratizing Data Science through Interactive Curation of ML Pipelines 2019 SIGMOD 0.0001274453
1,070 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.00012179575
1,278 Privacy Preserving Vertical Federated Learning for Tree-based Models 2020 VLDB 0.00011232568
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010208225
1,614 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010067153
1,669 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9324573e-05
1,686 Garlic: A New Flavor of Federated Query Processing for DB2 2002 SIGMOD 9.8662635e-05
1,692 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis 2018 SIGMOD 9.8526476e-05
1,780 Lazy Maintenance of Materialized Views 2007 VLDB 9.654014e-05
1,848 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.5075544e-05
1,927 Here are my Data Files. Here are my Queries. Where are my Results? 2011 CIDR 9.362697e-05
2,086 An Architecture for Recycling Intermediates in a Column-store 2009 SIGMOD 9.0677901e-05
2,087 LINVIEW: Incremental View Maintenance for Complex Analytical Queries 2014 SIGMOD 9.0661817e-05
2,186 Heterogeneity-aware Distributed Parameter Servers 2017 SIGMOD 8.8916253e-05
2,241 Query Optimization for Dynamic Imputation 2017 VLDB 8.7704255e-05
2,778 AIDA - Abstraction for Advanced In-Database Analytics 2018 VLDB 8.0261532e-05
3,168 HELIX: Accelerating Human-in-the-loop Machine Learning 2018 VLDB 7.5692753e-05
3,505 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2474175e-05
3,558 WANalytics: Analytics for a Geo-Distributed Data-Intensive World 2015 CIDR 7.2050361e-05
3,998 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 6.862274e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6537801e-05
4,613 Building High Throughput Permissioned Blockchain Fabrics: Challenges and Opportunities 2020 VLDB 6.4997362e-05
5,349 The NebulaStream Platform: Data and Application Management for the Internet of Things 2020 CIDR 6.1679453e-05
5,794 Optimizing Machine Learning Workloads in Collaborative Environments 2020 SIGMOD 5.9906542e-05
6,325 "Amnesia" - A Selection of Machine Learning Models That Can Forget User Data Very Fast 2020 CIDR 5.8121503e-05
6,444 Grizzly: Efficient Stream Processing Through Adaptive Query Compilation 2020 SIGMOD 5.7807677e-05
6,594 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 5.7394502e-05
7,942 State of Public and Private Blockchains: Myths and Reality 2019 SIGMOD 5.4198405e-05
Previous Page 1 / 1 Next

Semantically Similar Papers