DBScholar

Back to papers

ExDRa: Exploratory Data Science on Federated Raw Data

Summary: ExDRa enables exploratory data science on federated raw data: ad-hoc integration, intermediates reuse, and lifecycle optimization for partially accessible data. It adds a federated SystemDS backend for linear algebra, PS, and data prep to enable enterprise federated ML and privacy-aware data management. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6300
Venue
SIGMOD
Year
2021
Pagerank
5.5671645e-05
Overall Rank
7,687 | 47.27%
DOI
10.1145/3448016.3457549

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{baunsgaard_sigmod21,
        title = {{ExDRa: Exploratory Data Science on Federated Raw Data}},
        author = {Baunsgaard, Sebastian and Boehm, Matthias and Chaudhary, Ankit and Derakhshan, Behrouz and Geißelsöder, Stefan and Grulich, Philipp M. and Hildebrand, Michael and Innerebner, Kevin and Markl, Volker and Neubauer, Claus and Osterburg, Sarah and Ovcharenko, Olga and Redyuk, Sergey and Rieger, Tobias and Mahdiraji, Alireza Rezaei and Wrede, Sebastian Benjamin and Zeuch, Steffen},
        series = {{SIGMOD} '21},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3448016.3457549},
        url = {https://dl.acm.org/doi/10.1145/3448016.3457549},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 7 of 7 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 37 of 37 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
112 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00032801121
155 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028713176
259 Database Cracking 2007 CIDR 0.00023119313
452 An Architecture for Parallel Topic Models 2010 VLDB 0.00018146809
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016217563
640 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.00015409494
946 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013054126
963 The Data Civilizer System 2017 CIDR 0.00012935145
1,004 Democratizing Data Science through Interactive Curation of ML Pipelines 2019 SIGMOD 0.00012701932
1,070 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.0001232307
1,249 Privacy Preserving Vertical Federated Learning for Tree-based Models 2020 VLDB 0.00011495357
1,569 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010335423
1,644 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010132912
1,670 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis 2018 SIGMOD 0.00010045615
1,675 Garlic: A New Flavor of Federated Query Processing for DB2 2002 SIGMOD 0.00010035333
1,750 Lazy Maintenance of Materialized Views 2007 VLDB 9.8392477e-05
1,756 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.8172465e-05
1,927 Here are my Data Files. Here are my Queries. Where are my Results? 2011 CIDR 9.4703074e-05
2,059 LINVIEW: Incremental View Maintenance for Complex Analytical Queries 2014 SIGMOD 9.2471145e-05
2,147 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.0831495e-05
2,151 An Architecture for Recycling Intermediates in a Column-store 2009 SIGMOD 9.0784444e-05
2,162 Heterogeneity-aware Distributed Parameter Servers 2017 SIGMOD 9.0581831e-05
2,208 Query Optimization for Dynamic Imputation 2017 VLDB 8.9512455e-05
2,847 AIDA - Abstraction for Advanced In-Database Analytics 2018 VLDB 8.0583142e-05
3,146 HELIX: Accelerating Human-in-the-loop Machine Learning 2018 VLDB 7.7091591e-05
3,494 WANalytics: Analytics for a Geo-Distributed Data-Intensive World 2015 CIDR 7.3642231e-05
3,638 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2338361e-05
3,947 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 7.0040437e-05
4,240 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.809685e-05
4,510 Building High Throughput Permissioned Blockchain Fabrics: Challenges and Opportunities 2020 VLDB 6.6520543e-05
5,215 The NebulaStream Platform: Data and Application Management for the Internet of Things 2020 CIDR 6.3124972e-05
5,699 Optimizing Machine Learning Workloads in Collaborative Environments 2020 SIGMOD 6.1170243e-05
6,188 "Amnesia" - A Selection of Machine Learning Models That Can Forget User Data Very Fast 2020 CIDR 5.9483532e-05
6,349 Grizzly: Efficient Stream Processing Through Adaptive Query Compilation 2020 SIGMOD 5.9049304e-05
6,485 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 5.8657457e-05
7,775 State of Public and Private Blockchains: Myths and Reality 2019 SIGMOD 5.546517e-05
Previous Page 1 / 1 Next

Semantically Similar Papers