The VADA Architecture for Cost-Effective Data Wrangling
Summary: VADA is an extensible data-wrangling architecture that orchestrates extraction, cleaning, and integration via components guided by domain data. It enables feedback-driven refinement and user-priority tradeoffs, delivering results with less config. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Nikolaos Konstantinou (University of Manchester)
- 2. Martin Koehler (University of Manchester)
- 3. Edward Abel (University of Manchester)
- 4. Cristina Civili (University of Edinburgh)
- 5. Bernd Neumayr (University of Oxford)
- 6. Emanuel Sallinger (University of Oxford)
- 7. Alvaro A.A. Fernandes (University of Manchester)
- 8. Georg Gottlob (University of Oxford)
- 9. John A. Keane (University of Manchester)
- 10. Leonid Libkin (University of Edinburgh)
- 11. Norman W. Paton (University of Manchester)
BibTeX Citation
@inproceedings{konstantinou_sigmod17,
title = {{The VADA Architecture for Cost-Effective Data Wrangling}},
author = {Konstantinou, Nikolaos and Koehler, Martin and Abel, Edward and Civili, Cristina and Neumayr, Bernd and Sallinger, Emanuel and Fernandes, Alvaro A.A. and Gottlob, Georg and Keane, John A. and Libkin, Leonid and Paton, Norman W.},
series = {{SIGMOD} '17},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3035918.3058730},
url = {https://dl.acm.org/doi/10.1145/3035918.3058730},
year = {2017}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,050 | Capturing and Querying Fine-grained Provenance of Preprocessing Pipelines in Data Science | 2021 | VLDB | 5.5000099e-05 |
| 9,790 | Meta-Mappings for Schema Mapping Reuse | 2019 | VLDB | 5.220833e-05 |
| 9,850 | Materializing Knowledge Bases via Trigger Graphs | 2021 | VLDB | 5.2094004e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 963 | The Data Civilizer System | 2017 | CIDR | 0.00012935145 |
| 1,159 | Relational Transducers for Electronic Commerce | 1998 | PODS | 0.00011898742 |
| 1,947 | Data Wrangling: The Challenging Journey from the Wild to the Lake | 2015 | CIDR | 9.4326202e-05 |
| 2,913 | Constance: An Intelligent Data Lake System | 2016 | SIGMOD | 7.9684737e-05 |
| 4,201 | CLAMS: Bringing Quality to Data Lakes | 2016 | SIGMOD | 6.8355878e-05 |
| 4,546 | DataXFormer: An Interactive Data Transformation Tool | 2015 | SIGMOD | 6.6374215e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,978 | From Auto-tuning One Size Fits All to Self-designed and Learned Data-intensive Systems | 2019 | SIGMOD |
| 2 | 9,742 | Unified Data Analytics: State-of-the-art and Open Problems | 2022 | VLDB |
| 3 | 11,499 | Towards Auto-Generated Data Systems | 2023 | VLDB |
| 4 | 13,435 | Data Cleaning in the Era of Data Science: Challenges and Opportunities | 2021 | CIDR |
| 5 | 8,620 | Wisteria: Nurturing Scalable Data Cleaning Infrastructure | 2015 | VLDB |
| 6 | 7,878 | Dependency-Driven Analytics: a Compass for Uncharted Data Oceans | 2017 | CIDR |
| 7 | 10,131 | Towards Scalable Visual Data Wrangling via Direct Manipulation | 2026 | CIDR |
| 8 | 11,487 | To UDFs and Beyond: Demonstration of a Fully Decomposed Data Processor for General Data Wrangling Tasks | 2023 | VLDB |
| 9 | 6,181 | Just-In-Time Data Virtualization: Lightweight Data Management with ViDa | 2015 | CIDR |
| 10 | 1,947 | Data Wrangling: The Challenging Journey from the Wild to the Lake | 2015 | CIDR |