LucidScript: Bottom-up Standardization for Data Preparation
Summary: Bottom-up standardization that treats a general-purpose data-prep script as a sketch of intent and rewrites it into a simpler, standardized, analyzable form. LucidScript implements this to improve readability, verification, and engineering/statistical hygiene. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Eugenie Lai (Massachusetts Institute of Technology)
- 2. Yuze Lou (University of Michigan)
- 3. Brit Youngmann (Technion)
- 4. Michael Cafarella (Massachusetts Institute of Technology)
BibTeX Citation
@article{lai_vldb24,
title = {{LucidScript: Bottom-up Standardization for Data Preparation}},
author = {Lai, Eugenie and Lou, Yuze and Youngmann, Brit and Cafarella, Michael},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {12},
pages = {4317--4320},
doi = {10.14778/3685800.3685864},
url = {https://doi.org/10.14778/3685800.3685864},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,744 | Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks | 2020 | SIGMOD | 8.1781662e-05 |
| 4,412 | Auto-Transform: Learning-to-Transform by Patterns | 2020 | VLDB | 6.7168614e-05 |
| 4,875 | Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search | 2021 | VLDB | 6.4689177e-05 |
| 6,250 | Lightweight Inspection of Data Preprocessing in Native Machine Learning Pipelines | 2021 | CIDR | 5.9418015e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,072 | Lux: Always-on Visualization Recommendations for Exploratory Dataframe Workflows | 2022 | VLDB |
| 2 | 1,866 | Predictive Interaction for Data Transformation | 2015 | CIDR |
| 3 | 11,713 | From Papers to Practice: The openclean Open-Source Data Cleaning Library | 2021 | VLDB |
| 4 | 1,030 | Data-Driven Understanding and Refinement of Schema Mappings | 2001 | SIGMOD |
| 5 | 5,620 | DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python | 2021 | SIGMOD |
| 6 | 9,931 | C2Metadata: Automating the Capture of Data Transformations from Statistical Scripts in Data Documentation | 2019 | SIGMOD |
| 7 | 12,088 | Synthesizing Data Programs | 2015 | CIDR |
| 8 | 7,894 | A Grammar-based Entity Representation Framework for Data Cleaning | 2009 | SIGMOD |
| 9 | 2,744 | Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks | 2020 | SIGMOD |
| 10 | 11,496 | DataRinse: Semantic Transforms for Data preparation based on Code Mining | 2023 | VLDB |