DBScholar

Back to papers

Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks

Summary: Auto-Suggest learns to propose data prep steps by mining notebook-driven data manipulations. Crawls 4M GitHub Jupyter notebooks, replays steps to log inputs/outputs and decisions, using logs to learn data-driven prep recommendations that beat baselines. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h8cce4e8d55ae6cd8
Venue
SIGMOD
Year
2020
Pagerank
8.1787073e-05
Overall Rank
2,641 | 82.25%
DOI
10.1145/3318464.3389738

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{yan_sigmod20,
        title = {{Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks}},
        author = {Yan, Cong and He, Yeye},
        series = {{SIGMOD} '20},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3318464.3389738},
        url = {https://dl.acm.org/doi/10.1145/3318464.3389738},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 19 of 19 citing papers.

Rank Citing Paper Year Venue Pagerank
1,932 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 9.34643e-05
4,235 Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search 2021 VLDB 6.713221e-05
5,097 Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V 2023 VLDB 6.2742361e-05
5,167 Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples 2023 VLDB 6.2462501e-05
5,848 HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation 2023 SIGMOD 5.9720806e-05
6,190 Fine-Grained Lineage for Safer Notebook Interactions 2021 VLDB 5.8565828e-05
7,269 Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine 2025 VLDB 5.5694935e-05
7,812 Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table Representations 2024 SIGMOD 5.4486106e-05
8,153 Predicate Pushdown for Data Science Pipelines 2023 SIGMOD 5.3891786e-05
8,355 FEDEX: An Explainability Framework for Data Exploration Steps 2022 VLDB 5.3482186e-05
8,979 Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph 2023 VLDB 5.2443558e-05
10,418 BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree Search 2026 SIGMOD 4.9793485e-05
10,630 Data-Semantics-Aware Recommendation of Diverse Pivot Tables 2026 SIGMOD 4.9793485e-05
10,644 FlowPilot: A Suggestion System for Designing Scientific Workflows 2026 SIGMOD 4.9793485e-05
10,837 EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries 2026 VLDB 4.9793485e-05
11,598 Searching Data Lakes for Nested and Joined Data 2024 VLDB 4.9793485e-05
11,628 LucidScript: Bottom-up Standardization for Data Preparation 2024 VLDB 4.9793485e-05
11,805 DataRinse: Semantic Transforms for Data preparation based on Code Mining 2023 VLDB 4.9793485e-05
11,934 Leam: An Interactive System for In-situ Visual Text Analysis 2021 CIDR 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 25 of 25 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
86 XMark: A Benchmark for XML Data Management 2002 VLDB 0.00035810622
435 Mining Database Structure; Or, How to Build a Data Quality Browser 2002 SIGMOD 0.0001832766
594 Linear Road: A Stream Data Management Benchmark 2004 VLDB 0.00015823573
861 SnipSuggest: Context-Aware Autocompletion for SQL 2011 VLDB 0.00013402933
883 HoloDetect: Few-Shot Learning for Error Detection 2019 SIGMOD 0.00013268059
972 The Data Civilizer System 2017 CIDR 0.00012763234
1,094 BlinkFill: Semi-supervised Programming By Example for Syntactic String Transformations 2016 VLDB 0.0001205355
1,115 Foofah: Transforming Data By Example 2017 SIGMOD 0.00011964111
1,308 Automating Large-Scale Data Quality Verification 2018 VLDB 0.0001107886
1,337 Harvesting Relational Tables from Lists on the Web 2009 VLDB 0.00010987014
1,344 Detecting Data Errors: Where are we and what needs to be done? 2016 VLDB 0.00010956518
1,532 On Multi-Column Foreign Key Discovery 2010 VLDB 0.00010332696
1,871 Predictive Interaction for Data Transformation 2015 CIDR 9.464979e-05
2,371 SCODED: Statistical Constraint Oriented Data Error Detection 2020 SIGMOD 8.5615698e-05
2,612 Auto-Join: Joining Tables by Leveraging Transformations 2017 VLDB 8.2265197e-05
2,745 Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations 2018 VLDB 8.0668395e-05
2,748 Uni-Detect: A Unified Approach to Automated Error Detection in Tables 2019 SIGMOD 8.0610697e-05
2,960 Auto-Detect: Data-Driven Error Detection in Tables 2018 SIGMOD 7.8075103e-05
3,615 TEGRA: Table Extraction by Global Record Alignment 2015 SIGMOD 7.1591776e-05
3,912 Navigating the Data Lake with DATAMARAN: Automatically Extracting Structure from Log Datasets 2018 SIGMOD 6.9306061e-05
4,206 SEMA-JOIN: Joining Semantically-Related Tables Using Big Table Corpora 2015 VLDB 6.7348167e-05
5,135 Fast Foreign-Key Detection in Microsoft SQL Server PowerPivot for Excel 2014 VLDB 6.2587553e-05
6,015 WADaR: Joint Wrapper and Data Repair 2015 VLDB 5.9135138e-05
7,196 The TEXTURE Benchmark: Measuring Performance of Text Queries on a Relational DBMS 2005 VLDB 5.5896469e-05
8,144 Synthesizing Mapping Relationships Using Table Corpus 2017 SIGMOD 5.3910618e-05
Previous Page 1 / 1 Next

Semantically Similar Papers