Back to papers
Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks
Summary: Auto-Suggest learns to propose data prep steps by mining notebook-driven data manipulations. Crawls 4M GitHub Jupyter notebooks, replays steps to log inputs/outputs and decisions, using logs to learn data-driven prep recommendations that beat baselines.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h8cce4e8d55ae6cd8
Venue
SIGMOD
Year
2020
Pagerank
8.1751637e-05
Overall Rank
2,642 | 82.25%
DOI
10.1145/3318464.3389738
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{yan_sigmod20,
title = {{Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks}},
author = {Yan, Cong and He, Yeye},
series = {{SIGMOD} '20},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3318464.3389738},
url = {https://dl.acm.org/doi/10.1145/3318464.3389738},
year = {2020}
}
Incoming Citations (Sorted by Pagerank)
Showing 19 of 19 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
1,929
CHORUS: Foundation Models for Unified Data Discovery and Exploration
2024
VLDB
9.3551286e-05
4,235
Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search
2021
VLDB
6.710043e-05
5,100
Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V
2023
VLDB
6.271266e-05
5,168
Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples
2023
VLDB
6.2432932e-05
5,851
HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation
2023
SIGMOD
5.9692535e-05
6,193
Fine-Grained Lineage for Safer Notebook Interactions
2021
VLDB
5.8538103e-05
7,272
Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine
2025
VLDB
5.5668569e-05
7,819
Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table Representations
2024
SIGMOD
5.4460313e-05
8,159
Predicate Pushdown for Data Science Pipelines
2023
SIGMOD
5.3866275e-05
8,360
FEDEX: An Explainability Framework for Data Exploration Steps
2022
VLDB
5.3456868e-05
8,989
Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph
2023
VLDB
5.2418732e-05
10,430
BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree Search
2026
SIGMOD
4.9769913e-05
10,641
Data-Semantics-Aware Recommendation of Diverse Pivot Tables
2026
SIGMOD
4.9769913e-05
10,655
FlowPilot: A Suggestion System for Designing Scientific Workflows
2026
SIGMOD
4.9769913e-05
10,847
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries
2026
VLDB
4.9769913e-05
11,604
Searching Data Lakes for Nested and Joined Data
2024
VLDB
4.9769913e-05
11,634
LucidScript: Bottom-up Standardization for Data Preparation
2024
VLDB
4.9769913e-05
11,811
DataRinse: Semantic Transforms for Data preparation based on Code Mining
2023
VLDB
4.9769913e-05
11,940
Leam: An Interactive System for In-situ Visual Text Analysis
2021
CIDR
4.9769913e-05
Outgoing Citations (Sorted by Pagerank)
Showing 25 of 25 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
86
XMark: A Benchmark for XML Data Management
2002
VLDB
0.00035793969
435
Mining Database Structure; Or, How to Build a Data Quality Browser
2002
SIGMOD
0.0001831946
594
Linear Road: A Stream Data Management Benchmark
2004
VLDB
0.00015816384
862
SnipSuggest: Context-Aware Autocompletion for SQL
2011
VLDB
0.0001339689
884
HoloDetect: Few-Shot Learning for Error Detection
2019
SIGMOD
0.00013263269
972
The Data Civilizer System
2017
CIDR
0.00012757732
1,095
BlinkFill: Semi-supervised Programming By Example for Syntactic String Transformations
2016
VLDB
0.00012048043
1,115
Foofah: Transforming Data By Example
2017
SIGMOD
0.00011958625
1,308
Automating Large-Scale Data Quality Verification
2018
VLDB
0.00011073863
1,337
Harvesting Relational Tables from Lists on the Web
2009
VLDB
0.00010982201
1,344
Detecting Data Errors: Where are we and what needs to be done?
2016
VLDB
0.00010951939
1,533
On Multi-Column Foreign Key Discovery
2010
VLDB
0.00010328081
1,872
Predictive Interaction for Data Transformation
2015
CIDR
9.4606235e-05
2,372
SCODED: Statistical Constraint Oriented Data Error Detection
2020
SIGMOD
8.5575299e-05
2,613
Auto-Join: Joining Tables by Leveraging Transformations
2017
VLDB
8.2229082e-05
2,744
Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations
2018
VLDB
8.064168e-05
2,748
Uni-Detect: A Unified Approach to Automated Error Detection in Tables
2019
SIGMOD
8.0572603e-05
2,962
Auto-Detect: Data-Driven Error Detection in Tables
2018
SIGMOD
7.8038782e-05
3,615
TEGRA: Table Extraction by Global Record Alignment
2015
SIGMOD
7.1558238e-05
3,913
Navigating the Data Lake with DATAMARAN: Automatically Extracting Structure from Log Datasets
2018
SIGMOD
6.9274273e-05
4,204
SEMA-JOIN: Joining Semantically-Related Tables Using Big Table Corpora
2015
VLDB
6.7331936e-05
5,138
Fast Foreign-Key Detection in Microsoft SQL Server PowerPivot for Excel
2014
VLDB
6.2561301e-05
6,015
WADaR: Joint Wrapper and Data Repair
2015
VLDB
5.9107226e-05
7,198
The TEXTURE Benchmark: Measuring Performance of Text Queries on a Relational DBMS
2005
VLDB
5.5870272e-05
8,143
Synthesizing Mapping Relationships Using Table Corpus
2017
SIGMOD
5.3898955e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
12,055
Active Reinforcement Learning for Data Preparation: Learn2Clean with Human-In-The-Loop
2020
CIDR
2
7,273
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework
2025
VLDB
3
10,641
Data-Semantics-Aware Recommendation of Diverse Pivot Tables
2026
SIGMOD
4
8,989
Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph
2023
VLDB
5
4,426
Auto-Transform: Learning-to-Transform by Patterns
2020
VLDB
6
1,293
Finding Related Tables in Data Lakes for Interactive Data Science
2020
SIGMOD
7
10,878
DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
2026
VLDB
8
13,812
Towards Understanding Data Analysis Workflows using a Large Notebook Corpus
2019
SIGMOD
9
4,235
Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search
2021
VLDB
10
7,531
Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence
2025
VLDB