Back to papers
Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks
Summary: Auto-Suggest learns to propose data prep steps by mining notebook-driven data manipulations. Crawls 4M GitHub Jupyter notebooks, replays steps to log inputs/outputs and decisions, using logs to learn data-driven prep recommendations that beat baselines.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
6015
Venue
SIGMOD
Year
2020
Pagerank
8.1781662e-05
Overall Rank
2,744 | 81.18%
DOI
10.1145/3318464.3389738
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{yan_sigmod20,
title = {{Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks}},
author = {Yan, Cong and He, Yeye},
series = {{SIGMOD} '20},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3318464.3389738},
url = {https://dl.acm.org/doi/10.1145/3318464.3389738},
year = {2020}
}
Incoming Citations (Sorted by Pagerank)
Showing 18 of 18 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
2,242
CHORUS: Foundation Models for Unified Data Discovery and Exploration
2024
VLDB
8.8823802e-05
4,875
Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search
2021
VLDB
6.4689177e-05
4,990
Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V
2023
VLDB
6.4105738e-05
5,403
Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples
2023
VLDB
6.2306083e-05
6,071
Fine-Grained Lineage for Safer Notebook Interactions
2021
VLDB
5.9873765e-05
7,502
Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine
2025
VLDB
5.6029996e-05
8,177
HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation
2023
SIGMOD
5.4730821e-05
8,184
FEDEX: An Explainability Framework for Data Exploration Steps
2022
VLDB
5.4709725e-05
8,465
Predicate Pushdown for Data Science Pipelines
2023
SIGMOD
5.4194578e-05
8,797
Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table Representations
2024
SIGMOD
5.3702298e-05
9,629
Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph
2023
VLDB
5.2434488e-05
10,202
BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree Search
2026
SIGMOD
5.093636e-05
10,441
Data-Semantics-Aware Recommendation of Diverse Pivot Tables
2026
SIGMOD
5.093636e-05
10,457
FlowPilot: A Suggestion System for Designing Scientific Workflows
2026
SIGMOD
5.093636e-05
11,270
Searching Data Lakes for Nested and Joined Data
2024
VLDB
5.093636e-05
11,309
LucidScript: Bottom-up Standardization for Data Preparation
2024
VLDB
5.093636e-05
11,496
DataRinse: Semantic Transforms for Data preparation based on Code Mining
2023
VLDB
5.093636e-05
11,627
Leam: An Interactive System for In-situ Visual Text Analysis
2021
CIDR
5.093636e-05
Outgoing Citations (Sorted by Pagerank)
Showing 25 of 25 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
79
XMark: A Benchmark for XML Data Management
2002
VLDB
0.00036555628
432
Mining Database Structure; Or, How to Build a Data Quality Browser
2002
SIGMOD
0.00018572055
587
Linear Road: A Stream Data Management Benchmark
2004
VLDB
0.00016106078
846
SnipSuggest: Context-Aware Autocompletion for SQL
2011
VLDB
0.00013658824
946
HoloDetect: Few-Shot Learning for Error Detection
2019
SIGMOD
0.00013054126
963
The Data Civilizer System
2017
CIDR
0.00012935145
1,160
BlinkFill: Semi-supervised Programming By Example for Syntactic String Transformations
2016
VLDB
0.00011883163
1,165
Foofah: Transforming Data By Example
2017
SIGMOD
0.00011860616
1,316
Harvesting Relational Tables from Lists on the Web
2009
VLDB
0.00011181216
1,350
Automating Large-Scale Data Quality Verification
2018
VLDB
0.00011065626
1,351
Detecting Data Errors: Where are we and what needs to be done?
2016
VLDB
0.00011064851
1,521
On Multi-Column Foreign Key Discovery
2010
VLDB
0.00010506299
1,866
Predictive Interaction for Data Transformation
2015
CIDR
9.5914583e-05
2,728
Uni-Detect: A Unified Approach to Automated Error Detection in Tables
2019
SIGMOD
8.1995954e-05
2,978
Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations
2018
VLDB
7.9030989e-05
2,980
Auto-Join: Joining Tables by Leveraging Transformations
2017
VLDB
7.8975073e-05
3,004
SCODED: Statistical Constraint Oriented Data Error Detection
2020
SIGMOD
7.8608629e-05
3,147
Auto-Detect: Data-Driven Error Detection in Tables
2018
SIGMOD
7.7077175e-05
3,574
TEGRA: Table Extraction by Global Record Alignment
2015
SIGMOD
7.2958137e-05
3,894
Navigating the Data Lake with DATAMARAN: Automatically Extracting Structure from Log Datasets
2018
SIGMOD
7.0417782e-05
4,525
SEMA-JOIN: Joining Semantically-Related Tables Using Big Table Corpora
2015
VLDB
6.6456999e-05
5,099
Fast Foreign-Key Detection in Microsoft SQL Server PowerPivot for Excel
2014
VLDB
6.3646928e-05
5,915
WADaR: Joint Wrapper and Data Repair
2015
VLDB
6.0437843e-05
7,071
The TEXTURE Benchmark: Measuring Performance of Text Queries on a Relational DBMS
2005
VLDB
5.711485e-05
8,491
Synthesizing Mapping Relationships Using Table Corpus
2017
SIGMOD
5.414689e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
4,196
Automatic Example Queries for Ad Hoc Databases
2011
SIGMOD
2
11,746
Active Reinforcement Learning for Data Preparation: Learn2Clean with Human-In-The-Loop
2020
CIDR
3
10,931
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework
2025
VLDB
4
10,441
Data-Semantics-Aware Recommendation of Diverse Pivot Tables
2026
SIGMOD
5
9,629
Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph
2023
VLDB
6
4,412
Auto-Transform: Learning-to-Transform by Patterns
2020
VLDB
7
1,303
Finding Related Tables in Data Lakes for Interactive Data Science
2020
SIGMOD
8
13,493
Towards Understanding Data Analysis Workflows using a Large Notebook Corpus
2019
SIGMOD
9
4,875
Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search
2021
VLDB
10
9,380
Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence
2025
VLDB