Back to papers
Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks
Summary: Auto-Suggest learns to propose data prep steps by mining notebook-driven data manipulations. Crawls 4M GitHub Jupyter notebooks, replays steps to log inputs/outputs and decisions, using logs to learn data-driven prep recommendations that beat baselines.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h8cce4e8d55ae6cd8
Venue
SIGMOD
Year
2020
Pagerank
8.1787073e-05
Overall Rank
2,641 | 82.25%
DOI
10.1145/3318464.3389738
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{yan_sigmod20,
title = {{Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks}},
author = {Yan, Cong and He, Yeye},
series = {{SIGMOD} '20},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3318464.3389738},
url = {https://dl.acm.org/doi/10.1145/3318464.3389738},
year = {2020}
}
Incoming Citations (Sorted by Pagerank)
Showing 19 of 19 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
1,932
CHORUS: Foundation Models for Unified Data Discovery and Exploration
2024
VLDB
9.34643e-05
4,235
Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search
2021
VLDB
6.713221e-05
5,097
Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V
2023
VLDB
6.2742361e-05
5,167
Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples
2023
VLDB
6.2462501e-05
5,848
HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation
2023
SIGMOD
5.9720806e-05
6,190
Fine-Grained Lineage for Safer Notebook Interactions
2021
VLDB
5.8565828e-05
7,269
Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine
2025
VLDB
5.5694935e-05
7,812
Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table Representations
2024
SIGMOD
5.4486106e-05
8,153
Predicate Pushdown for Data Science Pipelines
2023
SIGMOD
5.3891786e-05
8,355
FEDEX: An Explainability Framework for Data Exploration Steps
2022
VLDB
5.3482186e-05
8,979
Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph
2023
VLDB
5.2443558e-05
10,418
BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree Search
2026
SIGMOD
4.9793485e-05
10,630
Data-Semantics-Aware Recommendation of Diverse Pivot Tables
2026
SIGMOD
4.9793485e-05
10,644
FlowPilot: A Suggestion System for Designing Scientific Workflows
2026
SIGMOD
4.9793485e-05
10,837
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries
2026
VLDB
4.9793485e-05
11,598
Searching Data Lakes for Nested and Joined Data
2024
VLDB
4.9793485e-05
11,628
LucidScript: Bottom-up Standardization for Data Preparation
2024
VLDB
4.9793485e-05
11,805
DataRinse: Semantic Transforms for Data preparation based on Code Mining
2023
VLDB
4.9793485e-05
11,934
Leam: An Interactive System for In-situ Visual Text Analysis
2021
CIDR
4.9793485e-05
Outgoing Citations (Sorted by Pagerank)
Showing 25 of 25 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
86
XMark: A Benchmark for XML Data Management
2002
VLDB
0.00035810622
435
Mining Database Structure; Or, How to Build a Data Quality Browser
2002
SIGMOD
0.0001832766
594
Linear Road: A Stream Data Management Benchmark
2004
VLDB
0.00015823573
861
SnipSuggest: Context-Aware Autocompletion for SQL
2011
VLDB
0.00013402933
883
HoloDetect: Few-Shot Learning for Error Detection
2019
SIGMOD
0.00013268059
972
The Data Civilizer System
2017
CIDR
0.00012763234
1,094
BlinkFill: Semi-supervised Programming By Example for Syntactic String Transformations
2016
VLDB
0.0001205355
1,115
Foofah: Transforming Data By Example
2017
SIGMOD
0.00011964111
1,308
Automating Large-Scale Data Quality Verification
2018
VLDB
0.0001107886
1,337
Harvesting Relational Tables from Lists on the Web
2009
VLDB
0.00010987014
1,344
Detecting Data Errors: Where are we and what needs to be done?
2016
VLDB
0.00010956518
1,532
On Multi-Column Foreign Key Discovery
2010
VLDB
0.00010332696
1,871
Predictive Interaction for Data Transformation
2015
CIDR
9.464979e-05
2,371
SCODED: Statistical Constraint Oriented Data Error Detection
2020
SIGMOD
8.5615698e-05
2,612
Auto-Join: Joining Tables by Leveraging Transformations
2017
VLDB
8.2265197e-05
2,745
Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations
2018
VLDB
8.0668395e-05
2,748
Uni-Detect: A Unified Approach to Automated Error Detection in Tables
2019
SIGMOD
8.0610697e-05
2,960
Auto-Detect: Data-Driven Error Detection in Tables
2018
SIGMOD
7.8075103e-05
3,615
TEGRA: Table Extraction by Global Record Alignment
2015
SIGMOD
7.1591776e-05
3,912
Navigating the Data Lake with DATAMARAN: Automatically Extracting Structure from Log Datasets
2018
SIGMOD
6.9306061e-05
4,206
SEMA-JOIN: Joining Semantically-Related Tables Using Big Table Corpora
2015
VLDB
6.7348167e-05
5,135
Fast Foreign-Key Detection in Microsoft SQL Server PowerPivot for Excel
2014
VLDB
6.2587553e-05
6,015
WADaR: Joint Wrapper and Data Repair
2015
VLDB
5.9135138e-05
7,196
The TEXTURE Benchmark: Measuring Performance of Text Queries on a Relational DBMS
2005
VLDB
5.5896469e-05
8,144
Synthesizing Mapping Relationships Using Table Corpus
2017
SIGMOD
5.3910618e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
12,049
Active Reinforcement Learning for Data Preparation: Learn2Clean with Human-In-The-Loop
2020
CIDR
2
7,270
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework
2025
VLDB
3
10,630
Data-Semantics-Aware Recommendation of Diverse Pivot Tables
2026
SIGMOD
4
8,979
Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph
2023
VLDB
5
4,424
Auto-Transform: Learning-to-Transform by Patterns
2020
VLDB
6
1,293
Finding Related Tables in Data Lakes for Interactive Data Science
2020
SIGMOD
7
10,869
DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
2026
VLDB
8
13,807
Towards Understanding Data Analysis Workflows using a Large Notebook Corpus
2019
SIGMOD
9
4,235
Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search
2021
VLDB
10
7,526
Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence
2025
VLDB