DBScholar

Back to papers

Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search

Summary: Auto-Pipeline synthesizes multi-step data-cleaning pipelines from input tables and a target schema, rather than examples. It exploits implicit constraints (FDs, keys) to guide reinforcement learning and search, recovering ~70% of real pipelines up to 10 steps. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
12619
Venue
VLDB
Year
2021
Pagerank
6.4689177e-05
Overall Rank
4,875 | 66.56%
DOI
10.14778/3476249.3476303

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{yang_vldb21,
        title = {{Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search}},
        author = {Yang, Junwen and He, Yeye and Chaudhuri, Surajit},
        journal = {PVLDB},
        series = {{VLDB} '21},
        volume = {14},
        number = {11},
        pages = {2563--2575},
        doi = {10.14778/3476249.3476303},
        url = {https://doi.org/10.14778/3476249.3476303},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 8 of 8 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 12 of 12 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers