DBScholar

Back to papers

Lachesis: Automatic Partitioning for UDF-Centric Analytics

Summary: Lachesis automates partitioning for UDF-centric analytics, where reusable computation and functional dependencies are scarce. It models workloads as analyzable workflow sub-computations and uses deep reinforcement learning to select partitioning features, optimizing shared storage across applications. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hf0140d5c584947a6
Venue
VLDB
Year
2021
Pagerank
5.5355684e-05
Overall Rank
7,388 | 50.35%
DOI
10.14778/3457390.3457392
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{zou_vldb21,
        title = {{Lachesis: Automatic Partitioning for UDF-Centric Analytics}},
        author = {Zou, Jia and Das, Amitabh and Barhate, Pratik and Iyengar, Arun and Yuan, Binhang and Jankov, Dimitrije and Jermaine, Chris},
        journal = {PVLDB},
        series = {{VLDB} '21},
        volume = {14},
        number = {8},
        pages = {1262--1275},
        doi = {10.14778/3457390.3457392},
        url = {https://doi.org/10.14778/3457390.3457392},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 32 of 32 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049821554
195 Integrating Vertical and Horizontal Partitioning into Automated Physical Database Design 2004 SIGMOD 0.00025619089
243 Automating Physical Database Design in a Parallel Database 2002 SIGMOD 0.00023349603
252 Database Cracking 2007 CIDR 0.00023101361
416 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.00018650998
495 Building Efficient Query Engines in a High-Level Language 2014 VLDB 0.00017363171
675 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00014880686
896 Froid: Optimization of Imperative Programs in a Relational Database 2018 VLDB 0.00013203085
1,073 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012163258
1,193 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.00011582619
1,884 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.4349703e-05
2,125 Tupleware: "Big" Data, Big Analytics, Small Clusters 2015 CIDR 9.0017828e-05
2,266 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.7248802e-05
2,476 Learning a Partitioning Advisor for Cloud Databases 2020 SIGMOD 8.407183e-05
2,688 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1262948e-05
2,704 Supporting Table Partitioning By Reference in Oracle 2008 SIGMOD 8.1112208e-05
2,753 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0513412e-05
3,241 Locality-aware Partitioning in Parallel Database Systems 2015 SIGMOD 7.4972383e-05
3,440 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.2954357e-05
3,455 Scaling Spark in the Real World: Performance and Usability 2015 VLDB 7.2850317e-05
3,544 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2108612e-05
3,725 Schema Management for Document Stores 2015 VLDB 7.067273e-05
4,288 Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics 2015 VLDB 6.6823776e-05
4,534 AdaptDB: Adaptive Partitioning for Distributed Joins 2017 VLDB 6.5535468e-05
5,292 Tensor Relational Algebra for Distributed Machine Learning System Design 2021 VLDB 6.1911633e-05
5,668 Skew-Aware Join Optimization for Array Databases 2015 SIGMOD 6.0433507e-05
6,725 Near-Optimal Distributed Band-Joins through Recursive Partitioning 2020 SIGMOD 5.6954425e-05
6,967 Incremental Elasticity For Array Databases 2014 SIGMOD 5.6301184e-05
7,638 MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs 2014 VLDB 5.4767699e-05
8,171 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 5.3826392e-05
8,226 Customizable and Scalable Fuzzy Join for Big Data 2019 VLDB 5.3736355e-05
9,753 PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development 2018 SIGMOD 5.1325223e-05
Previous Page 1 / 1 Next

Semantically Similar Papers