DBScholar

Back to papers

Lachesis: Automatic Partitioning for UDF-Centric Analytics

Summary: Lachesis automates partitioning for UDF-centric analytics, where reusable computation and functional dependencies are scarce. It models workloads as analyzable workflow sub-computations and uses deep reinforcement learning to select partitioning features, optimizing shared storage across applications. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hf0140d5c584947a6
Venue
VLDB
Year
2021
Pagerank
5.538189e-05
Overall Rank
7,386 | 50.35%
DOI
10.14778/3457390.3457392

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{zou_vldb21,
        title = {{Lachesis: Automatic Partitioning for UDF-Centric Analytics}},
        author = {Zou, Jia and Das, Amitabh and Barhate, Pratik and Iyengar, Arun and Yuan, Binhang and Jankov, Dimitrije and Jermaine, Chris},
        journal = {PVLDB},
        series = {{VLDB} '21},
        volume = {14},
        number = {8},
        pages = {1262--1275},
        doi = {10.14778/3457390.3457392},
        url = {https://doi.org/10.14778/3457390.3457392},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 32 of 32 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
195 Integrating Vertical and Horizontal Partitioning into Automated Physical Database Design 2004 SIGMOD 0.00025628849
243 Automating Physical Database Design in a Parallel Database 2002 SIGMOD 0.00023358891
253 Database Cracking 2007 CIDR 0.00023042111
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001865959
495 Building Efficient Query Engines in a High-Level Language 2014 VLDB 0.00017370758
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
896 Froid: Optimization of Imperative Programs in a Relational Database 2018 VLDB 0.00013209291
1,072 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012168947
1,193 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.0001158809
1,883 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.4391795e-05
2,123 Tupleware: "Big" Data, Big Analytics, Small Clusters 2015 CIDR 9.0060385e-05
2,264 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.7289107e-05
2,478 Learning a Partitioning Advisor for Cloud Databases 2020 SIGMOD 8.4079121e-05
2,687 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1301312e-05
2,704 Supporting Table Partitioning By Reference in Oracle 2008 SIGMOD 8.1149334e-05
2,754 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0534972e-05
3,239 Locality-aware Partitioning in Parallel Database Systems 2015 SIGMOD 7.5007569e-05
3,440 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.2986853e-05
3,455 Scaling Spark in the Real World: Performance and Usability 2015 VLDB 7.2884813e-05
3,545 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2134803e-05
3,723 Schema Management for Document Stores 2015 VLDB 7.0706017e-05
4,288 Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics 2015 VLDB 6.6855423e-05
4,533 AdaptDB: Adaptive Partitioning for Distributed Joins 2017 VLDB 6.5565658e-05
5,289 Tensor Relational Algebra for Distributed Machine Learning System Design 2021 VLDB 6.1940481e-05
5,666 Skew-Aware Join Optimization for Array Databases 2015 SIGMOD 6.0462129e-05
6,721 Near-Optimal Distributed Band-Joins through Recursive Partitioning 2020 SIGMOD 5.6981399e-05
6,966 Incremental Elasticity For Array Databases 2014 SIGMOD 5.6327836e-05
7,632 MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs 2014 VLDB 5.4793438e-05
8,165 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 5.3851885e-05
8,219 Customizable and Scalable Fuzzy Join for Big Data 2019 VLDB 5.3761699e-05
9,748 PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development 2018 SIGMOD 5.1349531e-05
Previous Page 1 / 1 Next

Semantically Similar Papers