DBScholar

Back to papers

Lachesis: Automatic Partitioning for UDF-Centric Analytics

Summary: Lachesis automates partitioning for UDF-centric analytics, where reusable computation and functional dependencies are scarce. It models workloads as analyzable workflow sub-computations and uses deep reinforcement learning to select partitioning features, optimizing shared storage across applications. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
12504
Venue
VLDB
Year
2021
Pagerank
5.6491618e-05
Overall Rank
7,311 | 49.85%
DOI
10.14778/3457390.3457392

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{zou_vldb21,
        title = {{Lachesis: Automatic Partitioning for UDF-Centric Analytics}},
        author = {Zou, Jia and Das, Amitabh and Barhate, Pratik and Iyengar, Arun and Yuan, Binhang and Jankov, Dimitrije and Jermaine, Chris},
        journal = {PVLDB},
        series = {{VLDB} '21},
        volume = {14},
        number = {8},
        pages = {1262--1275},
        doi = {10.14778/3457390.3457392},
        url = {https://doi.org/10.14778/3457390.3457392},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 32 of 32 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
199 Integrating Vertical and Horizontal Partitioning into Automated Physical Database Design 2004 SIGMOD 0.00025612088
246 Automating Physical Database Design in a Parallel Database 2002 SIGMOD 0.00023457421
259 Database Cracking 2007 CIDR 0.00023119313
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001888524
534 Building Efficient Query Engines in a High-Level Language 2014 VLDB 0.00017046514
660 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00015198804
894 Froid: Optimization of Imperative Programs in a Relational Database 2018 VLDB 0.00013367658
1,054 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012390673
1,229 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.00011578425
1,852 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.6134443e-05
2,094 Tupleware: "Big" Data, Big Analytics, Small Clusters 2015 CIDR 9.1819738e-05
2,239 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.8875753e-05
2,499 Learning a Partitioning Advisor for Cloud Databases 2020 SIGMOD 8.4993549e-05
2,642 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.3059948e-05
2,666 Supporting Table Partitioning By Reference in Oracle 2008 SIGMOD 8.2764713e-05
2,717 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.2102313e-05
3,200 Locality-aware Partitioning in Parallel Database Systems 2015 SIGMOD 7.642132e-05
3,411 Scaling Spark in the Real World: Performance and Usability 2015 VLDB 7.436229e-05
3,466 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.3909785e-05
3,605 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2640711e-05
3,655 Schema Management for Document Stores 2015 VLDB 7.2226714e-05
4,208 Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics 2015 VLDB 6.8319812e-05
4,456 AdaptDB: Adaptive Partitioning for Distributed Joins 2017 VLDB 6.692321e-05
5,195 Tensor Relational Algebra for Distributed Machine Learning System Design 2021 VLDB 6.3232105e-05
5,534 Skew-Aware Join Optimization for Array Databases 2015 SIGMOD 6.1831004e-05
6,596 Near-Optimal Distributed Band-Joins through Recursive Partitioning 2020 SIGMOD 5.828647e-05
6,831 Incremental Elasticity For Array Databases 2014 SIGMOD 5.7599208e-05
7,546 MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs 2014 VLDB 5.6024593e-05
8,001 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 5.508791e-05
8,088 Customizable and Scalable Fuzzy Join for Big Data 2019 VLDB 5.4900832e-05
9,572 PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development 2018 SIGMOD 5.2528121e-05
Previous Page 1 / 1 Next

Semantically Similar Papers