Database Paper Browser

Back to papers

Lachesis: Automatic Partitioning for UDF-Centric Analytics

Summary: Lachesis automates partitioning for UDF-centric analytics by modeling workloads as sub-computations to guide data partitioning. Deep RL selects sub-computations to partition, enabling automatic storage optimization across apps and improved productivity. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12317
Venue
VLDB
Year
2021
Pagerank
4.7199075e-05
Overall Rank
7,459 | 48.17%
DOI
10.14778/3457390.3457392

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 33 of 33 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
70 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00059744625
283 Integrating Vertical and Horizontal Partitioning into Automated Physical Database Design 2004 SIGMOD 0.00029024583
285 Automating Physical Database Design in a Parallel Database 2002 SIGMOD 0.00028978423
407 Database Cracking 2007 CIDR 0.00023941779
557 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.00020186115
701 Building Efficient Query Engines in a High-Level Language 2014 VLDB 0.00017893039
789 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00016602215
981 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00014866268
1,107 Froid: Optimization of Imperative Programs in a Relational Database 2018 VLDB 0.0001397627
1,751 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.00010673874
2,355 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.9727612e-05
2,410 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 8.8643562e-05
2,421 Tupleware: "Big" Data, Big Analytics, Small Clusters 2015 CIDR 8.8471956e-05
2,441 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.8106295e-05
2,823 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0593793e-05
2,985 Supporting Table Partitioning By Reference in Oracle 2008 SIGMOD 7.7758625e-05
3,066 Learning a Partitioning Advisor for Cloud Databases 2020 SIGMOD 7.6255556e-05
3,352 Schema Management for Document Stores 2015 VLDB 7.1838008e-05
3,536 Scaling Spark in the Real World: Performance and Usability 2015 VLDB 6.9938207e-05
3,825 Locality-aware Partitioning in Parallel Database Systems 2015 SIGMOD 6.7225803e-05
4,068 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 6.4748133e-05
4,171 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 6.3800823e-05
4,406 Declarative Recursive Computation on an RDBMS 2019 VLDB 6.2044305e-05
4,435 Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics 2015 VLDB 6.18493e-05
5,116 AdaptDB: Adaptive Partitioning for Distributed Joins 2017 VLDB 5.6805476e-05
5,836 Tensor Relational Algebra for Distributed Machine Learning System Design 2021 VLDB 5.3079723e-05
5,961 Skew-Aware Join Optimization for Array Databases 2015 SIGMOD 5.2510172e-05
6,618 Near-Optimal Distributed Band-Joins through Recursive Partitioning 2020 SIGMOD 4.9864636e-05
7,132 Incremental Elasticity For Array Databases 2014 SIGMOD 4.817779e-05
7,298 MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs 2014 VLDB 4.7651301e-05
8,005 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 4.6044087e-05
8,116 Customizable and Scalable Fuzzy Join for Big Data 2019 VLDB 4.5788731e-05
9,337 PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development 2018 SIGMOD 4.351469e-05
Previous Page 1 / 1 Next

Semantically Similar Papers