Back to papers
Tuplex: Data Science in Python at Native Code Speed
Summary: Tuplex JIT-compiles Python UDFs into end-to-end native code for data pipelines. It uses a dual-mode execution model: a fast path for the common case with exception paths for failures, yielding up to 91x speedups over Spark/Dask and near hand-tuned C++ performance.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h9c34eb13be105218
Venue
SIGMOD
Year
2021
Pagerank
9.6068397e-05
Overall Rank
1,803 | 87.88%
DOI
10.1145/3448016.3457244
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{spiegelberg_sigmod21,
title = {{Tuplex: Data Science in Python at Native Code Speed}},
author = {Spiegelberg, Leonhard and Yesantharao, Rahul and Schwarzkopf, Malte and Kraska, Tim},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457244},
url = {https://dl.acm.org/doi/10.1145/3448016.3457244},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 20 of 20 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
5,719
Accelerating Python UDFs in Vectorized Query Execution
2022
CIDR
6.0187779e-05
5,929
Self-Organizing Data Containers
2022
CIDR
5.9439349e-05
6,085
Dear User-Defined Functions, Inlining isn't working out so great for us. Let's try batching to make our relationship work. Sincerely, SQL
2024
CIDR
5.892166e-05
6,212
YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational Databases
2022
VLDB
5.8480745e-05
6,415
Containerized Execution of UDFs: An Experimental Evaluation
2022
VLDB
5.7926473e-05
6,742
Mitigating the Impedance Mismatch between Prediction Query Execution and Database Engine
2025
SIGMOD
5.6910432e-05
6,878
DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines
2022
CIDR
5.657878e-05
7,983
Efficient Execution of User-Defined Functions in SQL Queries
2023
VLDB
5.4131297e-05
8,153
Predicate Pushdown for Data Science Pipelines
2023
SIGMOD
5.3891786e-05
9,458
HyperBlocker: Accelerating Rule-based Blocking in Entity Resolution using GPUs
2025
VLDB
5.173446e-05
9,472
Query Compilation Without Regrets
2024
SIGMOD
5.1711207e-05
9,588
The Key to Effective UDF Optimization: Before Inlining, First Perform Outlining
2025
VLDB
5.156809e-05
9,656
BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach
2023
SIGMOD
5.1453267e-05
9,694
The UDFBench Benchmark for General-purpose UDF Queries
2025
VLDB
5.1396592e-05
10,046
YeSQL: Rich User-Defined Functions without the Overhead
2022
VLDB
5.091453e-05
10,416
Automating Database-Native Function Code Synthesis with LLMs
2026
SIGMOD
4.9793485e-05
11,168
UDFBench: A Tool for Benchmarking UDF Queries on SQL Engines
2025
SIGMOD
4.9793485e-05
11,177
Approximating Opaque Top-k Queries
2025
SIGMOD
4.9793485e-05
11,728
Udon: Efficient Debugging of User-Defined Functions in Big Data Systems with Line-by-Line Control
2023
SIGMOD
4.9793485e-05
11,797
To UDFs and Beyond: Demonstration of a Fully Decomposed Data Processor for General Data Wrangling Tasks
2023
VLDB
4.9793485e-05
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
21
Efficiently Compiling Efficient Query Plans for Modern Hardware
2011
VLDB
0.00056855599
23
Spark SQL: Relational Data Processing in Spark
2015
SIGMOD
0.00055406774
495
Building Efficient Query Engines in a High-Level Language
2014
VLDB
0.00017370758
605
Everything You Always Wanted to Know About Compiled and Vectorized Queries But Were Afraid to Ask
2018
VLDB
0.00015647561
1,444
An Architecture for Compiling UDF-centric Workflows
2015
VLDB
0.0001063181
1,720
A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data
2014
SIGMOD
9.7965659e-05
1,742
How to Architect a Query Compiler
2016
SIGMOD
9.7378418e-05
1,931
Instant Loading for Main Memory Databases
2013
VLDB
9.3474242e-05
2,195
Opening the Black Boxes in Data Flow Optimization
2012
VLDB
8.8781177e-05
2,249
Evaluating End-to-End Optimization for Data Analytics Applications in Weld
2018
VLDB
8.7549752e-05
2,419
How to Architect a Query Compiler, Revisited
2018
SIGMOD
8.4923505e-05
3,873
DBToaster: A SQL Compiler for High-Performance Delta Processing in Main-Memory Databases
2009
VLDB
6.9539394e-05
6,621
A Demonstration of DBWipes: Clean as You Query
2012
VLDB
5.7307884e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
12,285
The Best of Both Worlds: Big Data Programming with Both Productivity and Performance
2017
SIGMOD
2
2,053
tf.data: A Machine Learning Data Processing Framework
2021
VLDB
3
5,640
DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python
2021
SIGMOD
4
2,123
Tupleware: "Big" Data, Big Analytics, Small Clusters
2015
CIDR
5
1,193
Weld: A Common Runtime for High Performance Data Analytics
2017
CIDR
6
2,207
Magpie: Python at Speed and Scale using Cloud Backends
2021
CIDR
7
2,249
Evaluating End-to-End Optimization for Data Analytics Applications in Weld
2018
VLDB
8
5,719
Accelerating Python UDFs in Vectorized Query Execution
2022
CIDR
9
1,444
An Architecture for Compiling UDF-centric Workflows
2015
VLDB
10
10,047
Tuplex: Robust, Efficient Analytics When Python Rules
2019
VLDB