Back to papers
Tuplex: Data Science in Python at Native Code Speed
Summary: Tuplex JIT-compiles Python UDFs into end-to-end native code for data pipelines. It uses a dual-mode execution model: a fast path for the common case with exception paths for failures, yielding up to 91x speedups over Spark/Dask and near hand-tuned C++ performance.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 6136
- Venue
- SIGMOD
- Year
- 2021
- Pagerank
- 0.00010206514
- Overall Rank
- 1,884 | 86.91%
- DOI
-
10.1145/3448016.3457244
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 19 of 19 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 5,486 |
Containerized Execution of UDFs: An Experimental Evaluation |
2022 |
VLDB |
5.481452e-05 |
| 6,191 |
Accelerating Python UDFs in Vectorized Query Execution |
2022 |
CIDR |
5.1598046e-05 |
| 6,237 |
Self-Organizing Data Containers |
2022 |
CIDR |
5.1371094e-05 |
| 6,374 |
Dear User-Defined Functions, Inlining isn't working out so great for us. Let's try batching to make our relationship work. Sincerely, SQL |
2024 |
CIDR |
5.0874998e-05 |
| 6,377 |
Mitigating the Impedance Mismatch between Prediction Query Execution and Database Engine |
2025 |
SIGMOD |
5.0860948e-05 |
| 6,703 |
YeSQL: “You extend SQL” with Rich and Highly Performant User-Defined Functions in Relational Databases |
2022 |
VLDB |
4.9514593e-05 |
| 7,303 |
DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines |
2022 |
CIDR |
4.7632836e-05 |
| 8,580 |
Efficient Execution of User-Defined Functions in SQL Queries |
2023 |
VLDB |
4.4876382e-05 |
| 8,625 |
Predicate Pushdown for Data Science Pipelines |
2023 |
SIGMOD |
4.4784651e-05 |
| 9,331 |
BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach |
2023 |
SIGMOD |
4.351469e-05 |
| 9,348 |
The Key to Effective UDF Optimization: Before Inlining, First Perform Outlining |
2025 |
VLDB |
4.3504473e-05 |
| 9,717 |
YeSQL: Rich User-Defined Functions without the Overhead |
2022 |
VLDB |
4.2939577e-05 |
| 9,765 |
The UDFBench Benchmark for General-purpose UDF Queries |
2025 |
VLDB |
4.2815042e-05 |
| 9,846 |
HyperBlocker: Accelerating Rule-based Blocking in Entity Resolution using GPUs |
2025 |
VLDB |
4.2680295e-05 |
| 10,469 |
UDFBench: A Tool for Benchmarking UDF Queries on SQL Engines |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,481 |
Approximating Opaque Top-k Queries |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,972 |
Query Compilation Without Regrets |
2024 |
SIGMOD |
4.1905499e-05 |
| 11,215 |
Udon: Efficient Debugging of User-Defined Functions in Big Data Systems with Line-by-Line Control |
2023 |
SIGMOD |
4.1905499e-05 |
| 11,290 |
To UDFs and Beyond: Demonstration of a Fully Decomposed Data Processor for General Data Wrangling Tasks |
2023 |
VLDB |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 59 |
Efficiently Compiling Efficient Query Plans for Modern Hardware |
2011 |
VLDB |
0.0006445664 |
| 66 |
Spark SQL: Relational Data Processing in Spark |
2015 |
SIGMOD |
0.00061707583 |
| 701 |
Building Efficient Query Engines in a High-Level Language |
2014 |
VLDB |
0.00017893039 |
| 848 |
Everything You Always Wanted to Know About Compiled and Vectorized Queries But Were Afraid to Ask |
2018 |
VLDB |
0.00015933538 |
| 1,875 |
An Architecture for Compiling UDF-centric Workflows |
2015 |
VLDB |
0.00010243959 |
| 2,177 |
A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data |
2014 |
SIGMOD |
9.371335e-05 |
| 2,326 |
Instant Loading for Main Memory Databases |
2013 |
VLDB |
9.0271498e-05 |
| 2,383 |
How to Architect a Query Compiler |
2016 |
SIGMOD |
8.9198524e-05 |
| 2,616 |
Opening the Black Boxes in Data Flow Optimization |
2012 |
VLDB |
8.4457819e-05 |
| 2,843 |
How to Architect a Query Compiler, Revisited |
2018 |
SIGMOD |
8.0334687e-05 |
| 2,904 |
Evaluating End-to-End Optimization for Data Analytics Applications in Weld |
2018 |
VLDB |
7.9403097e-05 |
| 4,407 |
DBToaster: A SQL Compiler for High-Performance Delta Processing in Main-Memory Databases |
2009 |
VLDB |
6.2033567e-05 |
| 6,383 |
A Demonstration of DBWipes: Clean as You Query |
2012 |
VLDB |
5.0831604e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 11,790 |
The Best of Both Worlds: Big Data Programming with Both Productivity and Performance |
2017 |
SIGMOD |
4.1905499e-05 |
| 2,175 |
tf.data: A Machine Learning Data Processing Framework |
2021 |
VLDB |
9.3745231e-05 |
| 5,984 |
DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python |
2021 |
SIGMOD |
5.2400405e-05 |
| 2,421 |
Tupleware: "Big" Data, Big Analytics, Small Clusters |
2015 |
CIDR |
8.8471956e-05 |
| 1,751 |
Weld: A Common Runtime for High Performance Data Analytics |
2017 |
CIDR |
0.00010673874 |
| 2,955 |
Magpie: Python at Speed and Scale using Cloud Backends |
2021 |
CIDR |
7.8188583e-05 |
| 2,904 |
Evaluating End-to-End Optimization for Data Analytics Applications in Weld |
2018 |
VLDB |
7.9403097e-05 |
| 6,191 |
Accelerating Python UDFs in Vectorized Query Execution |
2022 |
CIDR |
5.1598046e-05 |
| 1,875 |
An Architecture for Compiling UDF-centric Workflows |
2015 |
VLDB |
0.00010243959 |
| 9,718 |
Tuplex: Robust, Efficient Analytics When Python Rules |
2019 |
VLDB |
4.2939577e-05 |