Back to papers
In-Database Data Imputation
Summary: In-database data imputation with MICE, using computation sharing and a ring abstraction to speed training. In-db learning of stochastic linear regression and Gaussian discriminant analysis for continuous/categorical imputation; PostgreSQL and DuckDB beat prior MICE and model-based methods by up to two orders of magnitude, preserving relationships.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
6941
Venue
SIGMOD
Year
2024
Pagerank
5.1815618e-05
Overall Rank
9,993 | 31.44%
DOI
10.1145/3639326
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{perini_sigmod24,
title = {{In-Database Data Imputation}},
author = {Perini, Massimo and Nikolic, Milos},
series = {{SIGMOD} '24},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3639326},
url = {https://dl.acm.org/doi/10.1145/3639326},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 35 of 35 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
106
The MADlib Analytics Library or MAD Skills, the SQL
2012
VLDB
0.00033539462
112
HoloClean: Holistic Data Repairs with Probabilistic Inference
2017
VLDB
0.00032801121
415
SystemML: Declarative Machine Learning on Spark
2016
VLDB
0.0001888524
518
Towards a Unified Architecture for in-RDBMS Analytics
2012
SIGMOD
0.00017167492
536
Learning Linear Regression Models over Factorized Joins
2016
SIGMOD
0.0001693369
715
Learning Generalized Linear Models Over Normalized Data
2015
SIGMOD
0.00014655327
960
Incremental Query Evaluation in a Ring of Databases
2010
PODS
0.00012945163
1,235
Towards Linear Algebra over Normalized Data
2017
VLDB
0.00011548457
1,250
Data Management in Machine Learning: Challenges, Techniques, and Systems
2017
SIGMOD
0.00011485301
1,323
Data Cleaning: Overview and Emerging Challenges
2016
SIGMOD
0.00011152602
1,340
Responsible Data Management
2020
VLDB
0.00011111667
1,350
Automating Large-Scale Data Quality Verification
2018
VLDB
0.00011065626
1,351
Detecting Data Errors: Where are we and what needs to be done?
2016
VLDB
0.00011064851
1,476
Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning
2020
VLDB
0.00010659277
2,123
Discovery of Genuine Functional Dependencies from Relational Data with Missing Values
2018
VLDB
9.1372798e-05
2,147
Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions
2021
VLDB
9.0831495e-05
2,208
Query Optimization for Dynamic Imputation
2017
VLDB
8.9512455e-05
2,545
Mind the Gap: An Experimental Evaluation of Imputation of Missing Values Techniques in Time Series
2020
VLDB
8.4401333e-05
2,769
A Layered Aggregate Engine for Analytics Workloads
2019
SIGMOD
8.1465406e-05
3,206
Incremental View Maintenance with Triple Lock Factorization Benefits
2018
SIGMOD
7.6367549e-05
3,386
Efficient and Effective Data Imputation with Influence Functions
2022
VLDB
7.453827e-05
4,034
Missing Value Imputation on Multidimensional Time Series
2021
VLDB
6.9446462e-05
4,488
Horizon: Scalable Dependency-driven Data Cleaning
2021
VLDB
6.668457e-05
4,609
F-IVM: Learning over Fast-Evolving Relational Data
2020
SIGMOD
6.6081313e-05
4,835
Adaptive Data Augmentation for Supervised Learning over Missing Data
2021
VLDB
6.486592e-05
5,175
Cleanits: A Data Cleaning System for Industrial Time Series
2019
VLDB
6.3332964e-05
5,378
Troubles with Nulls, Views from the Users
2022
VLDB
6.2406439e-05
5,518
Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks
2023
VLDB
6.1885722e-05
6,417
ORBITS: Online Recovery of Missing Values in Multiple Time Series Streams
2021
VLDB
5.8822847e-05
6,585
JoinBoost: Grow Trees Over Normalized Data Using Only SQL
2023
VLDB
5.8350362e-05
7,556
ReStore - Neural Data Completion for Relational Databases
2021
SIGMOD
5.5997742e-05
7,880
Learning Over Dirty Data Without Cleaning
2020
SIGMOD
5.5244204e-05
8,002
Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints
2020
SIGMOD
5.5085906e-05
8,105
Online Topic-Aware Entity Resolution Over Incomplete Data Streams
2021
SIGMOD
5.4860105e-05
9,653
CoClean: Collaborative Data Cleaning
2020
SIGMOD
5.2425585e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
11,258
Win-Win: On Simultaneous Clustering and Imputing over Incomplete Data
2024
VLDB
2
9,285
ZIP: Lazy Imputation during Query Processing
2024
VLDB
3
7,844
Missing Value Imputation for Multi-attribute Sensor Data Streams via Message Propagation
2024
VLDB
4
5,167
Enriching Data Imputation with Extensive Similarity Neighbors
2015
VLDB
5
8,002
Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints
2020
SIGMOD
6
6,642
Missing Data Imputation with Uncertainty-Driven Network
2024
SIGMOD
7
9,550
Data Imputation with Limited Data Redundancy Using Data Lakes
2025
VLDB
8
3,386
Efficient and Effective Data Imputation with Influence Functions
2022
VLDB
9
11,169
Certain and Approximately Certain Models for Statistical Learning
2024
SIGMOD
10
2,208
Query Optimization for Dynamic Imputation
2017
VLDB