Back to papers
In-Database Data Imputation
Summary: In-database data imputation with MICE, using computation sharing and a ring abstraction to speed training. In-db learning of stochastic linear regression and Gaussian discriminant analysis for continuous/categorical imputation; PostgreSQL and DuckDB beat prior MICE and model-based methods by up to two orders of magnitude, preserving relationships.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h928ab16d37b4852d
Venue
SIGMOD
Year
2024
Pagerank
5.0629036e-05
Overall Rank
10,185 | 31.55%
DOI
10.1145/3639326
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{perini_sigmod24,
title = {{In-Database Data Imputation}},
author = {Perini, Massimo and Nikolic, Milos},
series = {{SIGMOD} '24},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3639326},
url = {https://dl.acm.org/doi/10.1145/3639326},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 35 of 35 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
104
HoloClean: Holistic Data Repairs with Probabilistic Inference
2017
VLDB
0.00033676943
105
The MADlib Analytics Library or MAD Skills, the SQL
2012
VLDB
0.00033633007
416
SystemML: Declarative Machine Learning on Spark
2016
VLDB
0.00018650998
503
Towards a Unified Architecture for in-RDBMS Analytics
2012
SIGMOD
0.00017195428
521
Learning Linear Regression Models over Factorized Joins
2016
SIGMOD
0.00016923519
731
Learning Generalized Linear Models Over Normalized Data
2015
SIGMOD
0.00014400356
977
Incremental Query Evaluation in a Ring of Databases
2010
PODS
0.00012730241
1,044
Data Cleaning: Overview and Emerging Challenges
2016
SIGMOD
0.00012329478
1,162
Responsible Data Management
2020
VLDB
0.00011747599
1,223
Data Management in Machine Learning: Challenges, Techniques, and Systems
2017
SIGMOD
0.00011468426
1,257
Towards Linear Algebra over Normalized Data
2017
VLDB
0.0001130959
1,308
Automating Large-Scale Data Quality Verification
2018
VLDB
0.00011073863
1,342
Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning
2020
VLDB
0.00010963254
1,344
Detecting Data Errors: Where are we and what needs to be done?
2016
VLDB
0.00010951939
1,848
Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions
2021
VLDB
9.5075544e-05
2,127
Discovery of Genuine Functional Dependencies from Relational Data with Missing Values
2018
VLDB
8.9990594e-05
2,241
Query Optimization for Dynamic Imputation
2017
VLDB
8.7704255e-05
2,590
Mind the Gap: An Experimental Evaluation of Imputation of Missing Values Techniques in Time Series
2020
VLDB
8.2468536e-05
2,663
A Layered Aggregate Engine for Analytics Workloads
2019
SIGMOD
8.1542952e-05
3,114
Incremental View Maintenance with Triple Lock Factorization Benefits
2018
SIGMOD
7.6321464e-05
3,441
Efficient and Effective Data Imputation with Influence Functions
2022
VLDB
7.2946229e-05
4,132
Missing Value Imputation on Multidimensional Time Series
2021
VLDB
6.7856132e-05
4,348
Horizon: Scalable Dependency-driven Data Cleaning
2021
VLDB
6.6438522e-05
4,702
F-IVM: Learning over Fast-Evolving Relational Data
2020
SIGMOD
6.46033e-05
4,955
Adaptive Data Augmentation for Supervised Learning over Missing Data
2021
VLDB
6.3386868e-05
5,298
Cleanits: A Data Cleaning System for Industrial Time Series
2019
VLDB
6.1882634e-05
5,510
Troubles with Nulls, Views from the Users
2022
VLDB
6.0980926e-05
5,645
Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks
2023
VLDB
6.0509471e-05
6,552
JoinBoost: Grow Trees Over Normalized Data Using Only SQL
2023
VLDB
5.7475822e-05
6,553
ORBITS: Online Recovery of Missing Values in Multiple Time Series Streams
2021
VLDB
5.7475799e-05
7,710
ReStore - Neural Data Completion for Relational Databases
2021
SIGMOD
5.471539e-05
7,854
CoClean: Collaborative Data Cleaning
2020
SIGMOD
5.4380042e-05
7,880
Learning Over Dirty Data Without Cleaning
2020
SIGMOD
5.4330096e-05
8,172
Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints
2020
SIGMOD
5.3824434e-05
8,288
Online Topic-Aware Entity Resolution Over Incomplete Data Streams
2021
SIGMOD
5.3603804e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
11,593
Win-Win: On Simultaneous Clustering and Imputing over Incomplete Data
2024
VLDB
2
9,464
ZIP: Lazy Imputation during Query Processing
2024
VLDB
3
8,008
Missing Value Imputation for Multi-attribute Sensor Data Streams via Message Propagation
2024
VLDB
4
5,281
Enriching Data Imputation with Extensive Similarity Neighbors
2015
VLDB
5
8,172
Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints
2020
SIGMOD
6
6,780
Missing Data Imputation with Uncertainty-Driven Network
2024
SIGMOD
7
7,589
Data Imputation with Limited Data Redundancy Using Data Lakes
2025
VLDB
8
3,441
Efficient and Effective Data Imputation with Influence Functions
2022
VLDB
9
11,520
Certain and Approximately Certain Models for Statistical Learning
2024
SIGMOD
10
2,241
Query Optimization for Dynamic Imputation
2017
VLDB