Back to papers
In-Database Data Imputation
Summary: In-database data imputation with MICE, using computation sharing and a ring abstraction to speed training. In-db learning of stochastic linear regression and Gaussian discriminant analysis for continuous/categorical imputation; PostgreSQL and DuckDB beat prior MICE and model-based methods by up to two orders of magnitude, preserving relationships.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h928ab16d37b4852d
Venue
SIGMOD
Year
2024
Pagerank
5.0653015e-05
Overall Rank
10,181 | 31.55%
DOI
10.1145/3639326
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{perini_sigmod24,
title = {{In-Database Data Imputation}},
author = {Perini, Massimo and Nikolic, Milos},
series = {{SIGMOD} '24},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3639326},
url = {https://dl.acm.org/doi/10.1145/3639326},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 35 of 35 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
104
HoloClean: Holistic Data Repairs with Probabilistic Inference
2017
VLDB
0.00033690989
105
The MADlib Analytics Library or MAD Skills, the SQL
2012
VLDB
0.00033638251
415
SystemML: Declarative Machine Learning on Spark
2016
VLDB
0.0001865959
503
Towards a Unified Architecture for in-RDBMS Analytics
2012
SIGMOD
0.00017202276
521
Learning Linear Regression Models over Factorized Joins
2016
SIGMOD
0.00016929744
730
Learning Generalized Linear Models Over Normalized Data
2015
SIGMOD
0.00014406936
978
Incremental Query Evaluation in a Ring of Databases
2010
PODS
0.00012731074
1,043
Data Cleaning: Overview and Emerging Challenges
2016
SIGMOD
0.00012335114
1,162
Responsible Data Management
2020
VLDB
0.00011753159
1,255
Data Management in Machine Learning: Challenges, Techniques, and Systems
2017
SIGMOD
0.00011325762
1,256
Towards Linear Algebra over Normalized Data
2017
VLDB
0.00011314687
1,308
Automating Large-Scale Data Quality Verification
2018
VLDB
0.0001107886
1,342
Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning
2020
VLDB
0.00010968223
1,344
Detecting Data Errors: Where are we and what needs to be done?
2016
VLDB
0.00010956518
1,847
Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions
2021
VLDB
9.5120573e-05
2,125
Discovery of Genuine Functional Dependencies from Relational Data with Missing Values
2018
VLDB
9.0033119e-05
2,239
Query Optimization for Dynamic Imputation
2017
VLDB
8.7745792e-05
2,589
Mind the Gap: An Experimental Evaluation of Imputation of Missing Values Techniques in Time Series
2020
VLDB
8.2507594e-05
2,663
A Layered Aggregate Engine for Analytics Workloads
2019
SIGMOD
8.1581558e-05
3,112
Incremental View Maintenance with Triple Lock Factorization Benefits
2018
SIGMOD
7.6357579e-05
3,441
Efficient and Effective Data Imputation with Influence Functions
2022
VLDB
7.2980777e-05
4,131
Missing Value Imputation on Multidimensional Time Series
2021
VLDB
6.788827e-05
4,347
Horizon: Scalable Dependency-driven Data Cleaning
2021
VLDB
6.6469984e-05
4,699
F-IVM: Learning over Fast-Evolving Relational Data
2020
SIGMOD
6.4633894e-05
4,953
Adaptive Data Augmentation for Supervised Learning over Missing Data
2021
VLDB
6.3416889e-05
5,294
Cleanits: A Data Cleaning System for Industrial Time Series
2019
VLDB
6.1911943e-05
5,507
Troubles with Nulls, Views from the Users
2022
VLDB
6.1009808e-05
5,643
Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks
2023
VLDB
6.0538129e-05
6,550
JoinBoost: Grow Trees Over Normalized Data Using Only SQL
2023
VLDB
5.7503043e-05
6,551
ORBITS: Online Recovery of Missing Values in Multiple Time Series Streams
2021
VLDB
5.750302e-05
7,704
ReStore - Neural Data Completion for Relational Databases
2021
SIGMOD
5.4741304e-05
7,850
CoClean: Collaborative Data Cleaning
2020
SIGMOD
5.4405793e-05
7,875
Learning Over Dirty Data Without Cleaning
2020
SIGMOD
5.4355826e-05
8,166
Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints
2020
SIGMOD
5.3849926e-05
8,282
Online Topic-Aware Entity Resolution Over Incomplete Data Streams
2021
SIGMOD
5.3629192e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
11,587
Win-Win: On Simultaneous Clustering and Imputing over Incomplete Data
2024
VLDB
2
9,455
ZIP: Lazy Imputation during Query Processing
2024
VLDB
3
8,003
Missing Value Imputation for Multi-attribute Sensor Data Streams via Message Propagation
2024
VLDB
4
5,278
Enriching Data Imputation with Extensive Similarity Neighbors
2015
VLDB
5
8,166
Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints
2020
SIGMOD
6
6,775
Missing Data Imputation with Uncertainty-Driven Network
2024
SIGMOD
7
7,583
Data Imputation with Limited Data Redundancy Using Data Lakes
2025
VLDB
8
3,441
Efficient and Effective Data Imputation with Influence Functions
2022
VLDB
9
11,514
Certain and Approximately Certain Models for Statistical Learning
2024
SIGMOD
10
2,239
Query Optimization for Dynamic Imputation
2017
VLDB