Back to papers
BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees
Summary: BlinkML enables fast, approximate ML training with probabilistic guarantees that the approximate model matches full-model predictions. Supports any MLE-based model (GLMs, PPCA) and uses error-bounded sampling to deliver 6x–629x speedups while preserving decisions.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h5e3a956e639bb413
Venue
SIGMOD
Year
2019
Pagerank
5.9599042e-05
Overall Rank
5,877 | 60.51%
DOI
10.1145/3299869.3300077
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{park_sigmod19,
title = {{BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees}},
author = {Park, Yongjoo and Qing, Jingyi and Shen, Xiaoyang and Mozafari, Barzan},
series = {{SIGMOD} '19},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3299869.3300077},
url = {https://dl.acm.org/doi/10.1145/3299869.3300077},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 27 of 27 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
6
Pig Latin: A Not-So-Foreign Language for Data Processing
2008
SIGMOD
0.0010515896
105
The MADlib Analytics Library or MAD Skills, the SQL
2012
VLDB
0.00033633007
138
Join Synopses for Approximate Query Answering
1999
SIGMOD
0.00029618887
271
NoScope: Optimizing Neural Network Queries over Video at Scale
2017
VLDB
0.0002256866
335
The Aqua Approximate Query Answering System
1999
SIGMOD
0.000206533
416
SystemML: Declarative Machine Learning on Spark
2016
VLDB
0.00018650998
521
Learning Linear Regression Models over Factorized Joins
2016
SIGMOD
0.00016923519
579
Incremental Knowledge Base Construction Using DeepDive
2015
VLDB
0.00016083582
596
Wander Join: Online Aggregation via Random Walks
2016
SIGMOD
0.00015782051
654
Materialization Optimizations for Feature Selection Workloads
2014
SIGMOD
0.00015096817
731
Learning Generalized Linear Models Over Normalized Data
2015
SIGMOD
0.00014400356
772
VerdictDB: Universalizing Approximate Query Processing
2018
SIGMOD
0.0001409096
779
To Join or Not to Join? Thinking Twice about Joins before Feature Selection
2016
SIGMOD
0.00014048128
841
Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters
2016
SIGMOD
0.00013543
930
Dynamic Sample Selection for Approximate Query Processing
2003
SIGMOD
0.00013009255
1,167
DimmWitted: A Study of Main-Memory Statistical Analytics
2014
VLDB
0.0001172597
1,428
Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems
2014
SIGMOD
0.0001069161
1,614
Compressed Linear Algebra for Large-Scale Machine Learning
2016
VLDB
0.00010067153
1,916
The Analytical Bootstrap: a New Method for Fast Error Estimation in Approximate Query Processing
2014
SIGMOD
9.3822742e-05
2,028
Database Learning: Toward a Database that Becomes Smarter Every Time
2017
SIGMOD
9.1584244e-05
2,186
Heterogeneity-aware Distributed Parameter Servers
2017
SIGMOD
8.8916253e-05
2,587
Brainwash: A Data System for Feature Engineering
2013
CIDR
8.2523942e-05
3,213
Turbo-Charging Estimate Convergence in DBO
2009
VLDB
7.5304969e-05
4,852
Neighbor-Sensitive Hashing
2016
VLDB
6.3776281e-05
6,002
Approximate Query Engines: Commercial Challenges and Research Opportunities
2017
SIGMOD
5.9149725e-05
9,853
DimBoost: Boosting Gradient Boosting Decision Tree to Higher Dimensions
2018
SIGMOD
5.1188938e-05
12,221
Demonstration of VerdictDB, the Platform-Independent AQP System
2018
SIGMOD
4.9769913e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
3,793
SLiMFast: Guaranteed Results for Data Fusion and Source Reliability
2017
SIGMOD
2
6,656
Hindsight Logging for Model Training
2021
VLDB
3
12,045
FlashP: An Analytical Pipeline for Real-time Forecasting of Time-Series Relational Data
2021
VLDB
4
1,614
Compressed Linear Algebra for Large-Scale Machine Learning
2016
VLDB
5
3,873
SketchML: Accelerating Distributed Machine Learning with Data Sketches
2018
SIGMOD
6
4,881
Accelerating Generalized Linear Models with MLWeaving: A One-Size-Fits-All System for Any-Precision Learning
2019
VLDB
7
1,095
BlinkFill: Semi-supervised Programming By Example for Syntactic String Transformations
2016
VLDB
8
10,414
Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees
2026
SIGMOD
9
6,146
Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse
2018
VLDB
10
8,815
ApproxML: Efficient Approximate Ad-Hoc ML Models Through Materialization and Reuse
2019
VLDB