Back to papers
Sample Debiasing in the Themis Open World Database System
Summary: Themis is the first open-world DB that rebalances biased samples to approximate population-wide query results. It blends sample reweighting with Bayesian nets, using a priori population info to beat AQP and baselines with latency and robustness to gaps.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 5820
- Venue
- SIGMOD
- Year
- 2020
- Pagerank
- 6.2367043e-05
- Overall Rank
- 4,372 | 69.62%
- DOI
-
10.1145/3318464.3380606
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 7 of 7 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 17 of 17 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 47 |
Data Integration: A Theoretical Perspective |
2002 |
PODS |
0.00069691761 |
| 71 |
How Good Are Query Optimizers, Really? |
2016 |
VLDB |
0.00059446482 |
| 373 |
Selectivity Estimation using Probabilistic Models |
2001 |
SIGMOD |
0.00025354685 |
| 468 |
MauveDB: Supporting Model-based User Views in Database Systems |
2006 |
SIGMOD |
0.00022407392 |
| 1,574 |
Approximate Query Processing: No Silver Bullet |
2017 |
SIGMOD |
0.00011289028 |
| 1,978 |
Improved Selectivity Estimation by Combining Knowledge from Sampling and Synopses |
2018 |
VLDB |
9.8764627e-05 |
| 2,126 |
IDEBench: A Benchmark for Interactive Data Exploration |
2020 |
SIGMOD |
9.4814404e-05 |
| 2,177 |
A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data |
2014 |
SIGMOD |
9.371335e-05 |
| 2,385 |
Leveraging Aggregate Constraints For Deduplication |
2007 |
SIGMOD |
8.9167648e-05 |
| 2,424 |
The Analytical Bootstrap: a New Method for Fast Error Estimation in Approximate Query Processing |
2014 |
SIGMOD |
8.8415494e-05 |
| 2,583 |
Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee |
2016 |
SIGMOD |
8.4973431e-05 |
| 2,589 |
Database Learning: Toward a Database that Becomes Smarter Every Time |
2017 |
SIGMOD |
8.4868591e-05 |
| 3,944 |
AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics |
2018 |
SIGMOD |
6.6056349e-05 |
| 4,020 |
Revisiting Reuse for Approximate Query Processing |
2017 |
VLDB |
6.5209063e-05 |
| 5,877 |
Differential Privacy and the US Census |
2019 |
PODS |
5.2886439e-05 |
| 7,875 |
Probabilistic Database Summarization for Interactive Data Exploration |
2017 |
VLDB |
4.6262777e-05 |
| 8,108 |
NetCube: A Scalable Tool for Fast Data Mining and Compression |
2001 |
VLDB |
4.5808537e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 7,577 |
Pushing the Boundaries of Crowd-enabled Databases with Query-driven Schema Expansion |
2012 |
VLDB |
4.703047e-05 |
| 4,522 |
A Temporal-Probabilistic Database Model for Information Extraction |
2013 |
VLDB |
6.1109228e-05 |
| 5,206 |
ThalamusDB: Approximate Query Processing on Multi-Modal Data |
2024 |
SIGMOD |
5.625641e-05 |
| 12,033 |
A Social Network Database that Learns How to Answer Queries |
2013 |
CIDR |
4.1905499e-05 |
| 2,120 |
Using Probabilistic Models for Data Management in Acquisitional Environments |
2005 |
CIDR |
9.5017579e-05 |
| 3,084 |
Knowledge Expansion over Probabilistic Knowledge Bases |
2014 |
SIGMOD |
7.5967738e-05 |
| 4,347 |
On Biased Reservoir Sampling in the Presence of Stream Evolution |
2006 |
VLDB |
6.2588401e-05 |
| 9,206 |
Themis: A GPU-accelerated Relational Query Execution Engine |
2025 |
VLDB |
4.3695556e-05 |
| 8,330 |
THEMIS: Fairness in Federated Stream Processing under Overload |
2016 |
SIGMOD |
4.5391049e-05 |
| 6,230 |
Mosaic: A Sample-Based Database System for Open World Query Processing |
2020 |
CIDR |
5.1402482e-05 |