The Fast and the Private: Task-based Dataset Search
Summary: Mileena uses pre-computed semi-ring sketches to rapidly evaluate joins/unions that augment a requester's dataset for task-driven ML, enabling low-latency training and evaluation over large corpora. Introduces a Factorized Privacy Mechanism for scalable differential privacy with minimal utility loss, and integrates LLM-based transformation agents plus semi-ring extensions for causal discovery and treatment-effect estimation. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Zezhou Huang (Columbia University)
- 2. Jiaxiang Liu (Columbia University)
- 3. Haonan Wang (Columbia University)
- 4. Eugene Wu (Columbia University)
BibTeX Citation
@inproceedings{huang_cidr24,
address = {Amsterdam, Netherlands},
series = {{CIDR} '24},
title = {{The Fast and the Private: Task-based Dataset Search}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Huang, Zezhou and Liu, Jiaxiang and Wang, Haonan and Wu, Eugene},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,773 | Fair and Actionable Causal Prescription Ruleset | 2025 | SIGMOD | 5.2209769e-05 |
| 10,627 | Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution | 2026 | VLDB | 5.093636e-05 |
| 10,968 | Suna: Scalable Causal Confounder Discovery over Relational Data | 2025 | VLDB | 5.093636e-05 |
| 11,164 | An LDP Compatible Sketch for Securely Approximating Set Intersection Cardinalities | 2024 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 16 of 16 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,874 | Fast Search-By-Classification for Large-Scale Databases Using Index-Aware Decision Trees and Random Forests | 2023 | VLDB |
| 2 | 9,353 | On Efficient Approximate Queries over Machine Learning Models | 2023 | VLDB |
| 3 | 286 | Milvus: A Purpose-Built Vector Data Management System | 2021 | SIGMOD |
| 4 | 13,314 | SemExplorer: A User Interface for Semantic Approach to Customized Dataset Search | 2025 | SIGMOD |
| 5 | 10,638 | A Theoretical Framework for Distribution-Aware Dataset Search | 2025 | PODS |
| 6 | 10,720 | Finding What You’re Looking For: A Distribution-Aware Dataset Search Engine in Action | 2025 | SIGMOD |
| 7 | 10,627 | Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution | 2026 | VLDB |
| 8 | 8,908 | Privacy and Accuracy-Aware AI/ML Model Deduplication | 2025 | SIGMOD |
| 9 | 4,592 | Data Platform for Machine Learning | 2019 | SIGMOD |
| 10 | 7,914 | Saibot: A Differentially Private Data Search Platform | 2023 | VLDB |