Just can't get enough - Synthesizing Big Data
Summary: DBSynth automatically generates realistic, large-scale synthetic data from schema and sample data, extracting value-level features and Markov models. As an extension to PDGF, it enables fast, scalable generation across formats (CSV/JSON/XML/SQL) for big-data benchmarks (TPC-DI/BigBench) with automatic dictionaries and multi-core speedups. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Tilmann Rabl (University of Toronto)
- 2. Manuel Danisch (Bankmark UG)
- 3. Michael Frank (Bankmark UG)
- 4. Sebastian Schindler (Bankmark UG)
- 5. Hans-Arno Jacobsen (University of Toronto)
BibTeX Citation
@inproceedings{rabl_sigmod15,
title = {{Just can't get enough - Synthesizing Big Data}},
author = {Rabl, Tilmann and Danisch, Manuel and Frank, Michael and Schindler, Sebastian and Jacobsen, Hans-Arno},
series = {{SIGMOD} '15},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2723372.2735378},
url = {https://dl.acm.org/doi/10.1145/2723372.2735378},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,898 | Synthesizing Linked Data Under Cardinality and Integrity Constraints | 2021 | SIGMOD | 5.6506085e-05 |
| 8,034 | Dscaler: Synthetically Scaling A Given Relational Database | 2016 | VLDB | 5.401567e-05 |
| 8,206 | HYDRA: A Dynamic Big Data Regenerator | 2018 | VLDB | 5.3765919e-05 |
| 10,180 | Projection-Compliant Database Generation | 2022 | VLDB | 5.06553e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 106 | Quickly Generating Billion-Record Synthetic Databases | 1994 | SIGMOD | 0.00033518428 |
| 748 | QAGen: Generating Query-Aware Test Databases | 2007 | SIGMOD | 0.00014277246 |
| 956 | Flexible Database Generators | 2005 | VLDB | 0.0001286031 |
| 1,604 | Simple and Realistic Data Generation | 2006 | VLDB | 0.00010096875 |
| 1,702 | BigBench: Towards an Industry Standard Benchmark for Big Data Analytics | 2013 | SIGMOD | 9.8315163e-05 |
| 2,414 | Data Generation using Declarative Constraints | 2011 | SIGMOD | 8.5054535e-05 |
| 4,093 | Generating Databases for Query Workloads | 2010 | VLDB | 6.809504e-05 |
| 4,930 | TPC-DI: The First Industry Benchmark for Data Integration | 2014 | VLDB | 6.346494e-05 |
| 7,401 | Myriad: Scalable and Expressive Data Generation | 2012 | VLDB | 5.5338134e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,329 | Relational Data Synthesis using Generative Adversarial Networks: A Design Space Exploration | 2020 | VLDB |
| 2 | 9,035 | Generating Flexible Workloads for Graph Databases | 2016 | VLDB |
| 3 | 12,387 | Synthesizing Data Programs | 2015 | CIDR |
| 4 | 8,966 | Supporting Database Constraints in Synthetic Data Generation based on Generative Adversarial Networks | 2020 | SIGMOD |
| 5 | 6,081 | From Auto-tuning One Size Fits All to Self-designed and Learned Data-intensive Systems | 2019 | SIGMOD |
| 6 | 956 | Flexible Database Generators | 2005 | VLDB |
| 7 | 1,702 | BigBench: Towards an Industry Standard Benchmark for Big Data Analytics | 2013 | SIGMOD |
| 8 | 2,414 | Data Generation using Declarative Constraints | 2011 | SIGMOD |
| 9 | 106 | Quickly Generating Billion-Record Synthetic Databases | 1994 | SIGMOD |
| 10 | 9,206 | DataSynth: Generating Synthetic Data using Declarative Constraints | 2011 | VLDB |