Just can't get enough - Synthesizing Big Data
Summary: DBSynth automatically generates realistic, large-scale synthetic data from schema and sample data, extracting value-level features and Markov models. As an extension to PDGF, it enables fast, scalable generation across formats (CSV/JSON/XML/SQL) for big-data benchmarks (TPC-DI/BigBench) with automatic dictionaries and multi-core speedups. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Tilmann Rabl (University of Toronto)
- 2. Manuel Danisch (Bankmark UG)
- 3. Michael Frank (Bankmark UG)
- 4. Sebastian Schindler (Bankmark UG)
- 5. Hans-Arno Jacobsen (University of Toronto)
BibTeX Citation
@inproceedings{rabl_sigmod15,
title = {{Just can't get enough - Synthesizing Big Data}},
author = {Rabl, Tilmann and Danisch, Manuel and Frank, Michael and Schindler, Sebastian and Jacobsen, Hans-Arno},
series = {{SIGMOD} '15},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2723372.2735378},
url = {https://dl.acm.org/doi/10.1145/2723372.2735378},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,755 | Synthesizing Linked Data Under Cardinality and Integrity Constraints | 2021 | SIGMOD | 5.7830386e-05 |
| 7,866 | Dscaler: Synthetically Scaling A Given Relational Database | 2016 | VLDB | 5.5281143e-05 |
| 8,037 | HYDRA: A Dynamic Big Data Regenerator | 2018 | VLDB | 5.5025567e-05 |
| 9,985 | Projection-Compliant Database Generation | 2022 | VLDB | 5.1842045e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 105 | Quickly Generating Billion-Record Synthetic Databases | 1994 | SIGMOD | 0.00033877899 |
| 734 | QAGen: Generating Query-Aware Test Databases | 2007 | SIGMOD | 0.00014525584 |
| 938 | Flexible Database Generators | 2005 | VLDB | 0.00013089351 |
| 1,578 | Simple and Realistic Data Generation | 2006 | VLDB | 0.00010309264 |
| 1,693 | BigBench: Towards an Industry Standard Benchmark for Big Data Analytics | 2013 | SIGMOD | 9.9965799e-05 |
| 2,369 | Data Generation using Declarative Constraints | 2011 | SIGMOD | 8.682429e-05 |
| 3,994 | Generating Databases for Query Workloads | 2010 | VLDB | 6.9686731e-05 |
| 4,877 | TPC-DI: The First Industry Benchmark for Data Integration | 2014 | VLDB | 6.468584e-05 |
| 7,248 | Myriad: Scalable and Expressive Data Generation | 2012 | VLDB | 5.6635058e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,261 | Relational Data Synthesis using Generative Adversarial Networks: A Design Space Exploration | 2020 | VLDB |
| 2 | 8,867 | Generating Flexible Workloads for Graph Databases | 2016 | VLDB |
| 3 | 8,815 | Supporting Database Constraints in Synthetic Data Generation based on Generative Adversarial Networks | 2020 | SIGMOD |
| 4 | 12,088 | Synthesizing Data Programs | 2015 | CIDR |
| 5 | 5,978 | From Auto-tuning One Size Fits All to Self-designed and Learned Data-intensive Systems | 2019 | SIGMOD |
| 6 | 938 | Flexible Database Generators | 2005 | VLDB |
| 7 | 1,693 | BigBench: Towards an Industry Standard Benchmark for Big Data Analytics | 2013 | SIGMOD |
| 8 | 2,369 | Data Generation using Declarative Constraints | 2011 | SIGMOD |
| 9 | 105 | Quickly Generating Billion-Record Synthetic Databases | 1994 | SIGMOD |
| 10 | 9,028 | DataSynth: Generating Synthetic Data using Declarative Constraints | 2011 | VLDB |