Reducing Ambiguity in Json Schema Discovery
Summary: Reduces ambiguity in Json schema discovery for ad-hoc data and APIs with Jxplain, a heuristic-driven algorithm that constrains the schema space. Slightly slower than competitors but yields far more precise schemas, reducing validation false positives. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. William Spoth (State University of New York at Buffalo)
- 2. Oliver Kennedy (State University of New York at Buffalo)
- 3. Ying Lu (Oracle)
- 4. Beda Hammerschmidt (Oracle)
- 5. Zhen Hua Liu (Oracle)
BibTeX Citation
@inproceedings{spoth_sigmod21,
title = {{Reducing Ambiguity in Json Schema Discovery}},
author = {Spoth, William and Kennedy, Oliver and Lu, Ying and Hammerschmidt, Beda and Liu, Zhen Hua},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3452801},
url = {https://dl.acm.org/doi/10.1145/3448016.3452801},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,185 | Towards Theory for Real-World Data | 2022 | PODS | 5.2102846e-05 |
| 9,371 | Exploring Exploratory Querying | 2025 | VLDB | 5.1843659e-05 |
| 10,071 | ReCG: Bottom-Up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework | 2024 | VLDB | 5.0849989e-05 |
| 11,607 | Partition, Don’t Sort! Compression Boosters for Cloud Data Ingestion Pipelines | 2024 | VLDB | 4.9769913e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 50 | DataGuides: Enabling Query Formulation and Optimization in Semistructured Databases | 1997 | VLDB | 0.0004311045 |
| 1,259 | Extracting Schema from Semistructured Data | 1998 | SIGMOD | 0.00011301861 |
| 1,970 | Information-Theoretic Tools for Mining Database Structure from Large Data Sets | 2004 | SIGMOD | 9.2892451e-05 |
| 2,075 | Graceful Database Schema Evolution: the PRISM Workbench | 2008 | VLDB | 9.0810058e-05 |
| 2,797 | Inferring XML Schema Definitions from XML Data | 2007 | VLDB | 7.9933558e-05 |
| 3,244 | Inference of Concise DTDs from XML Data | 2006 | VLDB | 7.4956321e-05 |
| 3,725 | Schema Management for Document Stores | 2015 | VLDB | 7.067273e-05 |
| 4,091 | Automatic Generation of Normalized Relational Schemas from Nested Key-Value Data | 2016 | SIGMOD | 6.8116205e-05 |
| 4,478 | LegoDB: Customizing Relational Storage for XML Documents | 2002 | VLDB | 6.5802277e-05 |
| 7,460 | Closing the functional and Performance Gap between SQL and NoSQL | 2016 | SIGMOD | 5.5197855e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,091 | Automatic Generation of Normalized Relational Schemas from Nested Key-Value Data | 2016 | SIGMOD |
| 2 | 3,725 | Schema Management for Document Stores | 2015 | VLDB |
| 3 | 3,917 | JSON Tiles: Fast Analytics on Semi-Structured Data | 2021 | SIGMOD |
| 4 | 13,589 | Blaze: Compiling JSON Schema for 10x Faster Validation | 2026 | VLDB |
| 5 | 6,410 | Schemas and Types for JSON Data: from Theory to Practice | 2019 | SIGMOD |
| 6 | 10,323 | Witness Generation for JSON Schema | 2022 | VLDB |
| 7 | 3,354 | JSON: Data model, Query languages and Schema specification | 2017 | PODS |
| 8 | 11,049 | Streaming Validation of JSON Documents Against Schemas | 2026 | VLDB |
| 9 | 10,071 | ReCG: Bottom-Up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework | 2024 | VLDB |
| 10 | 12,079 | JSON Schema Matching: Empirical Observations | 2020 | SIGMOD |