Reducing Ambiguity in Json Schema Discovery
Summary: Reduces ambiguity in Json schema discovery for ad-hoc data and APIs with Jxplain, a heuristic-driven algorithm that constrains the schema space. Slightly slower than competitors but yields far more precise schemas, reducing validation false positives. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. William Spoth (State University of New York at Buffalo)
- 2. Oliver Kennedy (State University of New York at Buffalo)
- 3. Ying Lu (Oracle)
- 4. Beda Hammerschmidt (Oracle)
- 5. Zhen Hua Liu (Oracle)
BibTeX Citation
@inproceedings{spoth_sigmod21,
title = {{Reducing Ambiguity in Json Schema Discovery}},
author = {Spoth, William and Kennedy, Oliver and Lu, Ying and Hammerschmidt, Beda and Liu, Zhen Hua},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3452801},
url = {https://dl.acm.org/doi/10.1145/3448016.3452801},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,012 | Towards Theory for Real-World Data | 2022 | PODS | 5.3323444e-05 |
| 9,893 | ReCG: Bottom-Up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework | 2024 | VLDB | 5.1997534e-05 |
| 11,084 | Exploring Exploratory Querying | 2025 | VLDB | 5.093636e-05 |
| 11,274 | Partition, Don’t Sort! Compression Boosters for Cloud Data Ingestion Pipelines | 2024 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 48 | DataGuides: Enabling Query Formulation and Optimization in Semistructured Databases | 1997 | VLDB | 0.00044033592 |
| 1,253 | Extracting Schema from Semistructured Data | 1998 | SIGMOD | 0.00011475793 |
| 1,913 | Information-Theoretic Tools for Mining Database Structure from Large Data Sets | 2004 | SIGMOD | 9.4928569e-05 |
| 2,038 | Graceful Database Schema Evolution: the PRISM Workbench | 2008 | VLDB | 9.2727673e-05 |
| 2,795 | Inferring XML Schema Definitions from XML Data | 2007 | VLDB | 8.115421e-05 |
| 3,185 | Inference of Concise DTDs from XML Data | 2006 | VLDB | 7.6567423e-05 |
| 3,655 | Schema Management for Document Stores | 2015 | VLDB | 7.2226714e-05 |
| 4,014 | Automatic Generation of Normalized Relational Schemas from Nested Key-Value Data | 2016 | SIGMOD | 6.956076e-05 |
| 4,380 | LegoDB: Customizing Relational Storage for XML Documents | 2002 | VLDB | 6.7340205e-05 |
| 7,312 | Closing the functional and Performance Gap between SQL and NoSQL | 2016 | SIGMOD | 5.6480976e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,014 | Automatic Generation of Normalized Relational Schemas from Nested Key-Value Data | 2016 | SIGMOD |
| 2 | 3,655 | Schema Management for Document Stores | 2015 | VLDB |
| 3 | 4,069 | JSON Tiles: Fast Analytics on Semi-Structured Data | 2021 | SIGMOD |
| 4 | 13,293 | Blaze: Compiling JSON Schema for 10x Faster Validation | 2026 | VLDB |
| 5 | 6,283 | Schemas and Types for JSON Data: from Theory to Practice | 2019 | SIGMOD |
| 6 | 10,091 | Witness Generation for JSON Schema | 2022 | VLDB |
| 7 | 3,287 | JSON: Data model, Query languages and Schema specification | 2017 | PODS |
| 8 | 10,592 | Streaming Validation of JSON Documents Against Schemas | 2026 | VLDB |
| 9 | 9,893 | ReCG: Bottom-Up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework | 2024 | VLDB |
| 10 | 11,771 | JSON Schema Matching: Empirical Observations | 2020 | SIGMOD |