Reducing Ambiguity in Json Schema Discovery
Summary: Reduces ambiguity in Json schema discovery for ad-hoc data and APIs with Jxplain, a heuristic-driven algorithm that constrains the schema space. Slightly slower than competitors but yields far more precise schemas, reducing validation false positives. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. William Spoth (State University of New York at Buffalo)
- 2. Oliver Kennedy (State University of New York at Buffalo)
- 3. Ying Lu (Oracle)
- 4. Beda Hammerschmidt (Oracle)
- 5. Zhen Hua Liu (Oracle)
BibTeX Citation
@inproceedings{spoth_sigmod21,
title = {{Reducing Ambiguity in Json Schema Discovery}},
author = {Spoth, William and Kennedy, Oliver and Lu, Ying and Hammerschmidt, Beda and Liu, Zhen Hua},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3452801},
url = {https://dl.acm.org/doi/10.1145/3448016.3452801},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,176 | Towards Theory for Real-World Data | 2022 | PODS | 5.2127522e-05 |
| 9,362 | Exploring Exploratory Querying | 2025 | VLDB | 5.1868213e-05 |
| 10,066 | ReCG: Bottom-Up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework | 2024 | VLDB | 5.0874072e-05 |
| 11,601 | Partition, Don’t Sort! Compression Boosters for Cloud Data Ingestion Pipelines | 2024 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 50 | DataGuides: Enabling Query Formulation and Optimization in Semistructured Databases | 1997 | VLDB | 0.00043130126 |
| 1,258 | Extracting Schema from Semistructured Data | 1998 | SIGMOD | 0.00011306716 |
| 1,969 | Information-Theoretic Tools for Mining Database Structure from Large Data Sets | 2004 | SIGMOD | 9.2935435e-05 |
| 2,073 | Graceful Database Schema Evolution: the PRISM Workbench | 2008 | VLDB | 9.0852947e-05 |
| 2,797 | Inferring XML Schema Definitions from XML Data | 2007 | VLDB | 7.9969811e-05 |
| 3,242 | Inference of Concise DTDs from XML Data | 2006 | VLDB | 7.4991762e-05 |
| 3,723 | Schema Management for Document Stores | 2015 | VLDB | 7.0706017e-05 |
| 4,088 | Automatic Generation of Normalized Relational Schemas from Nested Key-Value Data | 2016 | SIGMOD | 6.8148416e-05 |
| 4,476 | LegoDB: Customizing Relational Storage for XML Documents | 2002 | VLDB | 6.5833438e-05 |
| 7,456 | Closing the functional and Performance Gap between SQL and NoSQL | 2016 | SIGMOD | 5.5223997e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,088 | Automatic Generation of Normalized Relational Schemas from Nested Key-Value Data | 2016 | SIGMOD |
| 2 | 3,723 | Schema Management for Document Stores | 2015 | VLDB |
| 3 | 3,916 | JSON Tiles: Fast Analytics on Semi-Structured Data | 2021 | SIGMOD |
| 4 | 13,583 | Blaze: Compiling JSON Schema for 10x Faster Validation | 2026 | VLDB |
| 5 | 6,407 | Schemas and Types for JSON Data: from Theory to Practice | 2019 | SIGMOD |
| 6 | 10,313 | Witness Generation for JSON Schema | 2022 | VLDB |
| 7 | 3,354 | JSON: Data model, Query languages and Schema specification | 2017 | PODS |
| 8 | 11,040 | Streaming Validation of JSON Documents Against Schemas | 2026 | VLDB |
| 9 | 10,066 | ReCG: Bottom-Up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework | 2024 | VLDB |
| 10 | 12,073 | JSON Schema Matching: Empirical Observations | 2020 | SIGMOD |