Debugging Missing Answers for Spark Queries over Nested Data with Breadcrumb
Summary: Breadcrumb provides query-based explanations for missing Spark results on nested data, pinpointing the operators responsible for the absence. It scales to big data and handles schema-semantic errors in nested/de-normalized queries, guiding fixes. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Ralf Diestelkämper
- 2. Seokki Lee
- 3. Boris Glavic
- 4. Melanie Herschel
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,839 | Why Not Yet: Fixing a Top-k Ranking that Is Not Fair to Individuals | 2023 | VLDB | 5.3073497e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 487 | Why Not? | 2009 | SIGMOD | 0.00022030123 |
| 593 | Lineage Tracing for General Data Warehouse Transformations | 2001 | VLDB | 0.0001952223 |
| 968 | Schema and Ontology Matching with COMA++ | 2005 | SIGMOD | 0.00014941789 |
| 6,974 | NLProveNAns: Natural Language Provenance for Non-Answers | 2018 | VLDB | 4.8725777e-05 |
| 7,678 | To Not Miss the Forest for the Trees - A Holistic Approach for Explaining Missing Answers over Nested Data | 2021 | SIGMOD | 4.6768157e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| Overall Rank | Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,872 | LEAP: A Low-cost Spark SQL Query Optimizer using Pairwise Comparison | 2025 | VLDB | 4.1905499e-05 |
| 11,408 | SparkCAD: Caching Anomalies Detector for Spark Applications | 2022 | VLDB | 4.1905499e-05 |
| 6,661 | Scalable Querying of Nested Data | 2021 | VLDB | 4.9663934e-05 |
| 8,582 | A Demonstration of DLBD: Database Logic Bug Detection System | 2023 | VLDB | 4.4868313e-05 |
| 1,481 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.00011715754 |
| 11,199 | QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark | 2023 | SIGMOD | 4.1905499e-05 |
| 3,207 | Big Data Analytics with Datalog Queries on Spark | 2016 | SIGMOD | 7.3847098e-05 |
| 5,108 | Debugging Big Data Analytics in Spark with BigDebug | 2017 | SIGMOD | 5.6872497e-05 |
| 11,667 | Capturing and Querying Structural Provenance in Spark with Pebble | 2019 | SIGMOD | 4.1905499e-05 |
| 7,678 | To Not Miss the Forest for the Trees - A Holistic Approach for Explaining Missing Answers over Nested Data | 2021 | SIGMOD | 4.6768157e-05 |