DBScholar

Back to papers

Jaql: A Scripting Language for Large Scale Semistructured Data Analysis

Summary: Jaql: declarative scripting for large-scale semistructured data on Hadoop MapReduce. JSON-like flexible model with schema-on-read; higher-order functions and reusable modules; varying abstraction with compiler rewrites for parallel execution. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
habb4aed386ff2088
Venue
VLDB
Year
2011
Pagerank
0.00012371105
Overall Rank
1,037 | 93.04%
DOI
10.14778/3402755.3402768

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{beyer_vldb11,
        title = {{Jaql: A Scripting Language for Large Scale Semistructured Data Analysis}},
        author = {Beyer, Kevin S. and Eltabakh, Mohamed and Ercegovac, Vuk and Gemulla, Rainer and Kanne, Carl-Christian and Ozcan, Fatma and Balmin, Andrey and Shekita, Eugene J.},
        journal = {PVLDB},
        series = {{VLDB} '11},
        volume = {4},
        number = {12},
        pages = {1272--1283},
        doi = {10.14778/3402755.3402768},
        url = {https://doi.org/10.14778/3402755.3402768},
        year = {2011}
}

Incoming Citations (Sorted by Pagerank)

Showing 27 of 27 citing papers.

Rank Citing Paper Year Venue Pagerank
1,082 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012118261
1,745 Sinew: A SQL System for Multi-Structured Data 2014 SIGMOD 9.7324334e-05
1,925 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3643089e-05
2,197 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 8.8740089e-05
2,437 Mison: A Fast JSON Parser for Data Analytics 2017 VLDB 8.4668658e-05
2,573 Minimal MapReduce Algorithms 2013 SIGMOD 8.2782871e-05
2,753 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0513412e-05
2,854 JSON Data Management – Supporting Schema-less Development in RDBMS 2014 SIGMOD 7.93571e-05
3,339 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.4045097e-05
3,440 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.2954357e-05
3,505 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2474175e-05
3,781 M3R: Increased Performance for In-Memory Hadoop Jobs 2012 VLDB 7.0223677e-05
4,069 Dynamically Optimizing Queries over Large Scale Data Platforms 2014 SIGMOD 6.8184364e-05
4,820 Continuous Cloud-Scale Query Optimization and Processing 2013 VLDB 6.3935932e-05
6,410 Schemas and Types for JSON Data: from Theory to Practice 2019 SIGMOD 5.7915684e-05
6,683 Semistructured Models, Queries and Algebras in the Big Data Era 2016 SIGMOD 5.7079679e-05
7,750 Rumble: Data Independence for Large Messy Data Sets 2021 VLDB 5.4585103e-05
7,870 Handling Environments in a Nested Relational Algebra with Combinators and an Implementation in a Verified Query Compiler 2017 SIGMOD 5.4340686e-05
8,840 Graph Data Models, Query Languages and Programming Paradigms 2018 VLDB 5.2656278e-05
9,264 QMapper for Smart Grid: Migrating SQL-based Application to Hive 2015 SIGMOD 5.2032182e-05
9,545 GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example 2023 SIGMOD 5.1600923e-05
9,698 Versatile Optimization of UDF-heavy Data Flows with Sofa 2014 SIGMOD 5.1377347e-05
10,071 ReCG: Bottom-Up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework 2024 VLDB 5.0849989e-05
11,714 dsJSON: A Distributed SQL JSON Processor 2023 SIGMOD 4.9769913e-05
12,557 Next Generation Data Analytics at IBM Research 2013 VLDB 4.9769913e-05
12,603 Declarative Error Management for Robust Data-Intensive Applications 2012 SIGMOD 4.9769913e-05
12,615 Surfacing Time-critical Insights from Social Media 2012 SIGMOD 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 11 of 11 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers