Large-scale Complex Analytics on Semi-structured Datasets using AsterixDB and Spark
Summary: Bridges AsterixDB and Spark for scalable analytics over semi-structured data. AsterixDB provides ingestion, indexing, geo-spatial, and fuzzy-text querying, while Spark enables downstream machine-learning and graph analysis. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Wail Y. Alkowaileet (King Abdullah University of Science and Technology; Massachusetts Institute of Technology)
- 2. Sattam Alsubaiee (King Abdullah University of Science and Technology; Massachusetts Institute of Technology)
- 3. Michael J. Carey (University of California Irvine)
- 4. Till Westmann (Couchbase)
- 5. Yingyi Bu (Couchbase)
BibTeX Citation
@article{alkowaileet_vldb16,
title = {{Large-scale Complex Analytics on Semi-structured Datasets using AsterixDB and Spark}},
author = {Alkowaileet, Wail Y. and Alsubaiee, Sattam and Carey, Michael J. and Westmann, Till and Bu, Yingyi},
journal = {PVLDB},
series = {{VLDB} '16},
volume = {9},
number = {13},
pages = {1585--1588},
doi = {10.14778/3007263.3007274},
url = {https://doi.org/10.14778/3007263.3007274},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,744 | Rumble: Data Independence for Large Messy Data Sets | 2021 | VLDB | 5.4610663e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 384 | HaLoop: Efficient Iterative Data Processing on Large Clusters | 2010 | VLDB | 0.0001948031 |
| 922 | AsterixDB: A Scalable, Open Source BDMS | 2014 | VLDB | 0.00013068048 |
| 1,875 | Storage Management in AsterixDB | 2014 | VLDB | 9.4576907e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,852 | An LSM-based Tuple Compaction Framework for Apache AsterixDB | 2020 | VLDB |
| 2 | 9,832 | [Demo] Low-latency Spark Queries on Updatable Data | 2019 | SIGMOD |
| 3 | 12,422 | StarDB: A Large-Scale DBMS for Strings | 2015 | VLDB |
| 4 | 7,930 | S2RDF: RDF Querying with SPARQL on Spark | 2016 | VLDB |
| 5 | 4,815 | PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes | 2021 | VLDB |
| 6 | 4,878 | ASTERIX: An Open Source System for "Big Data" Management and Analysis (Demo) | 2012 | VLDB |
| 7 | 1,875 | Storage Management in AsterixDB | 2014 | VLDB |
| 8 | 922 | AsterixDB: A Scalable, Open Source BDMS | 2014 | VLDB |
| 9 | 2,635 | Big Data Analytics with Datalog Queries on Spark | 2016 | SIGMOD |
| 10 | 9,676 | An IDEA: An Ingestion Framework for Data Enrichment in AsterixDB | 2019 | VLDB |