Data Ingestion for the Connected World
Summary: Positions traditional batch ETL as the bottleneck for timely analytics in connected/IoT settings and advocates a push-based streaming-ETL architecture to ensure scalable, correct ingestion. Implements this with Kafka + transactional S-Store + BigDAWG and a new ingestion-optimized time-series DB. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. John Meehan (Brown University)
- 2. Cansu Aslantas (Brown University)
- 3. Stan Zdonik (Brown University)
- 4. Nesime Tatbul (Intel; Massachusetts Institute of Technology)
- 5. Jiang Du (University of Toronto)
BibTeX Citation
@inproceedings{meehan_cidr17,
address = {Amsterdam, Netherlands},
series = {{CIDR} '17},
title = {{Data Ingestion for the Connected World}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Meehan, John and Aslantas, Cansu and Zdonik, Stan and Tatbul, Nesime and Du, Jiang},
year = {2017}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,123 | Monarch: Google’s Planet-Scale In-Memory Time Series Database | 2020 | VLDB | 7.7354933e-05 |
| 6,494 | Extract-Transform-Load for Video Streams | 2023 | VLDB | 5.8639535e-05 |
| 8,377 | Visual Exploration of Time Series Anomalies with Metro-Viz | 2019 | SIGMOD | 5.4388258e-05 |
| 9,366 | MorphStream: Adaptive Scheduling for Scalable Transactional Stream Processing on Multicores | 2023 | SIGMOD | 5.2789847e-05 |
| 9,508 | An IDEA: An Ingestion Framework for Data Enrichment in AsterixDB | 2019 | VLDB | 5.2581623e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 83 | H-Store: A High-Performance, Distributed Main Memory Transaction Processing System | 2008 | VLDB | 0.00036185259 |
| 148 | Gorilla: A Fast, Scalable, In-Memory Time Series Database | 2015 | VLDB | 0.00029250767 |
| 239 | Overview of SciDB: Large Scale Array Storage, Processing and Analysis | 2010 | SIGMOD | 0.00023674329 |
| 612 | Twitter Heron: Stream Processing at Scale | 2015 | SIGMOD | 0.0001573018 |
| 1,915 | S-Store: Streaming Meets Transaction Processing | 2015 | VLDB | 9.4884706e-05 |
| 2,449 | A Demonstration of the BigDAWG Polystore System | 2015 | VLDB | 8.5677637e-05 |
| 4,463 | Consistency in a Stream Warehouse | 2011 | CIDR | 6.6860367e-05 |
| 4,877 | TPC-DI: The First Industry Benchmark for Data Integration | 2014 | VLDB | 6.468584e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,494 | Extract-Transform-Load for Video Streams | 2023 | VLDB |
| 2 | 11,017 | Streaming View: An Efficient Data Processing Engine for Modern Real-time Data Warehouse of Alibaba Cloud | 2025 | VLDB |
| 3 | 3,082 | Analytics in Motion: High Performance Event-Processing AND Real-Time Analytics in the Same Database | 2015 | SIGMOD |
| 4 | 6,734 | Liquid: Unifying Nearline and Offline Big Data Integration | 2015 | CIDR |
| 5 | 12,006 | Query-able Kafka: An agile data analytics pipeline for mobile wireless networks | 2017 | VLDB |
| 6 | 9,258 | Data Stream Warehousing in Tidalrace | 2015 | CIDR |
| 7 | 2,449 | A Demonstration of the BigDAWG Polystore System | 2015 | VLDB |
| 8 | 3,532 | Scalable Distributed Stream Join Processing | 2015 | SIGMOD |
| 9 | 361 | The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing | 2015 | VLDB |
| 10 | 9,640 | Supporting Scalable Analytics with Latency Constraints | 2015 | VLDB |