DBScholar

Back to papers

GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example

Summary: GIO learns mapping from raw text to matrix/frame layouts from a sample and auto-generates a fast multi-threaded reader. Supports CSV/LibSVM/MatrixMarket and nested formats; mappings and readers editable; competitive with hand-written parsers. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6685
Venue
SIGMOD
Year
2023
Pagerank
5.2692207e-05
Overall Rank
9,436 | 35.27%
DOI
10.1145/3589265

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{fathollahzadeh_sigmod23,
        title = {{GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example}},
        author = {Fathollahzadeh, Saeed and Boehm, Matthias},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3589265},
        url = {https://dl.acm.org/doi/10.1145/3589265},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
6,118 ELEET: Efficient Learned Query Execution over Text and Tables 2024 VLDB 5.9698795e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 45 of 45 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
155 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028713176
180 Schema Mapping as Query Discovery 2000 VLDB 0.00026772368
297 Generic Schema Matching with Cupid 2001 VLDB 0.00022157284
330 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.0002104801
390 COMA - A system for flexible combination of schema matching approaches 2002 VLDB 0.00019382486
469 Clio Grows Up: From Research Prototype to Industrial Tool 2005 SIGMOD 0.00017973925
506 Schema Mappings, Data Exchange, and Metadata Management 2005 PODS 0.0001728171
882 HOT: A Height Optimized Trie Index for Main-Memory Database Systems 2018 SIGMOD 0.0001342403
963 The Data Civilizer System 2017 CIDR 0.00012935145
1,021 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012606673
1,030 Data-Driven Understanding and Refinement of Schema Mappings 2001 SIGMOD 0.00012545209
1,070 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.0001232307
1,138 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012023643
1,158 Schema and Ontology Matching with COMA++ 2005 SIGMOD 0.00011899424
1,497 Generic Schema Matching, Ten Years Later 2011 VLDB 0.00010568859
1,751 Sample-Driven Schema Mapping 2012 SIGMOD 9.838446e-05
1,756 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.8172465e-05
1,927 Here are my Data Files. Here are my Queries. Where are my Results? 2011 CIDR 9.4703074e-05
1,953 Clio: A Semi-Automatic Tool For Schema Mapping 2001 SIGMOD 9.4221557e-05
2,215 Characterizing Schema Mappings via Data Examples 2010 PODS 8.93781e-05
2,223 Sato: Contextual Semantic Type Detection in Tables 2020 VLDB 8.9189986e-05
2,389 Parallel Data Analysis Directly on Scientific File Formats 2014 SIGMOD 8.6439053e-05
2,391 Mison: A Fast JSON Parser for Data Analytics 2017 VLDB 8.6413407e-05
2,414 Filter Before You Parse: Faster Analytics on Raw Data with Sparser 2018 VLDB 8.6078841e-05
2,435 Parallel In-Situ Data Processing with Speculative Loading 2014 SIGMOD 8.5811022e-05
3,062 Adaptive Query Processing on RAW Data 2014 VLDB 7.8037446e-05
3,074 Speculative Distributed CSV Data Parsing for Big Data Analytics 2019 SIGMOD 7.7844208e-05
3,205 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6386536e-05
3,341 Designing and Refining Schema Mappings via Data Examples 2011 SIGMOD 7.5028917e-05
3,387 Data Profiling – A Tutorial 2017 SIGMOD 7.4534663e-05
3,638 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2338361e-05
4,069 JSON Tiles: Fast Analytics on Semi-Structured Data 2021 SIGMOD 6.9276175e-05
4,206 NoDB in Action: Adaptive Query Processing on Raw Data 2012 VLDB 6.8334568e-05
4,764 ReCache: Reactive Caching for Fast Analytics over Heterogeneous Data 2018 VLDB 6.51896e-05
5,051 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.385354e-05
5,274 EIRENE: Interactive Design and Refinement of Schema Mappings via Data Examples 2011 VLDB 6.2888391e-05
6,283 Schemas and Types for JSON Data: from Theory to Practice 2019 SIGMOD 5.9270536e-05
6,349 Grizzly: Efficient Stream Processing Through Adaptive Query Compilation 2020 SIGMOD 5.9049304e-05
7,217 ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data 2020 VLDB 5.668373e-05
7,687 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.5671645e-05
7,766 Scalable Structural Index Construction for JSON Analytics 2021 VLDB 5.549481e-05
8,217 Automated Migration of Hierarchical Data to Relational Tables using Programming-by-Example 2018 VLDB 5.4644927e-05
8,387 Rumble: Data Independence for Large Messy Data Sets 2021 VLDB 5.4364932e-05
10,091 Witness Generation for JSON Schema 2022 VLDB 5.1535135e-05
Previous Page 1 / 1 Next

Semantically Similar Papers