DBScholar

Back to papers

GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example

Summary: GIO learns mapping from raw text to matrix/frame layouts from a sample and auto-generates a fast multi-threaded reader. Supports CSV/LibSVM/MatrixMarket and nested formats; mappings and readers editable; competitive with hand-written parsers. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h18aaf8ef75f909bf
Venue
SIGMOD
Year
2023
Pagerank
5.1600923e-05
Overall Rank
9,545 | 35.85%
DOI
10.1145/3589265

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{fathollahzadeh_sigmod23,
        title = {{GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example}},
        author = {Fathollahzadeh, Saeed and Boehm, Matthias},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3589265},
        url = {https://dl.acm.org/doi/10.1145/3589265},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
5,208 ELEET: Efficient Learned Query Execution over Text and Tables 2024 VLDB 6.2254341e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 45 of 45 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055384955
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028568843
182 Schema Mapping as Query Discovery 2000 VLDB 0.00026265768
296 Generic Schema Matching with Cupid 2001 VLDB 0.00021857869
327 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.00020942751
407 COMA - A system for flexible combination of schema matching approaches 2002 VLDB 0.00019017823
481 Clio Grows Up: From Research Prototype to Industrial Tool 2005 SIGMOD 0.00017600067
517 Schema Mappings, Data Exchange, and Metadata Management 2005 PODS 0.0001694094
904 HOT: A Height Optimized Trie Index for Main-Memory Database Systems 2018 SIGMOD 0.00013170142
950 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012889569
972 The Data Civilizer System 2017 CIDR 0.00012757732
1,037 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012371105
1,050 Data-Driven Understanding and Refinement of Schema Mappings 2001 SIGMOD 0.00012292764
1,070 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.00012179575
1,146 Schema and Ontology Matching with COMA++ 2005 SIGMOD 0.00011810563
1,505 Generic Schema Matching, Ten Years Later 2011 VLDB 0.00010449986
1,669 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9324573e-05
1,776 Sample-Driven Schema Mapping 2012 SIGMOD 9.6631141e-05
1,927 Here are my Data Files. Here are my Queries. Where are my Results? 2011 CIDR 9.362697e-05
1,995 Clio: A Semi-Automatic Tool For Schema Mapping 2001 SIGMOD 9.2179352e-05
2,093 Sato: Contextual Semantic Type Detection in Tables 2020 VLDB 9.0593586e-05
2,246 Characterizing Schema Mappings via Data Examples 2010 PODS 8.7612699e-05
2,331 Parallel In-Situ Data Processing with Speculative Loading 2014 SIGMOD 8.6253336e-05
2,348 Filter Before You Parse: Faster Analytics on Raw Data with Sparser 2018 VLDB 8.5972138e-05
2,423 Parallel Data Analysis Directly on Scientific File Formats 2014 SIGMOD 8.4851588e-05
2,437 Mison: A Fast JSON Parser for Data Analytics 2017 VLDB 8.4668658e-05
2,735 Speculative Distributed CSV Data Parsing for Big Data Analytics 2019 SIGMOD 8.0731466e-05
3,071 Adaptive Query Processing on RAW Data 2014 VLDB 7.6789108e-05
3,103 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6456038e-05
3,397 Designing and Refining Schema Mappings via Data Examples 2011 SIGMOD 7.3332055e-05
3,419 Data Profiling – A Tutorial 2017 SIGMOD 7.3175997e-05
3,505 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2474175e-05
3,917 JSON Tiles: Fast Analytics on Semi-Structured Data 2021 SIGMOD 6.925328e-05
4,294 NoDB in Action: Adaptive Query Processing on Raw Data 2012 VLDB 6.6788179e-05
4,834 ReCache: Reactive Caching for Fast Analytics over Heterogeneous Data 2018 VLDB 6.3865472e-05
5,102 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.2705617e-05
5,403 EIRENE: Interactive Design and Refinement of Schema Mappings via Data Examples 2011 VLDB 6.1452376e-05
6,064 ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data 2020 VLDB 5.8954903e-05
6,410 Schemas and Types for JSON Data: from Theory to Practice 2019 SIGMOD 5.7915684e-05
6,444 Grizzly: Efficient Stream Processing Through Adaptive Query Compilation 2020 SIGMOD 5.7807677e-05
7,750 Rumble: Data Independence for Large Messy Data Sets 2021 VLDB 5.4585103e-05
7,843 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4406331e-05
7,931 Scalable Structural Index Construction for JSON Analytics 2021 VLDB 5.4230698e-05
8,387 Automated Migration of Hierarchical Data to Relational Tables using Programming-by-Example 2018 VLDB 5.3403083e-05
10,323 Witness Generation for JSON Schema 2022 VLDB 5.0356287e-05
Previous Page 1 / 1 Next

Semantically Similar Papers