DBScholar

Back to papers

GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example

Summary: GIO learns mapping from raw text to matrix/frame layouts from a sample and auto-generates a fast multi-threaded reader. Supports CSV/LibSVM/MatrixMarket and nested formats; mappings and readers editable; competitive with hand-written parsers. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h18aaf8ef75f909bf
Venue
SIGMOD
Year
2023
Pagerank
5.1599622e-05
Overall Rank
9,549 | 35.80%
DOI
10.1145/3589265

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{fathollahzadeh_sigmod23,
        title = {{GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example}},
        author = {Fathollahzadeh, Saeed and Boehm, Matthias},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3589265},
        url = {https://dl.acm.org/doi/10.1145/3589265},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
5,410 ELEET: Efficient Learned Query Execution over Text and Tables 2024 VLDB 6.1408676e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 45 of 45 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055406774
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028579704
182 Schema Mapping as Query Discovery 2000 VLDB 0.00026277842
296 Generic Schema Matching with Cupid 2001 VLDB 0.00021867512
327 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.0002095191
407 COMA - A system for flexible combination of schema matching approaches 2002 VLDB 0.00019026291
480 Clio Grows Up: From Research Prototype to Industrial Tool 2005 SIGMOD 0.00017608007
517 Schema Mappings, Data Exchange, and Metadata Management 2005 PODS 0.00016948466
907 HOT: A Height Optimized Trie Index for Main-Memory Database Systems 2018 SIGMOD 0.00013158824
950 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012895553
972 The Data Civilizer System 2017 CIDR 0.00012763234
1,037 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012376819
1,050 Data-Driven Understanding and Refinement of Schema Mappings 2001 SIGMOD 0.00012298141
1,069 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.00012185253
1,146 Schema and Ontology Matching with COMA++ 2005 SIGMOD 0.00011815973
1,505 Generic Schema Matching, Ten Years Later 2011 VLDB 0.00010454635
1,668 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9371612e-05
1,776 Sample-Driven Schema Mapping 2012 SIGMOD 9.6674087e-05
1,925 Here are my Data Files. Here are my Queries. Where are my Results? 2011 CIDR 9.3670488e-05
1,993 Clio: A Semi-Automatic Tool For Schema Mapping 2001 SIGMOD 9.2222424e-05
2,092 Sato: Contextual Semantic Type Detection in Tables 2020 VLDB 9.0626928e-05
2,244 Characterizing Schema Mappings via Data Examples 2010 PODS 8.7653687e-05
2,328 Parallel In-Situ Data Processing with Speculative Loading 2014 SIGMOD 8.62934e-05
2,347 Filter Before You Parse: Faster Analytics on Raw Data with Sparser 2018 VLDB 8.6012176e-05
2,422 Parallel Data Analysis Directly on Scientific File Formats 2014 SIGMOD 8.4891004e-05
2,436 Mison: A Fast JSON Parser for Data Analytics 2017 VLDB 8.4708399e-05
2,735 Speculative Distributed CSV Data Parsing for Big Data Analytics 2019 SIGMOD 8.0761736e-05
3,069 Adaptive Query Processing on RAW Data 2014 VLDB 7.6822697e-05
3,101 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.649219e-05
3,397 Designing and Refining Schema Mappings via Data Examples 2011 SIGMOD 7.3366398e-05
3,418 Data Profiling – A Tutorial 2017 SIGMOD 7.3210362e-05
3,505 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2508161e-05
3,916 JSON Tiles: Fast Analytics on Semi-Structured Data 2021 SIGMOD 6.9285787e-05
4,294 NoDB in Action: Adaptive Query Processing on Raw Data 2012 VLDB 6.6819514e-05
4,832 ReCache: Reactive Caching for Fast Analytics over Heterogeneous Data 2018 VLDB 6.3895407e-05
5,100 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.27349e-05
5,398 EIRENE: Interactive Design and Refinement of Schema Mappings via Data Examples 2011 VLDB 6.1481166e-05
6,063 ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data 2020 VLDB 5.8982522e-05
6,407 Schemas and Types for JSON Data: from Theory to Practice 2019 SIGMOD 5.7942821e-05
6,441 Grizzly: Efficient Stream Processing Through Adaptive Query Compilation 2020 SIGMOD 5.7834762e-05
7,744 Rumble: Data Independence for Large Messy Data Sets 2021 VLDB 5.4610663e-05
7,839 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4432099e-05
7,927 Scalable Structural Index Construction for JSON Analytics 2021 VLDB 5.425609e-05
8,381 Automated Migration of Hierarchical Data to Relational Tables using Programming-by-Example 2018 VLDB 5.3428083e-05
10,313 Witness Generation for JSON Schema 2022 VLDB 5.0379844e-05
Previous Page 1 / 1 Next

Semantically Similar Papers