Git is for Data
Summary: Argues Git's UX/ecosystem is ideal for ML dataset management but vanilla Git fails at scale; introduces XetHub, an extension that preserves Git semantics while enabling TB+ repositories. Demonstrates scalable, low‑friction reproducibility and integration with DevOps pipelines. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Yucheng Low (XetData Inc.)
- 2. Rajat Arya (XetData Inc.)
- 3. Ajit Banerjee (XetData Inc.)
- 4. Ann Huang (XetData Inc.)
- 5. Brian Ronan (XetData Inc.)
- 6. Hoyt Koepke (XetData Inc.)
- 7. Joseph Godlewski (XetData Inc.)
- 8. Zach Nation (XetData Inc.)
BibTeX Citation
@inproceedings{low_cidr23,
address = {Amsterdam, Netherlands},
series = {{CIDR} '23},
title = {{Git is for Data}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Low, Yucheng and Arya, Rajat and Banerjee, Ajit and Huang, Ann and Ronan, Brian and Koepke, Hoyt and Godlewski, Joseph and Nation, Zach},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,077 | DataHub: Collaborative Data Science & Dataset Version Management at Scale | 2015 | CIDR | 0.00012269438 |
| 2,000 | Data Management for Data Science: Towards Embedded Analytics | 2020 | CIDR | 9.3336258e-05 |
| 3,614 | Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML | 2020 | CIDR | 7.2568185e-05 |
| 4,592 | Data Platform for Machine Learning | 2019 | SIGMOD | 6.6139071e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,250 | Data Management in Machine Learning: Challenges, Techniques, and Systems | 2017 | SIGMOD |
| 2 | 3,402 | Collaborative Data Analytics with DataHub | 2015 | VLDB |
| 3 | 1,349 | Principles of Dataset Versioning: Exploring the Recreation/Storage Tradeoff | 2015 | VLDB |
| 4 | 1,147 | Data Management Challenges in Production Machine Learning | 2017 | SIGMOD |
| 5 | 13,469 | Towards Scalable Online Machine Learning Collaborations with OpenML | 2021 | VLDB |
| 6 | 2,657 | Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities | 2021 | SIGMOD |
| 7 | 9,245 | Towards Observability for Production Machine Learning Pipelines | 2022 | VLDB |
| 8 | 7,609 | Ease.ml/ci and Ease.ml/meter in Action: Towards Data Management for Statistical Generalization | 2019 | VLDB |
| 9 | 4,592 | Data Platform for Machine Learning | 2019 | SIGMOD |
| 10 | 1,077 | DataHub: Collaborative Data Science & Dataset Version Management at Scale | 2015 | CIDR |