DBScholar

Back to papers

Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis]

Summary: Empirical study across image, video, audio, text, and tabular FL using DataNoiseGenerator. Shows FL is more noise-sensitive than centralized learning: server aggregation amplifies divergent noisy-client updates, motivating decentralized data cleaning. (summarized by gpt-5.6-luna on Jul 26 2026)

Paper ID
h5b1df98ea9626e31
Venue
SIGMOD
Year
2026
Pagerank
4.9769913e-05
Overall Rank
10,526 | 29.26%
DOI
10.1145/3802124
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{hu_sigmod26,
        title = {{Understanding the Impact of Data Noise in Federated Learning: [Experiments \& Analysis]}},
        author = {Hu, Jinming and Gu, Jiahao and Ploch, Kenta and Wang, Hao and Wang, Jingxian and Wu, Wentao and Zhang, Qizhen},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3802124},
        url = {https://dl.acm.org/doi/10.1145/3802124},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033676943
188 A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification 2005 SIGMOD 0.00025861123
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017584249
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014687805
1,044 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00012329478
1,098 KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing 2015 SIGMOD 0.00012031983
1,722 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.7921604e-05
1,848 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.5075544e-05
3,308 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.4379101e-05
4,525 DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data 2023 SIGMOD 6.559044e-05
5,000 Descriptive and Prescriptive Data Cleaning 2014 SIGMOD 6.3194405e-05
5,574 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0773771e-05
5,645 Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks 2023 VLDB 6.0509471e-05
6,384 Qualitative Data Cleaning 2016 VLDB 5.8039496e-05
7,544 Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines 2023 SIGMOD 5.496984e-05
8,230 Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness 2024 VLDB 5.372223e-05
Previous Page 1 / 1 Next

Semantically Similar Papers