DBScholar

Back to papers

Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis]

Summary: Empirical study across image, video, audio, text, and tabular FL using DataNoiseGenerator. Shows FL is more noise-sensitive than centralized learning: server aggregation amplifies divergent noisy-client updates, motivating decentralized data cleaning. (summarized by gpt-5.6-luna on Jul 26 2026)

Paper ID
h5b1df98ea9626e31
Venue
SIGMOD
Year
2026
Pagerank
4.9793485e-05
Overall Rank
10,515 | 29.31%
DOI
10.1145/3802124

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{hu_sigmod26,
        title = {{Understanding the Impact of Data Noise in Federated Learning: [Experiments \& Analysis]}},
        author = {Hu, Jinming and Gu, Jiahao and Ploch, Kenta and Wang, Hao and Wang, Jingxian and Wu, Wentao and Zhang, Qizhen},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3802124},
        url = {https://dl.acm.org/doi/10.1145/3802124},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033690989
188 A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification 2005 SIGMOD 0.00025872962
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017590977
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014694048
1,043 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00012335114
1,099 KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing 2015 SIGMOD 0.00012037058
1,720 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.7965659e-05
1,847 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.5120573e-05
3,307 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.4414303e-05
4,524 DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data 2023 SIGMOD 6.5621504e-05
4,998 Descriptive and Prescriptive Data Cleaning 2014 SIGMOD 6.3207886e-05
5,572 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0802555e-05
5,643 Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks 2023 VLDB 6.0538129e-05
6,381 Qualitative Data Cleaning 2016 VLDB 5.8066959e-05
7,538 Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines 2023 SIGMOD 5.4995874e-05
8,224 Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness 2024 VLDB 5.3747673e-05
Previous Page 1 / 1 Next

Semantically Similar Papers