Private and efficient federated learning: empowering data-sensitive sectors

Loading...
Thumbnail Image

ORCID

Issue Date

Type

Electronic thesis
Thesis

Language

en_US

Degree

PhD

Research Projects

Organizational Units

Journal Issue

Alternative Title

Abstract

Federated learning is a decentralized machine learning approach that enables multiple clients to collaboratively train a global model by sharing only model updates (e.g., gradients or model parameters). Despite not transmitting raw data, federated learning remains vulnerable to privacy threats such as inference attacks, where adversaries may reconstruct sensitive information from the shared model updates. To mitigate these risks, differential privacy is employed by adding carefully calibrated noise to the shared components. However, differential privacy often comes with a trade-off in the model accuracy, and balancing privacy and accuracy in a federated learning setting remains a persistent challenge.This difficulty is further compounded by the diversity of the threat models. As different adversary models necessitate distinct differential privacy mechanisms, it is often unclear how to formally prove that an algorithm satisfies a desired privacy guarantee. Furthermore, the reality of heterogeneous data distributions and varying model architectures can exacerbate the privacy-accuracy trade-off. Consequently, there is a critical need for private federated learning frameworks that can adapt to these diverse architectural and data needs without sacrificing accuracy. To bridge these gaps, we introduce four novel privacy-preserving federated learning algorithms designed to effectively balance the trade-off in privacy and accuracy. By addressing these challenges, we demonstrate how federated learning can be made truly effective, private, and efficient for the most data-sensitive sectors. The first part of the thesis focuses on federated learning with vertically distributed data, or vertical federated learning (VFL). The goal is to protect the clients feature data against two different adversary models with differential privacy. We first present a VFL algorithm that integrates the cryptographic protocol multi-party computation to provide better tradeoff between data privacy and model accuracy. We introduce the novel concept of \textit{feature privacy} and provide privacy analysis throughout the training process. We discuss the relationship between privacy budget, convergence error, and communication cost of our algorithm. We next consider decentralized applications and propose a VFL framework that provides verifiability of the aggregation process. The proposed algorithm uses blockchain for transparent aggregation process and eliminates the need for a trusted server. Empirical studies show that both VFL algorithms achieve high accuracy under highly private differential privacy settings. In the second part of the thesis, we shift our focus to federated learning with large language models. We first consider the problem of data heterogeneity and data privacy in federated prompt learning, and we aim to balance the competing goals of personalization, generalization, and privacy. We propose a federated prompt learning algorithm that leverages a low-rank factorization scheme to capture generalization while maintaining a residual term to promote personalization. In experiments, the proposed approach mitigates the impact of privacy noise on the model accuracy while balancing the tradeoff between personalization and generalization. Next, we address the challenge of parameter-efficient fine-tuning large language models in a federated learning system under rigorous privacy settings. Existing parameter-efficient fine-tuning techniques such as low-rank adaptation bring flaw to federated learning systems which results in suboptimal utility under differential privacy. We introduce a federated fine-tuning method that periodically reinitializes the low-rank adapters to better capture the accurate aggregated update. The proposed method reduces the effect of DP noise on model utility by combining the privacy noise with the low-rank approximation scheme, which is supported by extensive results on various language understanding tasks.

Description

May2026
School of Science

Full Citation

Publisher

Rensselaer Polytechnic Institute, Troy, NY

Terms of Use

Journal

Volume

Issue

PubMed ID

DOI

ISSN

EISSN

Endorsement

Review

Supplemented By

Referenced By