Efficient post-processing techniques for foundation models

Loading...
Thumbnail Image

ORCID

https://orcid.org/0000-0001-7133-3775

Issue Date

Type

Electronic thesis
Thesis

Language

en_US

Degree

PhD

Research Projects

Organizational Units

Journal Issue

Alternative Title

Abstract

Foundation models, such as large language models (LLMs) and large vision models (LVMs), are often fine-tuned for various downstream tasks. Popular fine-tuning techniques include Reinforcement Learning from Human Feedback (RLHF) and supervised fine-tuning (SFT). Fine-tuning approaches, including RLHF and SFT, involve modifying model parameters, which is computationally expensive and may lead to unintended degradation of model performance, such as overfitting to specific biases or reducing response diversity. \noindent This thesis addresses three key challenges in the robustness and efficiency of foundation models:\begin{itemize} \item Learning interpretable disentangled representations (Chapter \ref{chap:pisco}): We propose a computationally cheap post-processing technique to separate style and content features in LVMs, leading to representations that improve out-of-distribution (OOD) generalization. \item Efficiently aligning LLMs (Chapter \ref{chap:aligners}): We introduce a lightweight pipeline that generates synthetic data to train aligners and inspectors, enabling on-demand alignment of any LLM. Inspectors are lightweight BERT models used to determine when an output needs to be aligned, and aligners are small LLMs used to perform alignment with respect to a particular targeted human value. We use AlpacaEval and PairRM, two standard automatic LLM evaluators, to establish the competitive advantages of our approach against baseline LLMs and alternative alignment procedures. \item Mitigating style-induced prompt brittleness in LLMs (Chapter \ref{chap:mof}): We present a novel mixture of formats (MOF) prompting strategy that incorporates diverse stylistic variations in few-shot examples, enhancing robustness across different models and task domains. MOF uses simple modifications to few-shot prompting, so it avoids the cost of post-processing of the LLM outputs that is incurred by other approaches to mitigating style-induced prompt brittleness.\end{itemize} Each of these approaches is designed as a post-processing technique, meaning they do not require modifying the original foundation model’s parameters. As a result, they are efficient and they provide performance improvements that are competitive with or exceed that of alternative solutions to these challenges.

Description

May2025
School of Science

Full Citation

Publisher

Rensselaer Polytechnic Institute, Troy, NY

Terms of Use

Journal

Volume

Issue

PubMed ID

DOI

ISSN

EISSN

Collections

Endorsement

Review

Supplemented By

Referenced By