Causal bandits with soft intervention: a unified framework

Loading...
Thumbnail Image

ORCID

https://orcid.org/0009-0003-1314-7773

Issue Date

Type

Electronic thesis
Thesis

Language

en_US

Degree

PhD

Research Projects

Organizational Units

Journal Issue

Alternative Title

Abstract

Causal bandits (CBs) provide a principled framework for designing sequential experiments to optimize the utility of causal networks. Interventions are actions that modify the causal mechanisms and affect the utility. By adaptively selecting interventions over time, CB algorithms gradually uncover causal relationships, improving experimental efficiency and cumulative utility. Despite strong theoretical progress, deploying CBs in real systems is often hindered by restrictive assumptions on (i) stationarity of causal mechanisms, (ii) prior knowledge of the causal graph, (iii) linearity of causal relationships, and (iv) direct observability of causal variables. This dissertation focuses on soft interventions, the most flexible type of intervention, and develops algorithms and regret bounds that systematically relax these assumptions, bringing CB theory closer to practical settings. First, this dissertation considers time-varying causal mechanisms (Chapter 3), moving beyond the common assumption that causal models remain fixed over time. We show that even tiny deviations over time can significantly affect the performance of existing CB algorithms, leading to linear regret growth. To handle this, we propose a robust CB algorithm for linear causal systems with unknown structural variations and prove that it achieves near-optimal regret. We show that the regret depends on the graph only through the in-degree and causal path length. Second, this dissertation addresses the unknown causal graph structure problem (Chapter 4), which goes beyond the typical setting where the graph is fully known in advance. We develop a CB algorithm for stochastic soft interventions on graphs with minimal structural information and post-intervention distributions. We provide both upper and lower regret bounds, showing how the lack of graph knowledge impacts performance. We also show the effectiveness of the proposed algorithm compared to existing algorithms. Our results also show that as the number of interaction rounds increases, the dependence of the algorithm on the number of nodes decreases. Third, this dissertation relaxes the assumption of linearity in causal models (Chapter 5) and investigates how regret grows with graph skeleton parameters and function complexity. We introduce a generalized CB framework that works with nonlinear causal models from Lipschitz-continuous function classes, including quadratic and neural networks. We also extend interventions to have arbitrary granularity. We provide general upper and lower regret bounds in terms of graph structure, two function class complexity (eluder dimension and covering number). For specific cases, such as neural networks, polynomial, and linear models, we further refine these bounds to show how nonlinearity affects CB performance. Fourth, this dissertation goes beyond the assumption in all previous CB literature that causal variables are directly observed by introducing reward-oriented causal representation learning (RO-CRL) (Chapter 6). In this setting, the causal variables are latent and only high-dimensional observations are available through an unknown transformation. RO-CRL integrates representation learning with sequential intervention design in causal networks, learning the latent graph and variables is needed but only to the coarsest extent necessary to optimize the downstream reward. We develop an adaptive exploration algorithm and show how the regret is impacted by the lack of latent observations. Finally, we outline future directions (Chapter 7), including: (i) extending RO-CRL to more general graph families and observation mappings; (ii) studying objectives beyond expected reward, incorporating higher-order moments and risk-sensitive criteria; and (iii) exploring task-oriented robotics applications, such as learning intervention policies for robot motion training from high-dimensional sensory observations. These directions aim to close the gap between theoretical CB models and their practical deployment in complex real-world systems.

Description

May2026
School of Engineering

Full Citation

Publisher

Rensselaer Polytechnic Institute, Troy, NY

Terms of Use

Journal

Volume

Issue

PubMed ID

DOI

ISSN

EISSN

Endorsement

Review

Supplemented By

Referenced By