Interpretable and generalizable machine learning methods for sequence modeling and decision making

Loading...
Thumbnail Image

ORCID

Issue Date

Type

Electronic thesis
Thesis

Language

en_US

Degree

PhD

Research Projects

Organizational Units

Journal Issue

Alternative Title

Abstract

Machine learning has achieved remarkable success across diverse applications. However, the black-box nature of deep neural networks obscures how predictions are generated. This lack of transparency hinders our understanding of internal model mechanisms and limits deployment in safety-critical domains. Concurrently, large pretrained models, which are also known as foundation models (FMs), have revolutionized natural language processing and computer vision. Yet, extending FMs to real-valued sequential data, such as time series, remains an ongoing challenge due to a scarcity of high-quality data and a lack of tailored modeling techniques. To address these dual challenges of opacity and scalability, this dissertation introduces novel machine learning methods for sequence modeling, with applications spanning time-series analysis, sequential decision-making, control systems, and robotics. The central idea of this work is that achieving interpretable and generalizable sequence modeling requires learning universal abstractions and concepts. Rather than training deep neural networks for end-to-end prediction, we propose predicting a set of abstracted, potentially interpretable concepts that distill essential information from the input data. By utilizing these concepts as an information bottleneck or as universal prototypes, final predictions are obtained from downstream models. This approach facilitates the development of scalable Neuro-Symbolic models that generate high-quality, distilled, and interpretable features. To demonstrate this framework, this dissertation addresses three essential tasks that together form a comprehensive workflow for sequential problems: First, we explore interpretable methods for time-series analysis. By treating time-domain shapes as interpretable concepts, we introduce a supervised approach for time-series classification using shapelets. This is subsequently generalized into a unified pretrained model that learns cross-domain shape-level concepts, and then further leveraged in the development of large-scale foundation models for general time-series analysis. Second, we study interpretable methods in dynamical systems modeling with neural networks and sequential decision-making. We demonstrate that models trained on time-series data can effectively represent dynamical systems and derive optimal control solutions. Furthermore, we establish that training accurate neural models to estimate dynamical systems requires explicit techniques, such as multi-step predictive training or physics-based regularization. Finally, we apply our concept-based modeling approach to robot learning, where embodied agents process high-dimensional observations to interact with the physical world. Building on our prior findings in time series foundation models, we present a novel method for learning universal, multi-scale latent actions from diverse robot manipulation videos. Specifically, we introduce a modular Vision-Language-Action architecture comprising an action tokenizer, an autoregressive coarse-to-fine policy, and a controller. By compressing visual dynamics into hierarchical latent actions, this framework infuses the concept of planning and encourages the emergence of reasoning ability in the unified latent action space. Ultimately, this dissertation demonstrates that concept-centric abstractions provide a unifying, interpretable, and scalable pathway for modeling complex sequential data across domains.

Description

May2026
School of Engineering

Full Citation

Publisher

Rensselaer Polytechnic Institute, Troy, NY

Terms of Use

Journal

Volume

Issue

PubMed ID

DOI

ISSN

EISSN

Endorsement

Review

Supplemented By

Referenced By