Interpretable and generalizable machine learning methods for sequence modeling and decision making
Loading...
Authors
ORCID
Other Contributors
Issue Date
Type
Electronic thesis
Thesis
Thesis
Language
en_US
Keywords
Degree
PhD
Alternative Title
Abstract
Machine learning has achieved remarkable success across diverse applications. However, the black-box nature of deep neural networks obscures how predictions are generated. This lack of transparency hinders our understanding of internal model mechanisms and limits deployment in safety-critical domains. Concurrently, large pretrained models, which are also known as foundation models (FMs), have revolutionized natural language processing and computer vision. Yet, extending FMs to real-valued sequential data, such as time series, remains an ongoing challenge due to a scarcity of high-quality data and a lack of tailored modeling techniques. To address these dual challenges of opacity and scalability, this dissertation introduces novel machine learning methods for sequence modeling, with applications spanning time-series analysis, sequential decision-making, control systems, and robotics. The central idea of this work is that achieving interpretable and generalizable sequence modeling requires learning universal abstractions and concepts. Rather than training deep neural networks for end-to-end prediction, we propose predicting a set of abstracted, potentially interpretable concepts that distill essential information from the input data. By utilizing these concepts as an information bottleneck or as universal prototypes, final predictions are obtained from downstream models. This approach facilitates the development of scalable Neuro-Symbolic models that generate high-quality, distilled, and interpretable features. To demonstrate this framework, this dissertation addresses three essential tasks that together form a comprehensive workflow for sequential problems: First, we explore interpretable methods for time-series analysis. By treating time-domain shapes as interpretable concepts, we introduce a supervised approach for time-series classification using shapelets. This is subsequently generalized into a unified pretrained model that learns cross-domain shape-level concepts, and then further leveraged in the development of large-scale foundation models for general time-series analysis. Second, we study interpretable methods in dynamical systems modeling with neural networks and sequential decision-making. We demonstrate that models trained on time-series data can effectively represent dynamical systems and derive optimal control solutions. Furthermore, we establish that training accurate neural models to estimate dynamical systems requires explicit techniques, such as multi-step predictive training or physics-based regularization. Finally, we apply our concept-based modeling approach to robot learning, where embodied agents process high-dimensional observations to interact with the physical world. Building on our prior findings in time series foundation models, we present a novel method for learning universal, multi-scale latent actions from diverse robot manipulation videos. Specifically, we introduce a modular Vision-Language-Action architecture comprising an action tokenizer, an autoregressive coarse-to-fine policy, and a controller. By compressing visual dynamics into hierarchical latent actions, this framework infuses the concept of planning and encourages the emergence of reasoning ability in the unified latent action space. Ultimately, this dissertation demonstrates that concept-centric abstractions provide a unifying, interpretable, and scalable pathway for modeling complex sequential data across domains.
Description
May2026
School of Engineering
School of Engineering
Full Citation
Publisher
Rensselaer Polytechnic Institute, Troy, NY
