Disentangling representation and policy complexity in reinforcement learning via dual mutual information penalties
Loading...
Authors
ORCID
Other Contributors
Issue Date
Type
Electronic thesis
Thesis
Thesis
Language
en_US
Keywords
Degree
MS
Alternative Title
Abstract
This thesis investigates information-theoretic regularization in reinforcement learning through the lens of mutual information penalties applied to two sequential stages of the perception-action pipeline: a state encoder and a policy. Motivated by cognitive science and neuroscience evidence that biological agents operate near optimal reward-complexity frontiers and that dopaminergic learning signals encode information-processing costs, we develop and evaluate a dual mutual information-regularized actor-critic algorithm in which the encoder penalty $\beta_1 I_t^c(S_t; X_t)$ and the policy penalty $\beta_2 I_t^\pi(X_t; A_t)$ are controlled independently.We first establish a single-penalty baseline incorporating a mutual information penalty on the state-action channel $I(S;A)$ into a standard actor-critic objective. Experiments in an open-room gridworld and a structurally demanding zoned corridor show that moderate regularization ($\beta \leq 1.0$) sustains high-reward behavior while extreme regularization ($\beta = 5.0$) degrades performance, with estimated mutual information resisting suppression throughout the moderate regime.
We then introduce an explicit categorical encoder mapping raw observations to discrete latent codes, and penalize both the encoder's information rate $I(S;X)$ and the policy's information rate $I(X;A)$ with independent coefficients. A $9 \times 9$ sweep over 81 configurations per environment reveals that the two penalties play qualitatively different roles. In both environments, the unregularized baseline spontaneously produces functional encoder representations without explicit regularization, and performance degrades primarily when $\beta_1 \geq 1.0$ or $\beta_2 \geq 1.0$. The decomposition of $I(S;A)$ into $I(S;X)$ and $I(X;A)$ exposes a cross-penalty coupling invisible to the single-penalty framework.
Strong policy compression can indirectly elevate encoder information use, while strong encoder compression can suppress policy information by destroying the state signal the policy exploits.
These results support the claim that dual mutual information penalties provide independent control over two qualitatively distinct stages of the perception-action pipeline.
Description
May2026
School of Science
School of Science
Full Citation
Publisher
Rensselaer Polytechnic Institute, Troy, NY
