Disentangling representation and policy complexity in reinforcement learning via dual mutual information penalties

Loading...
Thumbnail Image

ORCID

Issue Date

Type

Electronic thesis
Thesis

Language

en_US

Degree

MS

Research Projects

Organizational Units

Journal Issue

Alternative Title

Abstract

This thesis investigates information-theoretic regularization in reinforcement learning through the lens of mutual information penalties applied to two sequential stages of the perception-action pipeline: a state encoder and a policy. Motivated by cognitive science and neuroscience evidence that biological agents operate near optimal reward-complexity frontiers and that dopaminergic learning signals encode information-processing costs, we develop and evaluate a dual mutual information-regularized actor-critic algorithm in which the encoder penalty $\beta_1 I_t^c(S_t; X_t)$ and the policy penalty $\beta_2 I_t^\pi(X_t; A_t)$ are controlled independently.We first establish a single-penalty baseline incorporating a mutual information penalty on the state-action channel $I(S;A)$ into a standard actor-critic objective. Experiments in an open-room gridworld and a structurally demanding zoned corridor show that moderate regularization ($\beta \leq 1.0$) sustains high-reward behavior while extreme regularization ($\beta = 5.0$) degrades performance, with estimated mutual information resisting suppression throughout the moderate regime. We then introduce an explicit categorical encoder mapping raw observations to discrete latent codes, and penalize both the encoder's information rate $I(S;X)$ and the policy's information rate $I(X;A)$ with independent coefficients. A $9 \times 9$ sweep over 81 configurations per environment reveals that the two penalties play qualitatively different roles. In both environments, the unregularized baseline spontaneously produces functional encoder representations without explicit regularization, and performance degrades primarily when $\beta_1 \geq 1.0$ or $\beta_2 \geq 1.0$. The decomposition of $I(S;A)$ into $I(S;X)$ and $I(X;A)$ exposes a cross-penalty coupling invisible to the single-penalty framework. Strong policy compression can indirectly elevate encoder information use, while strong encoder compression can suppress policy information by destroying the state signal the policy exploits. These results support the claim that dual mutual information penalties provide independent control over two qualitatively distinct stages of the perception-action pipeline.

Description

May2026
School of Science

Full Citation

Publisher

Rensselaer Polytechnic Institute, Troy, NY

Terms of Use

Journal

Volume

Issue

PubMed ID

DOI

ISSN

EISSN

Endorsement

Review

Supplemented By

Referenced By