0

Factorization Regret mediates compositional generalization in latent space

Are there still barriers to generalization once all of the relevant variables are known? We consider the challenge of generalizing to novel combinations of task-relevant latent variables.

Preview
Year
2026
Hosting
Full text hostedCC0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2603.27134CC0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Are there still barriers to generalization once all of the relevant variables are known? We consider the challenge of generalizing to novel combinations of task-relevant latent variables. To explore this framework, we develop the Cognitive Gridworld, a stationary Partially Observable Markov Decision Process (POMDP) in which observations are generated jointly by multiple latent variables with parametric interactions. This setting allows us to describe Factorization Regret: an information-theoretic quantity that measures the contribution of latent variable interactions to task performance. Using this metric, we first analyze Recurrent Neural Networks (RNNs) that are explicitly provided with the interactions and find that Factorization Regret explains the accuracy gap between Echo State and Fully Trained networks. Additionally, our analysis uncovers a theoretically predicted failure mode, where confidence becomes decoupled from accuracy. These results suggest that utilizing the interactions between relevant variables is a non-trivial capability. We then address a harder regime where the interactions themselves must be learned by an embedding model. Learning how variables interact while simultaneously learning how to infer their values is a variational inference problem, often described as meta-learning. To explicitly disentangle variable inference and parameter estimation, we develop Representation Classification Chains (RCCs), a novel architecture which learns how latent variables interact through Reinforcement Learning (RL), from teaching signals provided only for a subset of training goal variables. Finally, we demonstrate the usefulness of RCCs in enabling generalization to novel combinations of latent variables through offline learning. In summary, we present a theoretically grounded setting for research, development and evaluation of goal-directed general intelligence.