All research
Adaptive learning

Continual Learning

An inference-oriented approach to cumulative learning using sequential Bayesian inference, predictive coding, uncertainty, and causal modularity.

Category
Research
Focus
Bayesian inference / predictive coding
Continual Learning project cover

Our project approaches continual learning as the problem of maintaining and recursively updating an evolving probabilistic knowledge state. Rather than treating learning primarily as a sequence of independent fine-tuning stages followed by mechanisms designed to suppress catastrophic forgetting, we are exploring sequential Bayesian inference as a guiding perspective: what has already been learned can be viewed as informing the prior through which new evidence is interpreted.

Recursive belief updating

Under the standard Bayesian formulation, incorporating the same body of evidence leads to a posterior determined by the accumulated evidence rather than its arbitrary presentation order. Order invariance holds at the level of ideal Bayesian inference but can be lost under practical approximation. Real systems operate with finite capacity, restricted posterior families, approximate representations, and iterative optimization. They therefore have to repeatedly compress and re-express their current state of knowledge, and those approximations can introduce distortions that accumulate over time.

This makes task-order dependence an interesting diagnostic for continual learning. If learning A, B, then C produces a substantially different final model from C, A, then B, even when the underlying body of information is equivalent, some of that difference may reflect the approximate inference process rather than the information itself. From this perspective, catastrophic forgetting and sequence sensitivity can be studied partly as consequences of imperfect recursive belief updating.

We do not assume that all order effects should disappear. Continual learning concerns genuinely non-stationary environments, where temporal structure may itself be informative and later observations may provide legitimate reasons to revise earlier beliefs. The distinction we are interested in is between meaningful adaptation to a changing environment and spurious dependence on task ordering introduced by approximation error.

Predictive coding and uncertainty

Predictive coding is our main computational substrate for investigating these ideas. Predictive-coding systems perform iterative inference over hierarchical latent representations using local prediction errors, bringing inference, representation formation, and learning into a closely coupled process.

Our probabilistic work asks how uncertainty might participate in those learning dynamics. A learner should ideally represent not only what it currently believes, but also how strongly those beliefs are supported. Well-supported structure may appropriately become less susceptible to incidental changes, while uncertain or weakly constrained structure may remain more available for adaptation. Under this interpretation, the familiar stability-plasticity trade-off can be investigated as an inference question.

Causal modularity

A complementary strand of the project investigates causal modularity. Whereas uncertainty may help characterize how strongly existing knowledge should constrain an update, modular structure provides a way of reasoning about where learning should take place. If different tasks rely on partially distinct mechanisms, restricting updates to the relevant parts of a system may reduce unnecessary interference elsewhere.

Our work has also contributed to Hierarchical Bayesian Causal Modular Learning (HiBaCaML), a continual-learning framework published at AGI-26, with members of our team as co-authors. HiBaCaML explores a complementary architectural approach built around a collection of structured modules and a sparse probabilistic controller that selects which modules participate in a given context.

Longer-term direction

We are interested in approaches that move beyond purely parameter-centric notions of memory toward function-space continual learning. Neural networks can realize similar predictive behavior through many different parameter configurations, so preserving particular weights is not necessarily equivalent to preserving the knowledge represented by the model.

More broadly, our aim is not to assume in advance that continual learning can be reduced to one mechanism. We are investigating whether retention, transfer, uncertainty, structural adaptation, and principled forgetting can be understood within a more coherent inference-oriented framework, and where that framing succeeds or breaks down empirically.