首页 > AI前沿 > Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

arXiv机器学习 2026-09-11 07:16 8 阅读 查看原文

World models let agents plan by predicting the consequences of their actions, but changes in the environment can make them inaccurate.

We study the problem of adapting a world model to an unknown test-time environment, drawn from a known environment family, using only a few episodes of interaction.

Existing approaches trade off computational cost against expressivity, i.e., the range of models a method can produce.

For example, in-context learning is computationally cheap but limited in expressivity, and gradient-based adaptation is expressive but computationally expensive.

We present CLAW (Context-conditioned Low-rank Adaptation of World models), which addresses this tradeoff by using a hypernetwork to generate low-rank (LoRA) adapters at test time.

During pretraining, we simulate adaptation to a variety of environments and jointly train the hypernetwork and base world model.

At test time, we freeze the base model and use a forward pass of the hypernetwork to generate adapters from a small batch of test-time transitions.

We evaluate CLAW in locomotion and manipulation environment families that vary in dynamics, embodiment, and reward.

We show that, using only seconds of test-time data, CLAW outperforms gradient-based adaptation and in-context learning during online adaptation.

We also show that CLAW avoids overfitting in data-scarce regimes, that its advantage comes from the expressive adapters rather than context conditioning, and that pretraining the hypernetwork jointly with the base model outperforms training it post hoc.