Reinforcement learning (RL) has emerged as a unifying framework for sequential decision making. Yet, its practical impact is curbed by three persistent limitations: prohibitive sample demands, poor generalization across tasks and domains, and a lack of interpretability that accompanies high-capacity function approximators. This research argues that these limitations share a common source-insufficient prior structure—and that they can be alleviated by modelling agents with learnable, discrete, and adaptive priors. We pursue this claim through a trilogy of works that together reshape the reward, the world model, and the skill library on which efficient learning depends.
The first work converts sparse, human-supplied scores for heavy-equipment operation into a dense, differentiable reward prior. By training a score predictor that both evaluates and guides the policy, we demonstrate safe exploration in environment interaction relative to hand-crafted heuristics. The second work introduces DART, a transformer-based architecture that tokenizes states, actions, and returns into a standard symbolic alphabet. This discrete world model captures long-range temporal dependencies while preserving the manipulability of tokens, delivering state-of-the-art sample efficiency on the Atari-100k benchmark without sacrificing final performance. The third work presents STRIDE, a dynamic vector-quantized VAE that autonomously expands a codebook of motor primitives as tasks accumulate. STRIDE retains previously learned behaviours with negligible degradation. It halves the adaptation time on challenging locomotion curricula, all while exposing an interpretable “skill trace” that allows practitioners to audit and debug decisions.
Collectively, these contributions substantiate a unified hypothesis: when knowledge about rewards, dynamics, and skills is cast into discrete, growing vocabularies, tabula-rasa search gives way to data-efficient, compositional, and transparent learning. Beyond empirical gains—up to a 2.6× speed-up across diverse domains—the work offers conceptual tools for reasoning about priors in RL, provides open-source implementations for community use, and outlines future directions toward multi-modal and human-editable prior structures. In doing so, it takes a decisive step toward RL systems that can be trusted to learn quickly, allowing transferability and explainability when it matters most.
Agarwal, P. (Author),
Andrews (Supervisor) &
Ebrahimi-Kahou (Co-supervisor),
24 Sept 2025Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering