Seminar talks at Penn State, Stanford, and Berkeley: "What Makes an RL Sample Transferable?"

In early October, I will give seminar talks on transfer reinforcement learning at three departments:

📅 October 1 — Department of Statistics, Penn State University
📅 October 5 — Department of Statistics, Stanford University
📅 October 6 — Department of Statistics, UC Berkeley

What Makes an RL Sample Transferable? A Statistical Theory of Transfer $Q$-Learning

In business, healthcare, and education, sequential decisions must often be learned for a target population with little data, while abundant data exist from related populations. Transfer learning is the natural remedy, and in regression the recipe is well understood: pool source and target samples, then correct for a low-complexity difference. This talk asks whether the same recipe works for reinforcement learning (RL), and answers: only sometimes. In a Markov decision process the regression response at each stage is not the observed reward but a Bellman target that embeds the source’s future value under the source’s transition. Pooling therefore introduces a bias that is absent from supervised transfer and that persists even when reward functions are identical across populations; it vanishes only at the final stage of a finite horizon. This observation leads to a precise notion of a transferable RL sample and to a re-weighted targeting procedure that restores Bellman alignment: source responses are re-targeted to the target’s value function and re-weighted by a transition density ratio, cascading backward through the horizon.

We establish convergence rates for the resulting transfer Q-learning estimator with deep neural network function approximation, expressed through two “knobs” of task discrepancy — the reward gap and the transition density ratio — and show exactly when source data accelerate learning over single-task rates. The same principle extends to stationary MDPs with iterative Q-learning, to online RL via one-step Bellman alignment, and to composite MDPs with transferable transition components. Experiments on synthetic environments and an ICU sepsis cohort from MIMIC-III confirm the theory.

Related papers: