On-policy learning
- Learn about behaviour policy from experience sampled from
Off-policy learning
- Learn about target policy from experience sampled from
- Learn ‘counterfactually’ about other things you could do: “what if…?”
- E.g., “What if I would turn left?” new observations, rewards?
- E.g., “What if I would play more defensively?” different win probability?
Evaluate target policy to compute or
While using behaviour policy to generate actions
Why is this important?
- Learn from observing humans or other agents (e.g., from logged data)
- Re-use experience from old policies (e.g., from your own past experience)
- Learn about multiple policies while following one policy
- Learn about greedy policy while following exploratory policy
Why one-step TD works off-policy
This is the argument behind off-policy Q learning, the replay buffer in Off-policy actor-critic, and Fitted Q Iteration.
A policy decides which pairs you visit. It does not decide what the environment does once are fixed. Writing out where a replayed transition comes from:
Condition on the you sampled and the first two factors drop out: . So is a genuine sample of the environment’s response to that action, whether a random policy, a human, or last week’s network chose it. The Bellman target
then has the right expectation for . The transition comes from the buffer, and the next action is asked fresh from the current (or the for Q-learning).
Coverage is the catch
The behaviour policy doesn’t bias the target, but it decides what you get to see. If never takes , the data says nothing about , and any value there is pure extrapolation by the network. Offline RL lives and dies by this: can be confidently wrong on unseen actions, and a policy that maximizes goes looking for exactly those errors. Hence behaviour constraints, conservatism, BC regularization, etc.
Why n-step returns break this
At one step, nothing chosen by sits between the action we condition on and the bootstrap. At two steps, one does:
means “take , then follow ”, but the replayed came from following . An -step return has of these unconditioned behaviour actions, which is why it’s only correct on-policy (see N step returns).
Two ways out:
- Reweight by : Importance sampling corrections.
- Change the question so those actions are conditioned on too: asks “what if I execute exactly this sequence?”, and the replayed rewards answer it with no bias. That’s Q-chunking.