On-policy learning

  • Learn about behaviour policy from experience sampled from

Off-policy learning

  • Learn about target policy from experience sampled from
  • Learn ‘counterfactually’ about other things you could do: “what if…?”
    • E.g., “What if I would turn left?” new observations, rewards?
    • E.g., “What if I would play more defensively?” different win probability?

Evaluate target policy to compute or
While using behaviour policy to generate actions

Why is this important?

  • Learn from observing humans or other agents (e.g., from logged data)
  • Re-use experience from old policies (e.g., from your own past experience)
  • Learn about multiple policies while following one policy
  • Learn about greedy policy while following exploratory policy

Why one-step TD works off-policy

This is the argument behind off-policy Q learning, the replay buffer in Off-policy actor-critic, and Fitted Q Iteration.

A policy decides which pairs you visit. It does not decide what the environment does once are fixed. Writing out where a replayed transition comes from:

Condition on the you sampled and the first two factors drop out: . So is a genuine sample of the environment’s response to that action, whether a random policy, a human, or last week’s network chose it. The Bellman target

then has the right expectation for . The transition comes from the buffer, and the next action is asked fresh from the current (or the for Q-learning).

Coverage is the catch

The behaviour policy doesn’t bias the target, but it decides what you get to see. If never takes , the data says nothing about , and any value there is pure extrapolation by the network. Offline RL lives and dies by this: can be confidently wrong on unseen actions, and a policy that maximizes goes looking for exactly those errors. Hence behaviour constraints, conservatism, BC regularization, etc.

Why n-step returns break this

At one step, nothing chosen by sits between the action we condition on and the bootstrap. At two steps, one does:

means “take , then follow ”, but the replayed came from following . An -step return has of these unconditioned behaviour actions, which is why it’s only correct on-policy (see N step returns).

Two ways out:

  • Reweight by : Importance sampling corrections.
  • Change the question so those actions are conditioned on too: asks “what if I execute exactly this sequence?”, and the replayed rewards answer it with no bias. That’s Q-chunking.