Notes written with Claude Code (Claude Opus 5.5) from a reading discussion.
Draw samples from a reference distribution , score them with some function , keep the winner:
The point worth remembering: the distribution of is itself a distribution. Call it . In RL with a behavior policy and , is a policy, and an improved one. In LLMs with the SFT model and a reward model, it’s the baseline that RLHF is often compared against.
Why it can’t drift far from
Every candidate came from , so selection can only reweight ; it can’t put mass where has none. Ignoring ties, wins when the other draws score lower, so
and the density ratio is capped: . The tighter known bound is
Notice what it doesn’t depend on: . You can swap the scorer every iteration (a -function being trained, say) and the bound still holds, as long as the procedure stays “sample from , pick one”. is the knob: is just ; larger is more aggressive optimization against , and more exposure to ‘s errors.
The bound is distributional. It doesn’t say any single is close to a typical sample of , only that the distribution of winners isn’t far from .
As a policy-improvement operator
In offline RL this gives a behavior constraint for free: is a BC policy, is , and is “improve on the data, but stay within its support”. It also solves the continuous problem from Continuous Actions by search instead of by an actor network.
Same family as other “reference distribution + value preference” operators:
- AWR: reweight dataset actions by and fit a new policy to them.
- CFGRL: tilt the generative sampling dynamics with guidance.
- Best-of-N: reweight by rank, and don’t fit anything. The selection runs every time you act, so you pay samples plus scores per decision. Training an actor to reproduce the improvement is the fix; see QC vs QC-FQL in Q-chunking, and the other workarounds in The bridges.