Notes written with Claude Code (Claude Opus 5.5) from a reading discussion.


Papers keep saying “flow / diffusion policies are hard for RL” and “Gaussian policies are bad”, which makes it sound like nothing works. Both statements are true, because they answer two different questions:

  1. Can the policy class represent the behavior data? Correlated action chunks, several distinct modes.
  2. Does it give RL something to work with? A log-probability, or an action that is a cheap differentiable function of the weights.
represents the dataRL handles
Gaussian (diagonal)poorly: unimodal, independent dimsboth, cheap
ACT-style (CVAE, L1, latent = prior mean at test time)partly: correlated chunks, but one modepathwise only (deterministic); add Gaussian noise for a likelihood, and you’re back to row 1
Flow / diffusionwellneither cheaply

So switching from an ACT-style to a flow policy fixes question 1 and breaks question 2. Every method below is a way to keep the flow’s answer to 1 while getting around 2.

Question 1: representing the data

Two separate failures, often blurred together:

  • Correlation within a chunk. A 5×7 action chunk’s entries are strongly coupled: step 3 should look like step 2. Independent per-dimension Gaussian noise wiggles each one separately, which is jitter, not exploration. A full covariance (Cholesky factor, ~630 numbers for 35 dims) or temporally correlated noise (DDPG’s Ornstein–Uhlenbeck noise) fixes this.
  • Multimodality. If the data is “LLLLL or RRRRR”, any Gaussian, with any covariance, is still one blob. The best fit sits in the middle, which may be the one thing nobody did. A covariance can’t fix this; an expressive generator can.

This matters for RL beyond imitation quality. The behavior constraint and the exploration both inherit from this model. A model that blurs the modes constrains the policy toward behavior that isn’t in the data and explores incoherently. That’s why the Gaussian chunk actor fails in Q-chunking.

Question 2: what RL needs from a policy

There are two standard ways to improve a policy, each needing one handle:

  • Likelihood: . Used by Policy Gradient, PPO ratios, SAC’s entropy term, and any explicit KL to a reference policy.
  • Pathwise gradient: cheap and differentiable, so . This is how DDPG, TD3, and SAC’s reparameterized actor learn.

A flow model has neither cheaply:

  • Likelihood: the sampler defines a real distribution, but requires integrating the divergence of the velocity field along the ODE path. That is expensive and noisy.
  • Pathwise: the action is the end of a 10-step ODE, so means backpropagating through every step, which is costly and unstable like backprop through time.

Note that “it’s trained by regression” is not the problem. ACT trains by regression too. What matters is what the trained model exposes.

A third, separate issue: the constraint needs a handle too

Independently of the model class, offline-to-online RL has to keep the policy near the data, because is extrapolating everywhere else (coverage). The usual constraint is a KL to the behavior policy, which needs log-probs: question 2 again. So with flows the constraint has to be rebuilt as well, not just the improvement step.

The bridges

Each method keeps the flow and replaces whichever handle it can’t get:

MethodReplacesHow
Best-of-N sampling (QC in Q-chunking)bothneeds only samples; the constraint comes free as a KL bound
One-step distilled actor (FQL, QC-FQL)pathwisetrain a single-pass actor near the flow; the constraint becomes via shared noise
Advantage-weighted flow matching (AWR-style)likelihoodweight the regression loss by instead of the log-likelihood
Guidance / conditioning (CFGRL, RECAP in Pi 0.6)bothimprove at sampling time by contrasting conditioned and unconditioned velocity
RL in the noise space (DSRL)bothkeep the flow fixed as a decoder; run ordinary Gaussian RL on its input noise , so every perturbation decodes to a coherent chunk

The last row also answers the covariance problem from question 1: noise in is already shaped by a model that knows the correlations and the modes.