Notes written with OpenAI Codex (GPT-5) from a reading discussion.


Suppose we have an offline dataset and want a policy better than the behavior in that dataset. AWR would fit a new policy while giving high-advantage actions larger training weights. CFGRL instead turns “this action is good” into a condition of a generative policy. It trains one model to generate both ordinary dataset actions and actions under that condition, then contrasts the two distributions during sampling. The contrast is the improvement direction; a guidance weight controls how strongly it is followed.

The entire method is:

  1. Give every dataset action an outcome condition. With a value function this can be the binary label ; with goal-conditioned BC it is a future state that the trajectory actually reached.

  2. Train one conditional flow policy on these examples. Randomly drop the condition on of examples, so the same network also learns unconditional BC.

  3. At inference, query the network once without the condition and once with the desired condition. At every flow step use

    For the binary version, the desired condition is ; for GCBC, it is the goal we want to reach.

At this is ordinary BC. At it is ordinary conditional BC. At it pushes past the conditional policy toward actions that are especially characteristic of the desired outcome. Crucially, is chosen while sampling, so one trained model provides the whole sweep.

That is enough to implement CFGRL. Everything else follows from explaining why the conditional-unconditional contrast is a policy-improvement direction, and why goal-conditioned BC already contains the needed signal without an explicit value function.

Conditioning secretly multiplies BC by an outcome-likelihood factor

The central observation is just Bayes’ rule applied to how the training data were generated. Let denote the dataset’s behavior policy, and let be an outcome associated with the sampled action. Their joint distribution is

If a conditional generative model perfectly learns the actions associated with outcome , it learns

This says something stronger than “conditioning filters the data.” The conditional policy is a product of:

  • the behavior prior , which keeps actions near the dataset; and
  • the outcome likelihood , which favors actions associated with the requested outcome.

If the outcome is chosen so that

for a nonnegative, nondecreasing function , then conditioning has already constructed the policy

It remains close to actions the behavior could produce, but reallocates probability toward actions with higher advantage. The improvement proof follows directly. At each state,

The reference policy has . Since grows with , the two quantities have nonnegative covariance, so the numerator is nonnegative. The conditioned policy therefore has nonnegative expected reference-policy advantage at every state, which is the standard condition for policy improvement.

This is the RL content of the method: choose a condition whose likelihood rises with advantage, and conditional density estimation becomes a regularized policy-improvement step.

Guidance exposes and amplifies the factor introduced by conditioning

Knowing that the conditional density is a useful product does not yet tell us how to sample from

This is where the generative-model machinery enters. A diffusion or flow sampler moves an action through a learned vector field. Products are convenient because the log-density gradient of a product is a sum:

Bayes’ rule gives the second term without training a separate outcome classifier:

So the difference between the conditional and unconditional model fields isolates the direction contributed by the outcome likelihood. Multiplying that difference by gives

This is ordinary Classifier-free guidance, except the condition means “good action” or “reaches this goal” rather than “matches this text prompt.” Guidance is useful here not merely because it sharpens a condition, but because the condition was constructed to encode advantage. Turning up CFG therefore turns up a policy-improvement factor.

Why no explicit optimality classifier appears

One could train directly and add its gradient to the behavior-policy field. That classifier would have to remain accurate on partially noised, potentially out-of-distribution actions; exploiting such a learned gradient is risky. Conditional and unconditional policy fields provide the same likelihood-ratio direction through Bayes’ rule, using only standard conditional generative-model training.

One network learns both ingredients before guidance is ever applied

The implementation uses Flow Matching. For each dataset action , sample Gaussian noise and time , interpolate

and regress the velocity network toward the straight conditional velocity:

There is no RL loss and no guidance weight in this objective. The outcome condition is replaced with on of examples, teaching the same parameters two distributions:

Sampling starts from Gaussian action noise. At every integration step the model is queried twice, the two velocities are combined with the chosen , and the action is advanced along the guided velocity. Training and sampling can be summarized as:

\begin{algorithm}
\begin{algorithmic}
\REQUIRE offline data $D$, an outcome-labeling rule, guidance strength $w$
\WHILE{training}
    \STATE Sample $(s,a)$ from $D$ and attach its outcome condition $o$.
    \STATE With probability $0.1$, replace $o$ with $\varnothing$.
    \STATE Sample $a_0$ from Gaussian noise and $t$ uniformly from $[0,1]$.
    \STATE Set $a_t \gets (1-t)a_0+ta$.
    \STATE Regress $v_\theta(a_t,t,s,o)$ toward $a-a_0$.
\ENDWHILE
\STATE At inference, initialize $a$ from Gaussian noise.
\WHILE{integrating the flow}
    \STATE $v_u \gets v_\theta(a,t,s,\varnothing)$
    \STATE $v_c \gets v_\theta(a,t,s,o_{\mathrm{desired}})$
    \STATE Advance $a$ using $v_u+w(v_c-v_u)$.
\ENDWHILE
\RETURN $a$
\end{algorithmic}
\end{algorithm}

The guidance identity above is exact for density scores, while the implementation predicts flow velocities. Prior work relates these parameterizations, and the same linear combination works empirically, but the exact score-space identity should be distinguished from this practical velocity-space transfer.

With a critic, “good” is just a label on every dataset action

In the ordinary offline-RL version, first learn and with IQL, then compute

Each transition receives the binary label

The flow policy is trained on every transition with an ordinary, equally weighted regression loss. The labels merely let it distinguish the distribution of nonnegative-advantage actions from the complete behavior distribution. During evaluation, asking for produces the conditional policy, and guidance beyond exaggerates the conditional-unconditional difference.

This explains both the similarity to and difference from advantage-weighted regression. AWR trains with

so high-advantage examples contribute larger gradients and the temperature is baked into the trained policy. CFGRL instead uses advantage to classify examples, keeps their training losses evenly weighted, and postpones the strength of reweighting until sampling. One trained network can therefore sweep without retraining.

The price is that CFGRL has not removed the RL problem: IQL still had to provide a trustworthy action ranking. CFGRL is the policy extraction step after value learning, not a replacement for learning values.

GCBC gets the same factor from trajectory structure instead of a critic

Goal-conditioned BC is the clever case because the dataset itself supplies an outcome label. Given a trajectory

sample a future offset

take , and train

This is hindsight relabeling, not a curriculum. The goal is a state the trajectory actually reached, not an intermediate waypoint invented by the learner. Each example says: “conditional on this rollout eventually reaching , this was the action taken earlier at .” At evaluation, is replaced by the goal we actually want.

Why sample the future offset geometrically? Because

so the probability of selecting as the hindsight future is proportional to its discounted visitation probability:

for the goal-reaching reward .

Now apply the same Bayes argument as before. GCBC examples are produced by first choosing an action under the behavior policy and then sampling a future goal reached after that action:

Therefore the conditional policy learned by perfect supervised training is

GCBC has thus already performed one product-policy update with . This satisfies the required monotonicity because, for fixed and ,

and the second term does not depend on .

GCBC is not estimating a queryable, calibrated -network. It directly learns the normalized product . But training the unconditional branch alongside it gives both

Their field difference isolates the direction contributed by the implicit goal-reaching value factor. Guidance simply turns that direction up:

This is why the goal-conditioned improvement is almost “free”: condition dropout already supplies unconditional BC, and the GCBC model already supplies the product with the implicit factor. The only new operation is combining their two vector fields differently at inference.

The operator improves a fixed baseline once

CFGRL is mostly an offline/off-policy policy-extraction method, though the mathematical operator itself is not tied to offline data. Its required input is an action-dependent outcome signal whose likelihood rises with action quality:

  • explicit advantage labels from a critic;
  • a future goal that the action helped reach; or
  • some other outcome condition with the same monotonic relationship.

Without such a signal, conditional and unconditional generation may still differ, but their difference has no reason to be an RL improvement direction.

The theory is anchored to one fixed reference policy and its or . A genuine iterative algorithm would need to repeat

using a newly valid signal after every update. With fixed offline data, later policies may also favor actions the dataset scarcely covers. The paper avoids policy reevaluation, new data collection, and repeated distribution shift by applying the operator only once.

Large is not a substitute for these iterations. It keeps amplifying the same -derived direction; a second policy-improvement step would recompute the direction under the changed policy.

What is and is not guaranteed as increases

For every fixed , is still nonnegative and nondecreasing, so the basic theorem establishes improvement over the original reference policy under its assumptions.

The paper additionally claims the stronger pairwise ordering for . Its appendix rewrites as a reweighting of by a function of , then invokes a policy-improvement lemma that would require this factor to be monotone in . That missing relationship is not established, and finite-MDP counterexamples exist. The safe interpretation is: controls how aggressively one fixed improvement factor is applied, but true return need not increase monotonically forever.

The experiments show a useful knob, not a universal RL solution

With explicit IQL advantages, CFGRL usually extracts a stronger policy than AWR across the ExORL and OGBench comparisons, while allowing the guidance sweep to happen without retraining. In the goal-conditioned experiments, improves over flow GCBC on most state and visual tasks, including pointmaze-giant ( success) and visual-cube-single (). Applying guidance at both levels of a hierarchical GCBC policy also produces large gains on several long-horizon tasks.

Some tasks are unchanged, remain at zero, or regress slightly. The useful range of ends when guidance pushes actions outside the region where the learned fields and optimality signal are reliable. Explicit CFGRL inherits critic errors; goal-conditioned CFGRL inherits the reachability and coverage of the behavior trajectories. The method also uses two network evaluations per flow step to obtain the conditional and unconditional fields.

The durable picture is therefore concrete: label actions by outcomes, train conditional and unconditional behavior in one generative policy, and use CFG to amplify the likelihood ratio between them. When that outcome likelihood increases with advantage, the familiar CFG direction becomes a policy-improvement direction. GCBC is the surprising special case where hindsight relabeling has already encoded the necessary value factor into the conditional action distribution.