Notes written with Claude Code (Claude Opus 5.5) from a reading discussion.
Offline-to-online RL on long-horizon, sparse-reward manipulation has two problems. Value propagates one step per backup, and exploration from a single-step policy is jittery and goes nowhere. The paper’s one idea: treat an -step action chunk as a single macro-action and do ordinary RL over chunks. The policy outputs and executes it open-loop. The critic is . Two things then come for free:
- One-step TD in the chunk MDP is an -step backup in the original MDP, without the off-policy bias of naive -step returns.
- A behavior constraint on chunks constrains short coherent behaviors from the data, not individually plausible actions. That gives structured exploration.
Concretely, for the main variant QC: from an offline dataset (later grown with online rollouts), train (a) a flow-matching BC model over action chunks, and (b) a chunked critic with TD. To act, sample chunks from , keep the one with the highest , and run its actions. There is no separate actor. The variant QC-FQL trains a one-step actor instead, so acting is a single forward pass.
Reading it backwards is probably closer to how it came about. Robot policies have predicted chunks since ACT. Given a model that outputs actions, the natural way to do RL with it is to make the chunk the unit of RL instead of un-chunking it. The macro-action view I had before reading the paper is the right mental model. The paper also quietly offers an answer to “why does chunking help robot policies”: the offline data is non-Markovian, and chunks capture that.

The recipe
Everything happens on chunks . A training sample from the buffer is . The critic target is
where comes from the same best-of- procedure used to act. QC’s whole loop, online phase (the offline phase is identical minus the environment):
\begin{algorithm}
\begin{algorithmic}
\REQUIRE flow BC policy $f_\xi(A \mid s)$, chunked critic $Q_\theta(s, A)$, dataset $\mathcal{D}$ (offline data)
\FOR{every environment step $t$}
\IF{$t \bmod h = 0$}
\STATE $A^1, \dots, A^N \sim f_\xi(\cdot \mid s_t)$
\STATE $A^\star \gets \arg\max_i Q_\theta(s_t, A^i)$ \COMMENT{a new chunk only every $h$ steps}
\ENDIF
\STATE execute the next action of $A^\star$, observe $r_t, s_{t+1}$, add to $\mathcal{D}$
\STATE update $f_\xi$ with the flow-matching loss on chunks from $\mathcal{D}$
\STATE update $Q_\theta$ with the chunked TD loss, $A^\star_{t+h}$ from best-of-$N$
\ENDFOR
\end{algorithmic}
\end{algorithm}
Note what is not there: no policy-gradient step, no actor. The flow is pure BC (on a dataset that keeps growing online), and all the RL lives in and in the argmax.
Why the chunked backup is unbiased when n-step isn’t
The paper’s equations put “biased” under the -step reward sum and “unbiased” under the identical-looking chunked one. The difference is only what is conditioned on.
One-step off-policy TD is fine because a transition is a sample of the environment’s response to , no matter who chose . Naive -step breaks that: actions were chosen by the behavior policy, yet claims chose them. (The general argument is in Why one-step TD works off-policy.)
Q-chunking changes the question so that those actions are inputs. asks “what if I execute exactly these actions, then follow ?“. The replayed rewards answer exactly that. At training time the chunk still comes from the replay buffer. It no longer matters who picked it, because every one of its actions is conditioned on.
It's not a clever n-step estimator
It’s plain one-step TD in an MDP whose action space is , with reward and transition . The faster value propagation is a side effect of that MDP having × fewer decision points.
The data isn’t Markov, but the environment is
The paper motivates the chunk-space behavior constraint with “offline data often exhibits non-Markovian structure”. At first that sounds like it undermines RL. It doesn’t, because two different things are being called Markov:
- The environment is a Markov MDP: . All the RL machinery needs only this.
- The behavior policy that produced the data need not be: a human or a scripted controller acts on intent, subtask, or a plan that isn’t in .
Multimodality alone isn’t the issue; a stochastic Markov policy can be 50% left / 50% right. The issue is temporal correlation. A demonstrator who starts going left keeps going left, so the chunk distribution is roughly . A single-step policy constrained to the per-step marginal can sample L, R, L, R, R, L, where every action looks fine on its own and the sequence is nothing the data ever did. A chunk policy matches the joint distribution, so it commits.
Is a chunk policy non-Markov? At primitive resolution, yes: depends on the chunk chosen at , not only on . At chunk resolution, is perfectly Markov, and is a valid induced transition. The Markov decision boundary moves from every step to every steps; nothing is broken.
So the constraint in chunk space says “stay among the short behaviors the data contains”: push in one direction, reach and close the gripper. Exploration then looks like skills instead of random dithering (the paper measures this: QC’s end-effector moves more per 5 steps and covers more states than the single-step BFN early in training). The paper frames this as the simplest hierarchical RL: the low-level skill is an open-loop action sequence, so the usual unstable bi-level optimization collapses into one RL problem.
Best-of-N is the policy, not a teacher
The behavior constraint for QC is implicit. Sampling chunks from and taking the argmax under induces a distribution over executed chunks, and that distribution satisfies
for any scorer. So there’s no actor that could drift outside the bound after training. Best-of- is the policy, used both to act and to produce in the TD target. RL never changes the flow’s weights. It changes , which changes which proposals win. (Why the bound holds, and the family resemblance to AWR and CFGRL: Best-of-N sampling.)
- Proposal vs policy. is the proposal distribution: it suggests plausible robot motions; chooses among them. Only at is the executed policy equal to .
- is the constraint knob. Small means close to BC; large means more aggressive improvement and more exposure to ‘s errors on rare samples.
- The reference moves. keeps training on the growing dataset online, so the constraint is always relative to the current data, not frozen to the offline prior.
- The cost is real: every chunk decision means flow samples (10 ODE steps each) plus critic calls, at inference too.
QC-FQL: put the improvement into the weights
QC-FQL is the answer to “why not just change the model weights?“. It’s FQL applied on (FQL is exactly the case).
There are three networks: the flow BC model (same as QC), the chunked critic , and a one-step actor that maps Gaussian noise to a full chunk in a single forward pass. “One-step” means one network evaluation instead of integrating the flow ODE, not one environment step. The actor loss is
where is what the flow produces by integrating from the same noise . So the "" term is not a difference of network outputs ( outputs a velocity). It is the L2 distance between the two final chunks generated from one . Feeding both the same noise defines one particular coupling of the two distributions, and is the minimum over couplings, so this loss upper-bounds . trades fidelity to the data for improvement.
Distillation and RL happen in the same update, not as “distill, then fine-tune”. That’s the contrast with Shortcut models: both replace a multi-step sampler with a one-step map, but a shortcut model wants student = teacher. Here the student is deliberately pulled off the teacher toward high . is better thought of as an amortized policy-improvement operator than as a distilled flow.
| where improvement happens | cost to act | |
|---|---|---|
| QC | search: best-of- at every decision | flow samples + calls |
| QC-FQL | inside ‘s weights, via + pull | one forward pass |
| CFGRL | inside the guided sampling dynamics | 2× flow passes per step |
FQL
FQL (Park et al., 2025) is the recipe above: a TD3+BC-style actor-critic whose BC term is distillation from a flow policy with shared noise. The paper doesn’t need anything more from it.
What the results lean on
Setup (snapshot)
- Tasks: 5 OGBench domains (scene, puzzle-3x3, cube-double/triple/quadruple; 5 tasks each), sparse reward, 1M–3M-transition play datasets (100M for cube-quadruple). Plus 3 robomimic tasks from multi-human data (300 trajectories each).
- Protocol: 1M offline gradient steps, then 1M online environment steps; same objective and same / in both phases.
- Networks: small MLPs (4×512), state-based, 10 flow steps. Not a pretrained VLA. by default, critic ensemble .
- Chunking is what does the work. BFN (the same best-of- flow method on single actions) and BFN-n / FQL-n (single-action critics with -step returns) are all clearly worse, most visibly on the hardest domains: after online training, QC reaches 64 / 73 on cube-triple / quadruple versus 41 / 36 for the best non-chunked baselines (RLPD, FQL-n).
- The expressive behavior model is required. Running chunked RL with an off-the-shelf Gaussian actor (RLPD-AC, even with a BC loss) mostly fails on the hard tasks. (Why a Gaussian is a poor behavior model yet a flow is hard to do RL on: RL with generative policies.)
- Chunk length is a real hyperparameter. For QC-FQL on cube-triple, is best, learns faster early but ends lower, and never succeeds: long open-loop chunks lose reactivity and are hard to predict. The authors leave choosing (or adaptive chunk boundaries) open.
- Open-loop chunks are a narrow slice of non-Markovian policies. Tasks that need tight feedback inside a chunk are a stated limitation.
- A larger critic ensemble () helps both QC and BFN a lot; a higher update-to-data ratio doesn’t help QC.