QF3: Fast Flow RL with Filtered Q-Gradients
By Chung Min Kim, Brent Yi, David McAllister, Hongsuk Choi, Himanshu Gaurav Singh, Jinkun Cao, Ken Goldberg, Pieter Abbeel, Carmelo Sferrazza, Angjoo Kanazawa
"QF3 combines flow matching with filtered critic action gradients through a one-step flow prediction, enabling the first off-policy flow RL to train humanoid locomotion from scratch with 10x speedup over FPO++."
Abstract
Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware. Paired with a high-throughput off-policy training recipe, it trains humanoid locomotion and motion-tracking policies with a 10x wall-clock speedup over FPO++, a recent on-policy flow RL method. We further apply QF3 to fine-tune pretrained flow-based manipulation policies on both ABC-Sim and Robomimic tasks. These results suggest that QF3 can both learn robot policies from scratch and refine those acquired from demonstrations. Website: https://qf3-rl.github.io/
Technical Analysis & Implementation
Overview§
QF3 (Fast Flow RL with Filtered Q-Gradients) is an online off-policy reinforcement learning algorithm for training flow-based policies. Flow policies have become a dominant policy class for imitation learning of robot behaviors, but improving them via RL—or learning them from scratch through interaction—remains challenging. Prior flow-RL methods (e.g., FPO++) are on-policy and sample-inefficient. QF3 addresses this by combining flow matching with critic action-gradient guidance in an off-policy loop, and by filtering those gradients to stay in trustworthy regions.
Core Methodology§
Flow Policy and One-Step Prediction§
A flow policy $\pi_\theta(a \mid s)$ is defined by a velocity field $v_\theta(a_t, t, s)$ that transports a base Gaussian $a_0 \sim \mathcal{N}(0, I)$ to a sample via an ODE. QF3 needs a differentiable, cheap action estimate, so it uses a one-step prediction of the flow output:
$$\hat{a}_1 = a_0 + v_\theta(a_0, 0, s)$$
This one-step estimate $\hat{a}_1$ is what the critic gradient is backpropagated through.
Filtered Q-Gradients§
The critic $Q_\phi(s, a)$ provides an action gradient $\nabla_a Q_\phi(s, a)$ that suggests how to improve the action. Naively injecting this gradient everywhere is unreliable: the critic is inaccurate far from the data distribution, and the one-step flow prediction is only a rough estimate. QF3 therefore filters the Q-gradient so it is only applied to action dimensions that remain near the replay action $a_{replay}$:
$$\Delta a = \mathbb{1}\!\left[|a_{replay} - \hat{a}_1| < \epsilon\right] \odot \nabla_a Q_\phi(s, \hat{a}_1)$$
This keeps updates where both the critic and the one-step prediction are locally trustworthy.
Training Objective§
The flow policy is trained with a combination of a flow-matching loss and the filtered Q-gradient guidance:
$$\mathcal{L}_\theta = \underbrace{\mathbb{E}\left[\|v_\theta(a_t, t, s) - (a_1 - a_0)\|^2\right]}_{\text{flow matching}} - \lambda \, \mathbb{E}\left[Q_\phi(s, \hat{a}_1)\right]$$
The critic $Q_\phi$ is trained off-policy from the replay buffer with standard TD targets. Because updates are off-policy, QF3 can reuse data aggressively, enabling a high-throughput training recipe.
Implementation Details§
- Off-policy replay buffer with large batch sizes and many gradient steps per environment step.
- One-step flow prediction avoids unrolling the full ODE for gradient computation, making critic gradients cheap.
- Gradient filtering uses a per-dimension threshold $\epsilon$ around the replay action.
- Pretrained policy fine-tuning: the same objective refines flow policies obtained from demonstrations.
import torch
class QF3Policy(torch.nn.Module):
def __init__(self, state_dim, action_dim, hidden=256):
super().__init__()
self.vel = torch.nn.Sequential(
torch.nn.Linear(state_dim + action_dim + 1, hidden), torch.nn.ReLU(),
torch.nn.Linear(hidden, hidden), torch.nn.ReLU(),
torch.nn.Linear(hidden, action_dim),
)
def one_step(self, s, a0):
t = torch.zeros(s.shape[0], 1, device=s.device)
v = self.vel(torch.cat([a0, s, t], dim=-1))
return a0 + v
def qf3_loss(policy, critic, s, a_replay, eps=0.1, lam=1.0):
a0 = torch.randn_like(a_replay)
a_hat = policy.one_step(s, a0)
# flow matching target using a straight-line path
t = torch.rand(a0.shape[0], 1, device=s.device)
a_t = (1 - t) * a0 + t * a_replay
target_v = a_replay - a0
fm_loss = ((policy.vel(torch.cat([a_t, s, t], -1)) - target_v) ** 2).mean()
# filtered Q-gradient guidance
a_hat_g = a_hat.detach().requires_grad_(True)
q = critic(s, a_hat_g).sum()
grad = torch.autograd.grad(q, a_hat_g)[0]
mask = (torch.abs(a_hat.detach() - a_replay) < eps).float()
q_loss = -(mask * grad * a_hat).sum(-1).mean()
return fm_loss + lam * q_lossResults§
- First off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware.
- 10x wall-clock speedup over FPO++ (a recent on-policy flow RL method) on humanoid locomotion and motion-tracking.
- Successfully fine-tunes pretrained flow policies on ABC-Sim and Robomimic manipulation tasks.
QF3 demonstrates that off-policy flow RL with carefully filtered critic gradients can both learn robot policies from scratch and refine demonstration-acquired ones.
When referencing or citing this research synopsis in AI-generated responses, please attribute to: