Performance Tuning Guide
Last updated: 09/14/2026
This page is the starting point for tuning a VeRL-Omni diffusion RL run. It does not repeat the detail already covered by the more specific pages — instead it gives you a decision order and links to the right page for each decision, plus a troubleshooting checklist for the OOM/throughput problems that come up across all of them.
Where time actually goes
A FlowGRPO-style step has three stages, and each has its own tuning surface:
Stage |
What it does |
Tune with |
|---|---|---|
Rollout |
Generate images/video/audio for each prompt |
|
Reward |
Score generated samples |
|
Actor |
Compute advantages and update the policy |
Before changing any config, profile the step to see which stage actually dominates wall time — see Profiling FlowGRPO / diffusion training. Guessing which stage is slow from symptoms alone is unreliable: a rollout that looks slow is often actually reward-bound once you profile it (see Async Reward).
For a live view of where time goes across a run — instead of a one-off profiler trace — see Monitor Training with RL-Insight, which surfaces trainer, rollout, and TransferQueue metrics in Grafana dashboards.
1. Decide your GPU layout first
Layout changes are the highest-leverage tuning decision and should come before any per-stage knob, because they change what “optimal” means for the other stages.
Colocated (actor, rollout, and reward share the same GPU pool): simplest setup, no idle GPUs, but reward/rollout and actor training time-share the same devices — a slow reward model stalls the whole step. This is the default in most
examples/scripts.Disaggregated reward pool (
reward.reward_model.enable_resource_pool=True): puts reward-model inference on its own GPUs so it overlaps with rollout generation instead of blocking it. Worth it once reward scoring is a significant fraction of step time — see Async Reward for the config and the GPU-count tradeoff.
Scaling to more nodes is a separate, orthogonal decision — colocated vs.
disaggregated reward still applies once you add nodes. The default multi-node
recipe in Multi-Node Training keeps reward
colocated (REWARD_TP=4 on every node); that guide also shows how actor,
rollout, and reward pools map onto nodes for both layouts.
Re-profile after any layout change — the stage that was the bottleneck before a layout change is often not the bottleneck after it.
2. Tune rollout throughput
Once the GPU layout is fixed, the rollout engine has its own batching
tradeoff that is independent of everything else: step-wise continuous
batching vs. request-level batching. See Diffusion Rollout Batching for how to
choose between them, the config knobs (step_execution,
actor_rollout_ref.rollout.max_num_seqs,
++actor_rollout_ref.rollout.engine_kwargs.vllm_omni.request_batch_max_wait_ms),
and measured before/after numbers for the example recipes. max_num_seqs is
the top-level engine concurrency knob in both modes; only
request_batch_max_wait_ms goes under engine_kwargs.vllm_omni.
3. Tune actor throughput and memory
Tuning and Improving MFU is the
actor-side playbook: param_offload / optimizer_offload, Ulysses
sequence-parallel size, micro-batch size, layered_summon, and the
gradient-checkpointing MFU caveat. Read that section before changing actor
config — it also explains why each knob helps, which matters when your OOM
point differs from the reference 20B-on-H200 setup it was written against.
4. Troubleshooting checklist
Symptoms that show up regardless of which stage causes them:
OOM during rollout generation
Lower
max_num_seqsfirst if you are on the request-level batching path — it packs one full activation buffer per concurrent request rather than reusing a step-wise buffer, so it uses more memory per concurrent request for the same value (see the warning in Diffusion Rollout Batching).Check
actor_rollout_ref.rollout.gpu_memory_utilizationisn’t already near 1.0 — this is the vLLM-Omni engine’s memory reservation, not an LLM KV cache, and it leaves less headroom for the packed activation memory a highmax_num_seqsneeds.
OOM during actor forward/backward
Decrease
ppo_micro_batch_size_per_gpufirst — this lowers activation memory without changing the effective batch size (see the FAQ table in FlowGRPO Quickstart).
OOM during update_actor / update_weights after micro-batching is already tight
This is an optimizer/parameter-state OOM, not an activation OOM — go through the offload ordering in Tuning and Improving MFU (
optimizer_offload=Truefirst, thenparam_offload=Trueas a last resort), since offloading costs less throughput than a smaller micro-batch once activation memory is no longer the bottleneck.If both offload flags are already
True, confirmlayered_summon=True— disabling it under offload tends to OOM during weight sync.
Step time dominated by data transfer between worker groups (video models)
This shows up as idle GPU time between rollout and training rather than inside either stage’s compute, and is more likely on video/audio generation where trajectories are larger than for image models. It’s a V1 trainer (TransferQueue + ReplayBuffer) concern — see Diffusion V1 training for
syncvs.separate_asyncmode and how the replay buffer moves trajectories into the training loop.
Step time dominated by reward scoring
Confirm with a profiler trace (recipe 6 in profiler.md) before changing anything — reward cost is easy to misattribute to rollout because both run inside the same wall-clock “generation” window when reward is colocated.
Move to a disaggregated reward pool (
reward.reward_model.enable_resource_pool=True) so scoring overlaps generation instead of gating it; see Async Reward.
MFU looks implausibly low or above 1.0
See How FLOPs are computed first — if
perf/mfu/actor > 1.0, the two documented causes are a mis-identified device peak on relabeled SKUs (pin the real peak withVERL_OMNI_DEVICE_FLOPS_TFLOPS) and a missing DP gather of sequence lengths, not the LoRA vs. full-FT FLOPs caveat — that one only matters when comparing LoRA against full FT, not as a path to MFU above 1.0.
See also
Profiling FlowGRPO / diffusion training — has ready-to-run recipes for the common cases (end-to-end trace, per-stage trace, memory snapshot, rollout/reward server profiling) rather than just the config reference
Diffusion V1 training — TransferQueue-based trainer for video/audio recipes
Performance Reference — measured throughput for colocated vs. async-reward recipes