(tuning_guide)= # Performance Tuning Guide Last updated: 09/14/2026 This page is the starting point for tuning a VeRL-Omni diffusion RL run. It does not repeat the detail already covered by the more specific pages — instead it gives you a decision order and links to the right page for each decision, plus a troubleshooting checklist for the OOM/throughput problems that come up across all of them. ## Where time actually goes A FlowGRPO-style step has three stages, and each has its own tuning surface: | Stage | What it does | Tune with | |---|---|---| | Rollout | Generate images/video/audio for each prompt | {ref}`rollout_batching` | | Reward | Score generated samples | [Async Reward](../algo/async_reward.md) | | Actor | Compute advantages and update the policy | [Tuning and Improving MFU](diffusion_mfu.md#tuning-and-improving-mfu) | Before changing any config, profile the step to see which stage actually dominates wall time — see [Profiling FlowGRPO / diffusion training](profiler.md). Guessing which stage is slow from symptoms alone is unreliable: a rollout that looks slow is often actually reward-bound once you profile it (see [Async Reward](../algo/async_reward.md#motivation)). For a live view of where time goes across a run — instead of a one-off profiler trace — see [Monitor Training with RL-Insight](../start/rl_insight.md), which surfaces trainer, rollout, and TransferQueue metrics in Grafana dashboards. ## 1. Decide your GPU layout first Layout changes are the highest-leverage tuning decision and should come before any per-stage knob, because they change what "optimal" means for the other stages. - **Colocated** (actor, rollout, and reward share the same GPU pool): simplest setup, no idle GPUs, but reward/rollout and actor training time-share the same devices — a slow reward model stalls the whole step. This is the default in most `examples/` scripts. - **Disaggregated reward pool** (`reward.reward_model.enable_resource_pool=True`): puts reward-model inference on its own GPUs so it overlaps with rollout generation instead of blocking it. Worth it once reward scoring is a significant fraction of step time — see [Async Reward](../algo/async_reward.md) for the config and the GPU-count tradeoff. Scaling to more nodes is a separate, orthogonal decision — colocated vs. disaggregated reward still applies once you add nodes. The default multi-node recipe in [Multi-Node Training](../start/multi_node_training.md) keeps reward colocated (`REWARD_TP=4` on every node); that guide also shows how actor, rollout, and reward pools map onto nodes for both layouts. Re-profile after any layout change — the stage that was the bottleneck before a layout change is often not the bottleneck after it. ## 2. Tune rollout throughput Once the GPU layout is fixed, the rollout engine has its own batching tradeoff that is independent of everything else: step-wise continuous batching vs. request-level batching. See {ref}`rollout_batching` for how to choose between them, the config knobs (`step_execution`, `actor_rollout_ref.rollout.max_num_seqs`, `++actor_rollout_ref.rollout.engine_kwargs.vllm_omni.request_batch_max_wait_ms`), and measured before/after numbers for the example recipes. `max_num_seqs` is the top-level engine concurrency knob in both modes; only `request_batch_max_wait_ms` goes under `engine_kwargs.vllm_omni`. ## 3. Tune actor throughput and memory [Tuning and Improving MFU](diffusion_mfu.md#tuning-and-improving-mfu) is the actor-side playbook: `param_offload` / `optimizer_offload`, Ulysses sequence-parallel size, micro-batch size, `layered_summon`, and the gradient-checkpointing MFU caveat. Read that section before changing actor config — it also explains *why* each knob helps, which matters when your OOM point differs from the reference 20B-on-H200 setup it was written against. ## 4. Troubleshooting checklist Symptoms that show up regardless of which stage causes them: **OOM during rollout generation** - Lower `max_num_seqs` first if you are on the request-level batching path — it packs one full activation buffer per concurrent request rather than reusing a step-wise buffer, so it uses more memory per concurrent request for the same value (see the warning in {ref}`rollout_batching`). - Check `actor_rollout_ref.rollout.gpu_memory_utilization` isn't already near 1.0 — this is the vLLM-Omni engine's memory reservation, not an LLM KV cache, and it leaves less headroom for the packed activation memory a high `max_num_seqs` needs. **OOM during actor forward/backward** - Decrease `ppo_micro_batch_size_per_gpu` first — this lowers activation memory without changing the effective batch size (see the FAQ table in [FlowGRPO Quickstart](../start/flowgrpo_quickstart.md#faq-tuning-oom-related-parameters)). **OOM during `update_actor` / `update_weights` after micro-batching is already tight** - This is an optimizer/parameter-state OOM, not an activation OOM — go through the offload ordering in [Tuning and Improving MFU](diffusion_mfu.md#tuning-and-improving-mfu) (`optimizer_offload=True` first, then `param_offload=True` as a last resort), since offloading costs less throughput than a smaller micro-batch once activation memory is no longer the bottleneck. - If both offload flags are already `True`, confirm `layered_summon=True` — disabling it under offload tends to OOM during weight sync. **Step time dominated by data transfer between worker groups (video models)** - This shows up as idle GPU time between rollout and training rather than inside either stage's compute, and is more likely on video/audio generation where trajectories are larger than for image models. It's a V1 trainer (TransferQueue + ReplayBuffer) concern — see [Diffusion V1 training](../start/diffusion_v1.md) for `sync` vs. `separate_async` mode and how the replay buffer moves trajectories into the training loop. **Step time dominated by reward scoring** - Confirm with a profiler trace (recipe 6 in [profiler.md](profiler.md)) before changing anything — reward cost is easy to misattribute to rollout because both run inside the same wall-clock "generation" window when reward is colocated. - Move to a disaggregated reward pool (`reward.reward_model.enable_resource_pool=True`) so scoring overlaps generation instead of gating it; see [Async Reward](../algo/async_reward.md). **MFU looks implausibly low or above 1.0** - See [How FLOPs are computed](diffusion_mfu.md#how-flops-are-computed) first — if `perf/mfu/actor > 1.0`, the two documented causes are a mis-identified device peak on relabeled SKUs (pin the real peak with `VERL_OMNI_DEVICE_FLOPS_TFLOPS`) and a missing DP gather of sequence lengths, not the LoRA vs. full-FT FLOPs caveat — that one only matters when comparing LoRA against full FT, not as a path to MFU above 1.0. ## See also - [Profiling FlowGRPO / diffusion training](profiler.md) — has ready-to-run recipes for the common cases (end-to-end trace, per-stage trace, memory snapshot, rollout/reward server profiling) rather than just the config reference - [Monitor Training with RL-Insight](../start/rl_insight.md) - [Diffusion V1 training](../start/diffusion_v1.md) — TransferQueue-based trainer for video/audio recipes - {ref}`diffusion_mfu` - {ref}`rollout_batching` - [Async Reward](../algo/async_reward.md) - [Multi-Node Training](../start/multi_node_training.md) - [Performance Reference](../algo/performance.md) — measured throughput for colocated vs. async-reward recipes