Optional Engine Backends
Last updated: 09/24/2026
VeRL-Omni defaults to FSDP2 as the training engine for the policy and reference models. The diffusion trainer and Qwen3-Omni Thinker can alternatively use VeOmni. The engine is selected at the Hydra command line — see examples/flowgrpo_trainer/qwen_image/run_qwen_image_ocr_veomni.sh for a complete recipe.
Installing VeOmni alongside vLLM 0.28.0
VeOmni 0.1.12’s gpu extra pins torch==2.11.0+cu130, which conflicts with the torch==2.13.0 pulled in by vllm==0.28.0. A plain uv pip install veomni[gpu]==0.1.12 therefore fails dependency resolution.
Install the base VeOmni package without dependency resolution so the existing
torch/vLLM stack is preserved, then add its media runtime dependencies
(CI runs the image-only VeOmni smoke and pulls librosa/soundfile/av/audioread
from the [audio]/[omni] extras; torchcodec is only needed when the VeOmni
engine decodes video). This is the shared CI install; Qwen3-Omni Thinker also
needs the kernels listed below.
uv pip install veomni==0.1.12 --no-deps
uv pip install torchcodec librosa soundfile av audioread
Verify the engine is importable:
python -c "import veomni; print('veomni', veomni.__version__)"
python -c "from veomni.distributed.offloading import load_model_to_gpu, load_optimizer, offload_model_to_cpu, offload_optimizer; print('VeOmni offloading helpers OK')"
The two-GPU Thinker backend check and two-step V1 smoke pass with torch 2.13
and vLLM 0.28, including forward, backward, optimizer updates, EP weight export
and full-weight rollout synchronization. These checks use tiny random weights;
they do not validate the full 30B checkpoint or convergence. The base
import/offloading checks above do not exercise these training paths.
The complete veomni[gpu] extra belongs in a separate environment compatible
with its torch pin; on the vLLM stack, install the required kernels individually.
Qwen3-Omni Thinker kernel dependencies
The VeOmni GSPO launcher uses the following native VeOmni defaults on every training node:
Selector |
Required package |
|---|---|
|
Local |
|
|
|
|
kernels Hub attention and vLLM’s internal attention package do not supply the
local flash_attn import used by VeOmni. The shared CI install and Docker
veomni target do not install local FA2; add it before using this recipe.
FA3, FA4 and Quack are not required by these defaults.
Start from the installed torch/vLLM environment and constrain its existing packages, including torch, Triton and CUDA runtime packages, while adding the missing dependencies:
uv pip freeze --exclude-editable > /tmp/verl-omni-thinker-constraints.txt
uv pip install -c /tmp/verl-omni-thinker-constraints.txt \
liger-kernel ninja packaging psutil setuptools wheel
For FA2, use a wheel matching the installed torch, CUDA, Python and platform,
or build from source against that environment. A source build requires a CUDA
development toolkit (nvcc) compatible with the installed torch build:
FLASH_ATTENTION_FORCE_BUILD=TRUE MAX_JOBS=4 \
uv pip install -c /tmp/verl-omni-thinker-constraints.txt \
--no-build-isolation --no-binary flash-attn --reinstall-package flash-attn flash-attn
These constraints prevent the added packages from replacing the installed stack; an incompatibility should fail resolution rather than change torch. The backend check passed on two GB200 GPUs with torch 2.13.0+cu130, locally built flash-attn 2.8.3.post1, liger-kernel 0.8.3, Triton 3.7.1 and Transformers 5.14.1 through normal package initialization. This uses a tiny random checkpoint, not the full 30B model. Do not install a torch 2.11 FA2 wheel into a torch 2.13 environment.
Run these checks on each training node. They load the FA2 extension and resolve
VeOmni’s native operator implementations, beyond merely importing veomni:
python - <<'PY'
import torch
import triton
from flash_attn import flash_attn_func, flash_attn_varlen_func
from veomni.arguments import OpsImplementationConfig
from veomni.ops import apply_ops_config
apply_ops_config(OpsImplementationConfig(moe_implementation="fused_triton"))
print("Thinker kernels imported:", torch.__version__, torch.version.cuda, triton.__version__)
PY
python -c "import vllm_omni; import verl_omni; print('Full package imports OK')"
Imports do not exercise GPU execution. Before treating the stack as validated, run both the backend check and the complete V1 smoke through normal package initialization, as described in the omni integration guide. The smoke allows 1800 seconds for rollout startup because a cold FlashInfer kernel build on GB200 can exceed vLLM-Omni’s default 600-second timeout.