Installation (ROCm)
Last updated: 09/10/2026
For NVIDIA GPU, see the GPU installation guide. For Ascend NPU, see the NPU installation guide.
Requirements
Python: Version >= 3.11
ROCm: Version >= 7.0
GPU: CDNA3 (
gfx942, MI300X / MI325X) or CDNA4 (gfx950, MI350X / MI355X)
Install
On ROCm the Docker image is the supported path; there is no bare-metal pip install.
git clone https://github.com/verl-project/verl-omni.git
cd verl-omni
docker build -f docker/Dockerfile.rocm -t verl-omni:rocm .
Compiling vLLM dominates the build. It defaults to both current CDNA generations; narrowing it to your card roughly halves the time:
docker build -f docker/Dockerfile.rocm -t verl-omni:rocm \
--build-arg PYTORCH_ROCM_ARCH=gfx950 .
Build context is controlled by the repo-root .dockerignore; keep large local folders such as .venv, data/, and checkpoints/ out of the context.
Build arguments
Argument |
Default |
Purpose |
|---|---|---|
|
|
ROCm PyTorch stack, vLLM checkout, hipcc |
|
|
GPU architectures to compile vLLM kernels for |
|
|
vLLM tag built from source |
|
verl commit to install |
Optional Dependencies
Extra |
Install |
When needed |
|---|---|---|
OCR reward |
|
Training with OCR-based reward |
Run
export REPO=/path/to/verl-omni
export WORKSPACE=$HOME
mkdir -p \
"$WORKSPACE/data" \
"$WORKSPACE/checkpoints" \
"$HOME/.cache/huggingface"
docker run -it --rm \
--network=host \
--ipc=host \
--device=/dev/kfd \
--device=/dev/dri \
--group-add=video \
--cap-add=SYS_PTRACE \
--security-opt=seccomp=unconfined \
--privileged \
--ulimit nofile=1048576:1048576 \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
-v "$REPO:/workspace/verl-omni" \
-v "$WORKSPACE/data:$WORKSPACE/data" \
-v "$WORKSPACE/checkpoints:$WORKSPACE/checkpoints" \
-v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
-e WORKSPACE="$WORKSPACE" \
-e HF_HOME=/root/.cache/huggingface \
-e WANDB_API_KEY="${WANDB_API_KEY:-}" \
-w /workspace/verl-omni \
verl-omni:rocm \
bash
Post-Installation Verification
The image verifies torch, vllm, verl, and vllm-omni at build time, but
only the checks that do not touch a device. Importing vllm_omni or verl_omni
pulls in vLLM’s cumem allocator (vllm.device_allocator.cumem), which asserts
that the GPU runtime library is loaded into the process — it is not until torch
has opened a device, so these cannot be checked during docker build. Run the
full set inside the container:
rocm-smi
python -c "import torch; print('torch', torch.__version__, '| HIP', torch.version.hip, '| GPU', torch.cuda.is_available())"
python -c "import vllm; print('vllm', vllm.__version__)"
python -c "from importlib.metadata import version; import vllm_omni; print('vllm-omni', version('vllm-omni'))"
python -c "import verl; print('verl', verl.__version__)"
python -c "import verl_omni; print('VeRL-Omni ready')"
Attention backend
Diffusion recipes default to FlashAttention 3
(attn_backend=_flash_3_varlen_hub), and the kernels-community hub repos
publish no ROCm build variant — the run fails at model init with
FileNotFoundError: Cannot find a build variant for this system. Nothing warns
first: fa_available() only checks that kernels is importable. Every
diffusion recipe on ROCm needs both overrides:
actor_rollout_ref.model.attn_backend=native \
actor_rollout_ref.rollout.rollout_attn_backend=TORCH_SDPA
Both are required together; validate_attention_consistency rejects a
mismatched pair. Omni recipes (main_omni) are unaffected — they select
attention via attn_implementation, which defaults to flash_attention_2 and
works on ROCm using the base image’s flash_attn build.
Build Your Own Docker Image
AMD GPU (ROCm) Dockerfile:
docker/Dockerfile.rocm
NVIDIA GPU images are documented in the GPU installation guide, and Ascend NPU images in the NPU installation guide.
Example: Qwen-Image FlowGRPO training in Docker
This walkthrough uses the OCR dataset and examples/flowgrpo_trainer/qwen_image/run_qwen_image_ocr_lora.sh.
1. Launch the interactive container
Use the docker run command from Run above.
2. Prepare the OCR dataset inside the container
export WORKSPACE=${WORKSPACE:-$HOME}
mkdir -p $WORKSPACE/data/ocr
# Obtain raw train.txt / test.txt from the Flow-GRPO repo:
# https://github.com/yifan123/flow_grpo/tree/main/dataset/ocr
# Place them under $WORKSPACE/data/ocr/, then preprocess:
python3 examples/flowgrpo_trainer/data_process/qwenimage_ocr.py \
--input_dir $WORKSPACE/data/ocr \
--output_dir $WORKSPACE/data/ocr/qwen_image
Qwen/Qwen-Image and the Qwen/Qwen3-VL-8B-Instruct reward model are fetched
from the Hub on first use; pre-download them with hf download to keep the
first step from timing out.
3. Optional: Set W&B credentials
export WANDB_API_KEY=<your_wandb_api_key>
Without a key the run aborts at logger init with
wandb.errors.UsageError: No API key configured. To run without W&B, append
trainer.logger=[console].
4. Run FlowGRPO training
ROCm requires the native/SDPA attention pair described in Attention backend; the script forwards extra Hydra overrides, so append them to the command:
bash examples/flowgrpo_trainer/qwen_image/run_qwen_image_ocr_lora.sh \
actor_rollout_ref.model.attn_backend=native \
actor_rollout_ref.rollout.rollout_attn_backend=TORCH_SDPA
Without those two overrides the run fails during model init with
FileNotFoundError: Cannot find a build variant for this system in kernels-community/flash-attn3.
The script launches python3 -m verl_omni.trainer.main_diffusion with FlowGRPO
vllm_omnirollout and OCR reward (compute_score_ocr). Checkpoints are written to:
checkpoints/flow_grpo/qwen_image_ocr_lora