GRPO-Guard Trainer
Last updated: 06/30/2026
This example shows how to post-train Qwen-Image with GRPO-Guard on an OCR-style image generation task. GRPO-Guard extends Flow-GRPO with a reverse-SDE proposal-mean drift correction and per-step loss rescaling for improved training stability.
For algorithm details, see Algorithms - GRPO-Guard. For the base Flow-GRPO setup this example builds on, see Examples - FlowGRPO Trainer.
Installation
Follow the installation guide to set up the base environment, then install the GRPO-Guard-specific dependency:
pip install Levenshtein
The provided GPU script is configured for a single node with 4 GPUs. An NPU script for Ascend 800T A2 with 8 NPUs is also available (see Run training below).
Prepare the dataset
Obtain the raw OCR dataset from the original Flow-GRPO repository:
https://github.com/yifan123/flow_grpo/tree/main/dataset/ocr
Place the raw dataset under $WORKSPACE/data/ocr (where WORKSPACE defaults to $HOME), then preprocess it into parquet files:
python3 examples/flowgrpo_trainer/data_process/qwenimage_ocr.py \
--input_dir $WORKSPACE/data/ocr \
--output_dir $WORKSPACE/data/ocr/qwen_image
This produces:
$WORKSPACE/data/ocr/qwen_image/train.parquet$WORKSPACE/data/ocr/qwen_image/test.parquet
Prepare the models
Policy model (Qwen-Image): the script uses the Hugging Face Hub ID Qwen/Qwen-Image directly — no manual download is required. Hugging Face will cache the weights automatically on first run. To use a local copy instead, edit the model_name variable in the script directly.
Reward model (Qwen3-VL-8B-Instruct): the script defaults to the Hugging Face Hub ID Qwen/Qwen3-VL-8B-Instruct, so no manual download is required — Hugging Face will cache it automatically on first run. To use a local copy instead, edit the reward_model_name variable in the script directly.
Run training
Launch the example from the repository root:
GPU (4 GPUs):
bash examples/grpoguard_trainer/qwen_image/run_qwen_image_ocr_lora.sh
NPU (8 NPUs, Atlas 800T A2):
The NPU script requires the CANN software stack. Before running, set the ASCEND_HOME_PATH environment variable (defaults to /usr/local/Ascend/cann-9.0.0).
bash examples/grpoguard_trainer/qwen_image/run_qwen_image_ocr_lora_npu.sh
The scripts run python3 -m verl_omni.trainer.main_diffusion with:
algorithm.adv_estimator=flow_grpoactor_rollout_ref.model.path=Qwen/Qwen-Imageactor_rollout_ref.model.lora_rank=64actor_rollout_ref.model.lora_alpha=128actor_rollout_ref.rollout.name=vllm_omniactor_rollout_ref.actor.diffusion_loss.loss_mode=grpo_guardactor_rollout_ref.actor.diffusion_loss.clip_ratio=2e-6actor_rollout_ref.rollout.algo.sde_type=sdereward.custom_reward_function.name=compute_score_ocr
Due to differences in memory capacity, the NPU and GPU configurations differ as follows:
Parameter |
GPU |
NPU |
|---|---|---|
|
gpu (default) |
|
|
default |
|
|
4 |
8 |
|
1 |
2 |
|
16 |
4 |
Logging
W&B logging is enabled by default in the example script:
export WANDB_API_KEY=<your_wandb_api_key>
The script sets:
trainer.logger='["console", "wandb"]'
trainer.project_name=grpo_guard
trainer.experiment_name=qwen_image_ocr_lora
Override these values on the command line if you want to log under a different project or run name.
Diffusion-specific metrics
See the Metrics Documentation for a full description of all diffusion-specific training metrics.