BAGEL-7B-MoT FlowGRPO training
Last updated: 07/06/2026
BAGEL-7B-MoT is a
Mixture-of-Transformers model supporting both image understanding and
generation. Unlike Qwen-Image, BAGEL is a non-diffusers model — it
cannot be loaded by diffusers and uses its own weight-loading path via
NonDiffusersModelBase. See
How to Integrate a Non-Diffusers Model for FlowGRPO Training
for the integration architecture.
Prerequisites
Install VeRL-Omni (see installation guide).
4 GPUs or 8 NPUs. Run commands from the repository root.
Download the checkpoint:
huggingface-cli download ByteDance-Seed/BAGEL-7B-MoT --local-dir ~/models/ByteDance-Seed/BAGEL-7B-MoT
OCR training
We use an OCR (optical character recognition) dataset that provides
ground-truth text for evaluating image-generation quality. Prompts are
stored in standard chat-message format for the agent loop (see
bagel_ocr.py).
Prepare the dataset
Preprocess the raw OCR data into parquet:
export WORKSPACE=${WORKSPACE:-$HOME}
python3 examples/flowgrpo_trainer/data_process/bagel_ocr.py \
--model_path ~/models/ByteDance-Seed/BAGEL-7B-MoT \
--input_dir ~/data/ocr \
--output_dir $WORKSPACE/data/ocr/bagel
This produces $WORKSPACE/data/ocr/bagel/train.parquet and
test.parquet.
Run training
For GPU:
bash examples/flowgrpo_trainer/bagel/run_bagel_ocr_lora.sh
For NPU:
bash examples/flowgrpo_trainer/bagel/run_bagel_ocr_lora_npu.sh
The launch script uses a Qwen3-VL-8B-Instruct
reward model with vLLM rollout (TP=4) and the genrm_ocr.py custom reward
function.
PickScore training
PickScore evaluates image-text alignment using a
CLIP-based model. The
reward function lives entirely in verl_omni/utils/reward_score/pickscore_reward.py
— there is no separate vLLM reward model deployment, so the GPU is shared
between the actor and the reward computation.
Prepare the dataset
The raw PickScore dataset (train.txt / test.txt) should be downloaded
from the flow_grpo repository.
Preprocess for BAGEL:
python3 examples/flowgrpo_trainer/data_process/bagel_pickscore.py \
--model_path ~/models/ByteDance-Seed/BAGEL-7B-MoT \
--input_dir ~/data/pickscore \
--output_dir $WORKSPACE/data/pickscore/bagel
This produces $WORKSPACE/data/pickscore/bagel/train.parquet and
test.parquet.
Run LoRA training
bash examples/flowgrpo_trainer/bagel/run_bagel_pickscore_lora.sh
Key configuration differences from OCR:
No
reward.reward_model.*flags — PickScore runs as a custom reward function on the rollout GPU.Higher
noise_level(1.3vs0.7) and SDE window (sde_window_size=2,range=[0,7]) to provide sufficient exploration for text-alignment learning.
Run full-weight (non-LoRA) training
A full-weight training variant is available that trains the entire
generation pathway (moe_gen parameters) while keeping the understanding
pathway frozen:
bash examples/flowgrpo_trainer/bagel/run_bagel_pickscore.sh
Key differences from the LoRA variant:
Aspect |
LoRA |
Full-weight |
|---|---|---|
Script |
|
|
Strategy |
default |
|
Trainable params |
Low-rank adapters on |
All |
|
64 / 128 |
N/A |
|
2 |
3 (more exploration for full-weight) |
Why FSDP2 is required. FSDP1 does not natively support mixed
requires_grad within a single wrapped module — some parameters frozen,
others trainable. FSDP2 handles this correctly and also reshards layer
parameters after forward, reducing peak memory during gradient
checkpointing. The understanding pathway (moe_und) is not a LoRA
wrapper replacement but simply has requires_grad=False set by the
configure_trainable_params hook.
Key differences from Qwen-Image
Aspect |
Qwen-Image |
BAGEL-7B-MoT |
|---|---|---|
Model loading |
diffusers |
Custom |
Architecture |
Auto-detected |
Explicit: |
Deploy config |
Not needed |
|
LoRA targets |
|
|
FSDP prefixes |
|
|
CFG |
Standard true CFG |
3-branch (gen / text-uncond / img-uncond) with global renormalisation |
Timestep convention |
|
Raw sigma with SD3-style shift of 3.0 |
Further reading
How to Integrate a Non-Diffusers Model for FlowGRPO Training — full integration guide using BAGEL as the worked example