Diffusion Serving Benchmark (Image/Video)
This folder contains an online-serving benchmark script for diffusion models. It sends requests to a vLLM OpenAI-compatible endpoint and reports throughput, latency percentiles, and optional SLO attainment.
The main entrypoint is:
benchmarks/diffusion/diffusion_benchmark_serving.py
1. Quick Start
- Start the server:
vllm serve Qwen/Qwen-Image --omni --port 8099
- Run a minimal benchmark:
python3 benchmarks/diffusion/diffusion_benchmark_serving.py \
--base-url http://localhost:8099 \
--model Qwen/Qwen-Image \
--task t2i \
--dataset vbench \
--num-prompts 5
Notes
- By default, image tasks talk to
http://<host>:<port>/v1/chat/completions; video tasks talk to/v1/videos. - If you run the server on another host or port, pass
--base-urlaccordingly.
2. Supported Datasets
The benchmark supports three dataset modes via --dataset:
vbench: Built-in prompt/data loader.trace: Heterogeneous request traces (each request can have different resolution/frames/steps).random: Synthetic prompts for quick smoke tests.
VBench dataset
vbench only provides prompt data (and image paths for i2v/i2i); it does not carry
per-request generation fields. In this mode, all requests share CLI values:
--width --height --num-frames --fps --num-inference-steps
(pass --width and --height together).
Example (t2v):
python3 benchmarks/diffusion/diffusion_benchmark_serving.py \
--base-url http://localhost:8099 \
--model Wan-AI/Wan2.2-T2V-A14B-Diffusers \
--task t2v \
--dataset vbench \
--num-prompts 50 \
--width 640 --height 480 \
--num-frames 81 --fps 16 \
--num-inference-steps 40
Note: vbench can also be used for other tasks such as t2i / i2v (and i2i). For t2i, the loader reuses VBench t2v text prompts; for i2v / i2i, it loads the VBench i2v dataset (with image paths).
If you use i2v/i2i bench datasets and need auto-download support, you may need:
uv pip install gdown
Trace dataset
Use --dataset trace to replay a trace file. The trace can specify per-request fields such as:
width,heightnum_frames(video)num_inference_stepsseed,fps- optional
slo_ms(per-request SLO target)
By default (when --dataset-path is not provided), the script downloads a default trace from
the HuggingFace dataset repo asukaqaqzz/Dit_Trace. The default filename can depend on --task
(e.g., t2v uses a video trace).
Current defaults:
--task t2i->sd3_trace.txt--task t2v->cogvideox_trace.txt
You can point to your own trace using --dataset-path.
3. Benchmark Parameters
Basic flags
--base-url: Server address;--endpointselects the path appended to this base URL.--model: The OpenAI-compatiblemodelfield.--endpoint: API endpoint path. Leading/is optional, e.g./v1/videosorv1/videos.--task: Task type (e.g.,t2i,t2v,i2i,i2v).--dataset: Dataset mode (vbench/trace/random).--num-prompts: Number of requests to send.
Common optional flags:
--output-file: Write metrics to a JSON file.--disable-tqdm: Disable the progress bar.
Resolution / frames / steps: CLI defaults vs dataset fields
Related flags: --width, --height, --num-frames, --fps, --num-inference-steps.
- For
vbench/random: these CLI flags act as global defaults for all generated requests. - For
trace: requests can carry their own fields (e.g.,width/height/num_frames/num_inference_steps), with overrides/fallbacks as below.
Precedence rules for trace (i.e., what actually gets sent):
width/height: if either--widthor--heightis explicitly set, it overrides per-request values from the trace; otherwise per-request values are used when present.num_frames: per-requestnum_framestakes precedence; otherwise fall back to--num-frames.num_inference_steps: per-requestnum_inference_stepstakes precedence; otherwise fall back to--num-inference-steps.
SLO, warmup, and max concurrency
Enable SLO evaluation with --slo.
- If a request in the trace already has
slo_ms, that value is used. - Otherwise, the script runs warmup requests to infer a base unit time, estimates
expected_msby linearly scaling with area/frames/steps, and then setsslo_ms = expected_ms * --slo-scale.
Warmup flags:
--warmup-requests: Number of warmup requests.--warmup-num-inference-steps: Steps used during warmup.--warmup-concurrency: Maximum concurrent warmup requests. Use this to warm the same batch shape as the measured run instead of warming only batch=1.- For
--task t2v: warmup requests are forced to usenum_frames=1to make warmup faster and less noisy.
Traffic / concurrency flags:
--request-rate: Target request rate (requests/second). If set toinf, the script sends all requests immediately.--max-concurrency: Max number of in-flight requests (default:1). This can hard-cap the achieved QPS: if it is too small, requests will queue behind the semaphore, and both achieved throughput and observed SLO attainment can be skewed.
Batched warmup note
For batched serving runs, warm the same in-flight shape you plan to measure.
For example, a run with --max-concurrency 8 should usually also use
--warmup-requests 8 --warmup-concurrency 8; otherwise the first measured
batch may still pay compile or CUDA-graph capture cost.
For a Qwen-Image continuous-batching replay example, see
performance_dashboard/qwen_image_serving_performance.md.