Files
wanshuiyin--auto-claude-cod…/docs/tutorials/diffusion_post_training_tutorial.html
2026-07-13 13:37:02 +08:00

1136 lines
106 KiB
HTML
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html lang="zh-CN">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Diffusion Post-Training 面试 Cheat Sheet</title>
<meta name="generator" content="ARIS render-html (academic, v1)">
<meta name="aris:source-path" content="docs/tutorials/diffusion_post_training_tutorial.md">
<meta name="aris:source-sha256" content="95b47c84420904cb38b68b156d25372d305084281e28fec71d19c405801efef8">
<meta name="aris:generated-at" content="2026-05-19 16:25 UTC">
<!-- MathJax 3 -->
<script>
window.MathJax = {
tex: { inlineMath: [['$', '$'], ['\\(', '\\)']], displayMath: [['$$', '$$'], ['\\[', '\\]']], processEscapes: true },
options: { skipHtmlTags: ['script', 'noscript', 'style', 'textarea', 'pre', 'code'] }
};
</script>
<script src="https://cdn.jsdelivr.net/npm/mathjax@3/es5/tex-mml-chtml.js" async></script>
<!-- highlight.js -->
<link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/highlightjs/cdn-release@11.9.0/build/styles/atom-one-light.min.css">
<script src="https://cdn.jsdelivr.net/gh/highlightjs/cdn-release@11.9.0/build/highlight.min.js"></script>
<script>document.addEventListener('DOMContentLoaded', () => hljs.highlightAll());</script>
<style>
:root {
--bg: #fdfcf7;
--bg-soft: #f4f1ea;
--bg-code: #f8f5ec;
--ink: #1a1a1a;
--ink-soft: #4a4a4a;
--ink-muted: #6b6b6b;
--primary: #1a4a8c;
--primary-soft: #2d6cb8;
--accent: #b8390e;
--warn: #b45309;
--warn-bg: #fef3c7;
--info-bg: #dbeafe;
--good-bg: #d1fae5;
--good: #065f46;
--bad-bg: #fee2e2;
--bad: #991b1b;
--border: #d6d0c0;
--border-soft: #e8e3d5;
}
* { box-sizing: border-box; }
html { scroll-behavior: smooth; }
body {
font-family: "Source Serif Pro", "Source Serif 4", "Crimson Pro", "Georgia", "Songti SC", "STSong", serif;
line-height: 1.65;
color: var(--ink);
background: var(--bg);
margin: 0;
padding: 0;
font-size: 16px;
}
.layout {
max-width: 1280px;
margin: 0 auto;
display: grid;
grid-template-columns: 260px 1fr;
gap: 48px;
padding: 40px 32px;
}
nav.toc {
position: sticky;
top: 24px;
align-self: start;
font-size: 13px;
max-height: calc(100vh - 48px);
overflow-y: auto;
border-right: 1px solid var(--border-soft);
padding-right: 16px;
}
nav.toc h3 {
margin: 0 0 12px;
font-size: 12px;
text-transform: uppercase;
letter-spacing: 0.08em;
color: var(--ink-muted);
font-weight: 600;
}
nav.toc ol { list-style: none; padding: 0; margin: 0; counter-reset: toc; }
nav.toc ol li { margin: 5px 0; counter-increment: toc; }
nav.toc ol li::before { content: counter(toc) ". "; color: var(--ink-muted); margin-right: 4px; }
nav.toc a {
color: var(--ink-soft);
text-decoration: none;
border-bottom: 1px dotted transparent;
}
nav.toc a:hover { color: var(--primary); border-bottom-color: var(--primary); }
nav.toc ul { list-style: none; padding-left: 14px; margin: 3px 0; font-size: 12px; }
nav.toc ul li::before { content: "→ "; color: var(--border); }
main { min-width: 0; }
header.hero {
border-bottom: 3px double var(--primary);
padding-bottom: 24px;
margin-bottom: 32px;
}
header.hero .eyebrow {
color: var(--accent);
font-size: 13px;
text-transform: uppercase;
letter-spacing: 0.12em;
font-weight: 600;
margin-bottom: 8px;
}
header.hero h1 {
font-size: 32px;
line-height: 1.2;
margin: 0 0 12px;
color: var(--ink);
font-weight: 700;
letter-spacing: -0.01em;
}
header.hero .subtitle {
font-size: 16px;
color: var(--ink-soft);
margin: 0 0 8px;
font-style: italic;
}
header.hero .byline {
font-size: 14px;
color: var(--ink-soft);
margin: 0 0 20px;
}
header.hero .byline strong {
color: var(--ink);
font-weight: 600;
}
header.hero .meta {
display: flex;
gap: 20px;
flex-wrap: wrap;
font-size: 12px;
color: var(--ink-muted);
border-top: 1px solid var(--border-soft);
padding-top: 14px;
}
header.hero .meta span strong { color: var(--ink-soft); }
header.hero .meta code {
font-family: "JetBrains Mono", "SF Mono", "Menlo", "Consolas", monospace;
font-size: 11px;
background: var(--bg-soft);
padding: 1px 5px;
border-radius: 3px;
border: 1px solid var(--border-soft);
}
h2 {
font-size: 24px;
margin: 44px 0 14px;
padding-bottom: 8px;
border-bottom: 1px solid var(--border);
color: var(--ink);
font-weight: 700;
}
h2 .num { color: var(--primary); font-weight: 600; margin-right: 8px; }
h3 { font-size: 19px; margin: 28px 0 10px; color: var(--primary); font-weight: 600; }
h4 { font-size: 16px; margin: 20px 0 8px; color: var(--ink); font-weight: 600; }
p { margin: 10px 0; }
ul, ol { padding-left: 22px; margin: 10px 0; }
ul li, ol li { margin: 4px 0; }
ul li::marker { color: var(--primary); }
strong { color: var(--accent); font-weight: 600; }
em { color: var(--ink-soft); }
a { color: var(--primary); }
a:hover { color: var(--accent); }
code:not(.hljs) {
font-family: "JetBrains Mono", "SF Mono", "Menlo", "Consolas", monospace;
font-size: 0.86em;
background: var(--bg-code);
padding: 1px 5px;
border-radius: 3px;
border: 1px solid var(--border-soft);
color: var(--accent);
}
pre {
background: #fafaf6;
border: 1px solid var(--border);
border-left: 4px solid var(--primary);
padding: 0;
overflow-x: auto;
border-radius: 4px;
margin: 14px 0;
}
pre code, pre code.hljs {
background: transparent !important;
display: block;
padding: 14px 18px !important;
font-size: 13px;
line-height: 1.55;
font-family: "JetBrains Mono", "SF Mono", "Menlo", monospace;
color: var(--ink);
}
pre.diagram {
background: #f9f6ed;
border-left: 4px solid var(--accent);
font-size: 12.5px;
line-height: 1.4;
}
.callout {
margin: 16px 0;
padding: 12px 16px;
border-radius: 4px;
border-left: 4px solid;
font-size: 15px;
}
.callout-title {
font-weight: 600;
margin-bottom: 6px;
font-size: 12px;
text-transform: uppercase;
letter-spacing: 0.06em;
}
.callout-info { background: var(--info-bg); border-left-color: var(--primary); }
.callout-info .callout-title { color: var(--primary); }
.callout-warn { background: var(--warn-bg); border-left-color: var(--warn); }
.callout-warn .callout-title { color: var(--warn); }
.callout-good { background: var(--good-bg); border-left-color: var(--good); }
.callout-good .callout-title { color: var(--good); }
.callout-bad { background: var(--bad-bg); border-left-color: var(--bad); }
.callout-bad .callout-title { color: var(--bad); }
table {
width: 100%;
border-collapse: collapse;
margin: 16px 0;
font-size: 14px;
border: 1px solid var(--border);
border-radius: 4px;
overflow: hidden;
}
thead { background: var(--primary); color: white; }
th, td {
text-align: left;
padding: 9px 12px;
border-bottom: 1px solid var(--border-soft);
vertical-align: top;
}
th { font-weight: 600; font-size: 13px; letter-spacing: 0.02em; }
tr:last-child td { border-bottom: none; }
tbody tr:nth-child(even) { background: var(--bg-soft); }
details.qa, details {
background: white;
border: 1px solid var(--border-soft);
border-radius: 6px;
margin: 10px 0;
padding: 0;
}
details summary {
cursor: pointer;
padding: 10px 14px;
font-weight: 600;
font-size: 14px;
color: var(--primary);
list-style: none;
user-select: none;
}
details summary::-webkit-details-marker { display: none; }
details summary::before {
content: "▸ ";
margin-right: 4px;
display: inline-block;
transition: transform 0.15s;
}
details[open] summary::before { transform: rotate(90deg); }
details[open] summary { border-bottom: 1px solid var(--border-soft); }
details > :not(summary) { padding: 10px 14px; }
details p:first-of-type { margin-top: 8px; }
mjx-container[display="true"] { margin: 12px 0 !important; }
footer.aris-footer {
margin-top: 60px;
padding-top: 20px;
border-top: 1px solid var(--border);
font-size: 12px;
color: var(--ink-muted);
}
footer.aris-footer a { color: var(--ink-muted); border-bottom: 1px dotted var(--border); }
@media (max-width: 900px) {
.layout { grid-template-columns: 1fr; gap: 20px; padding: 20px 16px; }
nav.toc {
position: static;
max-height: none;
border-right: none;
border-bottom: 1px solid var(--border-soft);
padding-right: 0;
padding-bottom: 14px;
}
header.hero h1 { font-size: 24px; }
h2 { font-size: 20px; }
}
@media print {
nav.toc { display: none; }
.layout { grid-template-columns: 1fr; padding: 0; }
body { background: white; }
header.hero { border-bottom-color: var(--ink); }
}
</style>
</head>
<body>
<div class="layout">
<nav class="toc">
<h3>Contents</h3>
<ol>
<li><a href="#0-tldr">§0 TL;DR</a>
</li>
<li><a href="#1-直觉为何-diffusion-post-training-难">§1 直觉:为何 diffusion post-training 难</a>
<ul>
<li><a href="#11-单步-vs-多步生成的本质差异">1.1 单步 vs 多步生成的本质差异</a></li>
<li><a href="#12-三条主线分类">1.2 三条主线分类</a></li>
<li><a href="#13-一句话直觉">1.3 一句话直觉</a></li>
<li><a href="#14-convention全文统一">1.4 Convention(全文统一)</a></li>
</ul>
</li>
<li><a href="#2-rl-for-diffusionddpo-与-dpok">§2 RL for DiffusionDDPO 与 DPOK</a>
<ul>
<li><a href="#21-把-denoising-当-mdpddpo-视角">2.1 把 denoising 当 MDPDDPO 视角)</a></li>
<li><a href="#22-ddpo-sf-score-function-算法">2.2 DDPO-SF (Score Function) 算法</a></li>
<li><a href="#23-ddpo-is-importance-sampling-ppo-style">2.3 DDPO-IS (Importance Sampling, PPO-style)</a></li>
<li><a href="#24-ddpo-的两个-reward-实验">2.4 DDPO 的两个 reward 实验</a></li>
<li><a href="#25-dpokkl-regularized-rl-for-diffusion">2.5 DPOKKL-regularized RL for diffusion</a></li>
<li><a href="#26-ddpo-失败模式与缓解">2.6 DDPO 失败模式与缓解</a></li>
</ul>
</li>
<li><a href="#3-direct-reward-fine-tuningdraft--alignprop--refl">§3 Direct Reward Fine-TuningDRaFT / AlignProp / ReFL</a>
<ul>
<li><a href="#31-核心-ideareward-当-differentiable-loss">3.1 核心 ideareward 当 differentiable loss</a></li>
<li><a href="#32-draft-clark-et-al-2024-iclr-arxiv-230917400">3.2 DRaFT (Clark et al. 2024 ICLR, arXiv 2309.17400)</a></li>
<li><a href="#33-alignprop-prabhudesai-et-al-arxiv-231003739-2023-10iclr-2024-venue-arxiv-后被-supersededwithdrawn">3.3 AlignProp (Prabhudesai et al. arXiv 2310.03739, 2023-10ICLR 2024 venue; arXiv 后被 superseded/withdrawn)</a></li>
<li><a href="#34-refl-xu-et-al-2023-neurips-arxiv-230405977">3.4 ReFL (Xu et al. 2023 NeurIPS, arXiv 2304.05977)</a></li>
<li><a href="#35-三者对比">3.5 三者对比</a></li>
<li><a href="#36-reward-hacking-in-direct-reward-backprop">3.6 Reward hacking in direct reward backprop</a></li>
</ul>
</li>
<li><a href="#4-preference-optimizationdiffusion-dpo-家族">§4 Preference OptimizationDiffusion-DPO 家族</a>
<ul>
<li><a href="#41-diffusion-dpo-wallace-et-al-2024-cvpr-arxiv-231112908">4.1 Diffusion-DPO (Wallace et al. 2024 CVPR, arXiv 2311.12908)</a></li>
<li><a href="#42-d3po-yang-et-al-2024-cvpr-arxiv-231113231">4.2 D3PO (Yang et al. 2024 CVPR, arXiv 2311.13231)</a></li>
<li><a href="#43-spo-liang-et-al-2024-arxiv-240604314">4.3 SPO (Liang et al. 2024, arXiv 2406.04314)</a></li>
<li><a href="#44-diffusion-kto-li-et-al-2024-neurips-arxiv-240404465">4.4 Diffusion-KTO (Li et al. 2024 NeurIPS, arXiv 2404.04465)</a></li>
<li><a href="#45-mapo-hong-et-al-2024-arxiv-240606424">4.5 MaPO (Hong et al. 2024, arXiv 2406.06424)</a></li>
<li><a href="#46-dpo-家族总览表">4.6 DPO 家族总览表</a></li>
</ul>
</li>
<li><a href="#5-flow-grpoflow-matching-的-rl">§5 Flow-GRPOFlow Matching 的 RL</a>
<ul>
<li><a href="#51-为什么-flow-matching-也要-post-training">5.1 为什么 Flow Matching 也要 post-training</a></li>
<li><a href="#52-flow-grpo-的两个核心-trick">5.2 Flow-GRPO 的两个核心 trick</a></li>
<li><a href="#53-flow-grpo-的-advantage-计算">5.3 Flow-GRPO 的 advantage 计算</a></li>
<li><a href="#54-flow-grpo-的-loss">5.4 Flow-GRPO 的 loss</a></li>
<li><a href="#55-vector-field-的-advantage-几何意义">5.5 vector field 的 advantage 几何意义</a></li>
<li><a href="#56-flow-grpo-实测结果">5.6 Flow-GRPO 实测结果</a></li>
</ul>
</li>
<li><a href="#6-code-patterns可读伪代码">§6 Code Patterns(可读伪代码)</a>
<ul>
<li><a href="#61-ddpo-reinforce-style-update">6.1 DDPO REINFORCE-style update</a></li>
<li><a href="#62-diffusion-dpo-loss">6.2 Diffusion-DPO loss</a></li>
<li><a href="#63-alignprop--draft-反传-with-checkpointing">6.3 AlignProp / DRaFT 反传 with checkpointing</a></li>
<li><a href="#64-spo-step-aware-preference-loss">6.4 SPO step-aware preference loss</a></li>
<li><a href="#65-flow-grpo-group-relative-advantage">6.5 Flow-GRPO group-relative advantage</a></li>
<li><a href="#66-combined-reward-signal">6.6 Combined reward signal</a></li>
</ul>
</li>
<li><a href="#7-reward-design--失败模式">§7 Reward Design &amp; 失败模式</a>
<ul>
<li><a href="#71-reward-model-选择">7.1 Reward model 选择</a></li>
<li><a href="#72-reward-hacking-gallerydiffusion-特色">7.2 Reward hacking gallerydiffusion 特色)</a></li>
<li><a href="#73-step-level-vs-trajectory-level-reward">7.3 Step-level vs trajectory-level reward</a></li>
<li><a href="#74-缓解-reward-hacking-的核心机制">7.4 缓解 reward hacking 的核心机制</a></li>
</ul>
</li>
<li><a href="#8-production-landscapesd3--flux-用了什么">§8 Production LandscapeSD3 / FLUX 用了什么</a>
<ul>
<li><a href="#81-公开论文--报告说了什么">8.1 公开论文 / 报告说了什么</a></li>
<li><a href="#82-sd3--flux-是否用了-post-training">8.2 SD3 / FLUX 是否用了 post-training</a></li>
<li><a href="#83-工业级-pipeline-假说">8.3 工业级 pipeline 假说</a></li>
</ul>
</li>
<li><a href="#9-vs-llm-rlhf-对比">§9 vs LLM RLHF 对比</a>
<ul>
<li><a href="#91-一表看完">9.1 一表看完</a></li>
<li><a href="#92-共同-lesson">9.2 共同 lesson</a></li>
<li><a href="#93-独有差异">9.3 独有差异</a></li>
</ul>
</li>
<li><a href="#10-25-高频面试题">§10 25 高频面试题</a>
<ul>
<li><a href="#l1-必会题10-题">L1 必会题(10 题)</a></li>
<li><a href="#l2-进阶题10-题">L2 进阶题(10 题)</a></li>
<li><a href="#l3-顶级-lab-题5-题">L3 顶级 lab 题(5 题)</a></li>
</ul>
</li>
<li><a href="#a-附录">§A 附录</a>
<ul>
<li><a href="#a1-关键论文清单含-arxiv-id">A.1 关键论文清单(含 arXiv ID)</a></li>
<li><a href="#a2-常用-reward-model-资源">A.2 常用 reward model 资源</a></li>
<li><a href="#a3-开源训练代码">A.3 开源训练代码</a></li>
<li><a href="#a4-工程踩坑清单">A.4 工程踩坑清单</a></li>
<li><a href="#a5-与-0-tldr-的呼应">A.5 与 §0 TL;DR 的呼应</a></li>
</ul>
</li>
</ol>
</nav>
<main>
<header class="hero">
<div class="eyebrow">Interview Prep · Diffusion Post-Training</div>
<h1>Diffusion Post-Training 面试 Cheat Sheet</h1>
<p class="subtitle">RL + Preference Optimization for Diffusion/Flow Models (DDPO/Diffusion-DPO/Flow-GRPO) + 25 高频题(L1 必会 · L2 进阶 · L3 顶级 lab)</p>
<p class="byline">By <strong>Ruofeng Yang (杨若峰), Shanghai Jiao Tong University</strong></p>
<div class="meta">
<span><strong>Source:</strong> <code>docs/tutorials/diffusion_post_training_tutorial.md</code></span>
<span><strong>SHA256:</strong> <code>95b47c844209</code></span>
<span><strong>Rendered:</strong> 2026-05-19 16:25 UTC</span>
</div>
</header>
<h2 id="0-tldr">§0 TL;DR</h2>
<div class="callout callout-info"><div class="callout-title">9 句话搞定 Diffusion Post-Training</div><p>一页拿下 RL/DPO/Flow-RL 全家桶(详见 §1–§10 推导)。</p></div>
<ol><li><strong>为什么难</strong>diffusion 是多步 $T$ 步生成(典型 20–50 步),reward 只在终态 $x_0$ 给一次 —— <strong>稀疏 terminal reward + 长 denoising trajectory + credit assignment</strong> 三件事叠加,比 LLM RLHF 多一个"轨迹积分"维度。</li><li><strong>三条主线</strong>(i) RL on denoising MDPDDPO / DPOK,把 $T$ 步 denoising 当 MDP);(ii) Direct reward backpropDRaFT / AlignProp / ReFL,把 reward 当 differentiable loss 沿 $T$ 步反传);(iii) Preference optimizationDiffusion-DPO / D3PO / SPO / Diffusion-KTO / MaPO,把 LLM DPO 家族搬到 diffusion)。</li><li><strong>DDPO (Black et al. 2024 ICLR, arXiv 2305.13301)</strong>denoising 视为 $T$-步 MDPstate $= (x_t, t, c)$action $= x_{t-1}$per-trajectory reward $R(x_0, c)$;用 REINFORCE 或 PPO-clip 更新 $\log p_\theta(x_{t-1} \mid x_t, c)$。</li><li><strong>AlignProp (Prabhudesai et al. 2024 ICLR, arXiv 2310.03739)</strong> &amp; <strong>DRaFT (Clark et al. 2024 ICLR, arXiv 2309.17400)</strong>reward $R$ 关于 $x_0$ 可导时,直接把 $R(x_0)$ 沿 $T$ 步 sampler <strong>反传</strong>到 $\theta$。<strong>关键工程问题</strong>:显存 $\mathcal{O}(T)$DRaFT-K / AlignProp 只回传最后 $K$ 步(典型 $K \in \{1, 5\}$),配合 gradient checkpointing 把显存压到 $\mathcal{O}(K)$。</li><li><strong>Diffusion-DPO (Wallace et al. 2024 CVPR, arXiv 2311.12908)</strong>:把 LLM 的 $\log\pi/\pi_\text{ref}$ 换成 diffusion 的 <strong>per-step ELBO surrogate</strong>——具体地,用 $-\|\epsilon - \epsilon_\theta(x_t, t)\|^2$ 作为 $\log p_\theta(x_0)$ 的一个 lower bound 项,对 $(y_w, y_l)$ 拼成 DPO contrastive。</li><li><strong>D3PO (Yang et al. 2024 CVPR, arXiv 2311.13231)</strong><strong>完全免 RM</strong>——直接把人类对生成图片的 thumbs up/down 信号代入 KL-regularized 最优解的 implicit reward;推导上与 DPO 平行,但放到 diffusion <strong>per-step Markov chain</strong> 上。</li><li><strong>SPO (Liang et al. 2024, arXiv 2406.04314)</strong>:观察到不同 denoising step 偏好不同(高噪 step 学构图,低噪 step 学细节),把 DPO 推广为 <strong>step-aware</strong>——每个 $t$ 单独采 in-step pair $(x_{t-1}^w, x_{t-1}^l)$loss 在 step 维度上加权。</li><li><strong>Flow-GRPO (Liu et al. 2025, arXiv 2505.05470)</strong>:第一个把 GRPO 搬到 Flow Matching 的工作。两个关键 trick<strong>ODE→SDE 等价转换</strong>让确定性 flow 变可探索的随机过程;<strong>denoising reduction</strong> 训练时减步、推理时全步。RL-tuned SD3.5-M 把 GenEval 从 63% 拉到 95%。</li><li><strong>Reward hacking is the real boss</strong>:过饱和颜色、构图单调、风格收敛、PickScore 高但人眼丑 —— 缓解靠 reward ensemble (HPSv2 + PickScore + ImageReward + CLIP-Score)、KL anchor (Diffusion-DPO 的 $\beta$)、early stop on reward plateau。SD3 / FLUX <strong>几乎不公开 post-training 细节</strong>,但社区主流认为 SD3.5 Turbo 系列、FLUX.1 dev 走的是 DPO + 蒸馏混合路线。</li></ol>
<div class="callout callout-good"><div class="callout-title">vs LLM RLHF 一句话对比</div><p>LLM RLHF 关心 "token-level credit assignment + KL anchor"diffusion post-training 关心 "denoising-step credit assignment + 显存爆炸 (backprop) 或 sample 爆炸 (RL)"。本质相同问题——稀疏 reward + 长轨迹——只是轨迹的物理含义换了。</p></div>
<h2 id="1-直觉为何-diffusion-post-training-难">§1 直觉:为何 diffusion post-training 难</h2>
<h3 id="11-单步-vs-多步生成的本质差异">1.1 单步 vs 多步生成的本质差异</h3>
<p>LLM 的 reward 一般也是 sequence-level,但 token 是离散、轨迹长 $L \sim 10^3$、词表中等大。diffusion 的"轨迹"是 $T$ 步 denoising,每步操作的是连续高维张量 $x_t \in \mathbb{R}^{C \times H \times W}$SDXL latent 是 $4\times128\times128 = 65536$ 维),$T$ 典型 2050。</p>
<table><thead><tr><th>维度</th><th>LLM RLHF</th><th>Diffusion Post-Training</th></tr></thead><tbody><tr><td>轨迹长度</td><td>$L$response token 数)</td><td>$T$denoising step 数,典型 2050</td></tr><tr><td>单步动作</td><td>离散 token</td><td>$\mathbb{R}^d$ 连续向量($d \sim 10^4$$10^5$</td></tr><tr><td>Reward 频率</td><td>通常只在终态</td><td>通常只在终态 $x_0$</td></tr><tr><td>探索性</td><td>sampling temperature / top-p</td><td>DDIM 是 deterministic,需要"加噪"才能探索(DDPO 用 stochastic DDPMFlow-GRPO 用 ODE→SDE</td></tr><tr><td>Reward 来源</td><td>trained RM (BT) / rule</td><td>trained image RM (ImageReward / HPSv2 / PickScore) / rule (object count / OCR)</td></tr><tr><td>显存瓶颈</td><td>4 副本 (policy + ref + RM + V)</td><td>1 副本 UNet/DiT,但<strong>直接反传时</strong>需存 $T$ 步 activation</td></tr></tbody></table>
<h3 id="12-三条主线分类">1.2 三条主线分类</h3>
<ul><li><strong>Line A (RL on denoising MDP)</strong>:把 diffusion 当成 RL 环境,<strong>不要求 reward 可导</strong>。代表:DDPO、DPOK、Flow-GRPO。</li><li><strong>Line B (Direct reward backprop)</strong>reward 关于 $x_0$ 可导时<strong>直接梯度下降</strong>,类似把 reward 当一个新 loss。代表:DRaFT、AlignProp、ReFL。</li><li><strong>Line C (Preference optimization, DPO-style)</strong>:把 LLM DPO 系移植到 diffusion<strong>不再 sample on-policy</strong>。代表:Diffusion-DPO、D3PO、SPO、Diffusion-KTO、MaPO。</li></ul>
<pre class="diagram"><code> 偏好/奖励信号
┌───────────────────────┼───────────────────────┐
│ │ │
reward 不可导 reward 可导 偏好对 (offline)
│ │ │
Line A: RL Line B: backprop Line C: DPO 家族
DDPO, DPOK, DRaFT, AlignProp, Diffusion-DPO,
Flow-GRPO ReFL D3PO, SPO, KTO, MaPO
│ │ │
最通用 显存友好 (K-step) off-policy 快
但 sample 贵 但要求可导 但需偏好数据</code></pre>
<h3 id="13-一句话直觉">1.3 一句话直觉</h3>
<div class="callout callout-info"><div class="callout-title">核心直觉</div></div>
<ul><li>Line A 把 $T$ 步 denoising 当 RL 轨迹:每步是一个 stochastic policy 输出。</li><li>Line B 把 reward 当 differentiable loss:用 $T$ 步反传,但显存 $O(T)$ 把人压垮。</li><li>Line C 把 reward 完全 bypass:用偏好对的 implicit reward = $\beta \log(p_\theta/p_\text{ref})$。</li></ul>
<h3 id="14-convention全文统一">1.4 Convention(全文统一)</h3>
<table><thead><tr><th>符号</th><th>含义</th></tr></thead><tbody><tr><td>$x_0$</td><td>干净图像(latent 或 pixel</td></tr><tr><td>$x_t,\; t = 0, \dots, T$</td><td>加噪样本;$x_T \approx \mathcal{N}(0, I)$</td></tr><tr><td>$\epsilon_\theta(x_t, t, c)$</td><td>UNet/DiT 预测的噪声(DDPM 参数化)</td></tr><tr><td>$v_\theta(t, x, c)$</td><td>Flow Matching 的 vector field</td></tr><tr><td>$c$</td><td>条件(text embedding / class</td></tr><tr><td>$p_\theta(x_{t-1} \mid x_t, c)$</td><td>reverse process 的单步条件分布</td></tr><tr><td>$R(x_0, c)$</td><td>终态 rewardscalar,可来自 RM 或 rule</td></tr><tr><td>$\pi_\text{ref}$ / $p_\text{ref}$</td><td>参考模型(一般是 SFT 后的 base)</td></tr><tr><td>$\beta$</td><td>KL/温度超参(与 LLM DPO 同义)</td></tr></tbody></table>
<h2 id="2-rl-for-diffusionddpo-与-dpok">§2 RL for DiffusionDDPO 与 DPOK</h2>
<h3 id="21-把-denoising-当-mdpddpo-视角">2.1 把 denoising 当 MDPDDPO 视角)</h3>
<p>Black et al. 2024 ICLR <em>Training Diffusion Models with Reinforcement Learning</em>arXiv 2305.13301)的关键观察:DDPM/DDIM 的 reverse process 本身就是一个<strong>有限 horizon MDP</strong></p>
<p>定义:</p>
<ul><li><strong>State</strong>: $s_t = (x_t, t, c)$,时间反向 $t = T, T-1, \dots, 1$</li><li><strong>Action</strong>: $a_t = x_{t-1}$(从 $p_\theta(\cdot \mid x_t, c)$ 采样得到)</li><li><strong>Transition</strong>: 确定性——$s_{t-1} = (x_{t-1}, t-1, c)$</li><li><strong>Reward</strong>: $r_t = 0$ for $t > 1$$r_1 = R(x_0, c)$terminal-only</li><li><strong>Policy</strong>: $\pi_\theta(a_t \mid s_t) = p_\theta(x_{t-1} \mid x_t, c)$</li></ul>
<p>策略梯度定理:</p>
<p>$$\nabla_\theta J(\theta) = \mathbb{E}_{\tau \sim p_\theta}\!\left[\sum_{t=1}^{T} \nabla_\theta \log p_\theta(x_{t-1} \mid x_t, c)\, R(x_0, c)\right]$$</p>
<p><strong>核心</strong>$\log p_\theta(x_{t-1} \mid x_t, c)$ 在 DDPM 中是 Gaussianlog-prob 可解析写出,所以梯度可直接算。</p>
<h3 id="22-ddpo-sf-score-function-算法">2.2 DDPO-SF (Score Function) 算法</h3>
<p>最朴素的版本(DDPO-SF, score function estimator):</p>
<ol><li><strong>采样阶段</strong> — 从 prompt $c$ 出发跑 $T$ 步 DDPM reverse,得到 trajectory $\tau = (x_T, x_{T-1}, \dots, x_0)$;计算 $R(x_0, c)$。</li><li><strong>更新阶段</strong> — 用 REINFORCE-style 梯度估计:</li></ol>
<p>$$\hat{g} = \frac{1}{N}\sum_{n=1}^{N} \sum_{t=1}^{T} \nabla_\theta \log p_\theta(x_{t-1}^{(n)} \mid x_t^{(n)}, c)\, (R^{(n)} - b)$$</p>
<p>其中 $b$ 是 baseline(典型用 batch mean reward)。</p>
<h3 id="23-ddpo-is-importance-sampling-ppo-style">2.3 DDPO-IS (Importance Sampling, PPO-style)</h3>
<p>DDPO 的实际推荐变体用 <strong>PPO-clip</strong> 在每步上做重要性比:</p>
<p>$$\rho_t = \frac{p_\theta(x_{t-1} \mid x_t, c)}{p_{\theta_\text{old}}(x_{t-1} \mid x_t, c)}, \qquad L^\text{CLIP}_t = \min\!\big(\rho_t R, \text{clip}(\rho_t, 1-\epsilon, 1+\epsilon) R\big)$$</p>
<div class="callout callout-warn"><div class="callout-title">Per-step ratio 而非 trajectory ratio</div><p>diffusion PPO 用 <strong>per-step</strong> importance ratio,不是把 $T$ 步乘起来。因为整条 trajectory 的 ratio 是 $T$ 个比值的积,方差爆炸;per-step clip 在每步独立 clamp 才稳。</p></div>
<h3 id="24-ddpo-的两个-reward-实验">2.4 DDPO 的两个 reward 实验</h3>
<p>Black et al. 用 DDPO + SD-1.5 在四个 reward 上跑:</p>
<table><thead><tr><th>Reward 类型</th><th>例子</th><th>信号性质</th></tr></thead><tbody><tr><td>Compressibility</td><td>JPEG file size</td><td>rule-based scalar</td></tr><tr><td>Aesthetic</td><td>LAION aesthetic predictor</td><td>trained MLP</td></tr><tr><td>Prompt-image alignment</td><td>CLIP-Score / LLaVA judge</td><td>VLM-based</td></tr><tr><td>Object presence</td><td>DETR / OWL-ViT count</td><td>rule-based</td></tr></tbody></table>
<p>实测在所有四个 reward 上,DDPO 比 reward-weighted regressionRWR baseline)涨幅明显,且能从 emoji 风格漂移到油画风格——证明 RL 真的在 explore 而不是简单 mode-seeking。</p>
<h3 id="25-dpokkl-regularized-rl-for-diffusion">2.5 DPOKKL-regularized RL for diffusion</h3>
<p>Fan et al. <strong>2023 NeurIPS</strong> <em>DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models</em>arXiv 2305.16381)和 DDPO 几乎同期,差别在<strong>显式 KL anchor</strong></p>
<p>$$\boxed{\;\max_\theta\; \mathbb{E}_{\tau \sim p_\theta}\!\left[R(x_0, c)\right] - \beta\, \mathbb{E}_c\!\left[\text{KL}\!\big(p_\theta(\cdot \mid c) \,\big\Vert\, p_\text{ref}(\cdot \mid c)\big)\right]\;}$$</p>
<p>KL 项展开到 per-step</p>
<p>$$\text{KL}(p_\theta \Vert p_\text{ref}) = \sum_{t=1}^{T} \mathbb{E}\!\left[\text{KL}\!\big(p_\theta(x_{t-1} \mid x_t, c) \,\Vert\, p_\text{ref}(x_{t-1} \mid x_t, c)\big)\right]$$</p>
<p>由于 DDPM 的 $p_\theta(x_{t-1} \mid x_t, c)$ 是 Gaussian<strong>两个 Gaussian 的 KL 有闭式</strong>per-step KL 直接算。DPOK 用 policy gradient + 这个 KL penalty,等价于 RLHF 的 "$\beta \log(\pi/\pi_\text{ref})$" 的 diffusion 版本。</p>
<div class="callout callout-info"><div class="callout-title">DPOK vs DDPO</div></div>
<ul><li>DDPO:纯 RLREINFORCE 或 PPO-clip),KL 隐式(通过 ratio clip)。</li><li>DPOK:显式 KL 项,与 LLM RLHF 的"reward 上加 KL"对应。</li><li>实测 DPOK 在 reward-prompt alignmentImageReward)上稳一些;DDPO 在 compressibility 等 rule-based reward 上更激进。</li></ul>
<h3 id="26-ddpo-失败模式与缓解">2.6 DDPO 失败模式与缓解</h3>
<table><thead><tr><th>现象</th><th>原因</th><th>缓解</th></tr></thead><tbody><tr><td>Reward 上升但 FID 暴跌</td><td>over-optimization on RM 盲点</td><td>KL penalty / LoRA fine-tune(防 base 漂移)</td></tr><tr><td>同一 prompt 收敛到单一构图</td><td>mode collapsepolicy 找到 RM 高分 mode</td><td>reward ensemble / early stop</td></tr><tr><td>高 reward 但人眼丑</td><td>RM scale 与 human 不对齐</td><td>多 RM 加权 + human eval 校准</td></tr><tr><td>Training 不稳</td><td>per-step ratio 在 $T$ 步上累计</td><td>per-step PPO-clip $\epsilon = 0.1$ 比 LLM 的 $0.2$ 更稳</td></tr></tbody></table>
<h2 id="3-direct-reward-fine-tuningdraft--alignprop--refl">§3 Direct Reward Fine-TuningDRaFT / AlignProp / ReFL</h2>
<h3 id="31-核心-ideareward-当-differentiable-loss">3.1 核心 ideareward 当 differentiable loss</h3>
<p>若 reward $R(x_0, c)$ 关于 $x_0$ <strong>可导</strong>(一般 CNN/ViT RM 都满足),且 diffusion sampler 是 differentiable,则可以<strong>直接对 $\theta$ 反传</strong></p>
<p>$$\theta \leftarrow \theta + \eta\, \nabla_\theta R\!\big(x_0(\theta), c\big), \quad x_0(\theta) = \text{Sample}_\theta^T(c)$$</p>
<p>其中 $\text{Sample}_\theta^T(c)$ 表示从 $x_T \sim \mathcal{N}(0,I)$ 出发跑 $T$ 步 reverse 得到 $x_0$。这是把 diffusion 整条 reverse trajectory 当成一个 <strong>giant differentiable computation graph</strong>end-to-end 优化 reward。</p>
<div class="callout callout-good"><div class="callout-title">优势</div><p>无需 sample variance;梯度信号方差远小于 REINFORCE-style RL。</p></div>
<div class="callout callout-bad"><div class="callout-title">代价</div><p>存 $T$ 步 activationvanilla 实现下显存 $\mathcal{O}(T \cdot M_\text{UNet})$,对 SDXL UNet $\approx$ 数百 GB<strong>完全不可训练</strong></p></div>
<h3 id="32-draft-clark-et-al-2024-iclr-arxiv-230917400">3.2 DRaFT (Clark et al. 2024 ICLR, arXiv 2309.17400)</h3>
<p><em>Directly Fine-Tuning Diffusion Models on Differentiable Rewards</em>。两个核心 trick</p>
<p><strong>Trick 1DRaFT-K,只回传最后 $K$ 步。</strong></p>
<p>完整的 chain ruledenoise step $\epsilon_\theta(x_t,t)$ 既影响下一步 $x_{t-1}$ 又<strong>直接</strong>依赖 $\theta$):</p>
<p>$$\nabla_\theta R(x_0) = \frac{\partial R}{\partial x_0} \cdot \sum_{t=1}^{K}\left(\prod_{s=1}^{t-1} \frac{\partial x_{s-1}}{\partial x_s}\right) \cdot \frac{\partial x_{t-1}}{\partial \theta}\bigg|_{\text{direct}}$$</p>
<p>其中 $\partial x_{t-1}/\partial\theta|_\text{direct}$ 是 step-$t$ 通过 $\epsilon_\theta(x_t,t)$ 直接对 $\theta$ 的偏导(不经过 $x_t \to x_t$ 的间接路径),$\prod_s$ 是 backward 时的 Jacobian product propagation。前 $T-K$ 步用 <code>torch.no_grad()</code> 跑,只在最后 $K$ 步保留 graphK=1 (DRaFT-1) 已能给出极强信号——最后一步对 $x_0$ 直接影响最大。autograd 自动累加所有 $K$ 步的 $\partial/\partial\theta|_\text{direct}$,所以代码只要 <code>loss.backward()</code> 即可。</p>
<p><strong>Trick 2LoRA fine-tune + 高学习率。</strong></p>
<p>只训 LoRA adapter$\sim$1% 参数),base UNet 冻结。配合 gradient checkpointing 把显存压到单卡 24GB 内可训。</p>
<p>伪代码:</p>
<pre><code class="language-python"># DRaFT-K 一步训练
x_t = torch.randn(B, C, H, W).to(device) # x_T
with torch.no_grad():
for t in range(T-1, K, -1): # 前 T-K 步无梯度
x_t = ddim_step(unet_lora, x_t, t, cond)
for t in range(K, 0, -1): # 最后 K 步要梯度
x_t = ddim_step(unet_lora, x_t, t, cond)
x_0 = x_t
reward = image_rm(x_0, prompt) # ImageReward / HPSv2
loss = -reward.mean() # 注意负号——最大化 reward
loss.backward() # 显存 O(K)
optimizer.step()</code></pre>
<div class="callout callout-warn"><div class="callout-title">DRaFT-1 等价于 REINFORCE 吗?</div><p>不等价。DRaFT-1 是<strong>真实 reparameterized gradient</strong>pathwise estimator),REINFORCE 是 score function estimator。前者方差极小但只在 reward 可导时可用;后者通用但方差大。$\nabla \log p$ vs $\partial x / \partial \theta$ 是两类不同的梯度估计。</p></div>
<h3 id="33-alignprop-prabhudesai-et-al-arxiv-231003739-2023-10iclr-2024-venue-arxiv-后被-supersededwithdrawn">3.3 AlignProp (Prabhudesai et al. arXiv 2310.03739, 2023-10ICLR 2024 venue; arXiv 后被 superseded/withdrawn)</h3>
<p><em>Aligning Text-to-Image Diffusion Models with Reward Backpropagation</em>——和 DRaFT 几乎并行的工作(2023 年底先后挂 arXiv),核心 idea 相同:<strong>reward backprop through denoising</strong>。差别:</p>
<table><thead><tr><th>维度</th><th>DRaFT</th><th>AlignProp</th></tr></thead><tbody><tr><td>截断</td><td>DRaFT-K,最后 $K$ 步保留梯度</td><td>随机选 $K$ 步保留梯度(randomized truncated BPTT</td></tr><tr><td>显存优化</td><td>gradient checkpointing</td><td>gradient checkpointing + LoRA</td></tr><tr><td>主推 reward</td><td>HPSv1, PickScore, Aesthetic</td><td>ImageReward, HPSv2, PickScore</td></tr><tr><td>Mode collapse 缓解</td><td>简单 KL anchor</td><td>$\text{LoRA scale}$ 退火 + early stop</td></tr></tbody></table>
<p>AlignProp 的关键贡献是把"为什么 reward backprop 可以工作"理论化——证明在 fixed-point 假设下,截断 BPTT 的梯度是真实梯度的有偏但低方差估计。</p>
<h3 id="34-refl-xu-et-al-2023-neurips-arxiv-230405977">3.4 ReFL (Xu et al. 2023 NeurIPS, arXiv 2304.05977)</h3>
<p><em>ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation</em>。这篇是 ImageReward 的<strong>原始论文</strong>,同时也提出了 ReFL (Reward Feedback Learning) 算法。</p>
<p>ReFL 与 DRaFT 思想接近,但发表更早。它把 reward 直接当 loss,并且<strong>只在一个随机选定的中间 step</strong> 用 reward 监督,相当于 DRaFT 思想的最早期实现:</p>
<p>$$\mathcal{L}_\text{ReFL} = \mathcal{L}_\text{simple} - \lambda \cdot \mathbb{E}_{t' \sim [t_\text{min}, t_\text{max}]}\!\big[R\big(\hat{x}_0(x_{t'}, t')\big)\big]$$</p>
<p>其中 $\hat{x}_0(x_{t'}, t') = (x_{t'} - \sqrt{1-\bar\alpha_{t'}}\epsilon_\theta(x_{t'}, t'))/\sqrt{\bar\alpha_{t'}}$ 是从单步 $\epsilon$-prediction 估的 $x_0$Tweedie 一步 unfold)。</p>
<p><strong>关键差异</strong>ReFL 是"单步反传 + $L_\text{simple}$ 同时训"DRaFT/AlignProp 是"多步反传 + 纯 reward loss"。ReFL 训练更稳但 reward 涨幅小,因为只看了 $x_0$ 的一步估计而非真实采样轨迹。</p>
<h3 id="35-三者对比">3.5 三者对比</h3>
<table><thead><tr><th>算法</th><th>反传策略</th><th>显存</th><th>reward 涨幅</th><th>稳定性</th></tr></thead><tbody><tr><td><strong>ReFL</strong> (Xu 2023)</td><td>单步 $\hat{x}_0$ + $L_\text{simple}$ 混合</td><td>$\mathcal{O}(1)$</td><td></td><td></td></tr><tr><td><strong>DRaFT-K</strong> (Clark 2024)</td><td>最后 $K$ 步 BPTT</td><td>$\mathcal{O}(K)$</td><td></td><td>中($K$ 大易过优化)</td></tr><tr><td><strong>AlignProp</strong> (Prabhudesai 2024)</td><td>随机选 $K$ 步 BPTT</td><td>$\mathcal{O}(K)$</td><td></td><td></td></tr></tbody></table>
<div class="callout callout-info"><div class="callout-title">显存粗算</div><p>SDXL UNet 约 2.6B 参数,单 forward 在 fp16 下需要约 68 GB activation$K = 5$ 时 $\sim$3040 GB$K = T = 50$ 时 $>$300 GB<strong>只能多机分片</strong>。这就是为什么 $K=1$ 实测最常用——精度损失可忽略但工程友好。</p></div>
<h3 id="36-reward-hacking-in-direct-reward-backprop">3.6 Reward hacking in direct reward backprop</h3>
<p>直接反传比 RL 更容易 hacking,因为梯度信号"太精确":</p>
<table><thead><tr><th>现象</th><th>例子</th></tr></thead><tbody><tr><td><strong>Over-saturation</strong></td><td>HPSv2 偏好高对比度 → 训练后图像饱和度爆表</td></tr><tr><td><strong>Style monotonicity</strong></td><td>ImageReward 训练数据有偏 → 所有 prompt 输出同一风格</td></tr><tr><td><strong>Trypophobia patterns</strong></td><td>某些 RM 偏好"细密纹理"model 学到密恐图案</td></tr><tr><td><strong>Mode collapse</strong></td><td>同一 prompt 的多次采样几乎一样</td></tr></tbody></table>
<p><strong>缓解</strong>reward ensemble (HPSv2 + PickScore + ImageReward 取 mean 或 min)、KL anchor、early stop、small LoRA scale。</p>
<h2 id="4-preference-optimizationdiffusion-dpo-家族">§4 Preference OptimizationDiffusion-DPO 家族</h2>
<h3 id="41-diffusion-dpo-wallace-et-al-2024-cvpr-arxiv-231112908">4.1 Diffusion-DPO (Wallace et al. 2024 CVPR, arXiv 2311.12908)</h3>
<p><em>Diffusion Model Alignment Using Direct Preference Optimization</em>。把 LLM DPO 移植到 diffusion 的关键挑战:<strong>diffusion 的 $\log p_\theta(x_0 \mid c)$ 没有闭式</strong>,要用 ELBO 代替。</p>
<h4 id="411-推导关键步骤">4.1.1 推导(关键步骤)</h4>
<p>LLM DPO 的核心是 KL-regularized RL 最优解:</p>
<p>$$\pi^*(y \mid x) \propto \pi_\text{ref}(y \mid x) \exp\!\big(r(x, y) / \beta\big)$$</p>
<p>反解 $r = \beta \log(\pi^*/\pi_\text{ref}) + \beta \log Z$,代入 Bradley-Terry$\log Z$ 消掉。</p>
<p>对 diffusion,把"sample $y$"替换为"sample trajectory $(x_T, \dots, x_0)$",最优解形式相同但 $\log p$ 用整条 trajectory 的 likelihood</p>
<p>$$\log p_\theta(x_{0:T} \mid c) = \log p(x_T) + \sum_{t=1}^{T} \log p_\theta(x_{t-1} \mid x_t, c)$$</p>
<p>这个 trajectory log-likelihood 是<strong>可解析</strong>的(每项都是 Gaussian log-prob)。<strong>但是</strong>:训练时若每个 update 都要跑完整 trajectory,计算成本爆炸。</p>
<p><strong>Wallace et al. 的 trick:用 ELBO surrogate。</strong></p>
<p>DDPM 的 $L_\text{simple}$ 是 $\log p_\theta(x_0)$ 的 (negative) ELBO 项之一,具体地:</p>
<p>$$-\log p_\theta(x_0 \mid c) \le L_\text{simple}(x_0, c, \theta) = \mathbb{E}_{t, \epsilon}\!\left[\|\epsilon - \epsilon_\theta(x_t, t, c)\|^2\right] + \text{const}$$</p>
<p>用 $-L_\text{simple}$ 作为 $\log p_\theta(x_0 \mid c)$ 的<strong>单 sample 估计</strong>(Jensen 不等式严格意义上给的是 lower bound,但作为 DPO 的 implicit reward 数值代理可用),代入 DPO 框架:</p>
<p>$$\boxed{\;\mathcal{L}_\text{Diff-DPO}(\theta) = -\mathbb{E}_{(x_0^w, x_0^l, c, t, \epsilon)}\log\sigma\!\left(-\beta T\!\left[\|\epsilon^w - \epsilon_\theta(x_t^w, t, c)\|^2 - \|\epsilon^w - \epsilon_\text{ref}(x_t^w, t, c)\|^2 - \|\epsilon^l - \epsilon_\theta(x_t^l, t, c)\|^2 + \|\epsilon^l - \epsilon_\text{ref}(x_t^l, t, c)\|^2\right]\right)\;}$$</p>
<div class="callout callout-info"><div class="callout-title">直觉读法</div><p>sigmoid 内部是"对 $y_w$policy 比 ref 更会去噪"减去"对 $y_l$policy 比 ref 更会去噪"。如果 policy 在 $y_w$ 上更准、在 $y_l$ 上更不准,差值正,loss 下降。</p></div>
<h4 id="412-实现细节">4.1.2 实现细节</h4>
<ul><li>$(x_0^w, x_0^l)$ 来自一个 prompt $c$ 下的人类偏好对(Pick-a-Pic 数据集是主力)。</li><li>训练时每步<strong>随机采 $t \in \{1, \dots, T\}$ 和 $\epsilon \sim \mathcal{N}(0,I)$</strong>,构造 $x_t^w = \sqrt{\bar\alpha_t} x_0^w + \sqrt{1-\bar\alpha_t}\epsilon$(同样的 $\epsilon$ 用在 $y_w$ 和 $y_l$ 上,做 paired noise)。</li><li>$\pi_\text{ref}$ 是冻结的 base UNet(一般用 SDXL 原始 checkpoint)。</li></ul>
<div class="callout callout-warn"><div class="callout-title">共享 noise 的关键性</div><p>论文强调 $x_t^w$ 和 $x_t^l$ 必须用<strong>同一个 $\epsilon$</strong>(即 paired noise),否则 $\beta$ 不再可比,loss 方差爆炸。这是 Diffusion-DPO 最容易踩的坑。</p></div>
<h4 id="413-结果sdxl-上">4.1.3 结果(SDXL 上)</h4>
<ul><li>在 PickScore / HPSv2 上稳定优于 SDXL base。</li><li>训练成本约 SFT 的 1.52x(要算两次 UNetpolicy + ref)。</li><li>比 DDPO 简单:完全 offline,不需要 sample。</li></ul>
<h3 id="42-d3po-yang-et-al-2024-cvpr-arxiv-231113231">4.2 D3PO (Yang et al. 2024 CVPR, arXiv 2311.13231)</h3>
<p><em>Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model</em>。和 Diffusion-DPO 几乎同时挂 arXiv2023-11),差别在<strong>推导路径</strong></p>
<ul><li><strong>Diffusion-DPO</strong>:先 KL-regularized RL → ELBO surrogate → DPO loss。</li><li><strong>D3PO</strong>:直接把 LLM DPO 的推导<strong>逐步搬到 diffusion 的 Markov chain</strong>——每个 denoising step 看作一个 MDP step,用相同的"反解 implicit reward + 代入 BT"框架。</li></ul>
<p>D3PO 最终 loss 形式与 Diffusion-DPO 几乎相同:</p>
<p>$$\mathcal{L}_\text{D3PO}(\theta) = -\mathbb{E}_{(\tau^w, \tau^l)} \log\sigma\!\left(\beta \sum_{t=1}^{T}\!\left[\log\frac{p_\theta(x_{t-1}^w \mid x_t^w, c)}{p_\text{ref}(x_{t-1}^w \mid x_t^w, c)} - \log\frac{p_\theta(x_{t-1}^l \mid x_t^l, c)}{p_\text{ref}(x_{t-1}^l \mid x_t^l, c)}\right]\right)$$</p>
<p>需要的是<strong>完整 trajectory pair</strong> $(\tau^w, \tau^l)$;如果偏好对只有 final image $(x_0^w, x_0^l)$,需先用 $q(x_{1:T} \mid x_0)$ 重构 trajectory(用 DDPM 的 forward q-sample)。</p>
<div class="callout callout-info"><div class="callout-title">Diffusion-DPO vs D3PO 实际差异</div><p>二者数学等价(在 ELBO surrogate 下,D3PO 的 trajectory log-ratio 退化为 Diffusion-DPO 的单步 $\epsilon$ 距离差)。<strong>实践中</strong></p></div>
<ul><li>Diffusion-DPO 用 single-$t$ 估计(更省),D3PO 用全 trajectory 求和(更精但贵)。</li><li>Diffusion-DPO 在 Pick-a-Pic 上稳,D3PO 在自采集 thumbs up/down 数据上稳。</li><li>工业部署主流用 Diffusion-DPO(计算简单)。</li></ul>
<h3 id="43-spo-liang-et-al-2024-arxiv-240604314">4.3 SPO (Liang et al. 2024, arXiv 2406.04314)</h3>
<p><em>Step-aware Preference Optimization: Aligning Preference with Denoising Performance at Each Step</em><strong>关键观察</strong>:不同 denoising step <strong>对图像的不同方面负责</strong></p>
<ul><li>高噪 step ($t \approx T$):决定 global 结构(构图、物体位置)。</li><li>低噪 step ($t \approx 0$):决定 local 细节(纹理、边缘)。</li></ul>
<p>如果用 Diffusion-DPO 的"单 $t$ 采样",相当于把所有 step 同等对待——但人类偏好在不同 step 上的"重要性"不同。</p>
<h4 id="431-spo-的两个修改">4.3.1 SPO 的两个修改</h4>
<p><strong>修改 1In-step preference</strong>——对同一 $x_t$<strong>独立采两个 $x_{t-1}^w, x_{t-1}^l$</strong>,由一个 step-wise reward model 判断"哪个 $x_{t-1}$ 在 step $t$ 上更好"。</p>
<p><strong>修改 2Step-aware weighting</strong>——SPO loss 在 step 维度上加权:</p>
<p>$$\mathcal{L}_\text{SPO}(\theta) = -\mathbb{E}_{t \sim w(t),\; x_t}\!\left[\log\sigma\!\left(\beta\!\log\frac{p_\theta(x_{t-1}^w \mid x_t, c)}{p_\text{ref}(x_{t-1}^w \mid x_t, c)} - \beta\!\log\frac{p_\theta(x_{t-1}^l \mid x_t, c)}{p_\text{ref}(x_{t-1}^l \mid x_t, c)}\right)\right]$$</p>
<p>其中 $w(t)$ 是 step 采样分布(典型 uniform 或更偏向中等 $t$)。</p>
<h4 id="432-in-step-reward-model">4.3.2 In-step reward model</h4>
<p>为了得到 in-step preference $(x_{t-1}^w, x_{t-1}^l)$SPO 训了一个<strong>step-wise reward model</strong> $R(x_{t-1}, x_t, c, t)$,判断"给定 $x_t$$x_{t-1}$ 在 step $t$ 上是好的过渡吗"。它不是直接打分 $x_{t-1}$ 的像素质量,而是估计<strong>在 step $t$ 上</strong>这个过渡是否会通向高质量 $x_0$。</p>
<div class="callout callout-good"><div class="callout-title">SPO 的关键收益</div><p>同一份偏好数据,SPO 的有效信号量 ×$T$ 倍(每个 prompt 在 $T$ 个 step 上都产生 pair)。实测在 PickScore / HPSv2 上比 Diffusion-DPO 涨 13 点。</p></div>
<h3 id="44-diffusion-kto-li-et-al-2024-neurips-arxiv-240404465">4.4 Diffusion-KTO (Li et al. 2024 NeurIPS, arXiv 2404.04465)</h3>
<p><em>Aligning Diffusion Models by Optimizing Human Utility</em>。LLM KTO (Ethayarajh 2024, arXiv 2402.01306) 的 diffusion 版。</p>
<p><strong>LLM KTO 核心 idea</strong>:用 Kahneman-Tversky prospect theory 替换 BT preference model,只需要 <strong>per-sample binary feedback</strong>thumbs up/down),<strong>不需要 pair</strong></p>
<p>$$L_\text{KTO} = \mathbb{E}_{x, y}\!\left[\lambda_y v\!\big(\beta \log\frac{\pi_\theta(y|x)}{\pi_\text{ref}(y|x)} - z_0(x)\big)\right]$$</p>
<p>其中 $v(\cdot)$ 是 prospect-theoretic 价值函数(thumbs up 用 $1 - \sigma(\cdot)$thumbs down 用 $\sigma(\cdot)$),$z_0$ 是 reference utility。</p>
<p>Diffusion-KTO 把 $\log(\pi_\theta/\pi_\text{ref})$ 替换为 Diffusion-DPO 的 $\epsilon$-distance ELBO surrogate</p>
<p>$$L_\text{Diff-KTO} = \mathbb{E}_{x_0, c, \text{label}}\!\left[\lambda_\text{label}\, v\!\left(\beta T \left[\|\epsilon - \epsilon_\text{ref}\|^2 - \|\epsilon - \epsilon_\theta\|^2\right] - z_0(c)\right)\right]$$</p>
<div class="callout callout-info"><div class="callout-title">Diffusion-KTO 的实用价值</div><p>工业场景大量 binary feedback(喜欢/不喜欢)远多于 paired comparisonKTO 让这部分数据可直接用。</p></div>
<h3 id="45-mapo-hong-et-al-2024-arxiv-240606424">4.5 MaPO (Hong et al. 2024, arXiv 2406.06424)</h3>
<p><em>Margin-aware Preference Optimization for Aligning Diffusion Models without Reference</em><strong>核心 idea</strong><strong>完全去掉 reference model</strong>——类似 LLM 的 SimPO 思想。</p>
<p>MaPO loss 同时优化两件事:</p>
<ol><li><strong>Likelihood margin</strong>$\log p_\theta(x_0^w) - \log p_\theta(x_0^l)$(用 ELBO surrogate $\|\epsilon - \epsilon_\theta\|^2$ 估计)。</li><li><strong>Likelihood of preferred</strong>$\log p_\theta(x_0^w)$ 本身要高(防止"两边都降")。</li></ol>
<p>$$\mathcal{L}_\text{MaPO}(\theta) = -\mathbb{E}\!\left[\log\sigma\!\big(\beta(\hat{\ell}_w - \hat{\ell}_l) - \gamma\big) + \alpha \hat{\ell}_w\right]$$</p>
<p>其中 $\hat{\ell} = -\|\epsilon - \epsilon_\theta(x_t, t, c)\|^2$ 是 likelihood surrogate$\gamma$ 是 margin$\alpha$ 是 likelihood term 权重。</p>
<p><strong>优势</strong></p>
<ul><li><strong>不需要 ref UNet</strong>:显存省一半(从 $2 \times M$ 到 $M$)。</li><li><strong>解决 reference mismatch</strong>:当 fine-tune 到新风格(reference 与目标分布差距大)时 Diffusion-DPO 训练崩溃,MaPO 稳定。</li><li><strong>训练快 15%</strong>(论文报告,5 domains 上验证)。</li></ul>
<h3 id="46-dpo-家族总览表">4.6 DPO 家族总览表</h3>
<table><thead><tr><th>方法</th><th>需要 ref?</th><th>偏好类型</th><th>显存</th><th>适用</th></tr></thead><tbody><tr><td><strong>Diffusion-DPO</strong> (Wallace 2024)</td><td></td><td>paired</td><td>2x</td><td>一般 alignment</td></tr><tr><td><strong>D3PO</strong> (Yang 2024)</td><td></td><td>paired or thumbs</td><td>2x</td><td>没 RM 时</td></tr><tr><td><strong>SPO</strong> (Liang 2024)</td><td>✅ + step-RM</td><td>per-step paired</td><td>2x + 小 step-RM</td><td>想榨干 step 信号</td></tr><tr><td><strong>Diffusion-KTO</strong> (Li 2024)</td><td></td><td>unpaired binary</td><td>2x</td><td>大量 thumbs 数据</td></tr><tr><td><strong>MaPO</strong> (Hong 2024)</td><td></td><td>paired</td><td>1x</td><td>风格 fine-tune / 显存紧</td></tr></tbody></table>
<h2 id="5-flow-grpoflow-matching-的-rl">§5 Flow-GRPOFlow Matching 的 RL</h2>
<h3 id="51-为什么-flow-matching-也要-post-training">5.1 为什么 Flow Matching 也要 post-training</h3>
<p>SD3 / FLUX / Lumina 全部转向 Flow MatchingRectified Flow),post-training 需求一样:</p>
<ul><li>提升 GenEval / DPG 等组合性 benchmark(颜色、计数、空间关系)。</li><li>提升 OCR / 文字渲染准确率。</li><li>提升 prompt-image alignmentVLM judge)。</li></ul>
<p>但 Flow Matching 是<strong>确定性 ODE</strong>$\dot x_t = v_\theta(t, x, c)$),DDPO/DPOK 假设的 stochastic transition 不存在——直接套 RL 框架会失败。</p>
<h3 id="52-flow-grpo-的两个核心-trick">5.2 Flow-GRPO 的两个核心 trick</h3>
<p>Liu et al. 2025 <em>Flow-GRPO: Training Flow Matching Models via Online RL</em>arXiv 2505.05470)解决了 Flow + RL 的两个根本问题:</p>
<h4 id="trick-1ode--sde-等价转换">Trick 1ODE → SDE 等价转换</h4>
<p>对 Rectified Flow 的 ODE $\dot x_t = v_\theta(t, x_t, c)$,构造一个<strong>等价的 SDE</strong></p>
<p>$$dx_t = \big[v_\theta(t, x_t, c) + \tfrac{1}{2}\sigma(t)^2 \nabla_x \log p_t(x_t)\big]\,dt + \sigma(t)\,dW_t$$</p>
<p><strong>关键性质</strong>Song et al. 2021 score SDE 框架):这条 SDE 的 marginal $p_t$ 与原 ODE <strong>完全相同</strong>。区别是 SDE 提供了<strong>随机探索</strong>($dW_t$ 噪声项),让 RL 可以 sample 不同 trajectory。</p>
<p>对 Flow Matching$\nabla_x \log p_t = -\epsilon/\sigma_t$(在 Gaussian path 下),可以从 $v_\theta$ 推得 score。把 $\sigma(t)$ 设为 schedule(典型 $\sigma(t) = \sqrt{1-t}$),就得到 Flow-GRPO 训练用的 SDE sampler。</p>
<div class="callout callout-info"><div class="callout-title">物理意义</div><p>加 $\sigma\,dW$ 让粒子在 marginal 不变的前提下"抖动"出多条 trajectory,于是同一 prompt 的 $G$ 次 sample 是真正不同的 → GRPO 的组内统计可算。</p></div>
<h4 id="trick-2denoising-reduction">Trick 2Denoising reduction</h4>
<p>GRPO 需要 sample 一组 $G$ 个 trajectory$G$ 典型 1632。Flow Matching 推理一般 2550 步,<strong>训练时 sample 一次 $\approx 25 G$ 次 forward</strong>,太贵。</p>
<p>Flow-GRPO 训练时用 <strong>fewer steps</strong>(如 10 步),推理时仍用 2550 步。SDE 在 schedule 上更均匀,少步训练的"探索质量"够用。具体地:</p>
<p>$$\text{Training}: T_\text{train} = 10, \quad \text{Inference}: T_\text{infer} = 28$$</p>
<p>实测在 GenEval / OCR / Aesthetic 上不掉点。</p>
<h3 id="53-flow-grpo-的-advantage-计算">5.3 Flow-GRPO 的 advantage 计算</h3>
<p>和 GRPO for LLM 完全平行——对同一 prompt $c$ sample $G$ 个 final image $\{x_0^{(1)}, \dots, x_0^{(G)}\}$,每个打 reward $r_i$,组内归一化:</p>
<p>$$\hat{A}_i = \frac{r_i - \text{mean}_{j}(r_j)}{\text{std}_j(r_j) + \epsilon}$$</p>
<p>整条 trajectory 内所有 step 共享 $\hat{A}_i$(同 LLM GRPO 的 per-token 共享)。</p>
<h3 id="54-flow-grpo-的-loss">5.4 Flow-GRPO 的 loss</h3>
<p>记 SDE Euler step 的 transition log-prob $\log p_\theta(x_{t-1} \mid x_t, c)$Gaussian),importance ratio $\rho_{i,t} = p_\theta / p_{\theta_\text{old}}$PPO-clip</p>
<p>$$L^\text{Flow-GRPO} = \mathbb{E}\!\left[\frac{1}{G}\sum_i\!\frac{1}{T_\text{train}}\!\sum_t \min\!\big(\rho_{i,t}\hat{A}_i, \text{clip}(\rho_{i,t}, 1-\epsilon, 1+\epsilon)\hat{A}_i\big) - \beta\, \text{KL}_{i,t}(p_\theta \Vert p_\text{ref})\right]$$</p>
<p>KL 仍用 K3 estimator(同 GRPO for LLM)。</p>
<h3 id="55-vector-field-的-advantage-几何意义">5.5 vector field 的 advantage 几何意义</h3>
<div class="callout callout-good"><div class="callout-title">L3 级理解</div></div>
<ul><li>LLM GRPO 的 advantage 在 token logit 空间上做 reweight</li><li>Flow-GRPO 的 advantage 直接 reweight $v_\theta$ 的<strong>方向修正</strong>——具体地,$\hat A_i > 0$ 时把 $v_\theta(t, x_t, c)$ 推向 trajectory $\tau_i$ 实际经过的方向 $(x_{t-1}^{(i)} - x_t^{(i)})/dt$。</li><li>这是 $v_\theta$ 空间的"方向梯度",等价于在 Gaussian path 下的 $\epsilon$-prediction 重要性加权。</li></ul>
<h3 id="56-flow-grpo-实测结果">5.6 Flow-GRPO 实测结果</h3>
<p>论文报告 SD3.5-M 上:</p>
<table><thead><tr><th>Benchmark</th><th>SD3.5-M base</th><th>Flow-GRPO</th></tr></thead><tbody><tr><td>GenEval overall</td><td>63%</td><td><strong>95%</strong></td></tr><tr><td>Visual text rendering</td><td>59%</td><td><strong>92%</strong></td></tr><tr><td>Aesthetic (Schuhmann)</td><td>5.8</td><td>6.1</td></tr></tbody></table>
<div class="callout callout-warn"><div class="callout-title">GenEval 95% 看起来过于完美</div><p>论文确实主张这个数字,但需注意 GenEval 测的是规则可验证的对象计数/颜色/空间关系,本身就是 RL 友好任务(reward 极规则化)。在更主观的 PartiPrompt / DPG 上涨幅是 510 点,更现实。</p></div>
<h2 id="6-code-patterns可读伪代码">§6 Code Patterns(可读伪代码)</h2>
<h3 id="61-ddpo-reinforce-style-update">6.1 DDPO REINFORCE-style update</h3>
<pre><code class="language-python">import torch
import torch.nn.functional as F
def ddpo_step(unet, ref_unet, scheduler, prompts, reward_fn,
T=20, B=4, lr=1e-5, beta=0.0):
&quot;&quot;&quot;
DDPO-SF 一步训练(REINFORCE + 可选 KL anchor)。
prompts: list of B text prompts
reward_fn: callable (x0_batch, prompts) -&gt; [B] scalar
&quot;&quot;&quot;
# ── 1. Rollout: sample G=B trajectories ──
x = torch.randn(B, 4, 64, 64, device=device)
traj_log_probs = []
with torch.set_grad_enabled(False): # rollout 不需要梯度
x_t = x
for t in reversed(range(T)):
# predict noise + sample x_{t-1}
eps_pred = unet(x_t, t, prompts)
mean, std = scheduler.step_mean_std(x_t, eps_pred, t)
x_tm1 = mean + std * torch.randn_like(mean) # stochastic transition
traj_log_probs.append((mean.detach(), std.detach(), x_tm1.detach()))
x_t = x_tm1
x_0 = x_t
# ── 2. Reward ──
R = reward_fn(x_0, prompts) # [B]
A = (R - R.mean()) / (R.std() + 1e-8) # batch baseline
# ── 3. Policy gradient: 重新 forward 取 log p_θ ──
x_t = x.detach()
loss_pg = 0.0
for t, (mean_old, std_old, x_tm1) in zip(reversed(range(T)), traj_log_probs):
eps_pred = unet(x_t, t, prompts) # 要梯度
mean, std = scheduler.step_mean_std(x_t, eps_pred, t)
# Gaussian log-prob
log_p = -0.5 * (((x_tm1 - mean) / std) ** 2).sum([1, 2, 3])
log_p -= std.log().sum([1, 2, 3])
loss_pg = loss_pg - (log_p * A).mean() # REINFORCE
if beta &gt; 0:
ref_eps = ref_unet(x_t, t, prompts).detach()
mean_ref, std_ref = scheduler.step_mean_std(x_t, ref_eps, t)
# Gaussian-Gaussian KL closed form
kl = ((mean - mean_ref) ** 2 / (2 * std_ref ** 2)
+ (std / std_ref) ** 2 / 2
- 0.5 - (std / std_ref).log()).sum([1, 2, 3])
loss_pg = loss_pg + beta * kl.mean()
x_t = x_tm1.detach()
return loss_pg</code></pre>
<div class="callout callout-warn"><div class="callout-title">DDPO 实现踩坑</div></div>
<ul><li>Rollout 用 <code>set_grad_enabled(False)</code>policy gradient pass 再开梯度,避免显存 $O(T)$。</li><li>DDPM stochastic transition 是关键:DDIM 是 deterministic,没有 $dW$ 维度可优化,<strong>DDPO 必须用 DDPM 或 DDIM-eta=1</strong></li><li>Batch baseline $A = (R - \bar R)/\sigma_R$ 比无 baseline 稳得多。</li><li>LoRA 训而非 full fine-tune,否则 base 漂移很快。</li></ul>
<h3 id="62-diffusion-dpo-loss">6.2 Diffusion-DPO loss</h3>
<pre><code class="language-python">def diffusion_dpo_loss(unet, ref_unet, scheduler,
x0_w, x0_l, prompt_embeds, beta=2000.0):
&quot;&quot;&quot;
Diffusion-DPO (Wallace 2024) 单步训练。
x0_w, x0_l: [B, 4, H, W] preferred / dispreferred latents
beta: 论文用 2000~5000(注意是 β·T 的合并系数,比 LLM DPO 大)
&quot;&quot;&quot;
B = x0_w.shape[0]
t = torch.randint(0, scheduler.num_train_timesteps, (B,), device=x0_w.device)
noise = torch.randn_like(x0_w) # paired noise!
xt_w = scheduler.add_noise(x0_w, noise, t)
xt_l = scheduler.add_noise(x0_l, noise, t)
# ── policy ε-prediction ──
eps_w = unet(xt_w, t, prompt_embeds)
eps_l = unet(xt_l, t, prompt_embeds)
# ── reference ε-prediction (frozen) ──
with torch.no_grad():
ref_eps_w = ref_unet(xt_w, t, prompt_embeds)
ref_eps_l = ref_unet(xt_l, t, prompt_embeds)
# ── ELBO surrogate: -‖ε - ε_θ‖² 是 log p_θ 的代理 ──
err_w_pol = ((noise - eps_w) ** 2).mean([1, 2, 3]) # [B]
err_w_ref = ((noise - ref_eps_w) ** 2).mean([1, 2, 3])
err_l_pol = ((noise - eps_l) ** 2).mean([1, 2, 3])
err_l_ref = ((noise - ref_eps_l) ** 2).mean([1, 2, 3])
# DPO log-ratio: smaller err = better likelihood
# log(π_θ/π_ref)(y_w) ≈ -(err_w_pol - err_w_ref)
diff_w = -(err_w_pol - err_w_ref)
diff_l = -(err_l_pol - err_l_ref)
inner = beta * (diff_w - diff_l)
loss = -F.logsigmoid(inner).mean()
with torch.no_grad():
margin = inner.mean()
accuracy = (inner &gt; 0).float().mean()
return loss, {&quot;margin&quot;: margin.item(), &quot;acc&quot;: accuracy.item()}</code></pre>
<div class="callout callout-warn"><div class="callout-title">β 量级注意</div><p>Diffusion-DPO 的 $\beta$ 比 LLM DPO 大几个数量级,因为它吸收了 $T$ 倍的累积项($\beta T$ 才是真正的"温度")。论文用 $\beta \in [2000, 5000]$LLM DPO 用 $\beta \in [0.05, 0.5]$。</p></div>
<h3 id="63-alignprop--draft-反传-with-checkpointing">6.3 AlignProp / DRaFT 反传 with checkpointing</h3>
<pre><code class="language-python">def alignprop_step(unet_lora, scheduler, prompts, reward_fn,
T=50, K=1, B=4, lr=1e-5):
&quot;&quot;&quot;
DRaFT-K / AlignProp 一步训练。
K: 最后 K 步保留梯度
显存 O(K)K=1 时与 SDXL 单步 forward 同量级
&quot;&quot;&quot;
x = torch.randn(B, 4, 64, 64, device=device)
# ── 前 T - K 步无梯度 ──
with torch.no_grad():
for t in reversed(range(K, T)):
eps = unet_lora(x, t, prompts)
x = scheduler.step_ddim(x, eps, t) # deterministic DDIM
# ── 最后 K 步要梯度 ──
for t in reversed(range(K)):
eps = unet_lora(x, t, prompts) # gradient ON
x = scheduler.step_ddim(x, eps, t)
x_0 = x
# ── 反传 ──
reward = reward_fn(x_0, prompts) # [B]
loss = -reward.mean() # 最大化 reward = 最小化 -reward
return loss
# 显存分析:
# K=1: ~24 GB on SDXL (single forward + grad)
# K=5: ~60 GB
# K=10: ~120 GB (需要 multi-GPU)
# K=T=50: ~600 GB (完全不可行)</code></pre>
<div class="callout callout-info"><div class="callout-title">K=1 已经够用?</div><p>是的。直觉:最后一步 $x_1 \to x_0$ 对 $x_0$ 的影响最大(前面 49 步的方差被压缩),所以反传 1 步的信号已经主导。Clark 2024 也实测 $K=1$ 与 $K=5$ 差距极小。</p></div>
<h3 id="64-spo-step-aware-preference-loss">6.4 SPO step-aware preference loss</h3>
<pre><code class="language-python">def spo_loss(unet, ref_unet, scheduler, step_rm,
x_t, t, prompt_embeds, beta=500.0):
&quot;&quot;&quot;
SPO (Liang 2024) in-step preference.
给定 x_t 和 t,独立采两个 x_{t-1},让 step_rm 判断 winner.
&quot;&quot;&quot;
# ── 1. 用 policy 采两个 x_{t-1} candidate(采样过程必须 no_grad,否则后续 DPO log-prob 会传梯度回到 sample)──
with torch.no_grad():
eps_sample = unet(x_t, t, prompt_embeds)
mean_s, std_s = scheduler.step_mean_std(x_t, eps_sample, t)
noise_a, noise_b = torch.randn_like(mean_s), torch.randn_like(mean_s)
x_a = mean_s + std_s * noise_a
x_b = mean_s + std_s * noise_b
# ── 2. step-wise reward model 判断 winner ──
with torch.no_grad():
r_a = step_rm(x_a, x_t, t, prompt_embeds) # [B]
r_b = step_rm(x_b, x_t, t, prompt_embeds)
winner = (r_a &gt; r_b).long() # [B], 1 if a wins
x_w = torch.where(winner.bool()[:, None, None, None], x_a, x_b).detach()
x_l = torch.where(winner.bool()[:, None, None, None], x_b, x_a).detach()
# ── 3. compute log p_θ / log p_ref 对 x_w, x_lgrad-aware forward)──
eps = unet(x_t, t, prompt_embeds)
mean, std = scheduler.step_mean_std(x_t, eps, t)
log_p_w = -0.5 * ((x_w - mean) / std).pow(2).sum([1, 2, 3])
log_p_l = -0.5 * ((x_l - mean) / std).pow(2).sum([1, 2, 3])
with torch.no_grad():
ref_eps = ref_unet(x_t, t, prompt_embeds)
ref_mean, ref_std = scheduler.step_mean_std(x_t, ref_eps, t)
log_pref_w = -0.5 * ((x_w - ref_mean) / ref_std).pow(2).sum([1, 2, 3])
log_pref_l = -0.5 * ((x_l - ref_mean) / ref_std).pow(2).sum([1, 2, 3])
inner = beta * ((log_p_w - log_pref_w) - (log_p_l - log_pref_l))
return -F.logsigmoid(inner).mean()</code></pre>
<h3 id="65-flow-grpo-group-relative-advantage">6.5 Flow-GRPO group-relative advantage</h3>
<pre><code class="language-python">def flow_grpo_step(flow_net, ref_flow, prompts, reward_fn,
G=16, T_train=10, sigma_fn=lambda t: (1 - t) ** 0.5,
eps_clip=0.2, beta=0.04):
&quot;&quot;&quot;
Flow-GRPO 一步训练。
G: 每个 prompt sample G 个 trajectory.
T_train: 训练用 SDE 步数(推理时另外用 28-50 步).
&quot;&quot;&quot;
P = len(prompts)
# 每个 prompt 重复 G 次
prompts_rep = sum([[p] * G for p in prompts], []) # [P*G]
# ── 1. SDE rollout: ODE→SDE 等价转换 ──
x_t = torch.randn(P * G, 4, 64, 64, device=device)
log_probs_old = [] # for PPO importance ratio
trajectory = [x_t.clone()]
with torch.no_grad():
for i in range(T_train):
t_now = 1.0 - i / T_train
t_next = 1.0 - (i + 1) / T_train
dt = t_next - t_now
sigma = sigma_fn(t_now)
v = flow_net(x_t, t_now, prompts_rep)
# SDE Euler: drift = v + 0.5 σ² ∇log p (PF-ODE → SDE 转换, Song 2021)
# !!! 重要:以下 drift 是简化教学版(placeholder),生产实现要按 Flow-GRPO 论文 Eq.(6)
# 正确地从 score = (data_pred - x_t)/σ_t² 推导,包含具体 Rectified Flow / EDM schedule.
# 真实部署请参考论文 + 官方 repo;此处 -v/σ 仅作 illustrative.
drift = v + 0.5 * sigma ** 2 * (-v / (sigma + 1e-6)) # placeholder, see paper Eq.(6)
noise = torch.randn_like(x_t)
x_next = x_t + drift * dt + sigma * noise * abs(dt) ** 0.5
# Gaussian log-prob (transition)
mean = x_t + drift * dt
std = sigma * abs(dt) ** 0.5
log_p = -0.5 * ((x_next - mean) / std).pow(2).sum([1, 2, 3])
log_probs_old.append(log_p)
x_t = x_next
trajectory.append(x_t.clone())
x_0 = x_t
# ── 2. Group-relative advantage ──
R = reward_fn(x_0, prompts_rep) # [P*G]
R = R.view(P, G)
mean_R = R.mean(dim=1, keepdim=True)
std_R = R.std(dim=1, keepdim=True) + 1e-8
A = ((R - mean_R) / std_R).view(P * G) # [P*G]
# ── 3. PPO-clip loss with KL ──
loss = 0.0
x_t = trajectory[0]
for i in range(T_train):
t_now = 1.0 - i / T_train
v = flow_net(x_t, t_now, prompts_rep) # grad ON
sigma = sigma_fn(t_now)
drift = v + 0.5 * sigma ** 2 * (-v / (sigma + 1e-6))
dt = -1.0 / T_train
mean = x_t + drift * dt
std = sigma * abs(dt) ** 0.5
log_p_new = -0.5 * ((trajectory[i+1] - mean) / std).pow(2).sum([1, 2, 3])
ratio = (log_p_new - log_probs_old[i]).exp()
surr1 = ratio * A
surr2 = ratio.clamp(1 - eps_clip, 1 + eps_clip) * A
loss = loss - torch.min(surr1, surr2).mean()
# K3 KL estimator
with torch.no_grad():
v_ref = ref_flow(x_t, t_now, prompts_rep)
drift_ref = v_ref + 0.5 * sigma ** 2 * (-v_ref / (sigma + 1e-6))
mean_ref = x_t + drift_ref * dt
log_p_ref = -0.5 * ((trajectory[i+1] - mean_ref) / std).pow(2).sum([1,2,3])
delta = log_p_ref - log_p_new
kl_k3 = (delta.exp() - delta - 1)
loss = loss + beta * kl_k3.mean()
x_t = trajectory[i + 1].detach()
return loss</code></pre>
<h3 id="66-combined-reward-signal">6.6 Combined reward signal</h3>
<pre><code class="language-python">def combined_reward(images, prompts, weights=None):
&quot;&quot;&quot;
多 reward 加权组合 — 缓解单 RM hacking.
&quot;&quot;&quot;
weights = weights or {&quot;image_reward&quot;: 0.4, &quot;hps_v2&quot;: 0.3,
&quot;pickscore&quot;: 0.2, &quot;clip_score&quot;: 0.1}
rewards = {}
rewards[&quot;image_reward&quot;] = image_reward_model(images, prompts) # [-1, 4]
rewards[&quot;hps_v2&quot;] = hps_v2(images, prompts) # [0, 1]
rewards[&quot;pickscore&quot;] = pickscore(images, prompts) # logits
rewards[&quot;clip_score&quot;] = clip_cosine(images, prompts) # [-1, 1]
# ── 各自 z-score 归一化(不同 reward scale 差异巨大)──
normed = {k: (v - v.mean()) / (v.std() + 1e-8) for k, v in rewards.items()}
# ── 加权 + length / safety penalty ──
R = sum(weights[k] * normed[k] for k in weights)
# NSFW penalty (rule-based)
nsfw_score = nsfw_detector(images) # [0, 1]
R = R - 5.0 * nsfw_score
return R</code></pre>
<div class="callout callout-warn"><div class="callout-title">多 reward 实操经验</div></div>
<ul><li><strong>每个 reward 单独 z-score</strong>:尺度差异巨大(HPSv2 ~0.25, ImageReward ~1.5, CLIP-Score ~0.3),不归一化等于让 ImageReward 主导。</li><li><strong>min 比 mean 更稳</strong><code>R = min(normed.values())</code> 能强制所有 RM 都满意,hacking 风险显著降低(reward ensemble 经典策略)。</li><li><strong>保留 rule-based safety override</strong>NSFW / 政治敏感 / 版权 reward 不能被 RL 优化掉。</li></ul>
<h2 id="7-reward-design--失败模式">§7 Reward Design &amp; 失败模式</h2>
<h3 id="71-reward-model-选择">7.1 Reward model 选择</h3>
<table><thead><tr><th>RM</th><th>来源</th><th>数据</th><th>scale</th><th>偏好特点</th></tr></thead><tbody><tr><td><strong>CLIP-Score</strong></td><td>OpenAI/LAION</td><td>4B image-text pair</td><td>$[-1, 1]$ cosine</td><td>text-image alignment 弱信号;倾向 caption 字面匹配</td></tr><tr><td><strong>ImageReward</strong> (Xu 2023 NeurIPS)</td><td>137K human pair</td><td>真实 prompt</td><td>$[-1, 4]$</td><td>aesthetic + alignment 综合;偏好高对比度</td></tr><tr><td><strong>HPSv2</strong> (Wu 2023 arXiv 2306.09341)</td><td>798K human pair</td><td>DiffusionDB-style</td><td>$[0, 1]$</td><td>综合 human preference;偏好饱和颜色</td></tr><tr><td><strong>PickScore</strong> (Kirstain 2023 NeurIPS)</td><td>Pick-a-Pic 1M pair</td><td>真实用户</td><td>logits</td><td>综合;倾向 trained-on-SDXL 风格</td></tr><tr><td><strong>PiCaR</strong> (rule)</td><td>OpenAI</td><td>counting/OCR</td><td>binary</td><td>rule-based, 不可 hack</td></tr></tbody></table>
<h3 id="72-reward-hacking-gallerydiffusion-特色">7.2 Reward hacking gallerydiffusion 特色)</h3>
<table><thead><tr><th>现象</th><th>视觉特征</th><th>原因</th></tr></thead><tbody><tr><td><strong>Over-saturation</strong></td><td>颜色饱和度 &gt;100%</td><td>HPSv2 / aesthetic 偏好鲜艳</td></tr><tr><td><strong>Center bias</strong></td><td>主体永远居中</td><td>RM 训练数据多为 centered subject</td></tr><tr><td><strong>Monotone composition</strong></td><td>不同 prompt 都用同一构图</td><td>mode collapse to RM 高分 mode</td></tr><tr><td><strong>Tryphobia-like patterns</strong></td><td>密集点状/孔状纹理</td><td>某些 RM 偏好"texture richness"</td></tr><tr><td><strong>Watermark hallucination</strong></td><td>角落出现 fake watermark</td><td>RM 训练数据含水印 → 学到"水印 = 真照片"</td></tr><tr><td><strong>Cartoon shift</strong></td><td>真实风格 prompt 输出动漫</td><td>RM 标注者偏好 anime</td></tr><tr><td><strong>Lighting overcooked</strong></td><td>后期 HDR 过强</td><td>aesthetic predictor 偏好后期重</td></tr></tbody></table>
<h3 id="73-step-level-vs-trajectory-level-reward">7.3 Step-level vs trajectory-level reward</h3>
<table><thead><tr><th>维度</th><th>Trajectory-level</th><th>Step-level</th></tr></thead><tbody><tr><td>信号位置</td><td>只在 $x_0$</td><td>每个 $t$ 都有</td></tr><tr><td>数据获取</td><td>易(一张图)</td><td>难(需要 step-wise RM 或 rollout</td></tr><tr><td>学习效率</td><td>低(稀疏)</td><td>高(dense</td></tr><tr><td>代表</td><td>DDPO, Diffusion-DPO</td><td>SPO (Liang 2024)</td></tr><tr><td>工程难度</td><td></td><td>高(要么训 step-RM,要么 PRM-shepherd 式 rollout</td></tr></tbody></table>
<div class="callout callout-info"><div class="callout-title">step-RM 怎么训</div><p>SPO 用一个 "given $x_t$ at step $t$, is $x_{t-1}$ a good transition?" 的 binary RM。训练数据:从 base UNet 跑多条 trajectory,用最终 $x_0$ 的 reward 反推每步的 step-reward (类似 Math-Shepherd 的 rollout-based PRM)。</p></div>
<h3 id="74-缓解-reward-hacking-的核心机制">7.4 缓解 reward hacking 的核心机制</h3>
<ol><li><strong>Reward ensemble</strong>:多 RM 取 min 或 meanHPSv2 + PickScore + ImageReward 是主流组合)。</li><li><strong>KL anchor</strong>DPO 的 $\beta$、DPOK 的显式 KL、Flow-GRPO 的 K3 KL term。</li><li><strong>LoRA scale</strong>full fine-tune 漂移快,LoRA scale 限制 reward hacking 上限。</li><li><strong>Early stop on reward plateau</strong>reward 涨 + FID 涨 = hacking 信号。</li><li><strong>Composite reward</strong>rule-based (object count, OCR) + neural RM (aesthetic, alignment) 加权。</li><li><strong>Adversarial RM</strong>:训 RM 时加 hacking 样本作为 negative。</li></ol>
<h2 id="8-production-landscapesd3--flux-用了什么">§8 Production LandscapeSD3 / FLUX 用了什么</h2>
<h3 id="81-公开论文--报告说了什么">8.1 公开论文 / 报告说了什么</h3>
<table><thead><tr><th>模型</th><th>Post-training?</th><th>公开内容</th></tr></thead><tbody><tr><td><strong>SD 1.5</strong></td><td>部分社区 DPO / DDPO LoRA</td><td>base 是纯 LDM;社区 fine-tune 多</td></tr><tr><td><strong>SDXL</strong></td><td>Stability AI 没明确 post-training</td><td>base + refiner; Pick-a-Pic + Diffusion-DPO 社区 LoRA 流行</td></tr><tr><td><strong>SDXL Turbo / ADD</strong></td><td>蒸馏为主</td><td>Adversarial Diffusion Distillation (2311.17042);本质是 1-step distill,不属 RL post-training</td></tr><tr><td><strong>SD3</strong> (Stable Diffusion 3, Esser et al. 2024 ICML)</td><td>base 用 Rectified Flow + MM-DiT</td><td>论文未公开 post-training;社区猜测有内部 DPO</td></tr><tr><td><strong>SD3.5 / SD3.5 Turbo</strong></td><td>有 distillpost-training 未公开</td><td>推测有 DPO + distill 混合</td></tr><tr><td><strong>FLUX.1 dev / pro</strong> (Black Forest Labs 2024)</td><td>未公开</td><td>社区猜测 DPO + distillpro 走 API 闭源</td></tr><tr><td><strong>DALL-E 3</strong> (OpenAI 2023)</td><td>"recaptioning + RLHF"</td><td>公开报告强调 prompt-faithful RLHF</td></tr><tr><td><strong>Imagen 3</strong> (Google 2024)</td><td>未公开</td><td>内部 alignment 流程</td></tr><tr><td><strong>DeepFloyd IF</strong></td><td>无 post-training</td><td>学术 base model</td></tr></tbody></table>
<h3 id="82-sd3--flux-是否用了-post-training">8.2 SD3 / FLUX 是否用了 post-training</h3>
<div class="callout callout-warn"><div class="callout-title">诚实回答</div><p>公开论文 / 技术报告<strong>都没有明说</strong>用 RL / DPO post-training。但有以下间接证据:</p></div>
<ul><li>SD3 论文 (arXiv 2403.03206) 的 "Improving Rectified Flow Transformers" 章节讨论 sampling + reflow,没提 reward fine-tune。</li><li>FLUX 完全没发论文,社区从 model card 推测有 distillationFLUX schnell 是 4-step 蒸馏版)。</li><li>Stability AI 在 SD3.5-Large 发布时提到 "fine-tuned with improved aesthetics",可能是 SFT 而非 RL。</li><li>DALL-E 3 论文 (OpenAI 2023) 明确说用了 caption-faithful RLHF。</li></ul>
<p><strong>业界共识</strong>(来自 HuggingFace 社区 + Reddit r/StableDiffusion):闭源大模型(FLUX pro, DALL-E 3, Midjourney v6+)有 reward-based fine-tune,但具体方法不公开;开源 baseSD3.5 base, FLUX dev base)公开训练 pipeline 不含 RL,但 Stability AI 内部 dev 版可能有。</p>
<h3 id="83-工业级-pipeline-假说">8.3 工业级 pipeline 假说</h3>
<pre class="diagram"><code> ┌─────────────────┐
│ LDM Pretrain │ 几百 M / B images, $L_simple$
└────────┬────────┘
┌────────▼────────┐
│ SFT on Curated │ 高质量 prompt-image pair
│ dataset │ (Aesthetic &gt; 6.0, no watermark)
└────────┬────────┘
┌────────────────┼────────────────┐
│ │
┌────────▼────────┐ ┌─────────▼────────┐
│ Diffusion-DPO │ │ DRaFT / AlignProp│
│ on Pick-a-Pic │ │ on multi-RM │
└────────┬────────┘ └─────────┬────────┘
│ │
└────────────────┬─────────────────┘
┌────────▼────────┐
│ Distillation │ 4-step / 1-step turbo
│ (ADD / LCM) │
└────────┬────────┘
Production</code></pre>
<div class="callout callout-info"><div class="callout-title">结论</div><p>Post-training 大概率发生在"SFT → distill"之间;具体方法工业界倾向 Diffusion-DPOoffline、稳定、不需要 sample),DDPO/DPOK 学术影响大但工程部署少。</p></div>
<h2 id="9-vs-llm-rlhf-对比">§9 vs LLM RLHF 对比</h2>
<h3 id="91-一表看完">9.1 一表看完</h3>
<table><thead><tr><th>维度</th><th>LLM RLHF (RLHF + DPO + GRPO)</th><th>Diffusion Post-Training</th></tr></thead><tbody><tr><td><strong>轨迹</strong></td><td>$L$ 个 token</td><td>$T$ 个 denoising step</td></tr><tr><td><strong>动作空间</strong></td><td>离散 vocab</td><td>连续 $\mathbb{R}^d$</td></tr><tr><td><strong>Reward 来源</strong></td><td>trained BT-RM / rule (math, code)</td><td>trained image RM (HPSv2/PickScore/ImageReward) / rule (count, OCR)</td></tr><tr><td><strong>Reward 稀疏度</strong></td><td>terminal only (response 末尾)</td><td>terminal only ($x_0$)</td></tr><tr><td><strong>On-policy 成本</strong></td><td>$L$ 次 forward</td><td>$T$ 次 forward + image RM forward</td></tr><tr><td><strong>Offline 方法</strong></td><td>DPO / IPO / KTO / SimPO / ORPO</td><td>Diffusion-DPO / D3PO / SPO / KTO / MaPO</td></tr><tr><td><strong>On-policy 方法</strong></td><td>PPO / GRPO / RLOO</td><td>DDPO / DPOK / Flow-GRPO</td></tr><tr><td><strong>直接 reward 反传</strong></td><td>❌(token 不可导)</td><td>✅(DRaFT / AlignProp / ReFL</td></tr><tr><td><strong>显存瓶颈</strong></td><td>4 副本 (policy + ref + RM + V)</td><td>1 副本(DPO/ $O(K)$DRaFT-K/ $O(T)$vanilla backprop</td></tr><tr><td><strong>典型 $\beta$</strong></td><td>$0.05 \sim 0.5$</td><td>$2000 \sim 5000$(吸收 $T$ 倍系数)</td></tr><tr><td><strong>典型 trajectory length</strong></td><td>$L \sim 10^3$ token</td><td>$T \sim 20$$50$ step</td></tr><tr><td><strong>Mode collapse 严重度</strong></td><td>中(vocab 大)</td><td><strong></strong>(连续空间易陷局部 mode</td></tr><tr><td><strong>Reward hacking 难度</strong></td><td>中(依赖 RM 质量)</td><td><strong></strong>(视觉 RM 比 BT-RM 更易被 hack</td></tr></tbody></table>
<h3 id="92-共同-lesson">9.2 共同 lesson</h3>
<ol><li><strong>KL anchor 是必须的</strong>:无 ref policy 的纯 RL 一定 reward hackLLM 是变长无内容,diffusion 是过饱和单一构图)。</li><li><strong>DPO 家族 &gt;&gt; on-policy RL(在工程性上)</strong>no sampling, no value model, offline——LLM 和 diffusion 都成立。</li><li><strong>Reward ensemble 反 hacking</strong>min-of-K RMs 是两个领域通用的缓解。</li><li><strong>Group-based advantage</strong>LLM 的 GRPO/RLOO 和 diffusion 的 Flow-GRPO 都通过组内统计绕过 value model。</li></ol>
<h3 id="93-独有差异">9.3 独有差异</h3>
<ul><li><strong>Diffusion 有"反传"选项</strong>reward 可导让 DRaFT/AlignProp 成立——LLM 因为 sampling 是离散的没有对应方法。</li><li><strong>Diffusion 有 "step-aware" preference</strong>denoising step 有明确语义(高噪管构图、低噪管细节),SPO 利用了这点;LLM token 没有这种自然分层。</li><li><strong>Diffusion 的 $\beta$ scale 大 1000x</strong>:因为 ELBO surrogate 吸收了 $T$ 倍 trajectory term。</li></ul>
<h2 id="10-25-高频面试题">§10 25 高频面试题</h2>
<p>按难度分 3 档:L1 = 多模态/diffusion 岗常问;L2 = research/alignment 方向会问;L3 = 顶级 lab 的硬核题。</p>
<h3 id="l1-必会题10-题">L1 必会题(10 题)</h3>
<details>
<summary>Q1. 为什么 diffusion 模型需要 post-trainingSFT 不够吗?</summary>
<ul><li>SFT 只能模仿正例("做得好的样子"),学不到<strong>对比信号</strong>A 比 B 好)。</li><li>Post-training 通过 reward / preference 提供对比信号,让模型在 alignment、aesthetic、prompt-faithful 维度都涨。</li><li>实测 Diffusion-DPO 在 PickScore 上 +5-10 点,远超继续 SFT。</li></ul>
<p>只说"提升画质"是浅;要说清"对比信号 vs 模仿信号"的差异。</p>
</details>
<details>
<summary>Q2. DDPO 把 diffusion 当成什么 MDPstate/action/reward 怎么定义?</summary>
<ul><li><strong>State</strong>: $s_t = (x_t, t, c)$</li><li><strong>Action</strong>: $a_t = x_{t-1}$(从 $p_\theta(\cdot \mid x_t, c)$ 采)</li><li><strong>Transition</strong>: 确定性 $s_{t-1} = (x_{t-1}, t-1, c)$</li><li><strong>Reward</strong>: $r_t = 0$ for $t > 1$$r_1 = R(x_0, c)$terminal-only</li></ul>
<p>说成"per-step reward"(错,只在终态);或不知道 transition 是确定性的(noise 是 action 自带的)。</p>
</details>
<details>
<summary>Q3. Diffusion-DPO 用什么代替 $\log \pi_\theta(y|x)$</summary>
<ul><li>用 ELBO surrogate$-\|\epsilon - \epsilon_\theta(x_t, t, c)\|^2$DDPM 的 $L_\text{simple}$)作为 $\log p_\theta(x_0)$ 的代理。</li><li>这是 $\log p_\theta$ 的(negative)下界项,方向正确。</li><li>配上 paired noise $\epsilon$$y_w$ 和 $y_l$ 共享同一 $\epsilon$)才稳。</li></ul>
<p>说用 $\log p(x_0)$ 解析式(错,diffusion 没闭式);忘记 paired noise。</p>
</details>
<details>
<summary>Q4. AlignProp 和 DRaFT 的核心 idea?显存为什么是 $\mathcal{O}(K)$</summary>
<ul><li><strong>核心</strong>reward $R(x_0)$ 可导 → 直接对 $\theta$ 反传,跳过 RL。</li><li>$T$ 步 sampler 是 differentiable computation graph<strong>vanilla 反传需存 $T$ 步 activation</strong> → $\mathcal{O}(T)$。</li><li><strong>DRaFT-K / AlignProp</strong>:前 $T-K$ 步用 <code>no_grad</code>,只在最后 $K$ 步保留梯度,显存压到 $\mathcal{O}(K)$。</li><li>典型 $K=1$ 已能给强信号。</li></ul>
<p>说 K=1 等于 REINFORCE(错,K=1 是 reparameterized gradient,方差远低于 REINFORCE)。</p>
</details>
<details>
<summary>Q5. 为什么 Diffusion-DPO 的 $\beta$ 比 LLM DPO 大 1000 倍?</summary>
<ul><li>LLM DPO: $\beta \in [0.05, 0.5]$</li><li>Diffusion-DPO: $\beta \in [2000, 5000]$</li><li>原因:diffusion 的"trajectory log-likelihood"是 $T$ 个 Gaussian log-prob 之和,单步 $\epsilon$ 距离差吸收了 $T$ 倍系数。<strong>实际有效温度</strong>是 $\beta T$。</li><li>也有 implementation 把 $T$ 显式分离,那时 $\beta$ 看起来与 LLM 同量级。</li></ul>
<p>说"diffusion 噪声大所以 β 大"(错,是 trajectory 长度的累计效应)。</p>
</details>
<details>
<summary>Q6. ImageReward / HPSv2 / PickScore 三者区别?</summary>
<table><thead><tr><th></th><th>ImageReward</th><th>HPSv2</th><th>PickScore</th></tr></thead><tbody><tr><td>数据规模</td><td>137K pair</td><td>798K pair</td><td>1M pair (Pick-a-Pic)</td></tr><tr><td>backbone</td><td>BLIP fine-tuned</td><td>CLIP fine-tuned</td><td>CLIP fine-tuned</td></tr><tr><td>scale</td><td>$[-1, 4]$</td><td>$[0, 1]$</td><td>logits</td></tr><tr><td>偏好</td><td>aesthetic + alignment</td><td>高对比度 + alignment</td><td>SDXL 风格</td></tr></tbody></table>
<p>工业上<strong>取 ensemble</strong>(最少两个)。</p>
<p>说三者都一样(错,scale 和偏好差异大)。</p>
</details>
<details>
<summary>Q7. DDPO 用 DDPM 还是 DDIM 采样?为什么?</summary>
<ul><li><strong>DDPM</strong>(或 DDIM-eta=1)—— 需要 stochastic transition。</li><li>DDIM-eta=0 是 deterministic,没有 noise term<strong>没有 action 可优化</strong> → policy gradient 等于 0。</li><li>类比:LLM RL 必须用 samplingtemperature &gt; 0),不能用 greedy.</li></ul>
<p>说 DDIM 也行(错,要 eta &gt; 0);不知道 stochasticity 是 RL 前提。</p>
</details>
<details>
<summary>Q8. Reward hacking 在 diffusion 上典型症状有哪些?</summary>
<ul><li><strong>Over-saturation</strong>(颜色饱和度暴增)—— HPSv2/aesthetic 偏好鲜艳。</li><li><strong>Center bias</strong>(主体永远居中)—— RM 训练数据偏 centered。</li><li><strong>Monotone composition</strong>(不同 prompt 同一构图)—— mode collapse。</li><li><strong>Watermark hallucination</strong>(角落 fake 水印)—— RM 训练数据含水印。</li><li><strong>Cartoon shift</strong>(写实 prompt 输出 anime)—— RM 标注者偏好。</li></ul>
<p>只说"过度优化"不具体;要能举至少 3 种具体视觉症状。</p>
</details>
<details>
<summary>Q9. Flow-GRPO 的 ODE→SDE 转换为什么必要?</summary>
<ul><li>Flow Matching 的 ODE $\dot x = v_\theta$ 是<strong>确定性</strong>的,给定 $x_T$ → $x_0$ 唯一。</li><li>RL 需要 stochastic policy 来 exploreODE 没有 sampling 维度。</li><li>ODE→SDE 加 $\sigma\, dW$ 噪声项,<strong>marginal $p_t$ 不变</strong>Anderson 1982),但每次 sample 路径不同 → 可 explore。</li></ul>
<p>不知道 marginal 不变(错,会以为 SDE 改变了 distribution)。</p>
</details>
<details>
<summary>Q10. Diffusion-DPO 训练时 $y_w$ 和 $y_l$ 的 noise 怎么处理?</summary>
<ul><li><strong>paired noise</strong>$x_t^w$ 和 $x_t^l$ 用<strong>同一个</strong> $\epsilon$(即 $x_t^w = \sqrt{\bar\alpha_t}x_0^w + \sqrt{1-\bar\alpha_t}\epsilon$$x_t^l$ 同理用同一 $\epsilon$)。</li><li>不 paired 时 $\beta$ 的尺度不再可比,loss 方差大幅上升,训练不稳。</li><li>这是 Diffusion-DPO 最容易忽视的实现细节。</li></ul>
<p>不知道 paired noise(错);或以为 $\epsilon$ 是 $\epsilon_\theta$ 的预测(错,这里 $\epsilon$ 是 q-sample 的 noise)。</p>
</details>
<h3 id="l2-进阶题10-题">L2 进阶题(10 题)</h3>
<details>
<summary>Q11. 推导 Diffusion-DPO loss(从 KL-regularized 最优解出发)。</summary>
<ol><li>KL-regularized 目标:$\max_p \mathbb{E}[R] - \beta\, \text{KL}(p \Vert p_\text{ref})$,最优解 $p^* \propto p_\text{ref} \exp(R/\beta)$。</li><li>反解 implicit reward$R(x_0, c) = \beta \log(p^*/p_\text{ref}) + \beta \log Z(c)$。</li><li>代入 BT$P(y_w \succ y_l) = \sigma(R_w - R_l)$$\log Z$ 在差中消掉。</li><li>替换 $p^* \to p_\theta$$\log p_\theta$ 用 ELBO surrogate$-L_\text{simple} = -\|\epsilon - \epsilon_\theta\|^2$(在 $x_t = q\text{-sample}(x_0, t, \epsilon)$ 处)。</li><li>期望对 $t \sim U(1, T)$ 取,loss 变成 $-\log\sigma(\beta T [\Delta_w - \Delta_l])$$\Delta_y = \|\epsilon - \epsilon_\text{ref}\|^2 - \|\epsilon - \epsilon_\theta\|^2$。</li></ol>
<p>直接背公式答不上来 $\log Z$ 为什么消掉;或不知道 ELBO surrogate 的来源。</p>
</details>
<details>
<summary>Q12. AlignProp 反传 $K$ 步显存 vs 性能怎么 trade-off</summary>
<ul><li>显存:$\mathcal{O}(K \cdot M_\text{UNet})$。SDXL 单 forward $\sim$8 GB activation$K=1 \to 24$ GB(含 weights + grad);$K=5 \to 60$ GB$K=10 \to 120$ GB。</li><li>性能:$K=1$ 实测已达 $K=5$ 的 95%$K \ge 5$ 在大多数 reward 上无显著提升。</li><li><strong>直觉</strong>:最后一步 $x_1 \to x_0$ 对终态影响最大,前面 49 步的方差被压缩。</li><li>工业上 $K=1$ 是标准选择(24GB 单卡可训)。</li></ul>
<p>只说"$K$ 越大越好"(错,性能曲线 saturate);不知道显存量级。</p>
</details>
<details>
<summary>Q13. DDPO 用 REINFORCE 和 PPO 区别?哪个 prod 常用?</summary>
<ul><li><strong>DDPO-SF (REINFORCE)</strong>$\hat g = \sum_t \nabla \log p_\theta \cdot (R - b)$,简单但方差大。</li><li><strong>DDPO-IS (PPO-clip)</strong>:用 importance ratio $\rho_t = p_\theta/p_{\theta_\text{old}}$,多次 update 同 batchclip $\rho_t$。</li><li><strong>per-step ratio</strong> 而非 trajectory ratio(避免 $T$ 个 ratio 累乘的方差爆炸)。</li><li>Prod 常用 PPO 形式:稳一些,sample efficiency 更高。</li></ul>
<p>说"trajectory-level ratio"(错,per-step);或不知道两个都是 DDPO。</p>
</details>
<details>
<summary>Q14. SPO 怎么得到 in-step preference pair?为什么需要 step-RM</summary>
<ul><li><strong>In-step</strong>:给定 $x_t$,独立采两个 $x_{t-1}^a, x_{t-1}^b$(用 policy 的 stochastic transition $p_\theta(\cdot \mid x_t)$ 采两次)。</li><li><strong>step-wise reward model</strong> $R_\text{step}(x_{t-1}, x_t, c, t)$ 判 winner。</li><li>step-RM 训练数据:base UNet rollout 多 trajectory,每步的 step-reward 由终态 reward 反推(类似 Math-Shepherd 的 rollout-based PRM)。</li><li>不用 step-RM 用终态 RM 也行,但要 rollout 到 $x_0$ 才能打分,贵 $T$ 倍。</li></ul>
<p>不知道 in-step pair 怎么得(错,要采两次);或不知道 step-RM 是 SPO 独有。</p>
</details>
<details>
<summary>Q15. Diffusion-DPO vs D3PO 实质差异?</summary>
<ul><li><strong>推导路径</strong>Diffusion-DPO 用 ELBO surrogate(单步 $\epsilon$-distance),D3PO 用完整 trajectory log-ratio。</li><li><strong>数学等价性</strong>:在 ELBO 下界 + 期望 over $t$ 下,D3PO 的 trajectory 形式退化为 Diffusion-DPO 的单步形式。</li><li><p><strong>实践差异</strong></p>
<ul><li>Diffusion-DPO 每步只算一次 UNet forwardpolicy + ref),便宜。</li><li>D3PO 严格意义上要算完整 trajectory $T$ 次 forward。</li></ul></li><li>工业部署主流用 Diffusion-DPO(便宜 + 稳)。</li></ul>
<p>说"完全不同"(错,理论等价);或不知道 D3PO 也是 DPO 家族。</p>
</details>
<details>
<summary>Q16. Flow-GRPO 的 denoising reduction 是什么?为什么不掉点?</summary>
<ul><li>训练时 SDE 用少步($T_\text{train} = 10$),推理时仍用全步($T_\text{infer} = 28$$50$)。</li><li><p><strong>经验上不大掉点的理由</strong>(注意:这是经验观察 + approximation,不是严格等价):</p>
<ul><li>SDE 的 <strong>连续 marginal</strong> $p_t$ 与离散步数无关;但 <strong>离散 sampler 的实际分布</strong> 与 step 数有关——少步是 discretization-error 较大的近似。所以严格说 "same marginal" 只在 continuous limit 成立。</li><li>RL 学的是 $v_\theta$ 的方向修正,<strong>方向信号</strong>与具体步数耦合较弱(这是经验观察)。</li><li>在 GenEval/OCR 这类 rule-based reward 上不掉点;在更主观 reward 上略掉但可接受。</li></ul></li><li><strong>省 sample 成本</strong>:训练每 prompt $G \cdot T_\text{train}$ 次 forward → 1/3 成本。</li></ul>
<p>不知道 marginal 不变(错);或以为 train/infer 必须同步数。</p>
</details>
<details>
<summary>Q17. MaPO 怎么去掉 reference modelloss 长什么样?</summary>
<ul><li>不用 $\log(p_\theta/p_\text{ref})$,直接用<strong>绝对 likelihood margin</strong></li></ul>
<p>$$\mathcal{L}_\text{MaPO} = -\log\sigma\!\big(\beta(\hat\ell_w - \hat\ell_l) - \gamma\big) + \alpha \hat\ell_w$$</p>
<ul><li>$\hat\ell = -\|\epsilon - \epsilon_\theta\|^2$ 是 likelihood surrogate。</li><li>$\gamma$ 是 margin(类似 SimPO),$\alpha\hat\ell_w$ 项防止"两边都降"。</li><li>显存省一半(无 ref UNet),训练快 15%,且解决 reference mismatch 问题(fine-tune 到风格差异大的目标时稳)。</li></ul>
<p>说去 ref 就完事(错,要加 likelihood term 防 degenerate);不知道 reference mismatch。</p>
</details>
<details>
<summary>Q18. Reward ensemble 为什么用 min 比 mean 好?</summary>
<ul><li>mean:被一个高分 RM 主导可能仍 hack。</li><li>min:要所有 RM 都同意"好"才给高 reward → hacking 必须同时骗过所有 RM,难度指数级上升。</li><li>等价于 conservative aggregation (Coste 2024 ICLR for LLM)diffusion 上同理。</li><li>代价:reward 偏保守,涨幅小。</li><li>工业上常用 <code>R = mean - k * std</code>(含 uncertainty penalty)作折中。</li></ul>
<p>只说"防 hacking"不知道为啥 min;或不知道这是 LLM 也用的 ensemble 策略。</p>
</details>
<details>
<summary>Q19. Diffusion-KTO 比 Diffusion-DPO 有什么独特优势?</summary>
<ul><li>只需 <strong>per-image binary feedback</strong>thumbs up/down),<strong>不需要 paired comparison</strong></li><li>工业场景大量用户 reaction(喜欢/不喜欢)远多于 paired comparison → KTO 让这部分数据可用。</li><li>prospect-theoretic 价值函数 $v(\cdot)$ 对正负 feedback 不对称(loss aversion)。</li><li>不需要"哪个更好"的标注成本。</li></ul>
<p>不知道 KTO 是 unpaired(错,这是 KTO 全部 idea);或不知道 prospect theory 来源。</p>
</details>
<details>
<summary>Q20. 为什么 diffusion 没有 token-level KL anchor,而是 trajectory-level</summary>
<ul><li>LLM 的 KL 是 per-token$\sum_t \log(\pi_\theta(y_t)/\pi_\text{ref}(y_t))$。</li><li>Diffusion 的 KL 是 per-stepper-denoising-step),不是 per-pixel$\sum_t \text{KL}(p_\theta(\cdot \mid x_t) \Vert p_\text{ref}(\cdot \mid x_t))$。</li><li>两个 Gaussian KL 有闭式:$\text{KL} = \frac{1}{2}\big[(\mu_\theta - \mu_\text{ref})^2/\sigma^2 + (\sigma_\theta/\sigma_\text{ref})^2 - 1 - 2\log(\sigma_\theta/\sigma_\text{ref})\big]$。</li><li>像素之间不独立(卷积/attention),所以 KL 是整张图 level 而非 per-pixel。</li></ul>
<p>说"per-pixel KL"(错,per-step);不知道 Gaussian KL 闭式。</p>
</details>
<h3 id="l3-顶级-lab-题5-题">L3 顶级 lab 题(5 题)</h3>
<details>
<summary>Q21. 推 Diffusion-DPO loss 从 reverse ELBO 出发,说清 ELBO surrogate 为何有效。</summary>
<ol><li>DDPM ELBO$\log p_\theta(x_0) \ge -\sum_{t=2}^T \text{KL}(q(x_{t-1}|x_t,x_0) \Vert p_\theta(x_{t-1}|x_t)) + \log p_\theta(x_0|x_1) - \text{KL}(q(x_T|x_0) \Vert p(x_T))$</li><li>化简(Ho 2020):$-\log p_\theta(x_0) \le L_\text{simple} + C$$L_\text{simple} = \mathbb{E}_{t,\epsilon}\|\epsilon - \epsilon_\theta(x_t,t)\|^2$。</li><li>KL-regularized 最优 $p^* \propto p_\text{ref}\exp(R/\beta)$,反解 $R = \beta\log(p^*/p_\text{ref}) + \beta\log Z$。</li><li>代入 BT$\log Z$ 消掉。</li><li>用 ELBO surrogate 代 $\log p$$\log p_\theta(x_0) \approx -L_\text{simple}$<strong>注意</strong>:这是上界的取负,作为单 sample 估计 — 严格意义上是 lower bound 的一个项,而非 $\log p$ 本身,但作为 DPO 的 implicit reward proxy 数值有效)。</li><li>期望对 $t$ 取,得到最终 loss。</li></ol>
<p><strong>为什么 ELBO surrogate 有效的更深一层</strong>DPO 的 implicit reward 是 $\beta\log(p_\theta/p_\text{ref})$,它只依赖<strong>相对</strong> likelihood。ELBO surrogate 的常数项($C$)在 $p_\theta$ 和 $p_\text{ref}$ 之间相消(两个模型用同一架构),只剩 $-\|\epsilon - \epsilon_\theta\|^2$ 的差。所以即使 ELBO 不是 $\log p$ 的紧 bound<strong>差异是可消的</strong></p>
<p>只能写出最终公式背不出推导链;或不知道常数项相消是关键。</p>
</details>
<details>
<summary>Q22. AlignProp 反传 $K$ 步显存 $\mathcal{O}(K)$ 是否真的无法绕过?</summary>
<p><strong>理论上</strong>可以,工程上很贵:</p>
<ol><li><p><strong>Gradient checkpointing</strong>:把 activation 的存储换成重算。每步 forward 不存 activation,反传时重新 forward 算 grad。</p>
<ul><li>显存:从 $\mathcal{O}(K \cdot M)$ 降到 $\mathcal{O}(\sqrt{K} \cdot M)$ + $\mathcal{O}(K \cdot \text{state})$。</li><li>代价:反传速度慢 2-3x。</li></ul></li><li><p><strong>Reversible ResNet</strong>:如果 UNet 用 reversible 架构(i-RevNet 风格),反传时从 output 反推 input,不存 activation。</p>
<ul><li>但 Stable Diffusion / SDXL UNet 不是 reversible。</li></ul></li><li><p><strong>Implicit gradient</strong>:通过 fixed-point 假设把 $\nabla_\theta$ 写成 implicit function theorem。</p>
<ul><li>需要 sampler 收敛到 fixed pointdiffusion 不满足。</li></ul></li><li><strong>Truncated backprop with control variates</strong>:DRaFT-K 已是这个方向;理论上加 control variates 可进一步降方差但不降内存。</li></ol>
<p><strong>实践答案</strong>$K=1$ + gradient checkpointing + LoRA 是工程最优解。$\mathcal{O}(K)$ 不可绕过的本质是 — sampler 不是 reversible computation。</p>
<p>只说 gradient checkpointing 不到位;不知道 reversibility 假设。</p>
</details>
<details>
<summary>Q23. Flow-GRPO 中 vector field $v_\theta$ 的 advantage 几何意义?</summary>
<p>GRPO 的 advantage 在 vector field 空间作用如下:</p>
<ol><li><strong>Group statistics</strong>:对同一 prompt $c$ sample $G$ 条 SDE trajectory,每条得到不同 $x_0^{(i)}$reward $r_i$ 给整条 trajectory 同一 advantage $\hat A_i = (r_i - \bar r)/\sigma_r$。</li><li><strong>沿 trajectory 的梯度</strong>$\nabla_\theta L = \sum_t \nabla_\theta \log p_\theta(x_{t-1}^{(i)} \mid x_t^{(i)}) \cdot \hat A_i$。在 Gaussian transition 下,$\log p \propto -(x_{t-1} - \mu_\theta)^2/(2\sigma^2)$,所以 $\nabla_\theta \log p \propto (x_{t-1} - \mu_\theta)\nabla_\theta \mu_\theta / \sigma^2$。</li><li><strong>$\mu_\theta$ 的物理含义</strong>:在 Flow Matching SDE 下,$\mu_\theta = x_t + (v_\theta + \frac{1}{2}\sigma^2 s_\theta) dt$$\nabla_\theta \mu_\theta \approx dt \cdot \nabla_\theta v_\theta$(忽略 score 项)。</li><li><strong>几何意义</strong>advantage $\hat A_i > 0$ 时,把 $v_\theta(t, x_t^{(i)})$ 朝 $(x_{t-1}^{(i)} - x_t^{(i)})/dt$ 方向推(即 trajectory 实际经过的方向);advantage $<0$ 朝相反方向推。</li><li><strong>vs ODE 视角</strong>:等价于在 vector field 空间做"组相对方向 reweight"——好的 trajectory 让 $v_\theta$ 在那个 $(t, x_t)$ 上指向它经过的方向,坏的反之。</li></ol>
<p>这是 vector field 上的 "reward-weighted importance sampling":每条 SDE trajectory 是 $v_\theta$ 的一次"提议方向"advantage 决定要不要 follow。</p>
<p>完全说不出几何就 0 分;说"reweighting"但不能说清在哪个空间也只值半分。</p>
</details>
<details>
<summary>Q24. SD3 / FLUX 是否真的用了 RL post-training?怎么判断?</summary>
<p><strong>诚实回答</strong>:公开论文 / 技术报告<strong>都没明说</strong>用 RL / DPO。但有以下线索:</p>
<ol><li><strong>SD3 论文 (arXiv 2403.03206)</strong>:只讨论 Rectified Flow + MM-DiT + reflow;没提 reward fine-tune。</li><li><strong>FLUX</strong>:完全没发论文,model card 只提"trained on a large image-text dataset"。</li><li><strong>DALL-E 3 (OpenAI 2023)</strong>:明确说用了 caption-faithful RLHFrewrite caption + RM)。</li><li><strong>业界共识</strong>:闭源大模型(FLUX pro, DALL-E 3, Midjourney v6+)几乎确定有 reward-based fine-tune,但具体方法不公开。</li></ol>
<p><strong>判断标准(black-box test</strong></p>
<ul><li>给同一 prompt 让模型生成 100 张,FID-100 / multi-mode 多样性低 → 可能是 RL/DPOmode collapse 信号)。</li><li>prompt-image alignment 在 GenEval 高分但 portrait 风格单一 → reward over-optimization 信号。</li><li>同一 model 对 "vibrant"/"colorful" prompt 反应过强 → HPSv2/aesthetic RM 痕迹。</li></ul>
<p><strong>结论</strong>FLUX 大概率有内部 DPO + distill 混合;SD3.5 推测有 SFT + 可能的 DPO。但<strong>没有公开证据</strong>——这道题的关键是答出"不公开但有间接证据",避免胡编技术细节。</p>
<p>如果直接答"SD3 用了 Diffusion-DPO"是错的(论文没说);要答"未公开但社区推测 + 列举证据"。</p>
</details>
<details>
<summary>Q25. 如果让你设计一个 diffusion post-training pipeline,从 $0$ 开始,你怎么选?</summary>
<p><strong>取决于约束</strong>。给一个 generic 推荐:</p>
<p><strong>Phase 1: 偏好数据收集</strong></p>
<ul><li>收集 paired preference (Pick-a-Pic 风格):成本高但 DPO 直接可用。</li><li>收集 binary feedback (thumbs up/down):成本低,用 Diffusion-KTO。</li><li>收集 rule-based ground truth (GenEval 类型 prompt + 自动 verifier):成本低,用 Flow-GRPO。</li></ul>
<p><strong>Phase 2: 算法选择</strong></p>
<ul><li><strong>首选 Diffusion-DPO</strong>:offline、稳、便宜、社区代码成熟(HuggingFace <code>diffusers</code> 直接支持)。</li><li><strong>如果 base 是 Flow Matching (SD3/FLUX)</strong>:用 Flow-GRPOrule-based reward 优先。</li><li><strong>如果 fine-tune 到新风格 / 显存紧</strong>:用 MaPO(去 ref,省一半显存)。</li><li><strong>如果 reward 可导且想榨干信号</strong>DRaFT-1 + LoRA,配 HPSv2 + PickScore ensemble。</li><li><strong>NOT 首选 DDPO</strong>on-policy sampling 太贵,工程复杂度高,性能 vs DPO 无显著优势。</li></ul>
<p><strong>Phase 3: Reward 设计</strong></p>
<ul><li><strong>Multi-RM ensemble</strong>min 或 mean - k·std):HPSv2 + PickScore + ImageReward。</li><li><strong>加 rule-based safety</strong>NSFW detector hard penalty。</li><li><strong>加 rule-based alignment</strong>GenEval 自动 verifierobject count, OCR)。</li><li><strong>每个 RM 独立 z-score 归一化</strong></li></ul>
<p><strong>Phase 4: 监控与 early stop</strong></p>
<ul><li>每 N 步算 reward + FID-100kreward 涨 + FID 涨 = hacking 信号。</li><li>KL budget 监控:$\text{KL}(p_\theta \Vert p_\text{ref}) > K_\text{target}$ 时 stop。</li><li>Human eval blind A/B (base vs RL) 每 1000 steps。</li></ul>
<p><strong>Phase 5: distill 衔接</strong></p>
<ul><li>Post-training 完后做 ADD / LCM 蒸馏到 4-step / 1-step。</li><li>注意 distill 可能消除部分 RL 增益,需要 distill-aware 后训。</li></ul>
<p>只答"用 Diffusion-DPO" 是浅;要答出"phase 分解 + 多 reward + 监控 + distill 衔接"才完整。</p>
</details>
<h2 id="a-附录">§A 附录</h2>
<h3 id="a1-关键论文清单含-arxiv-id">A.1 关键论文清单(含 arXiv ID)</h3>
<table><thead><tr><th>论文</th><th>一句话</th><th>arXiv</th><th>发表</th></tr></thead><tbody><tr><td><strong>DDPO</strong></td><td>Diffusion 当 MDPREINFORCE/PPO 训</td><td><a href="https://arxiv.org/abs/2305.13301">2305.13301</a></td><td>ICLR 2024</td></tr><tr><td><strong>DPOK</strong></td><td>KL-regularized RL for diffusion</td><td><a href="https://arxiv.org/abs/2305.16381">2305.16381</a></td><td>NeurIPS 2023</td></tr><tr><td><strong>DRaFT</strong></td><td>直接 reward 反传 $K$ 步</td><td><a href="https://arxiv.org/abs/2309.17400">2309.17400</a></td><td>ICLR 2024</td></tr><tr><td><strong>AlignProp</strong></td><td>reward backprop with randomized truncation</td><td><a href="https://arxiv.org/abs/2310.03739">2310.03739</a></td><td>ICLR 2024</td></tr><tr><td><strong>ImageReward / ReFL</strong></td><td>137K human pair RM + 单步 reward fine-tune</td><td><a href="https://arxiv.org/abs/2304.05977">2304.05977</a></td><td>NeurIPS 2023</td></tr><tr><td><strong>HPSv2</strong></td><td>798K human pair RM</td><td><a href="https://arxiv.org/abs/2306.09341">2306.09341</a></td><td>arXiv 2023</td></tr><tr><td><strong>PickScore (Pick-a-Pic)</strong></td><td>1M user pair, CLIP RM</td><td><a href="https://arxiv.org/abs/2305.01569">2305.01569</a></td><td>NeurIPS 2023</td></tr><tr><td><strong>Diffusion-DPO</strong></td><td>ELBO surrogate + DPO loss</td><td><a href="https://arxiv.org/abs/2311.12908">2311.12908</a></td><td>CVPR 2024</td></tr><tr><td><strong>D3PO</strong></td><td>trajectory-level DPO for diffusion</td><td><a href="https://arxiv.org/abs/2311.13231">2311.13231</a></td><td>CVPR 2024</td></tr><tr><td><strong>SPO</strong></td><td>step-aware preference + step-RM</td><td><a href="https://arxiv.org/abs/2406.04314">2406.04314</a></td><td>arXiv 2024</td></tr><tr><td><strong>Diffusion-KTO</strong></td><td>unpaired binary feedback (KTO for diffusion)</td><td><a href="https://arxiv.org/abs/2404.04465">2404.04465</a></td><td>NeurIPS 2024</td></tr><tr><td><strong>MaPO</strong></td><td>margin-aware, no ref</td><td><a href="https://arxiv.org/abs/2406.06424">2406.06424</a></td><td>arXiv 2024</td></tr><tr><td><strong>Flow-GRPO</strong></td><td>GRPO for Flow Matching via ODE→SDE</td><td><a href="https://arxiv.org/abs/2505.05470">2505.05470</a></td><td>arXiv 2025</td></tr><tr><td><strong>SD3 (Rectified Flow + MM-DiT)</strong></td><td>base model</td><td><a href="https://arxiv.org/abs/2403.03206">2403.03206</a></td><td>ICML 2024</td></tr><tr><td><strong>Constitutional AI (RLAIF 起源)</strong></td><td>AI feedback 替代 human</td><td><a href="https://arxiv.org/abs/2212.08073">2212.08073</a></td><td>arXiv 2022</td></tr><tr><td><strong>KTO (LLM)</strong></td><td>prospect theory alignment</td><td><a href="https://arxiv.org/abs/2402.01306">2402.01306</a></td><td>arXiv 2024</td></tr></tbody></table>
<h3 id="a2-常用-reward-model-资源">A.2 常用 reward model 资源</h3>
<ul><li><strong>ImageReward</strong>https://github.com/THUDM/ImageReward</li><li><strong>HPSv2</strong>https://github.com/tgxs002/HPSv2</li><li><strong>PickScore</strong>https://github.com/yuvalkirstain/PickScore</li><li><strong>CLIP</strong>OpenAI / OpenCLIP,多 backbone 可选</li></ul>
<h3 id="a3-开源训练代码">A.3 开源训练代码</h3>
<ul><li><strong>TRL (HuggingFace)</strong><code>diffusers</code> + <code>DPO Trainer</code> for Diffusion-DPO(最成熟)</li><li><strong>DDPO 原始仓库</strong>https://github.com/kvablack/ddpo-pytorch</li><li><strong>AlignProp</strong>https://github.com/mihirp1998/AlignProp</li><li><strong>DRaFT (Google research)</strong>https://github.com/clarkjkr/draftClark et al. 2024 ICLR</li><li><strong>MaPO</strong>https://github.com/mapo-t2i/mapo</li><li><strong>Flow-GRPO</strong>:通过论文 arXiv 2505.05470 找官方实现</li></ul>
<h3 id="a4-工程踩坑清单">A.4 工程踩坑清单</h3>
<table><thead><tr><th></th><th></th></tr></thead><tbody><tr><td>Diffusion-DPO 没 paired noise</td><td>$\epsilon$ for $x_t^w$ 和 $x_t^l$ 必须共享</td></tr><tr><td>DDPO 用 DDIM-eta=0</td><td>必须 eta&gt;0 或 DDPM,否则梯度为 0</td></tr><tr><td>AlignProp 显存爆炸</td><td>$K=1$ + gradient checkpoint + LoRA</td></tr><tr><td>Reward scale 不归一化</td><td>每个 RM 单独 z-score</td></tr><tr><td>RL 后 FID 暴跌</td><td>加 KL anchor 或 reward ensemble</td></tr><tr><td>$\beta$ 调不动</td><td>Diffusion-DPO 用 $\beta \in [2000, 5000]$,不是 LLM 的 0.1</td></tr><tr><td>Flow-GRPO 训练慢</td><td>用 denoising reduction ($T_\text{train} < T_\text{infer}$)</td></tr><tr><td>MaPO 训崩</td><td>$\alpha\hat\ell_w$ 项必须够大防 likelihood 一起降</td></tr><tr><td>Step-RM 训不起来</td><td>用 rollout-based 自动标注(类 Math-Shepherd</td></tr><tr><td>reward hacking 检测不到</td><td>同时监控 reward + FID + human blind A/B</td></tr></tbody></table>
<h3 id="a5-与-0-tldr-的呼应">A.5 与 §0 TL;DR 的呼应</h3>
<table><thead><tr><th>TL;DR 条</th><th>详见章节</th></tr></thead><tbody><tr><td>1. 为什么难</td><td>§1</td></tr><tr><td>2. 三条主线</td><td>§1.2</td></tr><tr><td>3. DDPO</td><td>§2.12.4</td></tr><tr><td>4. DRaFT / AlignProp</td><td>§3.23.3</td></tr><tr><td>5. Diffusion-DPO</td><td>§4.1</td></tr><tr><td>6. D3PO</td><td>§4.2</td></tr><tr><td>7. SPO</td><td>§4.3</td></tr><tr><td>8. Flow-GRPO</td><td>§5</td></tr><tr><td>9. Reward hacking</td><td>§3.6 + §7</td></tr></tbody></table>
<div class="callout callout-good"><div class="callout-title">学完 checkpoint</div></div>
<ul><li>能口述 Diffusion-DPO loss 形式 + paired noise 细节</li><li>能解释 AlignProp 为什么 $K=1$ 够用 + 显存 $\mathcal{O}(K)$</li><li>能写 DDPO 的 state/action/reward + per-step ratio</li><li>能讲 Flow-GRPO 的 ODE→SDE 转换为什么必要 + denoising reduction</li><li>知道 SD3/FLUX 是否用 RL 的诚实答案(公开未明说)</li></ul>
<footer class="aris-footer">
Generated by <a href="https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/blob/main/skills/render-html/SKILL.md">ARIS <code>/render-html</code></a> ·
source path <code>docs/tutorials/diffusion_post_training_tutorial.md</code> ·
SHA256 <code>95b47c844209</code> ·
generated at 2026-05-19 16:25 UTC.
This is a generated view — edit the source Markdown, then re-render.
</footer>
</main>
</div>
</body>
</html>