Files
2026-07-13 13:37:02 +08:00

1019 lines
92 KiB
HTML
Raw Permalink Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html lang="zh-CN">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>3D Generation 面试 Cheat Sheet</title>
<meta name="generator" content="ARIS render-html (academic, v1)">
<meta name="aris:source-path" content="docs/tutorials/3d_generation_tutorial.md">
<meta name="aris:source-sha256" content="beab28036da8f28744b8d64815a669d0fce44c978bd869145e41b15860659ccc">
<meta name="aris:generated-at" content="2026-05-19 07:54 UTC">
<!-- MathJax 3 -->
<script>
window.MathJax = {
tex: { inlineMath: [['$', '$'], ['\\(', '\\)']], displayMath: [['$$', '$$'], ['\\[', '\\]']], processEscapes: true },
options: { skipHtmlTags: ['script', 'noscript', 'style', 'textarea', 'pre', 'code'] }
};
</script>
<script src="https://cdn.jsdelivr.net/npm/mathjax@3/es5/tex-mml-chtml.js" async></script>
<!-- highlight.js -->
<link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/highlightjs/cdn-release@11.9.0/build/styles/atom-one-light.min.css">
<script src="https://cdn.jsdelivr.net/gh/highlightjs/cdn-release@11.9.0/build/highlight.min.js"></script>
<script>document.addEventListener('DOMContentLoaded', () => hljs.highlightAll());</script>
<style>
:root {
--bg: #fdfcf7;
--bg-soft: #f4f1ea;
--bg-code: #f8f5ec;
--ink: #1a1a1a;
--ink-soft: #4a4a4a;
--ink-muted: #6b6b6b;
--primary: #1a4a8c;
--primary-soft: #2d6cb8;
--accent: #b8390e;
--warn: #b45309;
--warn-bg: #fef3c7;
--info-bg: #dbeafe;
--good-bg: #d1fae5;
--good: #065f46;
--bad-bg: #fee2e2;
--bad: #991b1b;
--border: #d6d0c0;
--border-soft: #e8e3d5;
}
* { box-sizing: border-box; }
html { scroll-behavior: smooth; }
body {
font-family: "Source Serif Pro", "Source Serif 4", "Crimson Pro", "Georgia", "Songti SC", "STSong", serif;
line-height: 1.65;
color: var(--ink);
background: var(--bg);
margin: 0;
padding: 0;
font-size: 16px;
}
.layout {
max-width: 1280px;
margin: 0 auto;
display: grid;
grid-template-columns: 260px 1fr;
gap: 48px;
padding: 40px 32px;
}
nav.toc {
position: sticky;
top: 24px;
align-self: start;
font-size: 13px;
max-height: calc(100vh - 48px);
overflow-y: auto;
border-right: 1px solid var(--border-soft);
padding-right: 16px;
}
nav.toc h3 {
margin: 0 0 12px;
font-size: 12px;
text-transform: uppercase;
letter-spacing: 0.08em;
color: var(--ink-muted);
font-weight: 600;
}
nav.toc ol { list-style: none; padding: 0; margin: 0; counter-reset: toc; }
nav.toc ol li { margin: 5px 0; counter-increment: toc; }
nav.toc ol li::before { content: counter(toc) ". "; color: var(--ink-muted); margin-right: 4px; }
nav.toc a {
color: var(--ink-soft);
text-decoration: none;
border-bottom: 1px dotted transparent;
}
nav.toc a:hover { color: var(--primary); border-bottom-color: var(--primary); }
nav.toc ul { list-style: none; padding-left: 14px; margin: 3px 0; font-size: 12px; }
nav.toc ul li::before { content: "→ "; color: var(--border); }
main { min-width: 0; }
header.hero {
border-bottom: 3px double var(--primary);
padding-bottom: 24px;
margin-bottom: 32px;
}
header.hero .eyebrow {
color: var(--accent);
font-size: 13px;
text-transform: uppercase;
letter-spacing: 0.12em;
font-weight: 600;
margin-bottom: 8px;
}
header.hero h1 {
font-size: 32px;
line-height: 1.2;
margin: 0 0 12px;
color: var(--ink);
font-weight: 700;
letter-spacing: -0.01em;
}
header.hero .subtitle {
font-size: 16px;
color: var(--ink-soft);
margin: 0 0 8px;
font-style: italic;
}
header.hero .byline {
font-size: 14px;
color: var(--ink-soft);
margin: 0 0 20px;
}
header.hero .byline strong {
color: var(--ink);
font-weight: 600;
}
header.hero .meta {
display: flex;
gap: 20px;
flex-wrap: wrap;
font-size: 12px;
color: var(--ink-muted);
border-top: 1px solid var(--border-soft);
padding-top: 14px;
}
header.hero .meta span strong { color: var(--ink-soft); }
header.hero .meta code {
font-family: "JetBrains Mono", "SF Mono", "Menlo", "Consolas", monospace;
font-size: 11px;
background: var(--bg-soft);
padding: 1px 5px;
border-radius: 3px;
border: 1px solid var(--border-soft);
}
h2 {
font-size: 24px;
margin: 44px 0 14px;
padding-bottom: 8px;
border-bottom: 1px solid var(--border);
color: var(--ink);
font-weight: 700;
}
h2 .num { color: var(--primary); font-weight: 600; margin-right: 8px; }
h3 { font-size: 19px; margin: 28px 0 10px; color: var(--primary); font-weight: 600; }
h4 { font-size: 16px; margin: 20px 0 8px; color: var(--ink); font-weight: 600; }
p { margin: 10px 0; }
ul, ol { padding-left: 22px; margin: 10px 0; }
ul li, ol li { margin: 4px 0; }
ul li::marker { color: var(--primary); }
strong { color: var(--accent); font-weight: 600; }
em { color: var(--ink-soft); }
a { color: var(--primary); }
a:hover { color: var(--accent); }
code:not(.hljs) {
font-family: "JetBrains Mono", "SF Mono", "Menlo", "Consolas", monospace;
font-size: 0.86em;
background: var(--bg-code);
padding: 1px 5px;
border-radius: 3px;
border: 1px solid var(--border-soft);
color: var(--accent);
}
pre {
background: #fafaf6;
border: 1px solid var(--border);
border-left: 4px solid var(--primary);
padding: 0;
overflow-x: auto;
border-radius: 4px;
margin: 14px 0;
}
pre code, pre code.hljs {
background: transparent !important;
display: block;
padding: 14px 18px !important;
font-size: 13px;
line-height: 1.55;
font-family: "JetBrains Mono", "SF Mono", "Menlo", monospace;
color: var(--ink);
}
pre.diagram {
background: #f9f6ed;
border-left: 4px solid var(--accent);
font-size: 12.5px;
line-height: 1.4;
}
.callout {
margin: 16px 0;
padding: 12px 16px;
border-radius: 4px;
border-left: 4px solid;
font-size: 15px;
}
.callout-title {
font-weight: 600;
margin-bottom: 6px;
font-size: 12px;
text-transform: uppercase;
letter-spacing: 0.06em;
}
.callout-info { background: var(--info-bg); border-left-color: var(--primary); }
.callout-info .callout-title { color: var(--primary); }
.callout-warn { background: var(--warn-bg); border-left-color: var(--warn); }
.callout-warn .callout-title { color: var(--warn); }
.callout-good { background: var(--good-bg); border-left-color: var(--good); }
.callout-good .callout-title { color: var(--good); }
.callout-bad { background: var(--bad-bg); border-left-color: var(--bad); }
.callout-bad .callout-title { color: var(--bad); }
table {
width: 100%;
border-collapse: collapse;
margin: 16px 0;
font-size: 14px;
border: 1px solid var(--border);
border-radius: 4px;
overflow: hidden;
}
thead { background: var(--primary); color: white; }
th, td {
text-align: left;
padding: 9px 12px;
border-bottom: 1px solid var(--border-soft);
vertical-align: top;
}
th { font-weight: 600; font-size: 13px; letter-spacing: 0.02em; }
tr:last-child td { border-bottom: none; }
tbody tr:nth-child(even) { background: var(--bg-soft); }
details.qa, details {
background: white;
border: 1px solid var(--border-soft);
border-radius: 6px;
margin: 10px 0;
padding: 0;
}
details summary {
cursor: pointer;
padding: 10px 14px;
font-weight: 600;
font-size: 14px;
color: var(--primary);
list-style: none;
user-select: none;
}
details summary::-webkit-details-marker { display: none; }
details summary::before {
content: "▸ ";
margin-right: 4px;
display: inline-block;
transition: transform 0.15s;
}
details[open] summary::before { transform: rotate(90deg); }
details[open] summary { border-bottom: 1px solid var(--border-soft); }
details > :not(summary) { padding: 10px 14px; }
details p:first-of-type { margin-top: 8px; }
mjx-container[display="true"] { margin: 12px 0 !important; }
footer.aris-footer {
margin-top: 60px;
padding-top: 20px;
border-top: 1px solid var(--border);
font-size: 12px;
color: var(--ink-muted);
}
footer.aris-footer a { color: var(--ink-muted); border-bottom: 1px dotted var(--border); }
@media (max-width: 900px) {
.layout { grid-template-columns: 1fr; gap: 20px; padding: 20px 16px; }
nav.toc {
position: static;
max-height: none;
border-right: none;
border-bottom: 1px solid var(--border-soft);
padding-right: 0;
padding-bottom: 14px;
}
header.hero h1 { font-size: 24px; }
h2 { font-size: 20px; }
}
@media print {
nav.toc { display: none; }
.layout { grid-template-columns: 1fr; padding: 0; }
body { background: white; }
header.hero { border-bottom-color: var(--ink); }
}
</style>
</head>
<body>
<div class="layout">
<nav class="toc">
<h3>Contents</h3>
<ol>
<li><a href="#0-tldr-cheat-sheet">§0 TL;DR Cheat Sheet</a>
</li>
<li><a href="#1-三大表示的直觉对比">§1 三大表示的直觉对比</a>
</li>
<li><a href="#2-nerf体渲染原理推导必考">§2 NeRF:体渲染原理推导(必考)</a>
<ul>
<li><a href="#21-连续体渲染公式">2.1 连续体渲染公式</a></li>
<li><a href="#22-为什么是这个形式-从物理推导">2.2 为什么是这个形式?— 从物理推导</a></li>
<li><a href="#23-离散化alpha-compositing必考推导">2.3 离散化:$\alpha$-compositing**必考推导**</a></li>
<li><a href="#24-位置编码-gammap表示高频细节">2.4 位置编码 $\gamma(p)$:表示高频细节</a></li>
<li><a href="#25-hierarchical-sampling粗--细">2.5 Hierarchical Sampling:粗 → 细</a></li>
<li><a href="#26-nerf-训练代码核心-30-行">2.6 NeRF 训练代码(核心 30 行)</a></li>
<li><a href="#27-mip-nerf--mip-nerf-360抗锯齿">2.7 Mip-NeRF / Mip-NeRF 360(抗锯齿)</a></li>
<li><a href="#28-neus--volsdf体渲染--sdf导出-mesh-的关键">2.8 NeuS / VolSDF:体渲染 + SDF(导出 mesh 的关键)</a></li>
</ul>
</li>
<li><a href="#3-instant-ngp5-oom-加速必考">§3 Instant-NGP5+ OOM 加速(必考)</a>
<ul>
<li><a href="#31-核心-idea多分辨率-hash-网格">3.1 核心 idea:多分辨率 hash 网格</a></li>
<li><a href="#32-hash-function">3.2 Hash function</a></li>
<li><a href="#33-hash-collision-怎么消歧l3-高频追问">3.3 Hash collision 怎么消歧?(**L3 高频追问**)</a></li>
<li><a href="#34-instant-ngp-训练公式">3.4 Instant-NGP 训练公式</a></li>
<li><a href="#35-plenoxels--tensorf同期的-explicit-派">3.5 Plenoxels / TensoRF(同期的 explicit 派)</a></li>
</ul>
</li>
<li><a href="#4-3d-gaussian-splatting显式可微光栅化当前主力">§4 3D Gaussian Splatting:显式可微光栅化(**当前主力**)</a>
<ul>
<li><a href="#41-场景表示">4.1 场景表示</a></li>
<li><a href="#42-3d--2d-投影-jacobianl3-必考推导">4.2 3D → 2D 投影 Jacobian**L3 必考推导**</a></li>
<li><a href="#43-可微光栅化tile-based-front-to-back-alpha-blending">4.3 可微光栅化:tile-based front-to-back alpha-blending</a></li>
<li><a href="#44-3dgs-前向pytorch-reference-实现">4.4 3DGS 前向(PyTorch reference 实现)</a></li>
<li><a href="#45-自适应密度控制面试常问">4.5 自适应密度控制(**面试常问**)</a></li>
<li><a href="#46-2dgs--surfelssurface-aligned">4.6 2DGS / Surfelssurface-aligned</a></li>
<li><a href="#47-动态-4dgs">4.7 动态 4DGS</a></li>
</ul>
</li>
<li><a href="#5-mesh-提取marching-cubes--dmtet">§5 Mesh 提取:Marching Cubes / DMTet</a>
<ul>
<li><a href="#51-marching-cubes经典必考">5.1 Marching Cubes**经典必考**</a></li>
<li><a href="#52-differentiable-dmtet--flexicubes端到端学-mesh">5.2 Differentiable: DMTet / FlexiCubes(端到端学 mesh</a></li>
</ul>
</li>
<li><a href="#6-sds-loss用-2d-diffusion-监督-3ddreamfusion-系列">§6 SDS Loss:用 2D Diffusion 监督 3DDreamFusion 系列)</a>
<ul>
<li><a href="#61-问题设置">6.1 问题设置</a></li>
<li><a href="#62-setup">6.2 Setup</a></li>
<li><a href="#63-sds-gradient-推导l3-必考">6.3 SDS gradient 推导(**L3 必考**</a></li>
<li><a href="#64-为什么扔掉-jacobian-反而-work">6.4 为什么扔掉 Jacobian 反而 work</a></li>
<li><a href="#65-sds-副作用over-saturation--mode-collapse--janus">6.5 SDS 副作用:over-saturation / mode collapse / Janus</a></li>
<li><a href="#66-sds-代码核心-30-行">6.6 SDS 代码(核心 30 行)</a></li>
<li><a href="#67-vsd变分-sdsprolificdreamer-neurips-2023-spotlight">6.7 VSD:变分 SDS**ProlificDreamer**, NeurIPS 2023 Spotlight</a></li>
<li><a href="#68-sds-衍生家族mesh--3dgs--sds">6.8 SDS 衍生家族:mesh / 3DGS + SDS</a></li>
</ul>
</li>
<li><a href="#7-single-image--few-view-3d-生成">§7 Single-Image / Few-View 3D 生成</a>
<ul>
<li><a href="#71-zero-1-to-3-范式novel-view-via-diffusion">7.1 Zero-1-to-3 范式(novel view via diffusion</a></li>
<li><a href="#72-one-2-3-45--instantmesh--triposr--stable-fast-3d">7.2 One-2-3-45 / InstantMesh / TripoSR / Stable Fast 3D</a></li>
<li><a href="#73-lrm-triplane-表示面试高频">7.3 LRM Triplane 表示(**面试高频**</a></li>
</ul>
</li>
<li><a href="#8-3d-foundation-models2024-开源浪潮">§8 3D Foundation Models2024 开源浪潮)</a>
<ul>
<li><a href="#81-trellis-microsoft-2024-开源">8.1 Trellis (Microsoft 2024, 开源)</a></li>
<li><a href="#82-hunyuan3d-1---2-tencent-2024-25-开源">8.2 Hunyuan3D-1 / -2 (Tencent 2024-25, 开源)</a></li>
<li><a href="#83-clay-zhang-2024-siggraph">8.3 CLAY (Zhang 2024 SIGGRAPH)</a></li>
<li><a href="#84-对比表">8.4 对比表</a></li>
</ul>
</li>
<li><a href="#9-复杂度--资源对比">§9 复杂度 / 资源对比</a>
</li>
<li><a href="#10-与相关方法对比--embodied-ai-应用">§10 与相关方法对比 &amp; Embodied AI 应用</a>
<ul>
<li><a href="#101-3d-vs-2d-生成关键区别">10.1 3D-vs-2D 生成关键区别</a></li>
<li><a href="#102-embodied-ai--ar--vr-实战路线">10.2 Embodied AI / AR / VR 实战路线</a></li>
</ul>
</li>
<li><a href="#11-工程实战--易踩坑">§11 工程实战 &amp; 易踩坑</a>
<ul>
<li><a href="#111-colmap--sfm-前处理重建必经">11.1 COLMAP / SfM 前处理(重建必经)</a></li>
<li><a href="#112-数值稳定nerf3dgs-通用">11.2 数值稳定(NeRF/3DGS 通用)</a></li>
<li><a href="#113-多机分布式--评测指标">11.3 多机分布式 &amp; 评测指标</a></li>
</ul>
</li>
<li><a href="#12-25-高频面试题">§12 25 高频面试题</a>
<ul>
<li><a href="#l1-必会题任何-3d--vision-岗都会问">L1 必会题(任何 3D / vision 岗都会问)</a></li>
<li><a href="#l2-进阶题research-oriented-岗位">L2 进阶题(research-oriented 岗位)</a></li>
<li><a href="#l3-顶级-lab-题顶会--industry-研究岗">L3 顶级 lab 题(顶会 / industry 研究岗)</a></li>
</ul>
</li>
<li><a href="#a-附录代码完整骨架--参考文献">§A 附录:代码完整骨架 + 参考文献</a>
<ul>
<li><a href="#a1-完整-from-scratch-代码包含">A.1 完整 from-scratch 代码包含</a></li>
<li><a href="#a2-关键论文-reading-list">A.2 关键论文 reading list</a></li>
<li><a href="#a3-embodied-ai--ar--vr-常见追问">A.3 Embodied AI / AR / VR 常见追问</a></li>
</ul>
</li>
</ol>
</nav>
<main>
<header class="hero">
<div class="eyebrow">Interview Prep · 3D Generation / NeRF / Gaussian Splatting</div>
<h1>3D Generation 面试 Cheat Sheet</h1>
<p class="subtitle">NeRF / Instant-NGP / 3DGS / SDS / DreamFusion / Trellis + 25 高频题(L1 必会 · L2 进阶 · L3 顶级 lab)</p>
<p class="byline">By <strong>Ruofeng Yang (杨若峰), Shanghai Jiao Tong University</strong></p>
<div class="meta">
<span><strong>Source:</strong> <code>docs/tutorials/3d_generation_tutorial.md</code></span>
<span><strong>SHA256:</strong> <code>beab28036da8</code></span>
<span><strong>Rendered:</strong> 2026-05-19 07:54 UTC</span>
</div>
</header>
<h2 id="0-tldr-cheat-sheet">§0 TL;DR Cheat Sheet</h2>
<div class="callout callout-info"><div class="callout-title">9 句话搞定 3D Generation</div><p>Embodied AI / AR / VR 面试核心要点(详见后文 §1–§11 推导)。</p></div>
<ol><li><strong>三大表示</strong><strong>NeRF</strong>(隐式神经场 + 体渲染)、<strong>3DGS</strong>(显式 Gaussian 点云 + 光栅化)、<strong>Mesh / SDF</strong>(显式表面 / 隐式距离场)。重建质量与速度的 sweet spot3DGSKerbl 2023 SIGGRAPH Best Paper)。</li><li><strong>NeRF 核心公式</strong>$C(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\sigma(\mathbf{r}(t))\mathbf{c}(\mathbf{r}(t),\mathbf{d})\,dt$,其中 $T(t) = \exp\!\left(-\int_{t_n}^{t}\sigma(\mathbf{r}(s))\,ds\right)$ 是 transmittance。离散化得到 $\alpha$-compositing$C \approx \sum_i T_i (1-e^{-\sigma_i\delta_i})\mathbf{c}_i$。</li><li><strong>Instant-NGP</strong> (Müller 2022 SIGGRAPH)<strong>多分辨率 hash 网格</strong> + tiny MLP5+ OOM 加速;hash collision 由 MLP 在含碰撞表项上自动学习消歧(被 loss + 多尺度冗余共同压制)。</li><li><strong>3DGS 核心</strong>:场景表示为一组 3D Gaussian $\{\mu_i, \Sigma_i, \alpha_i, c_i(\mathbf{d})\}$<strong>可微光栅化</strong>通过将 3D 协方差用 Jacobian $J$ 投影到 2D$\Sigma' = J W \Sigma W^\top J^\top$,按深度排序后做 front-to-back alpha-blending。</li><li><strong>DreamFusion SDS</strong> (Poole et al. 2022 arXiv → ICLR 2023 Outstanding):用 pretrained 2D diffusion 监督 3D 表示:$\nabla_\theta \mathcal{L}_\text{SDS} = \mathbb{E}_{t,\epsilon}[w(t)(\epsilon_\phi(x_t;y,t)-\epsilon)\,\partial x/\partial \theta]$<strong>故意去掉 U-Net 对 $x_t$ 求导的 Jacobian 项</strong>,使得训练 simulation-free。代价:mode-seeking → over-saturation / Janus。</li><li><strong>VSD</strong> (Wang 2023 NeurIPS, ProlificDreamer):把 3D 参数 $\theta$ 视为 random variable $\mu(\theta)$<strong>最小化的是渲染加噪图像分布</strong>之间的 KL$\mathbb{E}_t\big[D_\text{KL}\big(q_\mu^t(x_t|y)\,\|\,p_\phi^t(x_t|y)\big)\big]$。梯度形式为 <strong>relative score</strong> $\nabla_\theta \approx (\epsilon_\phi(x_t;y,t) - \epsilon_\psi(x_t;y,t,\pi))\,\partial x/\partial\theta$,其中 $\epsilon_\psi$ 是 LoRA 微调的辅助 score。CFG 可从 100 降至 7.5。</li><li><strong>Single-image / Few-view 3D</strong>Zero-1-to-3 (Liu 2023 ICCV) 用 viewpoint conditioned diffusionSyncDreamer / MVDream 学多视图联合一致性;TripoSR / InstantMesh / Stable Fast 3D 把 image-to-mesh 推到 ≤3 秒。</li><li><strong>3D Foundation Models (2024-25 开源)</strong><strong>Trellis</strong> (Microsoft 2024) 用 structured latent + flow matching<strong>Hunyuan3D-2</strong> (Tencent 2025) shape→texture 两阶段;<strong>CLAY</strong> (Zhang 2024 SIGGRAPH) 大尺度 latent diffusion + 3DShape2VecSet。</li><li><strong>Embodied AI 关键应用</strong>Sim2Real 资产生成、NeRF/3DGS 作为可微 simulator、language-conditioned 3D affordance。<strong>面试常见交叉</strong>NeRF SLAM、Gaussian-Splat scene editing、3D 物理一致性。</li></ol>
<h2 id="1-三大表示的直觉对比">§1 三大表示的直觉对比</h2>
<p>3D 生成的第一选择题是 <strong>representation</strong>——选错了下游全废。</p>
<table><thead><tr><th></th><th>NeRF (Implicit Field)</th><th>3DGS (Explicit Point)</th><th>Mesh / SDF</th></tr></thead><tbody><tr><td><strong>存储</strong></td><td>MLP 权重 $f_\theta(\mathbf{x},\mathbf{d}) \to (\sigma, \mathbf{c})$</td><td>一堆 3D Gaussian $\{\mu_i, \Sigma_i, \alpha_i, c_i\}$</td><td>三角网格 / signed distance</td></tr><tr><td><strong>渲染</strong></td><td>Ray marching + 体积分(GPU 数百 ms/帧)</td><td>Differentiable rasterizationGPU 数 ms/帧)</td><td>Rasterization(实时)</td></tr><tr><td><strong>训练</strong></td><td>数百视图,数小时 (vanilla)</td><td>数十视图,10-30 分钟</td><td>需要 mesh + texture optim</td></tr><tr><td><strong>质量</strong></td><td>视图合成 SOTA</td><td>与 NeRF 持平甚至更好(PSNR 高)</td><td>受多边形分辨率制约</td></tr><tr><td><strong>编辑</strong></td><td>困难(神经场不可解释)</td><td>容易(点可移动、删除、合并)</td><td>容易(标准 DCC 流程)</td></tr><tr><td><strong>导出 mesh</strong></td><td>难(需 NeuS / Poisson</td><td>中等(2DGS / GSDF / SuGaR</td><td>自身就是 mesh</td></tr><tr><td><strong>下游适配</strong></td><td>物理仿真不友好</td><td>容易接 PBR、IsaacSim、URDF</td><td>标准 robot / AR/VR pipeline</td></tr></tbody></table>
<div class="callout callout-info"><div class="callout-title">面试直觉</div><p>Embodied AI 倾向 mesh / 3DGS(仿真器友好);AR/VR 看场景规模(前景小物 mesh,大场景 3DGS);视觉重建 SOTA 用 3DGS。NeRF 现在偏 research baseline,工业落地以 3DGS 为主。</p></div>
<h2 id="2-nerf体渲染原理推导必考">§2 NeRF:体渲染原理推导(必考)</h2>
<h3 id="21-连续体渲染公式">2.1 连续体渲染公式</h3>
<p>NeRF (Mildenhall 2020 ECCV <strong>Best Paper Honorable Mention</strong>) 把场景表示为 <strong>5D 神经场</strong> $f_\theta : (\mathbf{x}, \mathbf{d}) \to (\sigma, \mathbf{c})$</p>
<ul><li>输入:3D 位置 $\mathbf{x} \in \mathbb{R}^3$ + 视角方向 $\mathbf{d} \in \mathbb{S}^2$</li><li>输出:体密度 $\sigma \ge 0$(与方向无关)+ 颜色 $\mathbf{c} \in \mathbb{R}^3$(依赖方向,捕捉镜面反射)</li></ul>
<p>对相机射线 $\mathbf{r}(t) = \mathbf{o} + t\mathbf{d}$,沿 $t \in [t_n, t_f]$ 体积分得到像素颜色:</p>
<p>$$\boxed{\;C(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\,\sigma(\mathbf{r}(t))\,\mathbf{c}(\mathbf{r}(t),\mathbf{d})\,dt\;}$$</p>
<p>其中 <strong>transmittance</strong>(光线从 $t_n$ 到 $t$ 没被遮挡的概率):</p>
<p>$$\boxed{\;T(t) = \exp\!\left(-\int_{t_n}^{t}\sigma(\mathbf{r}(s))\,ds\right)\;}$$</p>
<h3 id="22-为什么是这个形式-从物理推导">2.2 为什么是这个形式?— 从物理推导</h3>
<p>考虑光线穿过参与介质(participating medium)。在 $[t, t+dt]$ 段内:</p>
<ul><li>被吸收 / 散射出射线的概率:$\sigma(\mathbf{r}(t))\,dt$</li><li>介质在该点发射的颜色贡献:$\mathbf{c}(\mathbf{r}(t),\mathbf{d})$</li></ul>
<p>令 $T(t)$ 为光线从 $t_n$ 到 $t$ 仍存活的概率。从 $t \to t + dt$,存活概率变化:</p>
<p>$$T(t+dt) = T(t)\,(1 - \sigma\,dt) \;\Rightarrow\; \frac{dT}{dt} = -\sigma(t)\,T(t)$$</p>
<p>这是一个一阶 ODE,初值 $T(t_n) = 1$,解出:</p>
<p>$$T(t) = \exp\!\left(-\int_{t_n}^{t}\sigma(\mathbf{r}(s))\,ds\right)$$</p>
<p>每个深度 $t$ 处贡献到像素的颜色 = <strong>存活概率 × 该处吸收概率 × 该处颜色</strong></p>
<p>$$dC = T(t)\,\sigma(t)\,\mathbf{c}(t)\,dt$$</p>
<p>积分即得 $C(\mathbf{r})$。</p>
<h3 id="23-离散化alpha-compositing必考推导">2.3 离散化:$\alpha$-compositing<strong>必考推导</strong></h3>
<p>实际无法做连续积分,把 $[t_n, t_f]$ 切成 $N$ 段,每段间距 $\delta_i = t_{i+1} - t_i$,假设段内 $\sigma, \mathbf{c}$ 为常数 $\sigma_i, \mathbf{c}_i$。</p>
<p><strong>段内 transmittance 衰减</strong>:在 $[t_i, t_{i+1}]$ 上 $T$ 满足 $dT/dt = -\sigma_i T$,所以</p>
<p>$$\frac{T(t_{i+1})}{T(t_i)} = e^{-\sigma_i \delta_i}$$</p>
<p>由此<strong>段间 transmittance</strong></p>
<p>$$T_i := T(t_i) = \prod_{j=1}^{i-1} e^{-\sigma_j \delta_j} = \exp\!\Big(\!-\!\sum_{j=1}^{i-1}\sigma_j\delta_j\Big)$$</p>
<p><strong>段内颜色贡献</strong>(积分而非简单矩形):</p>
<p>$$\int_{t_i}^{t_{i+1}} T(t)\sigma_i\mathbf{c}_i\,dt = T_i\,\mathbf{c}_i \int_0^{\delta_i} \sigma_i e^{-\sigma_i s}\,ds = T_i\,\mathbf{c}_i\,(1 - e^{-\sigma_i\delta_i})$$</p>
<p>记 $\alpha_i := 1 - e^{-\sigma_i \delta_i}$<strong>段不透明度</strong>),合并得 NeRF 离散公式:</p>
<p>$$\boxed{\;C(\mathbf{r}) \approx \sum_{i=1}^{N} T_i\,\alpha_i\,\mathbf{c}_i,\quad T_i = \prod_{j<i}(1-\alpha_j),\quad \alpha_i = 1 - e^{-\sigma_i\delta_i}\;}$$</p>
<p>这正是图形学 <strong>front-to-back alpha-compositing</strong> 公式。<strong>关键</strong>$\alpha_i = 1 - e^{-\sigma_i\delta_i}$ 而非 $\sigma_i\delta_i$;当 $\sigma_i\delta_i$ 小时近似相等(一阶 Taylor),但大时差异显著(饱和到 1 vs 线性发散)。</p>
<div class="callout callout-good"><div class="callout-title">物理一致性</div><p>$\sigma \ge 0$ 与 $\alpha = 1 - e^{-\sigma\delta} \in [0, 1)$ 保证颜色合成 <strong>永远在 [0, 1] 内</strong>;且无论 ray 怎么穿,$\sum T_i\alpha_i \le 1$(剩余 $T_{N+1}$ 给背景)。</p></div>
<h3 id="24-位置编码-gammap表示高频细节">2.4 位置编码 $\gamma(p)$:表示高频细节</h3>
<p>MLP 默认是低频偏置(NTK 分析),直接拟合 $f(\mathbf{x})$ 会糊。NeRF 用 <strong>positional encoding</strong> 提频:</p>
<p>$$\gamma(p) = \big(\sin(2^0\pi p),\cos(2^0\pi p),\sin(2^1\pi p),\cos(2^1\pi p),\dots,\sin(2^{L-1}\pi p),\cos(2^{L-1}\pi p)\big)$$</p>
<p>对 $\mathbf{x}$ 用 $L=10$60 维),对 $\mathbf{d}$ 用 $L=4$24 维)。后续 Tancik et al. 2020 "Fourier Features" 给了 NTK 解释:高频 $\sin/\cos$ 让 kernel 衰减更慢,使 MLP 可学高频。</p>
<h3 id="25-hierarchical-sampling粗--细">2.5 Hierarchical Sampling:粗 → 细</h3>
<ul><li><strong>Coarse 网络</strong>:均匀采样 $N_c = 64$ 点,渲染得到 weights $w_i = T_i \alpha_i$</li><li><strong>Fine 网络</strong>:归一化 $w$ 为 PDF,按重要性采 $N_f = 128$ 新点(<strong>重要性采样</strong>:表面附近权重大,应密集采样)</li><li>用 coarse + fine 总共 $N_c + N_f$ 点合成最终颜色</li><li>Loss$\mathcal{L} = \|C_c - C_\text{gt}\|^2 + \|C_f - C_\text{gt}\|^2$(两个网络都监督,coarse 提供 sampler 的 well-defined 梯度)</li></ul>
<h3 id="26-nerf-训练代码核心-30-行">2.6 NeRF 训练代码(核心 30 行)</h3>
<pre><code class="language-python">import torch
import torch.nn as nn
import torch.nn.functional as F
def positional_encoding(x: torch.Tensor, L: int) -&gt; torch.Tensor:
&quot;&quot;&quot; x: [..., D]; 返回 [..., D*2*L]NeRF γ(p) 不含原始 x &quot;&quot;&quot;
freqs = 2.0 ** torch.arange(L, device=x.device, dtype=x.dtype) * torch.pi
args = x.unsqueeze(-1) * freqs # [..., D, L]
pe = torch.stack([torch.sin(args), torch.cos(args)], dim=-1) # [..., D, L, 2]
return pe.flatten(-3) # [..., D*2*L]
def volume_render(sigma: torch.Tensor, color: torch.Tensor, t_vals: torch.Tensor,
ray_d: torch.Tensor):
&quot;&quot;&quot;
NeRF 离散 α-compositing
sigma: [B, N] 体密度 (&gt;= 0; 通常已经过 softplus / ReLU)
color: [B, N, 3] 颜色
t_vals: [B, N] ray 上采样点的 t 值(单调递增)
ray_d: [B, 3] ray 方向(用于把 Δt → 真实距离)
返回 C: [B, 3], weights: [B, N], depth: [B]
&quot;&quot;&quot;
# δ_i = t_{i+1} - t_i ;最后一段补 1e10(吸收掉远端到无穷的剩余)
deltas = t_vals[..., 1:] - t_vals[..., :-1]
delta_far = torch.full_like(deltas[..., :1], 1e10)
deltas = torch.cat([deltas, delta_far], dim=-1) # [B, N]
deltas = deltas * torch.norm(ray_d[:, None, :], dim=-1) # 把 t-间距换成真实欧氏距离
alpha = 1.0 - torch.exp(-sigma * deltas) # [B, N]
# T_i = ∏_{j&lt;i} (1 - α_j) —— 用 cumprodshift 一位让 T_1 = 1
T = torch.cumprod(torch.cat([torch.ones_like(alpha[..., :1]),
1.0 - alpha + 1e-10], dim=-1), dim=-1)[..., :-1]
weights = T * alpha # [B, N]
C = (weights[..., None] * color).sum(dim=-2) # [B, 3]
depth = (weights * t_vals).sum(dim=-1) # [B]
return C, weights, depth</code></pre>
<div class="callout callout-warn"><div class="callout-title">数值陷阱</div><p><code>1 - alpha</code> 在 alpha 接近 1 时会下溢到 0,连乘后 T 整体被 zeroed;加 <code>+ 1e-10</code> 防止 backward 时 <code>log(0)</code>。最后一段 <code>δ → 1e10</code> 强行把背景的 transmittance 推到 0,否则不可见区域的 ray 颜色会受未采样段污染。</p></div>
<h3 id="27-mip-nerf--mip-nerf-360抗锯齿">2.7 Mip-NeRF / Mip-NeRF 360(抗锯齿)</h3>
<p>vanilla NeRF 在低分辨率 / 缩放下走样严重(同一像素对应不同尺度的 cone,但 NeRF 当作 ray 处理)。</p>
<ul><li><strong>Mip-NeRF</strong> (Barron 2021 ICCV):把 ray 视为 cone(视锥),用 <strong>Integrated Positional Encoding (IPE)</strong>——对 cone 段内的 PE 做闭式 Gaussian 期望 $\mathbb{E}_{\mathbf{x}\sim\mathcal{N}(\mu,\Sigma)}[\gamma(\mathbf{x})]$。对频率 $\omega = 2^k\pi$ 而言,$\mathbb{E}[\sin\omega x] = \sin(\omega\mu)\,e^{-\frac{1}{2}\omega^\top\Sigma\,\omega}$<strong>高频系数被 cone 协方差 $\Sigma$ 通过 $e^{-\frac{1}{2}\omega^\top\Sigma\omega}$ 自动衰减</strong>,自然实现 multi-scale。</li><li><strong>Mip-NeRF 360</strong> (Barron 2022 CVPR)unbounded scene 用 contraction $f(x) = (2 - 1/\|x\|)\,x/\|x\|$ for $\|x\| > 1$,把无穷远压缩到 ball;加 distortion / proposal MLP 损失。</li></ul>
<h3 id="28-neus--volsdf体渲染--sdf导出-mesh-的关键">2.8 NeuS / VolSDF:体渲染 + SDF(导出 mesh 的关键)</h3>
<p>NeRF 是 density-based,提取 mesh 需选 $\sigma$ 阈值(不稳)。<strong>NeuS</strong> (Wang 2021 NeurIPS) 用 <strong>SDF $d(\mathbf{x})$</strong> 替换 density</p>
<p>$$\sigma(t) = \max\!\left(\frac{-\frac{d}{dt}\Phi_s(d(\mathbf{r}(t)))}{\Phi_s(d(\mathbf{r}(t)))},\; 0\right),\quad \Phi_s(d) = (1 + e^{-sd})^{-1}$$</p>
<p>其中 $\Phi_s$ 是 sigmoid$s$ 是可学习"sharpness"。性质:表面处($d=0$)权重峰值;可直接 Marching Cubes 提 meshmesh 是 $\{d = 0\}$)。</p>
<p><strong>VolSDF</strong> (Yariv 2021 NeurIPS) 用 Laplace CDF $\sigma = \alpha\,\Phi(-d/\beta)$,思想类似。</p>
<h2 id="3-instant-ngp5-oom-加速必考">§3 Instant-NGP5+ OOM 加速(必考)</h2>
<p>NeRF (vanilla) 训练一个场景要 1-2 天。<strong>Instant-NGP</strong> (Müller 2022 SIGGRAPH <strong>Best Paper</strong>) 5 秒就能拟合一个简单场景。</p>
<h3 id="31-核心-idea多分辨率-hash-网格">3.1 核心 idea:多分辨率 hash 网格</h3>
<p>把"密集网格 vs 大 MLP"换成"<strong>稀疏 hash 网格 + tiny MLP</strong>"。</p>
<ul><li>$L$ 层分辨率(如 $L = 16$),第 $\ell$ 层格点数 $N_\ell = \lfloor N_\min \cdot b^\ell \rfloor$,几何级数($b \approx 1.38$;论文典型取值 $N_\min = 16$$N_\max \in [512, 2048]$ 视场景大小而定)</li><li>每层用 <strong>hash function</strong> 把格点坐标映到固定大小的特征表($T = 2^{14}$$2^{24}$<strong>典型 $T = 2^{19} = 524288$</strong></li><li>查询点 $\mathbf{x}$:对每层做 8 角点三线性插值 → 拼成 $L \times F$ 维特征($F = 2$</li><li>喂给 <strong>tiny MLP</strong> ($2$ 层, hidden 64) 输出 $\sigma, \mathbf{c}$</li></ul>
<h3 id="32-hash-function">3.2 Hash function</h3>
<p>$$\text{hash}(\mathbf{x}) = \bigg(\bigoplus_{i=1}^{d} x_i \cdot \pi_i\bigg) \bmod T$$</p>
<p>$\pi_i$ 是大质数($\pi_1 = 1, \pi_2 = 2654435761, \pi_3 = 805459861$)。 $\oplus$ 是 XOR。这是一种 <strong>spatial hash</strong>:常用于物理仿真的 BVH。</p>
<h3 id="33-hash-collision-怎么消歧l3-高频追问">3.3 Hash collision 怎么消歧?(<strong>L3 高频追问</strong></h3>
<p>当 $N_\ell^d > T$fine level 必然发生),多个格点映到同一表项 → 冲突。Why does it still work?</p>
<ol><li><strong>Multi-resolution 冗余</strong>:粗 level 的特征 unique$N_\ell^d \le T$),细 level 提供补充。MLP 可从粗特征恢复结构,细 level 只负责 detail。</li><li><strong>稀疏性优先</strong>:大部分空间是空(NeRF 场景多数 voxel 是 background),有意义的 query 集中在表面附近,冲突的"有效格点对"很少。</li><li><strong>梯度自动消歧</strong>:训练时只有表面附近格点会有非零梯度(被 ray weights 加权)。空白区的"冲突 entry"得不到梯度信号,不污染表面 entry。</li><li><strong>MLP 后处理</strong>tiny MLP 在 $L \times F$ 拼接特征上学一个分类/回归,遇到冲突的表面点可用 <strong>其他 level 不冲突的特征</strong> disambiguate。</li></ol>
<div class="callout callout-info"><div class="callout-title">面试高分回答</div><p>"Hash collision 看似破坏 unique 性,但<strong>实际生效区域是稀疏的</strong>(场景的 thin surface 仅占 voxel 总数极小比例),冲突区域大概率是无监督信号的 background;即便表面也有冲突,多分辨率层的非冲突特征 + tiny MLP 也能学到一致输出。这是个 <strong>'lazy collision resolution'</strong>:与其代价昂贵地搞 perfect hash,不如用冗余 + 数据驱动消歧。"</p></div>
<h3 id="34-instant-ngp-训练公式">3.4 Instant-NGP 训练公式</h3>
<p>参数:hash table $\theta_\text{hash} \in \mathbb{R}^{L \times T \times F}$ + MLP weights $\theta_\text{MLP}$。Loss 仍是 photometric MSE,但训练 5 秒 vs NeRF 1 天的差别来自:</p>
<ul><li><strong>Tiny MLP</strong>:参数少 100×,前向快 ~50×</li><li><strong>Hash 表</strong>:稀疏激活,cache friendly</li><li><strong>CUDA kernel fuse</strong>tiny-cuda-nn 把 forward + backward 融合</li><li><strong>Occupancy grid</strong>:粗占用网格 skip 空白区采样,避免无效 query</li></ul>
<h3 id="35-plenoxels--tensorf同期的-explicit-派">3.5 Plenoxels / TensoRF(同期的 explicit 派)</h3>
<p><strong>Plenoxels</strong> (Fridovich-Keil 2022 CVPR):纯 voxel 网格 + 球谐 SH 系数 + density<strong>完全没 MLP</strong>,直接梯度下降到 voxel;速度类似 Instant-NGP 但显存大。<strong>TensoRF</strong> (Chen 2022 ECCV):把 4D tensor field 用 <strong>VM / CP 分解</strong> 压缩,参数量从 $O(N^4)$ 降到 $O(N)$ 或 $O(N^2)$。</p>
<h2 id="4-3d-gaussian-splatting显式可微光栅化当前主力">§4 3D Gaussian Splatting:显式可微光栅化(<strong>当前主力</strong></h2>
<p><strong>3DGS</strong> (Kerbl 2023 SIGGRAPH <strong>Best Paper</strong>) 解决了 NeRF 的两大痛:渲染慢、editing 难。</p>
<h3 id="41-场景表示">4.1 场景表示</h3>
<p>场景 = 一组 3D Gaussian $\{G_i\}$,每个:</p>
<ul><li><strong>均值</strong> $\mu_i \in \mathbb{R}^3$(位置)</li><li><strong>协方差</strong> $\Sigma_i \in \mathbb{R}^{3\times 3}$(形状),分解为 $\Sigma = R S S^\top R^\top$rotation $R$ + diag scaling $S$</li><li><strong>不透明度</strong> $\alpha_i \in [0, 1]$</li><li><strong>颜色</strong> $c_i(\mathbf{d})$ 用 SH 系数($\ell = 3$,每色 16 系数,共 48 参数)</li></ul>
<p>为什么用 $R S S R^\top$ 而不直接学 $\Sigma$?— 要保证 $\Sigma$ 正定。直接学 $\Sigma$ 矩阵在梯度下会跑出半正定锥;分解后只需保证 $R$ 正交(用四元数 $q$ 参数化)+ $S$ 正(用 $\exp(s)$ 参数化),自然满足。</p>
<h3 id="42-3d--2d-投影-jacobianl3-必考推导">4.2 3D → 2D 投影 Jacobian<strong>L3 必考推导</strong></h3>
<p>把 3D Gaussian splat 到屏幕上做光栅化,需要把 3D 协方差 $\Sigma$ 投影到 2D 协方差 $\Sigma'$。</p>
<p><strong>Step 1</strong>World → Camera:刚体变换 $W \in SE(3)$。$\Sigma_\text{cam} = W \Sigma W^\top$(这里 $W$ 取旋转部分;平移不影响协方差)。</p>
<p><strong>Step 2</strong>Camera → Screen:透视投影<strong>非线性</strong></p>
<p>$$\pi(\mathbf{x}) = \begin{pmatrix} f_x\,x/z \\ f_y\,y/z \end{pmatrix}$$</p>
<p>非线性映射的协方差近似用一阶 Taylor。在均值 $\mu_\text{cam} = (x, y, z)$ 处求 Jacobian</p>
<p>$$J = \frac{\partial \pi}{\partial \mathbf{x}}\bigg|_{\mu_\text{cam}} = \begin{pmatrix}\dfrac{f_x}{z} & 0 & -\dfrac{f_x\,x}{z^2}\\[2pt] 0 & \dfrac{f_y}{z} & -\dfrac{f_y\,y}{z^2}\end{pmatrix} \in \mathbb{R}^{2\times 3}$$</p>
<p><strong>Step 3</strong>2D 协方差(<strong>核心公式</strong>):</p>
<p>$$\boxed{\;\Sigma' = J\,W\,\Sigma\,W^\top\,J^\top \in \mathbb{R}^{2\times 2}\;}$$</p>
<p>推导:若 $\mathbf{x} \sim \mathcal{N}(\mu, \Sigma)$,则一阶近似 $\pi(\mathbf{x}) \approx \pi(\mu) + J(\mathbf{x} - \mu)$,所以 $\text{Cov}[\pi(\mathbf{x})] \approx J\,\Sigma_\text{cam}\,J^\top = J W \Sigma W^\top J^\top$。这就是 EWA splatting (Zwicker 2001) 的经典推论。</p>
<h3 id="43-可微光栅化tile-based-front-to-back-alpha-blending">4.3 可微光栅化:tile-based front-to-back alpha-blending</h3>
<p>像素 $\mathbf{p}$ 的颜色:</p>
<p>$$C(\mathbf{p}) = \sum_{i \in \mathcal{N}(\mathbf{p})}\,c_i\,\alpha_i\,G_i'(\mathbf{p}) \prod_{j < i}\big(1 - \alpha_j\,G_j'(\mathbf{p})\big)$$</p>
<p>其中 $G_i'(\mathbf{p}) = \exp\!\big(-\tfrac{1}{2}(\mathbf{p} - \mu_i')^\top \Sigma_i'^{-1} (\mathbf{p} - \mu_i')\big)$ 是 2D Gaussian 在像素的值,$\mathcal{N}(\mathbf{p})$ 是覆盖 $\mathbf{p}$ 的 Gaussian 按深度排序。</p>
<p><strong>关键工程</strong></p>
<ol><li><strong>Tile 划分</strong>:屏幕分成 $16\times 16$ tile,每 tile 内 Gaussian 按深度排序,并行渲染</li><li><strong>GPU sort</strong>:用 radix sort,按 <code>(tile_id, depth)</code> 复合键</li><li><strong>Front-to-back early stop</strong>:当累计 $\prod(1 - \alpha G') < 10^{-4}$ 时退出</li><li><strong>CUDA kernel</strong>:作者公开 <code>diff-gaussian-rasterization</code>,前向 + 反向都是 manual derivative</li></ol>
<h3 id="44-3dgs-前向pytorch-reference-实现">4.4 3DGS 前向(PyTorch reference 实现)</h3>
<pre><code class="language-python">def quat_to_rot(q: torch.Tensor) -&gt; torch.Tensor:
&quot;&quot;&quot; q: [N, 4] (w, x, y, z) already normalized; 返回 R: [N, 3, 3] &quot;&quot;&quot;
w, x, y, z = q.unbind(-1)
R = torch.stack([
1 - 2*(y*y + z*z), 2*(x*y - w*z), 2*(x*z + w*y),
2*(x*y + w*z), 1 - 2*(x*x + z*z), 2*(y*z - w*x),
2*(x*z - w*y), 2*(y*z + w*x), 1 - 2*(x*x + y*y),
], dim=-1).reshape(-1, 3, 3)
return R
def gaussian_splat_forward(
means3D: torch.Tensor, # [N, 3] Gaussian 中心 (world)
scales: torch.Tensor, # [N, 3] log-scale (取 exp 得真实 scale)
quats: torch.Tensor, # [N, 4] 四元数 (会归一化)
opacities: torch.Tensor, # [N, 1] σ(logit) → α
colors: torch.Tensor, # [N, 3] (这里简化为 RGB,不展 SH)
viewmat: torch.Tensor, # [4, 4] world→camera
K: torch.Tensor, # [3, 3] 内参 (fx, fy, cx, cy)
H: int, W: int,
):
&quot;&quot;&quot; 教学版前向:不做 tile sort / CUDA,仅展示数学。
实际生产用 gsplat / diff-gaussian-rasterization。 &quot;&quot;&quot;
N = means3D.shape[0]
device = means3D.device
# --- 1. World → Camera ---
homo = torch.cat([means3D, torch.ones(N, 1, device=device)], dim=-1)
mu_cam = (homo @ viewmat.T)[:, :3] # [N, 3]
z = mu_cam[:, 2].clamp(min=1e-4) # 防除零
# --- 2. 协方差 (3D) ---
q = quats / quats.norm(dim=-1, keepdim=True)
R = quat_to_rot(q) # [N, 3, 3]
S = torch.diag_embed(torch.exp(scales)) # [N, 3, 3]
cov3D = R @ S @ S.transpose(-1, -2) @ R.transpose(-1, -2) # [N, 3, 3]
# World→Cam 旋转部分 W_rot (3x3) 应用到协方差
W_rot = viewmat[:3, :3]
cov_cam = W_rot @ cov3D @ W_rot.T # [N, 3, 3]
# --- 3. 投影 Jacobian J (2x3) ---
fx, fy = K[0, 0], K[1, 1]
x_c, y_c, z_c = mu_cam[:, 0], mu_cam[:, 1], z
J = torch.zeros(N, 2, 3, device=device)
J[:, 0, 0] = fx / z_c
J[:, 0, 2] = -fx * x_c / z_c**2
J[:, 1, 1] = fy / z_c
J[:, 1, 2] = -fy * y_c / z_c**2
# --- 4. 2D 协方差 Σ&#x27; = J W Σ W^T J^T ---
cov2D = J @ cov_cam @ J.transpose(-1, -2) # [N, 2, 2]
cov2D = cov2D + 0.3 * torch.eye(2, device=device) # low-pass filter (anti-aliasing)
# --- 5. 2D 中心 (像素坐标) ---
cx, cy = K[0, 2], K[1, 2]
mu2D = torch.stack([fx * x_c / z_c + cx, fy * y_c / z_c + cy], dim=-1) # [N, 2]
# --- 6. 深度排序 (front to back) ---
depth = z_c
order = depth.argsort() # ascending z
mu2D, cov2D = mu2D[order], cov2D[order]
colors_o = colors[order]
alphas = torch.sigmoid(opacities[order]).squeeze(-1) # [N]
# --- 7. 像素遍历 (教学版用全图 loop;真实实现 tile + CUDA) ---
yy, xx = torch.meshgrid(torch.arange(H, device=device),
torch.arange(W, device=device), indexing=&#x27;ij&#x27;)
pix = torch.stack([xx, yy], dim=-1).float() # [H, W, 2]
img = torch.zeros(H, W, 3, device=device)
T_acc = torch.ones(H, W, device=device)
inv_cov = torch.linalg.inv(cov2D) # [N, 2, 2]
for i in range(mu2D.shape[0]):
diff = pix - mu2D[i] # [H, W, 2]
# 2D Gaussian: exp(-0.5 (diff)^T Σ^-1 (diff))
G = torch.exp(-0.5 * (diff @ inv_cov[i] * diff).sum(-1)) # [H, W]
contrib = alphas[i] * G # [H, W]
contrib = contrib.clamp(max=0.99) # 数值
img = img + (T_acc * contrib).unsqueeze(-1) * colors_o[i]
T_acc = T_acc * (1 - contrib)
if (T_acc &lt; 1e-4).all(): # early stop
break
return img</code></pre>
<div class="callout callout-warn"><div class="callout-title">教学版 vs 生产版差距</div><p>上面是 $O(N \cdot HW)$30k Gaussian + 800×800 图就要好几秒。真实 <code>gsplat</code> 是 (a) tile-based: 每 tile 只处理"接触此 tile"的 Gaussian(b) GPU radix sort 复合键;(c) 整个 forward / backward 全 manual CUDA1080p ≤ 10ms。</p></div>
<h3 id="45-自适应密度控制面试常问">4.5 自适应密度控制(<strong>面试常问</strong></h3>
<p>3DGS 初始用 SfM (COLMAP) 稀疏点云,但训练过程要"撒"出更多 Gaussian。</p>
<table><thead><tr><th>触发条件</th><th>操作</th><th>直觉</th></tr></thead><tbody><tr><td><strong>梯度大 + scale 小</strong></td><td><strong>clone</strong>(复制一份,沿梯度方向偏移)</td><td>"under-reconstruction"——这块区域缺细节</td></tr><tr><td><strong>梯度大 + scale 大</strong></td><td><strong>split</strong>(拆成 2 个小 Gaussianscale ÷ 1.6</td><td>"over-reconstruction"——一个大 Gaussian 覆盖了不该覆盖的区域</td></tr><tr><td><strong>opacity 接近 0</strong></td><td><strong>prune</strong>(删除)</td><td>该 Gaussian 没贡献,浪费显存</td></tr><tr><td><strong>每 3k iter</strong></td><td>reset opacities to 0.005</td><td>让 model 重新学透明度,防止 floater</td></tr></tbody></table>
<p>启发式条件:<code>gradient norm &gt; τ_pos</code>(如 $2 \times 10^{-4}$),<code>scale &gt; τ_scale</code>(场景尺度 1%)。</p>
<pre><code class="language-python">def densify_and_prune(gaussians, grad_thresh=2e-4, scale_thresh=0.01,
max_screen_size=None):
&quot;&quot;&quot; 教学版 densify 决策(简化;真实 gsplat 还有 screen-size 触发)。
假设 gaussians 暴露以下 1D 形状的字段(N = 当前 Gaussian 数量):
xyz_grad_accum: [N] ‖累积 xyz 梯度范数‖
denom: [N] 累积次数(防 /0)
scales: [N, 3] log-scale
opacities: [N] sigmoid 后 ∈ (0, 1)
screen_size: [N] 最近一次渲染的屏幕投影大小(可选)
grad_dir: [N, 3] 最近一次梯度方向(用于 clone offset
&quot;&quot;&quot;
grad_norm = gaussians.xyz_grad_accum / gaussians.denom.clamp(min=1) # [N]
mean_scale = gaussians.scales.exp().max(dim=-1).values # [N]
# CLONE:高梯度 + 小 scale —— 复制一份并沿梯度方向偏移;原始保留
clone_mask = (grad_norm &gt; grad_thresh) &amp; (mean_scale &lt;= scale_thresh) # [N]
gaussians.clone_at(clone_mask, offset=gaussians.grad_dir[clone_mask])
# SPLIT:高梯度 + 大 scale —— 拆成 2 个子高斯(scale ÷ 1.6),并在末尾删除原始
split_mask = (grad_norm &gt; grad_thresh) &amp; (mean_scale &gt; scale_thresh) # [N]
gaussians.split_at(split_mask, n=2, scale_div=1.6)
# PRUNE:低 opacity / 屏幕过大 / 已被 split 标记
# ⚠️ 注意:clone 增加的新 Gaussian 已 append 到末尾,长度变了;这里的 mask 仅作用于原 N 个
prune_mask = (gaussians.opacities[:split_mask.shape[0]] &lt; 0.005) | split_mask
if max_screen_size is not None:
prune_mask = prune_mask | (gaussians.screen_size[:split_mask.shape[0]] &gt; max_screen_size)
gaussians.remove_original(prune_mask) # 只删除原始 N 个里被标记的
gaussians.reset_grad_accum()
return gaussians</code></pre>
<div class="callout callout-info"><div class="callout-title">典型超参数</div><p>Kerbl 2023 paper: densify every 100 iters;最大 Gaussian 数 5e6;总训练 30k iter;约 30 分钟到几小时。</p></div>
<h3 id="46-2dgs--surfelssurface-aligned">4.6 2DGS / Surfelssurface-aligned</h3>
<p>3DGS 的 ellipsoid 不是 surface-aware;提 mesh 需 SuGaR / GSDF 后处理。<strong>2DGS</strong> (Huang 2024 SIGGRAPH) 把 3D ellipsoid <strong>退化成 2D disk</strong>(一个 axis = 0),直接对齐表面,更适合 normal / depth 监督,提 mesh 时质量明显更好。</p>
<h3 id="47-动态-4dgs">4.7 动态 4DGS</h3>
<p><strong>Dynamic 3DGS</strong> (Luiten 2024 3DV) 每帧独立 Gaussian + 物理 prior 连接;<strong>4DGS</strong> (Wu 2024 CVPR / Yang 2024 ICLR) 把 $\mu(t), \Sigma(t)$ 写成时间函数(MLP 或 spline);<strong>SC-GS</strong> (Huang 2024) 用 sparse control points 驱动密集 Gaussian(类似 LBS)。</p>
<h2 id="5-mesh-提取marching-cubes--dmtet">§5 Mesh 提取:Marching Cubes / DMTet</h2>
<p>NeRF / 3DGS 重建后,下游(仿真器、AR、3D 打印)经常需要 mesh。</p>
<h3 id="51-marching-cubes经典必考">5.1 Marching Cubes<strong>经典必考</strong></h3>
<p>输入:3D 标量场 $f(\mathbf{x})$density / SDF+ 阈值 $\tau$。输出:水平集 $\{f = \tau\}$ 的三角网格。</p>
<p><strong>算法骨架</strong></p>
<ol><li>把空间 voxel 化(每 voxel 8 角点)</li><li>对每个 voxel<strong>8 角点二值化</strong>$f > \tau$ 为 1, 否则 0)→ 256 种可能配置</li><li>查 lookup table:每个配置预定义了几条等值面三角片 + edge 上的顶点位置</li><li><strong>线性插值</strong> 找精确顶点:在 edge 两端 $\mathbf{a}, \mathbf{b}$ 之间,插值 $t = (\tau - f(\mathbf{a})) / (f(\mathbf{b}) - f(\mathbf{a}))$,顶点 $= \mathbf{a} + t(\mathbf{b} - \mathbf{a})$</li><li>合并所有 voxel 三角片 → 完整 mesh</li></ol>
<pre><code class="language-python">def marching_cubes_sketch(density: torch.Tensor, threshold: float):
&quot;&quot;&quot; 真实实现用 mcubes / scikit-image / pytorch3d;这里写思路 &quot;&quot;&quot;
from skimage.measure import marching_cubes
# density: [Nx, Ny, Nz]detach() 切断 autograd 图,cpu().numpy() 转 host
verts, faces, normals, _ = marching_cubes(
density.detach().cpu().numpy(),
level=threshold,
spacing=(1.0, 1.0, 1.0),
gradient_direction=&#x27;descent&#x27;, # 法线方向;descent = surface 朝低密度
)
return verts, faces, normals</code></pre>
<div class="callout callout-warn"><div class="callout-title">NeRF 提 mesh 坑</div><p>vanilla NeRF 没有"表面"概念,提 mesh 时阈值 $\tau$ 难选,且 floater 会被一起提出来。<strong>用 NeuS / VolSDF 提 mesh 才稳</strong>SDF 0 等值面定义良好)。</p></div>
<h3 id="52-differentiable-dmtet--flexicubes端到端学-mesh">5.2 Differentiable: DMTet / FlexiCubes(端到端学 mesh</h3>
<p>Marching Cubes 不可微(lookup table 离散)。</p>
<ul><li><strong>DMTet</strong> (Shen 2021 NeurIPS / Munkberg 2022 CVPR):用 <strong>deformable tetrahedral grid</strong>,每四面体 4 顶点 SDF + 位置 offset 可微。<strong>Marching Tetrahedra</strong> 替代 MCtopology 由 SDF sign 决定,几何由顶点位置决定,<strong>全程可微</strong></li><li><strong>FlexiCubes</strong> (Shen 2023 SIGGRAPH):泛化 dual marching cubes,引入额外可学参数(dual vertex offset / interpolation weight)解决 quality artifacts。</li></ul>
<p><strong>典型用法</strong>DreamFusion 之后的 Magic3D、Fantasia3D 用 DMTet 在 SDS 监督下学 mesh + texture。</p>
<h2 id="6-sds-loss用-2d-diffusion-监督-3ddreamfusion-系列">§6 SDS Loss:用 2D Diffusion 监督 3DDreamFusion 系列)</h2>
<h3 id="61-问题设置">6.1 问题设置</h3>
<p>我们想生成 3D 资产但 <strong>没有 3D 训练数据</strong>——3D 数据稀缺(ShapeNet ~5万件,Objaverse-XL 1000万件但质量参差)。<strong>Pretrained 2D diffusion</strong>Stable Diffusion, Imagen)海量。能否用 2D diffusion 当老师 supervise 3D</p>
<p><strong>DreamFusion</strong> (Poole et al. 2022 arXiv → <strong>ICLR 2023 Outstanding Paper</strong>) 提出 <strong>Score Distillation Sampling (SDS)</strong></p>
<h3 id="62-setup">6.2 Setup</h3>
<ul><li>3D 表示 $\theta$NeRF 参数 / DMTet vertices / 3DGS 点云)</li><li>可微渲染器 $g(\theta, \pi) \to x$$x$ 是图像($\pi$ 是相机视角)</li><li>Pretrained 2D diffusion $\epsilon_\phi(x_t; y, t)$$y$ 是文本 prompt</li></ul>
<p><strong>目标</strong>:让 $g(\theta, \pi)$ 看起来像"$y$ 的 photo",即 $g(\theta, \pi)$ 落在 diffusion 学到的数据流形上。</p>
<h3 id="63-sds-gradient-推导l3-必考">6.3 SDS gradient 推导(<strong>L3 必考</strong></h3>
<p>直觉:用 diffusion training loss 反传到 $\theta$。<strong>Naive 想法</strong>:把渲染图 $x = g(\theta, \pi)$ 当训练样本,最小化</p>
<p>$$\mathcal{L}_\text{diff}(\theta) = \mathbb{E}_{t, \epsilon}\Big[w(t)\big\|\epsilon_\phi(x_t; y, t) - \epsilon\big\|^2\Big],\quad x_t = \alpha_t x + \sigma_t \epsilon$$</p>
<p>对 $\theta$ 求梯度(链式法则):</p>
<p>$$\nabla_\theta \mathcal{L}_\text{diff} = \mathbb{E}\Big[w(t)\,2\big(\epsilon_\phi(x_t; y, t) - \epsilon\big)\,\underbrace{\frac{\partial \epsilon_\phi(x_t;y,t)}{\partial x_t}}_{\text{U-Net Jacobian}}\,\alpha_t\,\underbrace{\frac{\partial x}{\partial \theta}}_{\text{renderer Jacobian}}\Big]$$</p>
<p><strong>问题</strong>U-Net Jacobian $\partial \epsilon_\phi / \partial x_t$ 计算昂贵且数值差(diffusion 模型大且未训练 second-order 稳定)。</p>
<p><strong>SDS trick:直接扔掉 U-Net Jacobian</strong>,得到:</p>
<p>$$\boxed{\;\nabla_\theta \mathcal{L}_\text{SDS} \;=\; \mathbb{E}_{t, \epsilon}\Big[w(t)\,\big(\epsilon_\phi(x_t; y, t) - \epsilon\big)\,\frac{\partial x}{\partial \theta}\Big]\;}$$</p>
<p>(原 DreamFusion 论文写成 $\partial L/\partial \theta$ 形式;$\alpha_t$ 与常数 2 被吸收到 $w(t)$ 里。)</p>
<h3 id="64-为什么扔掉-jacobian-反而-work">6.4 为什么扔掉 Jacobian 反而 work</h3>
<p><strong>第一种解释(DreamFusion 原版,score 视角)</strong>$\epsilon_\phi(x_t; y, t)/\sigma_t \approx -\nabla_{x_t}\log p_\phi(x_t|y)$score)。 SDS gradient = <code>(predicted score - noise) × renderer Jacobian</code>,相当于把渲染图朝高概率方向推。</p>
<p><strong>第二种解释(mode-seeking</strong>SDS 等价最小化 $\mathbb{E}_t[D_\text{KL}(q(x_t|\theta) \,\|\, p_\phi(x_t|y))]$ 的某种 mode-seeking 形式:往 $p_\phi(\cdot|y)$ 的高概率区跑。</p>
<h3 id="65-sds-副作用over-saturation--mode-collapse--janus">6.5 SDS 副作用:over-saturation / mode collapse / Janus</h3>
<ul><li><strong>Over-saturation</strong>:颜色饱和,对比度过高("plastic-y look"</li><li><strong>Over-smoothing</strong>:细节糊</li><li><strong>Mode collapse</strong>:物体趋于"canonical" 单一形式</li><li><strong>Janus problem</strong>:3D 物体在不同视角都长出一张"前脸"(人脸出现在头后;动物两边都是头)</li></ul>
<p><strong>根因</strong>SDS 等价 mode-seeking KL,加上<strong>大 CFG 系数(DreamFusion 默认 100)才能逃出 mean-mode</strong>。CFG=100 把分布锐化到极端 → over-saturation。</p>
<div class="callout callout-warn"><div class="callout-title">面试高分点</div><p>SDS 公式形式上"丢掉 Jacobian"省了计算,但<strong>代价</strong>是隐式变成 mode-seeking KL,需要超大 CFG 来缓解 mean-seeking blur;超大 CFG 又导致 over-saturation。这是个<strong>信息论 trade-off</strong>simulation-free + computationally cheap = mode-seeking artifact。</p></div>
<h3 id="66-sds-代码核心-30-行">6.6 SDS 代码(核心 30 行)</h3>
<pre><code class="language-python">def sds_loss(
renderer, # θ → x (B, 3, H, W)
theta, # 3D 参数 (NeRF / 3DGS / DMTet)
prompt_emb, # 文本 embedding (cond) [B, L, D]
uncond_emb, # 文本 embedding (uncond / null) [B, L, D]
unet, # frozen 2D diffusion U-Net (e.g. SD); 返回 noise pred Tensor
alpha_cumprod: torch.Tensor, # [T_max] precomputed bar-alpha schedule
cfg_scale: float = 100.0,
t_range: tuple = (0.02, 0.98),
):
&quot;&quot;&quot; Score Distillation Sampling loss (DreamFusion).
约定: unet(x_t, t, encoder_hidden_states=emb) -&gt; [B, 3, H, W] noise pred.
如果用 diffusers 的 UNet2DConditionModel, 包一层取 .sample 即可。
返回的是 grad surrogate, 直接 backward 即可。 &quot;&quot;&quot;
x = renderer(theta) # [B, 3, H, W]
B = x.shape[0]
device, dtype = x.device, x.dtype
T_max = alpha_cumprod.shape[0] # 通常 1000
# 1. 采样 t 和 noiseforward 加噪
t = torch.randint(int(t_range[0] * T_max), int(t_range[1] * T_max),
(B,), device=device)
noise = torch.randn_like(x)
abar = alpha_cumprod.to(device=device, dtype=dtype)[t].view(B, 1, 1, 1)
x_t = abar.sqrt() * x + (1 - abar).sqrt() * noise
# 2. U-Net 预测 noise (cond / uncond)CFG 组合;关键:不对 diffusion 求导
with torch.no_grad():
eps_uncond = unet(x_t, t, encoder_hidden_states=uncond_emb)
eps_cond = unet(x_t, t, encoder_hidden_states=prompt_emb)
eps_pred = eps_uncond + cfg_scale * (eps_cond - eps_uncond)
# 3. SDS gradient: w(t)(ε_pred - ε) · ∂x/∂θ; w(t) = σ_t² 是常见选择
grad = ((1 - abar) * (eps_pred - noise)).detach()
# backward of (grad · x) 给出 grad · ∂x/∂θ
return (grad * x).sum() / B</code></pre>
<div class="callout callout-info"><div class="callout-title">训练 loop</div><p>每 iter 随机采视角 $\pi$,渲染 $x$,算 SDS loss,反传到 $\theta$。NeRF 表示要训 10k-100k 步(GPU 几小时);3DGS 表示(GaussianDreamer / DreamGaussian)几分钟到一小时。</p></div>
<h3 id="67-vsd变分-sdsprolificdreamer-neurips-2023-spotlight">6.7 VSD:变分 SDS<strong>ProlificDreamer</strong>, NeurIPS 2023 Spotlight</h3>
<p><strong>VSD</strong> (Wang 2023 NeurIPS) 把 SDS 视为"对 single $\theta$ 做点估计"的特殊情况,泛化为<strong>对 $\theta$ 分布 $\mu(\theta)$ 做变分推断</strong></p>
<h4 id="setup">Setup</h4>
<ul><li>把 3D 参数 $\theta$ 当 latent random variable$\mu(\theta)$ 是其分布</li><li>目标:让 rendered image distribution 与 diffusion 学到的 prior <strong>分布</strong>对齐(不是 mode 对齐)</li></ul>
<h4 id="objective-and-gradient">Objective and gradient</h4>
<p>ProlificDreamer 把目标写成 KL</p>
<p>$$\min_{\mu}\; D_\text{KL}\!\Big(q_\mu^t(x_t|y)\;\Big\|\;p_\phi^t(x_t|y)\Big),\quad t\sim\mathcal{U}[0,1]$$</p>
<p>其中 $q_\mu^t$ 是 "从 $\theta\sim\mu$ 渲染 + 加噪到 $t$" 诱导的分布。<strong>变分梯度</strong>Wang et al. 2023, Theorem 2 略写)给出关于 $\theta$ 的更新方向为 <strong>relative score</strong></p>
<p>$$\boxed{\;\nabla_\theta \mathcal{L}_\text{VSD} \;=\; \mathbb{E}_{t,\epsilon}\Big[\,w(t)\,\big(\epsilon_\phi(x_t;y,t) \;-\; \epsilon_\psi(x_t;y,t,\pi)\big)\,\frac{\partial x}{\partial \theta}\,\Big]\;}$$</p>
<p>对比 SDS:把 raw noise $\epsilon$ 替换成<strong>辅助 score</strong> $\epsilon_\psi$。$\epsilon_\psi$ 是 <strong>LoRA 微调</strong> 的 score network,在线最小化 score-matching loss 来跟踪当前 $q_\mu^t$ 的 score;它扮演的是 "variance-reduction baseline" 的角色(与 RL actor-critic 的 value baseline 同构)。<strong>注意</strong>:上面是 gradient form,不是平方-loss form;论文没有"先写一个 $\|\epsilon_\phi-\epsilon_\psi\|^2$ 标量 loss 再求导"的可实现形式——$\epsilon_\psi$ 依赖 $\mu$,那样会丢掉 KL 的关键项。</p>
<h4 id="直觉为什么缓解-over-saturation">直觉为什么缓解 over-saturation</h4>
<ul><li>SDS:把 rendered $x$ 推向 prior $p_\phi(\cdot|y)$ 的 modemean-mode → 需 CFG=100 → over-saturation</li><li>VSD:用辅助 score 学当前 rendered 分布的"自己的 mode",更新方向 = 从"我现在在哪"指向"prior 在哪"<strong>relative gradient</strong>),不需大 CFG 拉满。CFG 可降到 7.5(常规 diffusion 默认值),avoid extreme sharpening</li><li>实验上:VSD 颜色更自然,几何更复杂,可同时维护多个 modeProlificDreamer 给出 50k 步训练得 photorealistic Buddha 等)</li></ul>
<div class="callout callout-good"><div class="callout-title">VSD vs SDS 的关键认知</div><p>SDS 是 "<strong>single-point + mode-seeking</strong>"VSD 是 "<strong>particle / variational + relative score</strong>"。后者本质是给 SDS 加了个<strong>可学习 baseline</strong>$\epsilon_\psi$)减方差,思想上类比 RL actor-critic 的 value baseline。</p></div>
<h3 id="68-sds-衍生家族mesh--3dgs--sds">6.8 SDS 衍生家族:mesh / 3DGS + SDS</h3>
<table><thead><tr><th>方法</th><th>表示 / Stages</th><th>关键点</th></tr></thead><tbody><tr><td><strong>DreamFusion</strong></td><td>NeRF + SDS @ low-res</td><td>原版,提 mesh 难</td></tr><tr><td><strong>Magic3D</strong> (Lin 2023 CVPR)</td><td>Instant-NGP @ 64px → DMTet + SDS @ 512px</td><td><strong>两阶段</strong>:粗结构 → 高分辨率端到端 mesh</td></tr><tr><td><strong>Fantasia3D</strong> (Chen 2023 ICCV)</td><td>DMTet 几何 + PBR material</td><td>normal-as-input + 物理材质 BRDF</td></tr><tr><td><strong>DreamGaussian</strong> (Tang 2024 ICLR)</td><td>3DGS + SDS, ~2 分钟 / 物体</td><td>GPU 速度优势;mesh export + UV-Net texturing</td></tr><tr><td><strong>GaussianDreamer</strong> (Yi 2024 CVPR)</td><td>Point-E / Shap-E init → 3DGS + SDS</td><td>缓解 from-scratch 几何混乱</td></tr></tbody></table>
<h2 id="7-single-image--few-view-3d-生成">§7 Single-Image / Few-View 3D 生成</h2>
<p>更实用的设定:<strong>给一张图,生成 3D</strong></p>
<h3 id="71-zero-1-to-3-范式novel-view-via-diffusion">7.1 Zero-1-to-3 范式(novel view via diffusion</h3>
<p><strong>Zero-1-to-3</strong> (Liu 2023 ICCV):用 Objaverse 上 finetune Stable Diffusion,让它接收 (input view, target camera) → output novel view。</p>
<ul><li>Input:单图 $x$ + 相对相机位姿 $\Delta R, \Delta T$</li><li>Diffusion conditioningimage embedding (CLIP) + camera embedding (sinusoidal)</li><li>Output:在 $\Delta R, \Delta T$ 视角下的图</li></ul>
<p><strong>用法</strong>:给一个 input viewsample 16-32 个 novel view,再用 NeRF / 3DGS 重建。</p>
<p><strong>衍生</strong></p>
<ul><li><strong>Zero-1-to-3++</strong> (Shi 2023):固定生成 6 个 anchor view(北极视角 + 4 平视角 + 俯视角),减少 randomness</li><li><strong>SyncDreamer</strong> (Liu 2024 ICLR):在 latent 上<strong>联合</strong>预测多视图(cross-attention 让 views 看到彼此),保证 3D 一致</li><li><strong>MVDream</strong> (Shi 2024 ICLR)text-to-multi-view4 视图同时生成;后接 SDS 精化</li></ul>
<h3 id="72-one-2-3-45--instantmesh--triposr--stable-fast-3d">7.2 One-2-3-45 / InstantMesh / TripoSR / Stable Fast 3D</h3>
<table><thead><tr><th>方法</th><th>输入</th><th>输出</th><th>速度</th><th>关键</th></tr></thead><tbody><tr><td><strong>One-2-3-45</strong> (Liu 2023 NeurIPS)</td><td>单图</td><td>mesh</td><td>45 秒</td><td>Zero-1-to-3 → SparseNeuS</td></tr><tr><td><strong>One-2-3-45++</strong> (Liu 2024)</td><td>单图</td><td>mesh</td><td>60 秒</td><td>多视图 + SDF</td></tr><tr><td><strong>TripoSR</strong> (Tochilkin 2024, Stability+Tripo)</td><td>单图</td><td>NeRF/mesh</td><td>0.5-2 秒</td><td>LRM (Large Reconstruction Model) 风格 transformer</td></tr><tr><td><strong>InstantMesh</strong> (Xu 2024)</td><td>单图</td><td>mesh</td><td>3 秒</td><td>Zero-1-to-3++ 多视图 → sparse-view recon transformer</td></tr><tr><td><strong>Stable Fast 3D</strong> (SF3D, Stability 2024)</td><td>单图</td><td>textured mesh</td><td>~0.5 秒</td><td>TripoSR 后继;加 illumination disentangle + UV unwrap</td></tr></tbody></table>
<p><strong>LRM (Hong et al. 2023 arXiv → ICLR 2024) 设定</strong>:把图当 token + Plucker ray embeddingtransformer 输出 NeRF triplane。这是 TripoSR / InstantMesh 的母模型。</p>
<h3 id="73-lrm-triplane-表示面试高频">7.3 LRM Triplane 表示(<strong>面试高频</strong></h3>
<ul><li><strong>Triplane</strong> (Chan 2022 EG3D)3 个轴对齐 2D 平面(XY, YZ, XZ),共 $3 \times C \times N \times N$ 维</li><li>查询 3D 点 $(x, y, z)$:在每个平面双线性插值 → concat → 小 MLP → $(\sigma, \mathbf{c})$</li><li>优点:比 voxel grid 显存少($O(N^2)$ vs $O(N^3)$),比 hash grid 更 dense 适合 transformer 输出</li><li>LRM / TripoSR / InstantMesh 都让 transformer 直接 regress triplane tokens</li></ul>
<h2 id="8-3d-foundation-models2024-开源浪潮">§8 3D Foundation Models2024 开源浪潮)</h2>
<h3 id="81-trellis-microsoft-2024-开源">8.1 Trellis (Microsoft 2024, 开源)</h3>
<p><strong>Trellis</strong> (Xiang 2024 arXiv) 是首个尝试做"3D 的 Stable Diffusion"开源工作。</p>
<ul><li><strong>Structured Latent (SLAT)</strong>:把 3D 资产编码到 voxel 上的稀疏 latent grid——既保留空间结构(适合 sparse conv / sparse attention),又紧凑(仅 active voxel 存 latent</li><li><strong>3D VAE</strong>:把 mesh + texture (signed distance field 派生) → SLAT</li><li><strong>Flow matching prior</strong>:在 SLAT 上跑 rectified flowconditioned on text/image</li><li><strong>多 decoder</strong>:从 SLAT decode 出 NeRF / 3DGS / mesh 三种表示(同一 latent,可选输出格式)</li><li><strong>训练数据</strong>Objaverse-XL 子集 + 内部高质量集</li><li><strong>效果</strong>text-to-3D / image-to-3D,几秒到几十秒,质量超过 SDS 系列</li></ul>
<h3 id="82-hunyuan3d-1---2-tencent-2024-25-开源">8.2 Hunyuan3D-1 / -2 (Tencent 2024-25, 开源)</h3>
<p><strong>Hunyuan3D</strong><strong>shape-then-texture</strong> 两阶段路线。</p>
<ul><li><p><strong>Hunyuan3D-1</strong> (Yang 2024 arXiv)</p>
<ul><li>Stage 1: text/image → multi-view image (Zero-1-to-3 系)</li><li>Stage 2: multi-view → 3D mesh (LRM-like reconstructor)</li><li>几秒到几十秒输出 textured mesh</li></ul></li><li><p><strong>Hunyuan3D-2</strong> (Tencent 2025, arXiv 2501.12202)</p>
<ul><li><strong>Hunyuan3D-DiT</strong>geometry-only DiT 在 SDF latent 上生成 mesh</li><li><strong>Hunyuan3D-Paint</strong>multi-view PBR texture diffusionUV space refinement</li><li>高质量 PBR texture(实战可用于游戏 / VR 资产)</li></ul></li><li><strong>开源</strong>HuggingFace 上完整权重 + 推理代码</li></ul>
<h3 id="83-clay-zhang-2024-siggraph">8.3 CLAY (Zhang 2024 SIGGRAPH)</h3>
<ul><li><strong>3DShape2VecSet</strong> latent diffusion:把 mesh 表示为 vector set + cross-attention DiT</li><li>大规模训练(Objaverse-XL + 内部清洗集)</li><li>输出 SDF → marching cubes → mesh</li><li>加 PBR texture stage(类似 Hunyuan3D-2</li></ul>
<p><strong>Rodin</strong> (Microsoft 2023, 商业):早期 text-to-3D-avatar 产品级系统,diffusion on triplane,主打 character / avatar。</p>
<h3 id="84-对比表">8.4 对比表</h3>
<table><thead><tr><th>方法</th><th>表示</th><th>Prior</th><th>训练规模</th><th>开源</th></tr></thead><tbody><tr><td><strong>Trellis</strong></td><td>Structured Latent (SLAT) + 多 decoder</td><td>Rectified Flow</td><td>Objaverse-XL 子集</td><td></td></tr><tr><td><strong>Hunyuan3D-2</strong></td><td>SDF latent (Shape DiT) + UV texture diff</td><td>Diffusion</td><td>内部大规模集</td><td></td></tr><tr><td><strong>CLAY</strong></td><td>3DShape2VecSet</td><td>Diffusion</td><td>Objaverse-XL + 内部</td><td>部分</td></tr><tr><td><strong>Rodin</strong></td><td>Triplane</td><td>Diffusion</td><td>商业内部</td><td></td></tr><tr><td><strong>TripoSR / SF3D</strong></td><td>NeRF/mesh feedforward</td><td>无 prior,纯 regression</td><td>Objaverse 类</td><td></td></tr></tbody></table>
<div class="callout callout-info"><div class="callout-title">架构选择直觉</div><p>大 scene / general object 用 <strong>Trellis 风格 SLAT</strong>(保留空间结构);高质量 single mesh 用 <strong>CLAY 风格 vector set</strong>(紧凑、global attention);快速推理用 <strong>LRM/TripoSR feedforward</strong>(不做 diffusion,直接 regress)。</p></div>
<h2 id="9-复杂度--资源对比">§9 复杂度 / 资源对比</h2>
<table><thead><tr><th>方法</th><th>训练</th><th>推理 (一帧)</th><th>显存 (训练)</th><th>显存 (模型)</th></tr></thead><tbody><tr><td>NeRF vanilla</td><td>1-2 天</td><td>数秒</td><td>8 GB</td><td>&lt;10 MB MLP</td></tr><tr><td>Instant-NGP</td><td>5 秒 - 5 分钟</td><td>30 fps+</td><td>4-12 GB</td><td>100-500 MB hash</td></tr><tr><td>3DGS</td><td>10-30 分钟</td><td>100 fps+</td><td>6-24 GB</td><td>100 MB - 1 GB Gaussian</td></tr><tr><td>2DGS</td><td>与 3DGS 接近</td><td>与 3DGS 接近</td><td>类似</td><td>类似</td></tr><tr><td>DreamFusion (NeRF+SDS)</td><td>2 hr / 物体</td><td></td><td>12 GB</td><td>NeRF 本身</td></tr><tr><td>DreamGaussian (3DGS+SDS)</td><td>2 分钟 / 物体</td><td></td><td>8-16 GB</td><td></td></tr><tr><td>ProlificDreamer (VSD)</td><td>3-6 hr / 物体</td><td></td><td>24 GB</td><td></td></tr><tr><td>TripoSR feedforward</td><td>训练 50 GPU 天</td><td>0.5 秒 (A100)</td><td>inference 6 GB</td><td>1.5 GB</td></tr><tr><td>Trellis</td><td>训练 100+ GPU 天</td><td>数秒</td><td>inference 16 GB</td><td>数 GB</td></tr><tr><td>Hunyuan3D-2</td><td>训练大集群</td><td>数十秒</td><td>inference 24+ GB</td><td>多模型组合</td></tr></tbody></table>
<h2 id="10-与相关方法对比--embodied-ai-应用">§10 与相关方法对比 &amp; Embodied AI 应用</h2>
<h3 id="101-3d-vs-2d-生成关键区别">10.1 3D-vs-2D 生成关键区别</h3>
<table><thead><tr><th>维度</th><th>2D 生成 (Stable Diffusion)</th><th>3D 生成</th></tr></thead><tbody><tr><td><strong>数据量</strong></td><td>LAION-5B 50亿图</td><td>Objaverse-XL 1000万件(小 500×)</td></tr><tr><td><strong>数据格式</strong></td><td>图像(统一 RGB</td><td>mesh / SDF / point cloud / NeRF / 3DGS<strong>碎片化</strong></td></tr><tr><td><strong>训练 prior</strong></td><td>直接 train diffusion</td><td>用 2D diffusion 蒸馏 (SDS / Zero-1-to-3) <strong></strong> 用 3D-native diffusion (Trellis / CLAY)</td></tr><tr><td><strong>评测</strong></td><td>FID, CLIP score</td><td>Chamfer / IoU / PSNR (recon) + perceptual + user study</td></tr><tr><td><strong>下游</strong></td><td>直接出图</td><td>出资产 → 渲染 / 仿真 / 编辑</td></tr></tbody></table>
<h3 id="102-embodied-ai--ar--vr-实战路线">10.2 Embodied AI / AR / VR 实战路线</h3>
<table><thead><tr><th>任务</th><th>推荐表示</th><th>关键工具链 / 约束</th></tr></thead><tbody><tr><td><strong>Sim2Real 资产</strong></td><td>mesh (PBR)</td><td>Trellis / Hunyuan3D-2 → IsaacSim / MuJoCo</td></tr><tr><td><strong>室内大场景</strong></td><td>3DGS</td><td>COLMAP → 3DGSchunk-wise 用 VastGS / CityGS</td></tr><tr><td><strong>NeRF/3DGS as simulator</strong></td><td>NeRF / 3DGS + physics</td><td>DreamGaussian-Sim / Splatting Physics</td></tr><tr><td><strong>3D affordance / manipulation</strong></td><td>point cloud / 3DGS feature</td><td>OpenScene / LERF / RVT / 3D Diffuser Actor</td></tr><tr><td><strong>AR 物体扫描</strong></td><td>3DGS(光照真实 + 实时)</td><td>mobile 算力(PostShot / Luma),剪枝 / 量化</td></tr><tr><td><strong>VR 大场景</strong></td><td>3DGS (large-scale)</td><td>60 fps stereo + 6DoF</td></tr><tr><td><strong>Avatar</strong></td><td>mesh + LBS 或 3DGS avatar</td><td>实时表情 / 头发</td></tr><tr><td><strong>Object insertion</strong></td><td>mesh + PBR</td><td>环境光照一致(IBL</td></tr></tbody></table>
<div class="callout callout-warn"><div class="callout-title">Embodied AI 面试追问示例</div><p>"做 NeRF 物理仿真器最大挑战?" 要点:NeRF 是 radiance,没 mass / friction → 需手动叠物理 priormesh 提取有 floater → 碰撞检测难;可微但 backward 慢;<strong>业界更多用 3DGS / mesh 而非 vanilla NeRF</strong></p></div>
<h2 id="11-工程实战--易踩坑">§11 工程实战 &amp; 易踩坑</h2>
<h3 id="111-colmap--sfm-前处理重建必经">11.1 COLMAP / SfM 前处理(重建必经)</h3>
<p>输入多视图 → 输出内参 $K$ + 外参 $\{R_i, t_i\}$ + 稀疏点云;标准流程 SIFT → matching → incremental SfM → bundle adjustment。<strong>常见坑</strong>texture-less / 镜面物体 SfM 失败;动态物体污染外参。</p>
<h3 id="112-数值稳定nerf3dgs-通用">11.2 数值稳定(NeRF/3DGS 通用)</h3>
<table><thead><tr><th>问题</th><th>症状</th><th>修复</th></tr></thead><tbody><tr><td>Sigma 爆炸</td><td>floater 充斥空间</td><td>$\sigma$ 用 softplus 或 truncatedoccupancy grid skip</td></tr><tr><td>Alpha 饱和</td><td>1-α 下溢 → T 全 0</td><td><code>(1-α).clamp(min=1e-10)</code> 或 log-space cumprod</td></tr><tr><td>Gaussian 退化</td><td>极小 scale / 极大 anisotropy</td><td>clamp scale lower boundregularize anisotropy</td></tr><tr><td>Densify 爆炸</td><td>Gaussian 数量飙到内存上限</td><td>加 max gaussian 数;周期 prunereset opacity</td></tr><tr><td>SDS Janus</td><td>多视角脸 / 头</td><td>加 view-conditioning"front view" / "back view");MVDream</td></tr><tr><td>SDS over-sat</td><td>颜色饱和</td><td>CFG 降低;改用 VSD;或 negative prompt</td></tr></tbody></table>
<h3 id="113-多机分布式--评测指标">11.3 多机分布式 &amp; 评测指标</h3>
<p><strong>分布式</strong>NeRF / Instant-NGP / 3DGS 单 GPU 标准;大场景 3DGS 用 chunk-wise (VastGaussian, CityGaussian)SDS/VSD 每 iter 跑 2 次 SD forward8×A100 可显著提速;Trellis / Hunyuan3D 训练是大规模 multi-node DDP。</p>
<table><thead><tr><th>评测指标</th><th>用途</th><th>算法</th></tr></thead><tbody><tr><td><strong>PSNR / SSIM / LPIPS</strong></td><td>视图合成(重建)</td><td>与真实视图对比</td></tr><tr><td><strong>Chamfer Distance</strong></td><td>mesh 几何</td><td>两点云最近邻距离平均</td></tr><tr><td><strong>F-Score (3D)</strong></td><td>mesh / point</td><td>precision + recall under threshold</td></tr><tr><td><strong>CLIP Score / CLIP-R-Prec</strong></td><td>text-to-3D 对齐</td><td>render → CLIP 相似度 / 区分干扰 prompt</td></tr><tr><td><strong>User study</strong></td><td>最终质量</td><td>MTurk / lab-internal</td></tr></tbody></table>
<h2 id="12-25-高频面试题">§12 25 高频面试题</h2>
<p>按难度分 3 档(L1 必会 / L2 进阶 / L3 顶级 lab)。每题点开看答案要点 + 易踩坑。</p>
<h3 id="l1-必会题任何-3d--vision-岗都会问">L1 必会题(任何 3D / vision 岗都会问)</h3>
<details>
<summary>Q1.NeRF 体渲染公式?</summary>
<ul><li>$C(\mathbf{r}) = \int T(t)\sigma(\mathbf{r}(t))\mathbf{c}(\mathbf{r}(t),\mathbf{d})dt$</li><li>$T(t) = \exp(-\int_{t_n}^t\sigma\,ds)$ 是透射率</li><li>离散化 → $\alpha$-compositing$C \approx \sum T_i\alpha_i \mathbf{c}_i$$\alpha_i = 1 - e^{-\sigma_i\delta_i}$</li></ul>
<p>只写 $\sum \alpha_i \mathbf{c}_i$ 漏 $T_i$;或把 $\alpha_i$ 写成 $\sigma_i\delta_i$(一阶近似但严格错)。</p>
</details>
<details>
<summary>Q2.为什么 NeRF 要 positional encoding</summary>
<ul><li>MLP 默认低频偏置(NTK 分析)</li><li>$\gamma(p) = (\sin 2^k\pi p, \cos 2^k\pi p)_{k=0}^{L-1}$ 提供高频 basis</li><li>直接学 $(x,y,z) \to (\sigma,\mathbf{c})$ 出来的图糊;加 PE 后高频细节恢复</li></ul>
<p>误以为 PE 是给 MLP 加位置(其实是给空间频率谱),或弄反 $\mathbf{x}$ vs $\mathbf{d}$ 的频率级数($L=10$ vs $L=4$)。</p>
</details>
<details>
<summary>Q3.NeRF 的 hierarchical sampling 是什么?</summary>
<ul><li>两个网络:coarse + fine</li><li>coarse 均匀采 64 点,渲染得 weights $w_i = T_i\alpha_i$</li><li>把 $w$ 归一化为 PDF,按重要性采 128 个 fine 点(密集采在表面)</li><li>Loss 同时监督两网络</li></ul>
<p>说"只采一次更密集"——错过了 importance sampling 的核心。</p>
</details>
<details>
<summary>Q4.Instant-NGP 为什么比 NeRF 快 5+ OOM</summary>
<ul><li><strong>Hash 网格替代密集 grid</strong>:固定 $T$ 大小哈希表,cache-friendly</li><li><strong>Tiny MLP</strong> (2 层 hidden 64) 替代大 MLPNeRF 8 层 256</li><li><strong>多分辨率级联</strong> + <strong>occupancy grid</strong> skip 空白区采样</li><li><strong>CUDA fused kernel</strong>tiny-cuda-nn</li></ul>
<p>只说"用了哈希"——漏了 multi-resolution + tiny-MLP + occupancy skip 的组合贡献。</p>
</details>
<details>
<summary>Q5.3DGS 的"高斯"是怎么定义的?</summary>
<ul><li>每个 Gaussian $G_i = (\mu_i, \Sigma_i, \alpha_i, c_i(\mathbf{d}))$</li><li>$\mu \in \mathbb{R}^3$ 位置,$\Sigma \in \mathbb{R}^{3\times 3}$ 协方差</li><li>$\Sigma = R S S^\top R^\top$ 分解($R$ 用四元数,$S$ 用对角 + $\exp$),保证半正定</li><li>$c(\mathbf{d})$ 用球谐 SH 系数($\ell = 3$48 参数)</li></ul>
<p>只说"高斯分布"——漏了协方差参数化技巧 + SH color。</p>
</details>
<details>
<summary>Q6.3DGS 渲染怎么做?</summary>
<ul><li>把 3D Gaussian 投影到 2D$\Sigma' = JW\Sigma W^\top J^\top$</li><li>按深度排序</li><li>Front-to-back alpha-blending(与 NeRF $\alpha$-compositing 同源)</li><li>实际是 tile-based + CUDA radix sort</li></ul>
<p>只说"光栅化",不提投影 Jacobian / 排序 / alpha-blend。</p>
</details>
<details>
<summary>Q7.3DGS 的 densification 怎么做?</summary>
<ul><li>高梯度 + 小 scale → <strong>clone</strong>under-reconstruction</li><li>高梯度 + 大 scale → <strong>split</strong>over-reconstruction</li><li>低 opacity 或过大 screen-size → <strong>prune</strong></li><li>周期性 reset opacity 防 floater</li></ul>
<p>把 clone 和 split 弄反;忘了 reset 这步。</p>
</details>
<details>
<summary>Q8.NeRF vs 3DGS 对比?</summary>
<ul><li><strong>NeRF</strong>:隐式 (MLP),渲染慢(ray march),editing 难</li><li><strong>3DGS</strong>:显式(点云),渲染快(rasterize),editing 易</li><li><strong>质量</strong>3DGS PSNR 通常 ≥ NeRFNeRF 在体积效应(烟雾 / 半透明)更好</li><li><strong>业界趋势</strong>3DGS 主流,NeRF research-only</li></ul>
<p>把两者当不可比较的不同事物——其实都是 volumetric scene rep3DGS 是 explicit version of NeRF。</p>
</details>
<details>
<summary>Q9.Marching Cubes 是什么?</summary>
<ul><li>输入 3D 标量场 + 阈值,输出三角网格</li><li>每 voxel 8 角点二值化(高/低于阈值)→ 256 种 lookup table</li><li>Edge 上线性插值定顶点位置</li><li>不可微(lookup 离散)</li></ul>
<p>说"找等高线"——MC 是 3D,等高线是 2D Marching Squares 的事。</p>
</details>
<details>
<summary>Q10.SDS 大致是什么?</summary>
<ul><li>用 pretrained 2D diffusion (Stable Diffusion) 监督 3D 表示</li><li>渲染 $x = g(\theta, \pi)$,加噪 $x_t$,问 diffusion "这是 $y$ 的图吗"</li><li>gradient $\propto (\epsilon_\phi(x_t; y) - \epsilon)\cdot \partial x/\partial \theta$</li><li>DreamFusion (Poole et al. 2022 arXiv → ICLR 2023 Outstanding Paper) 提出</li></ul>
<p>只说"用 SD 训 NeRF",漏了 SDS gradient 的特殊形式(去掉 U-Net Jacobian)。</p>
</details>
<h3 id="l2-进阶题research-oriented-岗位">L2 进阶题(research-oriented 岗位)</h3>
<details>
<summary>Q11.推导 NeRF 连续积分 → 离散 $\alpha$-compositing。</summary>
<ul><li>$T$ 满足 $dT/dt = -\sigma T$,段内 $\sigma$ 常数 → $T(t_{i+1})/T(t_i) = e^{-\sigma_i\delta_i}$</li><li>段内颜色贡献 $\int_0^{\delta_i} T_i e^{-\sigma_i s}\sigma_i \mathbf{c}_i\,ds = T_i\mathbf{c}_i(1 - e^{-\sigma_i\delta_i})$</li><li>记 $\alpha_i = 1 - e^{-\sigma_i\delta_i}$,则 $C \approx \sum T_i\alpha_i \mathbf{c}_i$$T_i = \prod_{j<i}(1 - \alpha_j)$</li></ul>
<p>把 $\alpha_i$ 写成 $\sigma_i\delta_i$ 而非 $1 - e^{-\sigma_i\delta_i}$;或省了 ODE 求解过程。</p>
</details>
<details>
<summary>Q12.推导 3DGS 的 3D→2D 投影 Jacobian。</summary>
<ul><li>透视投影 $\pi(\mathbf{x}) = (f_x x/z, f_y y/z)$ 非线性</li><li>一阶 Taylor$\pi(\mathbf{x}) \approx \pi(\mu) + J(\mathbf{x}-\mu)$</li><li>$J = \partial\pi/\partial\mathbf{x}|_\mu = \begin{pmatrix} f_x/z & 0 & -f_x x/z^2 \\ 0 & f_y/z & -f_y y/z^2 \end{pmatrix}$</li><li>$\Sigma' = JW\Sigma W^\top J^\top$$W$ 是 world→cam 旋转)</li></ul>
<p>直接套 "covariance projection" 公式不推;或忘了 $W$ 这步(World→Cam 旋转)。</p>
</details>
<details>
<summary>Q13.Instant-NGP 的 hash collision 如何消歧?</summary>
<ul><li><strong>Multi-resolution 冗余</strong>:粗 level $N_\ell^d \le T$ 不冲突,细 level 才冲突;MLP 可从粗-fine 共同推</li><li><strong>稀疏激活</strong>:有效 supervision 集中在 surface 附近;空白区冲突 entry 无梯度</li><li><strong>MLP 后处理</strong>:在 $L\times F$ 拼接特征上学非线性融合,可 disambiguate</li><li>没有 explicit collision resolution;靠"lazy resolution by sparsity + redundancy"</li></ul>
<p>以为有 hash chaining 之类的传统消歧——实际是数据驱动 implicit 消歧。</p>
</details>
<details>
<summary>Q14.SDS gradient 漏了哪项 Jacobian?为什么?</summary>
<ul><li>Naive diffusion training grad$(\epsilon_\phi - \epsilon)\cdot \partial \epsilon_\phi/\partial x_t \cdot \alpha_t \cdot \partial x/\partial \theta$</li><li>SDS 把 $\partial \epsilon_\phi/\partial x_t$ <strong>U-Net Jacobian</strong> 扔掉</li><li>直觉:(1) 计算昂贵;(2) U-Net 没训练 second-order 稳定 → Jacobian 噪声大</li><li>代价:SDS 变成 mode-seeking KL,需要大 CFG (100) 才能逃 mean-mode → over-saturation</li></ul>
<p>只说"为了简化"不说后果。或不知道 mode-seeking 是 KL 方向决定的。</p>
</details>
<details>
<summary>Q15.VSD 如何缓解 SDS 的 over-saturation</summary>
<ul><li>SDS:拉向 prior $p_\phi$ 的 mode;需 CFG=100 强化 → over-saturation</li><li><strong>VSD</strong>:把 3D 参数 $\theta$ 视为 random variable $\mu(\theta)$,最小化 KL(rendered dist || prior)</li><li>引入<strong>辅助 score</strong> $\epsilon_\psi$LoRA 微调 SD)跟踪当前 $\mu$ 的 score</li><li>gradient = $(\epsilon_\phi - \epsilon_\psi)\cdot \partial x/\partial \theta$ —— <strong>relative score</strong>,不需大 CFG</li><li>类似 RL actor-critic 用 value baseline 减方差</li></ul>
<p>说 VSD 用 "variational" 但讲不清 $\epsilon_\psi$ 替代 raw noise 的角色。</p>
</details>
<details>
<summary>Q16.Zero-1-to-3 / SyncDreamer / MVDream 区别?</summary>
<ul><li><strong>Zero-1-to-3</strong> (Liu 2023 ICCV)input view + 相机 $\Delta R, \Delta T$ → single novel view;每次独立 sample</li><li><strong>Zero-1-to-3++</strong> (Shi 2023):固定 6 个 anchor view,一次出多张(减 randomness</li><li><strong>SyncDreamer</strong> (Liu 2024 ICLR):在 latent 上<strong>联合</strong>预测多视图,cross-attention 让 views 互看 → 一致性更好</li><li><strong>MVDream</strong> (Shi 2024 ICLR)text-to-multi-view(不需要 input image),4 视图同生成 + SDS 精化</li></ul>
<p>只说"都是 novel view"——漏了独立 vs 联合 vs text-only 这条主线。</p>
</details>
<details>
<summary>Q17.Mip-NeRF 怎么抗锯齿?</summary>
<ul><li>Vanilla NeRF 把像素当 ray;不同分辨率下同像素对应不同尺度 → aliasing</li><li><strong>Mip-NeRF</strong> 把像素当 cone(视锥),cone 段近似 anisotropic Gaussian</li><li><strong>IPE (Integrated Positional Encoding)</strong>$\mathbb{E}_{\mathbf{x}\sim\mathcal{N}(\mu,\Sigma)}[\gamma(\mathbf{x})]$ 有闭式解</li><li>高频系数被 $\Sigma$ 衰减 → multi-scale 自动平滑</li></ul>
<p>只说"用 cone",不讲 IPE 的高频衰减作用。</p>
</details>
<details>
<summary>Q18.NeuS vs vanilla NeRF 提 mesh 的差别?</summary>
<ul><li>vanilla NeRFdensity 没明确 surface,提 mesh 要选 $\sigma$ 阈值(不稳)</li><li><strong>NeuS</strong> (Wang 2021 NeurIPS):用 <strong>SDF $d(\mathbf{x})$</strong> 替换 density,定义 $\sigma$ via sigmoid 导数</li><li>表面 = $\{d = 0\}$<strong>良好定义</strong></li><li>Marching Cubes 直接对 SDF 跑,质量明显更好</li></ul>
<p>直接说"用 SDF",但不讲 NeuS 怎么把 SDF 接到 NeRF 体渲染里。</p>
</details>
<details>
<summary>Q19.LRM 系列(TripoSR / InstantMesh)核心?</summary>
<ul><li><strong>Triplane</strong> 表示:3 个轴对齐 2D 平面,$O(N^2)$ 显存</li><li>Transformer 把图像 token + Plucker ray embedding → regress triplane tokens</li><li>推理 feedforward(无 SDS / 无 iterative 优化),<strong>0.5-3 秒</strong> 出 3D</li><li>TripoSR (Stability+Tripo 2024) / InstantMesh (Xu 2024) / SF3D (2024) 都属此族</li></ul>
<p>把它们当成 SDS 系列——错,LRM 完全 feedforward;不算 distillation。</p>
</details>
<details>
<summary>Q20.3DGS 怎么提 mesh</summary>
<ul><li>vanilla 3DGS 不友好(ellipsoid 不是 surface</li><li><strong>SuGaR</strong> (Guédon 2024 CVPR)surface alignment loss + Poisson reconstruction</li><li><strong>2DGS</strong> (Huang 2024 SIGGRAPH):把 ellipsoid 退化为 2D disk,对齐表面 → MC 提 mesh 更稳</li><li><strong>GSDF</strong> (Yu 2024)joint train SDF head 与 3DGS</li></ul>
<p>说"直接 MC"——3DGS 没有 density 场,直接 MC 不 work;必须先 surface-align。</p>
</details>
<h3 id="l3-顶级-lab-题顶会--industry-研究岗">L3 顶级 lab 题(顶会 / industry 研究岗)</h3>
<details>
<summary>Q21.手推 NeRF 离散 $\alpha$-compositing。</summary>
<ul><li>ODE $dT/dt = -\sigma(t) T(t)$,初值 $T(t_n) = 1$ → $T(t) = \exp(-\int_{t_n}^t \sigma\,ds)$</li><li>段内 $[t_i, t_{i+1}]$ 上 $\sigma$ 常数 $= \sigma_i$,所以 $T(t_{i+1}) = T(t_i)e^{-\sigma_i\delta_i}$</li><li>段间累积 $T_i = T(t_i) = \prod_{j<i} e^{-\sigma_j\delta_j} = \prod_{j<i}(1 - \alpha_j)$其中 $\alpha_j = 1 - e^{-\sigma_j\delta_j}$</li><li>段内颜色贡献 $\int_{t_i}^{t_{i+1}} T(t)\sigma_i\mathbf{c}_i\,dt = \mathbf{c}_i T_i \int_0^{\delta_i}\sigma_i e^{-\sigma_i s}ds = T_i\mathbf{c}_i(1 - e^{-\sigma_i\delta_i}) = T_i\alpha_i\mathbf{c}_i$</li><li>合成 $C \approx \sum_i T_i\alpha_i \mathbf{c}_i$</li><li><strong>关键</strong>$\alpha_i = 1 - e^{-\sigma_i\delta_i}$ 严格 vs $\alpha_i \approx \sigma_i\delta_i$ 一阶近似(在 $\sigma\delta \ll 1$ 时一致)</li></ul>
<p>省略 ODE 推导直接套结论;或在 $\sigma\delta$ 大时用 $\sigma_i\delta_i$ 替 $\alpha_i$ 出错。</p>
</details>
<details>
<summary>Q22.Instant-NGP hash collision 如何被 MLP 自动消歧?</summary>
<ul><li><strong>冲突发生场景</strong>fine level 的格点数 $N_\ell^d > T$hash 表大小),多个格点映到同一 entry</li><li><strong>稀疏激活</strong>:场景大部分 voxel 是 background<strong>仅 surface 附近 voxel 有非零监督梯度</strong>——冲突的"两个 background entry"得不到信号,互不污染</li><li><strong>多分辨率冗余</strong>:粗 level $N_\ell^d \le T$ 保证 unique;细 level 提供 detail。即使细 level 冲突,粗 level 的非冲突特征已唯一识别该点</li><li><strong>MLP 后处理</strong>tiny MLP 在 $L\times F$ 拼接特征上学非线性融合,遇到冲突 entry 时可以用<strong>其他 level 的非冲突特征 disambiguate</strong></li><li><strong>梯度自动调节</strong>:训练时高梯度自然集中在 surface entry;冲突 entry 若同时在表面(罕见)会被 loss 推到妥协位置(取多次采样平均)</li><li><strong>物理直觉</strong>:与其代价昂贵地搞 perfect hash,不如允许冲突 + 用数据驱动 implicit 消歧("lazy collision resolution"</li></ul>
<p>说"哈希冲突由 MLP 解决"但讲不清是哪些 mechanism(稀疏性 + 多尺度 + MLP 非线性)共同作用。</p>
</details>
<details>
<summary>Q23.推导 3DGS 3D→2D 协方差投影 Jacobian。</summary>
<ul><li>World → Camera:刚体变换 $\mathbf{x}_\text{cam} = W\mathbf{x} + t$;协方差只受旋转影响,$\Sigma_\text{cam} = W\Sigma W^\top$</li><li>Camera → Screen:透视投影 $\pi(x, y, z) = (f_x x/z, f_y y/z)$ 非线性</li><li>在均值 $\mu_\text{cam}$ 处一阶 Taylor$\pi(\mathbf{x}) \approx \pi(\mu_\text{cam}) + J(\mathbf{x} - \mu_\text{cam})$$J = \partial\pi/\partial\mathbf{x}|_{\mu_\text{cam}}$</li><li>$J = \begin{pmatrix} f_x/z & 0 & -f_x x/z^2 \\ 0 & f_y/z & -f_y y/z^2 \end{pmatrix} \in \mathbb{R}^{2\times 3}$</li><li>$\text{Cov}[\pi(\mathbf{x})] = J\Sigma_\text{cam} J^\top = JW\Sigma W^\top J^\top \in \mathbb{R}^{2\times 2}$</li><li>这是 EWA splatting (Zwicker 2001) 的经典推论;3DGS 直接沿用</li><li>实际实现还加 $0.3 I$ low-pass filteranti-aliasing</li></ul>
<p>不会做一阶 Taylor 把非线性投影线性化;或漏了 World→Cam 那步。</p>
</details>
<details>
<summary>Q24.SDS gradient 丢掉哪项 Jacobian?为什么"反而 work"</summary>
<ul><li><p><strong>Naive diffusion training gradient</strong></p>
<p>$\nabla_\theta \mathcal{L}_\text{diff} = \mathbb{E}[w(t)\cdot 2(\epsilon_\phi - \epsilon)\cdot \underbrace{\partial \epsilon_\phi/\partial x_t}_{\text{U-Net Jacobian}}\cdot \alpha_t \cdot \partial x/\partial\theta]$</p></li><li><p><strong>SDS</strong> 扔掉 U-Net Jacobian $\partial \epsilon_\phi/\partial x_t$</p>
<p>$\nabla_\theta \mathcal{L}_\text{SDS} = \mathbb{E}[w(t)(\epsilon_\phi - \epsilon)\cdot \partial x/\partial\theta]$</p></li><li><p><strong>为什么扔掉 reasonable</strong></p>
<ul><li>U-Net Jacobian 计算昂贵(H×W×3 输入 → H×W×3 输出的 second-order</li><li>U-Net 没训练 second-order 稳定,Jacobian 数值差</li><li>$(\epsilon_\phi - \epsilon)$ 本身就是 score 的代理($\epsilon_\phi/\sigma_t \approx -\nabla_{x_t}\log p_\phi$),扔掉 Jacobian 相当于用 first-order score 信号</li></ul></li><li><strong>代价</strong>SDS 数学上等价 mode-seeking KL(往 prior 的 mode 跑),加大 CFG (100) 才能逃出 mean-blur</li><li><strong>症状</strong>over-saturation(颜色饱和)+ Janus(多视角同一面孔)+ over-smoothing(细节糊)</li></ul>
<p>只说"为简化扔了 Jacobian",不解释 mode-seeking 后果 + 为什么需要大 CFG。</p>
</details>
<details>
<summary>Q25.VSD 为什么能在小 CFG 下避免 over-saturation</summary>
<ul><li><strong>SDS 视角</strong>:把 $\theta$ 当点估计;gradient 拉向 $p_\phi(\cdot|y)$ mode;大 CFG 让 mode 更尖 → over-saturation</li><li><strong>VSD 视角</strong>:把 $\theta$ 当 random variable $\mu(\theta)$<strong>最小化的是渲染加噪图像分布</strong>之间的 KL$\mathbb{E}_t[D_\text{KL}(q_\mu^t(x_t|y) \,\|\, p_\phi^t(x_t|y))]$(不是直接在 $\theta$ 域上对一个 3D prior 求 KL</li><li>引入<strong>辅助 score</strong> $\epsilon_\psi$(用 LoRA 微调 Stable Diffusion)跟踪当前 rendered 分布的 score</li><li><p><strong>VSD gradient</strong></p>
<p>$\nabla_\theta \mathcal{L}_\text{VSD} = \mathbb{E}[w(t)(\epsilon_\phi - \epsilon_\psi)\cdot \partial x/\partial \theta]$</p>
<p><strong>relative score</strong>target prior score $-$ current rendered score</p></li><li><strong>几何直觉</strong>:从"我现在在哪"指向"prior 在哪"——是局部"梯度方向"而非全局 mode;不需大 CFG 锐化</li><li><strong>类比 RL</strong>actor-critic 用 value baseline 减方差;VSD 用 $\epsilon_\psi$ 作为 baseline 减 SDS 噪声</li><li><strong>效果(ProlificDreamer</strong>:CFG 可降至 7.5,颜色自然;几何更细;可维护多 mode(多样性)</li></ul>
<p>只说"VSD 引入变分推断"不讲 $\epsilon_\psi$ 角色 + relative gradient 视角。</p>
</details>
<h2 id="a-附录代码完整骨架--参考文献">§A 附录:代码完整骨架 + 参考文献</h2>
<h3 id="a1-完整-from-scratch-代码包含">A.1 完整 from-scratch 代码包含</h3>
<p><code>volume_render()</code> (NeRF α-compositing 含数值稳定) · <code>positional_encoding()</code> (γ(p) Fourier features) · <code>gaussian_splat_forward()</code> (3DGS 教学版前向 + 投影 Jacobian) · <code>densify_and_prune()</code> (3DGS densification 启发式) · <code>sds_loss()</code> (SDS gradient surrogate) · <code>marching_cubes_sketch()</code> (mesh 提取接口,用 scikit-image)。</p>
<h3 id="a2-关键论文-reading-list">A.2 关键论文 reading list</h3>
<ul><li><strong>NeRF 系</strong>Mildenhall 2020 ECCV (HM); Müller <strong>Instant-NGP</strong> SIGGRAPH 2022 Best; Barron <strong>Mip-NeRF</strong> / <strong>360</strong> ICCV 2021 / CVPR 2022; Wang <strong>NeuS</strong> + Yariv <strong>VolSDF</strong> NeurIPS 2021; Fridovich-Keil <strong>Plenoxels</strong> CVPR 2022; Chen <strong>TensoRF</strong> ECCV 2022.</li><li><strong>3DGS 系</strong>Kerbl <strong>3D Gaussian Splatting</strong> SIGGRAPH 2023 Best; Huang <strong>2D Gaussian Splatting</strong> SIGGRAPH 2024; Luiten <strong>Dynamic 3DGS</strong> 3DV 2024; Wu <strong>4DGS</strong> CVPR 2024; Guédon <strong>SuGaR</strong> CVPR 2024.</li><li><strong>Mesh / SDF</strong>Shen <strong>DMTet</strong> NeurIPS 2021 / <strong>FlexiCubes</strong> SIGGRAPH 2023.</li><li><strong>SDS 系</strong>Poole <strong>DreamFusion</strong> arXiv 2022.09 → ICLR 2023 Outstanding; Wang <strong>ProlificDreamer (VSD)</strong> NeurIPS 2023 Spotlight; Lin <strong>Magic3D</strong> CVPR 2023; Chen <strong>Fantasia3D</strong> ICCV 2023; Tang <strong>DreamGaussian</strong> ICLR 2024; Yi <strong>GaussianDreamer</strong> CVPR 2024.</li><li><strong>Single-image 3D</strong>Liu <strong>Zero-1-to-3</strong> ICCV 2023 / <strong>One-2-3-45</strong> NeurIPS 2023 / <strong>SyncDreamer</strong> ICLR 2024; Shi <strong>Zero-1-to-3++</strong> arXiv 2023 / <strong>MVDream</strong> ICLR 2024; Hong <strong>LRM</strong> arXiv 2023.11 → ICLR 2024; Tochilkin <strong>TripoSR</strong> arXiv 2024; Xu <strong>InstantMesh</strong> arXiv 2024; Boss <strong>Stable Fast 3D</strong> arXiv 2024.</li><li><strong>3D Foundation Models</strong>Xiang <strong>Trellis</strong> arXiv 2024 (Microsoft); Tencent <strong>Hunyuan3D-2</strong> arXiv 2501.12202 (2025); Zhang <strong>CLAY</strong> SIGGRAPH 2024.</li></ul>
<h3 id="a3-embodied-ai--ar--vr-常见追问">A.3 Embodied AI / AR / VR 常见追问</h3>
<p>3DGS 接物理引擎 → 先 2DGS / SuGaR 提 mesh → IsaacSim / MuJoCoNeRF 动态化 → 4DGS / D-NeRF / K-PlanesAR 实时 3DGS → mobile-friendly (PostShot, Luma) + 剪枝 / 量化;3D 数据不足 → Objaverse-XL (Trellis) 或 2D 蒸馏 (DreamFusion 系) 或 multi-view 启发式 (MVDream)。</p>
<hr />
<p><strong>3D Generation Quick Reference</strong> · 主要参考:Mildenhall 2020 (NeRF), Müller 2022 (Instant-NGP), Kerbl 2023 (3DGS), Poole 2022/ICLR 2023 (DreamFusion), Wang 2023 (VSD), Xiang 2024 (Trellis), Tencent 2025 (Hunyuan3D-2). 涵盖:NeRF 体渲染推导、Instant-NGP hash 网格、3DGS 投影 Jacobian、SDS / VSD 梯度推导、single-image 3D、3D foundation models。Embodied AI / AR / VR 必备。</p>
<footer class="aris-footer">
Generated by <a href="https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/blob/main/skills/render-html/SKILL.md">ARIS <code>/render-html</code></a> ·
source path <code>docs/tutorials/3d_generation_tutorial.md</code> ·
SHA256 <code>beab28036da8</code> ·
generated at 2026-05-19 07:54 UTC.
This is a generated view — edit the source Markdown, then re-render.
</footer>
</main>
</div>
</body>
</html>