Files
2026-07-13 13:37:02 +08:00

1118 lines
89 KiB
HTML
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html lang="zh-CN">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Quantization 面试 Cheat Sheet</title>
<meta name="generator" content="ARIS render-html (academic, v1)">
<meta name="aris:source-path" content="docs/tutorials/quantization_tutorial.md">
<meta name="aris:source-sha256" content="f04948e0f203c2972cb16b925ae326b7e46f34e691c037df51dbabe54f88a301">
<meta name="aris:generated-at" content="2026-05-19 09:46 UTC">
<!-- MathJax 3 -->
<script>
window.MathJax = {
tex: { inlineMath: [['$', '$'], ['\\(', '\\)']], displayMath: [['$$', '$$'], ['\\[', '\\]']], processEscapes: true },
options: { skipHtmlTags: ['script', 'noscript', 'style', 'textarea', 'pre', 'code'] }
};
</script>
<script src="https://cdn.jsdelivr.net/npm/mathjax@3/es5/tex-mml-chtml.js" async></script>
<!-- highlight.js -->
<link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/highlightjs/cdn-release@11.9.0/build/styles/atom-one-light.min.css">
<script src="https://cdn.jsdelivr.net/gh/highlightjs/cdn-release@11.9.0/build/highlight.min.js"></script>
<script>document.addEventListener('DOMContentLoaded', () => hljs.highlightAll());</script>
<style>
:root {
--bg: #fdfcf7;
--bg-soft: #f4f1ea;
--bg-code: #f8f5ec;
--ink: #1a1a1a;
--ink-soft: #4a4a4a;
--ink-muted: #6b6b6b;
--primary: #1a4a8c;
--primary-soft: #2d6cb8;
--accent: #b8390e;
--warn: #b45309;
--warn-bg: #fef3c7;
--info-bg: #dbeafe;
--good-bg: #d1fae5;
--good: #065f46;
--bad-bg: #fee2e2;
--bad: #991b1b;
--border: #d6d0c0;
--border-soft: #e8e3d5;
}
* { box-sizing: border-box; }
html { scroll-behavior: smooth; }
body {
font-family: "Source Serif Pro", "Source Serif 4", "Crimson Pro", "Georgia", "Songti SC", "STSong", serif;
line-height: 1.65;
color: var(--ink);
background: var(--bg);
margin: 0;
padding: 0;
font-size: 16px;
}
.layout {
max-width: 1280px;
margin: 0 auto;
display: grid;
grid-template-columns: 260px 1fr;
gap: 48px;
padding: 40px 32px;
}
nav.toc {
position: sticky;
top: 24px;
align-self: start;
font-size: 13px;
max-height: calc(100vh - 48px);
overflow-y: auto;
border-right: 1px solid var(--border-soft);
padding-right: 16px;
}
nav.toc h3 {
margin: 0 0 12px;
font-size: 12px;
text-transform: uppercase;
letter-spacing: 0.08em;
color: var(--ink-muted);
font-weight: 600;
}
nav.toc ol { list-style: none; padding: 0; margin: 0; counter-reset: toc; }
nav.toc ol li { margin: 5px 0; counter-increment: toc; }
nav.toc ol li::before { content: counter(toc) ". "; color: var(--ink-muted); margin-right: 4px; }
nav.toc a {
color: var(--ink-soft);
text-decoration: none;
border-bottom: 1px dotted transparent;
}
nav.toc a:hover { color: var(--primary); border-bottom-color: var(--primary); }
nav.toc ul { list-style: none; padding-left: 14px; margin: 3px 0; font-size: 12px; }
nav.toc ul li::before { content: "→ "; color: var(--border); }
main { min-width: 0; }
header.hero {
border-bottom: 3px double var(--primary);
padding-bottom: 24px;
margin-bottom: 32px;
}
header.hero .eyebrow {
color: var(--accent);
font-size: 13px;
text-transform: uppercase;
letter-spacing: 0.12em;
font-weight: 600;
margin-bottom: 8px;
}
header.hero h1 {
font-size: 32px;
line-height: 1.2;
margin: 0 0 12px;
color: var(--ink);
font-weight: 700;
letter-spacing: -0.01em;
}
header.hero .subtitle {
font-size: 16px;
color: var(--ink-soft);
margin: 0 0 8px;
font-style: italic;
}
header.hero .byline {
font-size: 14px;
color: var(--ink-soft);
margin: 0 0 20px;
}
header.hero .byline strong {
color: var(--ink);
font-weight: 600;
}
header.hero .meta {
display: flex;
gap: 20px;
flex-wrap: wrap;
font-size: 12px;
color: var(--ink-muted);
border-top: 1px solid var(--border-soft);
padding-top: 14px;
}
header.hero .meta span strong { color: var(--ink-soft); }
header.hero .meta code {
font-family: "JetBrains Mono", "SF Mono", "Menlo", "Consolas", monospace;
font-size: 11px;
background: var(--bg-soft);
padding: 1px 5px;
border-radius: 3px;
border: 1px solid var(--border-soft);
}
h2 {
font-size: 24px;
margin: 44px 0 14px;
padding-bottom: 8px;
border-bottom: 1px solid var(--border);
color: var(--ink);
font-weight: 700;
}
h2 .num { color: var(--primary); font-weight: 600; margin-right: 8px; }
h3 { font-size: 19px; margin: 28px 0 10px; color: var(--primary); font-weight: 600; }
h4 { font-size: 16px; margin: 20px 0 8px; color: var(--ink); font-weight: 600; }
p { margin: 10px 0; }
ul, ol { padding-left: 22px; margin: 10px 0; }
ul li, ol li { margin: 4px 0; }
ul li::marker { color: var(--primary); }
strong { color: var(--accent); font-weight: 600; }
em { color: var(--ink-soft); }
a { color: var(--primary); }
a:hover { color: var(--accent); }
code:not(.hljs) {
font-family: "JetBrains Mono", "SF Mono", "Menlo", "Consolas", monospace;
font-size: 0.86em;
background: var(--bg-code);
padding: 1px 5px;
border-radius: 3px;
border: 1px solid var(--border-soft);
color: var(--accent);
}
pre {
background: #fafaf6;
border: 1px solid var(--border);
border-left: 4px solid var(--primary);
padding: 0;
overflow-x: auto;
border-radius: 4px;
margin: 14px 0;
}
pre code, pre code.hljs {
background: transparent !important;
display: block;
padding: 14px 18px !important;
font-size: 13px;
line-height: 1.55;
font-family: "JetBrains Mono", "SF Mono", "Menlo", monospace;
color: var(--ink);
}
pre.diagram {
background: #f9f6ed;
border-left: 4px solid var(--accent);
font-size: 12.5px;
line-height: 1.4;
}
.callout {
margin: 16px 0;
padding: 12px 16px;
border-radius: 4px;
border-left: 4px solid;
font-size: 15px;
}
.callout-title {
font-weight: 600;
margin-bottom: 6px;
font-size: 12px;
text-transform: uppercase;
letter-spacing: 0.06em;
}
.callout-info { background: var(--info-bg); border-left-color: var(--primary); }
.callout-info .callout-title { color: var(--primary); }
.callout-warn { background: var(--warn-bg); border-left-color: var(--warn); }
.callout-warn .callout-title { color: var(--warn); }
.callout-good { background: var(--good-bg); border-left-color: var(--good); }
.callout-good .callout-title { color: var(--good); }
.callout-bad { background: var(--bad-bg); border-left-color: var(--bad); }
.callout-bad .callout-title { color: var(--bad); }
table {
width: 100%;
border-collapse: collapse;
margin: 16px 0;
font-size: 14px;
border: 1px solid var(--border);
border-radius: 4px;
overflow: hidden;
}
thead { background: var(--primary); color: white; }
th, td {
text-align: left;
padding: 9px 12px;
border-bottom: 1px solid var(--border-soft);
vertical-align: top;
}
th { font-weight: 600; font-size: 13px; letter-spacing: 0.02em; }
tr:last-child td { border-bottom: none; }
tbody tr:nth-child(even) { background: var(--bg-soft); }
details.qa, details {
background: white;
border: 1px solid var(--border-soft);
border-radius: 6px;
margin: 10px 0;
padding: 0;
}
details summary {
cursor: pointer;
padding: 10px 14px;
font-weight: 600;
font-size: 14px;
color: var(--primary);
list-style: none;
user-select: none;
}
details summary::-webkit-details-marker { display: none; }
details summary::before {
content: "▸ ";
margin-right: 4px;
display: inline-block;
transition: transform 0.15s;
}
details[open] summary::before { transform: rotate(90deg); }
details[open] summary { border-bottom: 1px solid var(--border-soft); }
details > :not(summary) { padding: 10px 14px; }
details p:first-of-type { margin-top: 8px; }
mjx-container[display="true"] { margin: 12px 0 !important; }
footer.aris-footer {
margin-top: 60px;
padding-top: 20px;
border-top: 1px solid var(--border);
font-size: 12px;
color: var(--ink-muted);
}
footer.aris-footer a { color: var(--ink-muted); border-bottom: 1px dotted var(--border); }
@media (max-width: 900px) {
.layout { grid-template-columns: 1fr; gap: 20px; padding: 20px 16px; }
nav.toc {
position: static;
max-height: none;
border-right: none;
border-bottom: 1px solid var(--border-soft);
padding-right: 0;
padding-bottom: 14px;
}
header.hero h1 { font-size: 24px; }
h2 { font-size: 20px; }
}
@media print {
nav.toc { display: none; }
.layout { grid-template-columns: 1fr; padding: 0; }
body { background: white; }
header.hero { border-bottom-color: var(--ink); }
}
</style>
</head>
<body>
<div class="layout">
<nav class="toc">
<h3>Contents</h3>
<ol>
<li><a href="#0-tldr-cheat-sheet">§0 TL;DR Cheat Sheet</a>
</li>
<li><a href="#1-直觉为什么-llm-需要量化为什么这么难">§1 直觉:为什么 LLM 需要量化、为什么这么难</a>
<ul>
<li><a href="#11-量化的本质">1.1 量化的本质</a></li>
<li><a href="#12-llm-量化为什么比-cnn-难">1.2 LLM 量化为什么比 CNN 难</a></li>
<li><a href="#13-三类主流方案">1.3 三类主流方案</a></li>
</ul>
</li>
<li><a href="#2-量化数学基础">§2 量化数学基础</a>
<ul>
<li><a href="#21-affine-quantizationuniform--linear-量化">2.1 Affine quantizationuniform / linear 量化)</a></li>
<li><a href="#22-量化误差与-snr">2.2 量化误差与 SNR</a></li>
<li><a href="#23-粒度granularity">2.3 粒度(granularity</a></li>
<li><a href="#24-rounding-模式">2.4 Rounding 模式</a></li>
<li><a href="#25-简洁可运行代码非对称-int8-per-channel-qdq">2.5 简洁可运行代码:非对称 INT8 per-channel q/dq</a></li>
</ul>
</li>
<li><a href="#3-llm-的-outlier-问题">§3 LLM 的 Outlier 问题</a>
<ul>
<li><a href="#31-经验观察dettmers-2022-llmint8">3.1 经验观察(Dettmers 2022, LLM.int8()</a></li>
<li><a href="#32-llmint8--第一个能落地的-8-bit-llm">3.2 LLM.int8() — 第一个能落地的 8-bit LLM</a></li>
<li><a href="#33-不同-outlier-形态gptq--awq--smoothquant-攻击的对象">3.3 不同 outlier 形态(GPTQ / AWQ / SmoothQuant 攻击的对象)</a></li>
</ul>
</li>
<li><a href="#4-gptq基于-obs-的最优-weight-量化必考推导">§4 GPTQ:基于 OBS 的最优 weight 量化(必考推导)</a>
<ul>
<li><a href="#41-问题设定">4.1 问题设定</a></li>
<li><a href="#42-二阶-taylor-展开">4.2 二阶 Taylor 展开</a></li>
<li><a href="#43-obs固定一个坐标到目标值后的最优-delta">4.3 OBS:固定一个坐标到目标值后的最优 $\delta$</a></li>
<li><a href="#44-工程加速cholesky-解-h-1-子矩阵">4.4 工程加速:Cholesky 解 $H^{-1}$ 子矩阵</a></li>
<li><a href="#45-gptq-style-伪代码必考写法">4.5 GPTQ-style 伪代码(必考写法)</a></li>
<li><a href="#46-gptq-vs-早期-round-to-nearest-rtn">4.6 GPTQ vs 早期 round-to-nearest (RTN)</a></li>
</ul>
</li>
<li><a href="#5-awqactivation-aware-weight-quantization">§5 AWQActivation-Aware Weight Quantization</a>
<ul>
<li><a href="#51-salient-channel-的识别">5.1 Salient channel 的识别</a></li>
<li><a href="#52-per-channel-scale-等价变换数学等价--量化误差等价">5.2 Per-channel scale 等价变换(数学等价 ≠ 量化误差等价)</a></li>
<li><a href="#53-与-gptq-的区别">5.3 与 GPTQ 的区别</a></li>
<li><a href="#54-awq-style-per-channel-scale-搜索代码">5.4 AWQ-style per-channel scale 搜索代码</a></li>
</ul>
</li>
<li><a href="#6-smoothquant把-activation-outlier-迁移到-weight-w8a8">§6 SmoothQuant:把 activation outlier 迁移到 weight (W8A8)</a>
<ul>
<li><a href="#61-核心数学">6.1 核心数学</a></li>
<li><a href="#62-migration-strength-alpha关键超参">6.2 Migration strength $\alpha$(关键超参)</a></li>
<li><a href="#63-等价变换不破坏-gemm必考">6.3 等价变换不破坏 GEMM(必考)</a></li>
<li><a href="#64-smoothquant-伪代码">6.4 SmoothQuant 伪代码</a></li>
<li><a href="#65-smoothquant-适用范围">6.5 SmoothQuant 适用范围</a></li>
</ul>
</li>
<li><a href="#7-旋转方法quip--quarot--spinquant">§7 旋转方法:QuIP / QuaRot / SpinQuant</a>
<ul>
<li><a href="#71-quip-chee-et-al-2023-neurips">7.1 QuIP (Chee et al. 2023, NeurIPS)</a></li>
<li><a href="#72-quarot-ashkboos-et-al-2024-neurips">7.2 QuaRot (Ashkboos et al. 2024 NeurIPS)</a></li>
<li><a href="#73-spinquant-liu-et-al-2024--iclr-2025-meta">7.3 SpinQuant (Liu et al. 2024 → ICLR 2025, Meta)</a></li>
</ul>
</li>
<li><a href="#8-低精度浮点fp8--mx--nvfp4">§8 低精度浮点:FP8 / MX / NVFP4</a>
<ul>
<li><a href="#81-ieee-754-style-浮点编码">8.1 IEEE 754-style 浮点编码</a></li>
<li><a href="#82-fp8-e4m3-bit-level-encoding-代码">8.2 FP8 E4M3 bit-level encoding 代码</a></li>
<li><a href="#83-mx-格式ocp--microsoft-2024">8.3 MX 格式(OCP / Microsoft 2024</a></li>
<li><a href="#84-nvfp4blackwell-2025-nvidia">8.4 NVFP4Blackwell 2025 NVIDIA</a></li>
<li><a href="#85-fp8-在大模型训练中的使用transformer-engine">8.5 FP8 在大模型训练中的使用(Transformer Engine</a></li>
</ul>
</li>
<li><a href="#9-kv-cache-量化">§9 KV Cache 量化</a>
<ul>
<li><a href="#91-kivi--kvquant-的关键观察">9.1 KIVI / KVQuant 的关键观察</a></li>
<li><a href="#92-kivi-liu-et-al-2024-icml">9.2 KIVI (Liu et al. 2024 ICML)</a></li>
<li><a href="#93-kvquant-hooper-et-al-2024-neurips">9.3 KVQuant (Hooper et al. 2024, NeurIPS)</a></li>
<li><a href="#94-qserve--qoq-lin-et-al-2024-mlsys-2025">9.4 QServe / QoQ (Lin et al. 2024 MLSys 2025)</a></li>
<li><a href="#95-kv-cache-quant-代码示意per-channel-k-per-token-v">9.5 KV cache quant 代码示意(per-channel K, per-token V</a></li>
</ul>
</li>
<li><a href="#10-qat-与训练时量化">§10 QAT 与训练时量化</a>
<ul>
<li><a href="#101-ste-straight-through-estimator">10.1 STE (Straight-Through Estimator)</a></li>
<li><a href="#102-llm-qat-liu-et-al-2023">10.2 LLM-QAT (Liu et al. 2023)</a></li>
<li><a href="#103-fp8-trainingtransformer-engine">10.3 FP8 TrainingTransformer Engine</a></li>
<li><a href="#104-bitnet-b158--b2">10.4 BitNet b1.58 / b2</a></li>
</ul>
</li>
<li><a href="#11-框架与生态对照">§11 框架与生态对照</a>
</li>
<li><a href="#12-25-高频面试题">§12 25 高频面试题</a>
<ul>
<li><a href="#l1必会题任何-ml-工程岗都会问">L1必会题(任何 ML 工程岗都会问)</a></li>
<li><a href="#l2进阶题research-oriented-岗位">L2进阶题(research-oriented 岗位)</a></li>
<li><a href="#l3高级变体顶级-lab--系统方向">L3高级变体(顶级 lab / 系统方向)</a></li>
</ul>
</li>
<li><a href="#a-附录完整的工程参考">§A 附录:完整的工程参考</a>
<ul>
<li><a href="#a1-主要论文-reference-list">A.1 主要论文 reference list</a></li>
<li><a href="#a2-一图速查选什么量化方案">A.2 一图速查:选什么量化方案</a></li>
<li><a href="#a3-量化-quick-reference-卡片">A.3 量化 quick reference 卡片</a></li>
<li><a href="#a4-sanity-check-checklist">A.4 Sanity check checklist</a></li>
</ul>
</li>
</ol>
</nav>
<main>
<header class="hero">
<div class="eyebrow">Interview Prep · Quantization for LLM Inference</div>
<h1>Quantization 面试 Cheat Sheet</h1>
<p class="subtitle">GPTQ / AWQ / FP8 / NVFP4 / SmoothQuant + 25 高频题(L1 必会 · L2 进阶 · L3 顶级 lab)</p>
<p class="byline">By <strong>Ruofeng Yang (杨若峰), Shanghai Jiao Tong University</strong></p>
<div class="meta">
<span><strong>Source:</strong> <code>docs/tutorials/quantization_tutorial.md</code></span>
<span><strong>SHA256:</strong> <code>f04948e0f203</code></span>
<span><strong>Rendered:</strong> 2026-05-19 09:46 UTC</span>
</div>
</header>
<h2 id="0-tldr-cheat-sheet">§0 TL;DR Cheat Sheet</h2>
<div class="callout callout-info"><div class="callout-title">8 句话搞定 LLM Quantization</div><p>一页拿下面试核心要点(详见后文 §2–§11 推导)。</p></div>
<ol><li><strong>Affine quantization 公式</strong>$q = \mathrm{round}(x / s) + z$,反量化 $\hat{x} = s\,(q - z)$。对称量化 $z = 0$;非对称量化 $z$ 把 zero-point 对齐到一个整数。</li><li><strong>粒度三档</strong>(scale 共享范围越小,精度越高、开销越大):per-tensor → per-channel → per-group(一行/一列内每 $g$ 个元素共享一个 scale$g = 32 / 64 / 128$ 主流)。</li><li><strong>LLM 量化痛点</strong>activation 存在 <strong>per-channel 系统性 outlier</strong>(少数 channel 量级是平均的 $50\text{-}100\times$),均匀量化必崩。Weight 分布相对平坦,weight-only 量化(GPTQ / AWQ)天生比 weight+act 容易。</li><li><strong>GPTQ (Frantar 2023 ICLR)</strong>:基于 OBS 推导的最优 weight update 公式 $\delta_{\mathbf{w}} = -\dfrac{w_q - \mathrm{quant}(w_q)}{[H^{-1}]_{qq}}\,[H^{-1}]_{q,:}$,逐列量化 + Cholesky 加速 + 128 列 block4-bit weight 几乎无损。</li><li><strong>AWQ (Lin 2024 MLSys)</strong>:观察 "1% salient weights drive most loss",按 activation 幅度选 salient channel<strong>per-channel scale $s_c$ 在 $w \to w\cdot s_c$ / $x \to x/s_c$ 下数学等价但量化误差降低</strong>grid search $s_c = \mathrm{mean}(|x_c|)^\alpha$。</li><li><strong>SmoothQuant (Xiao 2023 ICML)</strong>:把 activation outlier <strong>迁移到 weight</strong>——$Y = (X \mathrm{diag}(s)^{-1})(\mathrm{diag}(s)\,W)$,数学完全等价但 $X / s$ 平滑得多,得以做 W8A8。</li><li><strong>低精度浮点格式族</strong>FP8 (E4M3/E5M2, Hopper)、MX (OCP MXFP8/MXFP6/MXFP4, 32-elem block + E8M0 shared exp)、NVFP4 (Blackwell B100/B200, FP4 E2M1 + per-16-elem FP8 E4M3 scale + per-tensor FP32 scale)。Blackwell tensor core 原生支持 FP4 matmul。</li><li><strong>KV cache quant</strong>K 用 <strong>per-channel</strong>K 的 outlier 沿 channel 维稳定),V 用 <strong>per-token</strong>V outlier 沿 token 维变化)——KIVI / KVQuant 的基本设计。QServe 进一步把 W4A8KV4 整套量化做 SM89/SM90 kernel-level co-design。</li></ol>
<h2 id="1-直觉为什么-llm-需要量化为什么这么难">§1 直觉:为什么 LLM 需要量化、为什么这么难</h2>
<h3 id="11-量化的本质">1.1 量化的本质</h3>
<p>把高精度浮点数(FP16 / BF16)映射到低位宽整数(INT8 / INT4)或低精度浮点(FP8 / FP4),换取:</p>
<ul><li><strong>显存</strong>FP16 → INT4 直接 $4\times$ 压缩(70B 模型从 140 GB → 35 GB,单卡可推)。</li><li><strong>显存带宽</strong>decode 阶段是 memory-bound(每 token 读一遍全部 weights),位宽降一半延迟降一半。</li><li><strong>算力</strong>INT8 / FP8 / FP4 tensor core throughput 比 FP16 高 $2\text{-}8\times$Hopper / Blackwell 上更显著)。</li></ul>
<p>代价:<strong>精度损失</strong>。对 LLM 而言,损失主要来自两个源——activation outlier 和 weight 边界效应。</p>
<h3 id="12-llm-量化为什么比-cnn-难">1.2 LLM 量化为什么比 CNN 难</h3>
<p>CNN 时代的 INT8 PTQNVIDIA TensorRT 2017 那一套)几乎免费:CNN activation 分布近似 Gaussianper-tensor calibration 即可。LLM 上同样的方法用到 OPT-66B / LLaMA-70B 直接掉 5-10 个 PPL 点。原因:</p>
<ul><li><strong>Outlier scaling 现象</strong>Dettmers 2022, LLM.int8()):模型 ≥ 6.7B 之后,少数 channel(约 0.1%-1%)的 activation 量级显著高于其他($50\text{-}100\times$),且<strong>这些 channel 在不同 token / sample 上稳定地是同一批</strong></li><li><strong>Outlier 抢走 scale</strong>per-tensor scale 由 max 决定。少数 outlier 把 scale 顶得很大,大部分 "normal" channel 量化后落入很窄的整数区间(INT8 下只用 ±10 / ±127),有效 bit width 大幅缩减。</li><li><strong>Per-channel weight 易、per-channel activation 难</strong>activation 的 channel 维(hidden dim, $D$)是矩阵乘的内积维(K 维),per-channel scale 沿 K 维变化意味着不能直接在 GEMM 上 fuseweight 的 output-channel 维(N 维)天然 per-channel 可行(output dim 之间无相互内积)。</li></ul>
<h3 id="13-三类主流方案">1.3 三类主流方案</h3>
<table><thead><tr><th>方案</th><th>量化谁</th><th>代表方法</th><th>核心 trick</th></tr></thead><tbody><tr><td><strong>Weight-only PTQ</strong></td><td>W4/W8, A 保持 FP16</td><td>GPTQ, AWQ, QuIP, GGUF Q4_K</td><td>weight 易量化;用 calibration data 找最优量化误差补偿</td></tr><tr><td><strong>Weight + Activation PTQ</strong></td><td>W8A8</td><td>SmoothQuant, ZeroQuant, FP8</td><td>必须处理 activation outlier(迁移 / 旋转)</td></tr><tr><td><strong>Weight + Act + KV (低 bit)</strong></td><td>W4A8KV4 / W4A4</td><td>QuaRot, QServe, SpinQuant</td><td>Hadamard / 学习旋转把所有方向上的 outlier 打平</td></tr></tbody></table>
<div class="callout callout-warn"><div class="callout-title">decode 与 prefill 区别</div><p>Decode 是 memory-boundKV cache + weight 读取量主导),低 bit weight + KV 直接降延迟。Prefill 是 compute-bound$L^2$ 的 attention + 大 batch GEMM),W4A4 / FP8 才有意义;纯 weight-only 量化在 prefill 上几乎不省时间(甚至因 dequant overhead 反而变慢)。</p></div>
<h2 id="2-量化数学基础">§2 量化数学基础</h2>
<h3 id="21-affine-quantizationuniform--linear-量化">2.1 Affine quantizationuniform / linear 量化)</h3>
<p>把实数 $x \in [\alpha, \beta]$ 映射到整数 $q \in [Q_\min, Q_\max]$(如 INT8: $[-128, 127]$INT4: $[-8, 7]$)。</p>
<p>$$\boxed{\;q = \mathrm{clamp}\!\left(\mathrm{round}(x / s) + z,\; Q_\min,\; Q_\max\right),\quad \hat{x} = s\,(q - z)\;}$$</p>
<ul><li><strong>Scale</strong> $s = (\beta - \alpha) / (Q_\max - Q_\min)$<strong>zero-point</strong> $z = \mathrm{round}(Q_\min - \alpha/s)$。</li><li><strong>对称量化</strong>symmetric):$\alpha = -\beta$,强制 $z = 0$,反量化退化为 $\hat{x} = s\cdot q$。优点:no zero-point addition in GEMM;缺点:分布不对称时浪费 1 bit 表达范围。</li><li><strong>非对称量化</strong>asymmetric):保留 $z$,可精确覆盖任意 $[\alpha, \beta]$。代价:GEMM 内 $W\cdot X$ 展开时多 4 项(cross-zero-point 项),实际实现常用 "subtract zero-point first then GEMM" 把 cross 项归零。</li></ul>
<h3 id="22-量化误差与-snr">2.2 量化误差与 SNR</h3>
<p>舍入误差 $\epsilon = x - \hat{x}$ 在 $[-s/2, s/2]$ 上近似均匀分布,方差 $\mathbb{E}[\epsilon^2] = s^2/12$。信号 $\mathrm{Var}(x) = \sigma^2$,则量化 SNR</p>
<p>$$\mathrm{SNR} = 10 \log_{10} \frac{\sigma^2}{s^2/12} = 10 \log_{10}(12) + 20 \log_{10}\frac{\sigma}{s}$$</p>
<p>INT8 对 $\mathcal{N}(0, 1)$ 用 $\beta = 3\sigma$ 截断时,每减 1 bit SNR 降约 6 dB(每 bit 增加 1 倍量化 step → SNR 下降 $20\log_{10} 2 \approx 6$ dB)。但 LLM 实际衡量是 PPL / 任务指标,远比 SNR 复杂。</p>
<h3 id="23-粒度granularity">2.3 粒度(granularity</h3>
<table><thead><tr><th>粒度</th><th>Scale 共享范围</th><th>显存开销</th><th>精度</th></tr></thead><tbody><tr><td><strong>per-tensor</strong></td><td>整个张量一个 $s$</td><td>$O(1)$</td><td>最差(outlier 一坏全坏)</td></tr><tr><td><strong>per-channel</strong></td><td>weight 沿 output channel(行) / activation 沿 hidden(列)</td><td>$O(N)$</td><td>中等</td></tr><tr><td><strong>per-group</strong></td><td>每 $g$ 个元素共享一个 $s$(一般沿 input/K 维分组)</td><td>$O(NK/g)$</td><td>高($g = 32 / 64 / 128$</td></tr><tr><td><strong>per-token (act only)</strong></td><td>activation 每个 token (一行) 一个 $s$</td><td>$O(L)$</td><td>高,但每 step 算一次</td></tr></tbody></table>
<div class="callout callout-info"><div class="callout-title">Per-group 是 W4 量化的工业标准</div><p>GPTQ / AWQ / GGUF Q4_K 默认 group_size = 128:在 input dim 上每 128 个 weight 共享 scale+zero-point)。Storage 开销:W4 + group128INT4 quant + FP16 scale per 128 weights)≈ $4 + 16/128 = 4.125$ bits/weightgroup32 ≈ $4 + 16/32 = 4.5$ bits/weight,精度更高。</p></div>
<h3 id="24-rounding-模式">2.4 Rounding 模式</h3>
<ul><li><strong>Round-to-nearest, ties to even</strong> (RNE)IEEE 754 默认,无偏。LLM 量化一般默认用。</li><li><strong>Stochastic rounding</strong>$\mathrm{round}(x/s + u)$$u \sim \mathcal{U}[-0.5, 0.5)$。期望无偏,适合 QAT 反向传播。</li><li><strong>AdaRound</strong> (Nagel 2020 ICML):把舍入决策当作 $\{0, 1\}$ 二值优化变量,逐 layer 用代理 loss 学习("应该 round up 还是 down"),比 RTN 显著提升 PTQ。</li></ul>
<h3 id="25-简洁可运行代码非对称-int8-per-channel-qdq">2.5 简洁可运行代码:非对称 INT8 per-channel q/dq</h3>
<pre><code class="language-python">import torch
def quantize_per_channel_asym(x: torch.Tensor, n_bits: int = 8, channel_dim: int = 0):
&quot;&quot;&quot;
Asymmetric per-channel quantization.
Args:
x: float tensor, e.g. weight [out_features, in_features]
n_bits: target bit width (e.g. 8 -&gt; INT8)
channel_dim: which dim to share scale on (typically 0 for weight rows)
Returns:
q_int: int quantized tensor (still stored as int32 here for simplicity)
scale: [N] float scale per channel
zero_point: [N] int zero-point per channel
&quot;&quot;&quot;
q_min = -(1 &lt;&lt; (n_bits - 1)) # -128 for INT8
q_max = (1 &lt;&lt; (n_bits - 1)) - 1 # 127 for INT8
# Reduce over all dims except channel_dim
reduce_dims = [d for d in range(x.dim()) if d != channel_dim]
x_min = x.amin(dim=reduce_dims, keepdim=True)
x_max = x.amax(dim=reduce_dims, keepdim=True)
# Avoid degenerate (zero range) channels
eps = 1e-8
scale = (x_max - x_min).clamp(min=eps) / (q_max - q_min)
zero_point = (q_min - x_min / scale).round().clamp(q_min, q_max)
q = (x / scale + zero_point).round().clamp(q_min, q_max).to(torch.int32)
return q, scale, zero_point.to(torch.int32)
def dequantize_per_channel(q_int: torch.Tensor, scale: torch.Tensor, zero_point: torch.Tensor):
&quot;&quot;&quot; x_hat = scale * (q - zero_point) (per channel) &quot;&quot;&quot;
return scale * (q_int.to(scale.dtype) - zero_point.to(scale.dtype))
# Sanity check
if __name__ == &quot;__main__&quot;:
W = torch.randn(64, 1024) * 0.1
W[3, :] *= 10.0 # one outlier channel
q, s, z = quantize_per_channel_asym(W, n_bits=8, channel_dim=0)
W_hat = dequantize_per_channel(q, s, z)
err = (W - W_hat).abs().max().item()
print(f&quot;max abs err = {err:.4e} (should be ≤ max_scale/2)&quot;)</code></pre>
<div class="callout callout-warn"><div class="callout-title">PyTorch 早期版本的 `round` 是 banker&#x27;s rounding (RNE),与 CUDA / TensorRT 不一致</div><p>跨 backend 部署时务必用同一 rounding,否则同一份量化 weight 在两端产出会差 0.5 ULP。</p></div>
<h2 id="3-llm-的-outlier-问题">§3 LLM 的 Outlier 问题</h2>
<h3 id="31-经验观察dettmers-2022-llmint8">3.1 经验观察(Dettmers 2022, LLM.int8()</h3>
<p>模型规模超过 6.7B 后,每一层 attention input / FFN input activation 中,<strong>少数 hidden channel 的幅值是其他 channel 的 $50\text{-}100\times$</strong>,且:</p>
<ul><li><strong>稳定性</strong>:同一 channel 在所有 token、所有 sample 上都是 outlier(不是随机出现)。</li><li><strong>集中性</strong>:约 0.1%-1% 的 channel 贡献了几乎所有的最大值。</li><li><strong>后果</strong>per-tensor INT8 几乎所有 PPL 损失来自这些 channelper-channel weight + per-tensor activation INT8 在 &lt; 6.7B 模型上 OK,到 OPT-13B / 66B 上崩盘。</li></ul>
<h3 id="32-llmint8--第一个能落地的-8-bit-llm">3.2 LLM.int8() — 第一个能落地的 8-bit LLM</h3>
<p>Dettmers 2022 (NeurIPS) 的核心思路:<strong>Mixed-precision decomposition</strong>。在每一层把 activation 拆成两路:</p>
<ul><li><strong>Outlier path</strong>:选出 outlier channelthreshold $\alpha = 6.0$ 即 channel-max $> 6$),把对应的 weight 列和 activation 列<strong>保留 FP16</strong> 做矩阵乘。</li><li><strong>Normal path</strong>:其余 99% 的 channel 做 INT8 vector-wiseper-row activation × per-column weight)矩阵乘。</li><li><strong>合并</strong>:两路结果相加。</li></ul>
<p>数学等价于把矩阵分块:</p>
<p>$$Y = X W = \underbrace{X_O W_O}_{\text{FP16, outlier 列}} + \underbrace{X_N W_N}_{\text{INT8, 其余}}$$</p>
<p>效果:OPT-175B INT8 推理几乎零 PPL 损失,但 outlier path 的 FP16 GEMM 是 throughput 瓶颈(~10-15% 延迟),且 outlier mask 在每个 forward step 都要 detect。这促使后续工作(SmoothQuant / AWQ)转向"把 outlier 干掉而不是绕开"的思路。</p>
<h3 id="33-不同-outlier-形态gptq--awq--smoothquant-攻击的对象">3.3 不同 outlier 形态(GPTQ / AWQ / SmoothQuant 攻击的对象)</h3>
<table><thead><tr><th>Outlier 类型</th><th>沿哪一维稳定</th><th>谁能解决</th></tr></thead><tbody><tr><td><strong>Activation channel outlier</strong>hidden dim 维)</td><td>input channelK 维)</td><td>SmoothQuant(迁移到 weight, AWQper-channel scale 保护)</td></tr><tr><td><strong>Token outlier</strong>(少数 token 整行偏大)</td><td>sequence 维(L 维)</td><td>per-token activation quantzeroquant / SmoothQuant 默认)</td></tr><tr><td><strong>Weight outlier</strong>(少数 weight 偏大)</td><td>output dimN 维)</td><td>per-channel weight quant 即可吸收</td></tr><tr><td><strong>KV cache K-channel outlier</strong></td><td>K 的 head_dim 维</td><td>per-channel K quantKIVI / KVQuant</td></tr></tbody></table>
<h2 id="4-gptq基于-obs-的最优-weight-量化必考推导">§4 GPTQ:基于 OBS 的最优 weight 量化(必考推导)</h2>
<p>GPTQ (Frantar, Ashkboos, Hoefler, Alistarh, ICLR 2023) 是 weight-only PTQ 的工业标准。其数学基础是 <strong>Optimal Brain Surgeon (OBS)</strong>Hassibi &amp; Stork, NeurIPS 1992),把"删除/修改一个 weight 后如何最小化 loss 增量"推广到"量化一个 weight 后如何更新剩余 weight 以补偿"。</p>
<h3 id="41-问题设定">4.1 问题设定</h3>
<p>对单个 linear layer 输出 $Y = X W$$X \in \mathbb{R}^{B\times K}$ 来自 calibration set$W \in \mathbb{R}^{K \times N}$。量化目标:找 $\hat{W}$(每个元素是 INT4 / INT3)最小化:</p>
<p>$$\min_{\hat{W}} \|X W - X \hat{W}\|_F^2$$</p>
<p>这是关于 $\hat{W}$ 的 layer-wise reconstruction objective。注意:</p>
<ul><li>这是 <strong>layer-wise</strong>,不是 end-to-end lossPTQ 假设:layer-wise 重建好 → end-to-end loss 增量小,对线性 / 弱非线性架构基本成立)。</li><li>由于 GEMM 在 $N$ 维独立可分(每列 $w_j$ 自己一个最小化问题),把 $W$ 按列拆开是常见做法。下面对<strong>一列</strong> $w \in \mathbb{R}^K$ 推导。</li></ul>
<h3 id="42-二阶-taylor-展开">4.2 二阶 Taylor 展开</h3>
<p>设 $L(w) = \frac{1}{2}\|Xw - Xw^*\|^2$$w^*$ 是 FP16 原始 weight$w$ 待优化)。在 $w = w^*$ 处展开:</p>
<p>$$L(w^* + \delta) = \underbrace{L(w^*)}_{= 0} + \nabla L(w^*)^\top \delta + \frac{1}{2}\delta^\top H \delta + O(\|\delta\|^3)$$</p>
<p>由于 $L$ 在 $w^*$ 是全局极小,$\nabla L(w^*) = 0$。Hessian</p>
<p>$$H = X^\top X \in \mathbb{R}^{K\times K}$$</p>
<p>注意 $H$ 与 $w^*$ 无关(calibration data 算一次即可全列复用),且与 column index $j$ 无关(每列共享同一 $H$)。所以:</p>
<p>$$L(w^* + \delta) \approx \frac{1}{2}\delta^\top H \delta$$</p>
<h3 id="43-obs固定一个坐标到目标值后的最优-delta">4.3 OBS:固定一个坐标到目标值后的最优 $\delta$</h3>
<p>OBS 的关键问题:<strong>强制把 $w$ 的第 $q$ 个分量 $w_q$ 改成目标值 $w_q^{\mathrm{target}}$</strong>(在我们这里 $w_q^{\mathrm{target}} = \mathrm{Quant}(w_q)$),其他坐标如何调整才能最小化 $\delta^\top H \delta / 2$</p>
<p>约束 $e_q^\top \delta = w_q^{\mathrm{target}} - w_q^* := c_q$(其中 $e_q$ 是第 $q$ 个标准基)。用 Lagrangian</p>
<p>$$\mathcal{L}(\delta, \lambda) = \frac{1}{2}\delta^\top H \delta - \lambda (e_q^\top \delta - c_q)$$</p>
<p>求导:$\nabla_\delta \mathcal{L} = H\delta - \lambda e_q = 0 \Rightarrow \delta = \lambda H^{-1} e_q$。</p>
<p>代入约束 $e_q^\top \delta = c_q$</p>
<p>$$\lambda \cdot e_q^\top H^{-1} e_q = c_q \;\Rightarrow\; \lambda = \frac{c_q}{[H^{-1}]_{qq}}$$</p>
<p>所以:</p>
<p>$$\boxed{\;\delta^* = \frac{w_q^{\mathrm{target}} - w_q^*}{[H^{-1}]_{qq}}\; H^{-1} e_q\;}$$</p>
<p>即 $\delta^*$ 的第 $j$ 个分量 = $\dfrac{c_q}{[H^{-1}]_{qq}} \cdot [H^{-1}]_{jq}$。<strong>所有其他坐标的最优补偿全部来自 $H^{-1}$ 的第 $q$ 列</strong></p>
<p>最小损失增量:</p>
<p>$$\Delta L^* = \frac{1}{2}\delta^{*\top} H \delta^* = \frac{1}{2}\cdot \frac{c_q^2}{[H^{-1}]_{qq}}$$</p>
<div class="callout callout-good"><div class="callout-title">GPTQ 的最优 weight update(必考)</div><p>量化第 $q$ 列后,剩余未量化的所有列 $j > q$ 按下式更新一次:</p></div>
<p>$$w_j \leftarrow w_j - \frac{w_q - \mathrm{Quant}(w_q)}{[H^{-1}]_{qq}}\,[H^{-1}]_{jq}$$ 之后再量化 $q+1$ 列。这就是 GPTQ "iterative columnwise quantization" 的数学公式。</p>
<h3 id="44-工程加速cholesky-解-h-1-子矩阵">4.4 工程加速:Cholesky 解 $H^{-1}$ 子矩阵</h3>
<p>直接对每个 $q$ 求 $H^{-1}$ 的列代价 $O(K^2)$GPTQ 用 Cholesky 分解 $H^{-1} = U^\top U$$U$ 上三角,等价地 $H^{-1} = L L^\top$ 若取下三角 $L = U^\top$),扫到第 $q$ 列时只需 $U$ 的子矩阵,且更新可向量化。</p>
<p>GPTQ 的实际实现是把 $K$ 列按 <strong>block size = 128</strong> 分块,每个 block 内部按列 quantize 并 updateblock 间一次性 sync updateCholesky decomposition based)。整体复杂度:$O(K^3 + K \cdot K^2)$ per layer,对 7B 模型单 A100 上 ~30 分钟即可量化完。</p>
<h3 id="45-gptq-style-伪代码必考写法">4.5 GPTQ-style 伪代码(必考写法)</h3>
<pre><code class="language-python">import torch
@torch.no_grad()
def gptq_quantize_layer(
W: torch.Tensor, # [N, K] weight (one linear layer)
X: torch.Tensor, # [B, K] calibration input to this layer
n_bits: int = 4,
group_size: int = 128,
damp_percent: float = 0.01,
):
&quot;&quot;&quot;
Block-wise GPTQ for ONE linear weight matrix.
Hessian H = X^T X is shared across rows of W.
We quantize columns of W one-by-one (so we walk along K).
After quantizing column q, propagate the residual error
along H^-1 to all yet-unquantized columns j &gt; q.
&quot;&quot;&quot;
N, K = W.shape
device = W.device
H = X.t() @ X # [K, K]
# Damping: avoid singular H. Replace zero diag with mean diag * damp_percent.
mean_diag = torch.mean(torch.diag(H))
diag = torch.arange(K, device=device)
H[diag, diag] += damp_percent * mean_diag
# Cholesky on H^-1: GPTQ uses upper-triangular inverse Cholesky factor.
Hinv = torch.linalg.cholesky(torch.linalg.inv(H), upper=True)
Q = torch.zeros_like(W) # quantized W
W = W.clone() # we will mutate
q_min = -(1 &lt;&lt; (n_bits - 1))
q_max = (1 &lt;&lt; (n_bits - 1)) - 1
for q in range(K):
# Per-group scale: when entering a new group, recompute scale on W[:, q:q+group]
if q % group_size == 0:
end = min(q + group_size, K)
w_grp = W[:, q:end]
s = w_grp.abs().amax(dim=1, keepdim=True) / q_max
s = s.clamp(min=1e-8) # [N, 1]
w = W[:, q:q+1] # [N, 1] current column
# Quantize this column (symmetric, per-row scale `s`)
q_int = (w / s).round().clamp(q_min, q_max)
w_q = q_int * s # [N, 1] dequantized
Q[:, q:q+1] = w_q
# OBS-derived weight update on all columns to the right
err = (w - w_q) / Hinv[q, q] # [N, 1]
W[:, q+1:] -= err @ Hinv[q:q+1, q+1:] # [N, K-q-1]
return Q</code></pre>
<div class="callout callout-warn"><div class="callout-title">GPTQ 工程坑</div><p>(1) Calibration data 量:典型 128 samples × 2048 tokens,太少 Hessian 病态;(2) Damp 不能省,$H$ 经常有 zero diag(某些 input channel 在 calib set 上恒为 0);(3) Activation reorder<code>act_order=True</code>,按 $\mathrm{diag}(H)$ 降序量化)显著提升 W3 / W2 精度但增加 inference dequant 索引开销,AutoGPTQ 默认关。</p></div>
<h3 id="46-gptq-vs-早期-round-to-nearest-rtn">4.6 GPTQ vs 早期 round-to-nearest (RTN)</h3>
<table><thead><tr><th></th><th>RTN</th><th>GPTQ</th></tr></thead><tbody><tr><td>误差补偿</td><td>无(每个 weight 独立 round</td><td>有(OBS 公式向后传播误差)</td></tr><tr><td>Calibration</td><td>不需要</td><td>需要 128-512 samples</td></tr><tr><td>4-bit on LLaMA-7B</td><td>PPL +1.5</td><td>PPL +0.1</td></tr><tr><td>3-bit on LLaMA-7B</td><td>PPL +14 (崩)</td><td>PPL +0.7</td></tr><tr><td>时间</td><td>秒级</td><td>单卡 30-60 分钟(7B</td></tr></tbody></table>
<h2 id="5-awqactivation-aware-weight-quantization">§5 AWQActivation-Aware Weight Quantization</h2>
<p>AWQ (Lin, Tang, Tang, Yang, Chen, Wang, Xiao, Dang, Gan, Han, MLSys 2024) 的核心洞察:<strong>1% 的 "salient" weights 决定了几乎全部的量化损失</strong>——这些 salient weight 的输入 activation 幅值大。所以<strong>不应该独立量化所有 weight</strong>;要先把 salient channel 放大(量化前),等价地把对应 input scale 缩小,最终量化误差降低。</p>
<h3 id="51-salient-channel-的识别">5.1 Salient channel 的识别</h3>
<p>不是看 weight 自己的大小,而是看<strong>对应 activation channel 的幅值</strong></p>
<p>$$\text{salience}(c) = \mathrm{mean}_{x \sim \mathrm{calib}}\,|x_c|$$</p>
<p>把 channel 按 salience 排序,前 1% 是 "salient channel",对应 weight 列 $w_{\cdot, c}$ 是 "salient weight"。</p>
<h3 id="52-per-channel-scale-等价变换数学等价--量化误差等价">5.2 Per-channel scale 等价变换(数学等价 ≠ 量化误差等价)</h3>
<p>考虑 $Y = X W$$X \in \mathbb{R}^{B\times K}$$W \in \mathbb{R}^{K\times N}$。对每个 input channel $c$ 引入正 scale $s_c > 0$</p>
<p>$$Y = (X / S) (S \cdot W) = \tilde{X} \tilde{W}$$</p>
<p>其中 $S = \mathrm{diag}(s_1, \ldots, s_K)$$\tilde{X}_{:, c} = X_{:, c} / s_c$$\tilde{W}_{c, :} = s_c \cdot W_{c, :}$。<strong>在 FP16 下两者完全相等</strong>。但量化后:</p>
<ul><li>$\tilde{W} = S W$salient row $w_{c, :}$ 被乘以 $s_c$(变大),单 channel 量化更精细(per-channel scale 更小)。</li><li>$\tilde{X} = X / S$activation 没量化(weight-only PTQ 不动 activation),但下游若有 act quant 则误差也降。</li></ul>
<p><strong>问题</strong>:$s_c$ 怎么选?过大会让 salient row 太突出抢爆 scale;过小起不到保护效果。AWQ 给出 grid search</p>
<p>$$s_c = \mathrm{mean}(|x_c|)^\alpha,\quad \alpha \in \{0.0, 0.1, \ldots, 1.0\}$$</p>
<p>对每层独立 grid search $\alpha$,最小化 layer-wise MSE $\|Y_{\mathrm{fp}} - Y_{\mathrm{quant}}\|^2$。$\alpha = 0$ 退化为 RTN(无 scale);$\alpha = 1$ 直接用 mean act 做 scale;典型最优 $\alpha \in [0.4, 0.7]$。</p>
<h3 id="53-与-gptq-的区别">5.3 与 GPTQ 的区别</h3>
<table><thead><tr><th></th><th>GPTQ</th><th>AWQ</th></tr></thead><tbody><tr><td>数据依赖</td><td>Hessian $X^\top X$(需 calibration</td><td>$\text{mean}(\lvert x \rvert)$ per channel(也需 calibration</td></tr><tr><td>优化对象</td><td>所有 weight 的最优误差补偿</td><td>salient channel 的 scale</td></tr><tr><td>量化流程</td><td>iterative, columnwise + OBS update</td><td>one-shot scaling + RTN</td></tr><tr><td>时间</td><td>30-60 min (7B)</td><td>5-15 min (7B)</td></tr><tr><td>推理 dequant</td><td>可能需要 act reordering</td><td>无额外开销(scale 可吸收进 LayerNorm / W</td></tr><tr><td>与 W 重排兼容</td><td>较弱</td><td>强(scale 是 elementwise,不破坏 GEMM 结构)</td></tr></tbody></table>
<div class="callout callout-info"><div class="callout-title">AWQ 的工程亲和性</div><p>Per-channel $s_c$ 可以<strong>预 merge 进上游 LayerNorm / RMSNorm 的 weight</strong>$\mathrm{LN}(x) \cdot \gamma$ 中 $\gamma \leftarrow \gamma / s$,下游 $w \leftarrow s \cdot w$,运行时<strong>完全没有额外 elementwise 操作</strong>。这是 AWQ 比 SmoothQuant 更易部署的原因之一(SmoothQuant 也能 merge,但若 LN 后接 cat / residual 则不能)。</p></div>
<h3 id="54-awq-style-per-channel-scale-搜索代码">5.4 AWQ-style per-channel scale 搜索代码</h3>
<pre><code class="language-python">import torch
@torch.no_grad()
def awq_search_scale(
W: torch.Tensor, # [N, K] FP16 weight
X: torch.Tensor, # [B, K] calibration activations (post-LayerNorm)
n_bits: int = 4,
group_size: int = 128,
n_grid: int = 20,
):
&quot;&quot;&quot;
Search per-channel scale s in [0, 1] that minimizes layer-wise MSE
of dequantized W·X relative to FP16 W·X.
s_c = mean(|x_c|) ** alpha, alpha in {0, 1/n_grid, ..., 1}
&quot;&quot;&quot;
device = W.device
x_mean = X.abs().mean(dim=0).clamp(min=1e-5) # [K]
# Y_fp is the FP16 reference output (compute once)
Y_fp = X @ W.t() # [B, N]
q_min = -(1 &lt;&lt; (n_bits - 1))
q_max = (1 &lt;&lt; (n_bits - 1)) - 1
best_alpha, best_err = None, float(&quot;inf&quot;)
for i in range(n_grid + 1):
alpha = i / n_grid
s = x_mean.pow(alpha)
# Normalize s so that geometric mean is 1: keeps scale of W stable.
s = s / s.mean()
s = s.clamp(min=1e-4)
# Apply: W&#x27; = W · diag(s), X&#x27; = X / s
Wp = W * s.unsqueeze(0) # [N, K]
# Group-wise symmetric quantize Wp along K dim
Wq = _group_quant_dequant(Wp, n_bits, group_size, dim=-1, q_min=q_min, q_max=q_max)
# Equivalent activation rescaling: divide each input channel by s
Xp = X / s.unsqueeze(0)
Y_q = Xp @ Wq.t()
err = (Y_fp - Y_q).pow(2).mean().item()
if err &lt; best_err:
best_err, best_alpha, best_s = err, alpha, s
return best_alpha, best_s, best_err
def _group_quant_dequant(W, n_bits, g, dim, q_min, q_max):
&quot;&quot;&quot; Symmetric per-group quant-dequant along last dim (assumed K). Assumes K % g == 0. &quot;&quot;&quot;
N, K = W.shape
assert K % g == 0, f&quot;K={K} must be divisible by group_size={g}; pad W or pick g | K.&quot;
Wg = W.view(N, K // g, g) # [N, K/g, g]
s = Wg.abs().amax(dim=-1, keepdim=True) / q_max
s = s.clamp(min=1e-8)
Wq = (Wg / s).round().clamp(q_min, q_max) * s
return Wq.view(N, K)</code></pre>
<h2 id="6-smoothquant把-activation-outlier-迁移到-weight-w8a8">§6 SmoothQuant:把 activation outlier 迁移到 weight (W8A8)</h2>
<p>SmoothQuant (Xiao, Lin, Seznec, Wu, Demouth, Han, ICML 2023) 解决 weight + activation 同时量化(W8A8)的核心痛点:activation outlier 让 per-tensor INT8 崩盘。</p>
<h3 id="61-核心数学">6.1 核心数学</h3>
<p>对 $Y = X W$,引入<strong>对角 smoothing matrix</strong> $S = \mathrm{diag}(s_1, \ldots, s_K)$$s_c > 0$</p>
<p>$$Y = (X S^{-1})(S W) = \hat{X} \hat{W}$$</p>
<ul><li>$\hat{X}_{:,c} = X_{:,c} / s_c$activation 的 outlier channel 被压平。</li><li>$\hat{W}_{c,:} = s_c \cdot W_{c,:}$weight 的对应行被放大(吸收 outlier 量级)。</li></ul>
<p>数学上<strong>完全等价</strong>FP16 下逐元素相等);但量化后:</p>
<ul><li>$\hat{X}$ 的最大值显著下降,per-tensor INT8 的有效 bit width 提升。</li><li>$\hat{W}$ 因为 weight 本身<strong>分布平坦</strong><strong>可以 per-channel quant</strong>per output dim 维度,与 $S$ 操作的 K 维正交),少数 row 被放大无伤大雅。</li></ul>
<h3 id="62-migration-strength-alpha关键超参">6.2 Migration strength $\alpha$(关键超参)</h3>
<p>最优 $s_c$ 应平衡"activation 平滑了多少"和"weight 被放大多少"</p>
<p>$$\boxed{\;s_c = \dfrac{\max(|X_{:, c}|)^\alpha}{\max(|W_{c, :}|)^{1 - \alpha}}\;}$$</p>
<ul><li>$\alpha = 0$$s_c = 1/\max|W_c|$。$\hat X = X \cdot \max|W_c|$activation 被放大,量化反而更难),$\hat W = W / \max|W_c|$weight 被归一化,量化变易)。<strong>burden 全压在 activation 上</strong></li><li>$\alpha = 1$$s_c = \max|X_c|$。$\hat X = X / \max|X_c|$activation 被压平,量化变易),$\hat W = W \cdot \max|X_c|$weight 被放大,量化变难)。<strong>burden 全压在 weight 上</strong></li><li>$\alpha = 0.5$(默认):$s_c = \sqrt{\max|X_c| / \max|W_c|}$,两边各让一半。</li></ul>
<p>SmoothQuant 论文在 OPT / BLOOM 上扫 $\alpha \in [0.3, 0.7]$typically 0.5 即可。</p>
<h3 id="63-等价变换不破坏-gemm必考">6.3 等价变换不破坏 GEMM(必考)</h3>
<div class="callout callout-good"><div class="callout-title">为什么 SmoothQuant migration 不破坏 GEMM</div><p>等价变换 $Y = X W = (XS^{-1})(SW)$<strong>$S$ 是对角矩阵 (per-channel scale),对 $X$ 的列 / $W$ 的行做 elementwise rescale</strong></p></div>
<ul><li>数学等价:对角矩阵作用是 channelwise multiplication,与 GEMM 的内积顺序无关,最终输出 $Y$ 在 FP16 下逐元素相等。</li><li>工程实现:$S^{-1}$ 可以 fuse 进<strong>上一层的 LayerNorm weight</strong>$\gamma \leftarrow \gamma / s$),$S$ 可以 fuse 进<strong>本层 weight</strong>$W \leftarrow SW$,离线一次性),inference 时<strong>完全没有 elementwise overhead</strong></li><li>不破坏的关键:$S$ 是 diagonalrescale 在 K 维上每个 channel 独立。如果 $S$ 是 dense / rotation matrix,则需 explicit matmul,下面的 QuaRot / SpinQuant 选择了那个方向,但代价更大。</li></ul>
<h3 id="64-smoothquant-伪代码">6.4 SmoothQuant 伪代码</h3>
<pre><code class="language-python">import torch
@torch.no_grad()
def compute_smooth_scale(
X: torch.Tensor, # [B*L, K] per-token flattened activations (calibration)
W: torch.Tensor, # [N, K] weight
alpha: float = 0.5,
):
&quot;&quot;&quot;
SmoothQuant migration scale (per input channel c):
s_c = max(|x_c|)^alpha / max(|w_c|)^(1-alpha)
&quot;&quot;&quot;
x_max = X.abs().amax(dim=0) # [K] per channel
w_max = W.abs().amax(dim=0) # [K] per input ch of W
s = (x_max.pow(alpha) / w_max.pow(1 - alpha)).clamp(min=1e-5)
return s # [K]
@torch.no_grad()
def apply_smoothing(W: torch.Tensor, s: torch.Tensor, prev_ln_weight: torch.Tensor):
&quot;&quot;&quot;
Fuse smoothing into upstream LayerNorm weight and current layer W:
gamma_new = gamma / s (so output of LN becomes x / s)
W_new = W * s (broadcasts over output dim of W)
After this, FP16 forward is identical, but X and W are reshaped
such that simple per-tensor / per-channel quant works well.
&quot;&quot;&quot;
prev_ln_weight.div_(s) # in-place modify γ
W.mul_(s.unsqueeze(0)) # [N, K] broadcasts s along K dim
return W</code></pre>
<h3 id="65-smoothquant-适用范围">6.5 SmoothQuant 适用范围</h3>
<ul><li>✅ Decoder-only LLM (OPT, BLOOM, LLaMA family)W8A8 几乎无损(PPL +0.1)。</li><li>✅ FFN 和 attention input projectionK, V, Q):smoothing 可与上游 LN merge。</li><li>⚠️ Out projection 和 down projection:上游不是 LN(是 residual / attention output),smoothing 需要 explicit elementwise op,工程上 SmoothQuant 默认跳过这两层(保留 FP16 input)。</li><li>⚠️ INT4 activationW4A4 用 SmoothQuant 单独不够,需要配合 QuaRot / SpinQuant 旋转。</li></ul>
<h2 id="7-旋转方法quip--quarot--spinquant">§7 旋转方法:QuIP / QuaRot / SpinQuant</h2>
<p>SmoothQuant 用 <strong>diagonal</strong>per-channelscale 抑制 outlier;但 outlier 仍存在于某些 channel 子空间。<strong>Rotation methods</strong> 用一个 random / learned <strong>正交矩阵</strong> $R$ 把 outlier 在 hidden dim 上"打散",让分布更接近 Gaussian。</p>
<h3 id="71-quip-chee-et-al-2023-neurips">7.1 QuIP (Chee et al. 2023, NeurIPS)</h3>
<p>核心:用一个随机的 <strong>incoherence-inducing 矩阵</strong> $U$(例如随机 Hadamard 或 Householder)旋转 weight,使量化更友好。</p>
<ul><li><strong>Hadamard 变换</strong>$H \in \mathbb{R}^{d \times d}$$H_{ij} \in \{+1, -1\} / \sqrt{d}$,正交矩阵。</li><li>关键性质:随机 sign flip 后的 Hadamard 变换可证明把 weight 的 incoherence(最大 column $\ell_2$ norm 与 Frobenius norm 比)压到 $O(\sqrt{\log d / d})$ 级别。</li><li>$W' = U W V^\top$FP16 等价($U, V$ orthogonal);量化 $W'$ 比量化 $W$ 损失小(因为 incoherent)。</li></ul>
<p>代价:inference 时需保留 $U, V$ 的 matmul(一次 dense rotation)。Hadamard 变换有 fast algorithm$O(d \log d)$),但仍比 SmoothQuant 的对角 fuse 慢。</p>
<h3 id="72-quarot-ashkboos-et-al-2024-neurips">7.2 QuaRot (Ashkboos et al. 2024 NeurIPS)</h3>
<p>把 Hadamard 推广到 LLM 全栈:<strong>weight + activation + KV cache 全 INT4</strong>。核心:</p>
<ul><li>在每个 residual stream 进入 transformer block 前,<strong>乘上一个 Hadamard $H$</strong></li><li>Hadamard 是 orthogonal,可以"穿过" RMSNormRMSNorm 是 elementwise$\mathrm{RMSNorm}(Hx) \cdot \gamma = H \cdot \mathrm{RMSNorm}(x) \cdot (H \gamma)$ 不严格成立,但 QuaRot 用"online Hadamard"绕过)。</li><li>Activation 旋转后呈现更 Gaussian 的分布,<strong>outlier 被打散到所有维度</strong>INT4 activation 量化误差大幅下降。</li></ul>
<p>效果:LLaMA-2 70B 在 W4A4KV4 下 PPL 增量约 +0.5vs SmoothQuant 的 +5)。代价:每个 block 需要 1-2 次 Hadamard matmul(实际上 fast Hadamard transform 在 H100 上很便宜)。</p>
<h3 id="73-spinquant-liu-et-al-2024--iclr-2025-meta">7.3 SpinQuant (Liu et al. 2024 → ICLR 2025, Meta)</h3>
<p>把 QuaRot 的随机 Hadamard 换成<strong>学习的旋转矩阵</strong> $R_1, R_2, R_3, R_4$,分别作用于 residual stream / attention input / FFN input / KV cache。优化目标:layer-wise output MSE。</p>
<ul><li>$R_i \in SO(d)$(特殊正交群),用 Cayley parameterization 或 stochastic gradient on Stiefel manifold 优化。</li><li>比 QuaRot 多 ~0.5 PPL 改进,但训练时间增加(每个模型需 ~1 GPU 小时学 R)。</li></ul>
<div class="callout callout-info"><div class="callout-title">旋转方法 vs SmoothQuant</div><p>Smoothing 解决"channel 维 outlier"rotation 解决"channel-subspace outlier"。Rotation 更通用,但工程成本更高(dense matmul 不能 fuse 进 LN,需要 online compute 或 explicit kernel)。LLaMA-3 / Qwen-2 部署上 W4A8KV4 主流仍是 SmoothQuant + GPTQW4A4 才需要 QuaRot / SpinQuant 级别旋转。</p></div>
<h2 id="8-低精度浮点fp8--mx--nvfp4">§8 低精度浮点:FP8 / MX / NVFP4</h2>
<p>低精度浮点不是新事物(FP16 / BF16 已普遍),新的是 <strong>FP8 (E4M3 / E5M2)</strong> 在 Hopper 上原生 tensor core 支持,以及 <strong>MX / NVFP4</strong> 在 Blackwell 上的 block-scaled 浮点。</p>
<h3 id="81-ieee-754-style-浮点编码">8.1 IEEE 754-style 浮点编码</h3>
<p>一个浮点数 $x = (-1)^s \cdot (1 + m) \cdot 2^{e - \text{bias}}$normal)或 $x = (-1)^s \cdot m \cdot 2^{1 - \text{bias}}$subnormal)。</p>
<table><thead><tr><th>格式</th><th>Sign</th><th>Exp bits</th><th>Mantissa bits</th><th>Bias</th><th>Max</th><th>Min normal</th></tr></thead><tbody><tr><td>FP32</td><td>1</td><td>8</td><td>23</td><td>127</td><td>$\sim 3.4\times 10^{38}$</td><td>$\sim 1.2\times 10^{-38}$</td></tr><tr><td>FP16</td><td>1</td><td>5</td><td>10</td><td>15</td><td>65504</td><td>$\sim 6.1\times 10^{-5}$</td></tr><tr><td>BF16</td><td>1</td><td>8</td><td>7</td><td>127</td><td>$\sim 3.4\times 10^{38}$</td><td>$\sim 1.2\times 10^{-38}$</td></tr><tr><td><strong>FP8 E4M3</strong></td><td>1</td><td>4</td><td>3</td><td>7</td><td>448 (no Inf)</td><td>$\sim 1.5\times 10^{-2}$</td></tr><tr><td><strong>FP8 E5M2</strong></td><td>1</td><td>5</td><td>2</td><td>15</td><td>57344</td><td>$\sim 6.1\times 10^{-5}$</td></tr><tr><td><strong>FP4 E2M1</strong></td><td>1</td><td>2</td><td>1</td><td>1</td><td>6</td><td>1</td></tr></tbody></table>
<div class="callout callout-good"><div class="callout-title">E4M3 vs E5M2 forward/backward 选择</div><p>NVIDIA Transformer Engine 默认:</p></div>
<ul><li><strong>Forward (activation, weight)</strong>:用 <strong>E4M3</strong>——dynamic range 小但分辨率高 (mantissa 3 bit)。Activation 经 layer-wise scale 后落在 $[-448, 448]$ 内,mantissa 精度更重要。</li><li><strong>Backward (gradient)</strong>:用 <strong>E5M2</strong>——dynamic range 大但分辨率低 (mantissa 2 bit, 形态接近 FP16)。Gradient 量级跨越多个数量级,需要大动态范围。</li></ul>
<p>注意 E4M3 没有 Inf,只有 NaN(最大 normal 即 448);E5M2 有 Inf 和 NaN(与 FP16 类似)。这点 hardware 设计上有意区分。</p>
<h3 id="82-fp8-e4m3-bit-level-encoding-代码">8.2 FP8 E4M3 bit-level encoding 代码</h3>
<pre><code class="language-python">def fp8_e4m3_encode(x: float) -&gt; int:
&quot;&quot;&quot;
Encode a float into FP8 E4M3 8-bit pattern (returned as int 0..255).
Format: 1 sign | 4 exponent | 3 mantissa, bias = 7, no Inf, NaN = S.1111.111
Subnormals: exp = 0, value = (-1)^s * (mantissa/8) * 2^(1-7) = (-1)^s * (m/8) * 2^-6
Max normal: S.1111.110 -&gt; 448.0
&quot;&quot;&quot;
import math
if math.isnan(x):
return 0b0_1111_111
sign = 0 if x &gt;= 0 else 1
ax = abs(x)
if ax == 0:
return sign &lt;&lt; 7
if ax &gt;= 448.0:
return (sign &lt;&lt; 7) | 0b1111_110 # saturate to max normal
# Decompose ax = m * 2^e with m in [1, 2)
e = int(math.floor(math.log2(ax)))
m = ax / (2 ** e) # m in [1, 2)
# Adjust to FP8 E4M3 representation
biased_e = e + 7 # bias = 7
if biased_e &lt;= 0:
# Subnormal: shift mantissa right by (1 - biased_e), set exp = 0
shift = 1 - biased_e
m_int = int(round((m * 2 ** -shift) * 8)) # 3-bit mantissa
exp_bits = 0
else:
m_int = int(round((m - 1.0) * 8)) # 3-bit mantissa
if m_int == 8: # mantissa overflow
m_int = 0
biased_e += 1
if biased_e &gt;= 15: # exceeds max exp
return (sign &lt;&lt; 7) | 0b1111_110 # saturate
exp_bits = biased_e
return (sign &lt;&lt; 7) | (exp_bits &lt;&lt; 3) | m_int
def fp8_e4m3_decode(b: int) -&gt; float:
&quot;&quot;&quot; Decode an 8-bit FP8 E4M3 pattern (int 0..255) back to a Python float. &quot;&quot;&quot;
sign = -1.0 if (b &gt;&gt; 7) &amp; 1 else 1.0
exp_bits = (b &gt;&gt; 3) &amp; 0b1111
m_bits = b &amp; 0b111
if exp_bits == 0b1111 and m_bits == 0b111:
return float(&quot;nan&quot;)
if exp_bits == 0: # subnormal
return sign * (m_bits / 8.0) * (2 ** -6)
return sign * (1.0 + m_bits / 8.0) * (2 ** (exp_bits - 7))
# Sanity check: round-trip 448 (max normal)
assert abs(fp8_e4m3_decode(fp8_e4m3_encode(448.0)) - 448.0) &lt; 1e-6</code></pre>
<h3 id="83-mx-格式ocp--microsoft-2024">8.3 MX 格式(OCP / Microsoft 2024</h3>
<p>OCP (Open Compute Project) MX (Microscaling) 规范:把 32 个元素组成一个 block,共享一个 <strong>8-bit shared scale</strong>E8M0 格式,即 power-of-two scale);block 内每个元素用 FP4/FP6/FP8 编码。</p>
<table><thead><tr><th>MX 格式</th><th>Element type</th><th>Block size</th><th>Shared scale</th><th>总 bits/element</th></tr></thead><tbody><tr><td>MXFP8</td><td>FP8 (E5M2 or E4M3)</td><td>32</td><td>E8M0</td><td>$8 + 8/32 = 8.25$</td></tr><tr><td>MXFP6</td><td>FP6 (E3M2 or E2M3)</td><td>32</td><td>E8M0</td><td>$6 + 8/32 = 6.25$</td></tr><tr><td>MXFP4</td><td>FP4 (E2M1)</td><td>32</td><td>E8M0</td><td>$4 + 8/32 = 4.25$</td></tr><tr><td>MXINT8</td><td>INT8</td><td>32</td><td>E8M0</td><td>$8.25$</td></tr></tbody></table>
<p>E8M0 是 1 字节、纯指数(无 mantissa、无 sign)的 power-of-two scale$s = 2^{e - 127}$$e \in [0, 255]$。这种 scale 在 dequant 时是 bit shift(最便宜的硬件操作)。</p>
<h3 id="84-nvfp4blackwell-2025-nvidia">8.4 NVFP4Blackwell 2025 NVIDIA</h3>
<p>NVFP4 是 NVIDIA 在 BlackwellB100 / B200 / GB200)上推的 FP4 格式,与 OCP MXFP4 区别:</p>
<ul><li><strong>Element</strong>FP4 E2M1(与 MX 相同)。</li><li><strong>Block size</strong><strong>16</strong>(不是 32),更细粒度。</li><li><strong>Per-block scale</strong><strong>FP8 E4M3</strong>(而非 E8M0),保留 mantissascale 自身精度更高。</li><li><strong>Per-tensor scale</strong>:额外一个 FP32 全局 scaleblock scale 是 FP8 受 $\pm 448$ 限制,per-tensor scale 拉大动态范围)。</li></ul>
<p>总 bits/element$4 + 8/16 + \text{negligible per-tensor} \approx 4.5$ bits。Blackwell tensor core 原生支持 NVFP4 × NVFP4 matmulthroughput 在 B200 上号称 FP16 的 $\sim 8\times$。</p>
<div class="callout callout-warn"><div class="callout-title">NVFP4 ≠ MXFP4</div><p>工业界经常混用 "FP4" 字眼。NVFP4block=16, FP8 E4M3 scale, +FP32 tensor scale)是 NVIDIA Blackwell 专有;OCP MXFP4block=32, E8M0 scale)是开放规范,AMD MI350 / Intel Gaudi 3 部分支持。两者在数值精度和硬件路径上不通用。</p></div>
<h3 id="85-fp8-在大模型训练中的使用transformer-engine">8.5 FP8 在大模型训练中的使用(Transformer Engine</h3>
<p>NVIDIA Transformer Engine (TE) 在 Hopper / Blackwell 上用 FP8 训 LLM 的标准做法:</p>
<ul><li><strong>每个 GEMM 维护两个 amax history</strong>forward / backward 各一),每 8-16 step 更新一次 scale。</li><li><strong>Delayed scaling</strong>:用上一窗口的 amax 算 scale,避免阻塞当前 step。</li><li><strong>Hybrid precision</strong>linear weight 和 grad accumulator 仍 FP32 / BF16,只有 GEMM 的输入 cast 到 FP8。Optimizer stateAdam $m, v$)保持 FP32。</li><li><strong>Loss scaling</strong>:与 FP16 类似,但因 E5M2 dynamic range 已经接近 FP16loss scale 可设较小或省略。</li></ul>
<p>LLaMA-3 / DeepSeek-V3 系列 FP8 训练在 Hopper 上吞吐比 BF16 提升 ~1.5-2$\times$。</p>
<h2 id="9-kv-cache-量化">§9 KV Cache 量化</h2>
<p>LLM inference decode 阶段,per-sample KV cache 显存 = $L_\text{ctx} \cdot 2 \cdot n_\text{layers} \cdot H_\text{kv} \cdot d_\text{head} \cdot \text{bytes}$。LLaMA-2-70B (80 层, GQA $H_\text{kv}=8$, $d_\text{head}=128$, FP16, $L_\text{ctx}=4096$)</p>
<p>$$4096 \times 2 \times 80 \times 8 \times 128 \times 2\text{B} = 1.34\text{ GB / sample}$$</p>
<p>batch 64 → 86 GBA100 80GB 一卡塞不下,必须 KV 量化)。</p>
<h3 id="91-kivi--kvquant-的关键观察">9.1 KIVI / KVQuant 的关键观察</h3>
<p>经验:</p>
<ul><li><strong>K cache</strong> 的 outlier 沿 <strong>head_dim</strong> 维度 (channel 维) <strong>稳定出现</strong>(与 RoPE 编码的相位有关,某些 freq band 量级大)。</li><li><strong>V cache</strong> 的 outlier 沿 <strong>token</strong> 维度(sequence 维)出现,且每个 token 独立。</li></ul>
<p>所以最优粒度:</p>
<ul><li><strong>K</strong><strong>per-channel</strong> quant(每个 head_dim 一个 scale)。</li><li><strong>V</strong><strong>per-token</strong> quant(每个 sequence position 一个 scale)。</li></ul>
<h3 id="92-kivi-liu-et-al-2024-icml">9.2 KIVI (Liu et al. 2024 ICML)</h3>
<p>KIVI = "K per-channel + V per-token" + INT2 quant,搭配 sliding window outlier residual。流程:</p>
<ol><li>K cache update 时按 head_dim 维度计算 scale<strong>per-channel</strong>),quant 到 INT2/INT4。</li><li>V cache update 时按 token 维度计算 scale<strong>per-token</strong>),quant 到 INT2/INT4。</li><li>最近 $W$ 个 token 保留 FP16sliding window),避免 quant 噪声主导最新 attention。</li></ol>
<p>LLaMA-2-7B INT2 KVPPL 增量 $\sim 0.5$KV cache 本身 FP16→INT2 理论 $8\times$ 压缩,去除 scale / outlier residual 开销后实测 KIVI 论文报告 peak memory(含 weight + activation)约 $2.35\text{-}2.6\times$ 降低、batch size 可放大 $\sim 4\times$。</p>
<h3 id="93-kvquant-hooper-et-al-2024-neurips">9.3 KVQuant (Hooper et al. 2024, NeurIPS)</h3>
<p>进一步分析:</p>
<ul><li>Per-channel K 不够,<strong>Pre-RoPE quant</strong> 比 post-RoPE 更稳(RoPE 引入相位混合,破坏 channel 维结构)。</li><li>V 用 per-token + non-uniform quantdensity-aware)。</li></ul>
<p>效果:LLaMA-2-70B INT4 KV PPL 增量 ~0.04。</p>
<h3 id="94-qserve--qoq-lin-et-al-2024-mlsys-2025">9.4 QServe / QoQ (Lin et al. 2024 MLSys 2025)</h3>
<p>QServe 推出 <strong>W4A8KV4</strong> 全栈量化 + 自定义 GPU kernel。关键工程点:</p>
<ul><li><strong>W4A8 GEMM</strong>weight INT4, activation INT8。dequant 路径:每个 weight 在 register 内用 lookup table 展开成 INT8 后做 INT8×INT8 matmul,无 FP16 dequant 开销。</li><li><strong>KV4 attention</strong>K 和 V 都 INT4,与 INT8 query 做 mixed-precision dot product。</li><li><strong>QoQ (quattuor-octo-quattuor)</strong>4+8+4 命名,4-bit weight, 8-bit activation, 4-bit KV。</li></ul>
<p>QServe 在 A100 / H100 上比 vanilla TensorRT-LLM FP16 throughput 提升 1.2-3.5$\times$,端到端 LLaMA-3-70B-Instruct 解码达 1000+ tokens/s/H100。</p>
<h3 id="95-kv-cache-quant-代码示意per-channel-k-per-token-v">9.5 KV cache quant 代码示意(per-channel K, per-token V</h3>
<pre><code class="language-python">import torch
@torch.no_grad()
def quantize_kv_cache(
K: torch.Tensor, # [B, H_kv, L, d_head]
V: torch.Tensor, # [B, H_kv, L, d_head]
n_bits: int = 4,
):
&quot;&quot;&quot;
K: per-channel (head_dim) symmetric quant
V: per-token (sequence) symmetric quant
Returns int8-stored tensors plus scales.
For real deployment, pack two INT4 values into one INT8 byte.
&quot;&quot;&quot;
q_min = -(1 &lt;&lt; (n_bits - 1))
q_max = (1 &lt;&lt; (n_bits - 1)) - 1
# ---- K: per-channel along the last dim (d_head). Same scale across L, H_kv per-batch. ----
# Common choice: per (batch, head, channel) scale, reduce over L only.
s_K = K.abs().amax(dim=2, keepdim=True) / q_max # [B, H_kv, 1, d_head]
s_K = s_K.clamp(min=1e-8)
K_q = (K / s_K).round().clamp(q_min, q_max).to(torch.int8)
# ---- V: per-token, scale per (batch, head, token) reducing over d_head ----
s_V = V.abs().amax(dim=-1, keepdim=True) / q_max # [B, H_kv, L, 1]
s_V = s_V.clamp(min=1e-8)
V_q = (V / s_V).round().clamp(q_min, q_max).to(torch.int8)
return K_q, s_K, V_q, s_V
def dequantize_kv(K_q, s_K, V_q, s_V, dtype=torch.float16):
K = K_q.to(dtype) * s_K.to(dtype)
V = V_q.to(dtype) * s_V.to(dtype)
return K, V</code></pre>
<div class="callout callout-warn"><div class="callout-title">RoPE 前还是 RoPE 后量化?</div><p>学术界共识:<strong>RoPE 前量化 K</strong>KVQuant 主张)。原因:RoPE 是 frequency-band 上的 rotation,把 channel 维度上的 outlier"打散"到其他 dim,破坏 per-channel scale 的稳定性。RoPE 前每个 head_dim 的 outlier 是固定 channelpost-RoPE 则在每个 token 上不同。但 RoPE 前 quant 需要在 attention kernel 内 dequant 再做 RoPE,工程上 fuse 比较麻烦;折中方案:post-RoPE 但用更细 group_size(如 32)。</p></div>
<h2 id="10-qat-与训练时量化">§10 QAT 与训练时量化</h2>
<p>PTQ (Post-Training Quantization) 不动 weightQAT (Quantization-Aware Training) 在训练或 finetune 时模拟量化,让模型适应。</p>
<h3 id="101-ste-straight-through-estimator">10.1 STE (Straight-Through Estimator)</h3>
<p>Round / clamp 在数学上是不可导(round 的导数几乎处处为 0),反向传播无信号。<strong>STE</strong> 把量化-反量化函数 $\mathrm{QDQ}(x) = s\,(\mathrm{clamp}(\mathrm{round}(x/s), Q_\min, Q_\max))$(对称量化为例)的梯度近似为:</p>
<p>$$\frac{\partial \mathrm{QDQ}(x)}{\partial x} \;\overset{\text{STE}}{:=}\; \mathbf{1}\!\left[\,s\,Q_\min \le x \le s\,Q_\max\,\right]$$</p>
<p>即"前向用量化值,反向在 clipping 范围内 pass-through 梯度(饱和区梯度置 0"。这是 LSQ / DoReFa / PACT 等 QAT 方法的基础。</p>
<h3 id="102-llm-qat-liu-et-al-2023">10.2 LLM-QAT (Liu et al. 2023)</h3>
<ul><li><strong>自蒸馏 data</strong>teacher 是 FP16 模型本身生成的 sequence)做 QAT,避免下载额外训练数据。</li><li>在每个 forward step 内 simulate INT4 weight quant,反向用 STE。</li><li>适合 W4A8 / W4A4 finetune 几千 step 后 PPL 接近 FP16。</li></ul>
<p>代价:QAT 比 PTQ 慢 100-1000$\times$。Production 上 PTQ (GPTQ + AWQ) 已经够好,QAT 主要用于 &lt; 4-bitW2A4 / W1.58 ternary 等)。</p>
<h3 id="103-fp8-trainingtransformer-engine">10.3 FP8 TrainingTransformer Engine</h3>
<p>见 §8.5。FP8 training 是 QAT 的特例:训练全程用 FP8 GEMMscale 用 amax history 周期更新,loss / opt state 仍 FP32。</p>
<h3 id="104-bitnet-b158--b2">10.4 BitNet b1.58 / b2</h3>
<p>最近 (Ma et al. 2024) 微软推 <strong>BitNet b1.58</strong>weight 是 ternary $\{-1, 0, +1\}$$\log_2 3 \approx 1.58$ bits),activation INT8。需要 from-scratch QAT 训练(不能 PTQ 转换),3B 规模与 FP16 LLaMA 持平。这是当前最低 bit 的 production-ready LLM 量化方案。</p>
<h2 id="11-框架与生态对照">§11 框架与生态对照</h2>
<table><thead><tr><th>框架</th><th>量化方法支持</th><th>推理后端</th><th>典型用例</th></tr></thead><tbody><tr><td><strong>bitsandbytes</strong></td><td>LLM.int8(), NF4, FP4</td><td>PyTorch + Triton</td><td>HuggingFace transformers 集成;QLoRA finetune 必备</td></tr><tr><td><strong>AutoGPTQ</strong></td><td>GPTQ (W4, W3, W2)</td><td>ExLlama / Marlin kernels</td><td>4-bit 推理 ~2× FP16 throughput</td></tr><tr><td><strong>AutoAWQ</strong></td><td>AWQ (W4)</td><td>GEMM kernel</td><td>比 GPTQ 略快、精度相当</td></tr><tr><td><strong>llama.cpp / GGUF</strong></td><td>Q4_K, Q5_K, Q6_K, Q8_0, Q3_K, IQ2_XXS...</td><td>CPU + GPU + Metal + ROCm</td><td>端侧推理首选</td></tr><tr><td><strong>TensorRT-LLM</strong></td><td>INT8 SmoothQuant, FP8, W4A8, NVFP4</td><td>NVIDIA fused kernel</td><td>生产服务首选(Hopper / Blackwell</td></tr><tr><td><strong>vLLM</strong></td><td>GPTQ, AWQ, FP8, INT8 SmoothQuant</td><td>PagedAttention + 自定义 kernel</td><td>多用户 serving 首选</td></tr><tr><td><strong>SGLang</strong></td><td>GPTQ, AWQ, FP8, W4A8 KV</td><td>radix-tree + 自定义 kernel</td><td>latency-sensitive serving</td></tr><tr><td><strong>Transformer Engine</strong></td><td>FP8 training + inference</td><td>H100 / B100 cuBLAS</td><td>FP8 训练首选</td></tr></tbody></table>
<div class="callout callout-info"><div class="callout-title">Marlin kernel</div><p>Frantar 2024 的 W4A16 GEMM kernel,针对 Ampere / Ada / Hopper 设计,4-bit weight + FP16 activation 在 batch 1-32 上比 FP16 cuBLAS 快 1.5-2$\times$。vLLM / SGLang 默认 W4 路径。</p></div>
<h2 id="12-25-高频面试题">§12 25 高频面试题</h2>
<p>codex (gpt-5.5 xhigh) 顶级 lab 面试官视角列的,按难度分 3 档。每题点开看答案要点 + 易踩坑。</p>
<h3 id="l1必会题任何-ml-工程岗都会问">L1必会题(任何 ML 工程岗都会问)</h3>
<details>
<summary>Q1. Affine 量化的 quant / dequant 公式?</summary>
<ul><li>Quant$q = \mathrm{clamp}(\mathrm{round}(x/s) + z,\; Q_\min,\; Q_\max)$</li><li>Dequant$\hat{x} = s\,(q - z)$</li><li>对称量化 $z = 0$dequant 退化为 $\hat{x} = s\cdot q$</li></ul>
<p>把 round 和 clamp 顺序写反,或忘掉 zero-point 的减法。</p>
</details>
<details>
<summary>Q2. 对称 vs 非对称量化的取舍?</summary>
<ul><li>对称:$z = 0$GEMM 实现简单(无 cross-zero 项),但分布偏态时浪费 1 bit</li><li>非对称:精确覆盖任意 $[\alpha, \beta]$GEMM 需 fuse 掉 zero-point 项</li><li>LLM weight 一般近似零均值 → 对称即可;activation 可能偏(如 ReLU 输出非负)→ 非对称更好</li></ul>
<p>说"非对称一定更精确所以一定更好",忘了 GEMM cross 项工程开销。</p>
</details>
<details>
<summary>Q3. Per-tensor / per-channel / per-group 区别?</summary>
<ul><li>per-tensor:整张矩阵一个 scale,开销 $O(1)$,最差精度</li><li>per-channelweight 沿 output dim 每行一个 scaleor activation 沿 hidden dim</li><li>per-group:每 $g$ 个 weight 一个 scale$g = 32 / 64 / 128$,精度最高</li><li>Storage 影响:W4 + group128 ≈ 4.125 bits/weightgroup32 ≈ 4.5 bits/weight</li></ul>
<p>混淆"per-channel along which dim"——weight 是 output channel 安全(GEMM K 维独立),activation 沿 hiddenK 维)则不能直接 fuse 进 GEMM。</p>
</details>
<details>
<summary>Q4. LLM 量化为什么比 CNN 难?</summary>
<ul><li>LLM ≥ 6.7B 后出现 systematic activation outlier0.1%-1% channel 量级 $50\text{-}100\times$</li><li>Outlier 在不同 token / sample 上稳定(不是噪声,是结构)</li><li>per-tensor INT8 在小模型 OK,大模型崩 5-10 PPL 点</li><li>这是 LLM.int8() / SmoothQuant / AWQ 都在攻击的痛点</li></ul>
<p>说"LLM 量化和 CNN 一样"或者"只是规模大"。</p>
</details>
<details>
<summary>Q5. INT8 / INT4 / FP8 区别?</summary>
<ul><li>INT88-bit 整数 $[-128, 127]$,配 scale 表实数</li><li>INT44-bit 整数 $[-8, 7]$,必须配 group quant + 比较精细的 calibration</li><li>FP8 E4M31S/4E/3Mdynamic range $\pm 448$forward 用</li><li>FP8 E5M21S/5E/2Mdynamic range 与 FP16 相当,backward 用</li></ul>
<p>把 FP8 当 INT8 用;忘了 E4M3 没有 Inf 只有 NaN。</p>
</details>
<details>
<summary>Q6. GPTQ 是什么?它和 RTN 比好在哪?</summary>
<ul><li>GPTQ (Frantar 2023) = OBQ 在 LLM 上的高效化版本,基于 OBS 的 layer-wise PTQ</li><li>用 Hessian $H = X^\top X$ 信息,量化第 $q$ 列后<strong>更新剩余列</strong>补偿误差</li><li>W4 LLaMA-7BRTN PPL +1.5GPTQ +0.1</li><li>时间代价:单卡 30-60 分钟/7B</li></ul>
<p>只说"GPTQ 是 4-bit 量化",不提 OBS 误差传播。</p>
</details>
<details>
<summary>Q7. AWQ 的核心思路?</summary>
<ul><li>1% salient weight (input activation 大的 channel) 决定大部分量化损失</li><li>引入 per-input-channel scale $s_c$,做等价变换 $W \to s W$, $X \to X/s$</li><li>grid search $\alpha \in [0, 1]$$s_c = \mathrm{mean}(|x_c|)^\alpha$</li><li>比 GPTQ 快、与 Marlin / W4A16 kernel 配合好</li></ul>
<p>只说"AWQ 比 GPTQ 快",不提 activation-aware scale 等价变换的数学。</p>
</details>
<details>
<summary>Q8. SmoothQuant 解决什么?为什么 W8A8 不能直接量化?</summary>
<ul><li>W8A8 直接量化崩 in part because activation per-tensor scale 被 outlier 顶死</li><li>SmoothQuant$Y = (X/S)(SW)$$S$ 对角矩阵,<strong>等价变换</strong></li><li>$X/S$ 平滑、$SW$ 仍然 per-channel 可量</li><li>$s_c = \max|X_c|^\alpha / \max|W_c|^{1-\alpha}$$\alpha = 0.5$ 默认</li></ul>
<p>只说"SmoothQuant 是 W8A8",不解释 outlier migration 数学。</p>
</details>
<details>
<summary>Q9. KV cache 占多少显存?怎么算?</summary>
<ul><li>公式:$L_\text{ctx} \cdot 2 \cdot n_\text{layers} \cdot H_\text{kv} \cdot d_\text{head} \cdot \text{bytes}$</li><li>LLaMA-2-70B FP16 4K ctx$4096 \times 2 \times 80 \times 8 \times 128 \times 2 \approx 1.34$ GB / sample</li><li>batch 64 → 86 GB,需要 KV4 / KV8 才能装下</li><li>MQA / GQA 让 $H_\text{kv} \ll H$70B 是 GQA G=8,否则 vanilla MHA 是 10 GB / sample</li></ul>
<p>只说"KV cache 大",不会算具体数字。</p>
</details>
<details>
<summary>Q10. Bitsandbytes 的 NF4 和 INT4 区别?</summary>
<ul><li>INT4:均匀量化,16 个等间距 level</li><li>NF4 (Normal Float 4):非均匀,根据 standard normal 的分位数选 16 个 level</li><li>NF4 假设 weight $\sim \mathcal{N}(0, \sigma^2)$ 后归一化到 $[-1, 1]$,比 INT4 期望意义上更优</li><li>QLoRA (Dettmers 2023) 用 NF4 + double quantization</li></ul>
<p>说 NF4 是非整数所以更慢——错。它是 lookup table-based dequant,速度与 INT4 相当。</p>
</details>
<h3 id="l2进阶题research-oriented-岗位">L2进阶题(research-oriented 岗位)</h3>
<details>
<summary>Q11. GPTQ 最优 weight update 公式从 OBS 怎么推?</summary>
<ul><li>二阶 Taylor$L(w^* + \delta) \approx \frac{1}{2}\delta^\top H \delta$$H = X^\top X$</li><li>约束 $e_q^\top \delta = c_q := \mathrm{Quant}(w_q^*) - w_q^*$(注意符号:$c_q$ 是 quant 后减原值)</li><li>Lagrangian + KKT 得 $\delta^* = \lambda H^{-1} e_q$$\lambda = c_q / [H^{-1}]_{qq}$</li><li>所以剩余列更新 $w_j \mathrel{+}= (c_q / [H^{-1}]_{qq}) \cdot [H^{-1}]_{jq}$(等价于 §4.3 中用 $-(w_q-\mathrm{Quant}(w_q))/[H^{-1}]_{qq}\cdot[H^{-1}]_{jq}$ 的写法)</li></ul>
<p>只背公式不会推;或把 $H$ 当 weight 的 Hessian(错,是 input Hessian);或忘了量化第 $q$ 列后还要 propagate 误差到剩余列。</p>
</details>
<details>
<summary>Q12. SmoothQuant migration 为什么不破坏 GEMM?数学上证明。</summary>
<ul><li>$S = \mathrm{diag}(s_1, \ldots, s_K)$ 对角矩阵</li><li>$Y = X W = X S^{-1} \cdot S W = \hat{X} \hat{W}$ — 矩阵乘法关联律 + 对角矩阵可吸收</li><li>等价:$\hat{X}_{:, c} = X_{:, c} / s_c$$\hat{W}_{c, :} = s_c W_{c, :}$channelwise rescale,不改变 K 维内积结构)</li><li>工程:$S^{-1}$ fuse 进上游 LN weight$SW$ 离线 merge 一次,runtime 零开销</li></ul>
<p>只说"对角矩阵可以 fuse",不写出 $X S^{-1} \cdot S W$ 的代数等价。</p>
</details>
<details>
<summary>Q13. FP8 E4M3 vs E5M2 forward / backward 怎么选?为什么?</summary>
<ul><li><strong>Forward (W, A)</strong>E4M3 — 4E/3M, dynamic range $\pm 448$ 够覆盖经 layer scale 后的 weight/activation<strong>mantissa 多 1 bit 精度更高</strong></li><li><strong>Backward (gradient)</strong>E5M2 — 5E/2M, 与 FP16 相同 dynamic range<strong>gradient 量级跨越 $10^{-8}$ 到 $10^4$ 必须大动态范围</strong></li><li>E4M3 无 Inf 只有 NaN(最大 normal = 448),E5M2 有 Inf + NaN</li><li>NVIDIA Transformer Engine 默认采用此分工</li></ul>
<p>倒过来用(FP backward 用 E4M3)会 overflowgradient 经常 &gt; 448)。</p>
</details>
<details>
<summary>Q14. AWQ 与 GPTQ 哪个更快?精度差多少?</summary>
<ul><li>量化耗时:AWQ 更快(一次 grid search $\alpha$5-15 min/7B);GPTQ 慢(Hessian + Cholesky 迭代,30-60 min/7B</li><li>精度:W4 上几乎打平(LLaMA-7B 两者都 &lt; +0.2 PPL</li><li>推理:AWQ 的 scale 可 merge 进 LN weightruntime 零开销;GPTQ act_order=True 时有 reorder 索引开销</li><li>工程:AWQ 与 Marlin W4A16 kernel 配合好,vLLM 默认 W4 路径用 AWQ</li></ul>
<p>说"GPTQ 一定更准"——错,W4 上等价;说"AWQ 不需要 calibration"——错,需要 mean(|x|) per channel。</p>
</details>
<details>
<summary>Q15. INT8 量化后 GEMM 会有 cross zero-point 项,怎么消?</summary>
<ul><li>$\hat{x}_a = s_a (q_a - z_a)$$\hat{x}_b = s_b (q_b - z_b)$</li><li>$\hat{x}_a \hat{x}_b = s_a s_b (q_a q_b - z_a q_b - z_b q_a + z_a z_b)$</li><li>展开后有 4 项;通常预先<strong>让 weight 对称量化 $z_W = 0$</strong> 消两项</li><li>剩下 $- z_a \cdot q_b$ 项可以<strong>用一行 reduce sum 预算</strong>per-batch only),inference 时一次性减掉</li></ul>
<p>只说"对称量化",不解释如何处理 activation 非对称的 cross 项。</p>
</details>
<details>
<summary>Q16. PTQ vs QAT 区别?什么时候用 QAT?</summary>
<ul><li>PTQ:训练后 calibration + closed-form 量化(GPTQ, AWQ, SmoothQuant</li><li>QAT:训练或 finetune 时模拟量化(STE 反向)</li><li>LLM 工业现状:W8 / W4 PTQ 已足够(&lt; 0.2 PPL 损失),不需要 QAT</li><li>W2 / 1.58-bit BitNet 必须 from-scratch QATfinetune 后 W4A4 也常做 QAT</li></ul>
<p>说"QAT 一定更准所以总是用它"——成本 100-1000$\times$ PTQ,对 W8/W4 没必要。</p>
</details>
<details>
<summary>Q17. KV cache 量化 K 和 V 为什么粒度不同?</summary>
<ul><li>K 的 outlier 沿 <strong>head_dim</strong> 维稳定(特定 channel 大),所以 K 用 <strong>per-channel quant</strong></li><li>V 的 outlier 沿 <strong>token</strong> 维变化(每个 token 自己 magnitude),所以 V 用 <strong>per-token quant</strong></li><li>KIVI / KVQuant 都是这个设计</li><li>RoPE 前 quant K 更稳(post-RoPE 把 channel outlier 打散到不同 freq band</li></ul>
<p>说 "K 和 V 一视同仁 per-tensor"——这正是早期 KV cache 量化崩盘的原因。</p>
</details>
<details>
<summary>Q18. NVFP4 和 MXFP4 区别?</summary>
<ul><li>都是 FP4 E2M1 element type1S/2E/1M</li><li><strong>MXFP4</strong> (OCP)block size <strong>32</strong>shared scale <strong>E8M0</strong> (8-bit pure exponent, 即 $2^{e-127}$)</li><li><strong>NVFP4</strong> (NVIDIA Blackwell)block size <strong>16</strong>shared scale <strong>FP8 E4M3</strong>(带 mantissa),<strong>额外一个 per-tensor FP32 scale</strong></li><li>NVFP4 更细粒度、scale 精度更高,但 storage overhead 也大($4 + 8/16 \approx 4.5$ bits</li><li>Blackwell tensor core 原生支持 NVFP4MXFP4 需 AMD MI350 / Intel Gaudi 3</li></ul>
<p>混为一谈或说 "FP4 就是 INT4 加 sign"——错。</p>
</details>
<details>
<summary>Q19. LLM.int8() 的 mixed-precision decomposition 怎么做?</summary>
<ul><li>每层把 activation 拆两路:outlier path (channel-max &gt; 6, 保留 FP16) + normal path (其余, INT8 vector-wise quant)</li><li>数学:$Y = X_O W_O + X_N W_N$,两路独立 GEMM 后相加</li><li>Outlier mask 在每个 forward step detect(不能预先 baked</li><li>第一个能落地的 OPT-175B INT8 推理方案,PPL 几乎无损</li><li>缺点:FP16 outlier path 是 throughput 瓶颈(~10-15% 延迟),后续 SmoothQuant / AWQ 走"消灭 outlier"路线</li></ul>
<p>说 "LLM.int8() 是 pure INT8"——错,是 mixed-precision。</p>
</details>
<details>
<summary>Q20. STE (Straight-Through Estimator) 怎么用?为什么 work</summary>
<ul><li>Round 函数处处导数 0 或不存在,反向无信号</li><li>STE: $\partial \mathrm{Round}(x) / \partial x := 1$(在 clamp 范围内),范围外置 0</li><li>直觉:前向用量化值(discrete),反向当作恒等(pass gradient through</li><li>有偏估计但实践上 workLSQ (Esser 2020) 进一步学习 scalePACT 学习 clamp threshold</li><li>不 work 的情况:量化太激进(W2 from scratch),梯度方向显著偏离真实梯度,需要 BinaryConnect 等专门方法</li></ul>
<p>说 STE 是 unbiased estimator——错,是 biased but useful。</p>
</details>
<h3 id="l3高级变体顶级-lab--系统方向">L3高级变体(顶级 lab / 系统方向)</h3>
<details>
<summary>Q21. QuaRot / SpinQuant 的 Hadamard 旋转为什么能消 outlier</summary>
<ul><li>任意正交矩阵 $R$ 把 vector $x$ 变 $Rx$$\|Rx\|_2 = \|x\|_2$ 不变,但 $\max|x|$ 可以显著减小</li><li>Hadamard $H \in \{+1, -1\}^{d\times d} / \sqrt{d}$:把每个 channel 变成所有 channel 的 $\pm$ 等权平均,<strong>集中 outlier 被打散到所有维度</strong></li><li>数学:若 $x$ 有 $k \ll d$ 个 outlier$Hx$ 的 $\ell_\infty$ norm 约 $\sqrt{k/d} \cdot \max|x|$incoherence 性质)</li><li>QuaRot 在 RMSNorm 处穿过(用 online Hadamard 绕开 $\gamma$ 不通过 $H$ 的问题),SpinQuant 学习 $R$ 而非随机</li><li>W4A4 LLaMA-2-70BQuaRot PPL +0.5,比 SmoothQuant 的 +5 显著好</li></ul>
<p>说"旋转就是降维 PCA"——错,正交变换保 norm 不降维。</p>
</details>
<details>
<summary>Q22. QServe / QoQ (W4A8KV4) 为什么需要自定义 kernel</summary>
<ul><li>W4 weight + A8 activation 的 GEMM 在 stock cuBLAS / cuDNN 上没有直接路径</li><li>QServe 的 kernel:每个 W4 weight 在 register 里 dequant 到 INT8lookup table),然后做 INT8×INT8 matmul</li><li>关键优化:dequant + Tensor Core MMA fused 在同一个 warp instruction 内,<strong>避免 FP16 中间 buffer</strong></li><li>KV4 attentionK, V 都 INT4,与 INT8 query 做 mixed-precision dot,需要 attention kernel 内的 dequant 路径</li><li>端到端:LLaMA-3-70B H100 解码 1000+ tokens/s/GPU,比 FP16 TensorRT-LLM 快 1.2-3.5$\times$</li></ul>
<p>只说"W4A8 比 FP16 快",不解释为什么需要 kernel-level co-designstock GEMM 不支持 W4 input)。</p>
</details>
<details>
<summary>Q23. FP8 training 的 amax history / delayed scaling 是什么?</summary>
<ul><li>每个 GEMM 的输入 / 输出维护一个 amax history(最近 N 个 step 的 max absN = 16 典型)</li><li>Scale 用 history max 算($s = \max\text{history} / 448$ for E4M3),保证下一 window 不 overflow</li><li>"Delayed":用<strong>上一</strong>窗口的 amax 算<strong>当前</strong>窗口的 scale,避免阻塞 forward 等 amax 算出来</li><li>Cast 在 GEMM 入口:FP32 → FP8 用 scaleGEMM 输出累积 FP32 再 cast 出去</li><li>LLaMA-3 / DeepSeek-V3 FP8 train 比 BF16 提升 1.5-2$\times$ throughput</li></ul>
<p>把 delayed scaling 当成 loss scaling 同一个东西——loss scaling 是 backward 路径上抗 underflowdelayed scaling 是 per-GEMM forward/backward 的 amax 用法。</p>
</details>
<details>
<summary>Q24. NVFP4 的 per-tensor + per-block FP8 scale 双层结构为什么必要?</summary>
<ul><li>FP4 E2M1 max normal = 6,动态范围极窄</li><li>单层 per-block FP8 scale (E4M3, max 448):单 block 内可表 $\pm 6 \times 448 \approx \pm 2700$;但跨 block 仍受 FP8 scale 自己范围限制</li><li>LLM activation amax 经常 &gt; 2700 (outlier channel 出现 $10^4$ 级)per-block FP8 scale 不够</li><li><strong>额外 per-tensor FP32 scale</strong>:把整个 tensor 拉到 FP8 scale 的合理动态范围内,相当于先全局粗调再 block-wise 细调</li><li>类比:FP32 用 mantissa+exp 表大 rangeNVFP4 用 FP4 mantissa + FP8 block exp + FP32 tensor exp,三层 hierarchy</li></ul>
<p>只说 "NVFP4 = FP4 + scale"——错,是 FP4 + per-block FP8 + per-tensor FP32 三层。</p>
</details>
<details>
<summary>Q25. 设计一个 W4 量化方案给一个未知 LLM,你会怎么做?</summary>
<ul><li><strong>Step 1</strong>:跑 layer-wise activation profiling,看是否 ≥ 6.7B 有 systematic outlier;如果有,weight-only 必须配 AWQ scale 或 GPTQ Hessian</li><li><p><strong>Step 2</strong>:选 PTQ 方法:</p>
<ul><li>简单部署:AWQ (W4) + Marlin kernel,最佳精度-工程 trade-off</li><li>极致精度:GPTQ (W4) + act_order,多 0.5-1% throughput 但 &lt; 0.1 PPL</li><li>W4A8 / W4A4:必须额外 SmoothQuant / QuaRot</li></ul></li><li><p><strong>Step 3</strong>calibration data 选择:</p>
<ul><li>通用 LLM128 samples × 2048 tokens from C4 / WikiText</li><li>任务专用:用 in-domain dataPPL 差异显著(数学 task on math data 比 web data 好 2-3 PPL</li></ul></li><li><p><strong>Step 4</strong>:选 group_size</p>
<ul><li>W4 group128 默认(4.125 bits/weight, 良好精度)</li><li>W3 / W2 必须 group32 或更小</li></ul></li><li><strong>Step 5</strong>:验证:跑 PPL + 下游 task (MMLU / GSM8K)PPL diff &lt; 0.2 + task drop &lt; 1% 算 PASS</li></ul>
<p>只说"用 GPTQ 4-bit"——没解释如何选 calibration / group_size / kernel backend。</p>
</details>
<h2 id="a-附录完整的工程参考">§A 附录:完整的工程参考</h2>
<h3 id="a1-主要论文-reference-list">A.1 主要论文 reference list</h3>
<table><thead><tr><th>方法</th><th>论文</th><th>关键贡献</th></tr></thead><tbody><tr><td>LLM.int8()</td><td>Dettmers et al., NeurIPS 2022</td><td>Mixed-precision INT8 decomposition for outlier</td></tr><tr><td>GPTQ</td><td>Frantar et al., ICLR 2023</td><td>OBS-based W4 PTQ for LLM</td></tr><tr><td>SmoothQuant</td><td>Xiao et al., ICML 2023</td><td>Migrate activation outlier to weight (W8A8)</td></tr><tr><td>AWQ</td><td>Lin et al., MLSys 2024</td><td>Activation-aware per-channel scale (W4)</td></tr><tr><td>QuIP</td><td>Chee et al., NeurIPS 2023</td><td>Random Hadamard for weight incoherence</td></tr><tr><td>QuaRot</td><td>Ashkboos et al., NeurIPS 2024</td><td>Full Hadamard rotation, W4A4KV4</td></tr><tr><td>SpinQuant</td><td>Liu et al., ICLR 2025 (Meta)</td><td>Learned rotations $R_1$-$R_4$</td></tr><tr><td>OmniQuant</td><td>Shao et al., ICLR 2024</td><td>Learnable equivalent transforms</td></tr><tr><td>KIVI</td><td>Liu et al., ICML 2024</td><td>per-channel K + per-token V, INT2 KV</td></tr><tr><td>KVQuant</td><td>Hooper et al., NeurIPS 2024</td><td>Pre-RoPE quant K, non-uniform V</td></tr><tr><td>QServe / QoQ</td><td>Lin et al., MLSys 2025</td><td>W4A8KV4 GPU kernel co-design</td></tr><tr><td>LLM-QAT</td><td>Liu et al., 2023</td><td>Self-distillation QAT for W4</td></tr><tr><td>BitNet b1.58</td><td>Ma et al., 2024</td><td>Ternary weight LLM from scratch</td></tr><tr><td>FP8 Training</td><td>Micikevicius et al., 2022</td><td>E4M3 forward / E5M2 backward</td></tr><tr><td>MX formats</td><td>OCP / Microsoft, 2024</td><td>Block-scaled FP4/6/8 with E8M0</td></tr><tr><td>NVFP4</td><td>NVIDIA Blackwell, 2025</td><td>FP4 E2M1 + FP8 E4M3 block + FP32 tensor scale</td></tr></tbody></table>
<h3 id="a2-一图速查选什么量化方案">A.2 一图速查:选什么量化方案</h3>
<pre class="diagram"><code>
┌─────────────────────────┐
│ 是否能接受 PPL +0.5 ? │
└────────┬────────────────┘
┌─────┴─────┐
│ 能 │ 不能
↓ ↓
[W4 weight-only] [INT8 / FP8 W+A 量化]
GPTQ / AWQ SmoothQuant W8A8
+ Marlin kernel + per-channel W
群中典型 PPL: +0.1 PPL: +0.05
┌─────────────────────────┐
│ 是否做 INT4 activation?│
└────────┬────────────────┘
┌─────┴─────┐
│ 是 │ 否
↓ ↓
[W4A4] [W4A8KV4 / QoQ]
QuaRot / SpinQuant AWQ + KV4
+ Hadamard rotation 用 QServe kernel
PPL +0.5 (70B) PPL +0.2</code></pre>
<h3 id="a3-量化-quick-reference-卡片">A.3 量化 quick reference 卡片</h3>
<table><thead><tr><th>任务</th><th>推荐方案</th><th>框架</th></tr></thead><tbody><tr><td>单卡推理 LLaMA-2-70B</td><td>AWQ W4 + Marlin (vLLM / SGLang)</td><td>AutoAWQ</td></tr><tr><td>多卡推理 LLaMA-3-405B</td><td>SmoothQuant + GPTQ W4A8 + KV8</td><td>TensorRT-LLM</td></tr><tr><td>端侧 (Apple Silicon / CPU)</td><td>GGUF Q4_K_M / Q5_K_S</td><td>llama.cpp</td></tr><tr><td>训练 fp8 LLM</td><td>TE FP8 (E4M3/E5M2) + amax history</td><td>Transformer Engine</td></tr><tr><td>QLoRA finetune</td><td>NF4 + double quant + LoRA</td><td>bitsandbytes + peft</td></tr><tr><td>极致 throughput H100/B200 serving</td><td>QoQ W4A8KV4 / NVFP4</td><td>QServe / TensorRT-LLM</td></tr><tr><td>边缘 &lt; 100 MB model</td><td>BitNet b1.58 (1.58-bit) from scratch</td><td>自定义 / bitnet.cpp</td></tr></tbody></table>
<h3 id="a4-sanity-check-checklist">A.4 Sanity check checklist</h3>
<p>实战部署任何量化模型前,必跑:</p>
<ul><li>[ ] <strong>PPL on calibration domain</strong>:量化前后差 &lt; 0.2 算 OK</li><li>[ ] <strong>PPL on out-of-domain</strong>(重要!):差 &lt; 0.5 算 OK&gt; 1 重新选 calibration data</li><li>[ ] <strong>MMLU / GSM8K 等下游 task</strong>drop &lt; 1% 算 OK</li><li>[ ] <strong>Long context PPL</strong>(最考验 KV cache quant):4K / 8K / 32K 三档对比</li><li>[ ] <strong>生成质量主观评估</strong>:让人盲评 50 个 prompt 的输出,比 FP16 不显著差</li><li>[ ] <strong>Throughput / TTFT (Time To First Token)</strong>:实测 vs 理论数字差 &gt; 30% 说明 kernel 未充分优化</li><li>[ ] <strong>峰值显存</strong>:实测 vs nominal $\text{bits} \cdot \text{params} / 8 + \text{KV cache} + \text{activations}$</li></ul>
<p><strong>Quantization Quick Reference</strong> · 主要参考:Dettmers 2022 (LLM.int8()), Frantar 2023 (GPTQ), Xiao 2023 (SmoothQuant), Lin 2024 (AWQ), Ashkboos 2024 (QuaRot), Lin 2025 (QServe). 最后更新:2026-05。</p>
<footer class="aris-footer">
Generated by <a href="https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/blob/main/skills/render-html/SKILL.md">ARIS <code>/render-html</code></a> ·
source path <code>docs/tutorials/quantization_tutorial.md</code> ·
SHA256 <code>f04948e0f203</code> ·
generated at 2026-05-19 09:46 UTC.
This is a generated view — edit the source Markdown, then re-render.
</footer>
</main>
</div>
</body>
</html>