docs: make Chinese README the default
This commit is contained in:
@@ -1,458 +1,464 @@
|
||||
<!-- WEHUB_ZH_README -->
|
||||
> [!NOTE]
|
||||
> 本文档由 WeHub 基于上游 README 翻译整理,属于社区翻译,非官方中文文档。
|
||||
> [English](./README.en.md) · [原始项目](https://github.com/mlabonne/llm-course) · [上游 README](https://github.com/mlabonne/llm-course/blob/HEAD/README.md)
|
||||
> 原作者、版权与许可证归属以原始项目及本仓库 LICENSE 文件为准。
|
||||
|
||||
<div align="center">
|
||||
<img src="img/banner.png" alt="LLM Course">
|
||||
<p align="center">
|
||||
𝕏 <a href="https://twitter.com/maximelabonne">Follow me on X</a> •
|
||||
𝕏 <a href="https://twitter.com/maximelabonne">在 X 上关注我</a> •
|
||||
🤗 <a href="https://huggingface.co/mlabonne">Hugging Face</a> •
|
||||
💻 <a href="https://mlabonne.github.io/blog">Blog</a> •
|
||||
💻 <a href="https://mlabonne.github.io/blog">博客</a> •
|
||||
📙 <a href="https://packt.link/a/9781836200079">LLM Engineer's Handbook</a>
|
||||
</p>
|
||||
</div>
|
||||
<br/>
|
||||
|
||||
<a href="https://a.co/d/a2M67rE"><img align="right" width="25%" src="https://i.imgur.com/7iNjEq2.png" alt="LLM Engineer's Handbook Cover"/></a>The LLM course is divided into three parts:
|
||||
<a href="https://a.co/d/a2M67rE"><img align="right" width="25%" src="https://i.imgur.com/7iNjEq2.png" alt="LLM Engineer's Handbook Cover"/></a>LLM 课程分为三个部分:
|
||||
|
||||
1. 🧩 **LLM Fundamentals** is optional and covers fundamental knowledge about mathematics, Python, and neural networks.
|
||||
2. 🧑🔬 **The LLM Scientist** focuses on building the best possible LLMs using the latest techniques.
|
||||
3. 👷 **The LLM Engineer** focuses on creating LLM-based applications and deploying them.
|
||||
1. 🧩 **LLM Fundamentals** 为可选部分,涵盖数学、Python 和神经网络等基础知识。
|
||||
2. 🧑🔬 **The LLM Scientist** 专注于运用最新技术构建尽可能优秀的 LLM。
|
||||
3. 👷 **The LLM Engineer** 专注于创建基于 LLM 的应用并将其部署上线。
|
||||
|
||||
> [!NOTE]
|
||||
> Based on this course, I co-wrote the [LLM Engineer's Handbook](https://packt.link/a/9781836200079), a hands-on book that covers an end-to-end LLM application from design to deployment. The LLM course will always stay free, but you can support my work by purchasing this book.
|
||||
> 基于本课程,我与人合著了 [LLM Engineer's Handbook](https://packt.link/a/9781836200079), a hands-on book that covers an end-to-end LLM application from design to deployment. LLM 课程将始终保持免费,你也可以通过购买这本书来支持我的工作。
|
||||
|
||||
For a more comprehensive version of this course, check out the [DeepWiki](https://deepwiki.com/mlabonne/llm-course/).
|
||||
如需更全面的课程版本,请查看 [DeepWiki](https://deepwiki.com/mlabonne/llm-course/).
|
||||
|
||||
## 📝 Notebooks
|
||||
|
||||
A list of notebooks and articles I wrote about LLMs.
|
||||
我撰写的关于 LLM 的 Notebook 与文章列表。
|
||||
|
||||
<details>
|
||||
<summary>Toggle section (optional)</summary>
|
||||
<summary>展开/收起本节(可选)</summary>
|
||||
|
||||
### Tools
|
||||
|
||||
| Notebook | Description | Notebook |
|
||||
|----------|-------------|----------|
|
||||
| 🧐 [LLM AutoEval](https://github.com/mlabonne/llm-autoeval) | Automatically evaluate your LLMs using RunPod | <a href="https://colab.research.google.com/drive/1Igs3WZuXAIv9X0vwqiE90QlEPys8e8Oa?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🥱 LazyMergekit | Easily merge models using MergeKit in one click. | <a href="https://colab.research.google.com/drive/1obulZ1ROXHjYLn6PPZJwRR6GzgQogxxb?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🦎 LazyAxolotl | Fine-tune models in the cloud using Axolotl in one click. | <a href="https://colab.research.google.com/drive/1TsDKNo2riwVmU55gjuBgB1AXVtRRfRHW?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| ⚡ AutoQuant | Quantize LLMs in GGUF, GPTQ, EXL2, AWQ, and HQQ formats in one click. | <a href="https://colab.research.google.com/drive/1b6nqC7UZVt8bx4MksX7s656GXPM-eWw4?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🌳 Model Family Tree | Visualize the family tree of merged models. | <a href="https://colab.research.google.com/drive/1s2eQlolcI1VGgDhqWIANfkfKvcKrMyNr?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🚀 ZeroSpace | Automatically create a Gradio chat interface using a free ZeroGPU. | <a href="https://colab.research.google.com/drive/1LcVUW5wsJTO2NGmozjji5CkC--646LgC"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| ✂️ AutoAbliteration | Automatically abliteration models with custom datasets. | <a href="https://colab.research.google.com/drive/1RmLv-pCMBBsQGXQIM8yF-OdCNyoylUR1?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🧼 AutoDedup | Automatically deduplicate datasets using the Rensa library. | <a href="https://colab.research.google.com/drive/1o1nzwXWAa8kdkEJljbJFW1VuI-3VZLUn?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🧐 [LLM AutoEval](https://github.com/mlabonne/llm-autoeval) | 使用 RunPod 自动评估你的 LLM | <a href="https://colab.research.google.com/drive/1Igs3WZuXAIv9X0vwqiE90QlEPys8e8Oa?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🥱 LazyMergekit | 一键使用 MergeKit 轻松合并模型。 | <a href="https://colab.research.google.com/drive/1obulZ1ROXHjYLn6PPZJwRR6GzgQogxxb?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🦎 LazyAxolotl | 一键使用 Axolotl 在云端微调模型。 | <a href="https://colab.research.google.com/drive/1TsDKNo2riwVmU55gjuBgB1AXVtRRfRHW?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| ⚡ AutoQuant | 一键将 LLM 量化为 GGUF、GPTQ、EXL2、AWQ 和 HQQ 格式。 | <a href="https://colab.research.google.com/drive/1b6nqC7UZVt8bx4MksX7s656GXPM-eWw4?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🌳 Model Family Tree | 可视化合并模型的族谱树。 | <a href="https://colab.research.google.com/drive/1s2eQlolcI1VGgDhqWIANfkfKvcKrMyNr?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🚀 ZeroSpace | 使用免费的 ZeroGPU 自动创建 Gradio 聊天界面。 | <a href="https://colab.research.google.com/drive/1LcVUW5wsJTO2NGmozjji5CkC--646LgC"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| ✂️ AutoAbliteration | 使用自定义数据集自动进行 abliteration 处理模型。 | <a href="https://colab.research.google.com/drive/1RmLv-pCMBBsQGXQIM8yF-OdCNyoylUR1?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 🧼 AutoDedup | 使用 Rensa 库自动对数据集去重。 | <a href="https://colab.research.google.com/drive/1o1nzwXWAa8kdkEJljbJFW1VuI-3VZLUn?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
|
||||
### Fine-tuning
|
||||
|
||||
| Notebook | Description | Article | Notebook |
|
||||
|---------------------------------------|-------------------------------------------------------------------------|---------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| Fine-tune Llama 3.1 with Unsloth | Ultra-efficient supervised fine-tuning in Google Colab. | [Article](https://mlabonne.github.io/blog/posts/2024-07-29_Finetune_Llama31.html) | <a href="https://colab.research.google.com/drive/164cg_O7SV7G8kZr_JXqLd6VC7pd86-1Z?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune Llama 3 with ORPO | Cheaper and faster fine-tuning in a single stage with ORPO. | [Article](https://mlabonne.github.io/blog/posts/2024-04-19_Fine_tune_Llama_3_with_ORPO.html) | <a href="https://colab.research.google.com/drive/1eHNWg9gnaXErdAa8_mcvjMupbSS6rDvi"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune Mistral-7b with DPO | Boost the performance of supervised fine-tuned models with DPO. | [Article](https://mlabonne.github.io/blog/posts/Fine_tune_Mistral_7b_with_DPO.html) | <a href="https://colab.research.google.com/drive/15iFBr1xWgztXvhrj5I9fBv20c7CFOPBE?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune Mistral-7b with QLoRA | Supervised fine-tune Mistral-7b in a free-tier Google Colab with TRL. | | <a href="https://colab.research.google.com/drive/1o_w0KastmEJNVwT5GoqMCciH-18ca5WS?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune CodeLlama using Axolotl | End-to-end guide to the state-of-the-art tool for fine-tuning. | [Article](https://mlabonne.github.io/blog/posts/A_Beginners_Guide_to_LLM_Finetuning.html) | <a href="https://colab.research.google.com/drive/1Xu0BrCB7IShwSWKVcfAfhehwjDrDMH5m?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune Llama 2 with QLoRA | Step-by-step guide to supervised fine-tune Llama 2 in Google Colab. | [Article](https://mlabonne.github.io/blog/posts/Fine_Tune_Your_Own_Llama_2_Model_in_a_Colab_Notebook.html) | <a href="https://colab.research.google.com/drive/1PEQyJO1-f6j0S_XJ8DV50NkpzasXkrzd?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune Llama 3.1 with Unsloth | 在 Google Colab 中进行超高效率的监督微调(Supervised Fine-tuning)。 | [Article](https://mlabonne.github.io/blog/posts/2024-07-29_Finetune_Llama31.html) | <a href="https://colab.research.google.com/drive/164cg_O7SV7G8kZr_JXqLd6VC7pd86-1Z?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune Llama 3 with ORPO | 使用 ORPO 以单阶段方式实现更便宜、更快速的微调。 | [Article](https://mlabonne.github.io/blog/posts/2024-04-19_Fine_tune_Llama_3_with_ORPO.html) | <a href="https://colab.research.google.com/drive/1eHNWg9gnaXErdAa8_mcvjMupbSS6rDvi"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune Mistral-7b with DPO | 使用 DPO 提升监督微调模型的性能。 | [Article](https://mlabonne.github.io/blog/posts/Fine_tune_Mistral_7b_with_DPO.html) | <a href="https://colab.research.google.com/drive/15iFBr1xWgztXvhrj5I9fBv20c7CFOPBE?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune Mistral-7b with QLoRA | 在免费版 Google Colab 中使用 TRL 对 Mistral-7b 进行监督微调。 | | <a href="https://colab.research.google.com/drive/1o_w0KastmEJNVwT5GoqMCciH-18ca5WS?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune CodeLlama using Axolotl | 面向最先进微调工具的全流程指南。 | [Article](https://mlabonne.github.io/blog/posts/A_Beginners_Guide_to_LLM_Finetuning.html) | <a href="https://colab.research.google.com/drive/1Xu0BrCB7IShwSWKVcfAfhehwjDrDMH5m?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Fine-tune Llama 2 with QLoRA | 在 Google Colab 中监督微调 Llama 2 的分步指南。 | [Article](https://mlabonne.github.io/blog/posts/Fine_Tune_Your_Own_Llama_2_Model_in_a_Colab_Notebook.html) | <a href="https://colab.research.google.com/drive/1PEQyJO1-f6j0S_XJ8DV50NkpzasXkrzd?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
|
||||
### Quantization
|
||||
|
||||
| Notebook | Description | Article | Notebook |
|
||||
|---------------------------------------|-------------------------------------------------------------------------|---------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| Introduction to Quantization | Large language model optimization using 8-bit quantization. | [Article](https://mlabonne.github.io/blog/posts/Introduction_to_Weight_Quantization.html) | <a href="https://colab.research.google.com/drive/1DPr4mUQ92Cc-xf4GgAaB6dFcFnWIvqYi?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 4-bit Quantization using GPTQ | Quantize your own open-source LLMs to run them on consumer hardware. | [Article](https://mlabonne.github.io/blog/4bit_quantization/) | <a href="https://colab.research.google.com/drive/1lSvVDaRgqQp_mWK_jC9gydz6_-y6Aq4A?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Quantization with GGUF and llama.cpp | Quantize Llama 2 models with llama.cpp and upload GGUF versions to the HF Hub. | [Article](https://mlabonne.github.io/blog/posts/Quantize_Llama_2_models_using_ggml.html) | <a href="https://colab.research.google.com/drive/1pL8k7m04mgE5jo2NrjGi8atB0j_37aDD?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| ExLlamaV2: The Fastest Library to Run LLMs | Quantize and run EXL2 models and upload them to the HF Hub. | [Article](https://mlabonne.github.io/blog/posts/ExLlamaV2_The_Fastest_Library_to_Run%C2%A0LLMs.html) | <a href="https://colab.research.google.com/drive/1yrq4XBlxiA0fALtMoT2dwiACVc77PHou?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Introduction to Quantization | 使用 8 位量化(Quantization)优化大语言模型。 | [Article](https://mlabonne.github.io/blog/posts/Introduction_to_Weight_Quantization.html) | <a href="https://colab.research.google.com/drive/1DPr4mUQ92Cc-xf4GgAaB6dFcFnWIvqYi?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| 4-bit Quantization using GPTQ | 量化你自己的开源 LLM,以便在消费级硬件上运行。 | [Article](https://mlabonne.github.io/blog/4bit_quantization/) | <a href="https://colab.research.google.com/drive/1lSvVDaRgqQp_mWK_jC9gydz6_-y6Aq4A?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Quantization with GGUF and llama.cpp | 使用 llama.cpp 量化 Llama 2 模型,并将 GGUF 版本上传至 HF Hub。 | [Article](https://mlabonne.github.io/blog/posts/Quantize_Llama_2_models_using_ggml.html) | <a href="https://colab.research.google.com/drive/1pL8k7m04mgE5jo2NrjGi8atB0j_37aDD?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| ExLlamaV2: The Fastest Library to Run LLMs | 量化并运行 EXL2 模型,并将其上传至 HF Hub。 | [Article](https://mlabonne.github.io/blog/posts/ExLlamaV2_The_Fastest_Library_to_Run%C2%A0LLMs.html) | <a href="https://colab.research.google.com/drive/1yrq4XBlxiA0fALtMoT2dwiACVc77PHou?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
|
||||
### Other
|
||||
|
||||
| Notebook | Description | Article | Notebook |
|
||||
|---------------------------------------|-------------------------------------------------------------------------|---------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| Merge LLMs with MergeKit | Create your own models easily, no GPU required! | [Article](https://mlabonne.github.io/blog/posts/2024-01-08_Merge_LLMs_with_mergekit%20copy.html) | <a href="https://colab.research.google.com/drive/1_JS7JKJAQozD48-LhYdegcuuZ2ddgXfr?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Create MoEs with MergeKit | Combine multiple experts into a single frankenMoE | [Article](https://mlabonne.github.io/blog/posts/2024-03-28_Create_Mixture_of_Experts_with_MergeKit.html) | <a href="https://colab.research.google.com/drive/1obulZ1ROXHjYLn6PPZJwRR6GzgQogxxb?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Uncensor any LLM with abliteration | Fine-tuning without retraining | [Article](https://mlabonne.github.io/blog/posts/2024-06-04_Uncensor_any_LLM_with_abliteration.html) | <a href="https://colab.research.google.com/drive/1VYm3hOcvCpbGiqKZb141gJwjdmmCcVpR?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Improve ChatGPT with Knowledge Graphs | Augment ChatGPT's answers with knowledge graphs. | [Article](https://mlabonne.github.io/blog/posts/Article_Improve_ChatGPT_with_Knowledge_Graphs.html) | <a href="https://colab.research.google.com/drive/1mwhOSw9Y9bgEaIFKT4CLi0n18pXRM4cj?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Decoding Strategies in Large Language Models | A guide to text generation from beam search to nucleus sampling | [Article](https://mlabonne.github.io/blog/posts/2022-06-07-Decoding_strategies.html) | <a href="https://colab.research.google.com/drive/19CJlOS5lI29g-B3dziNn93Enez1yiHk2?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Merge LLMs with MergeKit | 轻松创建你自己的模型,无需 GPU! | [Article](https://mlabonne.github.io/blog/posts/2024-01-08_Merge_LLMs_with_mergekit%20copy.html) | <a href="https://colab.research.google.com/drive/1_JS7JKJAQozD48-LhYdegcuuZ2ddgXfr?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Create MoEs with MergeKit | 将多个专家合并为单个 frankenMoE | [Article](https://mlabonne.github.io/blog/posts/2024-03-28_Create_Mixture_of_Experts_with_MergeKit.html) | <a href="https://colab.research.google.com/drive/1obulZ1ROXHjYLn6PPZJwRR6GzgQogxxb?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Uncensor any LLM with abliteration | 无需重新训练即可微调 | [Article](https://mlabonne.github.io/blog/posts/2024-06-04_Uncensor_any_LLM_with_abliteration.html) | <a href="https://colab.research.google.com/drive/1VYm3hOcvCpbGiqKZb141gJwjdmmCcVpR?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Improve ChatGPT with Knowledge Graphs | 使用知识图谱增强 ChatGPT 的回答。 | [Article](https://mlabonne.github.io/blog/posts/Article_Improve_ChatGPT_with_Knowledge_Graphs.html) | <a href="https://colab.research.google.com/drive/1mwhOSw9Y9bgEaIFKT4CLi0n18pXRM4cj?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
| Decoding Strategies in Large Language Models | 从束搜索(beam search)到核采样(nucleus sampling)的文本生成指南 | [Article](https://mlabonne.github.io/blog/posts/2022-06-07-Decoding_strategies.html) | <a href="https://colab.research.google.com/drive/19CJlOS5lI29g-B3dziNn93Enez1yiHk2?usp=sharing"><img src="img/colab.svg" alt="Open In Colab"></a> |
|
||||
</details>
|
||||
|
||||
## 🧩 LLM Fundamentals
|
||||
## 🧩 LLM 基础
|
||||
|
||||
This section introduces essential knowledge about mathematics, Python, and neural networks. You might not want to start here but refer to it as needed.
|
||||
本节介绍数学、Python 与神经网络方面的必备知识。你可能不必从这里开始学习,但可以在需要时查阅。
|
||||
|
||||
<details>
|
||||
<summary>Toggle section (optional)</summary>
|
||||
<summary>展开/收起本节(可选)</summary>
|
||||
|
||||

|
||||
|
||||
### 1. Mathematics for Machine Learning
|
||||
### 1. 机器学习数学基础
|
||||
|
||||
Before mastering machine learning, it is important to understand the fundamental mathematical concepts that power these algorithms.
|
||||
在掌握机器学习之前,理解支撑这些算法的基本数学概念非常重要。
|
||||
|
||||
- **Linear Algebra**: This is crucial for understanding many algorithms, especially those used in deep learning. Key concepts include vectors, matrices, determinants, eigenvalues and eigenvectors, vector spaces, and linear transformations.
|
||||
- **Calculus**: Many machine learning algorithms involve the optimization of continuous functions, which requires an understanding of derivatives, integrals, limits, and series. Multivariable calculus and the concept of gradients are also important.
|
||||
- **Probability and Statistics**: These are crucial for understanding how models learn from data and make predictions. Key concepts include probability theory, random variables, probability distributions, expectations, variance, covariance, correlation, hypothesis testing, confidence intervals, maximum likelihood estimation, and Bayesian inference.
|
||||
- **线性代数(Linear Algebra)**:这对于理解许多算法至关重要,尤其是深度学习中的算法。核心概念包括向量、矩阵、行列式、特征值与特征向量、向量空间以及线性变换。
|
||||
- **微积分(Calculus)**:许多机器学习算法涉及连续函数的优化,这需要理解导数、积分、极限与级数。多元微积分以及梯度(gradient)概念同样重要。
|
||||
- **概率与统计(Probability and Statistics)**:这对于理解模型如何从数据中学习并做出预测至关重要。核心概念包括概率论、随机变量、概率分布、期望、方差、协方差、相关性、假设检验、置信区间、最大似然估计以及贝叶斯推断(Bayesian inference)。
|
||||
|
||||
📚 Resources:
|
||||
📚 学习资源:
|
||||
|
||||
- [3Blue1Brown - The Essence of Linear Algebra](https://www.youtube.com/watch?v=fNk_zzaMoSs&list=PLZHQObOWTQDPD3MizzM2xVFitgF8hE_ab): Series of videos that give a geometric intuition to these concepts.
|
||||
- [StatQuest with Josh Starmer - Statistics Fundamentals](https://www.youtube.com/watch?v=qBigTkBLU6g&list=PLblh5JKOoLUK0FLuzwntyYI10UQFUhsY9): Offers simple and clear explanations for many statistical concepts.
|
||||
- [Seeing Theory](https://seeing-theory.brown.edu/): A visual introduction to probability and statistics from Brown University.
|
||||
- [Immersive Linear Algebra](https://immersivemath.com/ila/learnmore.html): Another visual interpretation of linear algebra.
|
||||
- [Khan Academy - Linear Algebra](https://www.khanacademy.org/math/linear-algebra): Great for beginners as it explains the concepts in a very intuitive way.
|
||||
- [Khan Academy - Calculus](https://www.khanacademy.org/math/calculus-1): An interactive course that covers all the basics of calculus.
|
||||
- [Khan Academy - Probability and Statistics](https://www.khanacademy.org/math/statistics-probability): Delivers the material in an easy-to-understand format.
|
||||
- [3Blue1Brown - The Essence of Linear Algebra](https://www.youtube.com/watch?v=fNk_zzaMoSs&list=PLZHQObOWTQDPD3MizzM2xVFitgF8hE_ab):) 一系列视频,为这些概念提供几何直觉。
|
||||
- [StatQuest with Josh Starmer - Statistics Fundamentals](https://www.youtube.com/watch?v=qBigTkBLU6g&list=PLblh5JKOoLUK0FLuzwntyYI10UQFUhsY9):) 对许多统计概念提供简单清晰的讲解。
|
||||
- [Seeing Theory](https://seeing-theory.brown.edu/):) 布朗大学出品的概率与统计可视化入门。
|
||||
- [Immersive Linear Algebra](https://immersivemath.com/ila/learnmore.html):) 线性代数的另一种可视化解读。
|
||||
- [Khan Academy - Linear Algebra](https://www.khanacademy.org/math/linear-algebra):) 非常适合初学者,以非常直观的方式讲解概念。
|
||||
- [Khan Academy - Calculus](https://www.khanacademy.org/math/calculus-1):) 涵盖微积分所有基础内容的互动课程。
|
||||
- [Khan Academy - Probability and Statistics](https://www.khanacademy.org/math/statistics-probability):) 以易于理解的格式呈现教学内容。
|
||||
|
||||
---
|
||||
|
||||
### 2. Python for Machine Learning
|
||||
### 2. 机器学习 Python 编程
|
||||
|
||||
Python is a powerful and flexible programming language that's particularly good for machine learning, thanks to its readability, consistency, and robust ecosystem of data science libraries.
|
||||
Python 是一门强大而灵活的编程语言,凭借可读性、一致性与健壮的数据科学库生态,特别适合机器学习。
|
||||
|
||||
- **Python Basics**: Python programming requires a good understanding of the basic syntax, data types, error handling, and object-oriented programming.
|
||||
- **Data Science Libraries**: It includes familiarity with NumPy for numerical operations, Pandas for data manipulation and analysis, Matplotlib and Seaborn for data visualization.
|
||||
- **Data Preprocessing**: This involves feature scaling and normalization, handling missing data, outlier detection, categorical data encoding, and splitting data into training, validation, and test sets.
|
||||
- **Machine Learning Libraries**: Proficiency with Scikit-learn, a library providing a wide selection of supervised and unsupervised learning algorithms, is vital. Understanding how to implement algorithms like linear regression, logistic regression, decision trees, random forests, k-nearest neighbors (K-NN), and K-means clustering is important. Dimensionality reduction techniques like PCA and t-SNE are also helpful for visualizing high-dimensional data.
|
||||
- **Python 基础**:Python 编程需要扎实掌握基本语法、数据类型、错误处理以及面向对象编程。
|
||||
- **数据科学库**:需要熟悉用于数值运算的 NumPy、用于数据处理与分析的 Pandas,以及用于数据可视化的 Matplotlib 和 Seaborn。
|
||||
- **数据预处理**:包括特征缩放与归一化、处理缺失数据、异常值检测、类别数据编码,以及将数据划分为训练集、验证集与测试集。
|
||||
- **机器学习库**:熟练掌握 Scikit-learn 至关重要——该库提供大量监督学习与无监督学习算法。理解如何实现线性回归、逻辑回归、决策树、随机森林、k 近邻(k-nearest neighbors,K-NN)以及 K 均值聚类(K-means clustering)等算法也很重要。PCA 与 t-SNE 等降维技术也有助于可视化高维数据。
|
||||
|
||||
📚 Resources:
|
||||
📚 学习资源:
|
||||
|
||||
- [Real Python](https://realpython.com/): A comprehensive resource with articles and tutorials for both beginner and advanced Python concepts.
|
||||
- [freeCodeCamp - Learn Python](https://www.youtube.com/watch?v=rfscVS0vtbw): Long video that provides a full introduction into all of the core concepts in Python.
|
||||
- [Python Data Science Handbook](https://jakevdp.github.io/PythonDataScienceHandbook/): Free digital book that is a great resource for learning pandas, NumPy, Matplotlib, and Seaborn.
|
||||
- [freeCodeCamp - Machine Learning for Everybody](https://youtu.be/i_LwzRVP7bg): Practical introduction to different machine learning algorithms for beginners.
|
||||
- [Udacity - Intro to Machine Learning](https://www.udacity.com/course/intro-to-machine-learning--ud120): Free course that covers PCA and several other machine learning concepts.
|
||||
- [Real Python](https://realpython.com/):) 涵盖初学者与进阶 Python 概念的综合性文章与教程资源。
|
||||
- [freeCodeCamp - Learn Python](https://www.youtube.com/watch?v=rfscVS0vtbw):) 长视频课程,全面介绍 Python 的所有核心概念。
|
||||
- [Python Data Science Handbook](https://jakevdp.github.io/PythonDataScienceHandbook/):) 免费电子书,是学习 pandas、NumPy、Matplotlib 与 Seaborn 的优质资源。
|
||||
- [freeCodeCamp - Machine Learning for Everybody](https://youtu.be/i_LwzRVP7bg):) 面向初学者的实用机器学习算法入门。
|
||||
- [Udacity - Intro to Machine Learning](https://www.udacity.com/course/intro-to-machine-learning--ud120):) 免费课程,涵盖 PCA 及若干其他机器学习概念。
|
||||
|
||||
---
|
||||
|
||||
### 3. Neural Networks
|
||||
### 3. 神经网络
|
||||
|
||||
Neural networks are a fundamental part of many machine learning models, particularly in the realm of deep learning. To utilize them effectively, a comprehensive understanding of their design and mechanics is essential.
|
||||
神经网络是许多机器学习模型,尤其是深度学习领域的核心组成部分。要有效运用它们,必须全面理解其设计与运行机制。
|
||||
|
||||
- **Fundamentals**: This includes understanding the structure of a neural network, such as layers, weights, biases, and activation functions (sigmoid, tanh, ReLU, etc.)
|
||||
- **Training and Optimization**: Familiarize yourself with backpropagation and different types of loss functions, like Mean Squared Error (MSE) and Cross-Entropy. Understand various optimization algorithms like Gradient Descent, Stochastic Gradient Descent, RMSprop, and Adam.
|
||||
- **Overfitting**: Understand the concept of overfitting (where a model performs well on training data but poorly on unseen data) and learn various regularization techniques (dropout, L1/L2 regularization, early stopping, data augmentation) to prevent it.
|
||||
- **Implement a Multilayer Perceptron (MLP)**: Build an MLP, also known as a fully connected network, using PyTorch.
|
||||
- **基础概念**:包括理解神经网络的结构,例如层、权重、偏置以及激活函数(sigmoid、tanh、ReLU 等)。
|
||||
- **训练与优化**:熟悉反向传播(backpropagation)以及不同类型的损失函数,如均方误差(Mean Squared Error,MSE)与交叉熵(Cross-Entropy)。理解梯度下降(Gradient Descent)、随机梯度下降(Stochastic Gradient Descent)、RMSprop 与 Adam 等优化算法。
|
||||
- **过拟合(Overfitting)**:理解过拟合的概念(模型在训练数据上表现良好,但在未见数据上表现较差),并学习各种正则化技术(dropout、L1/L2 正则化、早停、数据增强)来防止过拟合。
|
||||
- **实现多层感知机(Multilayer Perceptron,MLP)**:使用 PyTorch 构建 MLP,也称为全连接网络(fully connected network)。
|
||||
|
||||
📚 Resources:
|
||||
📚 学习资源:
|
||||
|
||||
- [3Blue1Brown - But what is a Neural Network?](https://www.youtube.com/watch?v=aircAruvnKk): This video gives an intuitive explanation of neural networks and their inner workings.
|
||||
- [freeCodeCamp - Deep Learning Crash Course](https://www.youtube.com/watch?v=VyWAvY2CF9c): This video efficiently introduces all the most important concepts in deep learning.
|
||||
- [Fast.ai - Practical Deep Learning](https://course.fast.ai/): Free course designed for people with coding experience who want to learn about deep learning.
|
||||
- [Patrick Loeber - PyTorch Tutorials](https://www.youtube.com/playlist?list=PLqnslRFeH2UrcDBWF5mfPGpqQDSta6VK4): Series of videos for complete beginners to learn about PyTorch.
|
||||
- [3Blue1Brown - But what is a Neural Network?](https://www.youtube.com/watch?v=aircAruvnKk):) 该视频直观讲解了神经网络及其内部工作原理。
|
||||
- [freeCodeCamp - Deep Learning Crash Course](https://www.youtube.com/watch?v=VyWAvY2CF9c):) 该视频高效介绍了深度学习中所有最重要的概念。
|
||||
- [Fast.ai - Practical Deep Learning](https://course.fast.ai/):) 面向有编程经验、希望学习深度学习者的免费课程。
|
||||
- [Patrick Loeber - PyTorch Tutorials](https://www.youtube.com/playlist?list=PLqnslRFeH2UrcDBWF5mfPGpqQDSta6VK4):) 面向零基础初学者的 PyTorch 系列视频教程。
|
||||
|
||||
---
|
||||
|
||||
### 4. Natural Language Processing (NLP)
|
||||
### 4. 自然语言处理(Natural Language Processing,NLP)
|
||||
|
||||
NLP is a fascinating branch of artificial intelligence that bridges the gap between human language and machine understanding. From simple text processing to understanding linguistic nuances, NLP plays a crucial role in many applications like translation, sentiment analysis, chatbots, and much more.
|
||||
NLP 是人工智能中一个引人入胜的分支,它弥合人类语言与机器理解之间的鸿沟。从简单的文本处理到理解语言细微差别,NLP 在翻译、情感分析、聊天机器人等众多应用中发挥着关键作用。
|
||||
|
||||
- **Text Preprocessing**: Learn various text preprocessing steps like tokenization (splitting text into words or sentences), stemming (reducing words to their root form), lemmatization (similar to stemming but considers the context), stop word removal, etc.
|
||||
- **Feature Extraction Techniques**: Become familiar with techniques to convert text data into a format that can be understood by machine learning algorithms. Key methods include Bag-of-words (BoW), Term Frequency-Inverse Document Frequency (TF-IDF), and n-grams.
|
||||
- **Word Embeddings**: Word embeddings are a type of word representation that allows words with similar meanings to have similar representations. Key methods include Word2Vec, GloVe, and FastText.
|
||||
- **Recurrent Neural Networks (RNNs)**: Understand the working of RNNs, a type of neural network designed to work with sequence data. Explore LSTMs and GRUs, two RNN variants that are capable of learning long-term dependencies.
|
||||
- **文本预处理**:学习各种文本预处理步骤,如分词(tokenization,将文本拆分为词或句子)、词干提取(stemming,将词还原为词根形式)、词形还原(lemmatization,与词干提取类似但会考虑上下文)、停用词移除等。
|
||||
- **特征提取技术**:熟悉将文本数据转换为机器学习算法可理解格式的技术。主要方法包括词袋模型(Bag-of-words,BoW)、词频-逆文档频率(Term Frequency-Inverse Document Frequency,TF-IDF)以及 n-gram。
|
||||
- **词嵌入(Word Embeddings)**:词嵌入是一种词表示方法,使语义相近的词具有相近的表示。主要方法包括 Word2Vec、GloVe 与 FastText。
|
||||
- **循环神经网络(Recurrent Neural Networks,RNNs)**:理解 RNN 的工作原理——这是一种专为序列数据设计的神经网络。探索 LSTM 与 GRU,这两种能够学习长期依赖关系的 RNN 变体。
|
||||
|
||||
📚 Resources:
|
||||
📚 学习资源:
|
||||
|
||||
- [Lena Voita - Word Embeddings](https://lena-voita.github.io/nlp_course/word_embeddings.html): Beginner-friendly course about concepts related to word embeddings.
|
||||
- [RealPython - NLP with spaCy in Python](https://realpython.com/natural-language-processing-spacy-python/): Exhaustive guide about the spaCy library for NLP tasks in Python.
|
||||
- [Kaggle - NLP Guide](https://www.kaggle.com/learn-guide/natural-language-processing): A few notebooks and resources for a hands-on explanation of NLP in Python.
|
||||
- [Jay Alammar - The Illustration Word2Vec](https://jalammar.github.io/illustrated-word2vec/): A good reference to understand the famous Word2Vec architecture.
|
||||
- [Jake Tae - PyTorch RNN from Scratch](https://jaketae.github.io/study/pytorch-rnn/): Practical and simple implementation of RNN, LSTM, and GRU models in PyTorch.
|
||||
- [colah's blog - Understanding LSTM Networks](https://colah.github.io/posts/2015-08-Understanding-LSTMs/): A more theoretical article about the LSTM network.
|
||||
- [Lena Voita - Word Embeddings](https://lena-voita.github.io/nlp_course/word_embeddings.html):) 面向初学者的词嵌入相关概念课程。
|
||||
- [RealPython - NLP with spaCy in Python](https://realpython.com/natural-language-processing-spacy-python/):) 关于在 Python 中使用 spaCy 库完成 NLP 任务的详尽指南。
|
||||
- [Kaggle - NLP Guide](https://www.kaggle.com/learn-guide/natural-language-processing):) 若干 notebook 与资源,用于动手讲解 Python 中的 NLP。
|
||||
- [Jay Alammar - The Illustration Word2Vec](https://jalammar.github.io/illustrated-word2vec/):) 理解著名 Word2Vec 架构的优质参考资料。
|
||||
- [Jake Tae - PyTorch RNN from Scratch](https://jaketae.github.io/study/pytorch-rnn/):) 在 PyTorch 中实现 RNN、LSTM 与 GRU 模型的实用且简洁示例。
|
||||
- [colah's blog - Understanding LSTM Networks](https://colah.github.io/posts/2015-08-Understanding-LSTMs/):) 一篇关于 LSTM 网络的偏理论文章。
|
||||
</details>
|
||||
|
||||
## 🧑🔬 The LLM Scientist
|
||||
## 🧑🔬 LLM 科学家
|
||||
|
||||
This section of the course focuses on learning how to build the best possible LLMs using the latest techniques.
|
||||
本课程的这一部分侧重于学习如何运用最新技术构建尽可能优秀的 LLM。
|
||||
|
||||

|
||||
|
||||
### 1. The LLM Architecture
|
||||
### 1. LLM 架构
|
||||
|
||||
An in-depth knowledge of the Transformer architecture is not required, but it's important to understand the main steps of modern LLMs: converting text into numbers through tokenization, processing these tokens through layers including attention mechanisms, and finally generating new text through various sampling strategies.
|
||||
不必深入掌握 Transformer 架构,但理解现代 LLM 的主要步骤很重要:通过分词(tokenization)将文本转换为数字,通过包含注意力机制的各层处理这些 token,最后通过多种采样策略生成新文本。
|
||||
|
||||
- **Architectural overview**: Understand the evolution from encoder-decoder Transformers to decoder-only architectures like GPT, which form the basis of modern LLMs. Focus on how these models process and generate text at a high level.
|
||||
- **Tokenization**: Learn the principles of tokenization - how text is converted into numerical representations that LLMs can process. Explore different tokenization strategies and their impact on model performance and output quality.
|
||||
- **Attention mechanisms**: Master the core concepts of attention mechanisms, particularly self-attention and its variants. Understand how these mechanisms enable LLMs to process long-range dependencies and maintain context throughout sequences.
|
||||
- **Sampling techniques**: Explore various text generation approaches and their tradeoffs. Compare deterministic methods like greedy search and beam search with probabilistic approaches like temperature sampling and nucleus sampling.
|
||||
- **架构概览**:理解从编码器-解码器(encoder-decoder)Transformer 到仅解码器(decoder-only)架构(如 GPT)的演进,后者构成了现代 LLM 的基础。重点关注这些模型在宏观层面如何处理和生成文本。
|
||||
- **分词(Tokenization)**:学习分词原理——文本如何转换为 LLM 可处理的数值表示。探索不同分词策略及其对模型性能和输出质量的影响。
|
||||
- **注意力机制**:掌握注意力机制的核心概念,尤其是自注意力(self-attention)及其变体。理解这些机制如何使 LLM 能够处理长程依赖并在整个序列中保持上下文。
|
||||
- **采样技术**:探索各种文本生成方法及其权衡。将贪心搜索(greedy search)、束搜索(beam search)等确定性方法与温度采样(temperature sampling)、核采样(nucleus sampling)等概率方法进行比较。
|
||||
|
||||
📚 **References**:
|
||||
* [Visual intro to Transformers](https://www.youtube.com/watch?v=wjZofJX0v4M) by 3Blue1Brown: Visual introduction to Transformers for complete beginners.
|
||||
* [LLM Visualization](https://bbycroft.net/llm) by Brendan Bycroft: Interactive 3D visualization of LLM internals.
|
||||
* [nanoGPT](https://www.youtube.com/watch?v=kCc8FmEb1nY) by Andrej Karpathy: A 2h-long YouTube video to reimplement GPT from scratch (for programmers). He also made a video about [tokenization](https://www.youtube.com/watch?v=zduSFxRajkE).
|
||||
* [Attention? Attention!](https://lilianweng.github.io/posts/2018-06-24-attention/) by Lilian Weng: Historical overview to introduce the need for attention mechanisms.
|
||||
* [Decoding Strategies in LLMs](https://mlabonne.github.io/blog/posts/2023-06-07-Decoding_strategies.html) by Maxime Labonne: Provide code and a visual introduction to the different decoding strategies to generate text.
|
||||
📚 **参考资料**:
|
||||
* [Transformers 可视化入门](https://www.youtube.com/watch?v=wjZofJX0v4M) by 3Blue1Brown:面向零基础读者的 Transformers 可视化介绍。
|
||||
* [LLM 可视化](https://bbycroft.net/llm) by Brendan Bycroft:LLM 内部结构的交互式 3D 可视化。
|
||||
* [nanoGPT](https://www.youtube.com/watch?v=kCc8FmEb1nY) by Andrej Karpathy:长达 2 小时的 YouTube 视频,从零实现 GPT(面向程序员)。他还制作了一个关于 [分词(tokenization)](https://www.youtube.com/watch?v=zduSFxRajkE).
|
||||
* [Attention? Attention!](https://lilianweng.github.io/posts/2018-06-24-attention/) by Lilian Weng:历史综述,介绍注意力机制的必要性。
|
||||
* [LLM 中的解码策略](https://mlabonne.github.io/blog/posts/2023-06-07-Decoding_strategies.html) by Maxime Labonne:提供代码及不同文本生成解码策略的可视化介绍。
|
||||
|
||||
---
|
||||
### 2. Pre-Training Models
|
||||
### 2. 预训练模型
|
||||
|
||||
Pre-training is a computationally intensive and expensive process. While it's not the focus of this course, it's important to have a solid understanding of how models are pre-trained, especially in terms of data and parameters. Pre-training can also be performed by hobbyists at a small scale with <1B models.
|
||||
预训练是计算密集且成本高昂的过程。虽然本课程不以此为重点,但扎实理解模型如何预训练(尤其在数据和参数方面)很重要。爱好者也可以小规模预训练参数量小于 1B 的模型。
|
||||
|
||||
* **Data preparation**: Pre-training requires massive datasets (e.g., [Llama 3.1](https://arxiv.org/abs/2307.09288) was trained on 15 trillion tokens) that need careful curation, cleaning, deduplication, and tokenization. Modern pre-training pipelines implement sophisticated filtering to remove low-quality or problematic content.
|
||||
* **Distributed training**: Combine different parallelization strategies: data parallel (batch distribution), pipeline parallel (layer distribution), and tensor parallel (operation splitting). These strategies require optimized network communication and memory management across GPU clusters.
|
||||
* **Training optimization**: Use adaptive learning rates with warm-up, gradient clipping, and normalization to prevent explosions, mixed-precision training for memory efficiency, and modern optimizers (AdamW, Lion) with tuned hyperparameters.
|
||||
* **Monitoring**: Track key metrics (loss, gradients, GPU stats) using dashboards, implement targeted logging for distributed training issues, and set up performance profiling to identify bottlenecks in computation and communication across devices.
|
||||
* **数据准备**:预训练需要海量数据集(例如,[Llama 3.1](https://arxiv.org/abs/2307.09288) 在 15 万亿 token 上训练),需精心策划、清洗、去重和分词。现代预训练流水线会实施复杂过滤,以剔除低质量或有问题的内容。
|
||||
* **分布式训练**:组合多种并行策略:数据并行(批次分布)、流水线并行(层分布)和张量并行(运算拆分)。这些策略需要在 GPU 集群间优化网络通信和内存管理。
|
||||
* **训练优化**:使用带预热(warm-up)的自适应学习率、梯度裁剪和归一化以防止爆炸,采用混合精度训练以提高内存效率,并使用调优超参数的现代优化器(AdamW、Lion)。
|
||||
* **监控**:使用仪表盘跟踪关键指标(损失、梯度、GPU 统计),针对分布式训练问题实施定向日志记录,并设置性能剖析以识别跨设备的计算与通信瓶颈。
|
||||
|
||||
📚 **References**:
|
||||
* [FineWeb](https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1) by Penedo et al.: Article to recreate a large-scale dataset for LLM pretraining (15T), including FineWeb-Edu, a high-quality subset.
|
||||
* [RedPajama v2](https://www.together.ai/blog/redpajama-data-v2) by Weber et al.: Another article and paper about a large-scale pre-training dataset with a lot of interesting quality filters.
|
||||
* [nanotron](https://github.com/huggingface/nanotron) by Hugging Face: Minimalistic LLM training codebase used to make [SmolLM2](https://github.com/huggingface/smollm).
|
||||
* [Parallel training](https://www.andrew.cmu.edu/course/11-667/lectures/W10L2%20Scaling%20Up%20Parallel%20Training.pdf) by Chenyan Xiong: Overview of optimization and parallelism techniques.
|
||||
* [Distributed training](https://arxiv.org/abs/2407.20018) by Duan et al.: A survey about efficient training of LLM on distributed architectures.
|
||||
* [OLMo 2](https://allenai.org/olmo) by AI2: Open-source language model with model, data, training, and evaluation code.
|
||||
* [LLM360](https://www.llm360.ai/) by LLM360: A framework for open-source LLMs with training and data preparation code, data, metrics, and models.
|
||||
📚 **参考资料**:
|
||||
* [FineWeb](https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1) by Penedo et al.:重现用于 LLM 预训练的大规模数据集(15T)的文章,包含高质量子集 FineWeb-Edu。
|
||||
* [RedPajama v2](https://www.together.ai/blog/redpajama-data-v2) by Weber et al.:另一篇关于大规模预训练数据集的文章和论文,包含诸多有趣的质量过滤器。
|
||||
* [nanotron](https://github.com/huggingface/nanotron) by Hugging Face:用于训练 [SmolLM2](https://github.com/huggingface/smollm). 的极简 LLM 训练代码库
|
||||
* [并行训练](https://www.andrew.cmu.edu/course/11-667/lectures/W10L2%20Scaling%20Up%20Parallel%20Training.pdf) by Chenyan Xiong:优化与并行技术概览。
|
||||
* [分布式训练](https://arxiv.org/abs/2407.20018) by Duan et al.:关于在分布式架构上高效训练 LLM 的综述。
|
||||
* [OLMo 2](https://allenai.org/olmo) by AI2:开源语言模型,包含模型、数据、训练与评估代码。
|
||||
* [LLM360](https://www.llm360.ai/) by LLM360:开源 LLM 框架,提供训练与数据准备代码、数据、指标和模型。
|
||||
|
||||
---
|
||||
### 3. Post-Training Datasets
|
||||
### 3. 后训练数据集
|
||||
|
||||
Post-training datasets have a precise structure with instructions and answers (supervised fine-tuning) or instructions and chosen/rejected answers (preference alignment). Conversational structures are a lot rarer than the raw text used for pre-training, which is why we often need to process seed data and refine it to improve the accuracy, diversity, and complexity of the samples. More information and examples are available in my repo [💾 LLM Datasets](https://github.com/mlabonne/llm-datasets).
|
||||
后训练数据集具有精确结构:指令与答案(监督微调),或指令与优选/拒绝答案(偏好对齐)。对话结构远比预训练所用的原始文本稀少,因此我们常需处理种子数据并精炼样本,以提高准确性、多样性和复杂度。更多信息与示例见我的仓库 [💾 LLM Datasets](https://github.com/mlabonne/llm-datasets).
|
||||
|
||||
* **Storage & chat templates**: Because of the conversational structure, post-training datasets are stored in a specific format like ShareGPT or OpenAI/HF. Then, these formats are mapped to a chat template like ChatML or Alpaca to produce the final samples that the model is trained on.
|
||||
* **Synthetic data generation**: Create instruction-response pairs based on seed data using frontier models like GPT-4o. This approach allows for flexible and scalable dataset creation with high-quality answers. Key considerations include designing diverse seed tasks and effective system prompts.
|
||||
* **Data enhancement**: Enhance existing samples using techniques like verified outputs (using unit tests or solvers), multiple answers with rejection sampling, [Auto-Evol](https://arxiv.org/abs/2406.00770), Chain-of-Thought, Branch-Solve-Merge, personas, etc.
|
||||
* **Quality filtering**: Traditional techniques involve rule-based filtering, removing duplicates or near-duplicates (with MinHash or embeddings), and n-gram decontamination. Reward models and judge LLMs complement this step with fine-grained and customizable quality control.
|
||||
* **存储与对话模板**:由于对话结构,后训练数据集以 ShareGPT 或 OpenAI/HF 等特定格式存储。随后将这些格式映射到 ChatML 或 Alpaca 等对话模板,生成模型最终训练所用的样本。
|
||||
* **合成数据生成**:基于种子数据,使用 GPT-4o 等前沿模型创建指令-回答对。该方法可灵活、可扩展地构建高质量回答的数据集。关键考量包括设计多样化种子任务和有效的系统提示(system prompts)。
|
||||
* **数据增强**:使用经验证输出(单元测试或求解器)、带拒绝采样的多答案、[Auto-Evol](https://arxiv.org/abs/2406.00770), 思维链(Chain-of-Thought)、Branch-Solve-Merge、角色人设(personas)等技术增强现有样本。
|
||||
* **质量过滤**:传统技术包括基于规则的过滤、去除重复或近重复项(MinHash 或嵌入向量),以及 n-gram 去污染。奖励模型和评判 LLM 以细粒度、可定制的质量控制补充这一步骤。
|
||||
|
||||
📚 **References**:
|
||||
* [Synthetic Data Generator](https://huggingface.co/spaces/argilla/synthetic-data-generator) by Argilla: Beginner-friendly way of building datasets using natural language in a Hugging Face space.
|
||||
* [LLM Datasets](https://github.com/mlabonne/llm-datasets) by Maxime Labonne: Curated list of datasets and tools for post-training.
|
||||
* [NeMo-Curator](https://github.com/NVIDIA/NeMo-Curator) by Nvidia: Dataset preparation and curation framework for pre- and post-training data.
|
||||
* [Distilabel](https://distilabel.argilla.io/dev/sections/pipeline_samples/) by Argilla: Framework to generate synthetic data. It also includes interesting reproductions of papers like UltraFeedback.
|
||||
* [Semhash](https://github.com/MinishLab/semhash) by MinishLab: Minimalistic library for near-deduplication and decontamination with a distilled embedding model.
|
||||
* [Chat Template](https://huggingface.co/docs/transformers/main/en/chat_templating) by Hugging Face: Hugging Face's documentation about chat templates.
|
||||
📚 **参考资料**:
|
||||
* [Synthetic Data Generator](https://huggingface.co/spaces/argilla/synthetic-data-generator) by Argilla:在 Hugging Face Space 中用自然语言构建数据集的入门友好方式。
|
||||
* [LLM Datasets](https://github.com/mlabonne/llm-datasets) by Maxime Labonne:后训练数据集与工具的精选列表。
|
||||
* [NeMo-Curator](https://github.com/NVIDIA/NeMo-Curator) by Nvidia:用于预训练与后训练数据的数据集准备与策划框架。
|
||||
* [Distilabel](https://distilabel.argilla.io/dev/sections/pipeline_samples/) by Argilla:生成合成数据的框架,还包含 UltraFeedback 等论文的有趣复现。
|
||||
* [Semhash](https://github.com/MinishLab/semhash) by MinishLab:使用蒸馏嵌入模型进行近去重与去污染的极简库。
|
||||
* [Chat Template](https://huggingface.co/docs/transformers/main/en/chat_templating) by Hugging Face:Hugging Face 关于对话模板的文档。
|
||||
|
||||
---
|
||||
### 4. Supervised Fine-Tuning
|
||||
### 4. 监督微调
|
||||
|
||||
SFT turns base models into helpful assistants, capable of answering questions and following instructions. During this process, they learn how to structure answers and reactivate a subset of knowledge learned during pre-training. Instilling new knowledge is possible but superficial: it cannot be used to learn a completely new language. Always prioritize data quality over parameter optimization.
|
||||
监督微调(SFT)将基座模型转化为有用的助手,能够回答问题并遵循指令。在此过程中,模型学习如何组织答案,并重新激活预训练所学知识的子集。注入新知识是可能的,但较为肤浅:无法用于学习一门全新的语言。始终将数据质量置于参数优化之上。
|
||||
|
||||
- **Training techniques**: Full fine-tuning updates all model parameters but requires significant compute. Parameter-efficient fine-tuning techniques like LoRA and QLoRA reduce memory requirements by training a small number of adapter parameters while keeping base weights frozen. QLoRA combines 4-bit quantization with LoRA to reduce VRAM usage. These techniques are all implemented in the most popular fine-tuning frameworks: [TRL](https://huggingface.co/docs/trl/en/index), [Unsloth](https://docs.unsloth.ai/), and [Axolotl](https://axolotl.ai/).
|
||||
- **Training parameters**: Key parameters include learning rate with schedulers, batch size, gradient accumulation, number of epochs, optimizer (like 8-bit AdamW), weight decay for regularization, and warmup steps for training stability. LoRA also adds three parameters: rank (typically 16-128), alpha (1-2x rank), and target modules.
|
||||
- **Distributed training**: Scale training across multiple GPUs using DeepSpeed or FSDP. DeepSpeed provides three ZeRO optimization stages with increasing levels of memory efficiency through state partitioning. Both methods support gradient checkpointing for memory efficiency.
|
||||
- **Monitoring**: Track training metrics including loss curves, learning rate schedules, and gradient norms. Monitor for common issues like loss spikes, gradient explosions, or performance degradation.
|
||||
- **训练技术**:全量微调(Full fine-tuning)会更新所有模型参数,但需要大量算力。参数高效微调(Parameter-efficient fine-tuning)技术如 LoRA 和 QLoRA 通过训练少量适配器参数并保持基础权重冻结来降低内存需求。QLoRA 将 4-bit 量化与 LoRA 结合以降低 VRAM 使用。这些技术都在最流行的微调框架中实现:[TRL](https://huggingface.co/docs/trl/en/index), [Unsloth](https://docs.unsloth.ai/), and [Axolotl](https://axolotl.ai/).
|
||||
- **训练参数**:关键参数包括带调度器的学习率、批次大小、梯度累积、训练轮数(epochs)、优化器(如 8-bit AdamW)、用于正则化的权重衰减(weight decay),以及用于训练稳定性的预热步数(warmup steps)。LoRA 还增加了三个参数:秩(rank,通常为 16-128)、alpha(通常为 rank 的 1-2 倍)以及目标模块(target modules)。
|
||||
- **分布式训练**:使用 DeepSpeed 或 FSDP 在多块 GPU 上扩展训练。DeepSpeed 提供三个 ZeRO 优化阶段,通过状态分区实现逐级提升的内存效率。两种方法都支持梯度检查点(gradient checkpointing)以节省内存。
|
||||
- **监控**:跟踪训练指标,包括损失曲线、学习率调度和梯度范数。关注常见问题,如损失突增、梯度爆炸或性能下降。
|
||||
|
||||
📚 **References**:
|
||||
* [Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth](https://huggingface.co/blog/mlabonne/sft-llama3) by Maxime Labonne: Hands-on tutorial on how to fine-tune a Llama 3.1 model using Unsloth.
|
||||
* [Axolotl - Documentation](https://axolotl-ai-cloud.github.io/axolotl/) by Wing Lian: Lots of interesting information related to distributed training and dataset formats.
|
||||
* [Mastering LLMs](https://parlance-labs.com/education/) by Hamel Husain: Collection of educational resources about fine-tuning (but also RAG, evaluation, applications, and prompt engineering).
|
||||
* [LoRA insights](https://lightning.ai/pages/community/lora-insights/) by Sebastian Raschka: Practical insights about LoRA and how to select the best parameters.
|
||||
📚 **参考资料**:
|
||||
* [Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth](https://huggingface.co/blog/mlabonne/sft-llama3) by Maxime Labonne:使用 Unsloth 微调 Llama 3.1 模型的动手教程。
|
||||
* [Axolotl - Documentation](https://axolotl-ai-cloud.github.io/axolotl/) by Wing Lian:大量与分布式训练和数据集格式相关的实用信息。
|
||||
* [Mastering LLMs](https://parlance-labs.com/education/) by Hamel Husain:关于微调的教育资源合集(也涵盖 RAG、评估、应用和提示工程)。
|
||||
* [LoRA insights](https://lightning.ai/pages/community/lora-insights/) by Sebastian Raschka:关于 LoRA 的实用见解以及如何选择最佳参数。
|
||||
|
||||
---
|
||||
### 5. Preference Alignment
|
||||
### 5. 偏好对齐
|
||||
|
||||
Preference alignment is a second stage in the post-training pipeline, focused on aligning generated answers with human preferences. This stage was designed to tune the tone of LLMs and reduce toxicity and hallucinations. However, it has become increasingly important to also boost their performance and improve their usefulness. Unlike SFT, there are many preference alignment algorithms. Here, we'll focus on the three most important ones: DPO, GRPO, and PPO.
|
||||
偏好对齐是后训练流程中的第二阶段,专注于使生成答案与人类偏好对齐。该阶段旨在调整 LLM 的语气并减少毒性与幻觉。然而,提升性能与实用性也变得越来越重要。与 SFT 不同,偏好对齐算法有很多种。此处我们聚焦三种最重要的:DPO、GRPO 和 PPO。
|
||||
|
||||
- **Rejection sampling**: For each prompt, use the trained model to generate multiple responses, and score them to infer chosen/rejected answers. This creates on-policy data, where both responses come from the model being trained, improving alignment stability.
|
||||
- **[Direct Preference Optimization](https://arxiv.org/abs/2305.18290)** Directly optimizes the policy to maximize the likelihood of chosen responses over rejected ones. It doesn't require reward modeling, which makes it more computationally efficient than RL techniques but slightly worse in terms of quality. Great for creating chat models.
|
||||
- **Reward model**: Train a reward model with human feedback to predict metrics like human preferences. It can leverage frameworks like [TRL](https://huggingface.co/docs/trl/en/index), [verl](https://github.com/volcengine/verl), and [OpenRLHF](https://github.com/OpenRLHF/OpenRLHF) for scalable training.
|
||||
- **Reinforcement Learning**: RL techniques like [GRPO](https://arxiv.org/abs/2402.03300) and [PPO](https://arxiv.org/abs/1707.06347) iteratively update a policy to maximize rewards while staying close to the initial behavior. They can use a reward model or reward functions to score responses. They tend to be computationally expensive and require careful tuning of hyperparameters, including learning rate, batch size, and clip range. Ideal for creating reasoning models.
|
||||
- **拒绝采样(Rejection sampling)**:对每个提示,使用已训练模型生成多个回复并打分,以推断被选/被拒绝的答案。这会创建 on-policy 数据,其中两条回复均来自正在训练的模型,从而提升对齐稳定性。
|
||||
- **[Direct Preference Optimization](https://arxiv.org/abs/2305.18290)** 直接优化策略,以最大化被选回复相对被拒绝回复的似然。它不需要奖励建模,因此在计算上比 RL 技术更高效,但质量略逊一筹。非常适合创建聊天模型。
|
||||
- **奖励模型(Reward model)**:用人类反馈训练奖励模型,以预测人类偏好等指标。可借助 [TRL](https://huggingface.co/docs/trl/en/index), [verl](https://github.com/volcengine/verl), and [OpenRLHF](https://github.com/OpenRLHF/OpenRLHF) 等框架进行可扩展训练。
|
||||
- **强化学习(Reinforcement Learning)**:像 [GRPO](https://arxiv.org/abs/2402.03300) and [PPO](https://arxiv.org/abs/1707.06347) 这样的 RL 技术会迭代更新策略,在最大化奖励的同时尽量贴近初始行为。它们可使用奖励模型或奖励函数为回复打分。通常计算开销大,且需仔细调优超参数,包括学习率、批次大小和裁剪范围(clip range)。适合创建推理模型。
|
||||
|
||||
📚 **References**:
|
||||
* [Illustrating RLHF](https://huggingface.co/blog/rlhf) by Hugging Face: Introduction to RLHF with reward model training and fine-tuning with reinforcement learning.
|
||||
* [LLM Training: RLHF and Its Alternatives](https://magazine.sebastianraschka.com/p/llm-training-rlhf-and-its-alternatives) by Sebastian Raschka: Overview of the RLHF process and alternatives like RLAIF.
|
||||
* [Preference Tuning LLMs](https://huggingface.co/blog/pref-tuning) by Hugging Face: Comparison of the DPO, IPO, and KTO algorithms to perform preference alignment.
|
||||
* [Fine-tune with DPO](https://mlabonne.github.io/blog/posts/Fine_tune_Mistral_7b_with_DPO.html) by Maxime Labonne: Tutorial to fine-tune a Mistral-7b model with DPO and reproduce [NeuralHermes-2.5](https://huggingface.co/mlabonne/NeuralHermes-2.5-Mistral-7B).
|
||||
* [Fine-tune with GRPO](https://huggingface.co/learn/llm-course/en/chapter12/5) by Maxime Labonne: Practical exercise to fine-tune a small model with GRPO.
|
||||
* [DPO Wandb logs](https://wandb.ai/alexander-vishnevskiy/dpo/reports/TRL-Original-DPO--Vmlldzo1NjI4MTc4) by Alexander Vishnevskiy: It shows you the main DPO metrics to track and the trends you should expect.
|
||||
📚 **参考资料**:
|
||||
* [Illustrating RLHF](https://huggingface.co/blog/rlhf) by Hugging Face:介绍 RLHF,包括奖励模型训练与强化学习微调。
|
||||
* [LLM Training: RLHF and Its Alternatives](https://magazine.sebastianraschka.com/p/llm-training-rlhf-and-its-alternatives) by Sebastian Raschka:概述 RLHF 流程及 RLAIF 等替代方案。
|
||||
* [Preference Tuning LLMs](https://huggingface.co/blog/pref-tuning) by Hugging Face:比较 DPO、IPO 和 KTO 算法以进行偏好对齐。
|
||||
* [Fine-tune with DPO](https://mlabonne.github.io/blog/posts/Fine_tune_Mistral_7b_with_DPO.html) by Maxime Labonne:使用 DPO 微调 Mistral-7b 模型并复现 [NeuralHermes-2.5](https://huggingface.co/mlabonne/NeuralHermes-2.5-Mistral-7B). 的教程。
|
||||
* [Fine-tune with GRPO](https://huggingface.co/learn/llm-course/en/chapter12/5) by Maxime Labonne:使用 GRPO 微调小模型的实践练习。
|
||||
* [DPO Wandb logs](https://wandb.ai/alexander-vishnevskiy/dpo/reports/TRL-Original-DPO--Vmlldzo1NjI4MTc4) by Alexander Vishnevskiy:展示应跟踪的主要 DPO 指标及预期趋势。
|
||||
|
||||
---
|
||||
### 6. Evaluation
|
||||
### 6. 评估
|
||||
|
||||
Reliably evaluating LLMs is a complex but essential task guiding data generation and training. It provides invaluable feedback about areas of improvement, which can be leveraged to modify the data mixture, quality, and training parameters. However, it's always good to remember Goodhart's law: "When a measure becomes a target, it ceases to be a good measure."
|
||||
可靠地评估 LLM 是一项复杂但至关重要的任务,可指导数据生成与训练。它能提供关于改进方向的宝贵反馈,可用于调整数据配比、质量和训练参数。但始终要记住古德哈特定律(Goodhart's law):「当一项指标成为目标时,它就不再是好的指标。」
|
||||
|
||||
- **Automated benchmarks**: Evaluate models on specific tasks using curated datasets and metrics, like MMLU. It works well for concrete tasks but struggles with abstract and creative capabilities. It is also prone to data contamination.
|
||||
- **Human evaluation**: It involves humans prompting models and grading responses. Methods range from vibe checks to systematic annotations with specific guidelines and large-scale community voting (arena). It is more suited for subjective tasks and less reliable for factual accuracy.
|
||||
- **Model-based evaluation**: Use judge and reward models to evaluate model outputs. It highly correlates with human preferences but suffers from bias toward their own outputs and inconsistent scoring.
|
||||
- **Feedback signal**: Analyze error patterns to identify specific weaknesses, such as limitations in following complex instructions, lack of specific knowledge, or susceptibility to adversarial prompts. This can be improved with better data generation and training parameters.
|
||||
- **自动化基准测试**:使用精选数据集和指标(如 MMLU)在特定任务上评估模型。对具体任务效果较好,但在抽象与创造性能力上表现不足,也容易出现数据污染。
|
||||
- **人工评估**:由人类提示模型并对回复评分。方法从 vibe check 到带明确指南的系统化标注,再到大规模社区投票(arena)不等。更适合主观任务,在事实准确性上可靠性较低。
|
||||
- **基于模型的评估**:使用评判模型和奖励模型评估模型输出。与人类偏好高度相关,但存在偏向自身输出和评分不一致的问题。
|
||||
- **反馈信号**:分析错误模式以识别具体弱点,例如遵循复杂指令能力不足、缺乏特定知识,或易受对抗性提示影响。可通过改进数据生成与训练参数加以改善。
|
||||
|
||||
📚 **References**:
|
||||
* [LLM evaluation guidebook](https://huggingface.co/spaces/OpenEvals/evaluation-guidebook) by Hugging Face: Comprehensive guide about evaluation with practical insights.
|
||||
* [Open LLM Leaderboard](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard) by Hugging Face: Main leaderboard to compare LLMs in an open and reproducible way (automated benchmarks).
|
||||
* [Language Model Evaluation Harness](https://github.com/EleutherAI/lm-evaluation-harness) by EleutherAI: A popular framework for evaluating LLMs using automated benchmarks.
|
||||
* [Lighteval](https://github.com/huggingface/lighteval) by Hugging Face: Alternative evaluation framework that also includes model-based evaluations.
|
||||
* [Chatbot Arena](https://lmarena.ai/) by LMSYS: Elo rating of general-purpose LLMs, based on comparisons made by humans (human evaluation).
|
||||
📚 **参考资料**:
|
||||
* [LLM evaluation guidebook](https://huggingface.co/spaces/OpenEvals/evaluation-guidebook) by Hugging Face:关于评估的综合指南,含实用见解。
|
||||
* [Open LLM Leaderboard](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard) by Hugging Face:以开放、可复现方式比较 LLM 的主要排行榜(自动化基准测试)。
|
||||
* [Language Model Evaluation Harness](https://github.com/EleutherAI/lm-evaluation-harness) by EleutherAI:使用自动化基准测试评估 LLM 的流行框架。
|
||||
* [Lighteval](https://github.com/huggingface/lighteval) by Hugging Face:另一评估框架,也包含基于模型的评估。
|
||||
* [Chatbot Arena](https://lmarena.ai/) by LMSYS:基于人类对比评分的通用 LLM Elo 评级(人工评估)。
|
||||
|
||||
---
|
||||
### 7. Quantization
|
||||
### 7. 量化
|
||||
|
||||
Quantization is the process of converting the parameters and activations of a model to a lower precision. For example, weights stored using 16 bits can be converted into a 4-bit representation. This technique has become increasingly important to reduce the computational and memory costs associated with LLMs.
|
||||
量化是将模型参数和激活转换为更低精度的过程。例如,以 16 位存储的权重可转换为 4 位表示。该技术对降低与 LLM 相关的计算与内存成本日益重要。
|
||||
|
||||
* **Base techniques**: Learn the different levels of precision (FP32, FP16, INT8, etc.) and how to perform naïve quantization with absmax and zero-point techniques.
|
||||
* **GGUF & llama.cpp**: Originally designed to run on CPUs, [llama.cpp](https://github.com/ggerganov/llama.cpp) and the GGUF format have become the most popular tools to run LLMs on consumer-grade hardware. It supports storing special tokens, vocabulary, and metadata in a single file.
|
||||
* **GPTQ & AWQ**: Techniques like [GPTQ](https://arxiv.org/abs/2210.17323)/[EXL2](https://github.com/turboderp/exllamav2) and [AWQ](https://arxiv.org/abs/2306.00978) introduce layer-by-layer calibration that retains performance at extremely low bitwidths. They reduce catastrophic outliers using dynamic scaling, selectively skipping or re-centering the heaviest parameters.
|
||||
* **SmoothQuant & ZeroQuant**: New quantization-friendly transformations (SmoothQuant) and compiler-based optimizations (ZeroQuant) help mitigate outliers before quantization. They also reduce hardware overhead by fusing certain ops and optimizing dataflow.
|
||||
* **基础技术**:了解不同精度级别(FP32、FP16、INT8 等),以及如何使用 absmax 和 zero-point 技术进行朴素量化。
|
||||
* **GGUF & llama.cpp**:最初为在 CPU 上运行而设计,[llama.cpp](https://github.com/ggerganov/llama.cpp) 与 GGUF 格式已成为在消费级硬件上运行 LLM 的最流行工具。它支持在单个文件中存储特殊 token、词表和元数据。
|
||||
* **GPTQ & AWQ**:[GPTQ](https://arxiv.org/abs/2210.17323)/[EXL2](https://github.com/turboderp/exllamav2) and [AWQ](https://arxiv.org/abs/2306.00978) 等技术引入逐层校准,在极低比特宽度下仍能保持性能。它们通过动态缩放降低灾难性离群值影响,并有选择地跳过或重新居中最重参数。
|
||||
* **SmoothQuant & ZeroQuant**:新的量化友好变换(SmoothQuant)和基于编译器的优化(ZeroQuant)有助于在量化前缓解离群值。它们还通过融合特定算子并优化数据流来降低硬件开销。
|
||||
|
||||
📚 **References**:
|
||||
* [Introduction to quantization](https://mlabonne.github.io/blog/posts/Introduction_to_Weight_Quantization.html) by Maxime Labonne: Overview of quantization, absmax and zero-point quantization, and LLM.int8() with code.
|
||||
* [Quantize Llama models with llama.cpp](https://mlabonne.github.io/blog/posts/Quantize_Llama_2_models_using_ggml.html) by Maxime Labonne: Tutorial on how to quantize a Llama 2 model using llama.cpp and the GGUF format.
|
||||
* [4-bit LLM Quantization with GPTQ](https://mlabonne.github.io/blog/posts/4_bit_Quantization_with_GPTQ.html) by Maxime Labonne: Tutorial on how to quantize an LLM using the GPTQ algorithm with AutoGPTQ.
|
||||
* [Understanding Activation-Aware Weight Quantization](https://medium.com/friendliai/understanding-activation-aware-weight-quantization-awq-boosting-inference-serving-efficiency-in-10bb0faf63a8) by FriendliAI: Overview of the AWQ technique and its benefits.
|
||||
* [SmoothQuant on Llama 2 7B](https://github.com/mit-han-lab/smoothquant/blob/main/examples/smoothquant_llama_demo.ipynb) by MIT HAN Lab: Tutorial on how to use SmoothQuant with a Llama 2 model in 8-bit precision.
|
||||
* [DeepSpeed Model Compression](https://www.deepspeed.ai/tutorials/model-compression/) by DeepSpeed: Tutorial on how to use ZeroQuant and extreme compression (XTC) with DeepSpeed Compression.
|
||||
📚 **参考资料**:
|
||||
* [Introduction to quantization](https://mlabonne.github.io/blog/posts/Introduction_to_Weight_Quantization.html) by Maxime Labonne:量化概述,包括 absmax 与 zero-point 量化,以及带代码的 LLM.int8()。
|
||||
* [Quantize Llama models with llama.cpp](https://mlabonne.github.io/blog/posts/Quantize_Llama_2_models_using_ggml.html) by Maxime Labonne:使用 llama.cpp 与 GGUF 格式量化 Llama 2 模型的教程。
|
||||
* [4-bit LLM Quantization with GPTQ](https://mlabonne.github.io/blog/posts/4_bit_Quantization_with_GPTQ.html) by Maxime Labonne:使用 GPTQ 算法与 AutoGPTQ 量化 LLM 的教程。
|
||||
* [Understanding Activation-Aware Weight Quantization](https://medium.com/friendliai/understanding-activation-aware-weight-quantization-awq-boosting-inference-serving-efficiency-in-10bb0faf63a8) by FriendliAI:AWQ(Activation-Aware Weight Quantization,激活感知权重量化)技术及其优势概述。
|
||||
* [SmoothQuant on Llama 2 7B](https://github.com/mit-han-lab/smoothquant/blob/main/examples/smoothquant_llama_demo.ipynb) by MIT HAN Lab:在 8 位精度下将 SmoothQuant 用于 Llama 2 模型的教程。
|
||||
* [DeepSpeed Model Compression](https://www.deepspeed.ai/tutorials/model-compression/) by DeepSpeed:使用 DeepSpeed Compression 中的 ZeroQuant 与极端压缩(XTC)的教程。
|
||||
|
||||
---
|
||||
### 8. New Trends
|
||||
### 8. 新兴趋势
|
||||
|
||||
Here are notable topics that didn't fit into other categories. Some are established techniques (model merging, multimodal), but others are more experimental (interpretability, test-time compute scaling) and the focus of numerous research papers.
|
||||
以下是未归入其他类别的值得关注的主题。其中一些是成熟技术(模型合并、多模态),另一些则更具实验性(可解释性、测试时算力扩展),也是众多研究论文的焦点。
|
||||
|
||||
* **Model merging**: Merging trained models has become a popular way of creating performant models without any fine-tuning. The popular [mergekit](https://github.com/cg123/mergekit) library implements the most popular merging methods, like SLERP, [DARE](https://arxiv.org/abs/2311.03099), and [TIES](https://arxiv.org/abs/2311.03099).
|
||||
* **Multimodal models**: These models (like [CLIP](https://openai.com/research/clip), [Stable Diffusion](https://stability.ai/stable-image), or [LLaVA](https://llava-vl.github.io/)) process multiple types of inputs (text, images, audio, etc.) with a unified embedding space, which unlocks powerful applications like text-to-image.
|
||||
* **Interpretability**: Mechanistic interpretability techniques like Sparse Autoencoders (SAEs) have made remarkable progress to provide insights about the inner workings of LLMs. This has also been applied with techniques such as abliteration, which allow you to modify the behavior of models without training.
|
||||
* **Test-time compute**: Reasoning models trained with RL techniques can be further improved by scaling the compute budget during test time. It can involve multiple calls, MCTS, or specialized models like a Process Reward Model (PRM). Iterative steps with precise scoring significantly improve performance for complex reasoning tasks.
|
||||
* **Model merging(模型合并)**:合并已训练模型已成为一种流行的方式,无需任何微调即可创建高性能模型。流行的 [mergekit](https://github.com/cg123/mergekit) 库实现了最流行的合并方法,如 SLERP、[DARE](https://arxiv.org/abs/2311.03099), 与 [TIES](https://arxiv.org/abs/2311.03099).
|
||||
* **Multimodal models(多模态模型)**:这类模型(如 [CLIP](https://openai.com/research/clip), [Stable Diffusion](https://stability.ai/stable-image), 或 [LLaVA](https://llava-vl.github.io/)))以统一的嵌入空间处理多种类型的输入(文本、图像、音频等),从而解锁文生图等强大应用。
|
||||
* **Interpretability(可解释性)**:稀疏自编码器(Sparse Autoencoders,SAE)等机制可解释性技术取得了显著进展,为理解 LLM 内部运作提供洞见。这也已应用于 abliteration 等技术,使你无需训练即可修改模型行为。
|
||||
* **Test-time compute(测试时算力)**:通过强化学习(RL)技术训练的推理模型,可在测试时扩展算力预算以进一步提升。这可能涉及多次调用、MCTS,或 Process Reward Model(PRM)等专用模型。带精确评分的迭代步骤能显著提升复杂推理任务的性能。
|
||||
|
||||
📚 **References**:
|
||||
* [Merge LLMs with mergekit](https://mlabonne.github.io/blog/posts/2024-01-08_Merge_LLMs_with_mergekit.html) by Maxime Labonne: Tutorial about model merging using mergekit.
|
||||
* [Smol Vision](https://github.com/merveenoyan/smol-vision) by Merve Noyan: Collection of notebooks and scripts dedicated to small multimodal models.
|
||||
* [Large Multimodal Models](https://huyenchip.com/2023/10/10/multimodal.html) by Chip Huyen: Overview of multimodal systems and the recent history of this field.
|
||||
* [Unsensor any LLM with abliteration](https://huggingface.co/blog/mlabonne/abliteration) by Maxime Labonne: Direct application of interpretability techniques to modify the style of a model.
|
||||
* [Intuitive Explanation of SAEs](https://adamkarvonen.github.io/machine_learning/2024/06/11/sae-intuitions.html) by Adam Karvonen: Article about how SAEs work and why they make sense for interpretability.
|
||||
* [Scaling test-time compute](https://huggingface.co/spaces/HuggingFaceH4/blogpost-scaling-test-time-compute) by Beeching et al.: Tutorial and experiments to outperform Llama 3.1 70B on MATH-500 with a 3B model.
|
||||
📚 **参考资料**:
|
||||
* [Merge LLMs with mergekit](https://mlabonne.github.io/blog/posts/2024-01-08_Merge_LLMs_with_mergekit.html) by Maxime Labonne:使用 mergekit 进行模型合并的教程。
|
||||
* [Smol Vision](https://github.com/merveenoyan/smol-vision) by Merve Noyan:面向小型多模态模型的笔记本与脚本合集。
|
||||
* [Large Multimodal Models](https://huyenchip.com/2023/10/10/multimodal.html) by Chip Huyen:多模态系统概述及该领域近期发展史。
|
||||
* [Unsensor any LLM with abliteration](https://huggingface.co/blog/mlabonne/abliteration) by Maxime Labonne:将可解释性技术直接应用于修改模型风格。
|
||||
* [Intuitive Explanation of SAEs](https://adamkarvonen.github.io/machine_learning/2024/06/11/sae-intuitions.html) by Adam Karvonen:介绍 SAE 的工作原理及其对可解释性的意义。
|
||||
* [Scaling test-time compute](https://huggingface.co/spaces/HuggingFaceH4/blogpost-scaling-test-time-compute) by Beeching et al.:教程与实验,展示如何用 3B 模型在 MATH-500 上超越 Llama 3.1 70B。
|
||||
|
||||
## 👷 The LLM Engineer
|
||||
## 👷 LLM 工程师
|
||||
|
||||
This section of the course focuses on learning how to build LLM-powered applications that can be used in production, with a focus on augmenting models and deploying them.
|
||||
本课程章节聚焦学习如何构建可用于生产的 LLM 驱动应用,重点在于增强模型与部署。
|
||||
|
||||

|
||||
|
||||
### 1. Running LLMs
|
||||
### 1. 运行 LLM
|
||||
|
||||
Running LLMs can be difficult due to high hardware requirements. Depending on your use case, you might want to simply consume a model through an API (like GPT-4) or run it locally. In any case, additional prompting and guidance techniques can improve and constrain the output for your applications.
|
||||
由于硬件要求较高,运行 LLM 可能较为困难。根据你的用例,你可能只想通过 API(如 GPT-4)消费模型,或在本地运行。无论哪种方式,额外的提示(prompting)与引导技术都能改进并约束应用输出。
|
||||
|
||||
* **LLM APIs**: APIs are a convenient way to deploy LLMs. This space is divided between private LLMs ([OpenAI](https://platform.openai.com/), [Google](https://cloud.google.com/vertex-ai/docs/generative-ai/learn/overview), [Anthropic](https://docs.anthropic.com/claude/reference/getting-started-with-the-api), etc.) and open-source LLMs ([OpenRouter](https://openrouter.ai/), [Hugging Face](https://huggingface.co/inference-api), [Together AI](https://www.together.ai/), etc.).
|
||||
* **Open-source LLMs**: The [Hugging Face Hub](https://huggingface.co/models) is a great place to find LLMs. You can directly run some of them in [Hugging Face Spaces](https://huggingface.co/spaces), or download and run them locally in apps like [LM Studio](https://lmstudio.ai/) or through the CLI with [llama.cpp](https://github.com/ggerganov/llama.cpp) or [ollama](https://ollama.ai/).
|
||||
* **Prompt engineering**: Common techniques include zero-shot prompting, few-shot prompting, chain of thought, and ReAct. They work better with bigger models, but can be adapted to smaller ones.
|
||||
* **Structuring outputs**: Many tasks require a structured output, like a strict template or a JSON format. Libraries like [Outlines](https://github.com/outlines-dev/outlines) can be used to guide the generation and respect a given structure. Some APIs also support structured output generation natively using JSON schemas.
|
||||
* **LLM APIs**:API 是部署 LLM 的便捷方式。该领域分为私有 LLM([OpenAI](https://platform.openai.com/), [Google](https://cloud.google.com/vertex-ai/docs/generative-ai/learn/overview), [Anthropic](https://docs.anthropic.com/claude/reference/getting-started-with-the-api), 等)与开源 LLM([OpenRouter](https://openrouter.ai/), [Hugging Face](https://huggingface.co/inference-api), [Together AI](https://www.together.ai/), 等)。
|
||||
* **Open-source LLMs**: [Hugging Face Hub](https://huggingface.co/models) 是寻找 LLM 的绝佳去处。你可以直接在 [Hugging Face Spaces](https://huggingface.co/spaces), 中运行其中一些,或下载后在 [LM Studio](https://lmstudio.ai/) 等应用中本地运行,也可通过 CLI 使用 [llama.cpp](https://github.com/ggerganov/llama.cpp) 或 [ollama](https://ollama.ai/).
|
||||
* **Prompt engineering(提示工程)**:常用技术包括 zero-shot prompting、few-shot prompting、chain of thought 与 ReAct。它们在更大模型上效果更好,但也可适配较小模型。
|
||||
* **Structuring outputs(结构化输出)**:许多任务需要结构化输出,如严格模板或 JSON 格式。可使用 [Outlines](https://github.com/outlines-dev/outlines) 等库引导生成并遵循给定结构。部分 API 也支持通过 JSON schema 原生生成结构化输出。
|
||||
|
||||
📚 **References**:
|
||||
* [Run an LLM locally with LM Studio](https://www.kdnuggets.com/run-an-llm-locally-with-lm-studio) by Nisha Arya: Short guide on how to use LM Studio.
|
||||
* [Prompt engineering guide](https://www.promptingguide.ai/) by DAIR.AI: Exhaustive list of prompt techniques with examples
|
||||
* [Outlines - Quickstart](https://dottxt-ai.github.io/outlines/latest/quickstart/): List of guided generation techniques enabled by Outlines.
|
||||
* [LMQL - Overview](https://lmql.ai/docs/language/overview.html): Introduction to the LMQL language.
|
||||
📚 **参考资料**:
|
||||
* [Run an LLM locally with LM Studio](https://www.kdnuggets.com/run-an-llm-locally-with-lm-studio) by Nisha Arya:使用 LM Studio 的简短指南。
|
||||
* [Prompt engineering guide](https://www.promptingguide.ai/) by DAIR.AI:附带示例的提示技术详尽列表
|
||||
* [Outlines - Quickstart](https://dottxt-ai.github.io/outlines/latest/quickstart/): Outlines 支持的引导式生成技术列表。
|
||||
* [LMQL - Overview](https://lmql.ai/docs/language/overview.html): LMQL 语言介绍。
|
||||
|
||||
---
|
||||
### 2. Building a Vector Storage
|
||||
### 2. 构建向量存储
|
||||
|
||||
Creating a vector storage is the first step to building a Retrieval Augmented Generation (RAG) pipeline. Documents are loaded, split, and relevant chunks are used to produce vector representations (embeddings) that are stored for future use during inference.
|
||||
创建向量存储是构建检索增强生成(Retrieval Augmented Generation,RAG)流水线的第一步。文档被加载、切分,相关块用于生成向量表示(embeddings)并存储,以供推理时后续使用。
|
||||
|
||||
* **Ingesting documents**: Document loaders are convenient wrappers that can handle many formats: PDF, JSON, HTML, Markdown, etc. They can also directly retrieve data from some databases and APIs (GitHub, Reddit, Google Drive, etc.).
|
||||
* **Splitting documents**: Text splitters break down documents into smaller, semantically meaningful chunks. Instead of splitting text after *n* characters, it's often better to split by header or recursively, with some additional metadata.
|
||||
* **Embedding models**: Embedding models convert text into vector representations. Picking task-specific models significantly improves performance for semantic search and RAG.
|
||||
* **Vector databases**: Vector databases (like [Chroma](https://www.trychroma.com/), [Pinecone](https://www.pinecone.io/), [Milvus](https://milvus.io/), [FAISS](https://faiss.ai/), [Annoy](https://github.com/spotify/annoy), etc.) are designed to store embedding vectors. They enable efficient retrieval of data that is 'most similar' to a query based on vector similarity.
|
||||
* **Ingesting documents(文档摄入)**:文档加载器(document loaders)是便捷封装,可处理多种格式:PDF、JSON、HTML、Markdown 等。它们也可直接从部分数据库与 API(GitHub、Reddit、Google Drive 等)获取数据。
|
||||
* **Splitting documents(文档切分)**:文本切分器(text splitters)将文档拆分为更小、语义上有意义的块。相比在 *n* 个字符后切分,通常更宜按标题或递归切分,并附加一些元数据。
|
||||
* **Embedding models(嵌入模型)**:嵌入模型将文本转换为向量表示。选择面向特定任务的模型能显著提升语义搜索与 RAG 的性能。
|
||||
* **Vector databases(向量数据库)**:向量数据库(如 [Chroma](https://www.trychroma.com/), [Pinecone](https://www.pinecone.io/), [Milvus](https://milvus.io/), [FAISS](https://faiss.ai/), [Annoy](https://github.com/spotify/annoy), 等)专为存储嵌入向量而设计。它们基于向量相似度,高效检索与查询“最相似”的数据。
|
||||
|
||||
📚 **References**:
|
||||
* [LangChain - Text splitters](https://python.langchain.com/docs/how_to/#text-splitters): List of different text splitters implemented in LangChain.
|
||||
* [Sentence Transformers library](https://www.sbert.net/): Popular library for embedding models.
|
||||
* [MTEB Leaderboard](https://huggingface.co/spaces/mteb/leaderboard): Leaderboard for embedding models.
|
||||
* [The Top 7 Vector Databases](https://www.datacamp.com/blog/the-top-5-vector-databases) by Moez Ali: A comparison of the best and most popular vector databases.
|
||||
📚 **参考资料**:
|
||||
* [LangChain - Text splitters](https://python.langchain.com/docs/how_to/#text-splitters): LangChain 中实现的各类文本切分器列表。
|
||||
* [Sentence Transformers library](https://www.sbert.net/): 流行的嵌入模型库。
|
||||
* [MTEB Leaderboard](https://huggingface.co/spaces/mteb/leaderboard): 嵌入模型排行榜。
|
||||
* [The Top 7 Vector Databases](https://www.datacamp.com/blog/the-top-5-vector-databases) by Moez Ali:对最佳且最流行的向量数据库的比较。
|
||||
|
||||
---
|
||||
### 3. Retrieval Augmented Generation
|
||||
### 3. 检索增强生成(Retrieval Augmented Generation)
|
||||
|
||||
With RAG, LLMs retrieve contextual documents from a database to improve the accuracy of their answers. RAG is a popular way of augmenting the model's knowledge without any fine-tuning.
|
||||
借助 RAG,LLM 会从数据库检索上下文文档,以提高回答的准确性。RAG 是一种流行的方式,可在不进行任何微调的情况下扩充模型的知识。
|
||||
|
||||
* **Orchestrators**: Orchestrators like [LangChain](https://python.langchain.com/docs/get_started/introduction) and [LlamaIndex](https://docs.llamaindex.ai/en/stable/) are popular frameworks to connect your LLMs with tools and databases. The Model Context Protocol (MCP) introduces a new standard to pass data and context to models across providers.
|
||||
* **Retrievers**: Query rewriters and generative retrievers like CoRAG and HyDE enhance search by transforming user queries. Multi-vector and hybrid retrieval methods combine embeddings with keyword signals to improve recall and precision.
|
||||
* **Memory**: To remember previous instructions and answers, LLMs and chatbots like ChatGPT add this history to their context window. This buffer can be improved with summarization (e.g., using a smaller LLM), a vector store + RAG, etc.
|
||||
* **Evaluation**: We need to evaluate both the document retrieval (context precision and recall) and the generation stages (faithfulness and answer relevancy). It can be simplified with tools [Ragas](https://github.com/explodinggradients/ragas/tree/main) and [DeepEval](https://github.com/confident-ai/deepeval) (assessing quality).
|
||||
* **编排器(Orchestrators)**:像 [LangChain](https://python.langchain.com/docs/get_started/introduction) 和 [LlamaIndex](https://docs.llamaindex.ai/en/stable/) 这样的编排器,是连接 LLM 与工具和数据库的流行框架。模型上下文协议(Model Context Protocol, MCP)引入了一项新标准,用于跨提供商向模型传递数据和上下文。
|
||||
* **检索器(Retrievers)**:像 CoRAG 和 HyDE 这样的查询重写器与生成式检索器,通过转换用户查询来增强搜索。多向量与混合检索方法将嵌入与关键词信号相结合,以提高召回率和精确率。
|
||||
* **记忆(Memory)**:为记住先前的指令和回答,LLM 以及像 ChatGPT 这样的聊天机器人会将这段历史添加到其上下文窗口中。可通过摘要(例如使用较小的 LLM)、向量存储 + RAG 等方式改进该缓冲区。
|
||||
* **评估(Evaluation)**:我们需要同时评估文档检索(上下文精确率与召回率)和生成阶段(忠实度与答案相关性)。可使用 [Ragas](https://github.com/explodinggradients/ragas/tree/main) 和 [DeepEval](https://github.com/confident-ai/deepeval) 等工具简化评估(用于质量评估)。
|
||||
|
||||
📚 **References**:
|
||||
* [Llamaindex - High-level concepts](https://docs.llamaindex.ai/en/stable/getting_started/concepts.html): Main concepts to know when building RAG pipelines.
|
||||
* [Model Context Protocol](https://modelcontextprotocol.io/introduction): Introduction to MCP with motivate, architecture, and quick starts.
|
||||
* [Pinecone - Retrieval Augmentation](https://www.pinecone.io/learn/series/langchain/langchain-retrieval-augmentation/): Overview of the retrieval augmentation process.
|
||||
* [LangChain - Q&A with RAG](https://python.langchain.com/docs/tutorials/rag/): Step-by-step tutorial to build a typical RAG pipeline.
|
||||
* [LangChain - Memory types](https://python.langchain.com/docs/how_to/chatbots_memory/): List of different types of memories with relevant usage.
|
||||
* [RAG pipeline - Metrics](https://docs.ragas.io/en/stable/concepts/metrics/index.html): Overview of the main metrics used to evaluate RAG pipelines.
|
||||
📚 **参考资料**:
|
||||
* [Llamaindex - High-level concepts](https://docs.llamaindex.ai/en/stable/getting_started/concepts.html): 构建 RAG 流水线时需要了解的主要概念。
|
||||
* [Model Context Protocol](https://modelcontextprotocol.io/introduction): MCP 简介,包含动机、架构与快速入门。
|
||||
* [Pinecone - Retrieval Augmentation](https://www.pinecone.io/learn/series/langchain/langchain-retrieval-augmentation/): 检索增强流程概述。
|
||||
* [LangChain - Q&A with RAG](https://python.langchain.com/docs/tutorials/rag/): 构建典型 RAG 流水线的分步教程。
|
||||
* [LangChain - Memory types](https://python.langchain.com/docs/how_to/chatbots_memory/): 不同类型的记忆列表及其相关用法。
|
||||
* [RAG pipeline - Metrics](https://docs.ragas.io/en/stable/concepts/metrics/index.html): 用于评估 RAG 流水线的主要指标概述。
|
||||
|
||||
---
|
||||
### 4. Advanced RAG
|
||||
### 4. 高级 RAG(Advanced RAG)
|
||||
|
||||
Real-life applications can require complex pipelines, including SQL or graph databases, as well as automatically selecting relevant tools and APIs. These advanced techniques can improve a baseline solution and provide additional features.
|
||||
实际应用可能需要复杂的流水线,包括 SQL 或图数据库,以及自动选择相关工具和 API。这些高级技术可以改进基线方案,并提供额外功能。
|
||||
|
||||
* **Query construction**: Structured data stored in traditional databases requires a specific query language like SQL, Cypher, metadata, etc. We can directly translate the user instruction into a query to access the data with query construction.
|
||||
* **Tools**: Agents augment LLMs by automatically selecting the most relevant tools to provide an answer. These tools can be as simple as using Google or Wikipedia, or more complex, like a Python interpreter or Jira.
|
||||
* **Post-processing**: Final step that processes the inputs that are fed to the LLM. It enhances the relevance and diversity of documents retrieved with re-ranking, [RAG-fusion](https://github.com/Raudaschl/rag-fusion), and classification.
|
||||
* **Program LLMs**: Frameworks like [DSPy](https://github.com/stanfordnlp/dspy) allow you to optimize prompts and weights based on automated evaluations in a programmatic way.
|
||||
* **查询构建(Query construction)**:存储在传统数据库中的结构化数据需要特定的查询语言,如 SQL、Cypher、元数据等。我们可以通过查询构建,将用户指令直接翻译为查询以访问数据。
|
||||
* **工具(Tools)**:智能体(Agents)通过自动选择最相关的工具来提供答案,从而增强 LLM。这些工具可以很简单,例如使用 Google 或 Wikipedia;也可以更复杂,例如 Python 解释器或 Jira。
|
||||
* **后处理(Post-processing)**:对输入 LLM 的内容进行处理的最后一步。它通过重排序、[RAG-fusion](https://github.com/Raudaschl/rag-fusion), 和分类等方式,提升所检索文档的相关性与多样性。
|
||||
* **程序化 LLM(Program LLMs)**:像 [DSPy](https://github.com/stanfordnlp/dspy) 这样的框架,允许你以程序化方式,基于自动化评估来优化提示词和权重。
|
||||
|
||||
📚 **References**:
|
||||
* [LangChain - Query Construction](https://blog.langchain.dev/query-construction/): Blog post about different types of query construction.
|
||||
* [LangChain - SQL](https://python.langchain.com/docs/tutorials/sql_qa/): Tutorial on how to interact with SQL databases with LLMs, involving Text-to-SQL and an optional SQL agent.
|
||||
* [Pinecone - LLM agents](https://www.pinecone.io/learn/series/langchain/langchain-agents/): Introduction to agents and tools with different types.
|
||||
* [LLM Powered Autonomous Agents](https://lilianweng.github.io/posts/2023-06-23-agent/) by Lilian Weng: A more theoretical article about LLM agents.
|
||||
* [LangChain - OpenAI's RAG](https://blog.langchain.dev/applying-openai-rag/): Overview of the RAG strategies employed by OpenAI, including post-processing.
|
||||
* [DSPy in 8 Steps](https://dspy-docs.vercel.app/docs/building-blocks/solving_your_task): General-purpose guide to DSPy introducing modules, signatures, and optimizers.
|
||||
📚 **参考资料**:
|
||||
* [LangChain - Query Construction](https://blog.langchain.dev/query-construction/): 关于不同类型查询构建的博客文章。
|
||||
* [LangChain - SQL](https://python.langchain.com/docs/tutorials/sql_qa/): 关于如何使用 LLM 与 SQL 数据库交互的教程,涉及 Text-to-SQL 和可选的 SQL 智能体。
|
||||
* [Pinecone - LLM agents](https://www.pinecone.io/learn/series/langchain/langchain-agents/): 智能体与工具的介绍,涵盖不同类型。
|
||||
* [LLM Powered Autonomous Agents](https://lilianweng.github.io/posts/2023-06-23-agent/) by Lilian Weng:一篇关于 LLM 智能体的偏理论文章。
|
||||
* [LangChain - OpenAI's RAG](https://blog.langchain.dev/applying-openai-rag/): OpenAI 所采用的 RAG 策略概述,包括后处理。
|
||||
* [DSPy in 8 Steps](https://dspy-docs.vercel.app/docs/building-blocks/solving_your_task): DSPy 通用指南,介绍模块、签名与优化器。
|
||||
|
||||
---
|
||||
### 5. Agents
|
||||
### 5. 智能体(Agents)
|
||||
|
||||
An LLM agent can autonomously perform tasks by taking actions based on reasoning about its environment, typically through the use of tools or functions to interact with external systems.
|
||||
LLM 智能体可以基于对环境进行推理而自主执行任务,通常通过使用工具或函数与外部系统交互。
|
||||
|
||||
* **Agent fundamentals**: Agents operate using thoughts (internal reasoning to decide what to do next), action (executing tasks, often by interacting with external tools), and observation (analyzing feedback or results to refine the next step).
|
||||
* **Agent protocols**: [Model Context Protocol](https://modelcontextprotocol.io/) (MCP) is the industry standard for connecting agents to external tools and data sources with MCP servers and clients. More recently, [Agent2Agent](https://a2a-protocol.org/) (A2A) tries to standardize a common language for agent interoperability.
|
||||
* **Vendor frameworks**: Each major cloud model provider has its own agentic framework with [OpenAI SDK](https://openai.github.io/openai-agents-python/), [Google ADK](https://google.github.io/adk-docs/), and [Claude Agent SDK](https://platform.claude.com/docs/en/agent-sdk/overview) if you're particularly tied to one vendor.
|
||||
* **Other frameworks**: Agent development can be streamlined using different frameworks like [LangGraph](https://www.langchain.com/langgraph) (design and visualization of workflows) [LlamaIndex](https://docs.llamaindex.ai/en/stable/use_cases/agents/) (data-augmented agents with RAG), or custom solutions. More experimental frameworks include collaboration between different agents, such as [CrewAI](https://docs.crewai.com/introduction) (role-based team workflows) and [AutoGen](https://github.com/microsoft/autogen) (conversation-driven multi-agent systems).
|
||||
* **智能体基础(Agent fundamentals)**:智能体通过思考(内部推理以决定下一步做什么)、行动(执行任务,通常通过与外部工具交互)和观察(分析反馈或结果以优化下一步)来运作。
|
||||
* **智能体协议(Agent protocols)**:[Model Context Protocol](https://modelcontextprotocol.io/)(MCP)是通过 MCP 服务器与客户端将智能体连接到外部工具和数据源的行业标准。近来,[Agent2Agent](https://a2a-protocol.org/)(A2A)尝试为智能体互操作性标准化一种通用语言。
|
||||
* **厂商框架(Vendor frameworks)**:每个主要云模型提供商都有自己的智能体框架,包括 [OpenAI SDK](https://openai.github.io/openai-agents-python/),、[Google ADK](https://google.github.io/adk-docs/), 和 [Claude Agent SDK](https://platform.claude.com/docs/en/agent-sdk/overview);如果你特别绑定某一家厂商,可以选择对应方案。
|
||||
* **其他框架(Other frameworks)**:可使用不同框架简化智能体开发,例如 [LangGraph](https://www.langchain.com/langgraph)(工作流设计与可视化)、[LlamaIndex](https://docs.llamaindex.ai/en/stable/use_cases/agents/)(结合 RAG 的数据增强型智能体),或自定义方案。更具实验性的框架包括不同智能体之间的协作,例如 [CrewAI](https://docs.crewai.com/introduction)(基于角色的团队工作流)和 [AutoGen](https://github.com/microsoft/autogen)(对话驱动的多智能体系统)。
|
||||
|
||||
📚 **References**:
|
||||
* [Agents Course](https://huggingface.co/learn/agents-course/unit0/introduction): Popular course about AI agents made by Hugging Face.
|
||||
* [LangGraph](https://langchain-ai.github.io/langgraph/concepts/why-langgraph/): Overview of how to build AI agents with LangGraph.
|
||||
* [LlamaIndex Agents](https://docs.llamaindex.ai/en/stable/use_cases/agents/): Uses cases and resources to build agents with LlamaIndex.
|
||||
📚 **参考资料**:
|
||||
* [Agents Course](https://huggingface.co/learn/agents-course/unit0/introduction): Hugging Face 制作的关于 AI 智能体的热门课程。
|
||||
* [LangGraph](https://langchain-ai.github.io/langgraph/concepts/why-langgraph/): 如何使用 LangGraph 构建 AI 智能体的概述。
|
||||
* [LlamaIndex Agents](https://docs.llamaindex.ai/en/stable/use_cases/agents/): 使用 LlamaIndex 构建智能体的用例与资源。
|
||||
|
||||
---
|
||||
### 6. Inference optimization
|
||||
### 6. 推理优化(Inference optimization)
|
||||
|
||||
Text generation is a costly process that requires expensive hardware. In addition to quantization, various techniques have been proposed to maximize throughput and reduce inference costs.
|
||||
文本生成是一个成本高昂的过程,需要昂贵的硬件。除量化之外,还提出了多种技术,以最大化吞吐量并降低推理成本。
|
||||
|
||||
* **Flash Attention**: Optimization of the attention mechanism to transform its complexity from quadratic to linear, speeding up both training and inference.
|
||||
* **Key-value cache**: Understand the key-value cache and the improvements introduced in [Multi-Query Attention](https://arxiv.org/abs/1911.02150) (MQA) and [Grouped-Query Attention](https://arxiv.org/abs/2305.13245) (GQA).
|
||||
* **Speculative decoding**: Use a small model to produce drafts that are then reviewed by a larger model to speed up text generation. EAGLE-3 is a particularly popular solution.
|
||||
* **Flash Attention**:对注意力机制进行优化,将其复杂度从二次降为线性,从而加速训练与推理。
|
||||
* **键值缓存(Key-value cache)**:理解键值缓存,以及 [Multi-Query Attention](https://arxiv.org/abs/1911.02150)(MQA)和 [Grouped-Query Attention](https://arxiv.org/abs/2305.13245)(GQA)带来的改进。
|
||||
* **推测解码(Speculative decoding)**:使用小模型生成草稿,再由大模型审核,以加速文本生成。EAGLE-3 是一种特别流行的方案。
|
||||
|
||||
📚 **References**:
|
||||
* [GPU Inference](https://huggingface.co/docs/transformers/main/en/perf_infer_gpu_one) by Hugging Face: Explain how to optimize inference on GPUs.
|
||||
* [LLM Inference](https://www.databricks.com/blog/llm-inference-performance-engineering-best-practices) by Databricks: Best practices for how to optimize LLM inference in production.
|
||||
* [Optimizing LLMs for Speed and Memory](https://huggingface.co/docs/transformers/main/en/llm_tutorial_optimization) by Hugging Face: Explain three main techniques to optimize speed and memory, namely quantization, Flash Attention, and architectural innovations.
|
||||
* [Assisted Generation](https://huggingface.co/blog/assisted-generation) by Hugging Face: HF's version of speculative decoding. It's an interesting blog post about how it works with code to implement it.
|
||||
* [EAGLE-3 paper](https://arxiv.org/abs/2503.01840?utm_source=chatgpt.com): Introduces EAGLE-3 and reports speedups up to 6.5×.
|
||||
* [Speculators](https://github.com/vllm-project/speculators): Library made by vLLM for building, evaluating, and storing speculative decoding algorithms (e.g., EAGLE-3) for LLM inference.
|
||||
📚 **参考资料**:
|
||||
* [GPU Inference](https://huggingface.co/docs/transformers/main/en/perf_infer_gpu_one) by Hugging Face:讲解如何在 GPU 上优化推理。
|
||||
* [LLM Inference](https://www.databricks.com/blog/llm-inference-performance-engineering-best-practices) by Databricks:在生产环境中优化 LLM 推理的最佳实践。
|
||||
* [Optimizing LLMs for Speed and Memory](https://huggingface.co/docs/transformers/main/en/llm_tutorial_optimization) by Hugging Face:讲解优化速度与内存的三项主要技术,即量化、Flash Attention 和架构创新。
|
||||
* [Assisted Generation](https://huggingface.co/blog/assisted-generation) by Hugging Face:Hugging Face 版本的推测解码。这是一篇有趣的文章,讲解其工作原理,并附有实现代码。
|
||||
* [EAGLE-3 paper](https://arxiv.org/abs/2503.01840?utm_source=chatgpt.com): 介绍 EAGLE-3,并报告最高可达 6.5× 的加速。
|
||||
* [Speculators](https://github.com/vllm-project/speculators): 由 vLLM 开发的库,用于为 LLM 推理构建、评估和存储推测解码算法(例如 EAGLE-3)。
|
||||
|
||||
---
|
||||
### 7. Deploying LLMs
|
||||
### 7. 部署 LLM
|
||||
|
||||
Deploying LLMs at scale is an engineering feat that can require multiple clusters of GPUs. In other scenarios, demos and local apps can be achieved with much lower complexity.
|
||||
大规模部署 LLM 是一项工程挑战,可能需要多个 GPU 集群。在其他场景下,演示和本地应用可以用低得多的复杂度实现。
|
||||
|
||||
* **Local deployment**: Privacy is an important advantage that open-source LLMs have over private ones. Local LLM servers ([LM Studio](https://lmstudio.ai/), [Ollama](https://ollama.ai/), [oobabooga](https://github.com/oobabooga/text-generation-webui), [kobold.cpp](https://github.com/LostRuins/koboldcpp), etc.) capitalize on this advantage to power local apps.
|
||||
* **Demo deployment**: Frameworks like [Gradio](https://www.gradio.app/) and [Streamlit](https://docs.streamlit.io/) are helpful to prototype applications and share demos. You can also easily host them online, for example, using [Hugging Face Spaces](https://huggingface.co/spaces).
|
||||
* **Server deployment**: Deploying LLMs at scale requires cloud (see also [SkyPilot](https://skypilot.readthedocs.io/en/latest/)) or on-prem infrastructure and often leverages optimized text generation frameworks like [TGI](https://github.com/huggingface/text-generation-inference), [vLLM](https://github.com/vllm-project/vllm/tree/main), etc.
|
||||
* **Edge deployment**: In constrained environments, high-performance frameworks like [MLC LLM](https://github.com/mlc-ai/mlc-llm) and [mnn-llm](https://github.com/wangzhaode/mnn-llm/blob/master/README_en.md) can deploy LLM in web browsers, Android, and iOS.
|
||||
* **本地部署**:隐私是开源 LLM 相对闭源模型的一项重要优势。本地 LLM 服务器([LM Studio](https://lmstudio.ai/), [Ollama](https://ollama.ai/), [oobabooga](https://github.com/oobabooga/text-generation-webui), [kobold.cpp](https://github.com/LostRuins/koboldcpp), 等)利用这一优势为本地应用提供动力。
|
||||
* **演示部署**:像 [Gradio](https://www.gradio.app/) 和 [Streamlit](https://docs.streamlit.io/) 这样的框架有助于原型化应用并分享演示。你也可以轻松将它们托管到线上,例如使用 [Hugging Face Spaces](https://huggingface.co/spaces).
|
||||
* **服务器部署**:大规模部署 LLM 需要云端(另见 [SkyPilot](https://skypilot.readthedocs.io/en/latest/)) 或本地(on-prem)基础设施,并常借助 [TGI](https://github.com/huggingface/text-generation-inference), [vLLM](https://github.com/vllm-project/vllm/tree/main), 等优化的文本生成框架。
|
||||
* **边缘部署**:在资源受限的环境中,[MLC LLM](https://github.com/mlc-ai/mlc-llm) 和 [mnn-llm](https://github.com/wangzhaode/mnn-llm/blob/master/README_en.md) 等高性能框架可以在 Web 浏览器、Android 和 iOS 上部署 LLM。
|
||||
|
||||
📚 **References**:
|
||||
* [Streamlit - Build a basic LLM app](https://docs.streamlit.io/knowledge-base/tutorials/build-conversational-apps): Tutorial to make a basic ChatGPT-like app using Streamlit.
|
||||
* [HF LLM Inference Container](https://huggingface.co/blog/sagemaker-huggingface-llm): Deploy LLMs on Amazon SageMaker using Hugging Face's inference container.
|
||||
* [Philschmid blog](https://www.philschmid.de/) by Philipp Schmid: Collection of high-quality articles about LLM deployment using Amazon SageMaker.
|
||||
* [Optimizing latence](https://hamel.dev/notes/llm/inference/03_inference.html) by Hamel Husain: Comparison of TGI, vLLM, CTranslate2, and mlc in terms of throughput and latency.
|
||||
📚 **参考资料**:
|
||||
* [Streamlit - Build a basic LLM app](https://docs.streamlit.io/knowledge-base/tutorials/build-conversational-apps): 使用 Streamlit 制作基础 ChatGPT 风格应用的教程。
|
||||
* [HF LLM Inference Container](https://huggingface.co/blog/sagemaker-huggingface-llm): 使用 Hugging Face 推理容器在 Amazon SageMaker 上部署 LLM。
|
||||
* [Philschmid blog](https://www.philschmid.de/) Philipp Schmid 撰写:关于使用 Amazon SageMaker 部署 LLM 的高质量文章合集。
|
||||
* [Optimizing latence](https://hamel.dev/notes/llm/inference/03_inference.html) Hamel Husain 撰写:从吞吐量和延迟角度对比 TGI、vLLM、CTranslate2 和 mlc。
|
||||
|
||||
---
|
||||
### 8. Securing LLMs
|
||||
### 8. 保护 LLM 安全
|
||||
|
||||
In addition to traditional security problems associated with software, LLMs have unique weaknesses due to the way they are trained and prompted.
|
||||
除软件常见的传统安全问题外,LLM 因其训练和提示(prompting)方式而存在独特弱点。
|
||||
|
||||
* **Prompt hacking**: Different techniques related to prompt engineering, including prompt injection (additional instruction to hijack the model's answer), data/prompt leaking (retrieve its original data/prompt), and jailbreaking (craft prompts to bypass safety features).
|
||||
* **Backdoors**: Attack vectors can target the training data itself, by poisoning the training data (e.g., with false information) or creating backdoors (secret triggers to change the model's behavior during inference).
|
||||
* **Defensive measures**: The best way to protect your LLM applications is to test them against these vulnerabilities (e.g., using red teaming and checks like [garak](https://github.com/leondz/garak/)) and observe them in production (with a framework like [langfuse](https://github.com/langfuse/langfuse)).
|
||||
* **提示攻击(Prompt hacking)**:与提示工程相关的多种技术,包括提示注入(prompt injection,通过额外指令劫持模型回答)、数据/提示泄露(data/prompt leaking,检索其原始数据/提示)和越狱(jailbreaking,精心构造提示以绕过安全机制)。
|
||||
* **后门(Backdoors)**:攻击向量可直接针对训练数据,通过污染训练数据(例如植入虚假信息)或创建后门(秘密触发器,在推理时改变模型行为)。
|
||||
* **防御措施**:保护 LLM 应用的最佳方式是针对这些漏洞进行测试(例如使用红队演练以及 [garak](https://github.com/leondz/garak/)) 等检查工具),并在生产环境中进行观测(可借助 [langfuse](https://github.com/langfuse/langfuse)). 等框架)。
|
||||
|
||||
📚 **References**:
|
||||
* [OWASP LLM Top 10](https://owasp.org/www-project-top-10-for-large-language-model-applications/) by HEGO Wiki: List of the 10 most critical vulnerabilities seen in LLM applications.
|
||||
* [Prompt Injection Primer](https://github.com/jthack/PIPE) by Joseph Thacker: Short guide dedicated to prompt injection for engineers.
|
||||
* [LLM Security](https://llmsecurity.net/) by [@llm_sec](https://twitter.com/llm_sec): Extensive list of resources related to LLM security.
|
||||
* [Red teaming LLMs](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/red-teaming) by Microsoft: Guide on how to perform red teaming with LLMs.
|
||||
📚 **参考资料**:
|
||||
* [OWASP LLM Top 10](https://owasp.org/www-project-top-10-for-large-language-model-applications/) HEGO Wiki:LLM 应用中最常见的 10 项关键漏洞清单。
|
||||
* [Prompt Injection Primer](https://github.com/jthack/PIPE) Joseph Thacker 撰写:面向工程师的提示注入简明指南。
|
||||
* [LLM Security](https://llmsecurity.net/) [@llm_sec](https://twitter.com/llm_sec): 撰写:与 LLM 安全相关的广泛资源列表。
|
||||
* [Red teaming LLMs](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/red-teaming) Microsoft 撰写:如何使用 LLM 进行红队演练的指南。
|
||||
|
||||
---
|
||||
## Acknowledgements
|
||||
## 致谢
|
||||
|
||||
This roadmap was inspired by the excellent [DevOps Roadmap](https://github.com/milanm/DevOps-Roadmap) from Milan Milanović and Romano Roth.
|
||||
本路线图受到 Milan Milanović 和 Romano Roth 优秀的 [DevOps Roadmap](https://github.com/milanm/DevOps-Roadmap) 启发。
|
||||
|
||||
Special thanks to:
|
||||
特别感谢:
|
||||
|
||||
* Thomas Thelen for motivating me to create a roadmap
|
||||
* André Frade for his input and review of the first draft
|
||||
* Dino Dunn for providing resources about LLM security
|
||||
* Magdalena Kuhn for improving the "human evaluation" part
|
||||
* Odoverdose for suggesting 3Blue1Brown's video about Transformers
|
||||
* Everyone who contributed to the educational references in this course :)
|
||||
* Thomas Thelen 激励我创建这份路线图
|
||||
* André Frade 对初稿提供意见并进行审阅
|
||||
* Dino Dunn 提供 LLM 安全相关资源
|
||||
* Magdalena Kuhn 改进「人工评估」部分
|
||||
* Odoverdose 推荐 3Blue1Brown 关于 Transformer 的视频
|
||||
* 为本课程教育参考资料做出贡献的所有人 :)
|
||||
|
||||
*Disclaimer: I am not affiliated with any sources listed here.*
|
||||
*免责声明:本人与本文列出的任何来源均无隶属关系。*
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user