Files
wehub-resource-sync 91f715342e
Lint / quick-checks (push) Has been cancelled
Lint / flake8-py3 (push) Has been cancelled
docs: make Chinese README the default
2026-07-13 10:36:26 +00:00

282 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!-- WEHUB_ZH_README -->
> [!NOTE]
> 本文档由 WeHub 基于上游 README 翻译整理,属于社区翻译,非官方中文文档。
> [English](./README.en.md) · [原始项目](https://github.com/FunAudioLLM/CosyVoice) · [上游 README](https://github.com/FunAudioLLM/CosyVoice/blob/HEAD/README.md)
> 原作者、版权与许可证归属以原始项目及本仓库 LICENSE 文件为准。
![SVG Banners](https://svg-banners.vercel.app/api?type=origin&text1=CosyVoice🤠&text2=Text-to-Speech%20💖%20Large%20Language%20Model&width=800&height=210)
## 👉🏻 CosyVoice 👈🏻
**Fun-CosyVoice 3.0**: [Demos](https://funaudiollm.github.io/cosyvoice3/); [Paper](https://arxiv.org/pdf/2505.17589); [Modelscope](https://www.modelscope.cn/models/FunAudioLLM/Fun-CosyVoice3-0.5B-2512); [Huggingface](https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512); [CV3-Eval](https://github.com/FunAudioLLM/CV3-Eval)
**CosyVoice 2.0**: [Demos](https://funaudiollm.github.io/cosyvoice2/); [Paper](https://arxiv.org/pdf/2412.10117); [Modelscope](https://www.modelscope.cn/models/iic/CosyVoice2-0.5B); [HuggingFace](https://huggingface.co/FunAudioLLM/CosyVoice2-0.5B)
**CosyVoice 1.0**: [Demos](https://fun-audio-llm.github.io); [Paper](https://funaudiollm.github.io/pdf/CosyVoice_v1.pdf); [Modelscope](https://www.modelscope.cn/models/iic/CosyVoice-300M); [HuggingFace](https://huggingface.co/FunAudioLLM/CosyVoice-300M)
## 亮点🔥
**Fun-CosyVoice 3.0** 是一款基于大语言模型(LLM)的先进文本转语音(TTS)系统,在内容一致性、说话人相似度和韵律自然度方面均优于前代产品(CosyVoice 2.0)。它面向真实场景下的零样本多语言语音合成而设计。
### 核心特性
- **语言覆盖**:支持 9 种常见语言(中文、英语、日语、韩语、德语、西班牙语、法语、意大利语、俄语),18+ 种中文方言/口音(广东话、闽南语、四川话、东北话、陕西(Shan3xi)、山西(Shan1xi)、上海话、天津话、山东话、宁夏话、甘肃话等),同时支持多语言/跨语言零样本语音克隆。
- **内容一致性与自然度**:在内容一致性、说话人相似度和韵律自然度方面达到业界领先水平(state-of-the-art)。
- **发音修复(Pronunciation Inpainting**:支持中文拼音和英文 CMU 音素的发音修复,可控性更强,适合生产环境使用。
- **文本规范化**:无需传统前端模块,即可朗读数字、特殊符号及多种文本格式。
- **双向流式(Bi-Streaming)**:同时支持文本输入流式与音频输出流式,在保持高质量音频输出的同时,延迟可低至 150ms。
- **指令支持**:支持语言、方言、情感、语速、音量等多种指令。
## 路线图
- [x] 2025/12
- [x] 发布 Fun-CosyVoice3-0.5B-2512 基础模型、RL 模型及其训练/推理脚本
- [x] 发布 Fun-CosyVoice3-0.5B Modelscope Gradio 空间
- [x] 2025/08
- [x] 感谢 NVIDIA Yuekai Zhang 的贡献,新增 Triton TRT-LLM 运行时支持及 CosyVoice2 GRPO 训练支持
- [x] 2025/07
- [x] 发布 Fun-CosyVoice 3.0 评测集
- [x] 2025/05
- [x] 新增 CosyVoice2-0.5B vLLM 支持
- [x] 2024/12
- [x] 发布 25hz CosyVoice2-0.5B
- [x] 2024/09
- [x] 25hz CosyVoice-300M 基础模型
- [x] 25hz CosyVoice-300M 语音转换功能
- [x] 2024/08
- [x] 面向 LLM 稳定性的 Repetition Aware SamplingRAS)推理
- [x] 流式推理模式支持,包括用于 RTF 优化的 kv cache 和 sdpa
- [x] 2024/07
- [x] Flow matching 训练支持
- [x] 在 ttsfrd 不可用时支持 WeTextProcessing
- [x] FastAPI 服务端与客户端
## 评测
| Model | Open-Source | Model Size | test-zh<br>CER (%) ↓ | test-zh<br>SS (%) ↑ | test-en<br>WER (%) ↓ | test-en<br>SS (%) ↑ | test-hard<br>CER (%) ↓ | test-hard<br>SS (%) ↑ |
| :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: |
| Human | - | - | 1.26 | 75.5 | 2.14 | 73.4 | - | - |
| Seed-TTS | ❌ | - | 1.12 | 79.6 | 2.25 | 76.2 | 7.59 | 77.6 |
| MiniMax-Speech | ❌ | - | 0.83 | 78.3 | 1.65 | 69.2 | - | - |
| F5-TTS | ✅ | 0.3B | 1.52 | 74.1 | 2.00 | 64.7 | 8.67 | 71.3 |
| Spark TTS | ✅ | 0.5B | 1.2 | 66.0 | 1.98 | 57.3 | - | - |
| CosyVoice2 | ✅ | 0.5B | 1.45 | 75.7 | 2.57 | 65.9 | 6.83 | 72.4 |
| FireRedTTS2 | ✅ | 1.5B | 1.14 | 73.2 | 1.95 | 66.5 | - | - |
| Index-TTS2 | ✅ | 1.5B | 1.03 | 76.5 | 2.23 | 70.6 | 7.12 | 75.5 |
| VibeVoice-1.5B | ✅ | 1.5B | 1.16 | 74.4 | 3.04 | 68.9 | - | - |
| VibeVoice-Realtime | ✅ | 0.5B | - | - | 2.05 | 63.3 | - | - |
| HiggsAudio-v2 | ✅ | 3B | 1.50 | 74.0 | 2.44 | 67.7 | - | - |
| VoxCPM | ✅ | 0.5B | 0.93 | 77.2 | 1.85 | 72.9 | 8.87 | 73.0 |
| GLM-TTS | ✅ | 1.5B | 1.03 | 76.1 | - | - | - | - |
| GLM-TTS RL | ✅ | 1.5B | 0.89 | 76.4 | - | - | - | - |
| Fun-CosyVoice3-0.5B-2512 | ✅ | 0.5B | 1.21 | 78.0 | 2.24 | 71.8 | 6.71 | 75.8 |
| Fun-CosyVoice3-0.5B-2512_RL | ✅ | 0.5B | 0.81 | 77.4 | 1.68 | 69.5 | 5.44 | 75.0 |
## 安装
### 克隆与安装
- 克隆仓库
``` sh
git clone --recursive https://github.com/FunAudioLLM/CosyVoice.git
# If you failed to clone the submodule due to network failures, please run the following command until success
cd CosyVoice
git submodule update --init --recursive
```
- 安装 Conda:请参阅 https://docs.conda.io/en/latest/miniconda.html
- 创建 Conda 环境:
``` sh
conda create -n cosyvoice -y python=3.10
conda activate cosyvoice
pip install -r requirements.txt -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com
# If you encounter sox compatibility issues
# ubuntu
sudo apt-get install sox libsox-dev
# centos
sudo yum install sox sox-devel
```
### 模型下载
我们强烈建议您下载我们预训练的 `Fun-CosyVoice3-0.5B` `CosyVoice2-0.5B` `CosyVoice-300M` `CosyVoice-300M-SFT` `CosyVoice-300M-Instruct` 模型和 `CosyVoice-ttsfrd` 资源。
``` python
# modelscope SDK model download
from modelscope import snapshot_download
snapshot_download('FunAudioLLM/Fun-CosyVoice3-0.5B-2512', local_dir='pretrained_models/Fun-CosyVoice3-0.5B')
snapshot_download('iic/CosyVoice2-0.5B', local_dir='pretrained_models/CosyVoice2-0.5B')
snapshot_download('iic/CosyVoice-300M', local_dir='pretrained_models/CosyVoice-300M')
snapshot_download('iic/CosyVoice-300M-SFT', local_dir='pretrained_models/CosyVoice-300M-SFT')
snapshot_download('iic/CosyVoice-300M-Instruct', local_dir='pretrained_models/CosyVoice-300M-Instruct')
snapshot_download('iic/CosyVoice-ttsfrd', local_dir='pretrained_models/CosyVoice-ttsfrd')
# for oversea users, huggingface SDK model download
from huggingface_hub import snapshot_download
snapshot_download('FunAudioLLM/Fun-CosyVoice3-0.5B-2512', local_dir='pretrained_models/Fun-CosyVoice3-0.5B')
snapshot_download('FunAudioLLM/CosyVoice2-0.5B', local_dir='pretrained_models/CosyVoice2-0.5B')
snapshot_download('FunAudioLLM/CosyVoice-300M', local_dir='pretrained_models/CosyVoice-300M')
snapshot_download('FunAudioLLM/CosyVoice-300M-SFT', local_dir='pretrained_models/CosyVoice-300M-SFT')
snapshot_download('FunAudioLLM/CosyVoice-300M-Instruct', local_dir='pretrained_models/CosyVoice-300M-Instruct')
snapshot_download('FunAudioLLM/CosyVoice-ttsfrd', local_dir='pretrained_models/CosyVoice-ttsfrd')
```
可选地,您可以解压 `ttsfrd` 资源并安装 `ttsfrd` 包,以获得更好的文本规范化性能。
请注意,此步骤并非必需。若您未安装 `ttsfrd` 包,我们将默认使用 wetext。
``` sh
cd pretrained_models/CosyVoice-ttsfrd/
unzip resource.zip -d .
pip install ttsfrd_dependency-0.1-py3-none-any.whl
pip install ttsfrd-0.4.2-cp310-cp310-linux_x86_64.whl
```
### 基础用法
我们强烈建议使用 `Fun-CosyVoice3-0.5B` 以获得更好的性能。
请参照 `example.py` 中的代码,了解各模型的详细用法。
```sh
python example.py
```
#### vLLM 用法
CosyVoice2/3 现已支持 **vLLM 0.11.x+V1 引擎)** 和 **vLLM 0.9.0(旧版)**。
低于 0.9.0 的 vLLM 版本不支持 CosyVoice 推理,介于两者之间的版本(例如 0.10.x)尚未经过测试。
请注意,`vllm` 有许多特定要求。若您的硬件不支持 vLLM 且旧环境已损坏,您可以新建一个环境。
``` sh
conda create -n cosyvoice_vllm --clone cosyvoice
conda activate cosyvoice_vllm
# for vllm==0.9.0
pip install vllm==v0.9.0 transformers==4.51.3 numpy==1.26.4 -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com
# for vllm>=0.11.0
pip install vllm==v0.11.0 transformers==4.57.1 numpy==1.26.4 -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com
python vllm_example.py
```
#### 启动 Web 演示
你可以使用我们的 Web 演示页面,快速熟悉 CosyVoice。
详情请参阅演示网站。
``` python
# change iic/CosyVoice-300M-SFT for sft inference, or iic/CosyVoice-300M-Instruct for instruct inference
python3 webui.py --port 50000 --model_dir pretrained_models/CosyVoice-300M
```
#### 进阶用法
面向进阶用户,我们在 `examples/libritts` 中提供了训练与推理脚本。
#### 构建以供部署
可选地,如需进行服务部署,
你可以按以下步骤操作。
``` sh
cd runtime/python
docker build -t cosyvoice:v1.0 .
# change iic/CosyVoice-300M to iic/CosyVoice-300M-Instruct if you want to use instruct inference
# for grpc usage
docker run -d --runtime=nvidia -p 50000:50000 cosyvoice:v1.0 /bin/bash -c "cd /opt/CosyVoice/CosyVoice/runtime/python/grpc && python3 server.py --port 50000 --max_conc 4 --model_dir iic/CosyVoice-300M && sleep infinity"
cd grpc && python3 client.py --port 50000 --mode <sft|zero_shot|cross_lingual|instruct>
# for fastapi usage
docker run -d --runtime=nvidia -p 50000:50000 cosyvoice:v1.0 /bin/bash -c "cd /opt/CosyVoice/CosyVoice/runtime/python/fastapi && python3 server.py --port 50000 --model_dir iic/CosyVoice-300M && sleep infinity"
cd fastapi && python3 client.py --port 50000 --mode <sft|zero_shot|cross_lingual|instruct>
```
#### 使用 Nvidia TensorRT-LLM 进行部署
使用 TensorRT-LLM 加速 cosyvoice2 llm,相比 Hugging Face transformers 实现可获得约 4 倍加速。
快速开始:
``` sh
cd runtime/triton_trtllm
docker compose up -d
```
更多详情,请参阅 [这里](https://github.com/FunAudioLLM/CosyVoice/tree/main/runtime/triton_trtllm)
## 讨论与交流
你可以在 [Github Issues](https://github.com/FunAudioLLM/CosyVoice/issues). 上直接讨论
你也可以扫描二维码,加入我们的官方钉钉群。
<img src="./asset/dingding.png" width="250px">
## 致谢
1. 我们从 [FunASR](https://github.com/modelscope/FunASR). 借鉴了大量代码
2. 我们从 [FunCodec](https://github.com/modelscope/FunCodec). 借鉴了大量代码
3. 我们从 [Matcha-TTS](https://github.com/shivammehta25/Matcha-TTS). 借鉴了大量代码
4. 我们从 [AcademiCodec](https://github.com/yangdongchao/AcademiCodec). 借鉴了大量代码
5. 我们从 [WeNet](https://github.com/wenet-e2e/wenet). 借鉴了大量代码
## 引用
``` bibtex
@article{du2024cosyvoice,
title={Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens},
author={Du, Zhihao and Chen, Qian and Zhang, Shiliang and Hu, Kai and Lu, Heng and Yang, Yexin and Hu, Hangrui and Zheng, Siqi and Gu, Yue and Ma, Ziyang and others},
journal={arXiv preprint arXiv:2407.05407},
year={2024}
}
@article{du2024cosyvoice,
title={Cosyvoice 2: Scalable streaming speech synthesis with large language models},
author={Du, Zhihao and Wang, Yuxuan and Chen, Qian and Shi, Xian and Lv, Xiang and Zhao, Tianyu and Gao, Zhifu and Yang, Yexin and Gao, Changfeng and Wang, Hui and others},
journal={arXiv preprint arXiv:2412.10117},
year={2024}
}
@article{du2025cosyvoice,
title={CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training},
author={Du, Zhihao and Gao, Changfeng and Wang, Yuxuan and Yu, Fan and Zhao, Tianyu and Wang, Hao and Lv, Xiang and Wang, Hui and Shi, Xian and An, Keyu and others},
journal={arXiv preprint arXiv:2505.17589},
year={2025}
}
@inproceedings{lyu2025build,
title={Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice},
author={Lyu, Xiang and Wang, Yuxuan and Zhao, Tianyu and Wang, Hao and Liu, Huadai and Du, Zhihao},
booktitle={ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
pages={1--2},
year={2025},
organization={IEEE}
}
```
## 生态系统
CosyVoice 属于 **FunAudioLLM** 系列——一套完整的语音 AI 工具集:
| 项目 | 说明 | Stars |
|---------|-------------|-------|
| [FunASR](https://github.com/modelscope/FunASR) | 工业级语音识别——支持 50+ 种语言、说话人分离(speaker diarization)、流式识别 | [![](https://img.shields.io/github/stars/modelscope/FunASR?style=social)](https://github.com/modelscope/FunASR) |
| [Fun-ASR-Nano](https://github.com/FunAudioLLM/Fun-ASR) | 端到端基于 LLM 的 ASR——31 种语言、热词(hotwords)、vLLM 流式推理 | [![](https://img.shields.io/github/stars/FunAudioLLM/Fun-ASR?style=social)](https://github.com/FunAudioLLM/Fun-ASR) |
| [SenseVoice](https://github.com/FunAudioLLM/SenseVoice) | 超快 ASR + 情感识别 + 音频事件检测 | [![](https://img.shields.io/github/stars/FunAudioLLM/SenseVoice?style=social)](https://github.com/FunAudioLLM/SenseVoice) |
| [FunClip](https://github.com/modelscope/FunClip) | 基于语音识别的 AI 视频剪辑 | [![](https://img.shields.io/github/stars/modelscope/FunClip?style=social)](https://github.com/modelscope/FunClip) |
## 免责声明
以上内容仅供学术研究使用,旨在展示技术能力。部分示例来源于网络。如任何内容侵犯您的权益,请联系我们申请删除。