吞吐量
延迟
投机采样:
美杜莎:
-
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
-
OpenLLM: https://github.com/bentoml/OpenLLM
吞吐量
延迟
投机采样:
美杜莎:
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
OpenLLM: https://github.com/bentoml/OpenLLM