<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Triton on Tony老师的博客</title><link>https://blog.tanteng.space/tags/triton/</link><description>Recent content in Triton on Tony老师的博客</description><generator>Hugo</generator><language>zh</language><lastBuildDate>Fri, 12 Apr 2024 10:00:00 +0800</lastBuildDate><atom:link href="https://blog.tanteng.space/tags/triton/index.xml" rel="self" type="application/rss+xml"/><item><title>LLM 推理服务架构：vLLM、TGI 与 Triton 的工程对比</title><link>https://blog.tanteng.space/2024/04/llm-serving-architecture/</link><pubDate>Fri, 12 Apr 2024 10:00:00 +0800</pubDate><guid>https://blog.tanteng.space/2024/04/llm-serving-architecture/</guid><description>&lt;p&gt;2023 年我们给一个客服知识库做 RAG 接入，自以为把 Llama 2 装进 FastAPI 就算&amp;quot;上线了&amp;quot;。结果一上线就翻车：并发刚到 32，GPU 利用率就只剩 18%，P99 延迟 8 秒。问题不是模型不行，而是 LLM 推理的&amp;quot;显存管理&amp;quot;是个完全不同于传统推理服务的工程问题 —— KV Cache 巨大、请求长度不可预测、解码阶段每一步都要访问全部历史。&lt;/p&gt;
&lt;p&gt;通用 Web 服务的优化套路（线程池、连接池、批处理）在 LLM 推理里要么不适用，要么被彻底重做。2023 年以来，社区推出了三套成熟的推理服务：vLLM（伯克利系，学术派）、HuggingFace TGI（工业标准派）、NVIDIA Triton（GPU 厂商派）。三者的设计哲学、性能上限、运维成本差异很大，本文从 KV Cache 管理、调度策略、生态成熟度三个维度，给出工程取舍。&lt;/p&gt;</description></item></channel></rss>