<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>量化 on Tony老师的博客</title><link>https://blog.tanteng.space/tags/quantization/</link><description>Recent content in 量化 on Tony老师的博客</description><generator>Hugo</generator><language>zh</language><lastBuildDate>Tue, 25 Mar 2025 14:00:00 +0800</lastBuildDate><atom:link href="https://blog.tanteng.space/tags/quantization/index.xml" rel="self" type="application/rss+xml"/><item><title>LLM 推理优化：KV Cache、PagedAttention 与量化</title><link>https://blog.tanteng.space/2025/03/llm-inference-optimization/</link><pubDate>Tue, 25 Mar 2025 14:00:00 +0800</pubDate><guid>https://blog.tanteng.space/2025/03/llm-inference-optimization/</guid><description>&lt;blockquote&gt;
&lt;p&gt;上线一个 70B 模型，自以为把 transformers 包进 FastAPI 就算生产就绪。结果 P99 延迟 12 秒、显存爆掉、并发只有 4。问题不是模型不行，而是 LLM 推理的访存模式和传统 CNN 推理是两个世界——KV Cache 占显存、解码是 memory-bound、长度不可预测。&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;这是一篇 LLM 推理优化的&amp;quot;算法地图&amp;quot;。Phase 6 的 &lt;code&gt;llm-serving-architecture.md&lt;/code&gt; 讲了 vLLM/TGI/Triton 三套服务的工程对比；本文深入到&lt;strong&gt;推理算法层&lt;/strong&gt;，讲四个 10 倍速提升的技术：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;PagedAttention&lt;/strong&gt;：把 KV Cache 切成页，显存利用率从 ~30% 提到 ~95%&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;FlashAttention-2&lt;/strong&gt;：用 tiling 把 attention 的 HBM 读写从 O(N²) 降到 O(N)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;KV Cache 量化（KIVI）&lt;/strong&gt;：把已经生成的 KV 压到 2-bit，&lt;strong&gt;显存再砍 4 倍&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;模型权重量化（GPTQ/AWQ）&lt;/strong&gt;：把 70B 模型从 140GB 压到 20GB，&lt;strong&gt;单卡可跑&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Speculative Decoding&lt;/strong&gt;：用小模型草稿 + 大模型验收，&lt;strong&gt;无损 2-3× 加速&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>