<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Memory Bandwidth on Jeanphilo Blog</title><link>https://shio-chan-dev.github.io/jeanblog/zh/tags/memory-bandwidth/</link><description>Recent content in Memory Bandwidth on Jeanphilo Blog</description><generator>Hugo -- 0.165.0</generator><language>zh-cn</language><lastBuildDate>Wed, 26 Aug 2026 00:00:00 +0800</lastBuildDate><atom:link href="https://shio-chan-dev.github.io/jeanblog/zh/tags/memory-bandwidth/index.xml" rel="self" type="application/rss+xml"/><item><title>Speculative Decoding 为什么能加速：从并行验证到 GPU Memory-Bound</title><link>https://shio-chan-dev.github.io/jeanblog/zh/ai/llm/speculative-decoding-gpu-memory-bound/</link><pubDate>Wed, 26 Aug 2026 00:00:00 +0800</pubDate><guid>https://shio-chan-dev.github.io/jeanblog/zh/ai/llm/speculative-decoding-gpu-memory-bound/</guid><description>从自回归生成的串行依赖出发，解释 speculative decoding 如何用一次 Target forward 并行验证多个 Draft token，以及 acceptance rate、batch 和 KV Cache 如何决定实际收益。</description></item></channel></rss>