理解 PD 分离和分布式 KVCache 的几张图
图一:NVIDIA Dynamo, A Low-Latency Distributed Inference Framework for Scaling Reasoning AI Models. 图二:llm-d, a Kubernetes-native high-performance distributed LLM inference framework. 图三:Mooncake is a KVCache-centric disaggregated architecture for LLM serving. The core of Mooncake is the Transfer Engine(核心), which provides a unified interface for batched data transfer across various storage devices and network links. 图四:vLLM production stack, vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization. 图五:AIBrix is an open-source initiative designed to provide essential building blocks to construct scalable GenAI inference infrastructure.