June 24, 2026 FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling FlowKV 提出低成本低延迟的 KV Cache 传输方案和负载感知调度器,实现分离式 LLM 推理框架的 96% 传输延迟降低 llm-serving disaggregated-inference kv-cache load-balancing system-design nccl heterogeneous-gpu