June 24, 2026 Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving 面向 Kimi 的 KVCache 中心化解耦架构,通过分离 Prefill 和 Decode 集群实现 525% 吞吐量提升 llm-serving disaggregated-serving kv-cache system-design production-system moonshot-ai
June 23, 2026 Frontier: Towards Comprehensive and Accurate LLM Inference Simulation 面向现代 LLM 推理服务的离散事件模拟器,支持解耦架构、运行时优化和有状态工作负载,吞吐量误差低于 4% llm-serving simulation distributed-inference performance-modeling disaggregated-serving system-design