HyMCache: A KV Cache Framework for Multi-Turn LLM Serving with CXL-Hybrid Memory
CXL-Hybrid Memory-based remote KV-cache tier for scalable multi-turn LLM serving
6 posts tagged with "prefix-caching"
CXL-Hybrid Memory-based remote KV-cache tier for scalable multi-turn LLM serving
首个面向企业级LLM推理的开源高效KV缓存层,支持跨查询复用和PD分离架构
面向LLM多智能体工作流的高效前缀缓存管理框架,通过Agent Step Graph和工作流感知驱逐策略实现最高2.19倍加速
KV-Cache-centric memory management system for efficient embodied planning with static-dynamic memory construction, multi-hop recomputation, and layer-balanced loading
基于GPU虚拟内存管理的高效LLM推理张量结构
Agentic RL 正在重塑 LLM 后训练,但端到端训练时间被计算密集的多轮 rollout 主导,占总时间的 70% 以上。关键挑战: