June 24, 2026 TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale 提出基于 CXL 共享内存的机架级 KV Cache 系统 TraCT,消除 RDMA 网络跳数,实现 GPU-CXL 直接 DMA 传输,TTFT 降低 9.8x,吞吐量提升 1.6x llm-serving disaggregated-inference kv-cache cxl shared-memory distributed-inference system-design