June 24, 2026 Scalable Processing-Near-Memory for 1M-Token LLM Inference CXL-Enabled KV-Cache Management Beyond GPU Limits llm-inference kv-cache processing-near-memory cxl memory-management hardware-acceleration long-context