A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM Inference
Marvell Photonic Fabric Memory Appliance — 用无源光纤全互连拓扑替代电交换机,实现32TB共享DDR5内存的CXL内存池化
34 posts tagged with "memory-efficiency"
Marvell Photonic Fabric Memory Appliance — 用无源光纤全互连拓扑替代电交换机,实现32TB共享DDR5内存的CXL内存池化
CXL-Hybrid Memory-based remote KV-cache tier for scalable multi-turn LLM serving
AAFLOW+系统将KV缓存作为一等分发系统对象,实现零拷贝的多智能体工作流执行状态共享
面向AI加速器的LLM Token阶段优化框架FastTPS,包括全局KV缓存管理、序列维度扁平化、Fusion MLP等技术
面向 LLM 推理的 HBF(高带宽闪存)与 HBM GPU 软硬件协同设计:通过架构级 SRAM 缓存/预取、专用数据布局与异构存储管理层,将 HBF 的容量优势转化为吞吐与能效。
TriRoute — 单一轻量控制器联合决策注意力分辨率、FFN 专家选择与 KV-Cache 位宽,端到端可训练,单一预算旋钮扫出 Pareto 前沿,缓解跨轴路由坍缩级联
An analytical performance model (LIMINAL) that abstracts application requirements and hardware capabilities to systematically explore LLM inference performance limits across current, near-future, and hypothetical hardware
A software library that heuristically determines optimal tile sizes and queue configurations for TMA-based GPU kernels using the GPU Specification Table (GST) and Little's Law
A prefetch-aware (PA) warp scheduling policy that coordinates thread scheduling and data prefetching in GPGPUs to better tolerate long memory latencies
Position persistent sparse attention leveraging spatial coherence of high-attention tokens across Transformer layers for 2.1× decoding speedup
A quantitative analysis of FP8 E4M3 attention P-casting revealing P-collapse under attention sink and characterizing S=256 as the optimal static scaling factor
A general-purpose interactive text/image-to-video world model supporting camera navigation, scene revisits, and promptable events with 16 FPS streaming on 8x RTX 5090
覆盖分布式RL、RLHF系统、环境加速、采样优化、通信优化、内存优化等方向
面向资源受限环境的MoE模型高效推理系统,通过智能CPU-GPU协作实现最优执行策略
多模型LLM调度的实证研究,分析CPU-GPU卸载和抢占的性能影响,为下一代调度系统提供设计指导
通过专家卸载和动态放置实现MoE模型训练的高效扩展,支持67倍专家扩展和17.5倍吞吐量提升
面向GPU的3D稀疏卷积加速引擎
基于CPU-GPU-I/O流水线调度的高吞吐量MoE推理系统
通过细粒度专家卸载优化MoE大模型服务的延迟-内存权衡
通过灵活高效的卸载打破内存约束实现设备端LLM推理
基于模块级批处理的单GPU高吞吐量MoE推理系统
面向推理模型的冗余感知KV缓存压缩
通过自辅助推测解码实现高效MoE推理,最高4.30×吞吐量提升
通过O(1)主机缓存实现快速、实时的大模型自动扩展
通过级联规划器(EAM+EAP)实现高效专家预取,在资源受限设备上加速MoE推理,吞吐量提升65.13%
通过量化填充、专家内存池和层感知调度实现无回退的高效MoE推理,相比MoE-Infinity加速2.86倍
解耦敏感度与重要性:面向推理的KV缓存压缩框架
利用CPU计算实现灵活的LLM推理
MegaTrain:在单GPU上全精度训练100B+参数大语言模型
通过更好的并行性和工作分配实现更快的注意力
通过集体KV缓存共享扩展多智能体LLM服务
LLM推理中KV缓存管理策略的比较表征
GMLake 提出高效透明的 GPU 内存碎片整理方法,提升大模型训练效率。
PipeFill利用流水线并行训练中的气泡时间执行额外计算任务,提升GPU利用率。