All tags

moe

9 posts tagged with "moe"

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware

WiSP 将低资源 MoE serving 建模为专家权重与 KV 缓存两条 working-set 争夺同一 VRAM 的问题。核心是 routing-aware expert pager(基于 expert_map 的 LRU 分页,字节等价输出)加 MV-WSA 边际价值分配器,联合分配 expert scratch 与 KV pool。iso-VRAM 下相对 vLLM static offload 解码吞吐最高提升 1.95×,Jamba-52B/MiniMax-M2-229B 等 baseline 无法启动的模型上也能可服务,且揭示单流 decode 中 routing 预测只能省显存不能省延迟。