ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

gpu-memory-management

1 post tagged with "gpu-memory-management"

July 23, 2026

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

A runtime memory-management system for MoE LLM serving that dynamically quantizes expert weights at runtime to balance GPU memory between model weights and KV cache, using a quality-aware planner with offline sensitivity, online routing statistics, and prompt residuals.

moe llm-serving weight-quantization gpu-memory-management runtime-quantization pagedattention any-precision

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.