ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

distributionally-robust-optimization

1 post tagged with "distributionally-robust-optimization"

July 23, 2026

Robust KV Cache Management for LLM Serving under Output Token Length Uncertainty

A robust KV cache management framework using Wasserstein distributionally robust optimization (DRO) to jointly optimize GPU parallelism, cache reservation, request routing, and prefix caching under output token length uncertainty.

llm-serving kv-cache distributionally-robust-optimization gpu-clusters preemption-management throughput-latency

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.