ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

vllm

3 posts tagged with "vllm"

August 10, 2026

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure

首个将张量生命周期管理从计算栈中解耦的分布式服务层,通过 Artifact/Operation/Plan/Signal 四抽象实现 TaaS 范式,集成 vLLM/SGLang 后权重物化加速 60×、KV 缓存共享 TTFT 降低 87.5%

llm-infra tensor-management distributed-systems model-serving kv-cache taaS vllm sglang disaggregated-inference
July 24, 2026

LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention

A novel fused Indexer-TopK kernel that exploits the curse of dimensionality to accelerate sparse attention in long-context LLM inference, achieving up to 3.38x speedup on B200.

sparse-attention gpu-kernel llm-inference top-k-selection cuda vllm
June 24, 2026

Efficient Memory Management for Large Language Model Serving with PagedAttention

通过 PagedAttention 算法实现高效 KV cache 内存管理,构建 vLLM 高吞吐量 LLM 服务系统

llm-serving memory-management paged-attention kv-cache system-design production-system vllm

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.