ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

CUDA Kernel Fusion

1 post tagged with "CUDA Kernel Fusion"

May 1, 2026

TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference

首个在低延迟推理场景下为张量并行LLM推理实现高效计算通信重叠的系统,通过融合AllReduce-RMSNorm核和波感知智能分割达成最高1.28×加速

LLM Inference Tensor Parallelism Compute-Communication Overlap CUDA Kernel Fusion NVSHARP Multimem RMSNorm vLLM

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.