ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

cuda

3 posts tagged with "cuda"

July 24, 2026

LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention

A novel fused Indexer-TopK kernel that exploits the curse of dimensionality to accelerate sparse attention in long-context LLM inference, achieving up to 3.38x speedup on B200.

sparse-attention gpu-kernel llm-inference top-k-selection cuda vllm
July 4, 2026

PyGraph: Robust Compiler Support for CUDA Graphs in PyTorch

针对 PyTorch 中 CUDA Graph 部署困难的编译器框架。提出三项优化:CGCT(自动代码变换使 ML 程序兼容 CUDA Graph)、PI(参数间接寻址将数据拷贝转为指针拷贝,最高减少 99% 拷贝量)、SCG(基于成本效益分析的有选择部署)。在 25 个 ML 工作负载上,PyGraph 比 PyTorch2-CG 平均提速 29%,最高 3.36×,且在分布式 Tensor Parallelism 下表现更强。

compiler cuda gpu-optimization pytorch system-optimization
June 30, 2026

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding

PRR 推理运行时,通过 EMA 预测器+增量修复内核将动态稀疏注意力的选择-注意力依赖瓶颈降低 30%+

efficient-inference llm-serving sparse-attention kv-cache cuda

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.