ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

linear-attention

2 posts tagged with "linear-attention"

August 6, 2026

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

内核-运行时协同设计:通过闭式树验证、值域分块GPU核和因子化推测状态,将混合注意力LLM的树投机解码吞吐量提升至自回归解码的4.72×

tree-speculative-decoding hybrid-attention linear-attention gdn llm-inference sglang gpu-kernel-optimization closed-form-parallelization
July 23, 2026

SpecLA: Efficient Speculative Decoding for Linear-Attention Models

A speculative decoding runtime for stateful linear-attention models that verifies chains and trees with topology-aware kernels, stores compact factors to recover accepted states, and uses confidence pruning plus a target-aligned EAGLE-style drafter.

speculative-decoding linear-attention gdn recurrent-state gpu-kernels inference-acceleration triton

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.