LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention
A novel fused Indexer-TopK kernel that exploits the curse of dimensionality to accelerate sparse attention in long-context LLM inference, achieving up to 3.38x speedup on B200.