SpecLA: Efficient Speculative Decoding for Linear-Attention Models
A speculative decoding runtime for stateful linear-attention models that verifies chains and trees with topology-aware kernels, stores compact factors to recover accepted states, and uses confidence pruning plus a target-aligned EAGLE-style drafter.