July 23, 2026 Sol Video Inference Engine Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation video-diffusion inference-acceleration agent-framework quantization sparse-attention kernel-fusion token-pruning cache
December 2, 2024 Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking 通过动态输入剪枝和缓存感知掩码实现高效LLM推理,降低计算和内存开销。 token-pruning model-compression quantization sparse-attention llm-training