November 28, 2024 Marconi: Prefix Caching for the Era of Hybrid LLMs Marconi为混合架构LLM提出高效前缀缓存机制,支持注意力层和非注意力层的统一缓存管理。 transformer-variant llm-inference attention-mechanism long-context kv-cache-optimization