January 21, 2025 InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling 通过长和丰富上下文(LRC)建模,增强视频多模态大语言模型的细粒度细节感知和长时序结构捕获能力 video-understanding multimodal vision-language video-llm long-context token-compression