March 25, 2024 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 通过三阶段渐进训练和6B参数视频编码器,在70+视频理解任务上实现SOTA性能的视频基础模型 video-understanding multimodal vision-language foundation-model video-encoder