V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Internet-scale JEPA video model pretrained on 1M+ hours of video, achieving SOTA understanding benchmarks and enabling zero-shot robotic planning via action-conditioned post-training with <62h robot data