V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Next-gen JEPA model with dense predictive loss, deep self-supervision, and multi-modal tokenizers achieving SOTA on dense visual tasks (depth, segmentation, STA) while retaining global scene understanding