LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
First stable end-to-end JEPA that trains from raw pixels using only two loss terms (prediction + SIGReg), with ~15M params on single GPU, 48x faster planning than foundation-model world models