arXiv cs.AIPaper
LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
Simplifying self-supervised video pretraining to one encoder and one hyperparameter is the kind of efficiency win that matters for anyone training world models on tight compute budgets. If the collapse-free guarantee holds at scale, it could become a default recipe the way SimCLR-style objectives did for images. Worth tracking for infra and research teams working on video foundation models, not urgent for anyone else.