PhiZero: A World Model Built Around Physical Language
A physical world model that learns a compact discrete 'physical language' from in-the-wild videos, then uses a reason-then-render paradigm to generate and understand physically coherent video
12 posts tagged with "world-model"
A physical world model that learns a compact discrete 'physical language' from in-the-wild videos, then uses a reason-then-render paradigm to generate and understand physically coherent video
多块预测(MCP)训练框架,通过轻量辅助模块同时去噪多个未来视频块,加速世界动作模型训练与推理
First generative game engine combining multi-player control, real-time inference, complex physical interaction, and adversarial gameplay in KOF '97
一个因果视频生成世界模型,实现无限长、实时响应的交互式世界模拟,支持多种角色动作和环境事件,并引入Pilot-Director双智能体协作框架
LingBot-Video — 首个开源大规模 MoE 视频基础模型,面向具身智能的单流扩散 Transformer,稀疏 MoE + 五阶段课程预训练 + 多维奖励 GRPO 后训练
LingBot-World, an open-sourced world simulator stemming from video generation, featuring minute-level horizon, real-time interactivity, emergent memory, and three-stage evolutionary training pipeline
A comprehensive survey of world models from a robot-learning perspective, examining policy coupling, learned simulators for RL and evaluation, robotic video generation, and benchmarks/datasets across embodied applications
A test-time adaptation framework that performs lightweight updates within the closed loop of Model Predictive Control (MPC) to continuously recalibrate a JEPA world model during deployment
面向少步自回归世界模型推理的跨分块缓存加速
NVIDIA全模态世界模型,统一语言、图像、视频、音频和动作的理解与生成,为Physical AI提供通用骨干网络
全栈开源框架,将T2V/TI2V基础模型转换为相机可控的少步自回归视频世界模型
腾讯混元多模态3D世界模型:统一生成与重建