HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation
HAM-VLN 为 zero-shot VLN agent 提供决策耦合的层次化记忆系统,通过世界图存储导航历史,结合动态门控检索,在无需训练的情况下超越现有方法
17 posts tagged with "robotics"
HAM-VLN 为 zero-shot VLN agent 提供决策耦合的层次化记忆系统,通过世界图存储导航历史,结合动态门控检索,在无需训练的情况下超越现有方法
A streaming inference framework enabling real-time 50Hz control for Flow Matching Vision-Language-Action models via partitioned attention, AdaRMSNorm, and async pipeline
LingBot-Video — 首个开源大规模 MoE 视频基础模型,面向具身智能的单流扩散 Transformer,稀疏 MoE + 五阶段课程预训练 + 多维奖励 GRPO 后训练
Internet-scale JEPA video model pretrained on 1M+ hours of video, achieving SOTA understanding benchmarks and enabling zero-shot robotic planning via action-conditioned post-training with <62h robot data
Next-gen JEPA model with dense predictive loss, deep self-supervision, and multi-modal tokenizers achieving SOTA on dense visual tasks (depth, segmentation, STA) while retaining global scene understanding
LingBot-World, an open-sourced world simulator stemming from video generation, featuring minute-level horizon, real-time interactivity, emergent memory, and three-stage evolutionary training pipeline
A comprehensive survey of world models from a robot-learning perspective, examining policy coupling, learned simulators for RL and evaluation, robotic video generation, and benchmarks/datasets across embodied applications
从动作 tokenization 视角统一审视 VLA 模型的综述。提出 VLA 模块 + 动作 token 的统一框架,将现有 VLA 模型归为 8 种动作 token 类型(语言描述/代码/可供性/轨迹/目标状态/隐表示/原始动作/推理),逐一分析每种类型的优势、局限与未来方向,并提出分层架构趋势与从 VLA 模型到 VLA 智能体的演进路径。
面向真实世界部署的 VLA 全栈综述。系统梳理 VLA 三大挑战(数据稀缺 / 具身迁移 / 算力成本)、架构演进(CNN→Transformer→VLM→Diffusion/Flow→分层控制),并给出三类核心架构(sensorimotor 7 变体 / world model 3 模式 / affordance 3 模式)、模态处理(视觉/语言/动作/音频/触觉/3D)、训练范式(SL/SSL/RL + 预训练/后训练/梯度隔离/LoRA)、数据采集(遥操作/代理设备/人类视频)、公开数据集(Ego4D/OXE/DROID/AgiBot World)、机器人平台与评测基准(LIBERO/CALVIN/ManiSkill/SIMPLER/RoboArena)。最后给出面向实践者的 6 条建议与 8 个未来方向。
End-to-End VLA with Differentiable Latent Intent Bottleneck
具有空间理解和在线强化学习的视觉-语言-动作模型
通过潜在轨迹引导实现运动可控的视频生成
A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
统一的具身基础模型,通过DiT动作解码器将视觉-语言建模扩展到连续动作生成
A lightweight bridging paradigm that enables SOTA-level VLA performance with only 0.5B parameters, without robotic data pre-training
小米机器人团队发布的先进VLA模型,通过精心设计的训练方案和部署策略实现高性能实时执行
通过强化学习增强视觉-语言-动作框架,实现四足机器人的显式推理和连续控制