-
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-centric or video-centered. We present $ω$-0, a latent predictive whole-body world-
-
Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning and evaluation in diverse real-world environments remains costly and challenging. Action-conditioned world models offer a promising alternative, but they often suffer from limited action controllability and poor generalization to out-of-distribution
-
Learning from demonstration (LfD) provides a developmental framework through which robots can develop motor skills by observing and imitating human dynamics, reducing reliance on explicit programming to teach a skill to a robot. The resulting human-like robot motion is recognised as a key factor in building trust and enabling natural collaboration in human-robot interaction. This paper presents a
-
重新审视人类视频如何赋能机器人学习 作者丨邓哲敏 编辑丨齐铖湧 大模型时代,规模化的数据能带来能力跃迁,这一规律已被反复验证。当这套范式进入机器人领域,问题变得复杂。机器人需要的不仅是大量数据,而且是能够转化为动作能力的数据。人类视频近期成为研究者关注的新来源。它规模巨大、获取成本低、覆盖场景丰富,但它毕竟不是机器人数据,视频里没有机器人动作标签,人类手部运动也无法直接对应机器人的