OceanNav: Language-Conditioned World Models for Underwater Visual Navigation via Embodied Agents
编号:1428
访问权限:仅限参会人
更新:2026-09-01 00:34:37 浏览:0次
口头报告
摘要
Marine robots, operating as embodied agents, are essential for ocean science applications, but complex ocean environments—including communication bottlenecks, energy constraints, and hydrodynamic disturbances—severely degrade visual perception and physical interaction. While existing underwater vision-language-navigation (VLN) algorithms enable natural-language instruction following, traditional reactive policies fail to adapt to unpredictable ocean currents and intermittent signal loss, highlighting the urgent need for world models that equip embodied agents with predictive state estimation and forward-looking planning.
To address these challenges, we propose OceanNav, a language-conditioned world model framework for underwater visual navigation tasks. Designed for resource-constrained marine platforms, OceanNav unifies language grounding, predictive world modeling, and active physical control to navigate visually degraded environments. It grounds the agent's spatial awareness by fusing temporal visual features with multi-modal physical sensing—including inertial sensing (IMU), Doppler velocity log (DVL), and depth measurements—into a task-oriented latent belief state. Utilizing edge computing architecture, the embodied agent locally executes real-time topological planning and large-model inference without relying on high-bandwidth continuous communication. Unlike purely reactive systems, OceanNav utilizes natural language instructions alongside an initial ego-centric observation to anticipate future states via open-loop trajectory prediction. It predicts target visibility, relative pose, and observation uncertainty, ensuring continuous language-conditioned subgoal planning even when real-time visual observations are unreliable. By tightly coupling high-level semantic understanding with low-level physical dynamics, OceanNav provides a robust framework for autonomous navigation and inspection in complex marine environments.
稿件作者
Xiaoliang Fan
Fujian Ocean Innovation Center;School of Informatics, Xiamen University
发表评论