【中文】论文速递|Robostral Navigate:Mistral 的 8B 单目导航 VLM,超越深度多相机方案
Mistral AI 发布 Robostral Navigate,一个仅凭单目 RGB 视频预测路径点的 8B 视觉-语言导航模型:在图像空间而非机器人坐标系中输出,换机器人无需重新标定。训练数据来自 35 万仿真场景生成的 240 万条轨迹,前缀缓存训练把 token 消耗降低 22 倍(训练时间“从数月到数天”),树状注意力掩码保证视觉锚定,再用强化学习提升探索与恢复能力。R2R-CE 基准成功率 77.4%,比最佳单目方法高 10.5 个百分点,甚至比深度/多相机系统高 5.3 个百分点;RxR-CE 达 75.1%。
【EN】Paper Brief | Robostral Navigate: Mistral's 8B Monocular Navigation VLM Beats Depth and Multi-Camera Systems
Mistral AI's Robostral Navigate is an 8B vision-language model that predicts navigation waypoints from monocular RGB video alone, operating in image space rather than robot-specific coordinates so it transfers across robot types without recalibration. Trained on 2.4 million trajectories from 350k simulated scenes, with prefix-caching training cutting tokens 22x (training time from months to days), tree-based attention masking for visual grounding, and RL for exploration and recovery. It hits 77.4% success on R2R-CE — 10.5 points above the best monocular method and 5.3 points above depth/multi-camera systems — and 75.1% on RxR-CE.
来源 Source:
留言
NO REGISTRATION · 临时网名 + 邮箱即可开聊