【中文】论文速递|ReferTrack:先指代、后跟踪的具身视觉跟踪(腾讯)
腾讯提出 ReferTrack,解决移动智能体按自然语言描述跟踪目标的难题。现有视觉-语言-动作策略在抽象空间表征中推理、难以监督,ReferTrack 改为两阶段:先从候选框中“指代”锁定目标,再基于该显式决策规划跟踪路径点,并用滑动窗口的历史框队列维持时序上下文。在 EVT-Bench 上取得单目 SOTA:单目标 89.4%、干扰场景 73.3%、歧义场景 74.1%,追平甚至超过多相机基线,并在四足与人形机器人上完成真机验证。
【EN】Paper Brief | ReferTrack: Referring Then Tracking for Embodied Visual Tracking (Tencent)
Tencent's ReferTrack tackles embodied visual tracking, where a mobile agent must follow a target described in natural language. Instead of reasoning in hard-to-supervise abstract spatial representations, it splits the task in two: first refer, identifying the target among bounding-box options, then plan tracking waypoints from that grounded decision, maintaining temporal context via a sliding-window queue of past boxes. It sets single-camera state of the art on EVT-Bench (89.4% single-target, 73.3% distracted, 74.1% ambiguity), matching or beating several multi-camera baselines, with real-robot validation on legged and humanoid platforms.
来源 Source:
留言
NO REGISTRATION · 临时网名 + 邮箱即可开聊