Abstract
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and weakly aligned with explicit image-space detections. To address this, we introduce ReferTrack, a referring-then-tracking paradigm that grounds EVT using a single forward-facing camera. Our model first selects the target from an indexed set of bounding boxes, then decodes tracking waypoints conditioned on this image-grounded decision. To preserve target motion cues over time, ReferTrack maintains a sliding-window queue of previously selected bounding boxes, injecting their geometric features into the visual history via temporal-viewpoint-bbox indicator (TVBI) tokens. We further enhance target identification by co-training on a custom Refer-QA dataset. On EVT-Bench, ReferTrack achieves state-of-the-art single-view performance with success rates of 89.4%, 73.3%, and 74.1% on the single-target, distracted, and ambiguity tracking splits, respectively -- matching or even surpassing several multi-camera baselines on identification-heavy tasks. Finally, real-world deployments on legged and humanoid robots validate its robust sim-to-real transfer capabilities. Code is available at https://github.com/MedlarTea/referTrack.
Community
New embodied visual tracking SOTA on EVT Bench🏆
code, data and ckpt will be released on Github(https://github.com/MedlarTea/referTrack), welcome to try💪
An outstanding job !!!
impressive
support🤞🤞🤞
impressive
brilliant!!
impressive
Really good work👍
Goooooooood Work!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning (2026)
- ABot-N1: Toward a General Visual Language Navigation Foundation Model (2026)
- CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking (2026)
- SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation (2026)
- Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments (2026)
- FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models (2026)
- DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.20061 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper