Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning

Yihan Zhou1,2,3*, Rui Yan1*, Mingcong Li4, Zheyuan Huang4, Xu Yang1, Xueyang Guo1,2,3, Yilin Mo1,2,3★
1Department of Automation, Tsinghua University
2Beijing Key Laboratory of Embodied Intelligence Systems
3Institute for Embodied Intelligence and Robotics, Tsinghua University
4LingYu Robotics
*Equal Contribution ★Corresponding Author

Abstract

Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce eye-tracker-supervised gaze prompting, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with π0, gaze prompting increases mean success from 26.3% to 56.0% across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release GazeMani, a dataset of 1,200 teleoperated trajectories with synchronized gaze.

Model Architecture


Gaze Prompts Framework Overview

Experimental Results

Task π0 π0 + Aux-Loss π0 + Gaze-Prompt
Stack Bowls 19/50 30/50(+22%) 36/50(+34%)
Stack Cubes in Order 15/50 22/50(+14%) 26/50(+22%)
Bus Table 8/50 4/50(-8%) 22/50(+28%)
Fold Towel 21/50 10/50(-22%) 22/50(+2%)
Insert Toilet Paper 9/50 22/50(+26%) 36/50(+54%)
Catch Rolling Ball 7/50 12/50(+10%) 26/50(+38%)
Performance comparison on six real-world tasks. Each task is evaluated over fifty trials.

Real World Videos

We finetune π0 model and compare Gaze Prompts against Aux-Loss to evaluate our method, and we visualize sample rollouts below.




Task: Stack Yellow Bowls Click to view uncut single-take recordings of all rollouts

Failure: π0

Failure: π0 + Aux-Loss

Success: π0 + Gaze-Prompt

Task: Stack Red Bowls Click to view uncut single-take recordings of all rollouts

Failure: π0

Failure: π0 + Aux-Loss

Success: π0 + Gaze-Prompt

Task: Stack Cubes in Order Click to view uncut single-take recordings of all rollouts

Failure: π0

Failure: π0 + Aux-Loss

Success: π0 + Gaze-Prompt

Task: Bus Table Click to view uncut single-take recordings of all rollouts

Failure: π0

Failure: π0 + Aux-Loss

Success: π0 + Gaze-Prompt

Task: Fold Towel Click to view uncut single-take recordings of all rollouts

Failure: π0

Failure: π0 + Aux-Loss

Success: π0 + Gaze-Prompt

Task: Insert Toilet Paper Click to view uncut single-take recordings of all rollouts

Failure: π0

Failure: π0 + Aux-Loss

Success: π0 + Gaze-Prompt

Task: Catch Rolling Ball Click to view uncut single-take recordings of all rollouts

Failure: π0

Failure: π0 + Aux-Loss

Success: π0 + Gaze-Prompt