Failure: π0
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce eye-tracker-supervised gaze prompting, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with π0, gaze prompting increases mean success from 26.3% to 56.0% across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release GazeMani, a dataset of 1,200 teleoperated trajectories with synchronized gaze.
| Task | π0 | π0 + Aux-Loss | π0 + Gaze-Prompt |
|---|---|---|---|
| Stack Bowls | 19/50 | 30/50(+22%) | 36/50(+34%) |
| Stack Cubes in Order | 15/50 | 22/50(+14%) | 26/50(+22%) |
| Bus Table | 8/50 | 4/50(-8%) | 22/50(+28%) |
| Fold Towel | 21/50 | 10/50(-22%) | 22/50(+2%) |
| Insert Toilet Paper | 9/50 | 22/50(+26%) | 36/50(+54%) |
| Catch Rolling Ball | 7/50 | 12/50(+10%) | 26/50(+38%) |
We finetune π0 model and compare Gaze Prompts against Aux-Loss to evaluate our method, and we visualize sample rollouts below.
Failure: π0
Failure: π0 + Aux-Loss
Success: π0 + Gaze-Prompt
Failure: π0
Failure: π0 + Aux-Loss
Success: π0 + Gaze-Prompt
Failure: π0
Failure: π0 + Aux-Loss
Success: π0 + Gaze-Prompt
Failure: π0
Failure: π0 + Aux-Loss
Success: π0 + Gaze-Prompt
Failure: π0
Failure: π0 + Aux-Loss
Success: π0 + Gaze-Prompt
Failure: π0
Failure: π0 + Aux-Loss
Success: π0 + Gaze-Prompt
Failure: π0
Failure: π0 + Aux-Loss
Success: π0 + Gaze-Prompt