Social robots could soon understand where people are looking using ordinary cameras rather than expensive eye-tracking equipment. The technology could make natural robot interaction more affordable and practical in hospitals, schools, homes and workplaces. This is demonstrated by research conducted by AI scientist Linlin Cheng.
As robots move beyond research laboratories and into shared human spaces, being able to understand social cues will become increasingly important.
‘My research focuses on whether social robots can understand human gaze using only ordinary cameras, without relying on specialised eye-tracking hardware,’ Cheng says.
Traditional eye trackers can provide accurate measurements, but they are expensive, intrusive and hard to deploy outside controlled lab settings. ‘I therefore explored whether appearance-based gaze estimation, which predicts gaze direction from standard camera images, could offer a practical alternative,’ she continues.
The findings suggest it can.
According to Cheng: ‘Robots can get a useful sense of where someone is looking just from a regular camera, without needing special eye-tracking equipment. This means gaze-based interaction could become far more practical and affordable for real robots used in everyday places like hospitals or schools.’
Good at attention, less good at quick glances
The technology does have limitations. Rather than precisely tracking every eye movement, it is currently better at recognising broader patterns of attention.
‘It works well for picking up on someone's general pattern of attention, like whether they are looking at the robot or at a task in front of them,’ Cheng says, ‘particularly when the robot and person are between one and two metres apart.’
But it still struggles to catch quick glances or sudden shifts in attention, the research found. ‘Such rapid movements are a natural part of human behaviour but can be difficult for a camera-based system to capture.
That means the technology is not yet a replacement for highly accurate eye trackers in every situation. Its strength lies instead in providing robots with a useful, low-cost indication of where a person’s attention is focused.’
From medication trays to factory floors
Cheng’s research also shows how this information could be used by robots in practical situations.
‘I showed that a robot can use this kind of gaze tracking to notice when someone has finished a task, using only low-cost cameras rather than expensive lab equipment.’
That could open up applications in care, education and industry.
‘A care robot assisting an elderly person could tell whether they are looking at the medication tray or looking away, distracted or confused, and adjust its behaviour accordingly, without wearing any eye-tracking gear,’ Cheng explains.
In a workplace, the same principle could help robots understand when a person has completed an activity.
‘An assistive robot in a workshop could detect when someone has finished a task just by tracking their gaze shift away from it.’
Making robots cheaper and easier to use
Because the system relies on equipment that is already available on many robots, it could lower one of the barriers to bringing social robots into everyday environments.
‘This approach only needs an ordinary camera. It lowers the cost and complexity of deploying social robots, making near-term real-world use more realistic within the next few years rather than decades,’ Cheng says.
‘This is especially relevant as robots increasingly move from labs into shared human spaces, where affordable, natural interaction matters for trust and usability.’
According to Cheng, the potential users are broad: ‘Patients in hospitals, students in classrooms, and workers on factory floors or in warehouses could all benefit from robots that have a basic understanding of where people are directing their attention.’
Testing robots with real people
To establish how well the approach works, the research combined technical comparisons with experiments involving real people and robots.
First, Cheng ran comparison studies in which appearance-based gaze estimation models were tested against ground truth from eye trackers and manual human annotation, to see how closely their output matched more established methods.
She then moved into human-robot interaction experiments, in which participants interacted with real robots during structured tasks while their gaze was tracked and analysed.