Hybrid intelligence and visual expertise

We investigate how clinicians’ visual attention can inform image analysis, model training, and human–AI interaction.

How can a learning system use the way an expert examines an image, as well as the image itself?

Diagram linking a gaze scanpath and image patches to graph features.
Conceptual research schematic.

Method overview

From visual attention to image understanding

A sequence of fixation points records how an image is examined.

Schematic of the research approach; image patterns and measurements are illustrative.

Research overview

An expert’s interpretation involves a sequence of observations: where they look, how long they attend to a region, and how they move between findings. Our work studies these signals as a complementary source of information for medical image analysis. Eye tracking provides a record of visual attention; it does not, by itself, reveal a clinician’s reasoning.

In GazeGNN, image features and recorded gaze are represented together as a graph for chest X-ray classification. This allows the model to relate local visual information to the spatial and temporal structure of expert fixations. Related work examines how gaze can guide segmentation and interaction with foundation models.

From visual attention to image understanding

  • Record fixation locations and their sequence during image interpretation.
  • Associate gaze observations with image patches and learned visual features.
  • Evaluate models and interactions against task-specific image analysis outcomes.

Selected research sources

Research enquiries

Contact the lab