Research area

Spherical and Audio-Visual Intelligence

We study how people and machines attend to immersive 360° scenes, combining spherical geometry, temporal context, and spatial audio.

Three 360-degree scenes comparing the video frame, spatial audio energy, and viewer fixation density under ambisonic, mono, and muted audio.
Viewer fixation density changes with the audio condition. Across concert, driving, and conversation scenes, ambisonic spatial audio concentrates attention differently from mono or muted viewing. Figure from the Spherical Vision Transformers paper.

What we study

Omnidirectional video surrounds the viewer, so planar assumptions no longer hold. The same scene is distorted differently across an equirectangular projection, attention unfolds across viewports and time, and spatial sound can pull a viewer toward events outside the current field of view.

We study geometry-aware learning, visual attention, eye tracking, temporal saliency, and audio-visual fusion for immersive media.

Current questions

  • How should transformers account for spherical geometry and projection distortion?
  • How does spatial audio change visual attention in 360° environments?
  • Which temporal context is most useful for predicting gaze?
  • How can viewport-level predictions remain consistent on the sphere?
  • Which datasets and protocols best reflect natural immersive viewing?

Our approach

SalViT360 divides an omnidirectional scene into tangent viewports and introduces spherical position information with spatio-temporal attention. Its audio-visual extension aligns spatial sound with visual features. The accompanying YT360-EyeTracking dataset measures gaze under mute, mono, and ambisonic audio conditions, making it possible to study how sound actively guides visual attention.

The work links machine perception with human vision and supports downstream applications such as immersive video quality assessment and adaptive streaming.

From the lab

Selected publications

2025 · IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)

Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos

We introduce the YT360-EyeTracking dataset and spherical geometry-aware vision transformers that combine visual and spatial-audio cues for saliency prediction in 360-degree video.

2018 · IEEE Trans. Multimedia 20(7) (2018)

Spatio-Temporal Saliency Networks for Dynamic Saliency Prediction

We propose spatio-temporal saliency networks for predicting dynamic visual saliency in videos.

2013 · Journal of Vision (2013)

Visual saliency estimation by nonlinearly integrating features using region covariances

We propose a visual saliency estimation method that nonlinearly integrates features using region covariances.

Browse all publications