MARATTO

article

A Multi-Task Convolutional Neural Network for Gaze Estimation: A From-Scratch Approach

Abstract

Estimating gaze direction often suffers from the presence of distracting, non-gaze-related features within full-face images, which can hinder accurate prediction. In this article, we present GazeMTNet, a lightweight multi-task convolutional neural network that was created from scratch for the joint estimate of head posture and gaze direction. Our architecture makes use of the dependency to increase the accuracy of gaze estimation while minimizing computational complexity. Experiments on the EyeDiap and the MPIIGaze datasets show that GazeMTNet outperforms current techniques in terms of efficiency and accuracy, attaining remarkable performance, with mean angular errors of 4.29 and 3.20°, respectively. The 0.05 G FLOPs and 1.012 million parameters make our model ideal for embedded and real-time applications. These outcomes support the efficacy of multi-task learning when it comes to gaze estimation.

Research topics

  • Gaze Tracking and Assistive Technology
  • Vestibular and auditory disorders
  • Visual Attention and Saliency Detection

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/aiccsa66935.2025.11315435

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.