<!DOCTYPE article PUBLIC '-//NLM//DTD Journal Publishing DTD v2.1 20050630//EN' 'http://uploads.ingentaconnect.com/docs/dtd/ingenta-journalpublishing.dtd'>
<article article-type="research-article">
  <front>
    <journal-meta>
      <journal-id journal-id-type="aggregator">72010604</journal-id>
      <journal-title>Electronic Imaging</journal-title>
      <issn pub-type="ppub">2470-1173</issn><issn pub-type="epub"></issn>
      <publisher>
        <publisher-name>Society for Imaging Science and Technology</publisher-name>
        <publisher-loc>7003 Kilworth Lane, Springfield, VA 22151 USA</publisher-loc>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.2352/ISSN.2470-1173.2020.16.AVM-110</article-id>
      <article-id pub-id-type="sici">2470-1173(20200126)2020:16L.1101;1-</article-id>
      <article-id pub-id-type="publisher-id">ei_24701173_v2020n16_input/s14.xml</article-id>
      <article-id pub-id-type="other">/ist/ei/2020/00002020/00000016/art00013</article-id>
      <article-categories>
        <subj-group>
          <subject>Articles</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>End-to-End Multitask Learning for Driver Gaze and Head Pose Estimation</article-title>
      </title-group>
      <contrib-group>
        <contrib>
          <name>
            <surname>Ewaisha</surname>
            <given-names>Mahmoud</given-names>
          </name>
        </contrib>
        <contrib>
          <name>
            <surname>Shawarby</surname>
            <given-names>Marwa El</given-names>
          </name>
        </contrib>
        <contrib>
          <name>
            <surname>Abbas</surname>
            <given-names>Hazem</given-names>
          </name>
        </contrib>
        <contrib>
          <name>
            <surname>Sobh</surname>
            <given-names>Ibrahim</given-names>
          </name>
        </contrib>
      </contrib-group>
      <pub-date>
        <day>26</day>
        <month>01</month>
        <year>2020</year>
      </pub-date>
      <volume>2020</volume>
      <issue>16</issue>
      <fpage>110-1</fpage>
      <lpage>110-6</lpage>
      <permissions>
        <copyright-year>2020</copyright-year>
      </permissions>
      <abstract>
        <p>
          <italic>Modern automobiles accidents occur mostly due to inattentive behavior of drivers, which is why driver’s gaze estimation is becoming a critical component in automotive industry. Gaze estimation has introduced many challenges due to the nature of the surrounding environment like
 changes in illumination, or driver’s head motion, partial face occlusion, or wearing eye decorations. Previous work conducted in this field includes explicit extraction of hand-crafted features such as eye corners and pupil center to be used to estimate gaze, or appearance-based methods
 like Convolutional Neural Networks which implicitly extracts features from an image and directly map it to the corresponding gaze angle. In this work, a multitask Convolutional Neural Network architecture is proposed to predict subject’s gaze yaw and pitch angles, along with the head
 pose as an auxiliary task, making the model robust to head pose variations, without needing any complex preprocessing or hand-crafted feature extraction.Then the network’s output is clustered into nine gaze classes relevant in the driving scenario. The model achieves 95.8% accuracy on
 the test set and 78.2% accuracy in cross-subject testing, proving the model’s generalization capability and robustness to head pose variation.</italic>
        </p>
      </abstract>
      <kwd-group>
        <kwd>Gaze Estimation</kwd>
        <kwd>Appearance-based</kwd>
        <kwd>End-to-End</kwd>
        <kwd>Convolutional Neural Networks (CNNs)</kwd>
        <kwd>Multitask learning</kwd>
        <kwd>Driver Monitoring System</kwd>
        <kwd>Autonomous Driving</kwd>
      </kwd-group>
    </article-meta>
  </front>
</article>
