<?xml version="1.0"?>
                <!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "journalpublishing3.dtd">
                <article article-type="research-article" xmlns:mml="http://www.w3.org/1998/Math/MathML"
                xmlns:xlink="http://www.w3.org/1999/xlink"
                xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
                dtd-version="3.0">
                <front>
                    <journal-meta>
                    <journal-id journal-id-type="publisher-id">ei</journal-id>
                    <journal-title>Electronic Imaging</journal-title>
                    <issn pub-type="ppub">2470-1173</issn><issn pub-type="epub">2470-1173</issn>
                    <publisher>
                        <publisher-name>Society for Imaging Science and Technology</publisher-name>
                        <publisher-loc>IS&amp;T 7003 Kilworth Lane, Springfield, VA 22151 USA</publisher-loc>
                    </publisher>
                    </journal-meta>
                    <article-meta>
                    <article-id pub-id-type="doi">10.2352/EI.2024.36.17.AVM-111</article-id>
                    <article-id pub-id-type="publisher-id">AVM-111</article-id>
                    <article-categories>
                        <subj-group>
                        <subject>Proceedings</subject>
                        </subj-group>
                    </article-categories>
                    <title-group>
                        <article-title>Multi-Modal Pedestrian Detection Via Dual-Regressor and Object-Based Training for One-Stage Object Detection Network</article-title>
                    </title-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Wanchaitanawong</surname>
                            <given-names>Napat </given-names>
                           </name> <xref ref-type="aff" rid="aff1author1"/></contrib><aff id="aff1author1">Tokyo Institute of Technology, Japan</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Tanaka</surname>
                            <given-names>Masayuki </given-names>
                           </name> <xref ref-type="aff" rid="aff1author2"/></contrib><aff id="aff1author2">Tokyo Institute of Technology, Japan</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Shibata</surname>
                            <given-names>Takashi </given-names>
                           </name> <xref ref-type="aff" rid="aff2author3"/></contrib><aff id="aff2author3"> NTT Corporation,  Japan</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Okutomi</surname>
                            <given-names>Masatoshi </given-names>
                           </name> <xref ref-type="aff" rid="aff1author4"/></contrib><aff id="aff1author4">Tokyo Institute of Technology, Japan</aff></contrib-group><abstract>
                    <title>Abstract</title>
                    <p>Multi-modal pedestrian detection has been developed actively in the research field for the past few years. Multi-modal pedestrian detection with visible and thermal modalities outperforms visible-modal pedestrian detection by improving robustness to lighting effects and cluttered backgrounds because it can simultaneously use complementary information from visible and thermal frames. However, many existing multi-modal pedestrian detection algorithms assume that image pairs are perfectly aligned across those modalities. The existing methods often degrade the detection performance due to misalignment. This paper proposes a multi-modal pedestrian detection network for a one-stage detector enhanced by a dual-regressor and a new algorithm for learning multi-modal data, so-called object-based training. This study focuses on Single Shot MultiBox Detector (SSD), one of the most common one-stage detectors. Experiments demonstrate that the proposed method outperforms current state-of-the-art methods on artificial data with large misalignment and is comparable or superior to existing methods on existing aligned datasets.</p>
                    </abstract><pub-date>
                        <day>21</day>
                        <month>01</month>
                        <year>2024</year>
                        </pub-date><volume>36</volume>
                    <issue-acronym>AVM</issue-acronym>
                    <issue-title>Autonomous Vehicles and Machines 2024</issue-title>
                    <issue seq="111">17</issue>
                    <fpage>111-1</fpage>
                    <lpage>111-6</lpage>
                    <permissions>
                         <copyright-statement>© 2024, Society for Imaging Science and Technology</copyright-statement>
                        <copyright-year>2024</copyright-year>
                    </permissions><kwd-group><kwd>Convolution neural network</kwd><kwd>Deep-learning</kwd><kwd>Long-wave infrared image</kwd><kwd>Multispectral image</kwd><kwd>Pedestrian detection</kwd></kwd-group></article-meta>
                </front>
                </article>