<?xml version="1.0"?>
                <!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "journalpublishing3.dtd">
                <article article-type="research-article" xmlns:mml="http://www.w3.org/1998/Math/MathML"
                xmlns:xlink="http://www.w3.org/1999/xlink"
                xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
                dtd-version="3.0">
                <front>
                    <journal-meta>
                    <journal-id journal-id-type="publisher-id">ei</journal-id>
                    <journal-title>Electronic Imaging</journal-title>
                    <issn pub-type="ppub">2470-1173</issn><issn pub-type="epub">2470-1173</issn>
                    <publisher>
                        <publisher-name>Society for Imaging Science and Technology</publisher-name>
                        <publisher-loc>IS&amp;T 7003 Kilworth Lane, Springfield, VA 22151 USA</publisher-loc>
                    </publisher>
                    </journal-meta>
                    <article-meta>
                    <article-id pub-id-type="doi">10.2352/EI.2023.35.4.MWSF-372</article-id>
                    <article-id pub-id-type="publisher-id">MWSF-372</article-id>
                    <article-categories>
                        <subj-group>
                        <subject>Article</subject>
                        </subj-group>
                    </article-categories>
                    <title-group>
                        <article-title>Synthetic speech attribution using self supervised audio spectrogram transformer</article-title>
                    </title-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Yadav</surname>
                            <given-names>Amit Kumar Singh </given-names>
                           </name> <xref ref-type="aff" rid="aff1author1"/></contrib><aff id="aff1author1">Purdue University, United States</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Bartusiak</surname>
                            <given-names>Emily R.</given-names>
                           </name> <xref ref-type="aff" rid="aff1author2"/></contrib><aff id="aff1author2">Purdue University, United States</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Bhagtani</surname>
                            <given-names>Kratika </given-names>
                           </name> <xref ref-type="aff" rid="aff1author3"/></contrib><aff id="aff1author3">Purdue University, United States</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Delp</surname>
                            <given-names>Edward J.</given-names>
                           </name> <xref ref-type="aff" rid="aff1author4"/></contrib><aff id="aff1author4">Purdue University, United States</aff></contrib-group><abstract>
                    <title>Abstract</title>
                    <p>The ability to synthesize convincing human speech has become easier due to the availability of speech generation tools. This necessitates the development of forensics methods that can authenticate and attribute speech signals. In this paper, we examine a speech attribution task, which identifies the origin of a speech signal. Our proposed method known as Synthetic Speech Attribution Transformer (SSAT) converts speech signals into mel spectrograms and uses a self-supervised pretrained transformer for attribution. This transformer is pretrained on two large publicly available audio datasets: Audio Set and LibriSpeech. We finetune the pretrained transformer on three speech attribution datasets: the DARPA SemaFor Audio Attribution dataset, the ASVspoof2019 dataset, and the 2022 IEEE SP Cup dataset. SSAT achieves high closed-set accuracy on all datasets (99.8% on ASVspoof2019 dataset, 96.3% on SP Cup dataset, and 93.4% on DARPA SemaFor Audio Attribution dataset). We also investigate the methodâ€™s ability to generalize to unknown speech generation methods (open-set scenario). SSAT has high performance, achieving an open-set accuracy of 90.2% on the ASVspoof2019 dataset and 88.45% on DARPA SemaFor Audio Attribution dataset. Finally, we show that our approach is robust to typical compression rates used by YouTube for speech signals.</p>
                    </abstract><pub-date>
                        <day>16</day>
                        <month>1</month>
                        <year>2023</year>
                        </pub-date><volume>35</volume>
                    <issue-acronym>MWSF</issue-acronym>
                    <issue-title>Media Watermarking, Security, and Forensics 2023</issue-title>
                    <issue seq="372">4</issue>
                    <fpage>372-1</fpage>
                    <lpage>372-11</lpage>
                    <permissions>
                         <copyright-statement>© 2023, Society for Imaging Science and Technology</copyright-statement>
                        <copyright-year>2023</copyright-year>
                    </permissions><kwd-group><kwd>speech forensics</kwd><kwd>synthesized speech</kwd><kwd>attribution</kwd><kwd>transformers</kwd><kwd>mel spectrograms</kwd><kwd>machine learning</kwd><kwd>deep learning</kwd><kwd>media forensics</kwd></kwd-group></article-meta>
                </front>
                </article>