<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML"
         xmlns:xlink="http://www.w3.org/1999/xlink"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         article-type="research-article"
         dtd-version="3.0"><front>
      <journal-meta>
         <journal-id journal-id-type="publisher-id">jpi</journal-id>
         <journal-title-group>
            <journal-title>Journal of Perceptual Imaging</journal-title>
            <abbrev-journal-title abbrev-type="IST">J. Percept. Imaging</abbrev-journal-title>
            <abbrev-journal-title abbrev-type="publisher">J. Percept. Imaging</abbrev-journal-title>
         </journal-title-group>
         <issn pub-type="epub">2575-8144</issn>
         <publisher>
            <publisher-name>Society for Imaging Science and Technology</publisher-name>
         </publisher>
      </journal-meta>
      <article-meta>
         <article-id pub-id-type="publisher-id">000502</article-id>
         <article-id pub-id-type="doi">10.2352/J.Percept.Imaging.2022.5.000502</article-id>
         <article-id pub-id-type="manuscript">0156</article-id>
         <article-categories><subj-group subj-group-type="article-type"><subject>Regular Article</subject></subj-group>
         </article-categories>
         <title-group>
            <article-title>Controllable Medical Image Generation via GAN</article-title>
            <alt-title alt-title-type="short">Controllable medical image generation via GAN</alt-title>
         </title-group>
         <contrib-group content-type="all">
            <contrib contrib-type="author">
               <name>
                  <surname>Ren</surname>
                  <given-names>Zhihang</given-names>
               </name>
               <xref ref-type="aff" rid="jpi0156af1"/>
               <xref ref-type="aff" rid="jpi0156af2"/>
               <xref ref-type="aff" rid="jpi0156em1"/>
            </contrib>
            <contrib contrib-type="author">
               <name>
                  <surname>Yu</surname>
                  <given-names>Stella X.</given-names>
               </name>
               <xref ref-type="aff" rid="jpi0156af1"/>
               <xref ref-type="aff" rid="jpi0156af2"/>
            </contrib>
            <contrib contrib-type="author">
               <name>
                  <surname>Whitney</surname>
                  <given-names>David</given-names>
               </name>
               <xref ref-type="aff" rid="jpi0156af1"/>
               <xref ref-type="aff" rid="jpi0156af2"/>
               <xref ref-type="aff" rid="jpi0156af3"/>
               <xref ref-type="aff" rid="jpi0156af4"/>
            </contrib>
            <aff id="jpi0156af1">Vision Science Graduate Group, University of California, Berkeley, CA 94720, United States of America</aff>
            <aff id="jpi0156af2">International Computer Science Institute, Berkeley, CA 94720, United States of America</aff>
            <aff id="jpi0156af3">Department of Psychology, University of California, Berkeley, CA 94720, United States of America</aff>
            <aff id="jpi0156af4">Helen Wills Neuroscience Institute, University of California, Berkeley, CA 94720, United States of America</aff>
            <ext-link id="jpi0156em1" ext-link-type="email">peter.zhren@berkeley.edu</ext-link>
            <author-comment content-type="short-author-list">
               <p>Ren, Yu, and Whitney</p>
            </author-comment>
         </contrib-group>
         <pub-date pub-type="ppub">
            <month>03</month>
            <year>2022</year>
         </pub-date>
         <volume>5</volume>
         <issue seq="2">0</issue>
         <fpage>000502-1</fpage>
         <lpage>000502-15</lpage>
         <history>
            <date date-type="received">
               <day>15</day>
               <month>6</month>
               <year>2021</year>
            </date>
            <date date-type="accepted">
               <day>10</day>
               <month>1</month>
               <year>2022</year>
            </date>
         </history>
         <permissions>
            <copyright-statement>&#x00A9; Society for Imaging Science and Technology 2022</copyright-statement>
            <copyright-year>2022</copyright-year>
         </permissions>
         <abstract>
            <title>Abstract</title>
            <p>Medical image data is critically important for a range of disciplines, including medical image perception research, clinician training programs, and computer vision algorithms, among many other applications. Authentic medical image data, unfortunately, is relatively scarce for many of these uses. Because of this, researchers often collect their own data in nearby hospitals, which limits the generalizabilty of the data and findings. Moreover, even when larger datasets become available, they are of limited use because of the necessary data processing procedures such as de-identification, labeling, and categorizing, which requires significant time and effort. Thus, in some applications, including behavioral experiments on medical image perception, researchers have used naive artificial medical images (e.g., shapes or textures that are not realistic). These artificial medical images are easy to generate and manipulate, but the lack of authenticity inevitably raises questions about the applicability of the research to clinical practice. Recently, with the great progress in Generative Adversarial Networks (GAN), authentic images can be generated with high quality. In this paper, we propose to use GAN to generate authentic medical images for medical imaging studies. We also adopt a controllable method to manipulate the generated image attributes such that these images can satisfy any arbitrary experimenter goals, tasks, or stimulus settings. We have tested the proposed method on various medical image modalities, including mammogram, MRI, CT, and skin cancer images. The generated authentic medical images verify the success of the proposed method. The model and generated images could be employed in any medical image perception research.</p>
         </abstract>
         <counts>
            <page-count count="15"/>
         </counts>
         <custom-meta-group>
            <custom-meta>
               <meta-name>ccc</meta-name>
               <meta-value>2575-8144/2022/5/000502/15/$00.00</meta-value>
            </custom-meta>
            <custom-meta>
               <meta-name>printed</meta-name>
               <meta-value>Printed in the USA</meta-value>
            </custom-meta>
         </custom-meta-group>
      </article-meta>
   </front>
   <body><sec id="jpi0156us1">
         <label>1.</label>
         <title>Introduction</title>
         <p>Medical imaging has transformed modern medicine, allowing clinicians to noninvasively examine and diagnose patients with remarkable ease and speed. In recent years, there have been dramatic advances in the field of medical imaging technologies, ranging from MRI, CT, PET, photography, ultrasound, among many other techniques. These improvements are astounding, but it is worth noting that ultimately the data provided by these techniques requires critical human involvement in detection, selection, interpretation, and diagnosis. The imaging techniques themselves are not the only bottleneck for obtaining accurate diagnoses.</p>
         <p>Fortunately, along with the technological developments, there have also been concomitant advances in the application and use of these technologies. For instance, there is a recent surge in computer vision and medical image perception research, that require artificial (algorithmic) and human users respectively. In both machines and humans, there is a great deal of potential to improve the use of medical imaging in clinical practice. In addition to the more ambitious goals of automated diagnoses, filtering, or cuing clinicians&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib34">34</xref>, <xref ref-type="bibr" rid="jpi0156bib51">51</xref>, <xref ref-type="bibr" rid="jpi0156bib68">68</xref>, <xref ref-type="bibr" rid="jpi0156bib86">86</xref>], there are distinct and more pressing goals of improving clinicians&#x2019; medical image perception and decision-making&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib74">74</xref>, <xref ref-type="bibr" rid="jpi0156bib78">78</xref>, <xref ref-type="bibr" rid="jpi0156bib81">81</xref>] in the realms of training, error detection, diagnostic support, among others&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib79">79</xref>].</p>
         <p>To improve both machine and human medical image perception, it is necessary to have sufficient source data. Unfortunately, labeled and de-identified public medical imaging data is scarce. Sometimes researchers resort to collecting their own data from nearby hospitals, usually from local areas that cannot represent the broader population. Second, even if larger datasets are collected, the necessary data processing procedures such as data de-identification, labeling, and categorizing requires significant time and effort. For instance, in certain medical imaging tasks, such as lesion segmentation, in order to prepare the training data, it requires experts to perform meticulous annotations that are tedious and labor intensive&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib80">80</xref>]. Moreover, collected medical images are specific to each individual patient and it can be difficult to find specific images or image properties that satisfy certain desired experimental configurations&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib52">52</xref>]. Of course, due to intricate tissue structures, manipulating attributes of those collected medical images using traditional image processing methods is difficult or impossible, at least in a realistic manner.</p>
         <p>The data scarcity problem has presented a major challenge to research on medical image perception. At a broad level, medical image perception research studies the visual and cognitive processes that clinicians rely on to make decisions. As in other domains of human factors, the goal of understanding those mechanisms is to improve (i.e., guide, cue, facilitate, speed, etc) clinician performance. Recently, in many psychophysical experiments, artificial medical stimuli have been employed&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib46">46</xref>, <xref ref-type="bibr" rid="jpi0156bib52">52</xref>]. The artificial medical stimuli are often composed of simple shapes or textures with some form of noise background&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib46">46</xref>, <xref ref-type="bibr" rid="jpi0156bib52">52</xref>, <xref ref-type="bibr" rid="jpi0156bib77">77</xref>]. Related approaches involve using real medical images but superimposing clearly artificial &#x201C;targets&#x201D;&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib46">46</xref>, <xref ref-type="bibr" rid="jpi0156bib52">52</xref>]. An advantage of these approaches is that they are relatively easy to generate and control in a precise manner, which is important for studying the cognitive and perceptual systems of clinicians&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib46">46</xref>, <xref ref-type="bibr" rid="jpi0156bib52">52</xref>]. For example, the image attributes and &#x201C;targets&#x201D; are easy to manipulate such that researchers can perform shape morphing and background replacement. This level of stimulus control is necessary in perception research to study things like visual search for lesions, visual recognition of lesions, inattentional blindness, cognitive load and interference, etc. However, those artificial medical images are obviously inauthentic, completely unlike what clinicians routinely examine. Thus, the results of these experiments fall invariably within a shadow of a doubt about clinical applicability.</p>
         <p>Therefore, generating authentic and easily controllable medical images is critical for the entire field of medical image perception research. Alleviating constraints is only recently realistic, with the impressive development of deep learning in computer vision. For example, Generative Adversarial Network (GAN) is one of the promising models that have achieved great success on image generation tasks. GAN can generate high-quality authentic images with various categories&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib41">41</xref>, <xref ref-type="bibr" rid="jpi0156bib42">42</xref>, <xref ref-type="bibr" rid="jpi0156bib60">60</xref>], such as faces, cars, landscapes, and so on. Additionally, various methods can be applied to manipulate the attributes of Generative Adversarial Networks&#x2019; outputs&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib13">13</xref>, <xref ref-type="bibr" rid="jpi0156bib54">54</xref>, <xref ref-type="bibr" rid="jpi0156bib60">60</xref>].</p>
         <p>In this paper, we utilize Generative Adversarial Network (GAN) to generate authentic medical images (Figure&#x00A0;<xref ref-type="fig" rid="jpi0156fig1">1</xref>(a)). We also adopt a controllable approach to manipulate specific attributes of the generated images (Fig.&#x00A0;<xref ref-type="fig" rid="jpi0156fig1">1</xref>(b)). The proposed method is tested on various medical image modalities such as mammogram, MRI, CT, and skin cancer images. For example, via controllable generation, we can create authentic mammograms with desired tumor and breast shapes. We also recruited both expert clinicians and untrained participants to discriminate the authenticity of each image (real versus GAN generated) in an objective psychophysical experiment. Finally, we investigate the perceptual loss which is utilized in the controllable generation. Various experiments verify the success of the proposed controllable medical image generation model.</p>
         <fig id="jpi0156fig1"><label>Figure&#x00A0;1.</label>
            <caption id="jpi0156fc1">
               <p>Pipeline. Controllable medical image generation using the proposed GAN model. (a) Medical image generation: novel and authentic medical images can be generated from random latent codes <italic>z</italic>. (b) Attribute manipulation: desired attributes can be assembled together to satisfy certain experimental settings. Here, we use mammogram as an example medical modality. Real mammograms with tumor were utilized to train the proposed model. Our proposed model can be easily adapted to other medical modalities, such as MRI, CT, and skin cancer images.</p>
            </caption>
            <graphic id="jpi0156f1_online" content-type="online" xlink:href="jpi0156f1_online.jpg"/>
         </fig><p>Contributions: We propose a framework for controllable medical image generation with the following contributions.</p>
         <list id="jpi0156l1" list-type="bullet" prefix-word="bullet">
            <list-item id="jpi0156l1.1">
               <label>&#x2219;</label>
               <p>We propose to utilize Generative Adversarial Network (GAN) to generate medical images and verify the results on various medical image modalities such as mammogram, MRI, CT, and skin cancer images.</p>
            </list-item>
            <list-item id="jpi0156l1.2">
               <label>&#x2219;</label>
               <p>We adopt a controllable approach to manipulate the attributes of the generated images in order to meet certain experimental configurations.</p>
            </list-item>
            <list-item id="jpi0156l1.3">
               <label>&#x2219;</label>
               <p>We compare traditional similarity measurements with the perceptual metric in medical imaging.</p>
            </list-item>
         </list>
         <p>Although a shorter conference version of this paper appeared in&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib66">66</xref>], it was limited in scope and did not extend the model to multiple medical image modalities. This paper extends the model to MRI, CT, and skin cancer images. Moreover, this paper compares traditional similarity measurements with the perceptual metric in medical imaging.</p>
         <p></p>
      </sec>
      <sec id="jpi0156us2">
         <label>2.</label>
         <title>Related Work</title>
         <sec id="jpi0156us2-1">
            <label>2.1</label>
            <title>Convolutional Neural Networks</title>
            <p>The idea of Convolutional Neural Networks (CNN) stem from the discovery of the edge detector in cat&#x2019;s striate cortex&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib38">38</xref>]. Based on this finding, Fukushima&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib23">23</xref>] invented the first simple hierarchical, multilayered artificial neural network. After decades of development, LeCun et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib49">49</xref>] leveraged CNN for hand-written ZIP Code numbers recognition and trained the network end-to-end via gradient descent. This fully automatic image recognition model can be applied to many image categories and types. The great success is mainly attributed to the convolution operation, which can reveal the latent semantic information of an image, and the shared hierarchical kernels, which make the convolution shift-invariant. During training, the loss is computed based on specific metrics for certain tasks, updating the model parameters while it back propagates through the whole network.</p>
            <p>However, the computation is heavy, which limits the model&#x2019;s capacity and ability for high-resolution images. With the deployment of Graphical Processing Unit (GPU), CNNs&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib33">33</xref>, <xref ref-type="bibr" rid="jpi0156bib47">47</xref>, <xref ref-type="bibr" rid="jpi0156bib71">71</xref>&#x2013;<xref ref-type="bibr" rid="jpi0156bib73">73</xref>] have shown promise in computer vision tasks, such as image classification&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib33">33</xref>], object detection&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib27">27</xref>, <xref ref-type="bibr" rid="jpi0156bib28">28</xref>, <xref ref-type="bibr" rid="jpi0156bib64">64</xref>, <xref ref-type="bibr" rid="jpi0156bib65">65</xref>], and object segmentation&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib32">32</xref>]. Recently, many medical imaging tasks have been utilizing CNNs&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib22">22</xref>, <xref ref-type="bibr" rid="jpi0156bib44">44</xref>, <xref ref-type="bibr" rid="jpi0156bib70">70</xref>]. Compared to traditional image processing methods, CNNs have much better performance with much faster inference speed.</p>
         </sec>
         <sec id="jpi0156us2-2">
            <label>2.2</label>
            <title>Generative Adversarial Networks</title>
            <p>Generative Adversarial Networks are special Convolutional Neural Networks, which consist of two networks, the generator (G) and the discriminator&#x00A0;(D). These two networks are trained iteratively in an adversarial way where the generator (G) generates fake but authentic images to fool the discriminator and the discriminator (D) discriminates the real and fake images&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib29">29</xref>]. Using this promising computational model, high-quality images with various categories can be generated, such as faces, cars, and landscapes&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib41">41</xref>, <xref ref-type="bibr" rid="jpi0156bib42">42</xref>, <xref ref-type="bibr" rid="jpi0156bib60">60</xref>]. However, the initial GAN model&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib29">29</xref>] cannot generate sharp and recognizable images, and the training process is unstable. Later work improved the performance of GAN in different ways. Some papers focus on model architectures&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib13">13</xref>, <xref ref-type="bibr" rid="jpi0156bib54">54</xref>, <xref ref-type="bibr" rid="jpi0156bib58">58</xref>]. Others focus on improving the loss metrics and training strategies&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib2">2</xref>, <xref ref-type="bibr" rid="jpi0156bib9">9</xref>, <xref ref-type="bibr" rid="jpi0156bib30">30</xref>]. With these efforts, GAN training stability has improved, and GAN can generate low-resolution images with sufficient quality.</p>
            <p>Of late, numerous approaches for high-resolution image generation are also available. PGGAN&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib41">41</xref>] aims to train the standard GAN from coarse to fine scale. The parameters for low-resolution block are trained first. Then higher-resolution blocks are added on gradually with the corresponding parameters updated accordingly. Based on the same training strategy, StyleGAN&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib42">42</xref>, <xref ref-type="bibr" rid="jpi0156bib43">43</xref>] proposed to first map the original latent space <inline-formula><mml:math><mml:mrow><mml:mi mathvariant="script">Z</mml:mi></mml:mrow></mml:math></inline-formula> into the <inline-formula><mml:math><mml:mrow><mml:mi mathvariant="script">W</mml:mi></mml:mrow></mml:math></inline-formula> space through a non-linear mapping network. Then it is merged into the synthesis network via adaptive instance normalization (AdaIN) at each convolutional block&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib17">17</xref>, <xref ref-type="bibr" rid="jpi0156bib36">36</xref>]. This improves StyleGAN representations of scenes and details and allows it to produce authentic high-resolution images. In this paper, we adopt StyleGAN as our backbone for medical images generation. Moreover, a controllable approach is also utilized to manipulate the attributes of the generated images.</p>
            <p>In medical image applications, [<xref ref-type="bibr" rid="jpi0156bib22">22</xref>] DCGAN&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib61">61</xref>] and ACGAN [<xref ref-type="bibr" rid="jpi0156bib58">58</xref>] were utilized to generate CT liver lesion patches and boosted the liver lesion classification performance. Han et al.&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib31">31</xref>] deployed WGAN&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib30">30</xref>] to generate MR images for data augmentation and physician training. Nie et al.&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib57">57</xref>] used GAN to predict CT images from MR images. Cao et al.&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib10">10</xref>] proposed an Auto-GAN to synthesize missing modality for medical images. Moreover, GAN has been widely used for skin cancer image generation and purification&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib5">5</xref>&#x2013;<xref ref-type="bibr" rid="jpi0156bib7">7</xref>, <xref ref-type="bibr" rid="jpi0156bib26">26</xref>]. Our approach is different from aforementioned methods. In addition to purely generating new samples as GANs traditionally do, our method can also edit specific images via the encoder of our model.</p>
         </sec>
         <sec id="jpi0156us2-3">
            <label>2.3</label>
            <title>Perceptual Loss</title>
            <p>CNN features have already been utilized for calculating similarity for years. Ref.&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib1">1</xref>] proposed to use pre-trained AlexNet features for image quality measurement. Perceptual loss, which is also based on CNN features, was first proposed in Ref.&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib40">40</xref>] for style transfer&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib25">25</xref>] and super resolution tasks. Both are ill-posed problems. For style transfer, there is no absolute ground truth image for reference. For image super resolution, one low-resolution image can have many corresponding high-resolution images which can be down-sampled to the same low-resolution image. Thus, per-pixel metric is no longer suitable since semantic similarity matters. Recently, traditional similarity metrics, such as Structural Similarity Index Measure&#x00A0;(SSIM) and Peak Signal-to-Noise Ratio (PSNR), are found to be inconsistent with human perception, and a perceptual metric has been utilized to measure the semantic similarity in many papers&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib37">37</xref>, <xref ref-type="bibr" rid="jpi0156bib50">50</xref>, <xref ref-type="bibr" rid="jpi0156bib87">87</xref>, <xref ref-type="bibr" rid="jpi0156bib89">89</xref>]. In this paper, we use perceptual loss to regularize the encoder training and guide the latent code optimization in the encoding procedure.</p>
            <p></p>
         </sec>
      </sec>
      <sec id="jpi0156us3">
         <label>3.</label>
         <title>Method</title>
         <p>Here, we adapt the Generative Adversarial Network for medical image generation. In order to manipulate the image attributes, an encoder is added to encode certain image attributes into the latent code <italic>z</italic> which is the input of the GAN generator.</p>
         <p>Our proposed model is composed of two parts. The first part is the GAN which involves the generator (G) and the discriminator (D). The generator (G) will generate authentic (fake) images from the latent codes <italic>z</italic>, and try to fool the discriminator (D) during training. The discriminator (D) will discriminate whether the image is real (i.e. sampled from real images) or fake (i.e. generated from the generator), and try to beat the generator by distinguishing the fake images from the real ones. The second part of the model is the encoder (E), which can encode image attributes into the latent code <italic>z</italic>. This latent code can then be utilized to generate the image through the generator. Therefore, it can allow us to manipulate the generated image by manipulating the latent code through the encoder. The architecture is shown in Figure&#x00A0;<xref ref-type="fig" rid="jpi0156fig2">2</xref>.</p>
         <fig id="jpi0156fig2"><label>Figure&#x00A0;2.</label>
            <caption id="jpi0156fc2">
               <p>Architecture of proposed method. The architecture contains three sub-networks, the encoder (E), the generator (G), and the discriminator (D). The training has two phases. In the first phase, the generator and discriminator will be trained first without the encoder (E) via adversarial loss <italic>L</italic><sub>adversarial</sub>. In the second phase, the generator (G) will be fixed. The encoder (E) and discriminator (D) will be trained adversarially via the reconstruction loss <italic>L</italic><sub>reconstruction</sub>, the perceptual loss <italic>L</italic><sub>perceptual</sub>, and the adversarial loss <italic>L</italic><sub>adversarial</sub>. The dashed arrows indicate how to compute the corresponding loss metrics.</p>
            </caption>
            <graphic id="jpi0156f2_online" content-type="online" xlink:href="jpi0156f2_online.jpg"/>
         </fig><p>While training, the GAN part is first trained progressively&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib42">42</xref>] via adversarial loss <italic>L</italic><sub>Adversarial</sub>. The training process can be formulated as <disp-formula id="jpi0156eqn1"><label>(1)</label><mml:math><mml:mrow><mml:msub><mml:mrow><mml:mo>min</mml:mo></mml:mrow><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mo>max</mml:mo></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mtext>data</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>log</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>z</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>log</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula> where <italic>p</italic><sub>data</sub>(<italic>x</italic>) and <italic>q</italic>(<italic>z</italic>) indicate the real data distribution and the latent space distribution respectively, <italic>x</italic> is the sampled real image, <italic>z</italic> is the sampled latent code.</p>
         <p>Then, we train the encoder part. After training the GAN part, the generator (G) is fixed. While training the encoder network, traditional methods&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib3">3</xref>] regularize the encoder on the latent space, encouraging the encoder to encode the same latent codes for the corresponding generated images regardless of the reconstructed images. This method can degrade the reconstruction quality. Instead, we adopt the idea from In-domain GAN inversion&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib88">88</xref>], where the regularization of the encoder is on the image space. In particular, the encoded vector is passed into the generator (G) again and the regularization is on the reconstructed image. The L2 reconstruction loss <italic>L</italic><sub>Reconstruction</sub> and the perceptual loss&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib40">40</xref>] <italic>L</italic><sub>Perceptual</sub> are utilized for the regularization. Additionally, adversarial loss <italic>L</italic><sub>Adversarial</sub> is also utilized to guarantee that the reconstructed image looks authentic. The whole process can be summarized as follows <disp-formula id="jpi0156eqn2"><label>(2)</label><mml:math><mml:mrow><mml:mtable><mml:mtr><mml:mtd/><mml:mtd/><mml:mtd><mml:msub><mml:mrow><mml:mo>min</mml:mo></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2225;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2225;</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd/><mml:mtd><mml:mspace width="2em"/><mml:mo>&#x2212;</mml:mo><mml:mspace width="1em"/><mml:msub><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mtext>data</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>log</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula> <disp-formula id="jpi0156eqn3"><label>(3)</label><mml:math><mml:mrow><mml:mtable><mml:mtr><mml:mtd/><mml:mtd/><mml:mtd><mml:msub><mml:mrow><mml:mo>min</mml:mo></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mtext>data</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>log</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mtext>data</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>log</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd/><mml:mtd><mml:mspace width="2em"/><mml:mo>+</mml:mo><mml:mspace width="1em"/><mml:mfrac><mml:mrow><mml:mi>&#x03B3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mtext>data</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>&#x2225;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x2207;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula> where <italic>p</italic><sub>data</sub>(<italic>x</italic>) indicates the real data distribution, <italic>x</italic> is the real image, <italic>E</italic> represents the encoder, <italic>F</italic> represents the VGG feature extraction&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib71">71</xref>], and <italic>&#x03BB;</italic><sub>1</sub>, <italic>&#x03BB;</italic><sub>2</sub> and <italic>&#x03B3;</italic> are weights for the perceptual loss, the adversarial loss, and the gradient penalty&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib30">30</xref>].</p>
         <p>Since the inverse mapping via the encoder (E) will not always be perfect, in order to get the optimal inverse latent code, we apply another optimization on the latent code. This optimization will update the latent code based on the reconstruction loss and the perceptual loss within the neighborhood of the original encoded vector (the encoder regularization). The optimization process can be described as below <disp-formula id="jpi0156eqn4"><label>(4)</label><mml:math><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mtext>inv</mml:mtext></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mo>min</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2225;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2225;</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd/><mml:mtd><mml:mspace width="2em"/><mml:mo>+</mml:mo><mml:mspace width="1em"/><mml:msub><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2225;</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula> where <italic>z</italic><sup>inv</sup> is the optimized inverse code, <italic>&#x03BB;</italic><sub>3</sub> and <italic>&#x03BB;</italic><sub>4</sub> are weights for the perceptual loss, and the code reconstruction loss (i.e., the encoder regularization). This optimization metric can be computed using the whole image region (for image reconstruction) or the region of interest (for image manipulation).</p>
         <sec id="jpi0156us3-1">
            <label>3.1</label>
            <title>Medical Image Synthesis</title>
            <p>In general, informative images lie on a manifold. Through the GAN training, the generator (G) learns a transformation from the latent space to the image space, imitating the real image manifold of the training dataset. Thus, we can utilize this learned transformation to generate images authentic to the real images. First, the latent code <italic>z</italic> will be sampled from the latent space. Then, the output image <italic>x</italic> = <italic>G</italic>(<italic>z</italic>) is produced by the generator.</p>
            <p>Using the learned transformation, we can also generate similar medical images. As a manifold, the nearby images on the manifold are similar to each other. Therefore, we can sample a series of latent codes <italic>z</italic><sub><italic>i</italic></sub> on a closed path <italic>C</italic>, then passing these latent codes into the generator (G), we can obtain a series of gradually and continuously morphing images <italic>x</italic><sub><italic>i</italic></sub>: <disp-formula id="jpi0156eqn5"><label>(5)</label><mml:math><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mi>C</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
         </sec>
         <sec id="jpi0156us3-2">
            <label>3.2</label>
            <title>Attribute Manipulation</title>
            <p>While training the encoder (E), without the discriminator (D), the encoder and the generator form an autoencoder&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib53">53</xref>, <xref ref-type="bibr" rid="jpi0156bib63">63</xref>]. The training encourages the encoder to embed useful image attributes into the latent code. Since the generator is pretrained under the GAN, the generator has learned how to reconstruct the embedded image attributes with proper tissue context.</p>
            <p>In order to manipulate the image attributes, we first need to combine the desired image attributes into one assembled image <italic>x</italic><sup><italic>&#x2032;</italic></sup>. The combination can be achieved by merging image patches <italic>P</italic><sub><italic>i</italic></sub> which contain the desired image attributes: <disp-formula id="jpi0156eqn6"><label>(6)</label><mml:math><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo mathsize="big"> &#x22C3;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
            <p>Then this assembled image <italic>x</italic><sup><italic>&#x2032;</italic></sup> will be encoded by the encoder, <italic>z</italic><sup><italic>&#x2032;</italic></sup> = <italic>E</italic>(<italic>x</italic><sup><italic>&#x2032;</italic></sup>), obtaining the corresponding image attributes latent code <italic>z</italic><sup><italic>&#x2032;</italic></sup>. The generator will finally reconstruct those image attributes with proper tissue texture, <italic>x</italic><sub>reconstruct</sub> = <italic>G</italic>(<italic>z</italic><sup><italic>&#x2032;</italic></sup>).</p>
            <p>Since the image with all desired attributes may not exist on the image manifold, the reconstructed image may not have the exact desired attributes as we designed. The final optimization (shown in Eq.&#x00A0;(<xref ref-type="disp-formula" rid="jpi0156eqn4">4</xref>)) can be conducted on the region where the attributes need to be accurate. The pipeline for attribute manipulation is shown in Figure&#x00A0;<xref ref-type="fig" rid="jpi0156fig3">3</xref>.</p>
            <fig id="jpi0156fig3"><label>Figure&#x00A0;3.</label>
               <caption id="jpi0156fc3">
                  <p>Attribute manipulation pipeline. Firstly, the desired image attributes are combined by merging image patches that contain those attributes. Then, the corresponding latent code is produced by the encoder. The generator reconstructs the image with desired attributes. Finally, the desired image can be obtained after the final optimization.</p>
               </caption>
               <graphic id="jpi0156f3_online" content-type="online" xlink:href="jpi0156f3_online.jpg"/>
            </fig><p></p>
         </sec>
      </sec>
      <sec id="jpi0156us4">
         <label>4.</label>
         <title>Experiments and Results</title>
         <sec id="jpi0156us4-1">
            <label>4.1</label>
            <title>Implementation Details</title>
            <p>For the Generative Adversarial Network (GAN), we adopt StyleGAN&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib42">42</xref>]. The training is progressive. Starting from 8 &#x00D7; 8, the latter resolution blocks are added progressively after the previous blocks finish training. The output image resolution is 256 &#x00D7; 256. While training the encoder, the generator is fixed. Only the encoder and discriminator parameters are updated. For the perceptual loss, VGG&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib71">71</xref>] <italic>conv</italic>4_3 feature layer is utilized. As for the hyperparameters, <italic>&#x03BB;</italic><sub>1</sub> = 0.00005, <italic>&#x03BB;</italic><sub>2</sub> = 0.1, <italic>&#x03BB;</italic><sub>3</sub> = 0.00005, <italic>&#x03BB;</italic><sub>4</sub> = 2, and <italic>&#x03B3;</italic> = 10. We use the Adam optimizer [<xref ref-type="bibr" rid="jpi0156bib45">45</xref>] with <italic>&#x03B2;</italic><sub>1</sub> = 0.9 and <italic>&#x03B2;</italic><sub>2</sub> = 0.99. The learning rate is set to 0.0001. Pytorch is utilized for coding.</p>
            <p>For mammogram images, we use DDSM&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib8">8</xref>] dataset which contains 2,620 normal, benign, and malignant cases. Only the benign and malignant cases are utilized for training. For MRI images, we utilize fastMRI&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib84">84</xref>] multi-coil dataset which contains 7135 images. For CT images, DeepLesion&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib83">83</xref>] dataset is used. We utilize the abdomen image dataset which contains 14601 images. For skin cancer images, we use images from ISIC Archive (<uri>https://www.isic-archive.com/#!/topWithHeader/wideContentTop/main</uri>) which contains 69445 images in total.</p>
         </sec>
         <sec id="jpi0156us4-2">
            <label>4.2</label>
            <title>GAN Generated Results</title>
            <p>For different medical image modalities, we train the whole network separately using corresponding datasets. After the GAN part has been trained, we randomly sample latent codes <italic>z</italic> and pass them to the generator. The generated results for Mammogram, MRI, CT, and Skin Cancer are shown in Figure&#x00A0;<xref ref-type="fig" rid="jpi0156fig4">4</xref>. Compared to the real samples on the left, the generated samples on the right appear very similar, and this is seen across different medical image modalities. It is clear that the generator has learned the semantic statistics of the training dataset for different medical image modalities. The generator can generate authentic tissue texture, tissue distribution, tissue shapes, and color distributions. Moreover, the generator can not only reconstruct the original medical images, but also it can produce novel and authentic medical images which do not actually exist in the real world.</p>
            <fig id="jpi0156fig4"><label>Figure&#x00A0;4.</label>
               <caption id="jpi0156fc4">
                  <p>GAN generated results. The generated results for different medical image modalities. Comparing the real samples to the generated samples, it is clear that the generator has learned how to imitate tissue texture, tissue distribution, tissue shapes, and color distribution. The generator can produce authentic images (see below for psychophysical results confirming this).</p>
               </caption>
               <graphic id="jpi0156f4_online" content-type="online" xlink:href="jpi0156f4_online.jpg"/>
            </fig><p>Since the GAN training learns the manifold of the training dataset, we can also generate gradually and continuously morphing medical images for certain experiments. First, the latent codes need to be sampled from a closed path in the latent space. To do so, we randomly pick three anchor points in the latent space and calculate the interpolations between each pair of them. Then, passing those codes to the generator, we can obtain the gradually and continuously morphing medical images. The result is shown in Figure&#x00A0;<xref ref-type="fig" rid="jpi0156fig5">5</xref>. Due to the space limit, we only show three interpolations between each pair; arbitrarily fine grained interpolations can be created between any number of pairs.</p>
            <fig id="jpi0156fig5"><label>Figure&#x00A0;5.</label>
               <caption id="jpi0156fc5">
                  <p>Interpolation results. Here, we show a mammogram loop gradually changing among three anchor images. The mammograms between two of the anchor images are generated by passing the interpolated codes of those two anchor images to the trained generator. Any number of interpolated images between any pair of anchors can be created.</p>
               </caption>
               <graphic id="jpi0156f5_online" content-type="online" xlink:href="jpi0156f5_online.jpg"/>
            </fig></sec>
         <sec id="jpi0156us4-3">
            <label>4.3</label>
            <title>Attribute Manipulation</title>
            <p>Our proposed model can generate desired medical images by manipulating the image attributes. For illustration, we show how we generate mammograms with the desired lesion patch and desired breast shapes. The results are shown in Figure&#x00A0;<xref ref-type="fig" rid="jpi0156fig6">6</xref>.</p>
            <fig id="jpi0156fig6"><label>Figure&#x00A0;6.</label>
               <caption id="jpi0156fc6">
                  <p>Attribute manipulation results. The desired image attributes are combined by merging the corresponding image patches (in Column A and B) directly. Then, the encoder will encode the manipulated image attributes, and the generator will produce the output correspondingly. After the final optimization, it is clear that the proposed method can generate the mammograms with the desired lesion texture and breast shape (Column F), compared to the results from the traditional image blending method (Column D) and the proposed method without the final optimization (Column E).</p>
               </caption>
               <graphic id="jpi0156f6_online" content-type="online" xlink:href="jpi0156f6_online.jpg"/>
            </fig><p>First, we combine the desired image attributes, i.e. the lesion patch (Fig.&#x00A0;<xref ref-type="fig" rid="jpi0156fig6">6</xref>A) and shape templates (Fig.&#x00A0;<xref ref-type="fig" rid="jpi0156fig6">6</xref>B), by merging the lesion patch and shape templates directly. Then we encode these intermediate combined images (Fig.&#x00A0;<xref ref-type="fig" rid="jpi0156fig6">6</xref>C) using the encoder and pass the codes to the generator. The reconstructed images from the generator are shown in Fig.&#x00A0;<xref ref-type="fig" rid="jpi0156fig6">6</xref>(E) (without optimization). It is clear that the shapes are already the same as the shape templates and the overall texture is authentic. But the desired lesion texture is not maintained. After the last step of optimization over the lesion patch, as it is shown in Fig.&#x00A0;<xref ref-type="fig" rid="jpi0156fig6">6</xref>(F), the lesion texture is rendered. We also compare the results with the ones produced by a traditional image blending method. As it is shown in Fig.&#x00A0;<xref ref-type="fig" rid="jpi0156fig6">6</xref>(D), the transition region between the lesion texture and the shape template background is not natural. Our proposed method can maintain both the breast shape and the lesion texture while generating authentic tissue texture.</p>
         </sec>
         <sec id="jpi0156us4-4">
            <label>4.4</label>
            <title>Human Evaluation</title>
            <p>To verify the authenticity of the generated images for different medical image modalities, we conducted an online psychophysical experiment, recruiting both untrained participants (i.e. no knowledge of medical imaging) and experts (e.g. radiologists or practicing clinicians who routinely read radiographs).</p>
            <sec id="jpi0156us4-4-1">
               <label>4.4.1</label>
               <title>Participants</title>
               <p>Six untrained observers (3 females, age range: 22&#x2013;25) and seven experts (3 females, age range: 32&#x2013;39) participated in the mammogram online survey. Two experts were excluded from the mammogram online survey (one dropped out and the other gave the same response on every trial). Five untrained observers (3 females, age range: 23&#x2013;25) and seven experts (3 females, age range: 28&#x2013;40) participated in the CT online survey.</p>
               <p>All subjects reported to have normal or corrected-to-normal vision. Participants voluntarily participated and were offered $15 per hour as optional compensation. In our experience, radiologists typically refuse this modest compensation. The experiments were approved by the Institutional Review Board at the University of California, Berkeley. Participants provided informed consent.</p>
            </sec>
            <sec id="jpi0156us4-4-2">
               <label>4.4.2</label>
               <title>Stimuli</title>
               <p>For the mammogram online survey, 50 real mammograms and 50 fake (model generated) images were included. For the CT online survey, 50 real CT images and 50 fake CT images were presented. All the images were randomly selected from the corresponding data pools.</p>
            </sec>
            <sec id="jpi0156us4-4-3">
               <label>4.4.3</label>
               <title>Procedure</title>
               <p>The task was to rate each image from 0 (fake/generated image) to 10 (real image) in the data pool. Each individual image was shown for 5&#x00A0;s, and observers were asked to respond as quickly as possible. The experiment was self-paced, so observers viewed the stimuli as long as they wanted (up to 5&#x00A0;s), and they did not have time limit for giving responses. To ensure that participants did not randomly guess (or lapse), a small number of repetitive trials were also included in the online survey to establish a baseline test-retest reliability estimate. We compute the similarity among those repetitive trials.</p>
            </sec>
            <sec id="jpi0156us4-4-4">
               <label>4.4.4</label>
               <title>Results</title>
               <p>The results for mammogram and CT images in terms of the Receiver Operating Characteristic (ROC) curves are shown in Figure&#x00A0;<xref ref-type="fig" rid="jpi0156fig7">7</xref>. For both untrained participants and radiologists, in the context of mammogram and CT images, their performance curves are near the diagonal (i.e. the chance level performance region), indicating that the generated medical images appeared authentic. The area under the curve (AUC) can also confirm the chance-level performance. The mean AUCs are 0.52 (<italic>p</italic> = 0.395, permutation test) and 0.60 (<italic>p</italic> = 0.126, permutation test) for untrained participants and radiologists respectively in mammogram online survey. The mean AUCs are 0.42 (<italic>p</italic> = 0.888, permutation test) and 0.42 (<italic>p</italic> = 0.844, permutation test) for untrained participants and radiologists respectively in CT online survey. As shown in the permutation tests, the large <italic>p</italic>-values indicate that performance is not statistically different from random performance.</p>
               <fig id="jpi0156fig7"><label>Figure&#x00A0;7.</label>
                  <caption id="jpi0156fc7">
                     <p>Human evaluation results. Participant performance is shown in the Receiver Operating Characteristic (ROC) curves. It is clear that their performance is near chance level (curves near the diagonal region), indicating that the generated medical images are authentic. Here, <italic>P</italic><sub>1</sub> &#x2212; <italic>P</italic><sub><italic>N</italic></sub> and <italic>R</italic><sub>1</sub> &#x2212; <italic>R</italic><sub><italic>N</italic></sub> represent different untrained observers and experts in corresponding experiments.</p>
                  </caption>
                  <graphic id="jpi0156f7_online" content-type="online" xlink:href="jpi0156f7_online.jpg"/>
               </fig><p>Although the observers were not able to accurately discriminate real from fake images, this does not mean that observers randomly responded or failed to pay attention to the task. To confirm this, we calculated the test-retest reliability of each observers responses for repeated images. From the small number of repeated trials, the average test-retest similarity is 0.65, indicating &#x201C;good&#x201D; consistency. For a near-threshold task, the noise ceiling is not 1, and 0.65 is &#x201C;good&#x201D; in the sense that it is statistically reliable and significant&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib14">14</xref>, <xref ref-type="bibr" rid="jpi0156bib20">20</xref>, <xref ref-type="bibr" rid="jpi0156bib21">21</xref>]. The similarity is computed using Sokal-Michene metric&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib85">85</xref>]. It is noteworthy that observers can have high test-retest reliability despite low sensitivity (low AUC). The test-retest reliability indicates that observers tended to make the same judgments in repeated trials: they consistently confused some real (fake) images as being fake (real). This resulted in low sensitivity (low AUC) but consistent responses (&#x201C;good&#x201D; test-retest reliability).</p>
               <p>We have appended the results of MRI and Skin Cancer images in the Appendix&#x00A0;<xref ref-type="app" rid="jpi0156sB">B</xref> to avoid redundancy. Results indicate that the generated medical images appeared authentic.</p>
            </sec>
            <sec id="jpi0156us4-4-5">
               <label>4.4.5</label>
               <title>Limitations</title>
               <p>Online studies have a range of potential limitations&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib4">4</xref>]. However, it has been well documented in the literature that online studies can reveal even very subtle psychophysical phenomena reliably, and these methods are now established&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib15">15</xref>, <xref ref-type="bibr" rid="jpi0156bib62">62</xref>, <xref ref-type="bibr" rid="jpi0156bib67">67</xref>]. In our online experiment, variations in the environment or monitor settings that might occur could add noise to the data, but they wouldn&#x2019;t generate the high test retest reliability we found, or the consistent pattern of results. The growing literature on internet-based psychophysics is consistent with this&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib67">67</xref>]. Moreover, we believe that our data adds a unique perspective on this issue: the advantages of online experiments are pronounced in cases where subjects are rare and/or very expensive to recruit, as is the case with the experienced and highly trained radiologist observers reported here. Future studies should consider online data collection for medical image perception tasks, in order to broaden representation, diversity, and improve sample sizes.</p>
               <p>Another consideration with the experiments here is the images were viewed for a maximum of 5&#x00A0;s. The experiment was self-paced, and the participants could view the images as long as needed to make a choice, but this was limited to 5&#x00A0;s maximum viewing. There are both theoretical and empirical reasons that 5&#x00A0;s is likely to be sufficient for the task (see Appendix&#x00A0;<xref ref-type="app" rid="jpi0156sC">C</xref>), but it is conceivable that performance could change if observers were forced to view the images for prolonged periods of time. Future experiments should therefore examine the temporal integration of the visual processes that contribute to discrimination of near-metameric medical images.</p>
            </sec>
         </sec>
         <sec id="jpi0156us4-5">
            <label>4.5</label>
            <title>Perceptual Loss</title>
            <p>Currently, perceptual loss is utilized as a similarity metric in many computer vision tasks&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib37">37</xref>, <xref ref-type="bibr" rid="jpi0156bib50">50</xref>, <xref ref-type="bibr" rid="jpi0156bib87">87</xref>, <xref ref-type="bibr" rid="jpi0156bib89">89</xref>]. In this section, we investigate the perceptual loss as a similarity metric in medical imaging domain. We compare its results with the results of Structural Similarity Index Measure (SSIM) and Peak Signal-to-Noise Ratio (PSNR), which are two common similarity metrics.</p>
            <p>In the experiment, we utilize random samples from mammogram, MRI, CT, and skin cancer images as reference images (256 &#x00D7; 256). First, we apply traditional image distortions on those reference images, such as Gaussian blur, contrast distortion, geometric distortion, spatial shifting, and spatial rotation. Then, we calculate the similarity measurements for different outputs from traditional image distortions with respect to the reference images. Detailed computation algorithms can be found in Appendix&#x00A0;<xref ref-type="app" rid="jpi0156sA">A</xref>.</p>
            <p>For quantitative comparison, we show the similarity measurement results in the following tables. For the SSIM and PSNR metrics, the larger the measurement is, the more similar it is between the measured image and the reference image (indicating by  <italic>&#x2191;</italic>). For perceptual metric, the smaller the measurement is, the more similar it is between the measured image and the reference image (indicating by  <italic>&#x2193;</italic>). Tables&#x00A0;<xref ref-type="table" rid="jpi0156tabI">I</xref>,&#x00A0;<xref ref-type="table" rid="jpi0156tabII">II</xref>,&#x00A0;<xref ref-type="table" rid="jpi0156tabIII">III</xref> and&#x00A0;<xref ref-type="table" rid="jpi0156tabIV">IV</xref> show the similarity measurements for mammogram, MRI, CT, and skin cancer images respectively.</p>
            <table-wrap id="jpi0156tabI">
               <label>Table&#x00A0;I.</label>
               <caption id="jpi0156tcI">
                  <p>Similarity Measurements for Mammogram Images.</p>
               </caption>
               <table frame="void">
                  <colgroup>
                     <col align="center"/>
                     <col align="center"/>
                     <col align="center"/>
                  </colgroup>
                  <thead>
                     <tr>
                        <th align="center"/>
                        <th align="center">Image 1</th>
                        <th align="center">Image 2</th>
                     </tr>
                  </thead>
                  <tbody>
                     <tr>
                        <td align="center">SSIM <italic>&#x2191;</italic></td>
                        <td align="center"><bold>0.93</bold></td>
                        <td align="center">0.89</td>
                     </tr>
                     <tr>
                        <td align="center">PSNR (dB) <italic>&#x2191;</italic></td>
                        <td align="center"><bold>38.99</bold></td>
                        <td align="center">30.61</td>
                     </tr>
                     <tr>
                        <td align="center">Perceptual <italic>&#x2193;</italic></td>
                        <td align="center">0.96</td>
                        <td align="center"><bold>0.64</bold></td>
                     </tr>
                  </tbody>
               </table>
            </table-wrap><table-wrap id="jpi0156tabII">
               <label>Table&#x00A0;II.</label>
               <caption id="jpi0156tcII">
                  <p>Similarity Measurements for MRI Images.</p>
               </caption>
               <table frame="void">
                  <colgroup>
                     <col align="center"/>
                     <col align="center"/>
                     <col align="center"/>
                  </colgroup>
                  <thead>
                     <tr>
                        <th align="center"/>
                        <th align="center">Image 1</th>
                        <th align="center">Image 2</th>
                     </tr>
                  </thead>
                  <tbody>
                     <tr>
                        <td align="center">SSIM <italic>&#x2191;</italic></td>
                        <td align="center"><bold>0.84</bold></td>
                        <td align="center">0.68</td>
                     </tr>
                     <tr>
                        <td align="center">PSNR (dB) <italic>&#x2191;</italic></td>
                        <td align="center"><bold>34.42</bold></td>
                        <td align="center">33.01</td>
                     </tr>
                     <tr>
                        <td align="center">Perceptual <italic>&#x2193;</italic></td>
                        <td align="center">3.89</td>
                        <td align="center"><bold>2.96</bold></td>
                     </tr>
                  </tbody>
               </table>
            </table-wrap><table-wrap id="jpi0156tabIII">
               <label>Table&#x00A0;III.</label>
               <caption id="jpi0156tcIII">
                  <p>Similarity Measurements for CT Images.</p>
               </caption>
               <table frame="void">
                  <colgroup>
                     <col align="center"/>
                     <col align="center"/>
                     <col align="center"/>
                  </colgroup>
                  <thead>
                     <tr>
                        <th align="center"/>
                        <th align="center">Image 1</th>
                        <th align="center">Image 2</th>
                     </tr>
                  </thead>
                  <tbody>
                     <tr>
                        <td align="center">SSIM <italic>&#x2191;</italic></td>
                        <td align="center"><bold>0.54</bold></td>
                        <td align="center">0.24</td>
                     </tr>
                     <tr>
                        <td align="center">PSNR (dB) <italic>&#x2191;</italic></td>
                        <td align="center"><bold>31.04</bold></td>
                        <td align="center">29.91</td>
                     </tr>
                     <tr>
                        <td align="center">Perceptual <italic>&#x2193;</italic></td>
                        <td align="center">27.77</td>
                        <td align="center"><bold>7.42</bold></td>
                     </tr>
                  </tbody>
               </table>
            </table-wrap><table-wrap id="jpi0156tabIV">
               <label>Table&#x00A0;IV.</label>
               <caption id="jpi0156tcIV">
                  <p>Similarity Measurements for Skin Cancer Images.</p>
               </caption>
               <table frame="void">
                  <colgroup>
                     <col align="center"/>
                     <col align="center"/>
                     <col align="center"/>
                  </colgroup>
                  <thead>
                     <tr>
                        <th align="center"/>
                        <th align="center">Image 1</th>
                        <th align="center">Image 2</th>
                     </tr>
                  </thead>
                  <tbody>
                     <tr>
                        <td align="center">SSIM <italic>&#x2191;</italic></td>
                        <td align="center"><bold>0.87</bold></td>
                        <td align="center">0.74</td>
                     </tr>
                     <tr>
                        <td align="center">PSNR (dB) <italic>&#x2191;</italic></td>
                        <td align="center"><bold>36.26</bold></td>
                        <td align="center">31.58</td>
                     </tr>
                     <tr>
                        <td align="center">Perceptual <italic>&#x2193;</italic></td>
                        <td align="center">3.15</td>
                        <td align="center"><bold>1.99</bold></td>
                     </tr>
                  </tbody>
               </table>
            </table-wrap><p>For qualitative comparison, we compare the similarity measurement between Gaussian blur outputs (Figure&#x00A0;<xref ref-type="fig" rid="jpi0156fig8">8</xref> Image 1 Column) and the outputs from the rest of the traditional image distortions (Fig.&#x00A0;<xref ref-type="fig" rid="jpi0156fig8">8</xref> Image 2 Column). We first asked human participants to give their choices of the image, which was more similar to the reference image. The results are labeled with green check marks as shown in Fig.&#x00A0;<xref ref-type="fig" rid="jpi0156fig8">8</xref>. Then, according to the similarity measurements, we select the images which are preferred by SSIM/PSNR or perceptual loss metric. It is clear that SSIM and PSNR do not conform to human judgements. However, the similarity decisions from the perceptual loss metric are consistent with human judgements. Thus, the perceptual metric is more suitable for the similarity measurement in medical imaging area.</p>
            <fig id="jpi0156fig8"><label>Figure&#x00A0;8.</label>
               <caption id="jpi0156fc8">
                  <p>Which image is more similar to the reference?  Image 1 Column shows the distortion by Gaussian blur. Image 2 Column shows the distortions by contrast distortion, geometric distortion, spatial shifting, and spatial rotation respectively. The human judgements are marked using green ticks. It is clear that SSIM/PSNR results do not conform to human judgements, while perceptual metric do.</p>
               </caption>
               <graphic id="jpi0156f8_online" content-type="online" xlink:href="jpi0156f8_online.jpg"/>
            </fig><p></p>
         </sec>
      </sec>
      <sec id="jpi0156us5">
         <label>5.</label>
         <title>Discussion</title>
         <p>In this paper, we utilize Generative Adversarial Networks for medical image generation. Our results demonstrate generalizability of the proposed approach across different modalities, such as mammogram, MRI, CT, and skin cancer images. We also manipulate the generated images such that they contain desired attributes. Compared to traditional image blending methods which mainly edit locally, our proposed method not only embeds the desired image attributes but also edits the surrounding tissue texture accordingly to make the overall tissue texture distribution reasonable. Through adversarial training, the GAN model here learns an estimated manifold which is similar to the image manifold of the training dataset. This estimated manifold well-characterizes the semantic statistics of the training dataset, such as the tissue texture, tissue distribution, tissue shapes, and color distribution. Thus, once the contents of certain regions are altered, the GAN knows how to edit the surrounding region to match the semantic statistics of the training dataset, producing authentic manipulated images.</p>
         <p>Our model can generate a vast range of possible stimuli that accomplish a range of specific and controllable goals. For example, the model can output specific body part shapes, lesion types and locations, background and tissue textures, etc. Additionally, our model is capable of generating morphed medical images, gradually and continuously. In certain medical image perception tasks, such as visual search&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib16">16</xref>, <xref ref-type="bibr" rid="jpi0156bib82">82</xref>], visual detection and recognition&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib56">56</xref>], and decision making&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib75">75</xref>, <xref ref-type="bibr" rid="jpi0156bib76">76</xref>], this kind of controllable medical image stimuli can be very useful. The intrinsic problem using real medical image data is that individual differences are substantial: it is not realistic to collect gradually morphing medical images from real medical image data (e.g., finding a sequence of naturally occurring tumors that smoothly morph between shapes or textures is highly unlikely). Using our proposed method, we can generate any number of authentic medical image stimuli that gradually morph. Moreover, all the images are generated via interpolation, which allows us to control the grain of the morphing.</p>
         <p>For the perceptual loss metric, researchers&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib87">87</xref>] have determined that traditional similarity metrics, such as SSIM and PSNR, are not consistent with human perception of typical natural images. But deep neural network based perceptual metrics can, surprisingly, agree with human judgement. Through experiments, we arrive at the same conclusion in medical imaging domain as well; perceptual metrics preferred medical images and are more perceptually similar to the reference images compared to traditional similarity metrics. Thus, perceptual loss metric provides an important measurement of similarity in medical imaging. Using perceptual loss metric as the similarity measurement, we can also generate metamers for any specific medical image. The metamers are a cluster of perceptually similar images which have been widely used in perception researches.</p>
         <p>Medical image perception research is a rapidly growing field. Typical approaches directly or indirectly assume that computer vision will be an alternative to clinical practice. Our study introduces an additional but very different perspective, which is to use computer vision to improve research on medical image perception. Clinicians will not be replaced anytime soon (if ever). To help clinicians make better judgments, we need to understand clinician perception, cognition, and decision. That requires having stimuli (datasets) that are simultaneously realistic (from the perspective of clinicians) and also controllable. Without this, it will be impossible to make the connection between the cognitive mechanisms that clinicians possess, and their diagnostic success in their practice.</p>
         <p>Interestingly, the model and morphing approach we present here could be readily extended to three-dimensional volumetric images. Volumetric medical imaging is now gaining popularity as a standard practice in clinical setttings. The GAN model and morphing approach can be combined in future work to flexibly create volumetric data sets. Moreover, the GAN model is currently unconditioned. We can also change it to conditional GAN model such that changing certain part of the latent code (not through the encoder) can directly modify corresponding attributes of the output image.</p>
         <p></p>
      </sec>
      <sec id="jpi0156us6">
         <label>6.</label>
         <title>Conclusion</title>
         <p>In this paper, we propose usage of Generative Adversarial Network (GAN) for medical image generation. We tested our method on various medical image modalities such as mammogram, MRI, CT, and skin cancer images. Human evaluations verify the success of our method. We also adopt a controllable approach to manipulate the attributes of the generated images in order to meet certain experimental configurations. In the experiments, we successfully generate mammograms with the desired lesion texture and breast shape. The same approach can also be applied to MRI, CT, skin cancer images, and other medical imaging modalities. Finally, we compare traditional similarity measurements with the perceptual metric in medical imaging. We find that the perceptual metric performs better than the traditional similarity metrics such as SSIM and PSNR.</p>
      </sec>
   </body>
   <back>
      <app-group id="jpi0156app">
         <app id="jpi0156sA">
            <label>Appendix A.</label>
            <title>Similarity Measurements</title>
            <sec id="jpi0156appA-1">
               <label>A.1</label>
               <title>SSIM</title>
               <p>The Structural Similarity Index Measure (SSIM) is computed over various patches of an image. The measure between two patches <italic>x</italic> and <italic>y</italic> of the same size is: <disp-formula id="jpi0156eqn7"><label>(A1)</label><mml:math><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>I</mml:mi><mml:mi>M</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula> where <italic>&#x03BC;</italic><sub><italic>x</italic></sub> is the average of <italic>x</italic>, <italic>&#x03BC;</italic><sub><italic>y</italic></sub> is the average of <italic>y</italic>, <italic>&#x03C3;</italic><sub><italic>x</italic></sub><sup>2</sup> is the variance of <italic>x</italic>, <italic>&#x03C3;</italic><sub><italic>y</italic></sub><sup>2</sup> is the variance of <italic>y</italic>, <italic>&#x03C3;</italic><sub><italic>xy</italic></sub> is the covariance of <italic>x</italic> and <italic>y</italic>, <italic>c</italic><sub>1</sub> = (<italic>k</italic><sub>1</sub><italic>L</italic>)<sup>2</sup> and <italic>c</italic><sub>2</sub> = (<italic>k</italic><sub>2</sub><italic>L</italic>)<sup>2</sup> are two variables to stabilize the division with weak denominator with <italic>L</italic> = 2<sup>#bits&#x00A0;per&#x00A0;pixel</sup> &#x2212; 1, <italic>k</italic><sub>1</sub> = 0.01, and <italic>k</italic><sub>2</sub> = 0.03.</p>
            </sec>
            <sec id="jpi0156appA-2">
               <label>A.2</label>
               <title>PSNR</title>
               <p>Given a  <italic>m</italic> &#x00D7; <italic>n</italic> reference image <italic>I</italic> and its distorted version <italic>K</italic>, the PSNR is defined as: <disp-formula id="jpi0156eqn8"><label>(A2)</label><mml:math><mml:mrow><mml:mi>P</mml:mi><mml:mi>S</mml:mi><mml:mi>N</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mn>0</mml:mn><mml:msub><mml:mrow><mml:mo>log</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mn>0</mml:mn><mml:msub><mml:mrow><mml:mo>log</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula> where <italic>MAX</italic><sub><italic>I</italic></sub> is 255 for 8-bit images, and the MSE is computed as: <disp-formula id="jpi0156eqn9"><label>(A3)</label><mml:math><mml:mrow><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:mfrac><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo mathsize="big">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo mathsize="big"> &#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
            </sec>
            <sec id="jpi0156appA-3">
               <label>A.3</label>
               <title>Perceptual loss</title>
               <p>We utilize the same perceptual loss as&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib40">40</xref>]. The loss network is VGG-16&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib71">71</xref>]. For the reference image <italic>r</italic> and the distorted image <italic>x</italic>, the perceptual loss is defined as: <disp-formula id="jpi0156eqn10"><label>(A4)</label><mml:math><mml:mrow><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mtext>feat</mml:mtext></mml:mrow><mml:mrow><mml:mi>&#x03D5;</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mtext>style</mml:mtext></mml:mrow><mml:mrow><mml:mi>&#x03D5;</mml:mi><mml:mo>,</mml:mo><mml:mi>J</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula> where <italic>&#x03BB;</italic><sub><italic>c</italic></sub> and <italic>&#x03BB;</italic><sub><italic>s</italic></sub> are scalars. In the experiment, we set <italic>&#x03BB;</italic><sub><italic>c</italic></sub> = 1 and <italic>&#x03BB;</italic><sub><italic>s</italic></sub> = 1 &#x00D7; 10<sup>5</sup>. <italic>&#x03D5;</italic> represents the VGG network. <italic>l</italic><sub>feat</sub><sup><italic>&#x03D5;</italic>, <italic>j</italic></sup>(<italic>x</italic>, <italic>r</italic>) is the feature reconstruction loss. Let <italic>&#x03D5;</italic><sub><italic>j</italic></sub>(<italic>x</italic>) be the activation of the <italic>j</italic>th layer of the network <italic>&#x03D5;</italic> with a shape of <italic>C</italic><sub><italic>j</italic></sub> &#x00D7; <italic>H</italic><sub><italic>j</italic></sub> &#x00D7; <italic>W</italic><sub><italic>j</italic></sub>. The feature reconstruction loss is defined as: <disp-formula id="jpi0156eqn11"><label>(A5)</label><mml:math><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mtext>feat</mml:mtext></mml:mrow><mml:mrow><mml:mi>&#x03D5;</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x2225;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
               <p>The style reconstruction loss is defined as: <disp-formula id="jpi0156eqn12"><label>(A6)</label><mml:math><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mtext>style</mml:mtext></mml:mrow><mml:mrow><mml:mi>&#x03D5;</mml:mi><mml:mo>,</mml:mo><mml:mi>J</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2225;</mml:mo><mml:msubsup><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula> where <italic>G</italic><sub><italic>j</italic></sub><sup><italic>&#x03D5;</italic></sup>(<italic>x</italic>) is the Gram matrix with a shape of <italic>C</italic><sub><italic>j</italic></sub> &#x00D7; <italic>C</italic><sub><italic>j</italic></sub>. The elements of the Gram matrix can be computed as: <disp-formula id="jpi0156eqn13"><label>(A7)</label><mml:math><mml:mrow><mml:msubsup><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo mathsize="big"> &#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo mathsize="big"> &#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow></mml:math></disp-formula></p>
            </sec>
         </app>
         <app id="jpi0156sB">
            <label>Appendix B.</label>
            <title>Human Evaluation of Generated MRI and Skin Cancer Images</title>
            <p>We also collected human evaluation experiment data for MRI and Skin Cancer images. For the MRI experiment, four observers (1 expert, age range: 25&#x2013;39) participated. For the Skin Cancer experiment, five observers (1 expert, age range: 20&#x2013;39) participated. Unlike CT and mammogram image experiments, we could not recruit sufficient experts for MRI and Skin Cancer online surveys. All experiments were approved by the Institutional Review Board at UC Berkeley and the participants provided informed consent. Stimuli included 50 real and 50 corresponding fake images. All participants followed the same experimental procedures as described in Section&#x00A0;<xref ref-type="sec" rid="jpi0156us4-4-3">4.4.3</xref>.</p>
            <p>The results for MRI and Skin Cancer images in terms of the Receiver Operating Characteristic (ROC) curves are shown in Figure&#x00A0;<xref ref-type="fig" rid="jpi0156figB1">B.1</xref>. The mean area under the curves (AUCs) are 0.57 (<italic>p</italic> = 0.241, permutation test) and 0.62 (<italic>p</italic> = 0.123, permutation test) for MRI and Skin Cancer respectively. Although we did not have experts for these MRI and Dermatology tests, we did have one trained radiologist participate and their data echoes that of untrained observers, all of which are consistent with the CT and mammogram data.</p>
            <fig id="jpi0156figB1"><label>Figure&#x00A0;B.1.</label>
               <caption id="jpi0156fcB1">
                  <p>Human evaluation results for MRI and Skin Cancer images. Participant performance is shown in the Receiver Operating Characteristic (ROC) curves. It is clear that their performance is near chance level (curves near the diagonal region), indicating that the generated medical images were authentic. Here, <italic>P</italic><sub>1</sub> &#x2212; <italic>P</italic><sub><italic>N</italic></sub> represent different untrained observers and experts in corresponding experiments.</p>
               </caption>
               <graphic id="jpi0156fB1_online" content-type="online" xlink:href="jpi0156fB1_online.jpg"/>
            </fig></app>
         <app id="jpi0156sC">
            <label>Appendix C.</label>
            <title>Stimulus Duration Considerations</title>
            <p>There are both empirical and theoretical reasons for limiting the display to 5&#x00A0;s, and the empirical results confirm that 5&#x00A0;s was more than sufficient for observers to reach a reliable decision.</p>
            <p>Firstly, previous research has demonstrated that radiologists can reliably discriminate radiographs within 1&#x00A0;s&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib11">11</xref>, <xref ref-type="bibr" rid="jpi0156bib12">12</xref>, <xref ref-type="bibr" rid="jpi0156bib18">18</xref>, <xref ref-type="bibr" rid="jpi0156bib19">19</xref>, <xref ref-type="bibr" rid="jpi0156bib24">24</xref>, <xref ref-type="bibr" rid="jpi0156bib35">35</xref>, <xref ref-type="bibr" rid="jpi0156bib39">39</xref>, <xref ref-type="bibr" rid="jpi0156bib48">48</xref>, <xref ref-type="bibr" rid="jpi0156bib55">55</xref>, <xref ref-type="bibr" rid="jpi0156bib59">59</xref>]. In our experiment, we provide far more time than 1&#x00A0;s. Moreover, in self paced studies with static radiographs, radiologists often spend less than 5&#x00A0;s&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib69">69</xref>].</p>
            <p>Secondly, our results show that accuracy does not vary with decision time. The decision time is reported as the time from the first viewing of the page to the final &#x201C;submit&#x201D; click by the observer. This is a conservative estimate of the decision duration. The relation between the error and decision time is shown in Figure&#x00A0;<xref ref-type="fig" rid="jpi0156figC1">C.1</xref>. The fitted line reveals that error and decision time are not correlated; more time did not make observers more accurate. Moreover, in this experiment, 60.0% of the decisions were made before stimuli disappeared. Notably, the peak of response density does not occur at the 5&#x00A0;s boundary. It occurs around 2&#x00A0;s, which indicates that the 5&#x00A0;s stimulus duration limit does not induce pressure on participants&#x2019; decision.</p>
            <fig id="jpi0156figC1"><label>Figure&#x00A0;C.1.</label>
               <caption id="jpi0156fcC1">
                  <p>Error-Timing Relation. The scatter plot shows the raw data of participants&#x2019; error and their decision duration. We fit a linear function to reveal the relation between them. It is clear that the error and their decision time are not correlated. The bottom density distribution represents the distribution of participants&#x2019; decision time. The orange line indicates the time point when stimuli disappeared. In this experiment, 60.0% of the decisions were made before the stimuli disappeared.</p>
               </caption>
               <graphic id="jpi0156fC1_online" content-type="online" xlink:href="jpi0156fC1_online.jpg"/>
            </fig><p>Finally, the significant test-retest reliability demonstrates that observers were consistent in their responses. If exposure duration limited performance, it would add noise and that test-retest reliability would be low&#x00A0;[<xref ref-type="bibr" rid="jpi0156bib18">18</xref>].</p>
            <p>Together, all of these considerations suggest that the duration of image interpretation was probably not the limiting factor. From the examples here, it appears that scrutinizing the real and generated images for more than a few seconds does not make them appear more or less similar. This hints that the metameric quality of the images is not due to a time constraint. Nevertheless, we did not force observers to scrutinize the images for more than 5&#x00A0;s, and it is conceivable that forcing an extended viewing of the images could improve performance. For this reason, it will be valuable in future studies to examine the temporal integration of the visual processes that contribute to discrimination of near-metameric medical images.</p>
            <p></p>
         </app>
      </app-group>
      <ack>
         <title>Acknowledgment</title>
         <p>This work has been supported by National Institutes of Health (NIH) under grant # R01CA236793. We thank people who participated in the human evaluation and Min Zhou who recruited medical imaging experts from her hospital.</p>
         <p></p>
      </ack>
      <ref-list content-type="numerical">
         <title>References</title>
         <ref id="jpi0156bib1">
            <label>1</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Amirshahi</surname>
                     <given-names>S.&#x00A0;A.</given-names>
                  </name>
                  <name>
                     <surname>Pedersen</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Yu</surname>
                     <given-names>S.&#x00A0;X.</given-names>
                  </name>
               </person-group>
               <year>2016</year>
               <article-title>Image quality assessment by comparing CNN features between images</article-title>
               <source>J. Imaging Sci. Technol.</source>
               <volume>60</volume>
               <fpage>060410-1</fpage>
               <lpage>060410-10</lpage>
               <page-range>060410-1&#x2013;060410-10</page-range>
               <pub-id pub-id-type="doi">10.2352/J.ImagingSci.Technol.2016.60.6.060410</pub-id>
            </element-citation></ref>
         <ref id="jpi0156bib2">
            <label>2</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Arjovsky</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Chintala</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Bottou</surname>
                     <given-names>L.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Wasserstein generative adversarial networks</article-title>
               <source>Int&#x2019;l. Conf. on Machine Learning</source>
               <fpage>214</fpage>
               <lpage>223</lpage>
               <page-range>214&#x2013;23</page-range>
               <publisher-name>PMLR</publisher-name>
            </element-citation></ref>
         <ref id="jpi0156bib3">
            <label>3</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Bau</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Zhu</surname>
                     <given-names>J.-Y.</given-names>
                  </name>
                  <name>
                     <surname>Wulff</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Peebles</surname>
                     <given-names>W.</given-names>
                  </name>
                  <name>
                     <surname>Strobelt</surname>
                     <given-names>H.</given-names>
                  </name>
                  <name>
                     <surname>Zhou</surname>
                     <given-names>B.</given-names>
                  </name>
                  <name>
                     <surname>Torralba</surname>
                     <given-names>A.</given-names>
                  </name>
               </person-group>
               <year>2019</year>
               <article-title>Seeing what a gan cannot generate</article-title>
               <source>Proc. IEEE/CVF Int&#x2019;l. Conf. on Computer Vision</source>
               <fpage>4502</fpage>
               <lpage>4511</lpage>
               <page-range>4502&#x2013;11</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/ICCV.2019.00460</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib4">
            <label>4</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Birnbaum</surname>
                     <given-names>M.&#x00A0;H.</given-names>
                  </name>
               </person-group>
               <year>2004</year>
               <article-title>Human research and data collection via the internet</article-title>
               <source>Annu. Rev. Psychol.</source>
               <volume>55</volume>
               <fpage>803</fpage>
               <lpage>832</lpage>
               <page-range>803&#x2013;32</page-range>
               <pub-id pub-id-type="doi">10.1146/annurev.psych.55.090902.141601</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib5">
            <label>5</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Bisla</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Choromanska</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Berman</surname>
                     <given-names>R.&#x00A0;S.</given-names>
                  </name>
                  <name>
                     <surname>Stein</surname>
                     <given-names>J.&#x00A0;A.</given-names>
                  </name>
                  <name>
                     <surname>Polsky</surname>
                     <given-names>D.</given-names>
                  </name>
               </person-group>
               <year>2019</year>
               <article-title>Towards automated melanoma detection with deep learning: Data purification and augmentation</article-title>
               <source>Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops</source>
               <fpage>0</fpage>
               <lpage>0</lpage>
               <page-range>0&#x2013;</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPRW.2019.00330</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib6">
            <label>6</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Bissoto</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Perez</surname>
                     <given-names>F.</given-names>
                  </name>
                  <name>
                     <surname>Valle</surname>
                     <given-names>E.</given-names>
                  </name>
                  <name>
                     <surname>Avila</surname>
                     <given-names>S.</given-names>
                  </name>
               </person-group>
               <year>2018</year>
               <article-title>Skin lesion synthesis with generative adversarial networks</article-title>
               <source>OR 2.0 Context-Aware Operating Theaters, Computer Assisted Robotic Endoscopy, Clinical Image-Based Orocedures, and Skin Image Analysis</source>
               <fpage>294</fpage>
               <lpage>302</lpage>
               <page-range>294&#x2013;302</page-range>
               <publisher-name>Springer</publisher-name>
               <publisher-loc>Cham</publisher-loc>
               <pub-id pub-id-type="doi">10.1007/978-3-030-01201-4_32</pub-id>
            </element-citation></ref>
         <ref id="jpi0156bib7">
            <label>7</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Bissoto</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Valle</surname>
                     <given-names>E.</given-names>
                  </name>
                  <name>
                     <surname>Avila</surname>
                     <given-names>S.</given-names>
                  </name>
               </person-group>
               <year>2021</year>
               <article-title>Gan-based data augmentation and anonymization for skin-lesion analysis: A critical review</article-title>
               <source>Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>1847</fpage>
               <lpage>1856</lpage>
               <page-range>1847&#x2013;56</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPRW53098.2021.00204</pub-id>
            </element-citation></ref>
         <ref id="jpi0156bib8">
            <label>8</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Bowyer</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Kopans</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Kegelmeyer</surname>
                     <given-names>W.</given-names>
                  </name>
                  <name>
                     <surname>Moore</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Sallam</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Chang</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Woods</surname>
                     <given-names>K.</given-names>
                  </name>
               </person-group>
               <year>1996</year>
               <article-title>The digital database for screening mammography</article-title>
               <source>Third Int&#x2019;l. Workshop on Digital Mammography</source>
               <volume>Vol.&#x00A0;58</volume>
               <fpage>27</fpage>
               <publisher-name>Elsevier</publisher-name>
               <publisher-loc>Amsterdam</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib9">
            <label>9</label>
            <element-citation publication-type="other"><person-group person-group-type="author">
                  <name>
                     <surname>Brock</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Donahue</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Simonyan</surname>
                     <given-names>K.</given-names>
                  </name>
               </person-group>
               <comment>&#x201C;Large scale GAN training for high fidelity natural image synthesis,&#x201D; <italic>Int&#x2019;l. Conf. on Learning Representations</italic>. 2018</comment>
            </element-citation>
         </ref>
         <ref id="jpi0156bib10">
            <label>10</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Cao</surname>
                     <given-names>B.</given-names>
                  </name>
                  <name>
                     <surname>Zhang</surname>
                     <given-names>H.</given-names>
                  </name>
                  <name>
                     <surname>Wang</surname>
                     <given-names>N.</given-names>
                  </name>
                  <name>
                     <surname>Gao</surname>
                     <given-names>X.</given-names>
                  </name>
                  <name>
                     <surname>Shen</surname>
                     <given-names>D.</given-names>
                  </name>
               </person-group>
               <year>2020</year>
               <article-title>Auto-gan: self-supervised collaborative learning for medical image synthesis</article-title>
               <source>Proc. AAAI Conf. on Artificial Intelligence</source>
               <volume>Vol.&#x00A0;34</volume>
               <fpage>10486</fpage>
               <lpage>10493</lpage>
               <page-range>10486&#x2013;93</page-range>
               <publisher-name>AAAI</publisher-name>
               <publisher-loc>Palo Alto, CA</publisher-loc>
               <pub-id pub-id-type="doi">10.1609/AAAI.V34I07.6619</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib11">
            <label>11</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Carmody</surname>
                     <given-names>D.&#x00A0;P.</given-names>
                  </name>
                  <name>
                     <surname>Nodine</surname>
                     <given-names>C.&#x00A0;F.</given-names>
                  </name>
                  <name>
                     <surname>Kundel</surname>
                     <given-names>H.&#x00A0;L.</given-names>
                  </name>
               </person-group>
               <year>1980</year>
               <article-title>An analysis of perceptual and cognitive factors in radiographic interpretation</article-title>
               <source>Perception</source>
               <volume>9</volume>
               <fpage>339</fpage>
               <lpage>344</lpage>
               <page-range>339&#x2013;44</page-range>
               <pub-id pub-id-type="doi">10.1068/p090339</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib12">
            <label>12</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Carmody</surname>
                     <given-names>D.&#x00A0;P.</given-names>
                  </name>
                  <name>
                     <surname>Nodine</surname>
                     <given-names>C.&#x00A0;F.</given-names>
                  </name>
                  <name>
                     <surname>Kundel</surname>
                     <given-names>H.&#x00A0;L.</given-names>
                  </name>
               </person-group>
               <year>1981</year>
               <article-title>Finding lung nodules with and without comparative visual scanning</article-title>
               <source>Perception Psychophysics</source>
               <volume>29</volume>
               <fpage>594</fpage>
               <lpage>598</lpage>
               <page-range>594&#x2013;8</page-range>
               <pub-id pub-id-type="doi">10.3758/BF03207377</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib13">
            <label>13</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Chen</surname>
                     <given-names>X.</given-names>
                  </name>
                  <name>
                     <surname>Duan</surname>
                     <given-names>Y.</given-names>
                  </name>
                  <name>
                     <surname>Houthooft</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Schulman</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Sutskever</surname>
                     <given-names>I.</given-names>
                  </name>
                  <name>
                     <surname>Abbeel</surname>
                     <given-names>P.</given-names>
                  </name>
               </person-group>
               <year>2016</year>
               <article-title>Infogan: interpretable representation learning by information maximizing generative adversarial nets</article-title>
               <source>Advances in neural information processing systems</source>
               <volume>Vol. 29</volume>
               <fpage>2180</fpage>
               <lpage>2188</lpage>
               <page-range>2180&#x2013;8</page-range>
               <publisher-name>Morgan Kaufmann</publisher-name>
               <publisher-loc>San Francisco, CA</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib14">
            <label>14</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Cicchetti</surname>
                     <given-names>D.&#x00A0;V.</given-names>
                  </name>
               </person-group>
               <year>1994</year>
               <article-title>Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology</article-title>
               <source>Psychological Assess.</source>
               <volume>6</volume>
               <fpage>284</fpage>
               <pub-id pub-id-type="doi">10.1037/1040-3590.6.4.284</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib15">
            <label>15</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Crump</surname>
                     <given-names>M.&#x00A0;J.</given-names>
                  </name>
                  <name>
                     <surname>McDonnell</surname>
                     <given-names>J.&#x00A0;V.</given-names>
                  </name>
                  <name>
                     <surname>Gureckis</surname>
                     <given-names>T.&#x00A0;M.</given-names>
                  </name>
               </person-group>
               <year>2013</year>
               <article-title>Evaluating amazon&#x2019;s mechanical turk as a tool for experimental behavioral research</article-title>
               <source>PloS One</source>
               <volume>8</volume>
               <elocation-id content-type="artnum">e57410</elocation-id>
               <pub-id pub-id-type="doi">10.1371/journal.pone.0057410</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib16">
            <label>16</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Drew</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Evans</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>V&#x00F5;</surname>
                     <given-names>M.&#x00A0;L.-H.</given-names>
                  </name>
                  <name>
                     <surname>Jacobson</surname>
                     <given-names>F.&#x00A0;L.</given-names>
                  </name>
                  <name>
                     <surname>Wolfe</surname>
                     <given-names>J.&#x00A0;M.</given-names>
                  </name>
               </person-group>
               <year>2013</year>
               <article-title>Informatics in radiology: what can you see in a single glance and how might this guide visual search in medical images?</article-title>
               <source>Radiographics</source>
               <volume>33</volume>
               <fpage>263</fpage>
               <lpage>274</lpage>
               <page-range>263&#x2013;74</page-range>
               <pub-id pub-id-type="doi">10.1148/rg.331125023</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib17">
            <label>17</label>
            <element-citation publication-type="other"><person-group person-group-type="author">
                  <name>
                     <surname>Dumoulin</surname>
                     <given-names>V.</given-names>
                  </name>
                  <name>
                     <surname>Shlens</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Kudlur</surname>
                     <given-names>M.</given-names>
                  </name>
               </person-group>
               <comment>&#x201C;A learned representation for artistic style,&#x201D; Preprint arXiv:<ext-link ext-link-type="arxiv" xlink:href="http://arxiv.org/abs/1610.07629">1610.07629</ext-link>, (New York, NY, 2016)</comment>
            </element-citation>
         </ref>
         <ref id="jpi0156bib18">
            <label>18</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Evans</surname>
                     <given-names>K.&#x00A0;K.</given-names>
                  </name>
                  <name>
                     <surname>Georgian-Smith</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Tambouret</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Birdwell</surname>
                     <given-names>R.&#x00A0;L.</given-names>
                  </name>
                  <name>
                     <surname>Wolfe</surname>
                     <given-names>J.&#x00A0;M.</given-names>
                  </name>
               </person-group>
               <year>2013</year>
               <article-title>The gist of the abnormal: Above-chance medical decision making in the blink of an eye</article-title>
               <source>Psychonomic Bull. Rev.</source>
               <volume>20</volume>
               <fpage>1170</fpage>
               <lpage>1175</lpage>
               <page-range>1170&#x2013;5</page-range>
               <pub-id pub-id-type="doi">10.3758/s13423-013-0459-3</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib19">
            <label>19</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Evans</surname>
                     <given-names>K.&#x00A0;K.</given-names>
                  </name>
                  <name>
                     <surname>Haygood</surname>
                     <given-names>T.&#x00A0;M.</given-names>
                  </name>
                  <name>
                     <surname>Cooper</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Culpan</surname>
                     <given-names>A.-M.</given-names>
                  </name>
                  <name>
                     <surname>Wolfe</surname>
                     <given-names>J.&#x00A0;M.</given-names>
                  </name>
               </person-group>
               <year>2016</year>
               <article-title>A half-second glimpse often lets radiologists identify breast cancer cases even when viewing the mammogram of the opposite breast</article-title>
               <source>Proc. Natl. Acad. Sci.</source>
               <volume>113</volume>
               <fpage>10292</fpage>
               <lpage>10297</lpage>
               <page-range>10292&#x2013;7</page-range>
               <pub-id pub-id-type="doi">10.1073/pnas.1606187113</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib20">
            <label>20</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Fleiss</surname>
                     <given-names>J.&#x00A0;L.</given-names>
                  </name>
               </person-group>
               <year>1986</year>
               <article-title>Reliability of measurement</article-title>
               <source>The Design and Analysis of Clinical Experiments</source>
               <publisher-name>John Wiley &#x0026; Sons</publisher-name>
               <publisher-loc>Hoboken, NJ</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib21">
            <label>21</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Fleiss</surname>
                     <given-names>J.&#x00A0;L.</given-names>
                  </name>
               </person-group>
               <source>Design and Analysis of Clinical Experiments</source>
               <year>2011</year>
               <volume>Vol.&#x00A0;73</volume>
               <publisher-name>John Wiley &#x0026; Sons</publisher-name>
               <publisher-loc>Hoboken, NJ</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib22">
            <label>22</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Frid-Adar</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Diamant</surname>
                     <given-names>I.</given-names>
                  </name>
                  <name>
                     <surname>Klang</surname>
                     <given-names>E.</given-names>
                  </name>
                  <name>
                     <surname>Amitai</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Goldberger</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Greenspan</surname>
                     <given-names>H.</given-names>
                  </name>
               </person-group>
               <year>2018</year>
               <article-title>Gan-based synthetic medical image augmentation for increased cnn performance in liver lesion classification</article-title>
               <source>Neurocomputing</source>
               <volume>321</volume>
               <fpage>321</fpage>
               <lpage>331</lpage>
               <page-range>321&#x2013;31</page-range>
               <pub-id pub-id-type="doi">10.1016/j.neucom.2018.09.013</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib23">
            <label>23</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Fukushima</surname>
                     <given-names>K.</given-names>
                  </name>
               </person-group>
               <year>1979</year>
               <article-title>Neural network model for a mechanism of pattern recognition unaffected by shift in position-neocognitron</article-title>
               <source>IEICE Tech. Rep. A</source>
               <volume>62</volume>
               <fpage>658</fpage>
               <lpage>665</lpage>
               <page-range>658&#x2013;65</page-range>
            </element-citation>
         </ref>
         <ref id="jpi0156bib24">
            <label>24</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Gale</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Vernon</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Miller</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Worthington</surname>
                     <given-names>B.</given-names>
                  </name>
               </person-group>
               <year>1990</year>
               <article-title>Reporting in a flash</article-title>
               <source>Br. J. Radiol.</source>
               <volume>63</volume>
               <fpage>71</fpage>
            </element-citation>
         </ref>
         <ref id="jpi0156bib25">
            <label>25</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Gatys</surname>
                     <given-names>L.&#x00A0;A.</given-names>
                  </name>
                  <name>
                     <surname>Ecker</surname>
                     <given-names>A.&#x00A0;S.</given-names>
                  </name>
                  <name>
                     <surname>Bethge</surname>
                     <given-names>M.</given-names>
                  </name>
               </person-group>
               <year>2015</year>
               <article-title>Texture synthesis using convolutional neural networks</article-title>
               <source>Proc. 28th Int&#x2019;l. Conf. on Neural Information Processing Systems-Volume 1</source>
               <fpage>262</fpage>
               <lpage>270</lpage>
               <page-range>262&#x2013;70</page-range>
               <publisher-name>Morgan Kaufmann</publisher-name>
               <publisher-loc>San Francisco, CA</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib26">
            <label>26</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Ghorbani</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Natarajan</surname>
                     <given-names>V.</given-names>
                  </name>
                  <name>
                     <surname>Coz</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Liu</surname>
                     <given-names>Y.</given-names>
                  </name>
               </person-group>
               <year>2020</year>
               <article-title>Dermgan: Synthetic generation of clinical skin images with pathology</article-title>
               <source>Machine Learning for Health Workshop</source>
               <fpage>155</fpage>
               <lpage>170</lpage>
               <page-range>155&#x2013;70</page-range>
               <publisher-name>PMLR</publisher-name>
               <publisher-loc>Cambridge, MA</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib27">
            <label>27</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Girshick</surname>
                     <given-names>R.</given-names>
                  </name>
               </person-group>
               <year>2015</year>
               <article-title>Fast r-cnn</article-title>
               <source>Proc. IEEE Int&#x2019;l. Conf. on Computer Vision</source>
               <fpage>1440</fpage>
               <lpage>1448</lpage>
               <page-range>1440&#x2013;8</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/ICCV.2015.169</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib28">
            <label>28</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Girshick</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Donahue</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Darrell</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Malik</surname>
                     <given-names>J.</given-names>
                  </name>
               </person-group>
               <year>2014</year>
               <article-title>Rich feature hierarchies for accurate object detection and semantic segmentation</article-title>
               <source>Proc. IEEE Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>580</fpage>
               <lpage>587</lpage>
               <page-range>580&#x2013;7</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPR.2014.81</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib29">
            <label>29</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Goodfellow</surname>
                     <given-names>I.</given-names>
                  </name>
                  <name>
                     <surname>Pouget-Abadie</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Mirza</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Xu</surname>
                     <given-names>B.</given-names>
                  </name>
                  <name>
                     <surname>Warde-Farley</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Ozair</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Courville</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Bengio</surname>
                     <given-names>Y.</given-names>
                  </name>
               </person-group>
               <year>2014</year>
               <article-title>Generative adversarial nets</article-title>
               <source>Advances in Neural Information Processing Systems</source>
               <volume>27</volume>
            </element-citation>
         </ref>
         <ref id="jpi0156bib30">
            <label>30</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Gulrajani</surname>
                     <given-names>I.</given-names>
                  </name>
                  <name>
                     <surname>Ahmed</surname>
                     <given-names>F.</given-names>
                  </name>
                  <name>
                     <surname>Arjovsky</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Dumoulin</surname>
                     <given-names>V.</given-names>
                  </name>
                  <name>
                     <surname>Courville</surname>
                     <given-names>A. C.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Improved training of wasserstein gans</article-title>
               <source>Advances in Neural Information Processing Systems</source>
               <volume>30</volume>
            </element-citation>
         </ref>
         <ref id="jpi0156bib31">
            <label>31</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Han</surname>
                     <given-names>C.</given-names>
                  </name>
                  <name>
                     <surname>Hayashi</surname>
                     <given-names>H.</given-names>
                  </name>
                  <name>
                     <surname>Rundo</surname>
                     <given-names>L.</given-names>
                  </name>
                  <name>
                     <surname>Araki</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Shimoda</surname>
                     <given-names>W.</given-names>
                  </name>
                  <name>
                     <surname>Muramatsu</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Furukawa</surname>
                     <given-names>Y.</given-names>
                  </name>
                  <name>
                     <surname>Mauri</surname>
                     <given-names>G.</given-names>
                  </name>
                  <name>
                     <surname>Nakayama</surname>
                     <given-names>H.</given-names>
                  </name>
               </person-group>
               <year>2018</year>
               <article-title>Gan-based synthetic brain mr image generation</article-title>
               <source>2018 IEEE 15th Int&#x2019;l. Symposium on Biomedical Imaging (ISBI 2018)</source>
               <fpage>734</fpage>
               <lpage>738</lpage>
               <page-range>734&#x2013;8</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/ISBI.2018.8363678</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib32">
            <label>32</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>He</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Gkioxari</surname>
                     <given-names>G.</given-names>
                  </name>
                  <name>
                     <surname>Doll&#x00E1;r</surname>
                     <given-names>P.</given-names>
                  </name>
                  <name>
                     <surname>Girshick</surname>
                     <given-names>R.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Mask r-cnn</article-title>
               <source>Proc. IEEE Int&#x2019;l. Conf. on Computer Vision</source>
               <fpage>2961</fpage>
               <lpage>2969</lpage>
               <page-range>2961&#x2013;9</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/ICCV.2017.322</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib33">
            <label>33</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>He</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Zhang</surname>
                     <given-names>X.</given-names>
                  </name>
                  <name>
                     <surname>Ren</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Sun</surname>
                     <given-names>J.</given-names>
                  </name>
               </person-group>
               <year>2016</year>
               <article-title>Deep residual learning for image recognition</article-title>
               <source>Proc. IEEE Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>770</fpage>
               <lpage>778</lpage>
               <page-range>770&#x2013;8</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib34">
            <label>34</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Hesamian</surname>
                     <given-names>M.&#x00A0;H.</given-names>
                  </name>
                  <name>
                     <surname>Jia</surname>
                     <given-names>W.</given-names>
                  </name>
                  <name>
                     <surname>He</surname>
                     <given-names>X.</given-names>
                  </name>
                  <name>
                     <surname>Kennedy</surname>
                     <given-names>P.</given-names>
                  </name>
               </person-group>
               <year>2019</year>
               <article-title>Deep learning techniques for medical image segmentation: achievements and challenges</article-title>
               <source>J. Digit. Imaging</source>
               <volume>32</volume>
               <fpage>582</fpage>
               <lpage>596</lpage>
               <page-range>582&#x2013;96</page-range>
               <pub-id pub-id-type="doi">10.1007/s10278-019-00227-x</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib35">
            <label>35</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Houghton</surname>
                     <given-names>J.&#x00A0;P.</given-names>
                  </name>
                  <name>
                     <surname>Smoller</surname>
                     <given-names>B.&#x00A0;R.</given-names>
                  </name>
                  <name>
                     <surname>Leonard</surname>
                     <given-names>N.</given-names>
                  </name>
                  <name>
                     <surname>Stevenson</surname>
                     <given-names>M.&#x00A0;R.</given-names>
                  </name>
                  <name>
                     <surname>Dornan</surname>
                     <given-names>T.</given-names>
                  </name>
               </person-group>
               <year>2015</year>
               <article-title>Diagnostic performance on briefly presented digital pathology images</article-title>
               <source>J. Pathol. Inform.</source>
               <volume>6</volume>
               <pub-id pub-id-type="doi">10.4103/2153-3539.168517</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib36">
            <label>36</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Huang</surname>
                     <given-names>X.</given-names>
                  </name>
                  <name>
                     <surname>Belongie</surname>
                     <given-names>S.</given-names>
                  </name>
               </person-group>
               <article-title>Arbitrary style transfer in real-time with adaptive instance normalization</article-title>
               <source>Proc. IEEE Int&#x2019;l. Conf. on Computer Vision</source>
               <year>2017</year>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <fpage>1501</fpage>
               <lpage>1510</lpage>
               <page-range>1501&#x2013;10</page-range>
               <pub-id pub-id-type="doi">10.1109/ICCV.2017.167</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib37">
            <label>37</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Huang</surname>
                     <given-names>X.</given-names>
                  </name>
                  <name>
                     <surname>Liu</surname>
                     <given-names>M.-Y.</given-names>
                  </name>
                  <name>
                     <surname>Belongie</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Kautz</surname>
                     <given-names>J.</given-names>
                  </name>
               </person-group>
               <year>2018</year>
               <article-title>Multimodal unsupervised image-to-image translation</article-title>
               <source>Proc. European Conf. on Computer Vision (ECCV)</source>
               <fpage>172</fpage>
               <lpage>189</lpage>
               <page-range>172&#x2013;89</page-range>
               <publisher-name>Springer</publisher-name>
               <publisher-loc>Cham</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib38">
            <label>38</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Hubel</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Wiesel</surname>
                     <given-names>T.</given-names>
                  </name>
               </person-group>
               <year>1959</year>
               <article-title>Receptive fields of single neurones in the cat&#x2019;s striate cortex</article-title>
               <source>J. Physiol.</source>
               <volume>148</volume>
               <fpage>574</fpage>
               <pub-id pub-id-type="doi">10.1113/jphysiol.1959.sp006308</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib39">
            <label>39</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Jaarsma</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Jarodzka</surname>
                     <given-names>H.</given-names>
                  </name>
                  <name>
                     <surname>Nap</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>van Merrienboer</surname>
                     <given-names>J.&#x00A0;J.</given-names>
                  </name>
                  <name>
                     <surname>Boshuizen</surname>
                     <given-names>H.&#x00A0;P.</given-names>
                  </name>
               </person-group>
               <year>2014</year>
               <article-title>Expertise under the microscope: Processing histopathological slides</article-title>
               <source>Med. Educ.</source>
               <volume>48</volume>
               <fpage>292</fpage>
               <lpage>300</lpage>
               <page-range>292&#x2013;300</page-range>
               <pub-id pub-id-type="doi">10.1111/medu.12385</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib40">
            <label>40</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Johnson</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Alahi</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Fei-Fei</surname>
                     <given-names>L.</given-names>
                  </name>
               </person-group>
               <year>2016</year>
               <article-title>Perceptual losses for real-time style transfer and super-resolution</article-title>
               <source>European Conf. on Computer Vision</source>
               <fpage>694</fpage>
               <lpage>711</lpage>
               <page-range>694&#x2013;711</page-range>
               <publisher-name>Springer</publisher-name>
               <publisher-loc>Cham</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib41">
            <label>41</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Karras</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Aila</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Laine</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Lehtinen</surname>
                     <given-names>J.</given-names>
                  </name>
               </person-group>
               <article-title>Progressive growing of GANs for improved quality, stability, and variation</article-title>
               <source>Int&#x2019;l. Conf. on Learning Representations</source>
               <year>2018</year>
            </element-citation>
         </ref>
         <ref id="jpi0156bib42">
            <label>42</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Karras</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Laine</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Aila</surname>
                     <given-names>T.</given-names>
                  </name>
               </person-group>
               <year>2019</year>
               <article-title>A style-based generator architecture for generative adversarial networks</article-title>
               <source>Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>4401</fpage>
               <lpage>4410</lpage>
               <page-range>4401&#x2013;10</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPR.2019.00453</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib43">
            <label>43</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Karras</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Laine</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Aittala</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Hellsten</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Lehtinen</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Aila</surname>
                     <given-names>T.</given-names>
                  </name>
               </person-group>
               <year>2020</year>
               <article-title>Analyzing and improving the image quality of stylegan</article-title>
               <source>Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>8110</fpage>
               <lpage>8119</lpage>
               <page-range>8110&#x2013;9</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00813</pub-id>
            </element-citation></ref>
         <ref id="jpi0156bib44">
            <label>44</label>
            <element-citation publication-type="other"><person-group person-group-type="author">
                  <name>
                     <surname>Kayalibay</surname>
                     <given-names>B.</given-names>
                  </name>
                  <name>
                     <surname>Jensen</surname>
                     <given-names>G.</given-names>
                  </name>
                  <name>
                     <surname>van&#x00A0;der Smagt</surname>
                     <given-names>P.</given-names>
                  </name>
               </person-group>
               <comment>&#x201C;Cnn-based segmentation of medical imaging data,&#x201D; Preprint arXiv:<ext-link ext-link-type="arxiv" xlink:href="http://arxiv.org/abs/1701.03056">1701.03056</ext-link>, (2017)</comment>
            </element-citation>
         </ref>
         <ref id="jpi0156bib45">
            <label>45</label>
            <element-citation publication-type="other"><person-group person-group-type="author">
                  <name>
                     <surname>Kingma</surname>
                     <given-names>D.&#x00A0;P.</given-names>
                  </name>
                  <name>
                     <surname>Ba</surname>
                     <given-names>J.</given-names>
                  </name>
               </person-group>
               <comment>&#x201C;Adam: A method for stochastic optimization,&#x201D; Preprint arXiv:<ext-link ext-link-type="arxiv" xlink:href="http://arxiv.org/abs/1412.6980">1412.6980</ext-link>, (2014)</comment>
            </element-citation>
         </ref>
         <ref id="jpi0156bib46">
            <label>46</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Kompaniez-Dunigan</surname>
                     <given-names>E.</given-names>
                  </name>
                  <name>
                     <surname>Abbey</surname>
                     <given-names>C.&#x00A0;K.</given-names>
                  </name>
                  <name>
                     <surname>Boone</surname>
                     <given-names>J.&#x00A0;M.</given-names>
                  </name>
                  <name>
                     <surname>Webster</surname>
                     <given-names>M.&#x00A0;A.</given-names>
                  </name>
               </person-group>
               <year>2015</year>
               <article-title>Adaptation and visual search in mammographic images</article-title>
               <source>Attention Perception Psychophysics</source>
               <volume>77</volume>
               <fpage>1081</fpage>
               <lpage>1087</lpage>
               <page-range>1081&#x2013;7</page-range>
               <pub-id pub-id-type="doi">10.3758/s13414-015-0841-5</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib47">
            <label>47</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Krizhevsky</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Sutskever</surname>
                     <given-names>I.</given-names>
                  </name>
                  <name>
                     <surname>Hinton</surname>
                     <given-names>G.&#x00A0;E.</given-names>
                  </name>
               </person-group>
               <article-title>Imagenet classification with deep convolutional neural networks</article-title>
               <source>Advances in Neural Information Processing Systems</source>
               <year>2012</year>
               <volume>Vol. 25</volume>
               <publisher-name>Morgan Kaufmann</publisher-name>
               <publisher-loc>San Francisco, CA</publisher-loc>
               <fpage>1097</fpage>
               <lpage>1105</lpage>
               <page-range>1097&#x2013;105</page-range>
            </element-citation>
         </ref>
         <ref id="jpi0156bib48">
            <label>48</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Kundel</surname>
                     <given-names>H.&#x00A0;L.</given-names>
                  </name>
                  <name>
                     <surname>Nodine</surname>
                     <given-names>C.&#x00A0;F.</given-names>
                  </name>
               </person-group>
               <year>1975</year>
               <article-title>Interpreting chest radiographs without visual search</article-title>
               <source>Radiology</source>
               <volume>116</volume>
               <fpage>527</fpage>
               <lpage>532</lpage>
               <page-range>527&#x2013;32</page-range>
               <pub-id pub-id-type="doi">10.1148/116.3.527</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib49">
            <label>49</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>LeCun</surname>
                     <given-names>Y.</given-names>
                  </name>
                  <name>
                     <surname>Bottou</surname>
                     <given-names>L.</given-names>
                  </name>
                  <name>
                     <surname>Bengio</surname>
                     <given-names>Y.</given-names>
                  </name>
                  <name>
                     <surname>Haffner</surname>
                     <given-names>P.</given-names>
                  </name>
               </person-group>
               <year>1998</year>
               <article-title>Gradient-based learning applied to document recognition</article-title>
               <source>Proc. IEEE</source>
               <volume>86</volume>
               <fpage>2278</fpage>
               <lpage>2324</lpage>
               <page-range>2278&#x2013;324</page-range>
               <pub-id pub-id-type="doi">10.1109/5.726791</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib50">
            <label>50</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Ledig</surname>
                     <given-names>C.</given-names>
                  </name>
                  <name>
                     <surname>Theis</surname>
                     <given-names>L.</given-names>
                  </name>
                  <name>
                     <surname>Husz&#x00E1;r</surname>
                     <given-names>F.</given-names>
                  </name>
                  <name>
                     <surname>Caballero</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Cunningham</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Acosta</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Aitken</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Tejani</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Totz</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Wang</surname>
                     <given-names>Z.</given-names>
                  </name>
                  <name>
                     <surname>Shi</surname>
                     <given-names>W.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Photo-realistic single image super-resolution using a generative adversarial network</article-title>
               <source>Proc. IEEE Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>4681</fpage>
               <lpage>4690</lpage>
               <page-range>4681&#x2013;90</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPR.2017.19</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib51">
            <label>51</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Litjens</surname>
                     <given-names>G.</given-names>
                  </name>
                  <name>
                     <surname>Kooi</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Bejnordi</surname>
                     <given-names>B.&#x00A0;E.</given-names>
                  </name>
                  <name>
                     <surname>Setio</surname>
                     <given-names>A.&#x00A0;A.&#x00A0;A.</given-names>
                  </name>
                  <name>
                     <surname>Ciompi</surname>
                     <given-names>F.</given-names>
                  </name>
                  <name>
                     <surname>Ghafoorian</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Van Der&#x00A0;Laak</surname>
                     <given-names>J.&#x00A0;A.</given-names>
                  </name>
                  <name>
                     <surname>Van&#x00A0;Ginneken</surname>
                     <given-names>B.</given-names>
                  </name>
                  <name>
                     <surname>S&#x00E1;nchez</surname>
                     <given-names>C.&#x00A0;I.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>A survey on deep learning in medical image analysis</article-title>
               <source>Med. Image Anal.</source>
               <volume>42</volume>
               <fpage>60</fpage>
               <lpage>88</lpage>
               <page-range>60&#x2013;88</page-range>
               <pub-id pub-id-type="doi">10.1016/j.media.2017.07.005</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib52">
            <label>52</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Manassi</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Kristj&#x00E1;nsson</surname>
                     <given-names>&#x00C1;.</given-names>
                  </name>
                  <name>
                     <surname>Whitney</surname>
                     <given-names>D.</given-names>
                  </name>
               </person-group>
               <year>2019</year>
               <article-title>Serial dependence in a simulated clinical visual search task</article-title>
               <source>Sci. Rep.</source>
               <volume>9</volume>
               <fpage>1</fpage>
               <lpage>10</lpage>
               <page-range>1&#x2013;10</page-range>
               <pub-id pub-id-type="doi">10.1038/s41598-019-56315-z</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib53">
            <label>53</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Masci</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Meier</surname>
                     <given-names>U.</given-names>
                  </name>
                  <name>
                     <surname>Cire&#x015F;an</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Schmidhuber</surname>
                     <given-names>J.</given-names>
                  </name>
               </person-group>
               <year>2011</year>
               <article-title>Stacked convolutional auto-encoders for hierarchical feature extraction</article-title>
               <source>Int&#x2019;l. Conf. on Artificial Neural Networks</source>
               <fpage>52</fpage>
               <lpage>59</lpage>
               <page-range>52&#x2013;9</page-range>
               <publisher-name>Springer</publisher-name>
               <publisher-loc>Berlin, Heidelberg</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib54">
            <label>54</label>
            <element-citation publication-type="other"><person-group person-group-type="author">
                  <name>
                     <surname>Mirza</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Osindero</surname>
                     <given-names>S.</given-names>
                  </name>
               </person-group>
               <comment>&#x201C;Conditional generative adversarial nets,&#x201D; Preprint arXiv:<ext-link ext-link-type="arxiv" xlink:href="http://arxiv.org/abs/1411.1784">1411.1784</ext-link>, (2014)</comment>
            </element-citation>
         </ref>
         <ref id="jpi0156bib55">
            <label>55</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Mugglestone</surname>
                     <given-names>M.&#x00A0;D.</given-names>
                  </name>
                  <name>
                     <surname>Gale</surname>
                     <given-names>A.&#x00A0;G.</given-names>
                  </name>
                  <name>
                     <surname>Cowley</surname>
                     <given-names>H.&#x00A0;C.</given-names>
                  </name>
                  <name>
                     <surname>Wilson</surname>
                     <given-names>A.</given-names>
                  </name>
               </person-group>
               <year>1995</year>
               <article-title>Diagnostic performance on briefly presented mammographic images</article-title>
               <source>Proc. SPIE</source>
               <volume>2436</volume>
               <pub-id pub-id-type="doi">10.1117/12.206840</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib56">
            <label>56</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Nakashima</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Kobayashi</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Maeda</surname>
                     <given-names>E.</given-names>
                  </name>
                  <name>
                     <surname>Yoshikawa</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Yokosawa</surname>
                     <given-names>K.</given-names>
                  </name>
               </person-group>
               <year>2013</year>
               <article-title>Visual search of experts in medical image reading: the effect of training, target prevalence, and expert knowledge</article-title>
               <source>Frontiers Psychol.</source>
               <volume>4</volume>
               <fpage>166</fpage>
               <pub-id pub-id-type="doi">10.3389/fpsyg.2013.00166</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib57">
            <label>57</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Nie</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Trullo</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Lian</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Petitjean</surname>
                     <given-names>C.</given-names>
                  </name>
                  <name>
                     <surname>Ruan</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Wang</surname>
                     <given-names>Q.</given-names>
                  </name>
                  <name>
                     <surname>Shen</surname>
                     <given-names>D.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Medical image synthesis with context-aware generative adversarial networks</article-title>
               <source>Int&#x2019;l. Conf. on Medical Image Computing and Computer-Assisted Intervention</source>
               <fpage>417</fpage>
               <lpage>425</lpage>
               <page-range>417&#x2013;25</page-range>
               <publisher-name>Springer</publisher-name>
               <publisher-loc>Cham</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib58">
            <label>58</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Odena</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Olah</surname>
                     <given-names>C.</given-names>
                  </name>
                  <name>
                     <surname>Shlens</surname>
                     <given-names>J.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Conditional image synthesis with auxiliary classifier gans</article-title>
               <source>Int&#x2019;l. Conf. on Machine Learning</source>
               <fpage>2642</fpage>
               <lpage>2651</lpage>
               <page-range>2642&#x2013;51</page-range>
               <publisher-name>PMLR</publisher-name>
               <publisher-loc>Cambridge, MA</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib59">
            <label>59</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Oestmann</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Greene</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Kushner</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Bourgouin</surname>
                     <given-names>P.</given-names>
                  </name>
                  <name>
                     <surname>Linetsky</surname>
                     <given-names>L.</given-names>
                  </name>
                  <name>
                     <surname>Llewellyn</surname>
                     <given-names>H.</given-names>
                  </name>
               </person-group>
               <year>1988</year>
               <article-title>Lung lesions: correlation between viewing time and detection</article-title>
               <source>Radiology</source>
               <volume>166</volume>
               <fpage>451</fpage>
               <lpage>453</lpage>
               <page-range>451&#x2013;3</page-range>
               <pub-id pub-id-type="doi">10.1148/radiology.166.2.3336720</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib60">
            <label>60</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Park</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Liu</surname>
                     <given-names>M.-Y.</given-names>
                  </name>
                  <name>
                     <surname>Wang</surname>
                     <given-names>T.-C.</given-names>
                  </name>
                  <name>
                     <surname>Zhu</surname>
                     <given-names>J.-Y.</given-names>
                  </name>
               </person-group>
               <year>2019</year>
               <article-title>Semantic image synthesis with spatially-adaptive normalization</article-title>
               <source>Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>2337</fpage>
               <lpage>2346</lpage>
               <page-range>2337&#x2013;46</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPR.2019.00244</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib61">
            <label>61</label>
            <element-citation publication-type="other"><person-group person-group-type="author">
                  <name>
                     <surname>Radford</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Metz</surname>
                     <given-names>L.</given-names>
                  </name>
                  <name>
                     <surname>Chintala</surname>
                     <given-names>S.</given-names>
                  </name>
               </person-group>
               <comment>&#x201C;Unsupervised representation learning with deep convolutional generative adversarial networks,&#x201D; Preprint arXiv:<ext-link ext-link-type="arxiv" xlink:href="http://arxiv.org/abs/1511.06434">1511.06434</ext-link>, (2015)</comment>
            </element-citation>
         </ref>
         <ref id="jpi0156bib62">
            <label>62</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Rajananda</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Peters</surname>
                     <given-names>M.&#x00A0;A.</given-names>
                  </name>
                  <name>
                     <surname>Lau</surname>
                     <given-names>H.</given-names>
                  </name>
                  <name>
                     <surname>Odegaard</surname>
                     <given-names>B.</given-names>
                  </name>
               </person-group>
               <year>2018</year>
               <article-title>Visual psychophysics on the web: open-access tools, experiments, and results using online platforms</article-title>
               <source>J. Vis.</source>
               <volume>18</volume>
               <fpage>299</fpage>
               <lpage>299</lpage>
               <page-range>299&#x2013;</page-range>
               <pub-id pub-id-type="doi">10.1167/18.10.299</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib63">
            <label>63</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Ranzato</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Huang</surname>
                     <given-names>F.&#x00A0;J.</given-names>
                  </name>
                  <name>
                     <surname>Boureau</surname>
                     <given-names>Y.-L.</given-names>
                  </name>
                  <name>
                     <surname>LeCun</surname>
                     <given-names>Y.</given-names>
                  </name>
               </person-group>
               <year>2007</year>
               <article-title>Unsupervised learning of invariant feature hierarchies with applications to object recognition</article-title>
               <source>2007 IEEE Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>1</fpage>
               <lpage>8</lpage>
               <page-range>1&#x2013;8</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib64">
            <label>64</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Redmon</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Divvala</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Girshick</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Farhadi</surname>
                     <given-names>A.</given-names>
                  </name>
               </person-group>
               <year>2016</year>
               <article-title>You only look once: Unified, real-time object detection</article-title>
               <source>Proc. IEEE Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>779</fpage>
               <lpage>788</lpage>
               <page-range>779&#x2013;88</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib65">
            <label>65</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Ren</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>He</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Girshick</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Sun</surname>
                     <given-names>J.</given-names>
                  </name>
               </person-group>
               <year>2015</year>
               <article-title>Faster r-cnn: Towards real-time object detection with region proposal networks</article-title>
               <source>Advances in Neural Information Processing Systems</source>
               <volume>28</volume>
            </element-citation>
         </ref>
         <ref id="jpi0156bib66">
            <label>66</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Ren</surname>
                     <given-names>Z.</given-names>
                  </name>
                  <name>
                     <surname>Yu</surname>
                     <given-names>S.&#x00A0;X.</given-names>
                  </name>
                  <name>
                     <surname>Whitney</surname>
                     <given-names>D.</given-names>
                  </name>
               </person-group>
               <article-title>Controllable medical image generation via generative adversarial networks</article-title>
               <source>IS&#x0026;T Electronic Imaging: Human Vision and Electronic Imaging</source>
               <year>2021</year>
               <volume>Vol. 2021</volume>
               <publisher-name>IS&#x0026;T</publisher-name>
               <publisher-loc>Springfield, VA</publisher-loc>
               <fpage>112&#x2013;1</fpage>
               <lpage>112&#x2013;5</lpage>
               <page-range>112&#x2013;1&#x2013;5</page-range>
               <pub-id pub-id-type="doi">10.2352/ISSN.2470-1173.2021.11.HVEI-112</pub-id>
            </element-citation></ref>
         <ref id="jpi0156bib67">
            <label>67</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Semmelmann</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Weigelt</surname>
                     <given-names>S.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Online psychophysics: Reaction time effects in cognitive experiments</article-title>
               <source>Behav. Res. Methods</source>
               <volume>49</volume>
               <fpage>1241</fpage>
               <lpage>1260</lpage>
               <page-range>1241&#x2013;60</page-range>
               <pub-id pub-id-type="doi">10.3758/s13428-016-0783-4</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib68">
            <label>68</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Shen</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Wu</surname>
                     <given-names>G.</given-names>
                  </name>
                  <name>
                     <surname>Suk</surname>
                     <given-names>H.-I.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Deep learning in medical image analysis</article-title>
               <source>Annu. Rev. Biomed. Eng.</source>
               <volume>19</volume>
               <fpage>221</fpage>
               <lpage>248</lpage>
               <page-range>221&#x2013;48</page-range>
               <pub-id pub-id-type="doi">10.1146/annurev-bioeng-071516-044442</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib69">
            <label>69</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Sheridan</surname>
                     <given-names>H.</given-names>
                  </name>
                  <name>
                     <surname>Reingold</surname>
                     <given-names>E.&#x00A0;M.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>The holistic processing account of visual expertise in medical image perception: A review</article-title>
               <source>Frontiers Psychol.</source>
               <volume>8</volume>
               <fpage>1620</fpage>
               <pub-id pub-id-type="doi">10.3389/fpsyg.2017.01620</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib70">
            <label>70</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Shin</surname>
                     <given-names>H.-C.</given-names>
                  </name>
                  <name>
                     <surname>Roth</surname>
                     <given-names>H.&#x00A0;R.</given-names>
                  </name>
                  <name>
                     <surname>Gao</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Lu</surname>
                     <given-names>L.</given-names>
                  </name>
                  <name>
                     <surname>Xu</surname>
                     <given-names>Z.</given-names>
                  </name>
                  <name>
                     <surname>Nogues</surname>
                     <given-names>I.</given-names>
                  </name>
                  <name>
                     <surname>Yao</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Mollura</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Summers</surname>
                     <given-names>R.&#x00A0;M.</given-names>
                  </name>
               </person-group>
               <year>2016</year>
               <article-title>Deep convolutional neural networks for computer-aided detection: Cnn architectures, dataset characteristics and transfer learning</article-title>
               <source>IEEE Trans. Med. Imaging</source>
               <volume>35</volume>
               <fpage>1285</fpage>
               <lpage>1298</lpage>
               <page-range>1285&#x2013;98</page-range>
               <pub-id pub-id-type="doi">10.1109/TMI.2016.2528162</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib71">
            <label>71</label>
            <element-citation publication-type="other"><person-group person-group-type="author">
                  <name>
                     <surname>Simonyan</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Zisserman</surname>
                     <given-names>A.</given-names>
                  </name>
               </person-group>
               <comment>&#x201C;Very deep convolutional networks for large-scale image recognition,&#x201D; Preprint arXiv:<ext-link ext-link-type="arxiv" xlink:href="http://arxiv.org/abs/1409.1556">1409.1556</ext-link>, (2014)</comment>
            </element-citation>
         </ref>
         <ref id="jpi0156bib72">
            <label>72</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Szegedy</surname>
                     <given-names>C.</given-names>
                  </name>
                  <name>
                     <surname>Liu</surname>
                     <given-names>W.</given-names>
                  </name>
                  <name>
                     <surname>Jia</surname>
                     <given-names>Y.</given-names>
                  </name>
                  <name>
                     <surname>Sermanet</surname>
                     <given-names>P.</given-names>
                  </name>
                  <name>
                     <surname>Reed</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Anguelov</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Erhan</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Vanhoucke</surname>
                     <given-names>V.</given-names>
                  </name>
                  <name>
                     <surname>Rabinovich</surname>
                     <given-names>A.</given-names>
                  </name>
               </person-group>
               <year>2015</year>
               <article-title>Going deeper with convolutions</article-title>
               <source>Proc. IEEE Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>1</fpage>
               <lpage>9</lpage>
               <page-range>1&#x2013;9</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298594</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib73">
            <label>73</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Szegedy</surname>
                     <given-names>C.</given-names>
                  </name>
                  <name>
                     <surname>Vanhoucke</surname>
                     <given-names>V.</given-names>
                  </name>
                  <name>
                     <surname>Ioffe</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Shlens</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Wojna</surname>
                     <given-names>Z.</given-names>
                  </name>
               </person-group>
               <year>2016</year>
               <article-title>Rethinking the inception architecture for computer vision</article-title>
               <source>Proc. IEEE Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>2818</fpage>
               <lpage>2826</lpage>
               <page-range>2818&#x2013;26</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPR.2016.308</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib74">
            <label>74</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Trevi&#x00F1;o</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Birdsong</surname>
                     <given-names>G.</given-names>
                  </name>
                  <name>
                     <surname>Carrigan</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Choyke</surname>
                     <given-names>P.</given-names>
                  </name>
                  <name>
                     <surname>Drew</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Eckstein</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Fernandez</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Gallas</surname>
                     <given-names>B. D.</given-names>
                  </name>
                  <name>
                     <surname>Giger</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Hewitt</surname>
                     <given-names>S. M.</given-names>
                  </name>
                  <name>
                     <surname>Horowitz</surname>
                     <given-names>T. S.</given-names>
                  </name>
                  <name>
                     <surname>Jiang</surname>
                     <given-names>Y. V.</given-names>
                  </name>
                  <name>
                     <surname>Kudrick</surname>
                     <given-names>B.</given-names>
                  </name>
                  <name>
                     <surname>Martinez-Conde</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Mitroff</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Nebeling</surname>
                     <given-names>L.</given-names>
                  </name>
                  <name>
                     <surname>Saltz</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Samuelson</surname>
                     <given-names>F.</given-names>
                  </name>
                  <name>
                     <surname>Seltzer</surname>
                     <given-names>S. E.</given-names>
                  </name>
                  <name>
                     <surname>Shabestari</surname>
                     <given-names>B.</given-names>
                  </name>
                  <name>
                     <surname>Shankar</surname>
                     <given-names>L.</given-names>
                  </name>
                  <name>
                     <surname>Siegel</surname>
                     <given-names>E.</given-names>
                  </name>
                  <name>
                     <surname>Tilkin</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Trueblood</surname>
                     <given-names>J. S.</given-names>
                  </name>
                  <name>
                     <surname>Dyke</surname>
                     <given-names>A. L. V.</given-names>
                  </name>
                  <name>
                     <surname>Venkatesan</surname>
                     <given-names>A. M.</given-names>
                  </name>
                  <name>
                     <surname>Whitney</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Wolfe</surname>
                     <given-names>J. M.</given-names>
                  </name>
               </person-group>
               <year>2022</year>
               <article-title>Advancing research on medical image perception by strengthening multidisciplinary collaboration</article-title>
               <source>JNCI Cancer Spectrum</source>
               <volume>6</volume>
               <elocation-id content-type="artnum">pkab099</elocation-id>
               <pub-id pub-id-type="doi">10.1093/jncics/pkab099</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib75">
            <label>75</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Trueblood</surname>
                     <given-names>J.&#x00A0;S.</given-names>
                  </name>
                  <name>
                     <surname>Eichbaum</surname>
                     <given-names>Q.</given-names>
                  </name>
                  <name>
                     <surname>Seegmiller</surname>
                     <given-names>A.&#x00A0;C.</given-names>
                  </name>
                  <name>
                     <surname>Stratton</surname>
                     <given-names>C.</given-names>
                  </name>
                  <name>
                     <surname>O&#x2019;Daniels</surname>
                     <given-names>P.</given-names>
                  </name>
                  <name>
                     <surname>Holmes</surname>
                     <given-names>W.&#x00A0;R.</given-names>
                  </name>
               </person-group>
               <year>2021</year>
               <article-title>Disentangling prevalence induced biases in medical image decision-making</article-title>
               <source>Cognition</source>
               <volume>212</volume>
               <elocation-id content-type="artnum">104713</elocation-id>
               <pub-id pub-id-type="doi">10.1016/j.cognition.2021.104713</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib76">
            <label>76</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Trueblood</surname>
                     <given-names>J.&#x00A0;S.</given-names>
                  </name>
                  <name>
                     <surname>Holmes</surname>
                     <given-names>W.&#x00A0;R.</given-names>
                  </name>
                  <name>
                     <surname>Seegmiller</surname>
                     <given-names>A.&#x00A0;C.</given-names>
                  </name>
                  <name>
                     <surname>Douds</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Compton</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Szentirmai</surname>
                     <given-names>E.</given-names>
                  </name>
                  <name>
                     <surname>Woodruff</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Huang</surname>
                     <given-names>W.</given-names>
                  </name>
                  <name>
                     <surname>Stratton</surname>
                     <given-names>C.</given-names>
                  </name>
                  <name>
                     <surname>Eichbaum</surname>
                     <given-names>Q.</given-names>
                  </name>
               </person-group>
               <year>2018</year>
               <article-title>The impact of speed and bias on the cognitive processes of experts and novices in medical image decision-making</article-title>
               <source>Cogn. Res.: Princ. Implications</source>
               <volume>3</volume>
               <fpage>1</fpage>
               <lpage>14</lpage>
               <page-range>1&#x2013;14</page-range>
               <pub-id pub-id-type="doi">10.1186/s41235-017-0085-0</pub-id>
            </element-citation></ref>
         <ref id="jpi0156bib77">
            <label>77</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Wagner</surname>
                     <given-names>R.&#x00A0;F.</given-names>
                  </name>
                  <name>
                     <surname>Brown</surname>
                     <given-names>D.&#x00A0;G.</given-names>
                  </name>
               </person-group>
               <year>1985</year>
               <article-title>Unified snr analysis of medical imaging systems</article-title>
               <source>Phys. Med. Biol.</source>
               <volume>30</volume>
               <fpage>489</fpage>
               <pub-id pub-id-type="doi">10.1088/0031-9155/30/6/001</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib78">
            <label>78</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Waite</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Grigorian</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Alexander</surname>
                     <given-names>R.&#x00A0;G.</given-names>
                  </name>
                  <name>
                     <surname>Macknik</surname>
                     <given-names>S.&#x00A0;L.</given-names>
                  </name>
                  <name>
                     <surname>Carrasco</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Heeger</surname>
                     <given-names>D.&#x00A0;J.</given-names>
                  </name>
                  <name>
                     <surname>Martinez-Conde</surname>
                     <given-names>S.</given-names>
                  </name>
               </person-group>
               <year>2019</year>
               <article-title>Analysis of perceptual expertise in radiology&#x2013;current knowledge and a new perspective</article-title>
               <source>Frontiers Hum. Neurosci.</source>
               <volume>13</volume>
               <fpage>213</fpage>
               <pub-id pub-id-type="doi">10.3389/fnhum.2019.00213</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib79">
            <label>79</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Waite</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Scott</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Gale</surname>
                     <given-names>B.</given-names>
                  </name>
                  <name>
                     <surname>Fuchs</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Kolla</surname>
                     <given-names>S.</given-names>
                  </name>
                  <name>
                     <surname>Reede</surname>
                     <given-names>D.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Interpretive error in radiology</article-title>
               <source>Am. J. Roentgenol.</source>
               <volume>208</volume>
               <fpage>739</fpage>
               <lpage>749</lpage>
               <page-range>739&#x2013;49</page-range>
               <pub-id pub-id-type="doi">10.2214/AJR.16.16963</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib80">
            <label>80</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Willemink</surname>
                     <given-names>M.&#x00A0;J.</given-names>
                  </name>
                  <name>
                     <surname>Koszek</surname>
                     <given-names>W.&#x00A0;A.</given-names>
                  </name>
                  <name>
                     <surname>Hardell</surname>
                     <given-names>C.</given-names>
                  </name>
                  <name>
                     <surname>Wu</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Fleischmann</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Harvey</surname>
                     <given-names>H.</given-names>
                  </name>
                  <name>
                     <surname>Folio</surname>
                     <given-names>L.&#x00A0;R.</given-names>
                  </name>
                  <name>
                     <surname>Summers</surname>
                     <given-names>R.&#x00A0;M.</given-names>
                  </name>
                  <name>
                     <surname>Rubin</surname>
                     <given-names>D.&#x00A0;L.</given-names>
                  </name>
                  <name>
                     <surname>Lungren</surname>
                     <given-names>M.&#x00A0;P.</given-names>
                  </name>
               </person-group>
               <year>2020</year>
               <article-title>Preparing medical imaging data for machine learning</article-title>
               <source>Radiology</source>
               <volume>295</volume>
               <fpage>4</fpage>
               <lpage>15</lpage>
               <page-range>4&#x2013;15</page-range>
               <pub-id pub-id-type="doi">10.1148/radiol.2020192224</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib81">
            <label>81</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Williams</surname>
                     <given-names>L.&#x00A0;H.</given-names>
                  </name>
                  <name>
                     <surname>Drew</surname>
                     <given-names>T.</given-names>
                  </name>
               </person-group>
               <year>2019</year>
               <article-title>What do we know about volumetric medical image interpretation?: A review of the basic science and medical image perception literatures</article-title>
               <source>Cogn. Res. Princ. Implications</source>
               <volume>4</volume>
               <fpage>1</fpage>
               <lpage>24</lpage>
               <page-range>1&#x2013;24</page-range>
               <pub-id pub-id-type="doi">10.1186/s41235-018-0149-9</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib82">
            <label>82</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Wolfe</surname>
                     <given-names>J.&#x00A0;M.</given-names>
                  </name>
                  <name>
                     <surname>Horowitz</surname>
                     <given-names>T.&#x00A0;S.</given-names>
                  </name>
               </person-group>
               <year>2017</year>
               <article-title>Five factors that guide attention in visual search</article-title>
               <source>Nature Hum. Behav.</source>
               <volume>1</volume>
               <fpage>1</fpage>
               <lpage>8</lpage>
               <page-range>1&#x2013;8</page-range>
               <pub-id pub-id-type="doi">10.1038/s41562-016-0001</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib83">
            <label>83</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Yan</surname>
                     <given-names>K.</given-names>
                  </name>
                  <name>
                     <surname>Wang</surname>
                     <given-names>X.</given-names>
                  </name>
                  <name>
                     <surname>Lu</surname>
                     <given-names>L.</given-names>
                  </name>
                  <name>
                     <surname>Summers</surname>
                     <given-names>R.&#x00A0;M.</given-names>
                  </name>
               </person-group>
               <year>2018</year>
               <article-title>Deeplesion: automated mining of large-scale lesion annotations and universal lesion detection with deep learning</article-title>
               <source>J. Med. Imaging</source>
               <volume>5</volume>
               <elocation-id content-type="artnum">036501</elocation-id>
               <pub-id pub-id-type="doi">10.1117/1.JMI.5.3.036501</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib84">
            <label>84</label>
            <element-citation publication-type="other"><person-group person-group-type="author">
                  <name>
                     <surname>Zbontar</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Knoll</surname>
                     <given-names>F.</given-names>
                  </name>
                  <name>
                     <surname>Sriram</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Murrell</surname>
                     <given-names>T.</given-names>
                  </name>
                  <name>
                     <surname>Huang</surname>
                     <given-names>Z.</given-names>
                  </name>
                  <name>
                     <surname>Muckley</surname>
                     <given-names>M.&#x00A0;J.</given-names>
                  </name>
                  <name>
                     <surname>Defazio</surname>
                     <given-names>A.</given-names>
                  </name>
                  <name>
                     <surname>Stern</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Johnson</surname>
                     <given-names>P.</given-names>
                  </name>
                  <name>
                     <surname>Bruno</surname>
                     <given-names>M.</given-names>
                  </name>
                  <name>
                     <surname>Parente</surname>
                     <given-names>M.</given-names>
                  </name>
               </person-group>
               <comment>&#x201C;fastmri: An open dataset and benchmarks for accelerated mri,&#x201D; Preprint arXiv:<ext-link ext-link-type="arxiv" xlink:href="http://arxiv.org/abs/1811.08839">1811.08839</ext-link>, (2018)</comment>
            </element-citation>
         </ref>
         <ref id="jpi0156bib85">
            <label>85</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Zhang</surname>
                     <given-names>B.</given-names>
                  </name>
                  <name>
                     <surname>Srihari</surname>
                     <given-names>S.&#x00A0;N.</given-names>
                  </name>
               </person-group>
               <year>2003</year>
               <article-title>Properties of binary vector dissimilarity measures</article-title>
               <source>Proc. JCIS Int&#x2019;l. Conf. Computer Vision, Pattern Recognition, and Image Processing</source>
               <volume>Vol.&#x00A0;1</volume>
               <publisher-name>Springer</publisher-name>
               <publisher-loc>Cham</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib86">
            <label>86</label>
            <element-citation publication-type="journal"><person-group person-group-type="author">
                  <name>
                     <surname>Zhang</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Xie</surname>
                     <given-names>Y.</given-names>
                  </name>
                  <name>
                     <surname>Wu</surname>
                     <given-names>Q.</given-names>
                  </name>
                  <name>
                     <surname>Xia</surname>
                     <given-names>Y.</given-names>
                  </name>
               </person-group>
               <year>2019</year>
               <article-title>Medical image classification using synergic deep learning</article-title>
               <source>Med. Image Anal.</source>
               <volume>54</volume>
               <fpage>10</fpage>
               <lpage>19</lpage>
               <page-range>10&#x2013;9</page-range>
               <pub-id pub-id-type="doi">10.1016/j.media.2019.02.010</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib87">
            <label>87</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Zhang</surname>
                     <given-names>R.</given-names>
                  </name>
                  <name>
                     <surname>Isola</surname>
                     <given-names>P.</given-names>
                  </name>
                  <name>
                     <surname>Efros</surname>
                     <given-names>A.&#x00A0;A.</given-names>
                  </name>
                  <name>
                     <surname>Shechtman</surname>
                     <given-names>E.</given-names>
                  </name>
                  <name>
                     <surname>Wang</surname>
                     <given-names>O.</given-names>
                  </name>
               </person-group>
               <year>2018</year>
               <article-title>The unreasonable effectiveness of deep features as a perceptual metric</article-title>
               <source>Proc. IEEE Conf. on Computer Vision and Pattern Recognition</source>
               <fpage>586</fpage>
               <lpage>595</lpage>
               <page-range>586&#x2013;95</page-range>
               <publisher-name>IEEE</publisher-name>
               <publisher-loc>Piscataway, NJ</publisher-loc>
               <pub-id pub-id-type="doi">10.1109/CVPR.2018.00068</pub-id>
            </element-citation>
         </ref>
         <ref id="jpi0156bib88">
            <label>88</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Zhu</surname>
                     <given-names>J.</given-names>
                  </name>
                  <name>
                     <surname>Shen</surname>
                     <given-names>Y.</given-names>
                  </name>
                  <name>
                     <surname>Zhao</surname>
                     <given-names>D.</given-names>
                  </name>
                  <name>
                     <surname>Zhou</surname>
                     <given-names>B.</given-names>
                  </name>
               </person-group>
               <year>2020</year>
               <article-title>In-domain gan inversion for real image editing</article-title>
               <source>European Conf. on Computer Vision</source>
               <fpage>592</fpage>
               <lpage>608</lpage>
               <page-range>592&#x2013;608</page-range>
               <publisher-name>Springer</publisher-name>
               <publisher-loc>Cham</publisher-loc>
            </element-citation>
         </ref>
         <ref id="jpi0156bib89">
            <label>89</label>
            <element-citation publication-type="book"><person-group person-group-type="author">
                  <name>
                     <surname>Zhu</surname>
                     <given-names>J.-Y.</given-names>
                  </name>
                  <name>
                     <surname>Kr&#x00E4;henb&#x00FC;hl</surname>
                     <given-names>P.</given-names>
                  </name>
                  <name>
                     <surname>Shechtman</surname>
                     <given-names>E.</given-names>
                  </name>
                  <name>
                     <surname>Efros</surname>
                     <given-names>A.&#x00A0;A.</given-names>
                  </name>
               </person-group>
               <year>2016</year>
               <article-title>Generative visual manipulation on the natural image manifold</article-title>
               <source>European Conf. on Computer Vision</source>
               <fpage>597</fpage>
               <lpage>613</lpage>
               <page-range>597&#x2013;613</page-range>
               <publisher-name>Springer</publisher-name>
               <publisher-loc>Cham</publisher-loc>
            </element-citation>
         </ref>
      </ref-list>
   </back>
</article>