<?xml version="1.0"?>
                <!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "journalpublishing3.dtd">
                <article article-type="research-article" xmlns:mml="http://www.w3.org/1998/Math/MathML"
                xmlns:xlink="http://www.w3.org/1999/xlink"
                xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
                dtd-version="3.0">
                <front>
                    <journal-meta>
                    <journal-id journal-id-type="publisher-id">ei</journal-id>
                    <journal-title>Electronic Imaging</journal-title>
                    <issn pub-type="ppub">2470-1173</issn><issn pub-type="epub">2470-1173</issn>
                    <publisher>
                        <publisher-name>Society for Imaging Science and Technology</publisher-name>
                        <publisher-loc>IS&amp;T 7003 Kilworth Lane, Springfield, VA 22151 USA</publisher-loc>
                    </publisher>
                    </journal-meta>
                    <article-meta>
                    <article-id pub-id-type="doi">10.2352/EI.2025.37.12.HPCI-183</article-id>
                    <article-id pub-id-type="publisher-id">HPCI-183</article-id>
                    <article-categories>
                        <subj-group>
                        <subject>Proceedings Paper</subject>
                        </subj-group>
                    </article-categories>
                    <title-group>
                        <article-title>mTREE: Multi-level Text-guided Representation End-to-end Learning for Whole Slide Image Analysis</article-title>
                    </title-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Liu</surname>
                            <given-names>Quan </given-names>
                           </name> <xref ref-type="aff" rid="aff1author1"/></contrib><aff id="aff1author1">Vanderbilt University, US</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Deng</surname>
                            <given-names>Ruining </given-names>
                           </name> <xref ref-type="aff" rid="aff1author2"/></contrib><aff id="aff1author2">Vanderbilt University, US</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Cui</surname>
                            <given-names>Can </given-names>
                           </name> <xref ref-type="aff" rid="aff1author3"/></contrib><aff id="aff1author3">Vanderbilt University, US</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Yao</surname>
                            <given-names>Tianyuan </given-names>
                           </name> <xref ref-type="aff" rid="aff1author4"/></contrib><aff id="aff1author4">Vanderbilt University, US</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Yang</surname>
                            <given-names>Yuechen </given-names>
                           </name> <xref ref-type="aff" rid="aff1author5"/></contrib><aff id="aff1author5">Vanderbilt University, US</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Nath</surname>
                            <given-names>Vishwesh </given-names>
                           </name> <xref ref-type="aff" rid="aff2author6"/></contrib><aff id="aff2author6">NVIDIA</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Li</surname>
                            <given-names>Bingshan </given-names>
                           </name> <xref ref-type="aff" rid="aff1author7"/></contrib><aff id="aff1author7">Vanderbilt University, US</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Chen</surname>
                            <given-names>You </given-names>
                           </name> <xref ref-type="aff" rid="aff1author8"/></contrib><aff id="aff1author8">Vanderbilt University, US</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Tang</surname>
                            <given-names>Yucheng </given-names>
                           </name> <xref ref-type="aff" rid="aff2author9"/></contrib><aff id="aff2author9">NVIDIA</aff></contrib-group><contrib-group content-type="all"><contrib contrib-type="author"><name>
                            <surname>Huo</surname>
                            <given-names>Yuankai </given-names>
                           </name> <xref ref-type="aff" rid="aff1author10"/></contrib><aff id="aff1author10">Vanderbilt University, US</aff></contrib-group><abstract>
                    <title>Abstract</title>
                    <p>Multi-modal learning adeptly integrates visual and textual data, but its application to histopathology image and text analysis remains challenging, particularly with large, high-resolution images like gigapixel Whole Slide Images (WSIs). Current methods typically rely on manual region labeling or multi-stage learning to assemble local representations (e.g., patch-level) into global features (e.g., slide-level). However, there is no effective way to integrate multi-scale image representations with text data in a seamless end-to-end process. In this study, we introduce Multi-Level Text-Guided Representation End-to-End Learning (mTREE). This novel text-guided approach effectively captures multi-scale WSI representations by utilizing information from accompanying textual pathology information. mTREE innovatively combines – the localization of key areas (“global-tolocal”) and the development of a WSI-level image-text representation (“local-to-global”) – into a unified, end-to-end learning framework. In this model, textual information serves a dual purpose: firstly, functioning as an attention map to accurately identify key areas, and secondly, acting as a conduit for integrating textual features into the comprehensive representation of the image. Our study demonstrates the effectiveness of mTREE through quantitative analyses in two image-related tasks: classification and survival prediction, showcasing its remarkable superiority over baselines. Code and trained models are made available at https://github.com/hrlblab/mTREE.</p>
                    </abstract><pub-date>
                        <day>2</day>
                        <month>2</month>
                        <year>2025</year>
                        </pub-date><volume>37</volume>
                    <issue-acronym>HPCI</issue-acronym>
                    <issue-title>High Performance Computing for Imaging 2025</issue-title>
                    <issue seq="183">12</issue>
                    <fpage>183-1</fpage>
                    <lpage>183-7</lpage>
                    <permissions>
                         <copyright-statement>© 2025 Society for Imaging Science and Technology</copyright-statement>
                        <copyright-year>2025</copyright-year>
                    </permissions><kwd-group><kwd>Visual language model</kwd><kwd>Representation leaning</kwd><kwd>Prognosis analysis</kwd><kwd>Pathology</kwd></kwd-group></article-meta>
                </front>
                </article>