arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2607.26799v2 [cs.CV] 07 Aug 2026

[ orcid=0009-0004-2081-0964 ]

[ orcid=0009-0008-1245-5113 ]

mode = titlePRISM-Net for Breast DCE-MRI

PRISM-Net: Patient-specific reference-guided inter-breast symmetry matching for three-class breast DCE-MRI classification

Boya Zhang boyazhang@mail.nankai.edu.cn    Shuaiwen Zhou qdzsw666@bupt.edu.cn    Di Kong kd24@mails.tsinghua.edu.cn    Mingxu Wang wmxwork999@163.com    Wenbiao Du duwenbiao@bit.edu.cn    Yiman Zhong yeeman@buaa.edu.cn    Yuexin Duan duanyuexiner@163.com    Xiawei Yue yxw@mail.nankai.edu.cn    Liuquan Cheng 13910209982@139.com    Xiru Li 2468li@sina.com organization=Nankai University, addressline=No. 94 Weijin Road, Nankai District, city=Tianjin, postcode=300071, country=China organization=Zhongguancun Academy, city=Beijing, country=China organization=The First Medical Center of Chinese PLA General Hospital, city=Beijing, country=China organization=Beijing University of Posts and Telecommunications, city=Beijing, country=China organization=The Six Medical Center of Chinese PLA General Hospital, city=Beijing, country=China organization=Tsinghua University, city=Beijing, country=China organization=Beijing Institute of Technology, city=Beijing, country=China organization=Beihang University, city=Beijing, country=China
Abstract

Breast DCE-MRI classification of breast sides as no lesion, benign, or malignant is complicated by substantial patient-specific variation in breast anatomy, tissue composition, and physiological enhancement. Although radiologists routinely use the contralateral breast as an internal reference, computational bilateral comparison remains challenging because corresponding contralateral regions are not anatomically identical and bilateral differences are not necessarily pathological. We propose PRISM-Net, a registration-free bilateral framework that treats the contralateral breast as an approximate patient-specific reference. PRISM-Net learns complementary target-appearance and reference-conditioned deviation representations. Adaptive patch-level matching constructs content-dependent contralateral references in a shared feature space without voxel-wise registration, while local-contrast reweighting emphasizes focal deviations over diffuse bilateral mismatch. The resulting deviation features are integrated with intrinsic target-side appearance through residual cross-attention and aggregated across slices for breast-side classification. On ODELIA, PRISM-Net achieved Macro AUC, Micro AUC, and quadratic weighted kappa values of 84.11±2.3384.11\pm 2.33, 90.64±1.6190.64\pm 1.61, and 60.94±5.6460.94\pm 5.64 on the in-distribution test set, and 68.51±4.5468.51\pm 4.54, 80.74±2.6880.74\pm 2.68, and 43.45±7.1043.45\pm 7.10 on the held-out institution, respectively, yielding the highest point estimates among the evaluated methods across the primary metrics. PRISM-Net also demonstrated favorable performance on an independent institutional cohort and background-complexity subsets, while ablation experiments supported the contributions of adaptive reference matching and focal deviation modeling. These findings support patient-specific contralateral reference learning as a clinically grounded strategy for breast DCE-MRI classification under heterogeneous background conditions.

keywords
Breast DCE-MRI ,Patient specific contralateral reference learning ,Radiological prior knowledge ,Medical image classification ,Background parenchymal enhancement
credit: Conceptualization, Methodology, Investigation, Data curation, Validation, Formal analysis, Funding acquisition, Writing – original draftcredit: Conceptualization, Methodology, Software, Investigation, Formal analysis, Validation, Visualization, Data curation, Writing – original draftcredit: Methodology, Supervision, Conceptualization, Review and editingcredit: Data curation, Review and editingcredit: Data curation, Review and editingcredit: Data curation, Review and editingcredit: Data curation, Review and editingcredit: Data curation, Review and editingcredit: Supervision, Resources, Project administration, Conceptualization, Review and editingcredit: Supervision, Funding acquisition, Resources, Project administration, Review and editingThese authors contributed equally to this work.corresponding: Corresponding author. E-mail: 13910209982@139.comcorresponding: Corresponding author. E-mail: 2468li@sina.com

1 Introduction

Breast cancer remains a leading cause of cancer-related mortality among women worldwide [7, 35]. Dynamic contrast-enhanced breast magnetic resonance imaging (DCE-MRI) is highly sensitive for breast cancer detection and is widely used for screening, diagnosis, staging, and treatment assessment [1, 22, 17]. Nevertheless, its interpretation remains challenging because breast anatomy, tissue composition, physiological activity, and imaging appearance vary substantially across patients. These patient-specific background signals may resemble suspicious findings or reduce the conspicuity of true abnormalities, contributing to both over-calling and under-calling [9, 33, 6]. Separating lesion-related evidence from background-signal interference is therefore a fundamental problem in breast DCE-MRI interpretation.

Refer to caption
Figure 1: Representative clinical examples demonstrating bilateral comparison in breast DCE-MRI. Each row presents unilateral and bilateral views, with paired target-side and contralateral reference crops. (A) Typical mass lesion in a 5656-year-old woman. An irregular, mildly lobulated right upper-outer breast mass showed asymmetric heterogeneous enhancement. The lesion was assessed as BI-RADS 55, and biopsy confirmed invasive carcinoma. (B) Symmetric BPE mimic in a 2323-year-old woman with dense fibroglandular tissue and moderate symmetric BPE. Suspicious enhancement on the unilateral view had a corresponding contralateral pattern, supporting symmetric BPE; both breasts were assessed as BI-RADS 11(no suspicious enhancement). (C) Comparison-clarified NME in a 4444-year-old woman. Asymmetric progressive enhancement in the left upper breast comprised NME with multiple small satellite masses. Bilateral comparison showed no corresponding contralateral enhancement, clarifying the finding as NME. The lesion was assessed as BI-RADS 55, and pathology confirmed invasive carcinoma. (D) Background-obscured lesion in a 3939-year-old woman with dense fibroglandular tissue and moderate symmetric BPE. A 1010-mm right upper-inner breast mass could be overlooked among surrounding enhancement. The absent contralateral counterpart facilitated detection; the right breast was assessed as BI-RADS 33, and pathology confirmed fibroadenoma. BPE, background parenchymal enhancement; DCE-MRI, dynamic contrast-enhanced magnetic resonance imaging; NME, non-mass enhancement.

Bilateral comparison provides a clinically intuitive strategy for reducing background-signal interference in breast DCE-MRI. Also, previous computer-aided and multiparametric imaging studies indicated that bilateral asymmetry and contralateral healthy-tissue biomarkers can provide complementary diagnostic information [41, 23]. This prior is particularly informative in DCE-MRI, where contrast uptake occurs in both lesions and normal fibroglandular tissue. In the BI-RADS lexicon, fibroglandular tissue (FGT) characterizes the amount and distribution of glandular tissue, whereas background parenchymal enhancement (BPE) describes its physiological enhancement after contrast administration. FGT and BPE therefore provide complementary descriptors of background-signal complexity, which has been associated with reduced diagnostic performance [6, 3, 24, 5]. As illustrated in Fig. 1, using the contralateral breast as a patient-specific internal reference can help dismiss bilaterally similar mimics, clarify abnormalities without a contralateral counterpart, and reveal lesions obscured by surrounding background. The diagnostic value of bilateral comparison therefore lies not in asymmetry detection alone, but in leveraging shared physiological characteristics to approximate a patient-specific baseline for interpreting suspicious deviations.

Although deep learning studies in mammography and digital breast tomosynthesis have demonstrated the diagnostic value of bilateral information [34, 40, 42, 32], the clinical prior of bilateral comparison remains underused in deep learning models for breast MRI [13, 26, 11, 10, 31, 43]. At the modeling level, many methods remain lesion-centric, single-sided, or only implicitly bilateral. Fine-grained bilateral modeling requires correspondence between target regions and contralateral tissue. Conventional correspondence estimation often relies on deformable registration, which remains challenging in breast DCE-MRI because of nonrigid breast motion and enhancement-related temporal intensity changes [44, 38, 4, 18, 37]. At the task level, most breast MRI classification models have been developed in lesion-enriched binary settings. However, as breast MRI expands across screening and diagnostic applications, examinations without suspicious findings constitute a clinically relevant part of the case spectrum [27, 39]. The recently released multicenter ODELIA resource provides breast-side labels for no-lesion, benign, and malignant findings, yet studies in this setting remain scarce [28, 30]. Taken together, the central gap is the lack of an end-to-end breast MRI framework that jointly learns registration-free, spatially adaptive cross-breast correspondence, treats the contralateral breast as an approximate patient specific reference rather than a normal template, and distinguishes diagnostically relevant focal deviations from diffuse or bilaterally shared variation while preserving intrinsic target side appearance. This gap is particularly consequential for breast side three class classification, where the model must distinguish no lesion, benign, and malignant findings under heterogeneous anatomical and physiological backgrounds without assuming that a target lesion has already been identified.

To address this gap, we formulate bilateral DCE-MRI analysis as patient-specific reference-guided interpretation, in which the contralateral breast provides approximate internal context for identifying localized deviations from shared anatomical and physiological background patterns. We therefore propose PRISM-Net, a patient-specific reference-guided inter-breast symmetry matching for three-class breast DCE-MRI classification without explicit anatomical registration. PRISM-Net encodes the target and contralateral breasts using a shared-weight backbone to establish a common feature space. Bidirectional patch-level cross-breast matching then retrieves an adaptive contralateral reference for each target region under anatomical variability. The model integrates target features, matched references, directional differences, and cross-breast interactions to construct reference-conditioned representations of bilateral deviations. A local-contrast reweighting mechanism further suppresses diffuse or bilaterally shared variation while emphasizing diagnostically relevant focal asymmetries. The resulting deviation evidence is fused with intrinsic target-side appearance and aggregated across slices to classify each breast as no-lesion, benign, or malignant. By converting clinical bilateral comparison into a learnable diagnostic prior, PRISM-Net provides a unified framework for patient-specific reference modeling and breast-side classification..

We evaluate PRISM-Net across the public multicenter ODELIA dataset [30] and independent institutional DCE-MRI cohorts, all with breast-side labels for no-lesion, benign, and malignant findings. Bilateral no-lesion examinations supported by clinical BI-RADS assessments are retained to represent lesion-free settings, while dedicated BPE- and FGT-complexity cohorts capture challenging backgrounds. Together, this multi-cohort design provides a clinically oriented assessment of three-class performance, cross-cohort generalizability, and robustness to background-signal complexity.

The contributions of this work are as follows.

  1. 1.

    We formulate breast-side DCE-MRI classification as a patient-specific bilateral reference learning problem. The contralateral breast is treated as an approximate internal reference for distinguishing shared anatomical and physiological background patterns from localized lesion-related deviations across no-lesion, benign, and malignant findings.

  2. 2.

    We propose PRISM-Net, a registration-free bilateral framework that learns complementary target-appearance and reference-conditioned deviation representations. Its adaptive patch-level matching mechanism constructs content-dependent contralateral references in a shared feature space, avoiding fixed mirrored correspondence and deformable registration.

  3. 3.

    We introduce local-contrast focal deviation modeling and appearance-preserving fusion. The former emphasizes spatially localized target-to-reference deviations over diffuse bilateral mismatch, whereas the latter integrates these cues with intrinsic target-side appearance through a residual pathway for volumetric classification.

  4. 4.

    We evaluate PRISM-Net on the public multicenter ODELIA dataset, a held-out external ODELIA institution, and independent institutional cohorts including BPE- and FGT-complexity subsets. The results show favorable discrimination and ordinal agreement relative to the evaluated baselines, and ablation analyses support the contributions of adaptive contralateral matching and focal deviation modeling.

2 Related work

2.1 Deep learning in breast MRI diagnosis

Existing deep learning research in breast MRI has predominantly focused on lesion-level or binary diagnostic tasks, commonly using localized lesion regions or unilateral breast inputs.

Deep learning in breast MRI has predominantly focused on lesion-level characterization. CNNs, recurrent networks, projection-based models, and multi-task architectures have been used to distinguish benign from malignant lesions from manually segmented, radiologist-selected, or automatically detected regions. Representative approaches include multiparametric CNN classification, four-dimensional CNN–LSTM aggregation, deep-feature maximum-intensity projections, and joint lesion segmentation and classification [38, 4, 18, 19, 37]. DAE-CNN and its PBPK-informed extension further incorporate contrast-agent disentanglement and physiologically constrained data augmentation to improve lesion representation under limited training data [11, 10]. Despite differences in architecture and input representation, these methods generally assume that a target lesion has already been identified.

Existing breast MRI diagnostic models also predominantly adopt binary task formulations. Lesion-level studies commonly distinguish benign from malignant findings, whereas breast-level or volume-level models usually predict malignant versus non-malignant outcomes [29, 17]. Such formulations either exclude breasts without detectable lesions or combine no-lesion and benign cases within a single negative category. The ODELIA dataset and benchmark provide a clinically relevant task setting by assigning no-lesion, benign, and malignant labels to each breast in multicentre DCE-MRI examinations [28, 30]. Its Medical Slice Transformer baseline aggregates slice-level features from pre-contrast T1-weighted, first post-contrast subtraction, and T2-weighted volumes, establishing breast-side three-class diagnosis as a direct precedent [28].Because the two breasts are treated as separate samples, the model does not explicitly learn cross-breast interactions.

Taken together, current breast MRI deep-learning methods remain largely shaped by lesion-level inputs and binary label spaces. Model development would benefit from prediction units and label definitions that better reflect clinical diagnostic tasks. Breast-side three-class classification addresses this requirement by jointly distinguishing no-lesion, benign, and malignant states without presuming that a target lesion has already been identified.

2.2 Bilateral comparison and symmetry-aware learning in breast imaging

Bilateral comparison is a routine interpretive strategy in breast imaging, allowing local morphology and enhancement to be evaluated against a patient-specific internal reference. Because the contralateral breast may itself contain abnormalities and bilateral anatomy is affected by natural asymmetry, nonrigid deformation, positioning variation, and heterogeneous background enhancement, bilateral symmetry is an approximate prior and direct voxel-wise correspondence is unreliable [2, 16, 21, 24, 42]. Early breast MRI studies quantified bilateral symmetry using registration-free handcrafted three-dimensional descriptors or predefined whole-breast kinetic differences, but remained limited to global representations [2, 41].

In subsequent breast MRI deep learning studies, bilateral information has not been entirely absent, but it was often incorporated implicitly. Zhang et al. proposed a two-stage framework using Mask R-CNN for suspicious lesion detection followed by ResNet50 for malignancy classification in breast MRI [43]. In the detection stage, the model used pre-contrast images and subtraction images of the left and right breasts as inputs, allowing breast symmetry to be considered within the whole-breast context. However, the method did not explicitly establish correspondence between the two breasts, nor did it construct asymmetry features, a contralateral reference branch, bilateral attention, or a dedicated bilateral fusion module. The subsequent classification stage mainly focused on local candidate regions detected by Mask R-CNN and used DCE parametric maps for lesion characterization. Therefore, this type of work suggested that breast MRI deep learning had recognized the importance of bilateral information, but such information was mainly incorporated as global contextual input. Within BreastMRI-FCDD, which performed breast-level binary cancer detection on DCE-MRI subtraction maximum intensity projections, Oviedo et al. [31] further investigated FCDD-Symmetric, in which the contralateral breast representation defined the normal-class center for anomaly scoring. However, its projection-level, center-based formulation did not explicitly establish spatially adaptive correspondence between bilateral three-dimensional feature volumes or use spatially matched contralateral features for background signal baseline normalization.

In mammography, digital breast tomosynthesis, and other breast imaging modalities, bilateral information has been more explicitly embedded into deep learning models. Earlier studies introduced distortion-insensitive, logic-guided bilateral comparison and paired asymmetry modeling for cancer localization and classification, while BilAD subsequently extended bilateral asymmetry learning to DBT [34, 25, 12]. DisAsymNet attempted to disentangle asymmetric abnormalities from symmetric breast structures in bilateral mammograms, reconstructing abnormality-removed symmetric images to improve interpretability [40]. A recent bilateral information-guided Vision Transformer framework further proposed a registration-free strategy for mammography, directly concatenating bilateral images and introducing soft spatial prompts to model cross-breast structural differences without explicit image registration [42]. BiGAM-Net also used asymmetry-aware fusion and inspectable gating mechanisms to quantify model reliance on intrinsic lesion features and contralateral asymmetry evidence in paired breast imaging [14]. These studies collectively suggest that bilateral information can serve as a structural medical prior to improve diagnostic performance and interpretability.

These studies support the value of explicit bilateral reasoning, while their two-dimensional or quasi-two-dimensional formulations provide limited coverage of the deformable volumetric anatomy and heterogeneous enhancement of breast DCE-MRI. Registration-free, spatially adaptive cross-breast matching coupled with patient-specific background normalization therefore remains the central methodological gap.

3 Methods

PRISM-Net translates radiological bilateral comparison into a registration-free framework that conditions each target breast on an adaptively matched contralateral reference, as illustrated in Fig. 2. We first formulate bilateral breast-side classification and introduce the clinical rationale for patient-specific reference learning in Sec. 3.1. We then establish a shared bilateral representation space in Sec. 3.2 and construct adaptive patch level contralateral references without voxel-wise registration in Sec. 3.3. Next, Sec. 3.4 introduces focal bilateral-deviation modeling to attenuate diffuse mismatch and summarize localized reference-conditioned evidence. Finally, Sec. 3.5 integrates this evidence with target-side appearance and aggregates it across slices for breast-side three-class classification.

Refer to caption
Figure 2: Overall architecture of the proposed PRISM-Net for breast DCE-MRI classification.

3.1 Bilateral problem formulation

PRISM-Net is motivated by the clinical use of bilateral comparison in breast DCE-MRI. Rather than interpreting a target-side finding solely from its absolute appearance, radiologists compare it with contralateral tissue to determine whether it represents a meaningful deviation from patient-specific anatomical and physiological background patterns. Because the contralateral breast is neither anatomically identical to the target breast nor guaranteed to be lesion-free, we treat it as an approximate patient-specific reference rather than a normal template.

Operationalizing this reference-based interpretation raises three modeling challenges. First, corresponding bilateral regions are not necessarily located at identical mirrored coordinates because of nonrigid breast deformation, patient positioning, and anatomical asymmetry, making direct left-right subtraction unreliable. Second, bilateral mismatch may arise from moderate/marked BPE, heterogeneous/dense FGT, or other natural asymmetry rather than local pathology. Third, bilateral asymmetry alone is insufficient for three-class diagnosis, because distinguishing no-lesion, benign, and malignant findings also requires intrinsic target-side appearance and cross-slice volumetric context.

Accordingly, PRISM-Net learns two complementary representations for each breast: a target-appearance representation that preserves intrinsic diagnostic features, and a reference-conditioned deviation representation that characterizes how local target features differ from adaptively matched contralateral context.

Formally, let 𝐗L\mathbf{X}^{L} and 𝐗R\mathbf{X}^{R} denote the paired left and right breast DCE-MRI volumes from one examination, with 𝐗L,𝐗RD×H×W\mathbf{X}^{L},\mathbf{X}^{R}\in\mathbb{R}^{D\times H\times W}, where DD is the number of axial slices and HH and WW are the in-plane spatial dimensions. Let 𝒮={L,R}\mathcal{S}=\{L,R\} denote the set of breast sides. For each side a𝒮a\in\mathcal{S}, the diagnostic label is ya{0,1,2}y^{a}\in\{0,1,2\}, corresponding to no-lesion, benign, and malignant findings, respectively. Given the paired bilateral volumes, PRISM-Net learns a function fθf_{\theta} that jointly predicts the side-specific class probability distributions (𝝅L,𝝅R)=fθ(𝐗L,𝐗R)(\bm{\pi}^{L},\bm{\pi}^{R})=f_{\theta}(\mathbf{X}^{L},\mathbf{X}^{R}), where 𝝅a[0,1]3\bm{\pi}^{a}\in[0,1]^{3} and c=02πca=1\sum_{c=0}^{2}\pi_{c}^{a}=1 for each a𝒮a\in\mathcal{S}.

3.2 Shared bilateral representation

To make the approximate contralateral reference computationally usable, the bilateral volumes must first be mapped into a comparable semantic space. PRISM-Net horizontally flips the right breast volume, denoted by 𝐗¯R=Flip(𝐗R)\bar{\mathbf{X}}^{R}=\operatorname{Flip}(\mathbf{X}^{R}), to place both breasts in a common orientation. This operation provides only a coarse spatial prior and does not assume voxel-wise correspondence. The left volume 𝐗L\mathbf{X}^{L} and the flipped right volume 𝐗¯R\bar{\mathbf{X}}^{R} are then encoded slice-wise using a shared DINOv3 ViT-S/16 backbone [36]. Weight sharing ensures that bilateral features are represented in the same semantic space, thereby enabling subsequent feature-space matching.

For the dd-th slice and breast side a𝒮a\in\mathcal{S}, the shared encoder produces a global slice token and a matrix of local patch tokens:

𝐬daC,𝐏daN×C,\mathbf{s}_{d}^{a}\in\mathbb{R}^{C},\qquad\mathbf{P}_{d}^{a}\in\mathbb{R}^{N\times C}, (1)

where 𝐬da\mathbf{s}_{d}^{a} preserves the global target-side appearance, 𝐏da\mathbf{P}_{d}^{a} contains the corresponding local patch representations, CC is the feature dimension, and NN is the number of image patches. The patch-level representations are subsequently used to construct adaptive contralateral references, whereas the global slice tokens are retained for appearance-preserving classification.

3.3 Registration-free adaptive contralateral reference matching

With bilateral features represented in a common semantic space, the next step is to construct a patient-specific reference for each target patch. Because identical mirrored coordinates do not guarantee anatomical correspondence, direct patch-wise subtraction may confound pathological differences with variations in breast shape, positioning, and deformation. PRISM-Net therefore allows each target patch to retrieve an adaptive contralateral reference through soft feature-space matching within the paired axial slice, without requiring voxel-wise registration.

Before matching, unreliable regions are excluded using binary masks. Let 𝐌L,𝐌R{0,1}D×H×W\mathbf{M}^{L},\mathbf{M}^{R}\in\{0,1\}^{D\times H\times W} denote the exclusion masks corresponding to the original bilateral volumes, where 11 indicates a region excluded from matching. The right-side mask is flipped together with the right breast volume, 𝐌¯R=Flip(𝐌R)\bar{\mathbf{M}}^{R}=\operatorname{Flip}(\mathbf{M}^{R}), and the bilateral exclusion mask in the common orientation is defined as 𝐌U=𝐌L𝐌¯R\mathbf{M}^{U}=\mathbf{M}^{L}\vee\bar{\mathbf{M}}^{R}. For slice dd, the valid patch set is

𝒱d={i|max(u,v)𝒪iMd,u,vU=0},\mathcal{V}_{d}=\left\{i\;\middle|\;\max_{(u,v)\in\mathcal{O}_{i}}M^{U}_{d,u,v}=0\right\}, (2)

where 𝒪i\mathcal{O}_{i} denotes the spatial support of patch ii. This filtering ensures that bilateral matching is performed only between patches that are reliable on both sides.

For one matching direction, let a𝒮a\in\mathcal{S} denote the target side and let b𝒮b\in\mathcal{S}, bab\neq a, denote the contralateral reference side. Based on the valid patch-token matrices, the matching weights and adaptive contralateral references are computed as

𝐀dab\displaystyle\mathbf{A}_{d}^{a\leftarrow b} =softmax((𝐏d,𝒱da𝐖Q)(𝐏d,𝒱db𝐖K)dk),\displaystyle=\operatorname{softmax}\left(\frac{(\mathbf{P}_{d,\mathcal{V}_{d}}^{a}\mathbf{W}_{Q})(\mathbf{P}_{d,\mathcal{V}_{d}}^{b}\mathbf{W}_{K})^{\top}}{\sqrt{d_{k}}}\right), (3)
𝐑da\displaystyle\mathbf{R}_{d}^{a} =𝐀dab(𝐏d,𝒱db𝐖V),\displaystyle=\mathbf{A}_{d}^{a\leftarrow b}(\mathbf{P}_{d,\mathcal{V}_{d}}^{b}\mathbf{W}_{V}),

where dkd_{k} is the projected query–key dimension and 𝐖Q\mathbf{W}_{Q}, 𝐖K\mathbf{W}_{K}, and 𝐖V\mathbf{W}_{V} are learnable projection matrices. The softmax operation is applied over the valid contralateral patches, so each row of 𝐀dab\mathbf{A}_{d}^{a\leftarrow b} describes the correspondence distribution between one target patch and the candidate reference patches. Consequently, each row 𝐫d,ia\mathbf{r}_{d,i}^{a} of 𝐑da\mathbf{R}_{d}^{a} represents a soft, content-adaptive contralateral reference for the corresponding target token 𝐩d,ia\mathbf{p}_{d,i}^{a}, rather than the feature at a fixed mirrored location. The same matching operation is applied in both directions with shared parameters. These adaptively matched reference pairs provide the basis for bilateral relation modeling. However, their discrepancies are not uniformly diagnostic.

3.4 Focal bilateral-deviation modeling

A large target-to-reference difference does not necessarily signify pathology, as it may arise from natural asymmetry in breast morphology, FGT distribution, asymmetric BPE, hormonal or lactational influences, or inflammatory and edematous changes. Therefore, after obtaining the adaptive contralateral reference, PRISM-Net preserves both the target and reference representations and explicitly encodes their bilateral relationship.

For each valid patch i𝒱di\in\mathcal{V}_{d}, the bilateral relation representation is constructed as

𝐜d,ia\displaystyle\mathbf{c}_{d,i}^{a} =Concat(𝐩d,ia,𝐫d,ia,𝐩d,ia𝐫d,ia,𝐩d,ia𝐫d,ia),\displaystyle=\operatorname{Concat}\left(\mathbf{p}_{d,i}^{a},\mathbf{r}_{d,i}^{a},\mathbf{p}_{d,i}^{a}-\mathbf{r}_{d,i}^{a},\mathbf{p}_{d,i}^{a}\odot\mathbf{r}_{d,i}^{a}\right), (4)
𝐠d,ia\displaystyle\mathbf{g}_{d,i}^{a} =ϕ(𝐜d,ia),\displaystyle=\phi\left(\mathbf{c}_{d,i}^{a}\right),

where 𝐩d,ia\mathbf{p}_{d,i}^{a} preserves target-side appearance, 𝐫d,ia\mathbf{r}_{d,i}^{a} provides patient-specific contralateral context, their difference encodes the directional target-to-reference deviation, and their element-wise product captures feature-wise bilateral interaction. The projection layer ϕ()\phi(\cdot) maps the concatenated representation to a CC-dimensional relation feature 𝐠d,ia\mathbf{g}_{d,i}^{a}.

Although the relation representation captures bilateral differences, diffuse mismatch may reflect background variation rather than focal pathology. PRISM-Net therefore measures how strongly each patch-level discrepancy stands out from its local spatial neighborhood. Let 𝒩d(i)=𝒩(i)𝒱d\mathcal{N}_{d}(i)=\mathcal{N}(i)\cap\mathcal{V}_{d} denote the valid patches within the 3×33\times 3 neighborhood of patch ii. The local-contrast asymmetry score and reweighted relation feature are computed as

Ad,ia\displaystyle A_{d,i}^{a} =𝐩d,ia𝐫d,ia2,\displaystyle=\left\|\mathbf{p}_{d,i}^{a}-\mathbf{r}_{d,i}^{a}\right\|_{2}, (5)
A¯d,ia\displaystyle\bar{A}_{d,i}^{a} =1|𝒩d(i)|j𝒩d(i)Ad,ja,\displaystyle=\frac{1}{|\mathcal{N}_{d}(i)|}\sum_{j\in\mathcal{N}_{d}(i)}A_{d,j}^{a},
Fd,ia\displaystyle F_{d,i}^{a} =ReLU(Ad,iaA¯d,ia),\displaystyle=\operatorname{ReLU}\left(A_{d,i}^{a}-\bar{A}_{d,i}^{a}\right),
F~d,ia\displaystyle\widetilde{F}_{d,i}^{a} =Fd,iaminj𝒱dFd,jamaxj𝒱dFd,jaminj𝒱dFd,ja+ϵ,\displaystyle=\frac{F_{d,i}^{a}-\min_{j\in\mathcal{V}_{d}}F_{d,j}^{a}}{\max_{j\in\mathcal{V}_{d}}F_{d,j}^{a}-\min_{j\in\mathcal{V}_{d}}F_{d,j}^{a}+\epsilon},
𝐠^d,ia\displaystyle\widehat{\mathbf{g}}_{d,i}^{a} =F~d,ia𝐠d,ia.\displaystyle=\widetilde{F}_{d,i}^{a}\mathbf{g}_{d,i}^{a}.

This local-contrast formulation emphasizes discrepancies that are salient relative to their surroundings while attenuating spatially diffuse bilateral mismatch.

Because diagnostically relevant deviations may be sparse or distributed across multiple regions, direct global averaging could dilute their contribution. We therefore stack the reweighted relation features as 𝐆^da|𝒱d|×C\widehat{\mathbf{G}}_{d}^{a}\in\mathbb{R}^{|\mathcal{V}_{d}|\times C} and summarize them using a shared set of learnable asymmetry queries 𝐐AK×C\mathbf{Q}_{A}\in\mathbb{R}^{K\times C}:

𝐓da=CA(𝐐A,𝐆^da,𝐆^da),𝐓daK×C.\mathbf{T}_{d}^{a}=\operatorname{CA}\left(\mathbf{Q}_{A},\widehat{\mathbf{G}}_{d}^{a},\widehat{\mathbf{G}}_{d}^{a}\right),\qquad\mathbf{T}_{d}^{a}\in\mathbb{R}^{K\times C}. (6)

The complete relation-modeling module is applied bidirectionally with shared parameters. The resulting tokens remain side-specific because the target and contralateral reference roles are reversed in the two matching directions. The resulting side-specific asymmetry tokens summarize focal reference-conditioned deviations, but they do not replace the intrinsic appearance of the target breast. We therefore integrate them with the global target-side slice representations for final volumetric classification.

3.5 Appearance-preserving volumetric fusion and classification

The preceding module produces side-specific tokens that summarize focal bilateral deviations. For diagnosis, however, this information should complement rather than replace intrinsic target-side appearance, because symmetric findings are not necessarily normal and asymmetric findings are not necessarily pathological. Moreover, the contralateral breast may itself contain abnormalities. PRISM-Net therefore retains the target-side slice token as the primary appearance representation and introduces the bilateral-deviation tokens through a residual cross-attention pathway.

For side a𝒮a\in\mathcal{S} and slice dd, the reference-conditioned slice representation is computed as

𝐬~da=𝐬da+λCA(𝐬da,𝐓da,𝐓da),\widetilde{\mathbf{s}}_{d}^{a}=\mathbf{s}_{d}^{a}+\lambda\operatorname{CA}\left(\mathbf{s}_{d}^{a},\mathbf{T}_{d}^{a},\mathbf{T}_{d}^{a}\right), (7)

where λ\lambda is a fixed residual scaling hyperparameter that controls the contribution of the bilateral-deviation information. The residual scaling factor λ\lambda was fixed at 0.80.8 throughout training and evaluation. This residual formulation preserves the intrinsic target-side appearance while allowing the slice representation to be selectively conditioned on patient-specific contralateral context.

Because diagnostically relevant evidence may extend across multiple slices, the reference-conditioned slice representations are subsequently aggregated to capture volumetric context. To preserve slice order, a positional embedding 𝐞dC\mathbf{e}_{d}\in\mathbb{R}^{C} is added to each reference-conditioned slice token. A learnable fusion token 𝐮C\mathbf{u}\in\mathbb{R}^{C} is then appended to the resulting sequence, and a slice-level Transformer produces the side-specific volumetric representation:

𝐳a=Transformer([𝐬~1a+𝐞1,,𝐬~Da+𝐞D,𝐮])𝐮,\mathbf{z}^{a}=\operatorname{Transformer}\left([\widetilde{\mathbf{s}}_{1}^{a}+\mathbf{e}_{1},\ldots,\widetilde{\mathbf{s}}_{D}^{a}+\mathbf{e}_{D},\mathbf{u}]\right)_{\mathbf{u}}, (8)

where ()𝐮(\cdot)_{\mathbf{u}} denotes the output representation associated with the fusion token. The slice positional embeddings enable the Transformer to distinguish slice locations and model ordered inter-slice context. The same slice-level Transformer is applied to both breast sides with shared parameters.

Finally, a shared MLP classifier h()h(\cdot) maps the volumetric representation to side-specific logits a=h(𝐳a)3\bm{\ell}^{a}=h(\mathbf{z}^{a})\in\mathbb{R}^{3}, and the corresponding class probability distribution is 𝝅a=softmax(a)\bm{\pi}^{a}=\operatorname{softmax}(\bm{\ell}^{a}). The classifier parameters are shared between the left and right breasts. PRISM-Net is trained using the summed side-level cross-entropy loss:

=a𝒮CE(a,ya).\mathcal{L}=\sum_{a\in\mathcal{S}}\mathcal{L}_{\mathrm{CE}}\left(\bm{\ell}^{a},y^{a}\right). (9)

Thus, the left and right breast labels provide separate side-level supervision within a single bilateral forward pass.

4 Experiments and results

4.1 Dataset and cohort characteristics

Model development used the public ODELIA breast DCE-MRI dataset. Evaluation was performed using held-out ODELIA data and an independent institutional dataset from the General Hospital of the People’s Liberation Army. The selection workflow for the private institutional set is shown in Fig. 3.

Public ODELIA dataset. The public ODELIA dataset comprised 741741 examinations from 741741 women acquired at six European medical centers between December 20062006 and May 20242024 [30]. Each breast was independently labeled as no-lesion, benign, or malignant. 291291 examinations exhibited no lesions, 146146 examinations were diagnosed with benign but no malignant lesions, and 304304 examinations had malignant lesions. In total, the dataset includes 978978 breasts without lesions, 195195 with benign lesions, and 309309 with a malignant lesion. The original training (n = 408408) and validation (n = 103103) partitions were retained, while the test partition was split into an internal set (n = 130130) and a center-held-out set (n = 100100).

Private institutional dataset. An independent institutional cohort was retrospectively assembled to evaluate generalization and performance under complex background conditions. The primary candidate cohort was drawn from the 20252025 institutional archive. As examinations with moderate or marked BPE were underrepresented in this cohort, additional candidates meeting this BPE criterion were retrieved from the 20242024 and 20262026 archives. Examinations were reviewed by two breast radiologists and one clinician using DCE-MRI findings, BI-RADS assessments, radiology reports, and pathological results. Examinations were excluded for substantial treatment-related or postoperative changes, breast implants, or missing key DCE-MRI phases. After patient-level deduplication, 171171 examinations remained (135135 from 20252025 and 3636 BPE-enriched examinations from 20242024 and 20262026), yielding 342342 evaluable breast sides. Overlapping complexity subsets included 5555 examinations with moderate/marked BPE and 9292 with heterogeneous/dense FGT; 4949 met both criteria, as shown in Fig. 3. 7979 examinations exhibited no lesions, 3232 women were diagnosed with benign but no malignant lesions, and 6060 women had malignant lesions.

Breast-side reference standard The left and right breasts were treated as separate prediction targets and labeled as no-lesion, benign, or malignant. Labels were assigned using a pathology-prioritized hierarchy. Malignant or benign labels were assigned according to ipsilateral pathological findings. A no-lesion label was assigned when no ipsilateral enhancing lesion or pathological diagnosis was identified. For breasts with multiple findings, the most severe diagnosis determined the final label, with malignancy taking precedence over benignity and no-lesion. In total, the dataset includes 240240 breasts without lesions, 3838 with benign lesions, and 6464 with a malignant lesion. Cohort characteristics and acquisition protocols are summarized in Tables 1 and Supplementary Table S1.

Table 1: Demographic, breast background, and lesion characteristics of the final institutional breast DCE-MRI evaluation pool and its partially overlapping evaluation sets.
Characteristic Final pool 2025 set BPE set FGT set
Patients/examinations, nn 171 135 55 92
Evaluable breast sides, nn 342 270 110 184
Age, yearsa 47.5±13.047.5\pm 13.0 48.8±13.248.8\pm 13.2 42.0±11.042.0\pm 11.0 44.2±11.844.2\pm 11.8
BPE levela
Minimal/mild, nn (%) 116 (67.8) 116 (85.9) 0 (0.0) 43 (46.7)
Moderate/marked, nn (%) 55 (32.2) 19 (14.1) 55 (100.0) 49 (53.3)
FGT levela
Fatty/scattered, nn (%) 79 (46.2) 76 (56.3) 6 (10.9) 0 (0.0)
Heterogeneous/dense, nn (%) 92 (53.8) 59 (43.7) 49 (89.1) 92 (100.0)
Lesion lateralitya
No-lesion, nn (%) 79 (46.2) 66 (48.9) 23 (41.8) 22 (23.9)
Unilateral, nn (%) 82 (48.0) 62 (45.9) 27 (49.1) 64 (69.6)
Bilateral, nn (%) 10 (5.8) 7 (5.2) 5 (9.1) 6 (6.5)
Patient-level labela
No-lesion, nn (%) 79 (46.2) 66 (48.9) 23 (41.8) 22 (23.9)
Benign, nn (%) 32 (18.7) 23 (17.0) 13 (23.6) 28 (30.4)
Malignant, nn (%) 60 (35.1) 46 (34.1) 19 (34.5) 42 (45.7)
Breast-side labelb
No-lesion breast sides, nn (%) 240 (70.2) 194 (71.9) 73 (66.4) 108 (58.7)
Benign breast sides, nn (%) 38 (11.1) 28 (10.4) 17 (15.5) 33 (17.9)
Malignant breast sides, nn (%) 64 (18.7) 48 (17.8) 20 (18.2) 43 (23.4)
  • a

    Calculated at the patient/examination level. Age is presented as mean ±\pm standard deviation, and categorical variables as nn (%).

  • b

    Calculated at the individual breast-side level.

  • BPE, background parenchymal enhancement; FGT, fibroglandular tissue; DCE-MRI, dynamic contrast-enhanced magnetic resonance imaging.

Refer to caption
Figure 3: Flowchart of institutional patient selection and dataset construction. The candidate pool comprised 206206 patients, each contributing one bilateral DCE-MRI examination. It included a primary 20252025 institutional cohort of 166166 patients and an additional BPE-enriched cohort of 4040 patients from 20242024 and 20262026, and FGT-complexity membership was subsequently assigned within the final private evaluation pool according to FGT level. After the specified exclusions, the final private evaluation pool comprised 171171 patients and 342342 evaluable breast sides. Three partially overlapping evaluation sets were defined: the 20252025 private test set with 135135 patients and 270270 breast sides, the BPE-complexity set with 5555 patients and 110110 breast sides, and the FGT-complexity set with 9292 patients and 184184 breast sides. The dashed enclosure denotes the union of the two complexity sets, comprising 9898 patients. Among them, 6262 were also included in the 20252025 private test set, and 4949 satisfied both complexity criteria. BPE complexity was defined as moderate or marked BPE, whereas FGT complexity was defined as heterogeneous or dense FGT. DCE-MRI, dynamic contrast-enhanced magnetic resonance imaging; BPE, background parenchymal enhancement; FGT, fibroglandular tissue.

4.2 Experimental setup

4.2.1 Data preprocessing

A subtraction volume was generated from the pre-contrast and first post-contrast images, and each bilateral volume was split at the midline. The right breast was horizontally flipped to provide coarse bilateral correspondence without voxel-wise registration. Each breast was resized or cropped to a fixed shape, independently intensity-normalized, and masked to suppress non-breast regions. Paired breast volumes and masks were input jointly; no lesion-level annotations or manual registration were required.

4.2.2 Implementation details

PRISM-Net was implemented in PyTorch and optimized end-to-end using the summed cross-entropy losses of the left and right breasts. The shared DINOv3 ViT-S/16 backbone produced 384384-dimensional features. The slice fusion Transformer used eight attention heads, and the shared MLP classifier used a hidden dimension of 384384 with a dropout rate of 0.250.25. We used AdamW with a learning rate of 1.5×1051.5\times 10^{-5}, weight decay of 0.010.01, and a batch size of 88. The learning rate followed a cosine-decay schedule with 1010 warm-up epochs and a total schedule length of 120120 epochs. Training used mixed precision, with random flipping, rotation, noise perturbation, and Mixup (α=0.1\alpha=0.1, probability =0.5=0.5) applied only to the training set. Hyperparameters, checkpoints, and ordinal calibration thresholds were selected using the validation set, with early stopping after 3535 epochs without improvement. Final probabilities were obtained through validation-weighted checkpoint ensembling and test-time augmentation. All models were evaluated using identical preprocessing and evaluation protocols.

4.2.3 Baseline models

We compared PRISM-Net with representative breast MRI diagnostic models and bilateral imaging models, including BreastMRI-FCDD [31], PBPK DAE-CNN [10], DisAsymNet [40], BiGAM-Net [14], and the ODELIA Medical Slice Transformer baseline [28]. All models were adapted to the same breast-side three-class label space and evaluated on matched test cases.

4.2.4 Evaluation metrics

Performance was evaluated at the breast-side level using Macro AUC, Micro AUC, and QWK, which summarized class-balanced discrimination, overall discrimination, and ordinal agreement across no-lesion, benign, and malignant, respectively. All values were expressed on a percentage scale. Results are reported as the point estimate of the final checkpoint ensemble ± the standard deviation obtained from 2,0002,000 bootstrap resamples. PRISM-Net was compared with each baseline on matched cases using paired bootstrap testing with 10,00010,000 iterations and a two-sided significance threshold of p ¡ 0.050.05. In each bootstrap iteration, resampling was performed at the patient level, with the left and right breast-side predictions from the same patient kept together as one clustered unit. Cohort membership and reference labels were fixed across models, and no test label informed model selection, calibration, or threshold tuning.

4.3 Performance on the ODELIA public development dataset

Table 2: Performance comparison on the ODELIA public development dataset.
Method In-Distribution Out-of-Distribution
Macro AUC Micro AUC QWK Macro AUC Micro AUC QWK
BreastMRI-fcdd  [31] 59.90±6.1059.90\pm 6.10^{*} 70.20±9.7570.20\pm 9.75^{*} 10.77±9.7510.77\pm 9.75^{*} 54.11±4.0054.11\pm 4.00^{*} 73.01±5.4273.01\pm 5.42^{*} 2.61±2.832.61\pm 2.83^{*}
PBPK DAE-CNN  [10] 56.07±5.8456.07\pm 5.84^{*} 68.46±6.0768.46\pm 6.07^{*} 3.78±3.453.78\pm 3.45^{*} 50.36±3.6950.36\pm 3.69^{*} 65.80±2.3965.80\pm 2.39^{*} 2.24±4.012.24\pm 4.01^{*}
DisAsymNet  [40] 69.36±6.3469.36\pm 6.34^{*} 81.14±2.6481.14\pm 2.64^{*} 24.02±20.8524.02\pm 20.85^{*} 53.71±7.7153.71\pm 7.71 75.65±2.7475.65\pm 2.74 12.78±13.1712.78\pm 13.17^{*}
BiGAM-Net  [14] 62.32±5.8562.32\pm 5.85^{*} 72.43±6.8772.43\pm 6.87^{*} 10.86±12.1310.86\pm 12.13^{*} 50.00±3.0250.00\pm 3.02^{*} 70.87±4.7770.87\pm 4.77^{*} 2.96±3.652.96\pm 3.65^{*}
ODELIA  [28] 76.92±0.4876.92\pm 0.48^{*} 85.56±0.9185.56\pm 0.91^{*} 46.75±1.6246.75\pm 1.62^{*} 60.31±1.0760.31\pm 1.07 77.44±2.1777.44\pm 2.17 29.46±4.5829.46\pm 4.58^{*}
Ours 84.11±2.33\mathbf{84.11\pm 2.33} 90.64±1.61\mathbf{90.64\pm 1.61} 60.94±5.64\mathbf{60.94\pm 5.64} 68.51±4.54\mathbf{68.51\pm 4.54} 80.74±2.68\mathbf{80.74\pm 2.68} 43.45±7.10\mathbf{43.45\pm 7.10}
  • Values are reported as mean ±\pm SD. AUCs are reported as percentages, and QWK values are scaled by 100100. *p<0.05p<0.05 versus PRISM-Net.

Figure 4: ROC curves and confusion matrices of six models on the ODELIA in-distribution (ID) test set. The left panel shows the ROC curves, while the right panel presents the corresponding confusion matrices for three-class classification of no-lesion, benign, and malignant.
Refer to caption
(a) ROC curves.
Refer to caption
(b) Confusion matrices.
Figure 5: ROC curves and confusion matrices of six models on the RSH out-of-distribution (OOD) test set. The left panel shows the ROC curves, while the right panel presents the corresponding confusion matrices for three-class classification of no-lesion, benign, and malignant.
Refer to caption
(a) ROC curves.
Refer to caption
(b) Confusion matrices.

PRISM-Net achieved the best performance on both the ODELIA in-distribution and RSH out-of-distribution sets (Table 2). Its Macro AUC, Micro AUC, and QWK were 84.11±2.3384.11\pm 2.33, 90.64±1.6190.64\pm 1.61, and 60.94±5.6460.94\pm 5.64 in-distribution, and 68.51±4.5468.51\pm 4.54, 80.74±2.6880.74\pm 2.68, and 43.45±7.1043.45\pm 7.10 out-of-distribution. Relative to the strongest baseline, ODELIA, the corresponding gains were 7.197.19, 5.085.08, and 14.1914.19 points in-distribution and 8.208.20, 3.303.30, and 13.9913.99 points out-of-distribution. ROC curves and confusion matrices are shown in Figs. 4 and 5.

4.4 External validation on the private institutional test set

We compared PRISM-Net with DisAsymNet and ODELIA, which were the two strongest baseline methods on the ODELIA public development dataset. The independent 2025 institutional cohort was not used for training, validation, model selection, or tuning. PRISM-Net achieved the highest Macro AUC, Micro AUC, and QWK: 66.85±3.4766.85\pm 3.47, 86.38±1.8486.38\pm 1.84, and 38.85±6.4338.85\pm 6.43, respectively (Table 3). Its QWK exceeded ODELIA and DisAsymNet by 40.8640.86 and 26.4126.41 points, respectively. The corresponding ROC curves and confusion matrices are shown in panels (a) and (d) of Supplementary Fig. S1.

4.5 Performance on the background-complexity external test sets

The comparison was also restricted to DisAsymNet and ODELIA, the two best-performing baselines in the public-dataset evaluation. PRISM-Net also ranked first on both complexity cohorts (Table 3). On the BPE-complexity set, Macro AUC, Micro AUC, and QWK were 66.23±4.9066.23\pm 4.90, 82.66±2.9182.66\pm 2.91, and 31.48±9.3631.48\pm 9.36; on the FGT-complexity set, they were 67.71±3.3867.71\pm 3.38, 80.11±2.3780.11\pm 2.37, and 40.88±6.8340.88\pm 6.83. Corresponding ROC curves and confusion matrices are shown in panels (b,e) and (c,f) of Supplementary Fig. S1.

Table 3: External validation performance on the private institutional test set and background-complexity test sets.
Cohort Method Macro AUC Micro AUC QWK
2025 private external test set ODELIA  [28] 62.66±3.6362.66\pm 3.63 81.64±1.9581.64\pm 1.95^{*} 2.01±1.12-2.01\pm 1.12^{*}
DisAsymNet  [40] 63.14±4.2163.14\pm 4.21 82.62±2.1882.62\pm 2.18^{*} 12.44±5.4712.44\pm 5.47^{*}
PRISM-Net 66.85±3.47\mathbf{66.85\pm 3.47} 86.38±1.84\mathbf{86.38\pm 1.84} 38.85±6.43\mathbf{38.85\pm 6.43}
BPE-complexity test set ODELIA  [28] 64.87±5.6364.87\pm 5.63 78.55±3.2878.55\pm 3.28 1.91±2.191.91\pm 2.19^{*}
DisAsymNet  [40] 52.69±6.1552.69\pm 6.15 75.90±3.7575.90\pm 3.75^{*} 11.75±7.5511.75\pm 7.55^{*}
PRISM-Net 66.23±4.90\mathbf{66.23\pm 4.90} 82.66±2.91\mathbf{82.66\pm 2.91} 31.48±9.36\mathbf{31.48\pm 9.36}
FGT-complexity test set ODELIA  [28] 62.23±3.3062.23\pm 3.30 73.10±2.2673.10\pm 2.26^{*} 0.58±1.07-0.58\pm 1.07^{*}
DisAsymNet  [40] 64.64±4.1664.64\pm 4.16 75.53±2.4875.53\pm 2.48^{*} 15.21±5.5515.21\pm 5.55^{*}
PRISM-Net 67.71±3.38\mathbf{67.71\pm 3.38} 80.11±2.37\mathbf{80.11\pm 2.37} 40.88±6.83\mathbf{40.88\pm 6.83}
  • Values are mean ±\pm SD. AUCs are percentages; QWK values are scaled by 100100. *p<0.05p<0.05 PRISM-Net. BPE, background parenchymal enhancement; FGT, fibroglandular tissue.

4.6 Ablation studies

The full PRISM-Net was best in both settings (Table 4). Removing the bilateral asymmetry branch reduced QWK by 14.5614.56 points in-distribution and 13.4013.40 points out-of-distribution; removing focal score reweighting reduced QWK by 5.555.55 and 14.3814.38 points, respectively. Macro and Micro AUC also decreased for both variants, supporting the contribution of each component.

Table 4: Ablation study of the main components in PRISM-Net on the ODELIA public development dataset.
Variant In-Distribution Out-of-Distribution
Macro AUC Micro AUC QWK Macro AUC Micro AUC QWK
PRISM-Net w/o bilateral asymmetry branch 79.75±1.3779.75\pm 1.37^{*} 87.39±1.3487.39\pm 1.34^{*} 46.38±5.1746.38\pm 5.17^{*} 64.15±3.9264.15\pm 3.92^{*} 77.88±3.5377.88\pm 3.53^{*} 30.05±3.8830.05\pm 3.88^{*}
PRISM-Net w/o focal score reweighting 79.68±3.2279.68\pm 3.22^{*} 88.12±2.4588.12\pm 2.45^{*} 55.39±6.3955.39\pm 6.39 64.53±6.8364.53\pm 6.83^{*} 79.07±4.5879.07\pm 4.58 29.07±11.5329.07\pm 11.53^{*}
PRISM-Net 84.11±2.33\mathbf{84.11\pm 2.33} 90.64±1.61\mathbf{90.64\pm 1.61} 60.94±5.64\mathbf{60.94\pm 5.64} 68.51±4.54\mathbf{68.51\pm 4.54} 80.74±2.68\mathbf{80.74\pm 2.68} 43.45±7.10\mathbf{43.45\pm 7.10}
  • Values are mean ±\pm SD. AUCs are percentages; QWK values are scaled by 100100. *p<0.05p<0.05 versus PRISM-Net.

Second, we compared different feature extraction backbones while keeping the remaining bilateral learning framework unchanged. With the bilateral framework and evaluation protocol fixed, DINOv3 outperformed ResNet-18, ResNet-50, ViT, and DenseNet-121 on all six metrics (Table 5). DenseNet-121 was the second-best backbone. Relative to DenseNet-121, DINOv3 improved Macro AUC, Micro AUC, and QWK by 2.912.91, 1.481.48, and 6.746.74 points in-distribution and by 2.172.17, 0.770.77, and 3.293.29 points out-of-distribution.

Table 5: Ablation study of different feature extraction backbones on the ODELIA public development dataset.
Backbone In-Distribution Out-of-Distribution
Macro AUC Micro AUC QWK Macro AUC Micro AUC QWK
ResNet-18  [15] 80.40±1.2380.40\pm 1.23 88.65±1.3588.65\pm 1.35^{*} 51.29±3.4951.29\pm 3.49^{*} 65.49±4.3265.49\pm 4.32 80.34±2.8880.34\pm 2.88 35.84±5.8835.84\pm 5.88^{*}
ResNet-50  [15] 77.95±2.9077.95\pm 2.90^{*} 86.89±0.8186.89\pm 0.81^{*} 48.52±6.8648.52\pm 6.86^{*} 64.80±3.1764.80\pm 3.17 78.95±3.5478.95\pm 3.54 32.25±6.3532.25\pm 6.35^{*}
ViT  [8] 78.15±2.0678.15\pm 2.06^{*} 87.27±2.2387.27\pm 2.23^{*} 50.84±6.5150.84\pm 6.51 61.49±4.0361.49\pm 4.03^{*} 78.61±3.0178.61\pm 3.01 24.27±8.5824.27\pm 8.58^{*}
DenseNet-121  [20] 81.20±1.6781.20\pm 1.67 89.16±1.4689.16\pm 1.46 54.20±1.4854.20\pm 1.48 66.34±4.5166.34\pm 4.51 79.97±2.4779.97\pm 2.47 40.16±6.7940.16\pm 6.79
DINOv3  [36] 84.11±2.33\mathbf{84.11\pm 2.33} 90.64±1.61\mathbf{90.64\pm 1.61} 60.94±5.64\mathbf{60.94\pm 5.64} 68.51±4.54\mathbf{68.51\pm 4.54} 80.74±2.68\mathbf{80.74\pm 2.68} 43.45±7.10\mathbf{43.45\pm 7.10}
  • Values are mean ±\pm SD. AUCs are percentages; QWK values are scaled by 100100. *p<0.05p<0.05 versus PRISM-Net.

4.7 Interpretability and failure cases

Representative maps (Fig. 6) show that focal reweighting suppresses diffuse or background-related bilateral differences and concentrates responses on localized asymmetry. The no-lesion case showed diffuse activation, the benign case a mild localized response, and the malignant case a compact high-response region corresponding to the enhancing lesion. These maps provide feature-level explanations rather than lesion segmentations.

Refer to caption
Figure 6: Representative interpretability maps generated by PRISM-Net for no-lesion, benign, and malignant cases. Each row shows one representative case, and each column corresponds to a different stage of the bilateral asymmetry modeling process. From left to right, the columns show the target-side DCE-MRI slice, the relation feature before focal reweighting, the focal score map, the reweighted relation feature, the overlay of the reweighted relation feature on the target-side image, and the contralateral reference side. In the no-lesion case, the focal response is relatively diffuse and does not form a clear lesion-centered activation. In the benign case, the model produces mild and localized asymmetric responses. In the malignant case, the focal score and reweighted relation feature highlight a compact high-response region that spatially corresponds to the suspicious enhancing lesion. These visualizations indicate that PRISM-Net can suppress non-specific bilateral differences and emphasize diagnostically meaningful focal asymmetry.

We further analyzed representative failure cases to characterize the limitations and decision boundaries of PRISM-Net. As shown in Fig. 7, the errors mainly occurred in three challenging scenarios: benign lesions with malignant-like enhancement patterns, bilateral benign lesions that weakened the discriminative value of contralateral symmetry, and small low-suspicion lesions embedded within dense fibroglandular tissue. These cases suggest that the model may still be affected by overlapping enhancement appearances between benign and malignant lesions, confusing bilateral enhancement patterns, and limited lesion conspicuity in complex backgrounds.

Refer to caption
Figure 7: Failure-case interpretability and model decision boundaries. Each row shows one misclassified case. Columns show the target breast, bilateral relation feature, focal score, reweighted relation feature, heatmap, and contralateral breast. GT and PRED denote ground truth and prediction; 00, 11, and 22 represent no-lesion, benign, and malignant, respectively, with suffixes indicating breast side.Row 11 shows a benign BI-RADS 55 lesion predicted as malignant. The irregular, spiculated, heterogeneously enhancing lesion produced a strong focal response, while pathology revealed adenosis, sclerosing adenosis, ductal epithelial hyperplasia, and intraductal papilloma, illustrating overlap between benign proliferation and malignancy.Row 22 shows a benign case in a 3939-year-old woman that was predicted as no-lesion. The patient had multiple bilateral fibroadenomas, dense FGT, and minimal BPE. Enhancing masses were present in both breasts, and corresponding enhancement was also observed in the contralateral side, which weakened the discriminative value of bilateral asymmetry for lesion detection. The Row 33 shows an approximately 88-mm fibroadenoma predicted as no-lesion; its small size, circumscribed margin, and dense FGT corresponded to weak focal and reweighted relation responses. These cases highlight three decision boundaries: malignant-like benign proliferation, bilateral lesions with limited asymmetry, and small lesions obscured by complex background signal.

5 Discussion

This study developed PRISM-Net as a registration-free bilateral learning framework for three-class breast DCE-MRI classification. The main finding is that explicitly modeling the contralateral breast as a patient-specific bilateral reference representation improved both discriminative performance and ordinal diagnostic agreement across public, external, and background-complexity test settings. On the ODELIA in-distribution test set, PRISM-Net achieved a Macro AUC of 84.11±2.3384.11\pm 2.33, Micro AUC of 90.64±1.6190.64\pm 1.61, and QWK of 60.94±5.6460.94\pm 5.64. This advantage was preserved on the RSH out-of-distribution institution. These results suggest that PRISM-Net did not merely improve class separability, but also better preserved the ordered relationship among no-lesion, benign, and malignant.

The observed gains may reflect the way PRISM-Net uses bilateral information for patient-specific background normalization rather than simply treating the contralateral breast as an additional input. Radiological bilateral comparison benefits from shared patient-level physiological and acquisition conditions, but the contralateral breast cannot be assumed to be anatomically identical or disease-free, avoiding the assumption of anatomical symmetry or disease-free status. PRISM-Net therefore treats it as an approximate comparative context rather than a normal template. Shared-weight encoding maps both breasts into a comparable feature space, body-mask-guided patch filtering restricts matching to anatomically plausible breast tissue, and adaptive patch-level matching reduces dependence on voxel-wise registration despite differences in breast shape, positioning, and deformation. The interpretability analysis in Fig. 6 provides qualitative support for this mechanism. Before focal reweighting, the bilateral relation features contained relatively diffuse responses that could reflect both lesion-related asymmetry and nonspecific background variation. After reweighting, the responses became more spatially concentrated: the malignant example showed a compact high-response region spatially overlapping with the suspicious enhancing region, the no-lesion example lacked a distinct lesion-centered activation, and the benign example exhibited an intermediate pattern. These observations suggest that the bilateral focal response module may help reduce sensitivity to diffuse or bilaterally shared variations while emphasizing localized asymmetric features that are more relevant to classification, which is consistent with the model design. Nevertheless, these visualizations should be interpreted as qualitative feature-level explanations.

The background-complexity analyses further indicate that bilateral comparison may be particularly useful when unilateral enhancement patterns are confounded by complex parenchymal backgrounds. Moderate or marked BPE has previously been associated with reduced diagnostic performance on breast MRI [6]. In the BPE- and FGT-complexity test sets, PRISM-Net consistently achieved the highest performance among evaluated methods, with the most pronounced relative gains observed for QWK, as shown in Table 3. The confusion matrices in Supplementary Fig. S1 further illustrate fewer severe ordinal misclassifications. As shown in Fig. 6 benign case, dense FGT confounded radiological interpretation in a representative clinical case, leading to a BI-RADS 5 assessment. PRISM-Net, however, correctly classified the affected breast as benign, consistent with postoperative histopathology showing sclerosing adenosis with intraductal papilloma. These findings suggest that patient-specific bilateral reference modeling may provide valuable comparative information under complex background conditions. Physiological and hormonal factors, including menstrual cycle variation, pregnancy, and lactation, can contribute to increased BPE and heterogeneous background enhancement, potentially reducing lesion conspicuity and increasing diagnostic uncertainty. By incorporating contralateral breast information as an individualized comparative reference, PRISM-Net better capture clinically relevant asymmetric patterns beyond absolute enhancement characteristics in such challenging examinations. This capability may be relevant for lesions presenting as NME, where the absence of a discrete mass and substantial overlap between malignant, benign, and physiological enhancement patterns often complicate interpretation. Nevertheless, performance variation across these challenging sets may be influenced by limited sample sizes and the inherent uncertainty associated with complex background enhancement patterns. Further study in larger cohorts is required.

The ablation studies clarify not only the contribution of individual modules but also the design logic underlying PRISM-Net. As shown in Table 4, removing the bilateral asymmetry branch consistently degrades performance, indicating that the contralateral breast contributes more than additional image context. Instead, it serves as a patient-specific reference that enables the model to distinguish shared physiological enhancement patterns from localized abnormal deviations. This result supports adaptive reference matching in the feature space as a more appropriate formulation of bilateral comparison than unilateral analysis or simple bilateral feature aggregation. The contribution of focal score reweighting further suggests that bilateral differences are not uniformly informative. Diffuse enhancement variation and background related mismatch can generate nonspecific responses, whereas deviations that are locally salient are more likely to contain discriminative evidence. Focal reweighting therefore acts as an evidence selection mechanism that suppresses diffuse interference and concentrates the relational representation on diagnostically relevant asymmetry. The backbone comparison provides a complementary observation that reliable registration free correspondence depends on the semantic quality of the shared feature space, since more discriminative representations facilitate stable cross breast matching under anatomical, positional, and background variability. Taken together, these findings reveal a coherent computational pathway in PRISM-Net: robust feature encoding establishes a comparable bilateral representation space, adaptive matching constructs patient specific references, and focal reweighting identifies the most informative reference conditioned deviations. The performance gains therefore arise not merely from introducing bilateral input, but from structuring bilateral comparison into representation learning, correspondence estimation, and selective evidence extraction.

Several limitations should be acknowledged. First, although the study used a public development dataset, an out-of-distribution validation institution, and an independent private external test set, the private external data were collected from a single institution. Prospective multicenter validation is therefore needed to assess generalizability across centers, imaging devices, and acquisition protocols. Second, the registration-free bilateral design should be interpreted as a feature-space approximation of bilateral comparison rather than as explicit anatomical correspondence. Its effectiveness may be reduced when the contralateral breast provides an unreliable reference, which may be compromised by bilateral disease, marked asymmetric background parenchymal enhancement, implants, postoperative changes, motion artifacts, failed fat suppression, or asymmetric image quality. The observed failure cases were consistent with this boundary (Fig. 7). Third,the model was trained at the breast-side level without requiring lesion-level annotations. Although this weakly supervised setting is practical for large-scale clinical data, it limits direct evaluation of lesion localization. It also does not fully exploit complementary MRI information, including T2-weighted imaging, DWI, ADC maps, ultrafast kinetic features, and clinical risk factors.

Future work should extend the present retrospective evaluation through prospective multicenter studies, broader protocol heterogeneity, and reader studies in realistic clinical workflows. Further studies should also evaluate whether incorporating multiparametric MRI sequences, kinetic information, and clinical risk factors can improve diagnostic performance beyond early DCE-MRI alone. In addition, further comparison between asymmetry maps, model attention, and radiologist-annotated lesion locations may help clarify the spatial interpretability. Overall, the present findings support registration-free bilateral comparison as a potential strategy for improving AI-assisted breast DCE-MRI triage. Larger prospective multicenter studies are required to determine its clinical utility.

6 Conclusion

PRISM-Net presents a registration-free bilateral learning framework for breast-side three-class DCE-MRI classification that treats the contralateral breast as an approximate patient-specific reference rather than a normal template. Through adaptive patch-level matching and focal asymmetry reweighting, the framework leverages bilateral contextual information to characterize diagnostically relevant asymmetry under complex background variation. Across public, external, and background-complexity evaluations, PRISM-Net achieved superior performance compared with existing baselines, while ablation studies demonstrated the complementary contributions of bilateral reference modeling and focal asymmetry reweighting. These findings suggest that patient-specific bilateral comparison may provide a promising strategy for improving AI-assisted interpretation of challenging breast MRI examinations. Further prospective multicenter validation and reader studies are warranted to assess its potential clinical utility.

Acknowledgements

CRediT authorship contribution statement

Boya Zhang: Conceptualization, Methodology, Investigation, Data curation, Validation, Formal analysis, Funding acquisition, Writing – original draft. Shuaiwen Zhou: Conceptualization, Methodology, Software, Investigation, Formal analysis, Validation, Visualization, Data curation, Writing – original draft. Di Kong: Review & editing, Methodology, Supervision, Conceptualization. Mingxu Wang: Data curation, Review & editing. Wenbiao Du: Data curation, Review & editing. Yiman Zhong: Data curation, Review & editing. Yuexin Duan: Data curation, Review & editing. Xiawei Yue: Data curation, Review & editing. Liuquan Cheng: Supervision, Resources, Project administration, Conceptualization, Review & editing. Xiru Li: Supervision, Funding acquisition, Resources, Project administration, Review & editing.

Ethics statement

This study involved human subjects and was conducted in accordance with ethical standards. All procedures and protocols were approved by the Ethics Committee of the General Hospital of the People’s Liberation Army (Approval No. S2026-078-01).

Funding

This work is supported by the Zhongguancun Academy (Grant No. 02012411).

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data and code availability

The public dataset used for model training and evaluation is accessible through the official ODELIA data.The source code will be made publicly available upon acceptance of this article.

References

  • [1] K. A. Abdullah, S. Marziali, M. Nanaa, L. Escudero Sánchez, N. R. Payne, and F. J. Gilbert (2025) Deep learning-based breast cancer diagnosis in breast MRI: systematic review and meta-analysis. European Radiology 35 (8), pp. 4474–4489. External Links: Document Cited by: §1.
  • [2] R. Alterson and D. B. Plewes (2003) Bilateral symmetry analysis of breast MRI. Physics in Medicine and Biology 48 (20), pp. 3431–3443. External Links: Document Cited by: §2.2.
  • [3] American College of Radiology (2013) ACR BI-RADS Atlas: breast imaging reporting and data system. 5 edition, American College of Radiology, Reston, VA. External Links: ISBN 978-1-55903-016-8 Cited by: §1.
  • [4] N. Antropova, B. Huynh, H. Li, and M. L. Giger (2019) Breast lesion classification based on dynamic contrast-enhanced magnetic resonance image sequences with long short-term memory networks. Journal of Medical Imaging 6 (1), pp. 011002. External Links: Document Cited by: §1, §2.1.
  • [5] M. Bahl (2025) Does background parenchymal enhancement impact the diagnostic accuracy of breast MRI? revisiting the controversy. Radiology 315 (2), pp. e251328. External Links: Document Cited by: §1.
  • [6] S. Bechyna and P. A. T. Baltzer (2025) Impact of background parenchymal enhancement on diagnostic performance of breast MRI: a systematic review and meta-analysis. Radiology 315 (2), pp. e241919. External Links: Document Cited by: §1, §1, §5.
  • [7] F. Bray, M. Laversanne, H. Sung, J. Ferlay, R. L. Siegel, I. Soerjomataram, and A. Jemal (2024) Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: A Cancer Journal for Clinicians 74 (3), pp. 229–263. External Links: Document Cited by: §1.
  • [8] A. Dosovitskiy et al. (2021) An image is worth 16x16 words: transformers for image recognition at scale. In International Conference on Learning Representations, Cited by: Table 5.
  • [9] C. S. Giess, E. D. Yeh, S. Raza, and R. L. Birdwell (2014) Background parenchymal enhancement at breast MR imaging: normal patterns, diagnostic challenges, and potential for false-positive and false-negative interpretation. RadioGraphics 34 (1), pp. 234–247. External Links: Document Cited by: §1.
  • [10] M. Gravina, M. Maddaluno, S. Marrone, M. Sansone, R. Fusco, V. Granata, A. Petrillo, and C. Sansone (2024) A physiological-informed generative model for improving breast lesion classification in small DCE-MRI datasets. IEEE Journal of Biomedical and Health Informatics 28 (11), pp. 6764–6777. External Links: Document Cited by: §1, §2.1, §4.2.3, Table 2.
  • [11] M. Gravina, S. Marrone, M. Sansone, and C. Sansone (2021) DAE-CNN: exploiting and disentangling contrast agent effects for breast lesions classification in DCE-MRI. Pattern Recognition Letters 145, pp. 67–73. External Links: Document Cited by: §1, §2.1.
  • [12] Y. Guan, X. Wang, H. Li, Z. Zhang, X. Chen, O. Siddiqui, S. Nehring, and X. Huang (2020) Detecting asymmetric patterns and localizing cancers on mammograms. Patterns 1 (7), pp. 100106. External Links: Document Cited by: §2.2.
  • [13] W. Hao, J. Gong, S. Wang, H. Zhu, B. Zhao, and W. Peng (2020) Application of mri radiomics-based machine learning model to improve contralateral bi-rads 4 lesion assessment. Frontiers in oncology 10, pp. 531476. External Links: Document Cited by: §1.
  • [14] A. K. A. Haque, X. Li, Y. Li, and X. Yuan (2026) BiGAM-Net: a bilateral gated fusion framework for imbalanced breast imaging. Expert Systems with Applications 328, pp. 132746. External Links: Document Cited by: §2.2, §4.2.3, Table 2.
  • [15] K. He, X. Zhang, S. Ren, and J. Sun (2016) Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778. Cited by: Table 5, Table 5.
  • [16] S. Hennessey, E. Huszti, A. Gunasekura, A. Salleh, L. Martin, S. Minkin, S. Chavez, and N. F. Boyd (2014) Bilateral symmetry of breast tissue composition by magnetic resonance in young women and adults. Cancer Causes & Control 25 (4), pp. 491–497. External Links: Document Cited by: §2.2.
  • [17] L. Hirsch, E. J. Sutton, Y. Huang, B. Kayis, M. Hughes, D. Martinez, H. A. Makse, and L. C. Parra (2025) High-performance open-source AI for breast cancer detection and localization in MRI. Radiology: Artificial Intelligence 7 (5), pp. e240550. External Links: Document Cited by: §1, §2.1.
  • [18] A. Hizukuri, R. Nakayama, M. Nara, M. Suzuki, and K. Namba (2021) Computer-aided diagnosis scheme for distinguishing between benign and malignant masses on breast DCE-MRI images using deep convolutional neural network with bayesian optimization. Journal of Digital Imaging 34 (1), pp. 116–123. External Links: Document Cited by: §1, §2.1.
  • [19] Q. Hu, H. M. Whitney, H. Li, Y. Ji, P. Liu, and M. L. Giger (2021) Improved classification of benign and malignant breast lesions using deep feature maximum intensity projection MRI in breast cancer diagnosis using dynamic contrast-enhanced MRI. Radiology: Artificial Intelligence 3 (3), pp. e200159. External Links: Document Cited by: §2.1.
  • [20] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger (2017) Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4700–4708. Cited by: Table 5.
  • [21] S. A. Jansen, V. C. Lin, M. L. Giger, H. Li, G. S. Karczmar, and G. M. Newstead (2011) Normal parenchymal enhancement patterns in women undergoing mr screening of the breast. European Radiology 21 (7), pp. 1374–1382. External Links: Document Cited by: §2.2.
  • [22] C. K. Kuhl (2024) Abbreviated breast MRI: state of the art. Radiology 310 (3), pp. e221822. External Links: Document Cited by: §1.
  • [23] D. Leithner, J. V. Horvat, B. Bernard-Davila, T. H. Helbich, R. E. Ochoa-Albiztegui, D. F. Martinez, M. Zhang, S. B. Thakur, G. J. Wengert, A. Staudenherz, M. S. Jochelson, E. A. Morris, P. A. T. Baltzer, P. Clauser, P. Kapetas, and K. Pinker (2019) A multiparametric [18F]FDG PET/MRI diagnostic model including imaging biomarkers of the tumor and contralateral healthy breast tissue aids breast cancer diagnosis. European Journal of Nuclear Medicine and Molecular Imaging 46 (9), pp. 1878–1888. External Links: Document Cited by: §1.
  • [24] G. J. Liao, L. C. Henze Bancroft, R. M. Strigel, R. D. Chitalia, D. Kontos, L. Moy, S. C. Partridge, and H. Rahbar (2020) Background parenchymal enhancement on breast MRI: a comprehensive review. Journal of Magnetic Resonance Imaging 51 (1), pp. 43–61. External Links: Document Cited by: §1, §2.2.
  • [25] Y. Liu, Z. Zhou, S. Zhang, L. Luo, Q. Zhang, F. Zhang, X. Li, Y. Wang, and Y. Yu (2019) From unilateral to bilateral learning: detecting mammogram masses with contrasted bilateral network. In Medical Image Computing and Computer Assisted Intervention - MICCAI 2019, Lecture Notes in Computer Science, Vol. 11769, pp. 477–485. External Links: Document Cited by: §2.2.
  • [26] R. Lo Gullo, R. E. Ochoa-Albiztegui, J. Chakraborty, S. B. Thakur, M. Robson, M. S. Jochelson, K. Varela, D. Resch, S. Eskreis-Winkler, and K. Pinker (2024) Development of an MRI radiomic machine-learning model to predict triple-negative breast cancer based on fibroglandular tissue of the contralateral unaffected breast in breast cancer patients. Cancers 16 (20), pp. 3480. External Links: Document Cited by: §1.
  • [27] M. Marcon, M. H. Fuchsjäger, P. Clauser, and R. M. Mann (2024) ESR essentials: screening for breast cancer - general recommendations by EUSOBI. European Radiology 34 (10), pp. 6348–6357. External Links: Document Cited by: §1.
  • [28] G. Müller-Franzes, L. Escudero Sánchez, N. Payne, A. Athanasiou, M. Kalogeropoulos, A. Lopez, A. M. Soro Busto, J. Camps Herrero, N. Rasoolzadeh, T. Zhang, R. Mann, D. Jutz, M. Bode, C. Kuhl, W. Veldhuis, O. L. Saldanha, J. Zhu, J. N. Kather, D. Truhn, and F. J. Gilbert (2025) A european multi-center breast cancer MRI dataset. Note: arXiv preprint External Links: 2506.00474, Document, Link Cited by: §1, §2.1, §4.2.3, Table 2, Table 3, Table 3, Table 3.
  • [29] G. Müller-Franzes, F. Khader, R. Siepmann, T. Han, J. N. Kather, S. Nebelung, and D. Truhn (2025) Medical slice transformer for improved diagnosis and explainability on 3d medical images with DINOv2. Scientific Reports 15 (1), pp. 23979. External Links: Document Cited by: §2.1.
  • [30] ODELIA-AI (2025) ODELIA Challenge 2025. Note: Hugging Face DatasetsAccessed April 17, 2026 External Links: Link Cited by: §1, §1, §2.1, §4.1.
  • [31] F. Oviedo, A. S. Kazerouni, P. Liznerski, Y. Xu, M. Hirano, R. A. Vandermeulen, M. Kloft, E. Blum, A. M. Alessio, C. I. Li, W. B. Weeks, R. Dodhia, J. M. Lavista Ferres, H. Rahbar, and S. C. Partridge (2025) Cancer detection in breast MRI screening via explainable AI anomaly detection. Radiology 316 (1), pp. e241629. External Links: Document Cited by: §1, §2.2, §4.2.3, Table 2.
  • [32] G. Pelluet, M. Rizkallah, M. Tardy, and D. Mateus (2025) Enhancing breast cancer screening: unveiling explainable cross-view contributions in dual-view mammography with sparse bipartite graphs attention networks. Computerized Medical Imaging and Graphics 125, pp. 102620. External Links: Document Cited by: §1.
  • [33] K. M. Ray, K. Kerlikowske, I. V. Lobach, M. B. Hofmann, H. I. Greenwood, V. A. Arasu, N. M. Hylton, and B. N. Joe (2018) Effect of background parenchymal enhancement on breast MR imaging interpretive performance in community-based practices. Radiology 286 (3), pp. 822–829. External Links: Document Cited by: §1.
  • [34] D. Shimokawa, K. Takahashi, D. Kurosawa, E. Takaya, et al. (2023) Deep learning model for breast cancer diagnosis based on bilateral asymmetrical detection (BilAD) in digital breast tomosynthesis images. Radiological Physics and Technology 16 (1), pp. 20–27. External Links: Document Cited by: §1, §2.2.
  • [35] R. L. Siegel, A. N. Giaquinto, and A. Jemal (2024) Cancer statistics, 2024. CA: A Cancer Journal for Clinicians 74 (1), pp. 12–49. External Links: Document Cited by: §1.
  • [36] O. Siméoni et al. (2025) DINOv3. External Links: 2508.10104, Link Cited by: §3.2, Table 5.
  • [37] L. Sun, Y. Zhang, T. Liu, H. Ge, J. Tian, X. Qi, J. Sun, and Y. Zhao (2023) A collaborative multi-task learning method for BI-RADS category 4 breast lesion segmentation and classification of MRI images. Computer Methods and Programs in Biomedicine 240, pp. 107705. External Links: Document Cited by: §1, §2.1.
  • [38] D. Truhn, S. Schrading, C. Haarburger, H. Schneider, D. Merhof, and C. Kuhl (2019) Radiomic versus convolutional neural networks analysis for classification of contrast-enhancing lesions at multiparametric breast MRI. Radiology 290 (2), pp. 290–297. External Links: Document Cited by: §1, §2.1.
  • [39] E. Verburg, C. H. van Gils, B. H. M. van der Velden, M. F. Bakker, R. M. Pijnappel, W. B. Veldhuis, and K. G. A. Gilhuijs (2022) Deep learning for automated triaging of 4581 breast MRI examinations from the DENSE trial. Radiology 302 (1), pp. 29–36. External Links: Document Cited by: §1.
  • [40] X. Wang, T. Tan, Y. Gao, L. Han, T. Zhang, C. Lu, R. Beets-Tan, R. Su, and R. Mann (2023) DisAsymNet: disentanglement of asymmetrical abnormality on bilateral mammograms using self-adversarial learning. In Medical Image Computing and Computer Assisted Intervention - MICCAI 2023, Cham, pp. 57–67. External Links: Document Cited by: §1, §2.2, §4.2.3, Table 2, Table 3, Table 3, Table 3.
  • [41] Q. Yang, L. Li, J. Zhang, G. Shao, C. Zhang, and B. Zheng (2014) Computer-aided diagnosis of breast DCE-MRI images using bilateral asymmetry of contrast enhancement between two breasts. Journal of Digital Imaging 27 (1), pp. 152–160. External Links: Document Cited by: §1, §2.2.
  • [42] T. Zeng, Y. Zeng, Z. Zhang, X. Zhang, K. Ichiji, S. Chou, I. Bukovsky, J. Vrba, and N. Homma (2026) Bilateral information-guided diagnosis of breast masses in mammography using vision transformer. IEEE Journal of Biomedical and Health Informatics 30 (5), pp. 4226–4237. External Links: Document Cited by: §1, §2.2, §2.2.
  • [43] Y. Zhang, Y. Liu, K. Nie, J. Zhou, Z. Chen, J. Chen, X. Wang, B. Kim, R. Parajuli, R. S. Mehta, M. Wang, and M. Su (2023) Deep learning-based automatic diagnosis of breast cancer on MRI using mask R-CNN for detection followed by ResNet50 for classification. Academic Radiology 30 (Suppl 2), pp. S161–S171. External Links: Document Cited by: §1, §2.2.
  • [44] Y. Zheng, B. Wei, H. Liu, R. Xiao, and J. C. Gee (2015) Measuring sparse temporal-variation for accurate registration of dynamic contrast-enhanced breast MR images. Computerized Medical Imaging and Graphics 46, pp. 73–80. External Links: Document Cited by: §1.

Supplementary Material

Table S1: DCE-MRI acquisition protocols of the public and private cohorts.
Parameter Public multicenter cohort Private institutional cohort
Dataset ODELIA public breast MRI dataset Private institutional breast DCE-MRI cohort
Scanner vendor GE, Siemens, and Philips GE Discovery MR750
Field strength 1.5 T and 3.0 T 3.0 T
Breast coil Dedicated bilateral breast coils, 4–18 channels depending on center 8-channel dedicated breast coil
Sequence used for model input Dynamic T1-weighted sequence VIBRANT dynamic T1-weighted imaging
Imaging plane Axial Axial
TR/TE, ms 4.4–7.1 / 1.7–4.6 for most 3D protocols 7.7 / 4.3
Flip angle 9.8–18 for most 3D protocols 10
Field of view, mm 201–427 320
Acquisition matrix 256–672 320 ×\times 320
Slice thickness, mm 1.0–3.1 1.0
Post-contrast phases 2–7 4
Contrast dose 0.1 mmol/kg 0.1 mmol/kg
Injection rate 1–3 mL/s 2 mL/s
Saline flush 25–30 mL 20 mL
  • Public-cohort values are summarized as ranges across centers from the original dataset supplementary material. DCE-MRI, dynamic contrast-enhanced magnetic resonance imaging; TE, echo time; TR, repetition time.

Refer to caption
(a) 2025 private external test set.
Refer to caption
(b) BPE-complexity test set.
Refer to caption
(c) FGT-complexity test set.
Refer to caption
(d) Confusion matrices on the 2025 private external test set.
Refer to caption
(e) Confusion matrices on the BPE-complexity test set.
Refer to caption
(f) Confusion matrices on the FGT-complexity test set.
Figure S1: ROC curves and confusion matrices of PRISM-Net, DisAsymNet, and ODELIA on the private institutional cohorts. Panels (a)–(c) show the ROC curves on the 2025 private external test set, the BPE-complexity test set, and the FGT-complexity test set, respectively. Panels (d)–(f) present the corresponding confusion matrices for three-class classification.