arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2410.19907v5 [astro-ph.GA] 26 Feb 2025

Weak-lensing Mass Reconstruction of Galaxy Clusters with a Convolutional Neural Network - II: Application to Next-Generation Wide-Field Surveys

Astropy [3, 4, 5], DisPerSE [53, 54], Matplotlib [29], NumPy [23], SExtractor [9], SciPy [60], Tensorflow [1], TreeCorr [34, 33]
Sangjun Cha Affiliation: Department of Astronomy, Yonsei University, 50 Yonsei-ro, Seoul 03722, Korea    M. James Jee Affiliation: Department of Astronomy, Yonsei University, 50 Yonsei-ro, Seoul 03722, Korea Affiliation: Department of Physics and Astronomy, University of California, Davis, One Shields Avenue, Davis, CA 95616, USA Corresponding author: M. James Jee    Sungwook E. Hong Affiliation: Korea Astronomy and Space Science Institute, 776 Daedeok-daero, Yuseong-gu, Daejeon 34055, Korea Affiliation: Astronomy Campus, University of Science and Technology, 776 Daedeok-daero, Yuseong-gu, Daejeon 34055, Korea    Sangnam Park Affiliation: Physics Department & Natural Science Research Institute, University of Seoul, 163 Seoulsiripdaero, Dongdaemun-gu, Seoul 02504, Korea    Dongsu Bak Affiliation: Physics Department & Natural Science Research Institute, University of Seoul, 163 Seoulsiripdaero, Dongdaemun-gu, Seoul 02504, Korea    Taehwan Kim Affiliation: Artificial Intelligence Graduate School, UNIST, Ulsan, Republic of Korea
Abstract

Traditional weak-lensing mass reconstruction techniques suffer from various artifacts, including noise amplification and the mass-sheet degeneracy. In Hong et al. (2021), we demonstrated that many of these pitfalls of traditional mass reconstruction can be mitigated using a deep learning approach based on a convolutional neural network (CNN). In this paper, we present our improvements and report on the detailed performance of our CNN algorithm applied to next-generation wide-field observations. Assuming the field of view (3.5×3.53\fdg 5\times 3\fdg 5) and depth (27 mag at 5σ5\sigma) of the Vera C. Rubin Observatory, we generated training datasets of mock shear catalogs with a source density of 33 arcmin-2 from cosmological simulation ray-tracing data. We find that the current CNN method provides high-fidelity reconstructions consistent with the true convergence field, restoring both small and large-scale structures. In addition, the cluster detection utilizing our CNN reconstruction achieves 75\sim 75% completeness down to 1014M\sim 10^{14}M_{\sun}. We anticipate that this CNN-based mass reconstruction will be a powerful tool in the Rubin era, enabling fast and robust wide-field mass reconstructions on a routine basis.

I Introduction

Weak lensing (WL) is a powerful method for reconstructing the mass distributions of galaxy clusters without any dynamical assumptions [37, 14, 51, 26, 35, 46, 58, 20, 39, 2, e.g.,]. Despite this unique merit, WL mass reconstruction faces several critical limitations, including the mass-sheet degeneracy, boundary effect, noise amplification, resolution loss, and finite-field effect [19, 6, 50, 12, e.g.,].

The mass-sheet degeneracy arises because the shear field is invariant under a certain transformation of the mass density. The finite-field/boundary effect occurs because the shear field near the edges of the observed region is influenced by the mass outside the reconstruction field, for which we have no constraining data. The noise amplification is caused because the mass reconstruction is an ill-posed inversion problem where the relation between the observed shear and underlying mass is nonlinear. The resolution loss is inevitable because the number density of the sources is finite and thus the shear field needs to be smoothed. In general, these issues happen together, and their combined effects are present in mass reconstructions, leading to numerous artifacts and difficulties in interpretation.

One of the notable efforts to mitigate some of the aforementioned artifacts is mass reconstruction through the maximum likelihood approach. In this approach, a grid representing either mass or potential is set up on the reconstruction field, and the parameters defining the grid are iteratively determined by maximizing the likelihood relating mass and shear [7, 55, 22, 38]. However, maximizing the likelihood alone leads to overfitting, and thus it is necessary to impose a prior to regularize the result. Some methods utilizing the entropy of the mass grid as a regularization channel have been demonstrated to be effective in overcoming some pitfalls of traditional mass reconstructions [14, 51, 13, 17, 16, e.g.,]. One of the major drawbacks of this approach is its large computational cost. Even a moderately sized grid of 100×100 requires 10,000 free parameters, making mass reconstruction computationally expensive and slow when using a traditional CPU-based approach. Hence, it is difficult to apply the method for routine use to deep wide-field (WF) WL data from next-generation facilities such as the Vera C. Rubin Observatory, the Euclid Mission, and the Nancy Grace Roman telescope.

Deep learning is poised to revolutionize many fields of astronomy. In particular, its unparalleled ability to solve ill-posed nonlinear inversion problems holds tremendous potential in mitigating the limitations of traditional approaches in image denoising, restoration, deconvolution, super-resolution, and more [52, 42, 57, 49, e.g.,]. Certainly, WL mass reconstruction is a promising field that stands to gain from the rapid advancements in deep learning.

Recently, we proposed a convolutional neural network (CNN) architecture for WL mass reconstruction [27, hereafter H21]. The CNN-based mass reconstruction was shown to significantly outperform traditional mass reconstructions in restoring mass density distributions, recovering projected halo masses, suppressing spurious substructures, and predicting mass centroids.

In this paper, we present our improved CNN architecture and report on the detailed performance when the method is applied to next-generation WF observations. The improvement stems from two factors. First, we enhanced our CNN algorithm to be more compact and efficient, enabling it to preserve the resolution of mass maps for capturing substructures in WF observations. Second, we updated the loss function to ensure that the resulting mass reconstruction captures not only high-contrast individual clusters but also diffuse large-scale structures within the reconstruction field. Among several next-generation WL surveys, we target Vera C. Rubin’s Legacy Survey of Space and Time (LSST) and generate training data for its field of view (FOV) (3.5×3.53\fdg 5\times 3\fdg 5) with the 10-year nominal depth of 27 mag [32].

The paper is organized as follows. In §II, we present our CNN architecture and the methods for generating mock WL data. §III shows the reconstruction results. In §IV, we discuss the fidelity of the results with a few metrics. We summarize in §V. Unless stated otherwise, we assume a flat Λ\LambdaCDM cosmology with the dimensionless Hubble parameter h=0.7h=0.7 and the matter density parameter ΩM=0.3\Omega_{M}=0.3 throughout the paper.

II Method

II.1 Basic Weak Lensing Theory

In this section, we provide a brief overview of the WL theory. For more details on gravitational lensing, readers can refer to review papers [8, 25, e.g.,].

The distortion of source galaxies in the WL regime can be expressed as 𝐱=𝐀𝐱\mathbf{x^{\prime}}=\mathbf{Ax}, where 𝐱\mathbf{x} (𝐱\mathbf{x^{\prime}}) is the position in the source (image) plane. The matrix 𝐀\mathbf{A} is expressed as:

𝐀=(1κ)(1g1g2g21+g1).\mathbf{A}=(1-\kappa)\begin{pmatrix}1-g_{1}&-g_{2}\\ -g_{2}&1+g_{1}\end{pmatrix}. (1)

In Equation 1, κ\kappa denotes the convergence, and g1(2)g_{1(2)} is the first (second) component of the reduced shear g=(g12+g22)1/2g=({g_{1}}^{2}+{g_{2}}^{2})^{1/2}. The reduced shear gg is computed from the relation g=γ/(1κ)g=\gamma/(1-\kappa), where γ\gamma is the shear. The convergence κ\kappa is defined as:

κ=ΣΣc,\kappa=\frac{\Sigma}{\Sigma_{c}}, (2)

where Σ\Sigma is the surface mass density, and the critical surface mass density Σc\Sigma_{c} is given by:

Σc=c2Ds4πGDdDds.\Sigma_{c}=\frac{c^{2}D_{s}}{4\pi GD_{d}D_{ds}}. (3)

Here, cc is the speed of light, Dd(s)D_{d(s)} is the angular diameter distance to the lens (source), and DdsD_{ds} is the angular diameter distance between the lens and the source. γ\gamma and κ\kappa are related by the following convolution:

𝜸(𝐱)=1π𝐃(𝐱𝐱)κ(𝐱)d𝐱,\bm{\gamma}\mathbf{(x)}=\frac{1}{\pi}\int\mathbf{D}(\mathbf{x}-\mathbf{x^{\prime}})\kappa(\mathbf{x^{\prime}})d\mathbf{x^{\prime}}, (4)

where the convolution kernel 𝐃\mathbf{D} is:

𝐃(𝐱)=1(x1𝐢x2)2.\mathbf{D}\mathbf{(x)}=-\frac{1}{(x_{1}-\mathbf{i}x_{2})^{2}}. (5)

II.2 Generation of the Training Data

We follow the procedure outlined in H21 for generating the training data. We use the publicly available convergence map from the MassiveNuS cosmological simulation [44]11 1 http://columbialensing.org. From the convergence map datasets at five different source redshifts (z=0.5,1.0,1.5,2.0,and2.5z=0.5,1.0,1.5,2.0,~\rm and~2.5), we select the dataset created for the source plane at z=2.5z=2.5. Note that this z=2.5z=2.5 source redshift of the convergence map is different from the source redshift that we assume for creating the WL galaxy shape catalog described below. In H21, we selected the convergence map for the source plane at z=1.5; however, we find that using the dataset for the source plane at z=2.5z=2.5 improves the reconstruction quality. Due to the higher lensing efficiency, the convergence map at z=2.5z=2.5 contains richer structures. One potential concern is the risk of bias, as the source plane at z=2.5z=2.5 is significantly higher than the typical value (z1z\sim 1) achieved with ground-based data. However, our CNN model is designed to learn the relationship between the shear catalog and the corresponding underlying mass distribution, rather than the characteristic convergence pattern at a specific source redshift. Consequently, we find that this particular choice of source redshift for our training dataset does not introduce a significant bias in the mass reconstruction. The dataset consists of 10,000 convergence maps covering a field size of 3.5×3.53\fdg 5\times 3\fdg 5 with a resolution of 512×512512\times 512 (0.410\sim 0\farcm 410 per grid cell).

When the nominal 10-year mission is completed, the LSST is expected to cover the entire southern sky with a depth of 27 mag. From the Subaru Hyper-Suprime Cam (HSC) archival images with the expected median seeing (0.7\raise 0.81805pt\hbox{$\scriptstyle\sim$}0.7\arcsec) of the LSST, we find that the average source number density is \scriptstyle\sim33 arcmin2{\rm arcmin}^{-2}. Since the Subaru/HSC imaging data share many properties with the Rubin Observatory, we adopt this number density as the average source density for our mock data [30]. The noise in the shape of each source galaxy is comprised of the shape dispersion and measurement components. As in H21, we adopt σe=0.24\sigma_{e}=0.24 as the intrinsic shape dispersion. For the measurement error, we utilize the relation between galaxy magnitude and measurement error determined from an analysis of Subaru/HSC data in HyeongHan et al. [30].

We divide the 10000 convergence maps into 7000 for training, 2800 for validation, and 200 for testing. Data augmentation is employed to increase the number of training data by applying four rotations and two-axis flips, but we do not use random crops as done in H21. The resulting number of convergence maps for training is 7000×(4+2)=420007000\times(4+2)=42000. Within the 3.\fdg5 x 3.\fdg5 field, there are no preferred locations for galaxy clusters unlike in H21.

II.3 CNN Model Architecture

Table 1: Architecture of our convolutional neural network.
Order Layer Filter Size Output Size
1 Input - - (3, 512, 512)
2 Conv2D-0 (9, 9) - (8, 512, 512)
Conv2D-1 (25, 25) - (8, 512, 512)
Conv2D-2 (49, 49) - (8, 512, 512)
3 Concatenate - Conv2D-0 (24, 512, 512)
Conv2D-1
Conv2D-2
4 Dropout-01 - - (24, 512, 512)
5 Conv2D-3 (49, 49) - (24, 464, 464)
6 TransConv2D-0 (49, 49) - (24, 512, 512)
7 Multiply - Concatenate (24, 512, 512)
8 Dropout-11 - - (24, 512, 512)
9 Conv2D-4 (49, 49) - (24, 464, 464)
10 TransConv2D-1 (49, 49) - (24, 512, 512)
11 Conv2D-5 (1, 1) - (1, 512, 512)
12 Output - - (1, 512, 512)

Note. — 1Dropout rate is 25 percent. 2The number of free parameters in our CNN model is 5.6×1065.6\times 10^{6}.

Refer to caption
Figure 1: Example of three-channel input layers. The left and middle panels show the ϵ1\epsilon_{1} and ϵ2\epsilon_{2} channels, respectively, while the right panel displays the Δϵ\Delta_{\epsilon} channel.

Table 1 presents the architecture of our new CNN model. The input layer is the 3×512×5123\times 512\times 512 array consisting of the three (ϵ1,ϵ2\epsilon_{1},\epsilon_{2}, and Δϵ\Delta\epsilon) channels. The first two channels are the average ellipticity components within the grid cell, while the last channel is the average ellipticity measurement error. Figure 1 presents an example of the three-channel input layers. Unlike in H21, we did not smooth the input layers with a Gaussian kernel22 2 H21 needed to smooth the input layers because the source density per pixel was much lower than in the current setup. . To mitigate overfitting associated with empty pixels, we applied smoothing to the input layers. However, in this study, we did not smooth the input layers because most pixels contain sufficient ellipticity information.

Our CNN model is composed of 8 two-dimensional (2D) convolutional layers, whereas the previous model used 11 2D convolutional layers. Instead of increasing the depth of our CNN model, we improved it by employing different filter sizes in the initial layers and concatenating them. This approach allows our CNN model to capture both small and large-scale structures with shallower layers than H21, making it more compact and efficient. Additionally, unlike in H21, we maintain the input layer size (512×512512\times 512) throughout the CNN operations to minimize resolution loss. The number of free parameters in our CNN model is 5.6×1065.6\times 10^{6}.

Similarly to the approach in H21, a skip connection is added after each combination of the convolution and transposed convolution layers. This measure is motivated by the residual neural network [24, ResNet;], which is shown to improve the overall performance of the CNN model by reducing the problem of vanishing gradients. The activation function for the convolutional layer is the Leaky Rectified Linear Activation (LReLU) function, which has a slope of 0.2 when x<0x<0. To prevent overfitting, we applied Dropout layers [56] after the Concatenate and Multiply layers, randomly excluding neurons during training. In this study, we excluded 25 percent of neurons per convolution layer. For training our CNN model, the Adam optimizer [40] was used with a learning rate of 10510^{-5}. We ran 1000 epochs and adopted gradient clipping for the optimizer [62] to prevent exploding gradients.

II.4 Loss Function

In H21, we utilized a weighted mean-squared-error loss function inspired by the focal loss [43]. This approach emphasized the importance of higher κ\kappa values. The loss function in H21 is given by:

L=𝒙ωf(𝒙)[κpred(𝒙)κtruth(𝒙)]2,L=\sum_{\bm{x}}\omega_{f}(\bm{x})[\kappa_{pred}(\bm{x})-\kappa_{truth}(\bm{x})]^{2}, (6)

where ωf(𝒙)\omega_{f}(\bm{x}) is a weight function defined as:

ωf(𝒙)=1+|κtruth|max(κtruth).\omega_{f}(\bm{x})=1+\frac{|\kappa_{truth}|}{{\rm max}(\kappa_{truth})}. (7)

In this study, we aim to reconstruct both high-contrast individual clusters and more diffuse large-scale structures. The ratio of high (\gtrsim0.2) to low (\lesssim0.1) κ\kappa pixels in our WF dataset is lower than that in H21, rendering the original weighting scheme less effective for our purposes. Thus, we have modified the loss function to better suit the broader range of structures as follows:

L=𝒙ωf(𝒙)6[κpred(𝒙)κtruth(𝒙)]2,L=\sum_{\bm{x}}{\omega_{f}(\bm{x})}^{6}[\kappa_{pred}(\bm{x})-\kappa_{truth}(\bm{x})]^{2}, (8)

where ωf(𝒙)\omega_{f}(\bm{x}) is the weight defined in Equation 7.

This revision amplifies the influence of high-κ\kappa pixels more than the original weighting scheme. The choice of the exponent 6 is empirically determined to provide optimal balance and training stability, enhancing the model’s sensitivity to high-κ\kappa pixels while preventing overfitting. By increasing the weight of high-κ\kappa pixels, the model is better tailored to accurately predict these pixels while still maintaining its ability to reconstruct low-κ\kappa pixels without overfitting.

III Result

To evaluate the performance of our CNN model, we performed mass reconstruction using 200 test datasets that were not used either in training or validation. Additionally, for comparison, we generated mass maps employing the H21 CNN model and the KS93 method [37]. For the KS93 analysis, we use a smoothing scale of σ150\sigma\sim 150\arcsec, which is a common choice for smoothing shear data [31, e.g.,]. When displaying the results from the H21 method, we applied bicubic interpolation to match the dimension of the current results, as the H21 implementation provides lower-resolution outputs.

III.1 Visual Analysis

Refer to caption
Figure 2: Comparison of the κ\kappa maps derived from the truth, our CNN model, the KS93 method, and the H21 across five test datasets. The top row displays the truth convergence maps, the second row shows the maps reconstructed by our CNN model, the third row presents the maps reconstructed using the model in H21, and the bottom row exhibits the maps from the KS93 method. LL indicates the loss values derived from Equation 8. Each panel displays a field of 3.5×3.53\fdg 5\times 3\fdg 5.

Figure 2 compares the mass reconstruction results from the current CNN, H21, and KS93 methods for five randomly selected fields from the test dataset. The κ\kappa values from our CNN model range from 0.05\raise 0.81805pt\hbox{$\scriptstyle\sim$}-0.05 to 0.3\raise 0.81805pt\hbox{$\scriptstyle\sim$}0.3, consistent with the true convergence maps. Notably, the reconstructed maps from our CNN model exhibit good agreement in features with the true κ\kappa values for both low (\lesssim0.1) and high (\gtrsim0.2) κ\kappa values, marking an improvement over the results from the H21 model. In contrast, the KS93 reconstructions tend to significantly overestimate the κ\kappa values mainly due to the mass-sheet degeneracy. We also found that the residuals between the truth and reconstructed mass maps from our CNN model do not present any significant systematic pattern.

The model from H21 tends to smooth out small-scale structures, a trend discernible across all samples in Figure 2. We attribute this to the reduction of the input channels in the first convolutional layer. While this approach was introduced to enhance training efficiency, the reduction makes it difficult to capture small-scale structures. The KS93 results show correlations with the large-scale structures of the true mass distribution but significantly smooth out finer details. Of course, one can consider decreasing the kernel size for KS93. However, when we decreased the kernel size to restore the compactness of the highest peaks, the result showed numerous spurious fluctuations across the field. This is not surprising because the reduced kernel size is inadequate in low density regions.

Finally, the superiority of our new CNN model can also be evaluated by the loss LL (see the value in the lower right corner of each panel in Figure 2) evaluated with Equation 8. The LL value from our new CNN results spans the range 200210200-210, which is 1020%10-20\% lower than those from the H21 results. The KS93 result shows significantly higher LL values.

III.2 Direct Comparison of Convergence Values

Refer to caption
Figure 3: Correlation between the truth and the reconstructed κ\kappa values pixel-by-pixel. The upper panel (middle and lower panel) shows the correlation for our CNN model (model in H21 and KS93 method). The x-axis and y-axis indicate the truth and the reconstructed κ\kappa values, respectively. The solid green (dashed magenta) lines represent perfect correlations (total least squares regression results).
Figure 4: PDF of κ\kappa. The black solid line indicates the PDF of κ\kappa of the truth maps. The red (blue and green) solid line represents the probability distribution of κ\kappa from our CNN model (KS93 method and model in H21).

Following the qualitative visual analysis in §III.1, we present a quantitative evaluation by directly comparing convergence values. Figure 3 displays pixel-by-pixel correlations between the truth and reconstructed κ\kappa maps. We find that the current CNN results exhibit a stronger correlation, with slopes of 0.47\raise 0.81805pt\hbox{$\scriptstyle\sim$}0.47 and 0.70\raise 0.81805pt\hbox{$\scriptstyle\sim$}0.70 for the H21 and current models, respectively. In contrast, the slope deviates substantially from unity for the KS93 mass reconstruction.

Figure 4 compares the probability distribution function (PDF) of κ\kappa values. Once again, it is evident that the distribution derived from the current CNN model matches the truth more closely than the H21 model. In particular, the agreement is significantly improved in the recovery of high κ\kappa values (κ>0.1\kappa>0.1). The distribution is considerably wider in the KS93 reconstruction.

III.3 κκ\kappa-\kappa Auto-Correlation Functions

Figure 5: κκ\kappa-\kappa auto-correlation functions. The black (red) solid line shows the median of κκ\kappa-\kappa correlation functions from the 200 truth (reconstructed) convergence maps. The shaded regions indicate 1σ\sigma uncertainties. The power is well recovered at both small and large scales. The discrepancy at 2\lesssim 2 pixels (0.8\lesssim 0.8 arcmin) is due to the inevitable resolution loss in the reconstruction caused by the finite source density.

In §III.1 and §III.2, we evaluated the fidelity of our CNN mass reconstruction by comparing 2D mass maps, pixel-by-pixel values, and κ\kappa distributions. Here, we assess the similarity in structures between the reconstructed map and truth by comparing their κκ\kappa-\kappa auto-correlation functions.

To calculate the auto-correlation function, we used TreeCorr 33 3 https://github.com/rmjarvis/TreeCorr [34, 33] on 200 mass maps from the truth and the reconstruction. Figure 5 displays the results. We find that the correlation is well recovered at both small and large scales. The discrepancy at 2\lesssim 2 pixels (0.8\lesssim 0.8 arcmin) is due to the inevitable resolution loss in the reconstruction caused by the finite source density.

III.4 Cluster Detection

Refer to caption
Figure 6: Projected cluster masses within 1 Mpc radius between the truth and reconstruction maps from our CNN model. The black solid (magenta dashed) line indicates a perfect correlation (linear regression) between the truth and reconstructed projected masses. The magenta shade shows the uncertainty.
Figure 7: Completeness of the cluster detection. The x-axis and y-axis indicate the cluster mass for detection and completeness, respectively. The red (blue and green) line and shaded region represent the median, and 16th and 84th percentiles of the completeness from our CNN model (KS93 method and model in H21).

One of the key merits of future WL observations is their ability to detect galaxy clusters solely based on projected mass maps. This method offers significant advantages over traditional approaches such as galaxy overdensity mapping, Sunyaev-Zeldovich effect observations, and X-ray surveys because the signals directly correlate with their masses.

We assess the cluster detection capability of our CNN model by the following. First, we identify cluster candidates as local peaks on both the truth and reconstructed maps. We then compute projected masses within a 1 Mpc radius for each identified cluster and evaluate the detection performance by comparing the projected mass and completeness. Clusters from the truth and CNN results were matched based on a distance criterion, which we set to 6 pixels (2.4\sim 2\farcm 4).

Figure 6 compares the truth and CNN reconstructed projected cluster masses. The slope from the linear regression (dashed magenta line) is very close to the one-to-one correlation (solid black line), indicating that the convergence values obtained directly from the CNN maps serve as accurate mass proxies.

In Figure 7, we present the completeness of cluster detection, defined as the ratio of the number of detections in the reconstructed mass maps to those in the true convergence maps, as a function of the cluster mass. All three methods achieve 100%100\% completeness for clusters with masses larger than 3.5×1014M\raise 0.81805pt\hbox{$\scriptstyle\sim$}3.5\times 10^{14}M_{\odot}. Our CNN model maintains a completeness of 75%\raise 0.81805pt\hbox{$\scriptstyle\sim$}75\% down to 1014M\raise 0.81805pt\hbox{$\scriptstyle\sim$}10^{14}M_{\odot}. The KS93 completeness is much lower, yielding 35%\raise 0.81805pt\hbox{$\scriptstyle\sim$}35\% at 1014M\raise 0.81805pt\hbox{$\scriptstyle\sim$}10^{14}M_{\odot}. The completeness measured from the H21 result is 10\raise 0.81805pt\hbox{$\scriptstyle\sim$}10% lower than that of the current version in the 2.5×1014M\lesssim 2.5\times 10^{14}M_{\odot} regime.

In the upcoming era of the Vera C. Rubin Observatory, the Roman Space Telescope, and the Euclid mission, WF WL mass reconstructions will be performed routinely. By leveraging the unique property of WL, which eliminates the need for dynamical assumptions (i.e., it is only influenced by the gravitational potential), our CNN model would be an effective and independent tool for detecting galaxy cluster candidates. This capability enables a range of studies, such as the cluster mass function reconstruction [47, 18, 11] and cosmology [61, 59, 28]. In addition, the independently obtained galaxy cluster candidate catalogs will synergize with other cluster surveys across various wavelengths [10, 41, 15].

IV Discussion

IV.1 Null Test

Figure 8: κ\kappa distribution from the 1000 null fields. The results from the null test have the κ\kappa values within 0.02κ0.04-0.02\lesssim\kappa\lesssim 0.04.

In the evaluation of machine learning (ML) models, a null test serves as a critical validation step to ensure that the model does not produce false positive signals from random noise fields. Since H21 placed clusters always near the center of the field in the training stage, a moderate level of bias was observed for null input data.

In order to investigate the same issue for the current model, we generated 1000 mock WL observations from null fields (i.e., κ=0\kappa=0) with the same observational properties (i.e., including measurement and shape noise) as the training data and reconstructed the corresponding mass maps. Figure 8 shows the κ\kappa distributions from these mass maps. The κ\kappa values range between -0.04 and 0.06, and the distribution is positively skewed, with a mode around 0.01-0.01. However, compared to H21, the distribution is relatively symmetric, and the bias level is very low. This improvement is somewhat expected, as the current training dataset is produced from random cosmological ray tracing fields. Our null test provides confidence in the robustness of the model and its ability to produce reliable mass reconstructions under realistic observational conditions.

IV.2 Boundary & Finite-Field Effect

Refer to caption
Figure 9: Comparison between the truth and the reconstructed convergence maps using masked WL observations. Same samples are used as Figure 2.

Traditional WL mass reconstructions are affected by boundary and finite-field effects due to the lack of information beyond the mass reconstruction field. In §III.1, our CNN results do not exhibit these artifacts for square boundaries.

Real observations often have irregular observational footprints, which can introduce biases in mass reconstruction. To investigate these effects under more realistic conditions, we reconstruct κ\kappa maps with masked input data. We adopt the observational footprint of the Subaru/HSC survey of the Coma cluster presented in HyeongHan et al. [30] as our mask. Figure 9 shows the resulting mass reconstruction.

While all three methods can restore the overall mass distributions, we find that the current CNN results are the most satisfactory. Again, the KS93 maps show highly inaccurate convergence values. The H21 result is significantly better that the KS93 result; However, its convergence values are systematically lower than the true values. This experiment demonstrates that the current CNN model can be applied to real observations, where complex boundary shapes are present.

IV.3 Application to Real Observations: Mass Reconstruction of the Coma Cluster Field

Refer to caption
Figure 10: Mass reconstruction of the Coma cluster field. The left panel shows the reconstructed mass map based on the current CNN model, while the right panel overlays the filament candidates (pink solid) detected with DisPerSE on the same mass map. The convergence field closely follows the Coma cluster substructures traced by its member galaxies. See text for the discussion on the filament detection.
Figure 11: Azimuthal distribution of the convergence in the Coma field. The red solid line indicates the mean κ\kappa within each azimuthal bin. The black dashed (dotted) line shows the direction of the intracluster filament (candidate) reported by [30].

While ML has shown remarkable promise in various fields, it is not without its pitfalls. A significant challenge arises when models that perform exceptionally well on training data fail to generalize to real-world applications. ML models are often sensitive to the quality and representativeness of the training data. If the training set does not adequately reflect the diversity and complexity of real-world scenarios, the model may become biased or overly simplistic. This lack of robustness can result in poor performance in practical applications, where conditions may vary significantly from those encountered during training. Additionally, changes in the underlying data distribution, known as dataset shift, can further exacerbate these issues, rendering a once-effective model ineffective.

To ensure that the current CNN model excels not only on training datasets but also on real-world data, we experimented with actual WF WL observations. We present the mass reconstruction of the Coma cluster in the left panel of Figure 10. We used the WL shear catalog from [30], which measured WL signals from Subaru/Hyper-Suprime Cam imaging data covering approximately 3.5×3.53\fdg 5\times 3\fdg 5 centered on its two BCGs. The Coma cluster, being a very massive cluster, may serve as an excellent real-world test case, as its characteristics are expected to differ significantly from those of the average cluster in our training dataset.

As the true mass distribution in the Coma field is unknown, we cannot assess its fidelity using the same metrics presented in §III. Instead, we provide comparisons with previous results as follows. First, we find that the mass estimate derived from our convergence field is consistent with previous studies. Fitting an NFW profile provides M200=7.43.12.4×1014MM_{200}=7.4_{-3.1}^{2.4}\times 10^{14}M_{\sun}, which aligns with literature values [21, 48, e.g.,]. Second, the reconstruction reveals a number of significant substructures, some of which were also reported by previous low-resolution WL studies. Third, the mass distribution is well traced by those of the Coma cluster galaxies and intracluster light [48, 36, e.g.,].

Most notably, the CNN mass map reveals some of its intracluster filaments reported by [30]. Intracluster filaments are the terminal ends of intercluster filaments that stretch over tens of megaparsecs. Due to their low contrast, their detection has been considered elusive. [30] identified the Coma cluster’s intracluster filaments not through mass reconstruction, but using two novel methods: the matched-filter technique and the shear peak statistic [45]. Here, we independently detect the intracluster filament of the Coma based on the current CNN mass map as follows. First, we ran DisPerSE44 4 https://www2.iap.fr/users/sousbie/web/html/indexd41d.html [53, 54] on the CNN mass map. To set the threshold cut for filament detection, we adopt a persistence above the 3-sigma level, where persistence indicates the absolute difference between the saddle and maximum critical points. The solid pink lines in the right panel of Figure 10 depict the resulting filaments. The three main branches originating from the cluster center is clearly detected and aligns well with the result from [30]. Second, we measured the azimuthal excess of the convergence with respect to the background value. Figure 11 shows that this method too reveals the three main filaments identified by DisPerSE, which is also consistent with the findings of [30] (compare with Figure 2 of their paper). The consistency in intracluster filament detection showcases the utility of CNN-based WL mass reconstruction for identifying large-scale structures in future WL surveys, including the detection of cosmic webs. We refer readers to HyeongHan et al. [30] for further scientific details on the Coma cluster.

V Summary & Conclusion

In this paper, we have presented an improved multi-layered CNN model for WF WL mass reconstruction and evaluated its performance using mock next-generation WF observations. Our enhancements include a more compact and efficient CNN architecture and a refined loss function that targets the recovery of both high-contrast individual clusters and diffuse large-scale structures.

We find that the current CNN model significantly outperforms the H21 version across several metrics. Visual comparisons of the 2D convergence maps demonstrate that both large-scale structures and high-density cluster peaks have been remarkably well restored. Pixel-by-pixel comparisons reveal a stronger correlation, and the convergence distribution obtained from the current CNN model aligns closely with the truth. However, improvements are still needed at both extremes of the distribution.

Additionally, we demonstrate that our current CNN model serves as a reliable tool for detecting galaxy clusters based solely on gravitational lensing data. The detection completeness reaches 75% at 1014M\raise 0.81805pt\hbox{$\scriptstyle\sim$}10^{14}M_{\odot} and 90% at 3×1014M\raise 0.81805pt\hbox{$\scriptstyle\sim$}3\times 10^{14}M_{\odot}. The projected cluster masses estimated from the mass map show a near one-to-one correlation with the true values.

Our CNN model maintains robust performance even when the field boundaries are irregular. While the convergence scales near the field boundaries are severely distorted in the H21 result, artifacts in our new mass reconstruction are negligible.

Finally, we use Subaru/HSC WF WL data of the Coma cluster to verify the robustness of our CNN model when applied to real-world data. The Coma cluster, as a very massive system, provides an excellent real-world test case since its characteristics are expected to differ significantly from those of the average cluster in our training dataset. We find that the mass estimates and distributions are highly consistent with literature values and luminous tracers. We also demonstrate that the CNN mass map alone can be used to identify intracluster filaments reported by independent methods such as the matched-filter statistic and shear-peak count.

Current and future WL surveys will cover a significant fraction of the sky, necessitating high-fidelity mass reconstruction on a routine basis with arbitrary observational footprints. We anticipate that our CNN method will be a valuable tool, enabling fast and robust WF mass reconstructions.

We thank the Columbia Lensing group (http://columbialensing.org) for their simulations available. The creation of these simulations is supported through grants NSF AST-1210877, NSF AST-140041, and NASA ATP-80NSSC18K1093. We thank New Mexico State University (USA) and Instituto de Astrofisica de Andalucia CSIC (Spain) for hosting the Skies & Universes site for cosmological simulation products. This work was supported by the Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT; No.2021-0-02068, Artificial Intelligence Innovation Hub). M. J. Jee acknowledges support for the current research from the National Research Foundation (NRF) of Korea under the programs 2022R1A2C1003130 and RS-2023-00219959. SC acknowledges this research was supported by Basic Science Research Program through the National Research Foundation of Korea (NRF) funded by the Ministry of Education (No. RS-2024-00413036). S.E.H. is supported by the Korea Astronomy and Space Science Institute grant funded by the Korea government (MSIT) (No. 2024186901). SP and DB were supported in part by NRF Grant RS-2023-00208011, and by Basic Science Research Program through NRF funded by the Ministry of Education (2018R1A6A1A06024).

References

  • [1] Abadi, M., Agarwal, A., Barham, P., et al. 2015, TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems, software available from tensorflow.org
  • [2] Ahn, E., Jee, M. J., Lee, W., Joo, H., & ZuHone, J. 2024, ApJ, 973, 79
  • [3] Astropy Collaboration, Robitaille, T. P., Tollerud, E. J., et al. 2013, A&A, 558, A33
  • [4] Astropy Collaboration, Price-Whelan, A. M., Sipőcz, B. M., et al. 2018, AJ, 156, 123
  • [5] Astropy Collaboration, Price-Whelan, A. M., Lim, P. L., et al. 2022, ApJ, 935, 167
  • [6] Bartelmann, M. 1995, A&A, 303, 643
  • [7] Bartelmann, M., Narayan, R., Seitz, S., & Schneider, P. 1996, ApJ, 464, L115
  • [8] Bartelmann, M., & Schneider, P. 2001, Phys. Rep., 340, 291
  • [9] Bertin, E., & Arnouts, S. 1996, A&AS, 117, 393
  • [10] Bleem, L. E., Stalder, B., de Haan, T., et al. 2015, ApJS, 216, 27
  • [11] Böhringer, H., Chon, G., & Fukugita, M. 2017, A&A, 608, A65
  • [12] Bradač, M., Lombardi, M., & Schneider, P. 2004, A&A, 424, 13
  • [13] Bradač, M., Schneider, P., Lombardi, M., & Erben, T. 2005, A&A, 437, 39
  • [14] Bridle, S. L., Hobson, M. P., Lasenby, A. N., & Saunders, R. 1998, MNRAS, 299, 895
  • [15] Bulbul, E., Liu, A., Kluge, M., et al. 2024, A&A, 685, A106
  • [16] Cha, S., HyeongHan, K., Scofield, Z. P., Joo, H., & Jee, M. J. 2024, ApJ, 961, 186
  • [17] Cha, S., & Jee, M. J. 2022, ApJ, 931, 127
  • [18] Despali, G., Giocoli, C., Angulo, R. E., et al. 2016, MNRAS, 456, 2486
  • [19] Falco, E. E., Gorenstein, M. V., & Shapiro, I. I. 1985, ApJ, 289, L1
  • [20] Finner, K., Jee, M. J., Golovich, N., et al. 2017, ApJ, 851, 46
  • [21] Gavazzi, R., Adami, C., Durret, F., et al. 2009, A&A, 498, L33
  • [22] Geiger, B., & Schneider, P. 1999, MNRAS, 302, 118
  • [23] Harris, C. R., Millman, K. J., van der Walt, S. J., et al. 2020, Nature, 585, 357
  • [24] He, K., Zhang, X., Ren, S., & Sun, J. 2016, in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770–778
  • [25] Hoekstra, H., Bartelmann, M., Dahle, H., et al. 2013, Space Sci. Rev., 177, 75
  • [26] Hoekstra, H., Franx, M., & Kuijken, K. 2000, ApJ, 532, 88
  • [27] Hong, S. E., Park, S., Jee, M. J., Bak, D., & Cha, S. 2021, ApJ, 923, 266
  • [28] Hung, D., Lemaux, B. C., Gal, R. R., et al. 2021, MNRAS, 502, 3942
  • [29] Hunter, J. D. 2007, Computing in Science & Engineering, 9, 90
  • [30] HyeongHan, K., Jee, M. J., Cha, S., & Cho, H. 2024a, Nature Astronomy, doi:10.1038/s41550-023-02164-w
  • [31] HyeongHan, K., Jee, M. J., Lee, W., et al. 2024b, arXiv e-prints, arXiv:2405.00115
  • [32] Ivezić, Ž., Kahn, S. M., Tyson, J. A., et al. 2019, ApJ, 873, 111
  • [33] Jarvis, M. 2015, TreeCorr: Two-point correlation functions, Astrophysics Source Code Library, record ascl:1508.007, ascl:1508.007
  • [34] Jarvis, M., Bernstein, G., & Jain, B. 2004, MNRAS, 352, 338
  • [35] Jee, M. J., White, R. L., Ford, H. C., et al. 2006, ApJ, 642, 720
  • [36] Jiménez-Teja, Y., Román, J., HyeongHan, K., et al. 2024, arXiv e-prints, arXiv:2412.15328
  • [37] Kaiser, N., & Squires, G. 1993, ApJ, 404, 441
  • [38] Khiabanian, H., & Dell’Antonio, I. P. 2008, ApJ, 684, 794
  • [39] Kim, J., Jee, M. J., Hughes, J. P., et al. 2021, ApJ, 923, 101
  • [40] Kingma, D. P., & Ba, J. 2014, arXiv e-prints, arXiv:1412.6980
  • [41] Knowles, K., Cotton, W. D., Rudnick, L., et al. 2022, A&A, 657, A56
  • [42] Lanusse, F., Mandelbaum, R., Ravanbakhsh, S., et al. 2021, MNRAS, 504, 5543
  • [43] Lin, T.-Y., Goyal, P., Girshick, R., He, K., & Dollar, P. 2017, in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
  • [44] Liu, J., Bird, S., Zorrilla Matilla, J. M., et al. 2018, J. Cosmology Astropart. Phys, 2018, 049
  • [45] Maturi, M., & Merten, J. 2013, A&A, 559, A112
  • [46] Medezinski, E., Broadhurst, T., Umetsu, K., et al. 2007, ApJ, 663, 717
  • [47] Mehrtens, N., Romer, A. K., Nichol, R. C., et al. 2016, MNRAS, 463, 1929
  • [48] Okabe, N., Futamase, T., Kajisawa, M., & Kuroshima, R. 2014, ApJ, 784, 90
  • [49] Park, H., Jo, Y., Kang, S., Kim, T., & Jee, M. J. 2024, ApJ, 972, 45
  • [50] Seitz, S., & Schneider, P. 1996, A&A, 305, 383
  • [51] Seitz, S., Schneider, P., & Bartelmann, M. 1998, A&A, 337, 325
  • [52] Shirasaki, M., Yoshida, N., & Ikeda, S. 2019, Phys. Rev. D, 100, 043527
  • [53] Sousbie, T. 2011, MNRAS, 414, 350
  • [54] Sousbie, T., Pichon, C., & Kawahara, H. 2011, MNRAS, 414, 384
  • [55] Squires, G., & Kaiser, N. 1996, ApJ, 473, 65
  • [56] Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. 2014, Journal of Machine Learning Research, 15, 1929
  • [57] Sweere, S. F., Valtchanov, I., Lieu, M., et al. 2022, MNRAS, 517, 4054
  • [58] Umetsu, K., Broadhurst, T., Zitrin, A., Medezinski, E., & Hsu, L.-Y. 2011, ApJ, 729, 127
  • [59] Vikhlinin, A., Kravtsov, A. V., Burenin, R. A., et al. 2009, ApJ, 692, 1060
  • [60] Virtanen, P., Gommers, R., Oliphant, T. E., et al. 2020, Nature Methods, 17, 261
  • [61] Warren, M. S., Abazajian, K., Holz, D. E., & Teodoro, L. 2006, ApJ, 646, 881
  • [62] Zhang, J., He, T., Sra, S., & Jadbabaie, A. 2020, in International Conference on Learning Representations