-
Carrier scattering considerations and thermoelectric power factors of half-Heuslers
Authors:
Rajeev Dutt,
Bhawna Sahni,
Yao Zhao,
Yuji Go,
Saff E Awal Akhtar,
Ankit Kumar,
Sumit Kukreti,
Patrizio Graziosi,
Zhen Li,
Neophytos Neophytou
Abstract:
The electronic and thermoelectric (TE) transport properties of 13 n-type and p-type half-Heusler alloys are computationally examined using Boltzmann transport. The electronic scattering times resulting from all relevant phonon interactions and ionized impurity scattering (IIS) are fully accounted for using ab initio extracted parameters. We find that at room temperature the average peak TE power f…
▽ More
The electronic and thermoelectric (TE) transport properties of 13 n-type and p-type half-Heusler alloys are computationally examined using Boltzmann transport. The electronic scattering times resulting from all relevant phonon interactions and ionized impurity scattering (IIS) are fully accounted for using ab initio extracted parameters. We find that at room temperature the average peak TE power factors (PF) of all materials we examine reside between 5 and 10 mW/mK$^2$. We also find that IIS in combination with the long range polar optical phonon (POP) scattering are more influential in determining the electronic transport and PF over all other non-polar phonon interactions (acoustic and optical phonon transport). In fact, the combination of POP and IIS determines the thermoelectric power factor of the half-Heuslers examined on average by about 65\%. The results highlight the crucial impact of Coulombic scattering process (POP and IIS) on the TE properties of half-Heusler alloys and provide profound insight for understanding transport, which can be applied widely in other complex bandstructure materials. In terms of computation expense, the computationally cheaper POP and IIS provide an acceptable first-order estimate of the power factor of these materials, while the non-polar contributions, which require more expensive ab initio calculations, could be of secondary importance.
△ Less
Submitted 23 April, 2026;
originally announced April 2026.
-
Robust and Generalizable Background Subtraction on Images of Calorimeter Jets using Unsupervised Generative Learning
Authors:
Yeonju Go,
Dmitrii Torbunov,
Yi Huang,
Shuhang Li,
Timothy Rinn,
Haiwang Yu,
Brett Viren,
Meifeng Lin,
Yihui Ren,
Dennis Perepelitsa,
Jin Huang
Abstract:
Accurate separation of signal from background is one of the main challenges for precision measurements across high-energy and nuclear physics. Conventional supervised learning methods are insufficient here because the required paired signal and background examples are impossible to acquire in real experiments. Here, we introduce an unsupervised unpaired image-to-image translation neural network th…
▽ More
Accurate separation of signal from background is one of the main challenges for precision measurements across high-energy and nuclear physics. Conventional supervised learning methods are insufficient here because the required paired signal and background examples are impossible to acquire in real experiments. Here, we introduce an unsupervised unpaired image-to-image translation neural network that learns to separate the signal and background from the input experimental data using cycle-consistency principles. We demonstrate the efficacy of this approach using images composed of simulated calorimeter data from the sPHENIX experiment, where physics signals (jets) are immersed in the extremely dense and fluctuating heavy-ion collision environment. Our method outperforms conventional subtraction algorithms in fidelity and overcomes the limitations of supervised methods. Furthermore, we evaluated the model's robustness in an out-of-distribution test scenario designed to emulate modified jets as in real experimental data. The model, trained on a simpler dataset, maintained its high fidelity on a more realistic, highly modified jet signal. This work represents the first use of unsupervised unpaired generative models for full detector jet background subtraction and offers a path for novel applications in real experimental data, enabling high-precision analyses across a wide range of imaging-based experiments.
△ Less
Submitted 27 October, 2025;
originally announced October 2025.
-
TPCpp-10M: Simulated proton-proton collisions in a Time Projection Chamber for AI Foundation Models
Authors:
Shuhang Li,
Yi Huang,
David Park,
Xihaier Luo,
Haiwang Yu,
Yeonju Go,
Christopher Pinkenburg,
Yuewei Lin,
Shinjae Yoo,
Joseph Osborn,
Christof Roland,
Jin Huang,
Yihui Ren
Abstract:
Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field is often limited by the lack of openly available large scale datasets, as well as standardized evaluation tasks and metrics. Furthermore, the specialized knowledge and software typically required to process particle physi…
▽ More
Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field is often limited by the lack of openly available large scale datasets, as well as standardized evaluation tasks and metrics. Furthermore, the specialized knowledge and software typically required to process particle physics data pose significant barriers to interdisciplinary collaboration with the broader machine learning community.
This work introduces a large, openly accessible dataset of 10 million simulated proton-proton collisions, designed to support self-supervised training of foundation models. To facilitate ease of use, the dataset is provided in a common NumPy format. In addition, it includes 70,000 labeled examples spanning three well defined downstream tasks: track finding, particle identification, and noise tagging, to enable systematic evaluation of the foundation model's adaptability.
The simulated data are generated using the Pythia Monte Carlo event generator at a center of mass energy of sqrt(s) = 200 GeV and processed with Geant4 to include realistic detector conditions and signal emulation in the sPHENIX Time Projection Chamber at the Relativistic Heavy Ion Collider, located at Brookhaven National Laboratory.
This dataset resource establishes a common ground for interdisciplinary research, enabling machine learning scientists and physicists alike to explore scaling behaviors, assess transferability, and accelerate progress toward foundation models in nuclear and high energy physics. The complete simulation and reconstruction chain is reproducible with the sPHENIX software stack. All data and code locations are provided under Data Accessibility.
△ Less
Submitted 6 September, 2025;
originally announced September 2025.
-
Variable Rate Neural Compression for Sparse Detector Data
Authors:
Yi Huang,
Yeonju Go,
Jin Huang,
Shuhang Li,
Xihaier Luo,
Thomas Marshall,
Joseph Osborn,
Christopher Pinkenburg,
Yihui Ren,
Evgeny Shulga,
Shinjae Yoo,
Byung-Jun Yoon
Abstract:
High-energy large-scale particle colliders generate data at extraordinary rates. Developing real-time high-throughput data compression algorithms to reduce data volume and meet the bandwidth requirement for storage has become increasingly critical. Deep learning is a promising technology that can address this challenging topic. At the newly constructed sPHENIX experiment at the Relativistic Heavy…
▽ More
High-energy large-scale particle colliders generate data at extraordinary rates. Developing real-time high-throughput data compression algorithms to reduce data volume and meet the bandwidth requirement for storage has become increasingly critical. Deep learning is a promising technology that can address this challenging topic. At the newly constructed sPHENIX experiment at the Relativistic Heavy Ion Collider, a Time Projection Chamber (TPC) serves as the main tracking detector, which records three-dimensional particle trajectories in a volume of a gas-filled cylinder. In terms of occupancy, the resulting data flow can be very sparse reaching $10^{-3}$ for proton-proton collisions. Such sparsity presents a challenge to conventional learning-free lossy compression algorithms, such as SZ, ZFP, and MGARD. In contrast, emerging deep learning-based models, particularly those utilizing convolutional neural networks for compression, have outperformed these conventional methods in terms of compression ratios and reconstruction accuracy. However, research on the efficacy of these deep learning models in handling sparse datasets, like those produced in particle colliders, remains limited. Furthermore, most deep learning models do not adapt their processing speeds to data sparsity, which affects efficiency. To address this issue, we propose a novel approach for TPC data compression via key-point identification facilitated by sparse convolution. Our proposed algorithm, BCAE-VS, achieves a $75\%$ improvement in reconstruction accuracy with a $10\%$ increase in compression ratio over the previous state-of-the-art model. Additionally, BCAE-VS manages to achieve these results with a model size over two orders of magnitude smaller. Lastly, we have experimentally verified that as sparsity increases, so does the model's throughput.
△ Less
Submitted 18 November, 2024;
originally announced November 2024.
-
Effectiveness of denoising diffusion probabilistic models for fast and high-fidelity whole-event simulation in high-energy heavy-ion experiments
Authors:
Yeonju Go,
Dmitrii Torbunov,
Timothy Rinn,
Yi Huang,
Haiwang Yu,
Brett Viren,
Meifeng Lin,
Yihui Ren,
Jin Huang
Abstract:
Artificial intelligence (AI) generative models, such as generative adversarial networks (GANs), variational auto-encoders, and normalizing flows, have been widely used and studied as efficient alternatives for traditional scientific simulations. However, they have several drawbacks, including training instability and inability to cover the entire data distribution, especially for regions where dat…
▽ More
Artificial intelligence (AI) generative models, such as generative adversarial networks (GANs), variational auto-encoders, and normalizing flows, have been widely used and studied as efficient alternatives for traditional scientific simulations. However, they have several drawbacks, including training instability and inability to cover the entire data distribution, especially for regions where data are rare. This is particularly challenging for whole-event, full-detector simulations in high-energy heavy-ion experiments, such as sPHENIX at the Relativistic Heavy Ion Collider and Large Hadron Collider experiments, where thousands of particles are produced per event and interact with the detector. This work investigates the effectiveness of Denoising Diffusion Probabilistic Models (DDPMs) as an AI-based generative surrogate model for the sPHENIX experiment that includes the heavy-ion event generation and response of the entire calorimeter stack. DDPM performance in sPHENIX simulation data is compared with a popular rival, GANs. Results show that both DDPMs and GANs can reproduce the data distribution where the examples are abundant (low-to-medium calorimeter energies). Nonetheless, DDPMs significantly outperform GANs, especially in high-energy regions where data are rare. Additionally, DDPMs exhibit superior stability compared to GANs. The results are consistent between both central and peripheral centrality heavy-ion collision events. Moreover, DDPMs offer a substantial speedup of approximately a factor of 100 compared to the traditional Geant4 simulation method.
△ Less
Submitted 30 January, 2025; v1 submitted 23 May, 2024;
originally announced June 2024.