-
ePIC Early Science Report
Authors:
D. Abbott,
N. Abdelrahman,
S. Abhijit,
I. Abualrob,
R. B. Achari,
J. Adam,
L. Adamczyk,
K. Adkins,
A. Affolder,
K. Agarwal,
J. Agarwala,
N. Agrawal,
C. A. Aidala,
W. Akers,
A. Al-bataineh,
S. N. Alam,
M. Alekseev,
P. R. Altieri,
J. -S. Alvarado Gallenao,
S. B. L. Amar,
R. Ammendola,
I. Amos Cali,
G. An,
D. Anderson,
E. Anderssen
, et al. (774 additional authors not shown)
Abstract:
This Early Science Report from the ePIC Collaboration outlines the compelling physics program achievable during the first years of operation of the Electron-Ion Collider (EIC), prior to the establishment of the full design luminosity and energy range. The analyses are based on realistic early-running beam configurations and detailed Geant4 ePIC detector simulations, hit digitization and data recon…
▽ More
This Early Science Report from the ePIC Collaboration outlines the compelling physics program achievable during the first years of operation of the Electron-Ion Collider (EIC), prior to the establishment of the full design luminosity and energy range. The analyses are based on realistic early-running beam configurations and detailed Geant4 ePIC detector simulations, hit digitization and data reconstruction. The projected studies from the physics working groups of ePIC span inclusive, semi-inclusive, exclusive, diffractive and tagging, as well as jet and heavy flavor measurements in both electron-proton and electron-ion collisions. Even before the collider reaches its full design performance, these measurements will constrain parton distribution functions in nucleons and nuclei, access transverse-momentum-dependent and spin-dependent observables, probe gluon dynamics in nuclei, and initiate a program of imaging of quarks and gluons. Each measurement is directly connected to the core science pillars of the EIC, identified in the 2018 report by the National Academy of Sciences: understanding the origin of the nucleon mass, unraveling the spin structure of the nucleon, and exploring the emergent properties of dense gluonic matter. The results presented here provide examples that demonstrate that the early years of EIC running with ePIC will deliver novel world-leading insights into Quantum Chromodynamics. In addition, the early science program will establish measurement and analysis methodologies that will pave the way to the subsequent full EIC physics program.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
Strong invariants and Tverberg numbers in convexity spaces
Authors:
Minho Cho,
Andreas F. Holmsen,
Attila Jung,
Hong Liu
Abstract:
Helly, Carathéodory, and Radon numbers encode three kinds of finite certificates in a convexity space: for the emptiness of an intersection, for membership in a convex hull, and for the existence of intersecting hulls. We study exact versions of these certificates, in which a subfamily must preserve the whole intersection or a subset must preserve the whole hull. Our first main result shows that,…
▽ More
Helly, Carathéodory, and Radon numbers encode three kinds of finite certificates in a convexity space: for the emptiness of an intersection, for membership in a convex hull, and for the existence of intersecting hulls. We study exact versions of these certificates, in which a subfamily must preserve the whole intersection or a subset must preserve the whole hull. Our first main result shows that, for finite configurations in an arbitrary convexity space, five a priori different boundedness conditions are equivalent: VC-dimension, strong Helly number, strong Carathéodory number, comatching number, and strong Radon number (with the expected additive-one shift). We also obtain equivalent layered Tverberg-type decompositions and colorful consequences.
The common mechanism is exposed by the bipartite incidence graph between points and a generating family. For finite spaces, the unique minimal generator yields a natural dual convexity space; we characterize double dualization and prove that the strong parameters are duality invariant. The same model gives a polynomial-size, $O(t^4)$, realization of Bukh's counterexample to the Calder-Eckhoff partition conjecture.
Finally, we obtain the first Tverberg bound for separable convexity spaces that is simultaneously linear in the number of parts and polynomial in the Radon number. If an $S_3$-separable convexity space has Helly number $h$ and its halfspaces have VC-dimension $d$, then $r_t=O(dh\log h)\,t$; in particular, Radon number $r$ gives $r_t=O(r^2\log r)\,t$. The bound attains the weak-Eckhoff scale $O(rt)$ whenever the Helly number is bounded. For axis-parallel box convexity in $\mathbb{R}^k$, gives the optimal order $r_t=O(rt)$ uniformly in every dimension. This appears to be the first dimension-uniform estimate of weak-Eckhoff order for box convexity, whereas the previous direct theory was confined to dimension three.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
Anatomy-Privileged Distillation with Token Routing for MRI-Based Prediction of Perineural Invasion
Authors:
Hyunsu Go,
Youngung Han,
Kyeonghun Kim,
Junga Kim,
Dohyun Kweon,
Jinyong Jun,
Sungha Park,
Anna Jung,
Induk Um,
Yului Jeong,
Suah Park,
Jina Jeong,
Pa Hong,
Woo Kyoung Jeong,
Won Jae Lee,
Ken Ying-Kai Liao,
Hyuk-Jae Lee,
Nam-Joon Kim
Abstract:
Perineural invasion (PNI) is associated with poor postoperative outcomes in intrahepatic cholangiocarcinoma, but it is confirmed by surgical pathology. Existing preoperative imaging models often rely on radiologist-defined variables, contrast-enhanced imaging, or manual annotations. We propose an anatomy-privileged teacher--student framework for patient-level PNI prediction from T2-weighted MRI. D…
▽ More
Perineural invasion (PNI) is associated with poor postoperative outcomes in intrahepatic cholangiocarcinoma, but it is confirmed by surgical pathology. Existing preoperative imaging models often rely on radiologist-defined variables, contrast-enhanced imaging, or manual annotations. We propose an anatomy-privileged teacher--student framework for patient-level PNI prediction from T2-weighted MRI. During training, the teacher uses MRI with tumor and liver masks to learn dense token routing, and the student distills this guidance to retain and aggregate informative tokens under a fixed budget. Anatomical supervision is restricted to training, and the deployed model does not require masks at inference. In 155 patients, the proposed method achieved the highest mean AUROC of 0.750 among matched MRI-only baselines evaluated under the same protocol, with 1.43 GFLOPs and 8.02 ms per case on a Jetson Orin Nano Super Developer Kit.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
SpikeDS: Dual Sparsity Spikformer for Perineural Invasion Prediction in 3D MRI
Authors:
Induk Um,
Youngung Han,
Kyeonghun Kim,
Yului Jeong,
Jina Jeong,
Hyunsu Go,
Dohyun Kweon,
Sungha Park,
Junga Kim,
Anna Jung,
Suah Park,
Hyuk-Jae Lee,
Pa Hong,
Woo Kyoung Jeong,
Won Jae Lee,
Ken Ying-Kai Liao,
Nam-Joon Kim
Abstract:
Perineural invasion (PNI) is associated with poor prognosis in cholangiocarcinoma (CCA). However, its detection from 3D MRI remains challenging due to the subtle and spatially heterogeneous imaging signatures at the tumor periphery. Capturing such spatially sparse cues necessitates volumetric analysis of 3D MRI, but existing deep learning approaches incur prohibitive computational costs on volumet…
▽ More
Perineural invasion (PNI) is associated with poor prognosis in cholangiocarcinoma (CCA). However, its detection from 3D MRI remains challenging due to the subtle and spatially heterogeneous imaging signatures at the tumor periphery. Capturing such spatially sparse cues necessitates volumetric analysis of 3D MRI, but existing deep learning approaches incur prohibitive computational costs on volumetric medical images, limiting their clinical deployment. We propose Dual Sparsity Spikformer (SpikeDS), a spiking neural network architecture that jointly exploits activation sparsity from binary spike communication and spatial sparsity from window pruning based on firing rates. SpikeDS introduces Dual Sparsity Spiking Attention (DSSA), which combines two complementary mechanisms. The first is Window-based Expert Mixture Spiking Attention (W-EMSA), which selectively applies attention only to salient windows identified by their firing rates. The second is Cross-Window Spiking Self-Attention (CW-SSA), which enables global context exchange through an asymmetric scheme in which pruned windows still contribute as key-value sources. Evaluated on a clinical cohort of 139 CCA patients via 5-fold cross-validation, SpikeDS achieves an AUC of 0.753 while consuming only 14.4 mJ, surpassing the best baseline in both AUC and energy efficiency. These results suggest that dual sparsity provides an effective hardware-aware strategy for improving the efficiency of 3D spiking transformers without compromising diagnostic performance.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Adaptive Routing for Efficient Diffusion Transformer-Based PNI Prediction
Authors:
Youngung Han,
Dohyun Kweon,
Kyeonghun Kim,
Hyunsu Go,
Jina Jeong,
Suah Park,
Induk Um,
Junga Kim,
Anna Jung,
Yului Jeong,
Sungha Park,
Jinyong Jun,
Pa Hong,
Woo Kyoung Jeong,
Won Jae Lee,
Ken Ying-Kai Liao,
Hyuk-Jae Lee,
Nam-Joon Kim
Abstract:
Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. However, its preoperative prediction from magnetic resonance imaging (MRI) remains challenging due to subtle imaging features that extend beyond tumor boundaries into surrounding regions. Conventional convolutional neural networks are limited in capturing long-range spatial dependencies. Transformer-based architecture…
▽ More
Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. However, its preoperative prediction from magnetic resonance imaging (MRI) remains challenging due to subtle imaging features that extend beyond tumor boundaries into surrounding regions. Conventional convolutional neural networks are limited in capturing long-range spatial dependencies. Transformer-based architectures improve global modeling of volumetric MRI by aggregating spatially distributed contextual cues, yet capturing subtle and noise-sensitive patterns in peritumoral regions remains challenging. Diffusion-based classifiers offer an alternative formulation by leveraging denoising-based class scoring to better capture such subtle patterns. However, these approaches introduce substantial computational overhead due to the combination of transformer-based modeling and iterative denoising processes. To address these challenges, we formulate PNI prediction as a diffusion-based classification problem and implement the denoising network using a transformer-based representation. To improve computational efficiency, we introduce adaptive routing across attention heads, spatial tokens, and MLP width. Experimental results demonstrate that the proposed approach achieves an AUC of 0.731 with 257.57 GFLOPs.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
The VC dimension of partial concept classes via Radon's theorem
Authors:
Grigory Ivanov,
Attila Jung,
Márton Naszódi
Abstract:
Following Alon, Hanneke, Holzman, and Moran (FOCS 2021), we define a partial concept class (PCC) as a family of partial functions \(f: V\to\{0,1,\ast\}\); equivalently, its concepts partition the ground set into black ($f^{-1}(1)$), grey ($f^{-1}(\ast)$), and white parts ($f^{-1}(0)$). Its VC dimension is defined by shattering sets on which the value $\ast$ is not taken. We study two geometric PCC…
▽ More
Following Alon, Hanneke, Holzman, and Moran (FOCS 2021), we define a partial concept class (PCC) as a family of partial functions \(f: V\to\{0,1,\ast\}\); equivalently, its concepts partition the ground set into black ($f^{-1}(1)$), grey ($f^{-1}(\ast)$), and white parts ($f^{-1}(0)$). Its VC dimension is defined by shattering sets on which the value $\ast$ is not taken. We study two geometric PCCs in real Banach spaces, both with a margin \(δ>0\): expanded half-spaces, where the grey part is a strip of width at least \(δ\) adjacent to a half-space, and expanded balls, where the grey part is an annulus of width \(δ\) around a unit radius ball.
Our main results are dimension-free upper bounds on the VC dimension of the PCC of expanded balls in \(L_p\parenthμ\), \(1\le p<\infty\), including the non-Euclidean and algorithmically particularly relevant case \(\ell^d_1\). These bounds depend on the margin and on the radii, but not on the ambient dimension or the underlying measure space. These are extensions of the work of Bourneuf, Charbit, and Thomassé (FOCS 2025) who studied the PCC of expanded balls in Euclidean space, that is, $\ell_2^d$. We also prove lower bounds on the VC dimension that match the upper bounds in terms of the margin parameter $δ$. Finally, we derive a Dense Neighborhood Lemma in \(L_p\)-spaces, again extending the known Euclidean results.
Our method relies on the linearization of the distance through a map into a space of non-trivial Rademacher type, and then the use of a balanced signed-sum estimate, or a no-dimensional Radon theorem. The arguments rely on ideas from functional analysis that are clearly explained for the non-expert in that field.
△ Less
Submitted 12 July, 2026;
originally announced July 2026.
-
ProsMAE: Multi-Source MAE Pretraining for ISUP Grade Classification
Authors:
Anna Jung,
Kyeonghun Kim,
Youngung Han,
Eunseob Choi,
Jiwon Yang,
Ken Ying-Kai Liao,
Hyuk-Jae Lee,
Nam-Joon Kim
Abstract:
Whole slide images (WSIs) provide rich diagnostic information for computational pathology, but their gigapixel scale, stain variation, scanner differences, tissue artifacts, and limited expert annotation make robust model training challenging. This paper presents a multi-source Masked Autoencoder (MAE) framework, named ProsMAE, for histopathology representation learning. Tiles from Prostate cANcer…
▽ More
Whole slide images (WSIs) provide rich diagnostic information for computational pathology, but their gigapixel scale, stain variation, scanner differences, tissue artifacts, and limited expert annotation make robust model training challenging. This paper presents a multi-source Masked Autoencoder (MAE) framework, named ProsMAE, for histopathology representation learning. Tiles from Prostate cANcer graDe Assessment (PANDA), CAncer MEtastases in LYmph nOdes challeNge 2017 (CAMELYON17), and BReAst Carcinoma Subtyping (BRACS) are used for ProsMAE pretraining to expose the encoder to diverse tissue morphology and acquisition conditions. The learned encoder is transferred for International Society of Urological Pathology (ISUP) grade classification through ProsCLS, using a frozen encoder and a linear classification head. ProsMAE achieved a higher mean validation quadratic weighted kappa (QWK) than the vanilla MAE frozen linear-probe baseline under the evaluated disjoint PANDA split. Repeated-split evaluation remains necessary to further establish robustness across split compositions.
△ Less
Submitted 9 July, 2026;
originally announced July 2026.
-
Trigger system for the Payload for Ultrahigh Energy Observations (PUEO) balloon-borne neutrino detector
Authors:
Q. Abarr,
J. Alfaro,
P. Allison,
J. Alvarez-Muñiz,
T. Anderson,
H. Barnett,
A. Basharina-Freshville,
J. J. Beatty,
L. Beaufore,
D. Z. Besson,
M. Betts,
R. Bose,
D. Braun,
B. Chamanbahar,
P. Chen,
Y. Chen,
J. M. Clem,
T. Coakley,
A. Connolly,
K. Couberly,
L. Cremonesi,
A. Cummings,
P. Dasgupta,
C. Deaconu,
J. Flaherty
, et al. (45 additional authors not shown)
Abstract:
The Payload for Ultrahigh Energy Observations (PUEO) is a NASA balloon-borne instrument for the detection of ultra-high energy (UHE) neutrinos with energies above $10^{17.5}~\textrm{eV}$ via either the Askaryan effect or geomagnetic emissions from an upward-going air shower. The main instrument trigger system for PUEO is a fully digital supersample rate beamformer based on 24 Xilinx Radio Frequenc…
▽ More
The Payload for Ultrahigh Energy Observations (PUEO) is a NASA balloon-borne instrument for the detection of ultra-high energy (UHE) neutrinos with energies above $10^{17.5}~\textrm{eV}$ via either the Askaryan effect or geomagnetic emissions from an upward-going air shower. The main instrument trigger system for PUEO is a fully digital supersample rate beamformer based on 24 Xilinx Radio Frequency System-on-a-Chip (RFSoC) digitizers sampling 192 channels operating at $3~\textrm{GSa/s}$ and a system clock frequency of $375~\textrm{MHz}$. The trigger implements frequency band conditioning, dynamic radio-frequency interference (RFI) rejection, and matched filtering, with significant emphasis on optimization to reduce both the power and resource usage while maintaining sensitivity. The system implements 48 total synthetic antenna beams with up to 8 antennas each, covering a $\sim25^\circ$ range in zenith and $\sim60^\circ$ range in azimuth. Preflight testing demonstrated a trigger performance of a minimum signal-to-noise ratio (SNR) of $\sim1.5$ using simulated signals while consuming between $5-7~\textrm{W}$ in the trigger logic.
△ Less
Submitted 7 July, 2026; v1 submitted 6 July, 2026;
originally announced July 2026.
-
$α$-TCAV: A Unified Framework for Testing with Concept Activation Vectors
Authors:
Ekkehard Schnoor,
Jawher Said,
Malik Tiomoko,
Wojciech Samek,
Alexander Jung
Abstract:
Concept Activation Vectors (CAVs) are a fundamental tool for concept-based explainability in deep learning, yet their practical utility is limited by statistical instability. We analyze the stochastic nature of CAVs and the Testing with CAVs (TCAV) method, deriving the distributions of major CAV classes including PatternCAV, FastCAV, and ridge regression-based CAVs. We then identify a fundamental…
▽ More
Concept Activation Vectors (CAVs) are a fundamental tool for concept-based explainability in deep learning, yet their practical utility is limited by statistical instability. We analyze the stochastic nature of CAVs and the Testing with CAVs (TCAV) method, deriving the distributions of major CAV classes including PatternCAV, FastCAV, and ridge regression-based CAVs. We then identify a fundamental flaw in the standard TCAV score: its reliance on a discontinuous indicator function induces non-decaying variance in critical regimes. To address this, we introduce $α$-TCAV, a generalized framework that replaces the indicator with a parameterized smooth function, yielding a unified probabilistic formulation that subsumes both TCAV and Multi-TCAV. We characterize the induced distributions of sensitivity scores and different TCAV variants, showing that established state-of-the-art choices lack theoretical justification. We provide principled guidance on tuning the parameter in $α$-TCAV -- either to imitate Multi-TCAV at substantially lower computational cost, or to obtain a calibrated Bayes-optimal probabilistic measure of a concept's influence. Finally, our analysis yields practical recommendations that challenge established routines: most notably, allocating the full sampling budget to a single CAV rather than splitting it across several.
△ Less
Submitted 15 May, 2026;
originally announced May 2026.
-
Piercing all maximum cliques in hypergraphs
Authors:
Andreas Holmsen,
Attila Jung,
Balázs Keszegh,
Dániel G. Simon,
Gábor Tardos
Abstract:
Graphs whose maximum clique size exceeds half of the total number of vertices satisfy a classical property: the family of their maximum sized cliques can be pierced by a single vertex. This result dates back to a 1965 theorem by Hajnal. Motivated by this theorem, Jung, Keszegh, Pálvölgyi, and Yuditsky recently conjectured that an analogous result should hold for hypergraphs of larger uniformity, w…
▽ More
Graphs whose maximum clique size exceeds half of the total number of vertices satisfy a classical property: the family of their maximum sized cliques can be pierced by a single vertex. This result dates back to a 1965 theorem by Hajnal. Motivated by this theorem, Jung, Keszegh, Pálvölgyi, and Yuditsky recently conjectured that an analogous result should hold for hypergraphs of larger uniformity, with an appropriate constant replacing the threshold $1/2$.
In this paper we refute this conjecture in a strong form. We show that for any constant $c<1$ and integers $k\ge 3$ and $t\ge 1$, there exist $k$-uniform hypergraphs $G$ whose maximum clique size exceeds $c|V(G)|$, yet the family of maximum size cliques of $G$ cannot be pierced by $t$ vertices. This demonstrates that no universal constant threshold guarantees bounded piercing number for maximum cliques in uniform hypergraphs.
We discuss further questions concerning the relationship between clique size and piercing maximum cliques in hypergraphs, and introduce a geometric variant of the problem using Helly's Theorem.
△ Less
Submitted 6 August, 2026; v1 submitted 23 April, 2026;
originally announced April 2026.
-
MATHENA: Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy
Authors:
Kyeonghun Kim,
Jaehyung Park,
Youngung Han,
Anna Jung,
Seongbin Park,
Sumin Lee,
Jiwon Yang,
Jiyoon Han,
Subeen Lee,
Junsu Lim,
Hyunsu Go,
Eunseob Choi,
Hyeonseok Jung,
Soo Yong Kim,
Woo Kyoung Jeong,
Won Jae Lee,
Pa Hong,
Hyuk-Jae Lee,
Ken Ying-Kai Liao,
Nam-Joon Kim
Abstract:
Dental diagnosis from Orthopantomograms (OPGs) requires coordination of tooth detection, caries segmentation (CarSeg), anomaly detection (AD), and dental developmental staging (DDS). We propose Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy (MATHENA), a unified framework leveraging Mamba's linear-complexity State Space Models (SSM) to address all…
▽ More
Dental diagnosis from Orthopantomograms (OPGs) requires coordination of tooth detection, caries segmentation (CarSeg), anomaly detection (AD), and dental developmental staging (DDS). We propose Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy (MATHENA), a unified framework leveraging Mamba's linear-complexity State Space Models (SSM) to address all four tasks. MATHENA integrates MATHE, a multi-resolution SSM-driven detector with four-directional Vision State Space (VSS) blocks for O(N) global context modeling, generating per-tooth crops. These crops are processed by HENA, a lightweight Mamba-UNet with a triple-head architecture and Global Context State Token (GCST). In the triple-head architecture, CarSeg is first trained as an upstream task to establish shared representations, which are then frozen and reused for downstream AD fine-tuning and DDS classification via linear probing, enabling stable, efficient learning. We also curate PARTHENON, a benchmark comprising 15,062 annotated instances from ten datasets. MATHENA achieves 93.78% mAP@50 in tooth detection, 90.11% Dice for CarSeg, 88.35% for AD, and 72.40% ACC for DDS.
△ Less
Submitted 1 April, 2026;
originally announced April 2026.
-
Wayfinder: Automated Operating System Specialization
Authors:
Alexander Jung,
Cezar Crăciunoiu,
Nikolaos Karaolidis,
Hugo Lefeuvre,
Daniel Oñoro Rubio,
Felipe Huici,
Charalampos Rotsos,
Pierre Olivier
Abstract:
Specializing an OS to optimize the performance of a particular application is typically a manual process that requires great expertise. Specialization through configuration lends itself well to automation; however, it is challenging due to the sheer size of the configuration space of modern OSes, the difficulty to quantify that space, the long time it takes to evaluate a configuration, and the lar…
▽ More
Specializing an OS to optimize the performance of a particular application is typically a manual process that requires great expertise. Specialization through configuration lends itself well to automation; however, it is challenging due to the sheer size of the configuration space of modern OSes, the difficulty to quantify that space, the long time it takes to evaluate a configuration, and the large number of invalid configurations. Hence, existing attempts at specializing OSes automatically are limited to switching features on and off to minimize memory consumption or attack surface, and cannot target metrics such as performance.
We present Wayfinder, a framework specializing the configuration of OSes completely automatically and without expert knowledge. It can specialize all aspects of an OS configuration (compile-/boot-/run-time) towards any quantifiable performance, resource consumption, or security metric, for an application processing a given workload on a given hardware setup. Wayfinder consists of an automated OS benchmarking platform, and a neural network-based search algorithm driving the specialization process. This is achieved by learning on the fly which configuration parameters and values impact performance the most, and which ones lead to runtime failures. Optionally, a model pre-trained on one application can be reused to accelerate the specialization of related applications. We evaluate Wayfinder on two OSes, four applications, and two target metrics: Wayfinder fully automatically identifies specialized configurations with up to 24% application performance improvement and 8.5% memory usage reduction compared to default configurations. We highlight the benefits of our neural network, reaching good solutions faster than competing approaches (random and Bayesian), and successfully transferring knowledge between related applications.
△ Less
Submitted 24 March, 2026;
originally announced March 2026.
-
Interpretable Multiple Myeloma Prognosis with Observational Medical Outcomes Partnership Data
Authors:
Salma Rachidi,
Aso Bozorgpanah,
Eric Fey,
Alexander Jung
Abstract:
Machine learning (ML) promises better clinical decision-making, yet opaque model behavior limits the adoption in healthcare. We propose two novel regularization techniques for ensuring the interpretability of ML models trained on real-world data. In particular, we consider the prediction of five-year survival for multiple myeloma patients using clinical data from Helsinki University Hospital. To e…
▽ More
Machine learning (ML) promises better clinical decision-making, yet opaque model behavior limits the adoption in healthcare. We propose two novel regularization techniques for ensuring the interpretability of ML models trained on real-world data. In particular, we consider the prediction of five-year survival for multiple myeloma patients using clinical data from Helsinki University Hospital. To ensure the interpretability of the trained models, we use two alternative constructions for a penalty term used for regularization. The first one penalizes deviations from the predictions obtained from an interpretable logistic regression method with two manually chosen features. The second construction requires consistency of model predictions with the revised international staging system (R-ISS). We verify the usefulness of the proposed regularization techniques in numerical experiments using data from 812 patients. They achieve an accuracy up to 0.721 on a test set and SHAP values show that the models rely on the selected important features.
△ Less
Submitted 20 March, 2026;
originally announced March 2026.
-
Shapes are not enough: CONSERVAttack and its use for finding vulnerabilities and uncertainties in machine learning applications
Authors:
Philip Bechtle,
Lucie Flek,
Philipp Alexander Jung,
Akbar Karimi,
Timo Saala,
Alexander Schmidt,
Matthias Schott,
Philipp Soldin,
Christopher Wiebusch,
Ulrich Willemsen
Abstract:
In High Energy Physics, as in many other fields of science, the application of machine learning techniques has been crucial in advancing our understanding of fundamental phenomena. Increasingly, deep learning models are applied to analyze both simulated and experimental data. In most experiments, a rigorous regime of testing for physically motivated systematic uncertainties is in place. The numeri…
▽ More
In High Energy Physics, as in many other fields of science, the application of machine learning techniques has been crucial in advancing our understanding of fundamental phenomena. Increasingly, deep learning models are applied to analyze both simulated and experimental data. In most experiments, a rigorous regime of testing for physically motivated systematic uncertainties is in place. The numerical evaluation of these tests for differences between the data on the one side and simulations on the other side quantifies the effect of potential sources of mismodelling on the machine learning output. In addition, thorough comparisons of marginal distributions and (linear) feature correlations between data and simulation in "control regions" are applied. However, the guidance by physical motivation, and the need to constrain comparisons to specific regions, does not guarantee that all possible sources of deviations have been accounted for. We therefore propose a new adversarial attack - the CONSERVAttack - designed to exploit the remaining space of hypothetical deviations between simulation and data after the above mentioned tests. The resulting adversarial perturbations are consistent within the uncertainty bounds - evading standard validation checks - while successfully fooling the underlying model. We further propose strategies to mitigate such vulnerabilities and argue that robustness to adversarial effects must be considered when interpreting results from deep learning in particle physics.
△ Less
Submitted 8 April, 2026; v1 submitted 14 March, 2026;
originally announced March 2026.
-
Machine Learning on Heterogeneous, Edge, and Quantum Hardware for Particle Physics (ML-HEQUPP)
Authors:
Julia Gonski,
Jenni Ott,
Shiva Abbaszadeh,
Sagar Addepalli,
Matteo Cremonesi,
Jennet Dickinson,
Giuseppe Di Guglielmo,
Erdem Yigit Ertorer,
Lindsey Gray,
Ryan Herbst,
Christian Herwig,
Tae Min Hong,
Benedikt Maier,
Maryam Bayat Makou,
David Miller,
Mark S. Neubauer,
Cristián Peña,
Dylan Rankin,
Seon-Hee,
Seo,
Giordon Stark,
Alexander Tapper,
Audrey Corbeil Therrien,
Ioannis Xiotidis,
Keisuke Yoshihara
, et al. (99 additional authors not shown)
Abstract:
The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for scientific discovery demands real-time inference and decision-making, intelligent data reduction, and efficient processing architectures beyond current capabilitie…
▽ More
The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for scientific discovery demands real-time inference and decision-making, intelligent data reduction, and efficient processing architectures beyond current capabilities. Crucial to the success of this experimental paradigm are several emerging technologies, such as artificial intelligence and machine learning (AI/ML), silicon microelectronics, and the advent of quantum algorithms and processing. Their intersection includes areas of research such as low-power and low-latency devices for edge computing, heterogeneous accelerator systems, reconfigurable hardware, novel codesign and synthesis strategies, readout for cryogenic or high-radiation environments, and analog computing. This white paper presents a community-driven vision to identify and prioritize research and development opportunities in hardware-based ML systems and corresponding physics applications, contributing towards a successful transition to the new data frontier of fundamental science.
△ Less
Submitted 24 July, 2026; v1 submitted 24 February, 2026;
originally announced February 2026.
-
Building an AI-native Research Ecosystem for Experimental Particle Physics: A Community Vision
Authors:
Thea Klaeboe Aarrestad,
Alaa Abdelhamid,
Haider Abidi,
Jahred Adelman,
Jennifer Adelman-McCarthy,
Shuchin Aeron,
Garvita Agarwal,
Usman Ali,
Cristiano Alpigiani,
Omar Alterkait,
Mohamed Aly,
Oz Amram,
Saeed Ansari Fard,
Aram Apyan,
John Arrington,
Marvin Ascencio-Sosa,
Mohammad Atif,
Aneesha Avasthi,
Muhammad Bilal Azam,
Bhim Bam,
Joshua Barrow,
Rainer Bartoldus,
Amit Bashyal,
Aashwin Basnet,
Ayse Bat
, et al. (435 additional authors not shown)
Abstract:
Experimental particle physics seeks to understand the universe by probing its fundamental particles and forces and exploring how they govern the large-scale processes that shape cosmic evolution. This whitepaper presents a vision for how Artificial Intelligence (AI) can accelerate discovery in this field. We outline grand challenges that must be addressed to enable transformative breakthroughs and…
▽ More
Experimental particle physics seeks to understand the universe by probing its fundamental particles and forces and exploring how they govern the large-scale processes that shape cosmic evolution. This whitepaper presents a vision for how Artificial Intelligence (AI) can accelerate discovery in this field. We outline grand challenges that must be addressed to enable transformative breakthroughs and describe how current and planned experimental facilities can implement this vision to advance our understanding of the vast and complex physical world from the smallest to the largest scales. We show how facilities currently under construction, such as the HL-LHC, DUNE and soon EIC, can both benefit from and serve as proving grounds for this vision, while also enabling a longer-term goal for how future experiments -- like FCC-ee at CERN, IceCube-Gen2, a Muon Collider in the U.S., and smaller to mid-scale projects -- can be fully AI-native. We describe how a truly national-scale collaboration, jointly managed across large funding partners, and involving both DOE laboratories and universities, can make this happen.
△ Less
Submitted 19 February, 2026;
originally announced February 2026.
-
Nonparametric Distribution Regression Re-calibration
Authors:
Ádám Jung,
Domokos M. Kelen,
András A. Benczúr
Abstract:
A key challenge in probabilistic regression is ensuring that predictive distributions accurately reflect true empirical uncertainty. Minimizing overall prediction error often encourages models to prioritize informativeness over calibration, producing narrow but overconfident predictions. However, in safety-critical settings, trustworthy uncertainty estimates are often more valuable than narrow int…
▽ More
A key challenge in probabilistic regression is ensuring that predictive distributions accurately reflect true empirical uncertainty. Minimizing overall prediction error often encourages models to prioritize informativeness over calibration, producing narrow but overconfident predictions. However, in safety-critical settings, trustworthy uncertainty estimates are often more valuable than narrow intervals. Realizing the problem, several recent works have focused on post-hoc corrections; however, existing methods either rely on weak notions of calibration (such as PIT uniformity) or impose restrictive parametric assumptions on the nature of the error. To address these limitations, we propose a novel nonparametric re-calibration algorithm based on conditional kernel mean embeddings, capable of correcting calibration error without restrictive modeling assumptions. For efficient inference with real-valued targets, we introduce a novel characteristic kernel over distributions that can be evaluated in $\mathcal{O}(n \log n)$ time for empirical distributions of size $n$. We demonstrate that our method consistently outperforms prior re-calibration approaches across a diverse set of regression benchmarks and model classes.
△ Less
Submitted 2 August, 2026; v1 submitted 13 February, 2026;
originally announced February 2026.
-
Anisotropic magnon transport in an antiferromagnetic trilayer heterostructure: is BiFeO$_3$ an altermagnet?
Authors:
Sajid Husain,
Maya Ramesh,
Qian Song,
Sergei Prokhorenko,
Shashank Kumar Ojha,
Surya Narayan Panda,
Xinyan Li,
Yousra Nahas,
Yogesh Kumar,
Pushpendra Gupta,
Tenzin Chang,
Alan Ji-in Jung,
Rogério de Sousa,
James G. Analytis,
Lane W. Martin,
Zhi Yao,
Sang-Wook Cheong,
Laurent Bellaiche,
Manuel Bibes,
Darrell G. Schlom,
Ramamoorthy Ramesh
Abstract:
Magnons provide a route to ultra-fast transport and non-destructive readout of spin-based information transfer. Here, we report magnon transport and its emergent anisotropic nature in BiFeO$_3$ layers confined between ultrathin layers of the antiferromagnet LaFeO$_3$. Due to the confined state, BiFeO$_3$ serves as an efficient magnon transmission channel as well as a magnetoelectric knob by which…
▽ More
Magnons provide a route to ultra-fast transport and non-destructive readout of spin-based information transfer. Here, we report magnon transport and its emergent anisotropic nature in BiFeO$_3$ layers confined between ultrathin layers of the antiferromagnet LaFeO$_3$. Due to the confined state, BiFeO$_3$ serves as an efficient magnon transmission channel as well as a magnetoelectric knob by which to control the stack by means of an electric field. We discuss the mechanism of the anisotropic spin transport based on the interaction between the antiferromagnetic order and the electric field. This allows us to manipulate and amplify the spin transport in such a confined geometry. Furthermore, lower crystal symmetric and suppression of the spin cycloid in ultrathin BiFeO$_3$ stabilizes a non-trivial antiferromagnetic state exhibiting symmetry-protected spin-split bands that provide the non-trivial sign inversion of the spin current, which is a characteristic of an altermagnet. This work provides an understanding of the anisotropic spin transport in complex antiferromagnetic heterostructures where ferroelectricity and altermagnetism coexist, paving the way for a new route to realize electric-field control of a novel state of magnetism.
△ Less
Submitted 7 January, 2026;
originally announced January 2026.
-
Evaluation of Grid-based Uncertainty Propagation for Collaborative Self-Calibration in Indoor Positioning Systems
Authors:
Paul Schwarzbach,
Andrea Jung
Abstract:
Radio-based localization systems conventionally require stationary reference points (e.g. anchors) with precisely surveyed positions, making deployment time-consuming and costly. This paper presents an empirical evaluation of collaborative self-calibration for Ultra-Wideband (UWB) networks, extending a discrete Bayesian approach based on grid-based uncertainty propagation. The enhanced algorithm r…
▽ More
Radio-based localization systems conventionally require stationary reference points (e.g. anchors) with precisely surveyed positions, making deployment time-consuming and costly. This paper presents an empirical evaluation of collaborative self-calibration for Ultra-Wideband (UWB) networks, extending a discrete Bayesian approach based on grid-based uncertainty propagation. The enhanced algorithm reduces measurement availability requirements while maintaining positioning accuracy through probabilistic state estimation. We validate the approach using real-world data from controlled indoor UWB network experiments with 12 nodes in a static environment. Experimental evaluation demonstrates 0.28~m mean ranging error under line-of-sight conditions and 1.11~m overall ranging error across mixed propagation scenarios, achieving sub-meter positioning accuracy. Results demonstrate the algorithm's robustness to measurement noise and partial connectivity scenarios typical in industrial deployments. The findings contribute to automated UWB network initialization for indoor positioning applications, reducing infrastructure dependency compared to manual anchor calibration procedures.
△ Less
Submitted 20 April, 2026; v1 submitted 13 November, 2025;
originally announced November 2025.
-
MiniFool -- Physics-Constraint-Aware Minimizer-Based Adversarial Attacks in Deep Neural Networks
Authors:
Lucie Flek,
Oliver Janik,
Philipp Alexander Jung,
Akbar Karimi,
Timo Saala,
Alexander Schmidt,
Matthias Schott,
Philipp Soldin,
Matthias Thiesmeyer,
Christopher Wiebusch,
Ulrich Willemsen
Abstract:
In this paper, we present a new algorithm, MiniFool, that implements physics-inspired adversarial attacks for testing neural network-based classification tasks in particle and astroparticle physics. While we initially developed the algorithm for the search for astrophysical tau neutrinos with the IceCube Neutrino Observatory, we apply it to further data from other science domains, thus demonstrati…
▽ More
In this paper, we present a new algorithm, MiniFool, that implements physics-inspired adversarial attacks for testing neural network-based classification tasks in particle and astroparticle physics. While we initially developed the algorithm for the search for astrophysical tau neutrinos with the IceCube Neutrino Observatory, we apply it to further data from other science domains, thus demonstrating its general applicability. Here, we apply the algorithm to the well-known MNIST data set and furthermore, to Open Data data from the CMS experiment at the Large Hadron Collider. The algorithm is based on minimizing a cost function that combines a $χ^2$ based test-statistic with the deviation from the desired target score. The test statistic quantifies the probability of the perturbations applied to the data based on the experimental uncertainties. For our studied use cases, we find that the likelihood of a flipped classification differs for both the initially correctly and incorrectly classified events. When testing changes of the classifications as a function of an attack parameter that scales the experimental uncertainties, the robustness of the network decision can be quantified. Furthermore, this allows testing the robustness of the classification of unlabeled experimental data.
△ Less
Submitted 16 June, 2026; v1 submitted 3 November, 2025;
originally announced November 2025.
-
Federated k-Means over Networks
Authors:
Xu Yang,
Salvatore Rastelli,
Alexander Jung
Abstract:
We study federated clustering, where interconnected devices collaboratively cluster the data points of private local datasets. Focusing on hard clustering via the k-means principle, we formulate federated k-means as an instance of generalized total variation minimization (GTVMin). This leads to a federated k-means algorithm in which each device updates its local cluster centroids by solving a regu…
▽ More
We study federated clustering, where interconnected devices collaboratively cluster the data points of private local datasets. Focusing on hard clustering via the k-means principle, we formulate federated k-means as an instance of generalized total variation minimization (GTVMin). This leads to a federated k-means algorithm in which each device updates its local cluster centroids by solving a regularized k-means problem with a regularizer that enforces consistency between neighbouring devices. The resulting algorithm is privacy-friendly, as only aggregated information is exchanged.
△ Less
Submitted 28 January, 2026; v1 submitted 10 October, 2025;
originally announced October 2025.
-
ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Models
Authors:
Youngeun Kim,
Youjia Zhang,
Huiling Liu,
Aecheon Jung,
Sunwoo Lee,
Sungeun Hong
Abstract:
Large Vision-Language Models (VLMs) enable strong multimodal reasoning but incur heavy inference costs from redundant visual tokens. Token pruning alleviates this issue, yet existing approaches face limitations. Attention-based methods rely on raw attention scores, which are often unstable across layers and heads and can lead to redundant selections. Diversity-based methods improve robustness by s…
▽ More
Large Vision-Language Models (VLMs) enable strong multimodal reasoning but incur heavy inference costs from redundant visual tokens. Token pruning alleviates this issue, yet existing approaches face limitations. Attention-based methods rely on raw attention scores, which are often unstable across layers and heads and can lead to redundant selections. Diversity-based methods improve robustness by selecting tokens far apart in feature space, but risk dropping regions needed for accurate prediction. We propose ZOO-Prune, a training-free framework built on the intuition that highly sensitive tokens have a stronger influence on the model's output and capture complementary visual cues rather than redundant ones. To achieve this, we estimate token sensitivity using zeroth-order perturbations at the lightweight projection layer. This measures how small random perturbations affect the projected features and enables efficient approximation of each token's influence without backpropagation. Extensive experiments across multiple VLMs and benchmarks show that ZOO-Prune consistently outperforms prior methods while pruning up to 94.4% of tokens without sacrificing accuracy. Our method also improves efficiency, reaching up to 2.30x faster end-to-end inference compared to the baseline.
△ Less
Submitted 20 March, 2026; v1 submitted 29 September, 2025;
originally announced September 2025.
-
Concept activation vectors: a unifying view and adversarial attacks
Authors:
Ekkehard Schnoor,
Malik Tiomoko,
Jawher Said,
Alex Jung,
Wojciech Samek
Abstract:
Concept Activation Vectors (CAVs) are a tool from explainable AI, offering a promising approach for understanding how human-understandable concepts are encoded in a model's latent spaces. They are computed from hidden-layer activations of inputs belonging either to a concept class or to non-concept examples. Adopting a probabilistic perspective, the distribution of the (non-)concept inputs induces…
▽ More
Concept Activation Vectors (CAVs) are a tool from explainable AI, offering a promising approach for understanding how human-understandable concepts are encoded in a model's latent spaces. They are computed from hidden-layer activations of inputs belonging either to a concept class or to non-concept examples. Adopting a probabilistic perspective, the distribution of the (non-)concept inputs induces a distribution over the CAV, making it a random vector in the latent space. This enables us to derive mean and covariance for different types of CAVs, leading to a unified theoretical view. This probabilistic perspective also reveals a potential vulnerability: CAVs can strongly depend on the rather arbitrary non-concept distribution, a factor largely overlooked in prior work. We illustrate this with a simple yet effective adversarial attack, underscoring the need for a more systematic study.
△ Less
Submitted 27 January, 2026; v1 submitted 26 September, 2025;
originally announced September 2025.
-
Graph-Regularized Learning of Gaussian Mixture Models
Authors:
Shamsiiat Abdurakhmanova,
Alex Jung
Abstract:
We present a graph-regularized learning of Gaussian Mixture Models (GMMs) in distributed settings with heterogeneous and limited local data. The method exploits a provided similarity graph to guide parameter sharing among nodes, avoiding the transfer of raw data. The resulting model allows for flexible aggregation of neighbors' parameters and outperforms both centralized and locally trained GMMs i…
▽ More
We present a graph-regularized learning of Gaussian Mixture Models (GMMs) in distributed settings with heterogeneous and limited local data. The method exploits a provided similarity graph to guide parameter sharing among nodes, avoiding the transfer of raw data. The resulting model allows for flexible aggregation of neighbors' parameters and outperforms both centralized and locally trained GMMs in heterogeneous, low-sample regimes.
△ Less
Submitted 17 September, 2025;
originally announced September 2025.
-
Dynamic Rank Adjustment for Accurate and Efficient Neural Network Training
Authors:
Hyuntak Shin,
Aecheon Jung,
Sungeun Hong,
Sunwoo Lee
Abstract:
Low-rank training methods reduce the number of trainable parameters by re-parameterizing the weights with matrix decompositions (e.g., singular value decomposition). However, enforcing a fixed low-rank structure caps the rank of the weight matrices and can hinder the model's ability to learn complex patterns. Furthermore, the effective rank of the model's weights tends to decline during training,…
▽ More
Low-rank training methods reduce the number of trainable parameters by re-parameterizing the weights with matrix decompositions (e.g., singular value decomposition). However, enforcing a fixed low-rank structure caps the rank of the weight matrices and can hinder the model's ability to learn complex patterns. Furthermore, the effective rank of the model's weights tends to decline during training, and this drop is accelerated when the model is reparameterized into a low-rank structure. In this study, we argue that strategically interleaving full-rank training epochs within low-rank training epochs can effectively restore the rank of the model's weights. Based on our findings, we propose a general dynamic-rank training framework that is readily applicable to a wide range of neural-network tasks. We first describe how to adjust the rank of weight matrix to alleviate the inevitable rank collapse that arises during training, and then present extensive empirical results that validate our claims and demonstrate the efficacy of the proposed framework. Our empirical study shows that the proposed method achieves almost the same computational cost as SVD-based low-rank training while achieving a comparable accuracy to full-rank training across various benchmarks.
△ Less
Submitted 14 October, 2025; v1 submitted 12 August, 2025;
originally announced August 2025.
-
On the symmetry behind duality
Authors:
Marco Abbadini,
Achim Jung
Abstract:
Dualities such as Stone duality and the duality between sober spaces and spatial frames hinge on an interaction between open sets and compact saturated sets. In several important classes of spaces-Stone spaces, spectral spaces, and stably compact spaces-this interaction forms a perfect symmetry, reflected dually as order self-duality. But the class of sober spaces, despite being central to Stone-l…
▽ More
Dualities such as Stone duality and the duality between sober spaces and spatial frames hinge on an interaction between open sets and compact saturated sets. In several important classes of spaces-Stone spaces, spectral spaces, and stably compact spaces-this interaction forms a perfect symmetry, reflected dually as order self-duality. But the class of sober spaces, despite being central to Stone-like dualities, exhibits only a partial symmetry between openness and compactness.
This raises a central question: can we enlarge the setting enough to recover a perfect symmetry, while still retaining sober spaces and preserving the conditions that make the sober-spatial-frame duality work?
We answer this question affirmatively. We introduce ko-spaces, whose families of open and compact saturated sets satisfy the compatibility needed for duality, and bi-dcpos, a pointfree companion generalizing both spatial frames and continuous domains. We prove that the categories of ko-spaces and distributive bi-dcpos are equivalent (and dually equivalent, too), and that each category carries a symmetry in the form of a self-duality. On spaces, this extends de Groot duality; on domains, it extends Lawson duality.
Classical results fall out as special cases: the sober-spatial-frame duality reappears inside our symmetric framework, and continuous domains acquire a presentation akin to that of d-frames. Our work suggests that an appropriate home for Stone-like duality is a fully symmetric two-sorted world in which openness and compactness play on equal footing.
△ Less
Submitted 9 July, 2026; v1 submitted 24 July, 2025;
originally announced July 2025.
-
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
Authors:
Mika Markus Müller,
Konstantin Lübeck,
Alexander Louis-Ferdinand Jung,
Jannik Steinmetz,
Oliver Bringmann
Abstract:
Artificial Intelligence (AI) algorithms, such as Deep Neural Networks (DNNs), have become an important tool for a wide range of applications, from computer vision to natural language processing. However, the computational complexity of DNN inference poses a significant challenge, particularly for processing on resource-constrained edge devices. One promising approach to address this challenge is t…
▽ More
Artificial Intelligence (AI) algorithms, such as Deep Neural Networks (DNNs), have become an important tool for a wide range of applications, from computer vision to natural language processing. However, the computational complexity of DNN inference poses a significant challenge, particularly for processing on resource-constrained edge devices. One promising approach to address this challenge is the exploitation of sparsity in DNN operator weights.
In this work, we present FlexiSAGA, an architecturally configurable and dataflow-flexible AI hardware accelerator for the sparse and dense processing of general matrix multiplications (GEMMs). FlexiSAGA supports seven different sparse and dense dataflows, enabling efficient processing of resource intensive DNN operators. Additionally, we propose a DNN pruning method specifically tailored towards the FlexiSAGA architecture, allowing for near-optimal processing of dense and sparse convolution and fully-connected operators, facilitating a DNN/HW co-design flow. Our results show a whole DNN sparse-over-dense inference speedup ranging from 1.41 up to 4.28, outperforming commercial and literature-reported accelerator platforms.
△ Less
Submitted 2 June, 2025;
originally announced June 2025.
-
Federated Learning: From Theory to Practice
Authors:
A. Jung
Abstract:
This book offers a hands-on introduction to building and understanding federated learning (FL) systems. FL enables multiple devices -- such as smartphones, sensors, or local computers -- to collaboratively train machine learning (ML) models, while keeping their data private and local. It is a powerful solution when data cannot or should not be centralized due to privacy, regulatory, or technical r…
▽ More
This book offers a hands-on introduction to building and understanding federated learning (FL) systems. FL enables multiple devices -- such as smartphones, sensors, or local computers -- to collaboratively train machine learning (ML) models, while keeping their data private and local. It is a powerful solution when data cannot or should not be centralized due to privacy, regulatory, or technical reasons. The book is designed for students, engineers, and researchers who want to learn how to design scalable, privacy preserving FL systems. Our main focus is on personalization: enabling each device to train its own model while still benefiting from collaboration with relevant devices. This is achieved by leveraging similarities between (the learning tasks associated with) devices that are encoded by the weighted edges (or links) of a federated learning network (FL network). The key idea is to represent real-world FL systems as networks of devices, where nodes correspond to device and edges represent communication links and data similarities between them. The training of personalized models for these devices can be naturally framed as a distributed optimization problem. This optimization problem is referred to as generalized total variation minimization (GTVMin) and ensures that devices with similar learning tasks learn similar model parameters. Our approach is both mathematically principled and practically motivated. While we introduce some advanced ideas from optimization theory and graph-based learning, we aim to keep the book accessible. Readers are guided through the core ideas step by step, with intuitive explanations.
△ Less
Submitted 10 June, 2025; v1 submitted 25 May, 2025;
originally announced May 2025.
-
Future Circular Collider Feasibility Study Report: Volume 2, Accelerators, Technical Infrastructure and Safety
Authors:
M. Benedikt,
F. Zimmermann,
B. Auchmann,
W. Bartmann,
J. P. Burnet,
C. Carli,
A. Chancé,
P. Craievich,
M. Giovannozzi,
C. Grojean,
J. Gutleber,
K. Hanke,
A. Henriques,
P. Janot,
C. Lourenço,
M. Mangano,
T. Otto,
J. Poole,
S. Rajagopalan,
T. Raubenheimer,
E. Todesco,
L. Ulrici,
T. Watson,
G. Wilkinson,
A. Abada
, et al. (1439 additional authors not shown)
Abstract:
In response to the 2020 Update of the European Strategy for Particle Physics, the Future Circular Collider (FCC) Feasibility Study was launched as an international collaboration hosted by CERN. This report describes the FCC integrated programme, which consists of two stages: an electron-positron collider (FCC-ee) in the first phase, serving as a high-luminosity Higgs, top, and electroweak factory;…
▽ More
In response to the 2020 Update of the European Strategy for Particle Physics, the Future Circular Collider (FCC) Feasibility Study was launched as an international collaboration hosted by CERN. This report describes the FCC integrated programme, which consists of two stages: an electron-positron collider (FCC-ee) in the first phase, serving as a high-luminosity Higgs, top, and electroweak factory; followed by a proton-proton collider (FCC-hh) at the energy frontier in the second phase.
FCC-ee is designed to operate at four key centre-of-mass energies: the Z pole, the WW production threshold, the ZH production peak, and the top/anti-top production threshold - delivering the highest possible luminosities to four experiments. Over 15 years of operation, FCC-ee will produce more than 6 trillion Z bosons, 200 million WW pairs, nearly 3 million Higgs bosons, and 2 million top anti-top pairs. Precise energy calibration at the Z pole and WW threshold will be achieved through frequent resonant depolarisation of pilot bunches. The sequence of operation modes remains flexible.
FCC-hh will operate at a centre-of-mass energy of approximately 85 TeV - nearly an order of magnitude higher than the LHC - and is designed to deliver 5 to 10 times the integrated luminosity of the HL-LHC. Its mass reach for direct discovery extends to several tens of TeV. In addition to proton-proton collisions, FCC-hh is capable of supporting ion-ion, ion-proton, and lepton-hadron collision modes.
This second volume of the Feasibility Study Report presents the complete design of the FCC-ee collider, its operation and staging strategy, the full-energy booster and injector complex, required accelerator technologies, safety concepts, and technical infrastructure. It also includes the design of the FCC-hh hadron collider, development of high-field magnets, hadron injector options, and key technical systems for FCC-hh.
△ Less
Submitted 25 April, 2025;
originally announced May 2025.
-
Future Circular Collider Feasibility Study Report: Volume 3, Civil Engineering, Implementation and Sustainability
Authors:
M. Benedikt,
F. Zimmermann,
B. Auchmann,
W. Bartmann,
J. P. Burnet,
C. Carli,
A. Chancé,
P. Craievich,
M. Giovannozzi,
C. Grojean,
J. Gutleber,
K. Hanke,
A. Henriques,
P. Janot,
C. Lourenço,
M. Mangano,
T. Otto,
J. Poole,
S. Rajagopalan,
T. Raubenheimer,
E. Todesco,
L. Ulrici,
T. Watson,
G. Wilkinson,
P. Azzi
, et al. (1439 additional authors not shown)
Abstract:
Volume 3 of the FCC Feasibility Report presents studies related to civil engineering, the development of a project implementation scenario, and environmental and sustainability aspects. The report details the iterative improvements made to the civil engineering concepts since 2018, taking into account subsurface conditions, accelerator and experiment requirements, and territorial considerations. I…
▽ More
Volume 3 of the FCC Feasibility Report presents studies related to civil engineering, the development of a project implementation scenario, and environmental and sustainability aspects. The report details the iterative improvements made to the civil engineering concepts since 2018, taking into account subsurface conditions, accelerator and experiment requirements, and territorial considerations. It outlines a technically feasible and economically viable civil engineering configuration that serves as the baseline for detailed subsurface investigations, construction design, cost estimation, and project implementation planning. Additionally, the report highlights ongoing subsurface investigations in key areas to support the development of an improved 3D subsurface model of the region.
The report describes development of the project scenario based on the 'avoid-reduce-compensate' iterative optimisation approach. The reference scenario balances optimal physics performance with territorial compatibility, implementation risks, and costs. Environmental field investigations covering almost 600 hectares of terrain - including numerous urban, economic, social, and technical aspects - confirmed the project's technical feasibility and contributed to the preparation of essential input documents for the formal project authorisation phase. The summary also highlights the initiation of public dialogue as part of the authorisation process. The results of a comprehensive socio-economic impact assessment, which included significant environmental effects, are presented. Even under the most conservative and stringent conditions, a positive benefit-cost ratio for the FCC-ee is obtained. Finally, the report provides a concise summary of the studies conducted to document the current state of the environment.
△ Less
Submitted 25 April, 2025;
originally announced May 2025.
-
Future Circular Collider Feasibility Study Report: Volume 1, Physics, Experiments, Detectors
Authors:
M. Benedikt,
F. Zimmermann,
B. Auchmann,
W. Bartmann,
J. P. Burnet,
C. Carli,
A. Chancé,
P. Craievich,
M. Giovannozzi,
C. Grojean,
J. Gutleber,
K. Hanke,
A. Henriques,
P. Janot,
C. Lourenço,
M. Mangano,
T. Otto,
J. Poole,
S. Rajagopalan,
T. Raubenheimer,
E. Todesco,
L. Ulrici,
T. Watson,
G. Wilkinson,
P. Azzi
, et al. (1439 additional authors not shown)
Abstract:
Volume 1 of the FCC Feasibility Report presents an overview of the physics case, experimental programme, and detector concepts for the Future Circular Collider (FCC). This volume outlines how FCC would address some of the most profound open questions in particle physics, from precision studies of the Higgs and EW bosons and of the top quark, to the exploration of physics beyond the Standard Model.…
▽ More
Volume 1 of the FCC Feasibility Report presents an overview of the physics case, experimental programme, and detector concepts for the Future Circular Collider (FCC). This volume outlines how FCC would address some of the most profound open questions in particle physics, from precision studies of the Higgs and EW bosons and of the top quark, to the exploration of physics beyond the Standard Model. The report reviews the experimental opportunities offered by the staged implementation of FCC, beginning with an electron-positron collider (FCC-ee), operating at several centre-of-mass energies, followed by a hadron collider (FCC-hh). Benchmark examples are given of the expected physics performance, in terms of precision and sensitivity to new phenomena, of each collider stage. Detector requirements and conceptual designs for FCC-ee experiments are discussed, as are the specific demands that the physics programme imposes on the accelerator in the domains of the calibration of the collision energy, and the interface region between the accelerator and the detector. The report also highlights advances in detector, software and computing technologies, as well as the theoretical tools /reconstruction techniques that will enable the precision measurements and discovery potential of the FCC experimental programme. This volume reflects the outcome of a global collaborative effort involving hundreds of scientists and institutions, aided by a dedicated community-building coordination, and provides a targeted assessment of the scientific opportunities and experimental foundations of the FCC programme.
△ Less
Submitted 25 April, 2025;
originally announced May 2025.
-
Quantum Information meets High-Energy Physics: Input to the update of the European Strategy for Particle Physics
Authors:
Yoav Afik,
Federica Fabbri,
Matthew Low,
Luca Marzola,
Juan Antonio Aguilar-Saavedra,
Mohammad Mahdi Altakach,
Nedaa Alexandra Asbah,
Yang Bai,
Hannah Banks,
Alan J. Barr,
Alexander Bernal,
Thomas E. Browder,
Paweł Caban,
J. Alberto Casas,
Kun Cheng,
Frédéric Déliot,
Regina Demina,
Antonio Di Domenico,
Michał Eckstein,
Marco Fabbrichesi,
Benjamin Fuks,
Emidio Gabrielli,
Dorival Gonçalves,
Radosław Grabarczyk,
Michele Grossi
, et al. (46 additional authors not shown)
Abstract:
Some of the most astonishing and prominent properties of Quantum Mechanics, such as entanglement and Bell nonlocality, have only been studied extensively in dedicated low-energy laboratory setups. The feasibility of these studies in the high-energy regime explored by particle colliders was only recently shown and has gathered the attention of the scientific community. For the range of particles an…
▽ More
Some of the most astonishing and prominent properties of Quantum Mechanics, such as entanglement and Bell nonlocality, have only been studied extensively in dedicated low-energy laboratory setups. The feasibility of these studies in the high-energy regime explored by particle colliders was only recently shown and has gathered the attention of the scientific community. For the range of particles and fundamental interactions involved, particle colliders provide a novel environment where quantum information theory can be probed, with energies exceeding by about 12 orders of magnitude those employed in dedicated laboratory setups. Furthermore, collider detectors have inherent advantages in performing certain quantum information measurements, and allow for the reconstruction of the state of the system under consideration via quantum state tomography. Here, we elaborate on the potential, challenges, and goals of this innovative and rapidly evolving line of research and discuss its expected impact on both quantum information theory and high-energy physics.
△ Less
Submitted 8 October, 2025; v1 submitted 31 March, 2025;
originally announced April 2025.
-
Helly-type theorems for monotone properties of boxes
Authors:
Nóra Frankl,
Attila Jung
Abstract:
We present a unified approach to prove Helly-type theorems for monotone properties of boxes, such as having large volume or containing points from a given set. As a corollary, we obtain new proofs for several earlier results regarding specific monotone properties. Our results generalise to $H$-convex sets as well.
We present a unified approach to prove Helly-type theorems for monotone properties of boxes, such as having large volume or containing points from a given set. As a corollary, we obtain new proofs for several earlier results regarding specific monotone properties. Our results generalise to $H$-convex sets as well.
△ Less
Submitted 28 March, 2025;
originally announced March 2025.
-
Task Vector Quantization for Memory-Efficient Model Merging
Authors:
Youngeun Kim,
Seunghwan Lee,
Aecheon Jung,
Bogon Ryu,
Sungeun Hong
Abstract:
Model merging enables efficient multi-task models by combining task-specific fine-tuned checkpoints. However, storing multiple task-specific checkpoints requires significant memory, limiting scalability and restricting model merging to larger models and diverse tasks. In this paper, we propose quantizing task vectors (i.e., the difference between pre-trained and fine-tuned checkpoints) instead of…
▽ More
Model merging enables efficient multi-task models by combining task-specific fine-tuned checkpoints. However, storing multiple task-specific checkpoints requires significant memory, limiting scalability and restricting model merging to larger models and diverse tasks. In this paper, we propose quantizing task vectors (i.e., the difference between pre-trained and fine-tuned checkpoints) instead of quantizing fine-tuned checkpoints. We observe that task vectors exhibit a narrow weight range, enabling low precision quantization (e.g., 4 bit) within existing task vector merging frameworks. To further mitigate quantization errors within ultra-low bit precision (e.g., 2 bit), we introduce Residual Task Vector Quantization, which decomposes the task vector into a base vector and offset component. We allocate bits based on quantization sensitivity, ensuring precision while minimizing error within a memory budget. Experiments on image classification and dense prediction show our method maintains or improves model merging performance while using only 8% of the memory required for full-precision checkpoints.
△ Less
Submitted 7 August, 2025; v1 submitted 10 March, 2025;
originally announced March 2025.
-
The MACIV multiscale seismic experiments in the French Massif Central (2023-2027): deployment, data quality and availability
Authors:
Coralie Aubert,
Guilhem Scheiblin,
Anne Paul,
Hélène Pauchet,
Aurélien Mordret,
Vincent Baudot,
Sébastien Chevrot,
Nicolas Cluzel,
Isabelle Douste-Bacqué,
Franck Grimaud,
Axel Jung,
Stéphane Mercier,
Piel Pawlowski,
Sandrine Roussel,
Thierry Souriot,
Nikolai M. Shapiro,
Matthieu Sylvander,
Benjamin Vial,
David Wolyniec
Abstract:
In the framework of the MACIV project, a consortium of French laboratories has deployed a temporary seismic network of 100 broadband stations in the French Massif Central (FMC) for 3-4 years (2023-2027). The project aims at imaging the crust and upper mantle of the FMC to better assess the sources of volcanism, and the impacts of the Variscan inheritance or the Cenozoic rift system on volcanic sys…
▽ More
In the framework of the MACIV project, a consortium of French laboratories has deployed a temporary seismic network of 100 broadband stations in the French Massif Central (FMC) for 3-4 years (2023-2027). The project aims at imaging the crust and upper mantle of the FMC to better assess the sources of volcanism, and the impacts of the Variscan inheritance or the Cenozoic rift system on volcanic systems. A large-scale array of 35 broadband stations covers the entire FMC and complements the permanent networks to reach a homogeneous coverage with ~35 km spacing. This network, with XP code, is the French contribution to AdriaArray. The XP array is complemented with 3 quasi-linear north-south, east-west and northwest-southeast profiles with inter-station spacing of 5-20 km, making up the XF network of 65 stations. The profiles cross volcanic areas and the main Variscan structures. We describe the experimental setup designed to optimize the performance/cost ratio and minimize the number of field visits, the deployment, the state-of-health monitoring, the data management and the data quality control strategies, outcomes of our 15-years' experience with major temporary seismic experiments in France and neighboring countries, including AlpArray. We also show some preliminary results including hypocenter locations and receiver function analysis. The 2 broadband arrays will be supplemented in 2025 by a month-long deployment of 3 large-N dense arrays of 625 3-C short-period nodes. These dense arrays will complete our multi-scale seismic experiment and illuminate active faults and possible plumbing systems of the youngest volcanoes.
△ Less
Submitted 7 March, 2025;
originally announced March 2025.
-
Cosmic Rays and the Askaryan Effect Reveal Subsurface Structure and Buried Ice on the Moon
Authors:
E. S. Costello,
R. R. Ghent,
A. Romero-Wolf,
P. W. Gorham,
P. G. Lucey,
C. J. Tai Udovicic,
P. Linton,
A. Ludwig,
K. McBride,
C. Miki,
E. Oberla,
J. Rolla,
A. Jung
Abstract:
We present the first full-wavelength numerical simulations of the electric field generated by cosmic ray impacts into the Moon. Billions of cosmic rays fall onto the Moon every year. Ultra-high energy cosmic ray impacts produce secondary particle cascades within the regolith and subsequent coherent, widebandwidth, linearly-polarized radio pulses by the Askaryan Effect. Observations of the cosmic r…
▽ More
We present the first full-wavelength numerical simulations of the electric field generated by cosmic ray impacts into the Moon. Billions of cosmic rays fall onto the Moon every year. Ultra-high energy cosmic ray impacts produce secondary particle cascades within the regolith and subsequent coherent, widebandwidth, linearly-polarized radio pulses by the Askaryan Effect. Observations of the cosmic ray particle shower radio emissions can reveal subsurface structure on the Moon and enable the broad and deep prospecting necessary to confirm or refute the existence of polar ice deposits. Our simulations show that the radio emissions and reflections could reveal ice layers as thin as 10 cm and buried under regolith as deep as 9 m. The Askaryan Effect presents a novel and untapped opportunity for characterizing buried lunar ice at unprecedented depths and spatial scales.
△ Less
Submitted 6 March, 2025;
originally announced March 2025.
-
Membrane phononic crystals for high-Qm mechanical defect modes in piezoelectric aluminum nitride
Authors:
Anastasiia Ciers,
Laurentius Radit Nindito,
Alexander Jung,
Hannes Pfeifer,
Armin Dadgar,
Andre Strittmatter,
Witlef Wieczorek
Abstract:
Nanomechanical resonators with exceptionally low dissipation are advancing mechanics-based sensors and quantum technologies. The key for these advances is the engineering of localized phononic modes that are well-isolated from the environment, i.e., that exhibit a high mechanical quality factor, Qm. Membrane phononic crystals fabricated from strained thin films can realize high-Qm single or multip…
▽ More
Nanomechanical resonators with exceptionally low dissipation are advancing mechanics-based sensors and quantum technologies. The key for these advances is the engineering of localized phononic modes that are well-isolated from the environment, i.e., that exhibit a high mechanical quality factor, Qm. Membrane phononic crystals fabricated from strained thin films can realize high-Qm single or multiple localized phononic defect modes at MHz frequencies. These defect modes can be efficiently interfaced with out-of-plane light or coupled to a microwave quantum circuit, enabling readout and control of their motion. When membrane phononic crystals are fabricated from a crystalline film, they could offer built-in functionality. We demonstrate a membrane phononic crystal realized in a strained 90 nm-thin film of aluminum nitride (AlN), which is a crystalline piezoelectric material. We engineer a high-Qm localized phononic defect mode at 1.8 MHz with a Qxf-product of 1.5x10^13 Hz at room temperature. In future devices, the built-in piezoelectricity of AlN can be utilized for direct coupling to qubits or in-situ tuning of mechanical mode frequencies, defect mode couplings, or acoustic bandgaps, which can be used as building blocks of tunable phononic circuits or low-noise sensors.
△ Less
Submitted 25 June, 2025; v1 submitted 31 January, 2025;
originally announced January 2025.
-
3D image based stochastic micro-structure modelling of foams for simulating elasticity
Authors:
Anne Jung,
Claudia Redenbach,
Katja Schladitz,
Sarah Staub
Abstract:
Image acquisition techniques such as micro-computed tomography are nowadays widely available. Quantitative analysis of the resulting 3D image data enables geometric characterization of the micro-structure of materials. Stochastic geometry models can be fit to the observed micro-structures. By alteration of the model parameters, virtual micro-structures with modified geometry can be generated. Nume…
▽ More
Image acquisition techniques such as micro-computed tomography are nowadays widely available. Quantitative analysis of the resulting 3D image data enables geometric characterization of the micro-structure of materials. Stochastic geometry models can be fit to the observed micro-structures. By alteration of the model parameters, virtual micro-structures with modified geometry can be generated. Numerical simulation of elastic properties in realizations of these models yields deeper insight on the influence of particular micro-structural features. Ultimately, this allows for an optimization of the micro-structure geometry for particular applications. Here, we present this workflow at the example of open cell foams. Applicability is demonstrated using an aluminum alloy foam sample. The structure observed in a micro-computed tomography image is modeled by the edge system of a random Laguerre tessellation generated by a system of closely packed spheres. Elastic moduli are computed in the binarized micro-CT image of the foam as well as in realizations of the model. They agree well with the results of a compression test on the real material.
△ Less
Submitted 27 January, 2025;
originally announced January 2025.
-
Enforcing Fundamental Relations via Adversarial Attacks on Input Parameter Correlations
Authors:
Timo Saala,
Lucie Flek,
Alexander Jung,
Akbar Karimi,
Alexander Schmidt,
Matthias Schott,
Philipp Soldin,
Christopher Wiebusch
Abstract:
Correlations between input parameters play a crucial role in many scientific classification tasks, since these are often related to fundamental laws of nature. For example, in high energy physics, one of the common deep learning use-cases is the classification of signal and background processes in particle collisions. In many such cases, the fundamental principles of the correlations between obser…
▽ More
Correlations between input parameters play a crucial role in many scientific classification tasks, since these are often related to fundamental laws of nature. For example, in high energy physics, one of the common deep learning use-cases is the classification of signal and background processes in particle collisions. In many such cases, the fundamental principles of the correlations between observables are often better understood than the actual distributions of the observables themselves. In this work, we present a new adversarial attack algorithm called Random Distribution Shuffle Attack (RDSA), emphasizing the correlations between observables in the network rather than individual feature characteristics. Correct application of the proposed novel attack can result in a significant improvement in classification performance - particularly in the context of data augmentation - when using the generated adversaries within adversarial training. Given that correlations between input features are also crucial in many other disciplines. We demonstrate the RDSA effectiveness on six classification tasks, including two particle collision challenges (using CERN Open Data), hand-written digit recognition (MNIST784), human activity recognition (HAR), weather forecasting (Rain in Australia), and ICU patient mortality (MIMIC-IV), demonstrating a general use case beyond fundamental physics for this new type of adversarial attack algorithms.
△ Less
Submitted 9 January, 2025;
originally announced January 2025.
-
IAM: Enhancing RGB-D Instance Segmentation with New Benchmarks
Authors:
Aecheon Jung,
Soyun Choi,
Junhong Min,
Sungeun Hong
Abstract:
Image segmentation is a vital task for providing human assistance and enhancing autonomy in our daily lives. In particular, RGB-D segmentation-leveraging both visual and depth cues-has attracted increasing attention as it promises richer scene understanding than RGB-only methods. However, most existing efforts have primarily focused on semantic segmentation and thus leave a critical gap. There is…
▽ More
Image segmentation is a vital task for providing human assistance and enhancing autonomy in our daily lives. In particular, RGB-D segmentation-leveraging both visual and depth cues-has attracted increasing attention as it promises richer scene understanding than RGB-only methods. However, most existing efforts have primarily focused on semantic segmentation and thus leave a critical gap. There is a relative scarcity of instance-level RGB-D segmentation datasets, which restricts current methods to broad category distinctions rather than fully capturing the fine-grained details required for recognizing individual objects. To bridge this gap, we introduce three RGB-D instance segmentation benchmarks, distinguished at the instance level. These datasets are versatile, supporting a wide range of applications from indoor navigation to robotic manipulation. In addition, we present an extensive evaluation of various baseline models on these benchmarks. This comprehensive analysis identifies both their strengths and shortcomings, guiding future work toward more robust, generalizable solutions. Finally, we propose a simple yet effective method for RGB-D data integration. Extensive evaluations affirm the effectiveness of our approach, offering a robust framework for advancing toward more nuanced scene understanding.
△ Less
Submitted 3 January, 2025;
originally announced January 2025.
-
SyMerge: From Non-Interference to Synergistic Merging via Single-Layer Adaptation
Authors:
Aecheon Jung,
Seunghwan Lee,
Dongyoon Han,
Sungeun Hong
Abstract:
Model merging combines independently trained models into a single multi-task model. However, most existing approaches focus primarily on avoiding task interference. We argue that its greater potential lies in enabling task synergy, where tasks actively improve one another. We identify cross-task performance, defined by compatibility between encoders and predictors across tasks, as a key indicator…
▽ More
Model merging combines independently trained models into a single multi-task model. However, most existing approaches focus primarily on avoiding task interference. We argue that its greater potential lies in enabling task synergy, where tasks actively improve one another. We identify cross-task performance, defined by compatibility between encoders and predictors across tasks, as a key indicator of merge quality. We demonstrate that adapting only a single task-specific layer is sufficient to induce such synergy. This study proposes SyMerge, a lightweight framework that jointly optimizes merging coefficients and a single task-specific layer. We adopt an expert-guided self-labeling objective, providing stable supervision beyond entropy minimization. Intriguingly, we further show that SyMerge successfully merges models trained from different initializations, a regime where standard methods break down. Our minimalist yet principled method achieves state-of-the-art results across vision, dense prediction, and NLP benchmarks. Our code is available at https://aim-skku.github.io/SyMerge
△ Less
Submitted 22 May, 2026; v1 submitted 26 December, 2024;
originally announced December 2024.
-
Machine Learning-Assisted Measurement of Lepton-Jet Azimuthal Angular Asymmetries in Deep-Inelastic Scattering at HERA
Authors:
The H1 collaboration,
V. Andreev,
M. Arratia,
A. Baghdasaryan,
A. Baty,
K. Begzsuren,
A. Bolz,
V. Boudry,
G. Brandt,
D. Britzger,
A. Buniatyan,
L. Bystritskaya,
A. J. Campbell,
K. B. Cantun Avila,
K. Cerny,
V. Chekelian,
Z. Chen,
J. G. Contreras,
J. Cvach,
J. B. Dainton,
K. Daum,
A. Deshpande,
C. Diaconu,
A. Drees,
G. Eckerlin
, et al. (119 additional authors not shown)
Abstract:
In deep-inelastic positron-proton scattering, the lepton-jet azimuthal angular asymmetry is measured using data collected with the H1 detector at HERA. When the average transverse momentum of the lepton-jet system, $\lvert \vec{P}_\perp \rvert $, is much larger than the total transverse momentum of the system, $\lvert \vec{q}_\perp \rvert$, the asymmetry between parallel and antiparallel configura…
▽ More
In deep-inelastic positron-proton scattering, the lepton-jet azimuthal angular asymmetry is measured using data collected with the H1 detector at HERA. When the average transverse momentum of the lepton-jet system, $\lvert \vec{P}_\perp \rvert $, is much larger than the total transverse momentum of the system, $\lvert \vec{q}_\perp \rvert$, the asymmetry between parallel and antiparallel configurations, $\vec{P}_\perp$ and $\vec{q}_\perp$, is expected to be generated by initial and final state soft gluon radiation and can be predicted using perturbation theory. Quantifying the angular properties of the asymmetry therefore provides an additional test of the strong force. Studying the asymmetry is important for future measurements of intrinsic asymmetries generated by the proton's constituents through Transverse Momentum Dependent (TMD) Parton Distribution Functions (PDFs), where this asymmetry constitutes a dominant background. Moments of the azimuthal asymmetries are measured using a machine learning method for unfolding that does not require binning.
△ Less
Submitted 21 December, 2024; v1 submitted 18 December, 2024;
originally announced December 2024.
-
Bumblebee: Foundation Model for Particle Physics Discovery
Authors:
Andrew J. Wildridge,
Jack P. Rodgers,
Ethan M. Colbert,
Yao yao,
Andreas W. Jung,
Miaoyuan Liu
Abstract:
Bumblebee is a foundation model for particle physics discovery, inspired by BERT. By removing positional encodings and embedding particle 4-vectors, Bumblebee captures both generator- and reconstruction-level information while ensuring sequence-order invariance. Pre-trained on a masked task, it improves dileptonic top quark reconstruction resolution by 10-20% and excels in downstream tasks, includ…
▽ More
Bumblebee is a foundation model for particle physics discovery, inspired by BERT. By removing positional encodings and embedding particle 4-vectors, Bumblebee captures both generator- and reconstruction-level information while ensuring sequence-order invariance. Pre-trained on a masked task, it improves dileptonic top quark reconstruction resolution by 10-20% and excels in downstream tasks, including toponium discrimination (AUROC 0.877) and initial state classification (AUROC 0.625). The flexibility of Bumblebee makes it suitable for a wide range of particle physics applications, especially the discovery of new particles.
△ Less
Submitted 10 December, 2024;
originally announced December 2024.
-
A note on infinite versions of $(p,q)$-theorems
Authors:
Attila Jung,
Dömötör Pálvölgyi
Abstract:
We prove that fractional Helly and $(p,q)$-theorems imply $(\aleph_0,q)$-theorems in an entirely abstract setting. We give a plethora of applications, including reproving almost all earlier $(\aleph_0,q)$-theorems about geometric hypergraphs that were proved recently. Some of the corollaries are new results, for example, we prove that if $\mathcal{F}$ is an infinite family of convex compact sets i…
▽ More
We prove that fractional Helly and $(p,q)$-theorems imply $(\aleph_0,q)$-theorems in an entirely abstract setting. We give a plethora of applications, including reproving almost all earlier $(\aleph_0,q)$-theorems about geometric hypergraphs that were proved recently. Some of the corollaries are new results, for example, we prove that if $\mathcal{F}$ is an infinite family of convex compact sets in $\mathbb{R}^d$ and among every $\aleph_0$ of the sets some $d+1$ contain a point in their intersection with integer coordinates, then all the members of $\mathcal{F}$ can be hit with finitely many points with integer coordinates.
△ Less
Submitted 5 December, 2024;
originally announced December 2024.
-
Continuous Domains for Function Spaces Using Spectral Compactification
Authors:
Amin Farjudian,
Achim Jung
Abstract:
We introduce a continuous domain for function spaces over topological spaces which are not core-compact. Notable examples of such topological spaces include the real line with the upper limit topology, which is used in solution of initial value problems with temporal discretization, and various infinite dimensional Banach spaces which are ubiquitous in functional analysis and solution of partial d…
▽ More
We introduce a continuous domain for function spaces over topological spaces which are not core-compact. Notable examples of such topological spaces include the real line with the upper limit topology, which is used in solution of initial value problems with temporal discretization, and various infinite dimensional Banach spaces which are ubiquitous in functional analysis and solution of partial differential equations. If a topological space $\mathbb{X}$ is not core-compact and $\mathbb{D}$ is a non-singleton bounded-complete domain, the function space $[\mathbb{X} \to \mathbb{D}]$ is not a continuous domain. To construct a continuous domain, we consider a spectral compactification $\mathbb{Y}$ of $\mathbb{X}$ and relate $[\mathbb{X} \to \mathbb{D}]$ with the continuous domain $[\mathbb{Y} \to \mathbb{D}]$ via a Galois connection. This allows us to perform computations in the native structure $[\mathbb{X} \to \mathbb{D}]$ while computable analysis is performed in the continuous domain $[\mathbb{Y} \to \mathbb{D}]$, with the left and right adjoints used for moving between the two function spaces.
△ Less
Submitted 7 December, 2024; v1 submitted 11 November, 2024;
originally announced November 2024.
-
Engineering Trustworthy AI: A Developer Guide for Empirical Risk Minimization
Authors:
Diana Pfau,
Alexander Jung
Abstract:
AI systems increasingly shape critical decisions across personal and societal domains. While empirical risk minimization (ERM) drives much of the AI success, it typically prioritizes accuracy over trustworthiness, often resulting in biases, opacity, and other adverse effects. This paper discusses how key requirements for trustworthy AI can be translated into design choices for the components of ER…
▽ More
AI systems increasingly shape critical decisions across personal and societal domains. While empirical risk minimization (ERM) drives much of the AI success, it typically prioritizes accuracy over trustworthiness, often resulting in biases, opacity, and other adverse effects. This paper discusses how key requirements for trustworthy AI can be translated into design choices for the components of ERM. We hope to provide actionable guidance for building AI systems that meet emerging standards for trustworthiness of AI.
△ Less
Submitted 6 November, 2024; v1 submitted 25 October, 2024;
originally announced October 2024.
-
Thickness dependence of the mechanical properties of piezoelectric high-$Q_m$ nanomechanical resonators made from aluminium nitride
Authors:
Anastasiia Ciers,
Alexander Jung,
Joachim Ciers,
Laurentius Radit Nindito,
Hannes Pfeifer,
Armin Dadgar,
Jürgen Bläsing,
André Strittmatter,
Witlef Wieczorek
Abstract:
Nanomechanical resonators with high quality factors (\Qm{}) enable mechanics-based quantum technologies, in particular quantum sensing and quantum transduction. High-\Qm{} nanomechanical resonators in the kHz to MHz frequency range can be realized in tensile-strained thin films that allow the use of dissipation dilution techniques to drastically increase \Qm{}. In our work, we study the material p…
▽ More
Nanomechanical resonators with high quality factors (\Qm{}) enable mechanics-based quantum technologies, in particular quantum sensing and quantum transduction. High-\Qm{} nanomechanical resonators in the kHz to MHz frequency range can be realized in tensile-strained thin films that allow the use of dissipation dilution techniques to drastically increase \Qm{}. In our work, we study the material properties of tensile-strained piezoelectric films made from aluminium nitride (AlN). We characterize crystalline AlN films with a thickness ranging from \SI{45}{\nano\meter} to \SI{295}{\nano\meter}, which are directly grown on Si(111) by metal-organic vapour-phase epitaxy. We report on the crystal quality and surface roughness, the piezoelectric response, and the residual and released stress of the AlN thin films. Importantly, we determine the intrinsic quality factor of the films at room temperature in high vacuum. We fabricate and characterize AlN nanomechanical resonators that exploit dissipation dilution to enhance the intrinsic quality factor by utilizing the tensile strain in the film. We find that AlN nanomechanical resonators below \SI{200}{\nano\meter} thickness exhibit the highest \Qf{}-product, on the order of $10^{12}$\,Hz. We discuss possible strategies to optimize the material growth that should lead to devices that reach even higher \Qf{}-products. This will pave the way for future advancements of optoelectromechanical quantum devices made from tensile-strained piezoelectric AlN.
△ Less
Submitted 13 January, 2025; v1 submitted 4 October, 2024;
originally announced October 2024.
-
Evaluating WAIC and PSIS-LOO for Bayesian Diagnostic Classification Model Selection
Authors:
Ae Kyong Jung,
Jonathan Templin
Abstract:
Bayesian diagnostic classification models (Bayesian DCMs) are effective for diagnosing students' skills. Research on the evaluation of relative model fit indices for DCMs using Bayesian estimation, however, is deficient. This study introduces the performance of Bayesian relative model fit indices, the widely applicable information criterion (WAIC) and leave-one-out cross-validation using Pareto-sm…
▽ More
Bayesian diagnostic classification models (Bayesian DCMs) are effective for diagnosing students' skills. Research on the evaluation of relative model fit indices for DCMs using Bayesian estimation, however, is deficient. This study introduces the performance of Bayesian relative model fit indices, the widely applicable information criterion (WAIC) and leave-one-out cross-validation using Pareto-smoothed importance sampling (PSIS-LOO), in comparison to simpler and more widely used deviance information criterion (DIC). The simulation study evaluates the performance of WAIC and PSIS-LOO by detecting the true model with varying sample sizes, item qualities, and prior information levels. The results of the study indicate that WAIC and PSIS-LOO primarily favored the generating model; however, occasional inconsistencies were observed. This study recommends using WAIC and PSIS-LOO when the data is assumed to follow a simpler model and the models are estimated under uninformative priors, and DIC when the data is assumed to follow a more complex model.
△ Less
Submitted 3 October, 2024;
originally announced October 2024.
-
Coherent Dipolar Coupling between Magnetoelastic Waves and Nitrogen Vacancy Centers
Authors:
Adi Jung,
Samuel Margueron,
Ausrine Bartasyte,
Sayeef Salahuddin
Abstract:
We experimentally demonstrate coherent Rabi oscillations of Nitrogen Vacancy (NV) centers by magnetoelastic waves. The coupling is consistent with dipolar stray field drive from spin-wave modes in a ferromagnetic film, and displays a significant improvement in Radio Frequency power efficiency relative to other methods of microwave excitation. Further, it demonstrates coherent coupling with NV cent…
▽ More
We experimentally demonstrate coherent Rabi oscillations of Nitrogen Vacancy (NV) centers by magnetoelastic waves. The coupling is consistent with dipolar stray field drive from spin-wave modes in a ferromagnetic film, and displays a significant improvement in Radio Frequency power efficiency relative to other methods of microwave excitation. Further, it demonstrates coherent coupling with NV centers over mm-scale distances from the microwave excitation source. By utilizing a piezoelectric-magnetostrictive heterostucture, where magnetoelastic waves can be launched by an applied voltage, a pure voltage driven coherent drive of the NV centers is achieved. This voltage driven, magnetoelastic excitation enables a new approach to couple with two level quantum states that is not reliant on long spin-wave coherence lengths.
△ Less
Submitted 18 September, 2024; v1 submitted 16 September, 2024;
originally announced September 2024.
-
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
Authors:
Konstantin Lübeck,
Alexander Louis-Ferdinand Jung,
Felix Wedlich,
Mika Markus Müller,
Federico Nicolás Peccia,
Felix Thömmes,
Jannik Steinmetz,
Valentin Biermaier,
Adrian Frischknecht,
Paul Palomero Bernardo,
Oliver Bringmann
Abstract:
Implementing Deep Neural Networks (DNNs) on resource-constrained edge devices is a challenging task that requires tailored hardware accelerator architectures and a clear understanding of their performance characteristics when executing the intended AI workload. To facilitate this, we present an automated generation approach for fast performance models to accurately estimate the latency of a DNN ma…
▽ More
Implementing Deep Neural Networks (DNNs) on resource-constrained edge devices is a challenging task that requires tailored hardware accelerator architectures and a clear understanding of their performance characteristics when executing the intended AI workload. To facilitate this, we present an automated generation approach for fast performance models to accurately estimate the latency of a DNN mapped onto systematically modeled and concisely described accelerator architectures. Using our accelerator architecture description method, we modeled representative DNN accelerators such as Gemmini, UltraTrail, Plasticine-derived, and a parameterizable systolic array. Together with DNN mappings for those modeled architectures, we perform a combined DNN/hardware dependency graph analysis, which enables us, in the best case, to evaluate only 154 loop kernel iterations to estimate the performance for 4.19 billion instructions achieving a significant speedup. We outperform regression and analytical models in terms of mean absolute percentage error (MAPE) compared to simulation results, while being several magnitudes faster than an RTL simulation.
△ Less
Submitted 13 September, 2024;
originally announced September 2024.