-
A Bayesian approach to the long-baseline neutrino oscillation sensitivity of DUNE
Authors:
DUNE Collaboration,
S. Abbaslu,
F. Abd Alrahman,
A. Abed Abud,
R. Acciarri,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
K. Adhikari,
C. Adriano,
K. Agudelo-Jaramillo,
F. Akbar,
F. Alemanno,
N. S. Alex,
L. Aliaga Soplin,
A. Alqaisi,
O. Alterkait,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
R. M. Amarinei,
P. Amedo
, et al. (1262 additional authors not shown)
Abstract:
The sensitivity of the Deep Underground Neutrino Experiment (DUNE) to neutrino oscillation is evaluated using a Bayesian Markov Chain Monte Carlo (MCMC) approach. This analysis uses the same underlying sensitivity inputs as previous DUNE studies [Eur. Phys. J. C 80, 978 (2020)], and therefore does not present updated DUNE sensitivities, but instead explores the additional inferences accessible usi…
▽ More
The sensitivity of the Deep Underground Neutrino Experiment (DUNE) to neutrino oscillation is evaluated using a Bayesian Markov Chain Monte Carlo (MCMC) approach. This analysis uses the same underlying sensitivity inputs as previous DUNE studies [Eur. Phys. J. C 80, 978 (2020)], and therefore does not present updated DUNE sensitivities, but instead explores the additional inferences accessible using a Bayesian approach. We present four-dimensional posterior probability distributions of the oscillation parameters, highlighting the breadth of correlation in the parameter space of interest, especially between $\sin^2 θ_{23}$ and $\sin^2 θ_{13}$. We exploit the flexibility of the Bayesian framework to incorporate parameter constraints post hoc and assess the impact of applying a reactor short-baseline $θ_{13}$ constraint. A significant increase in the sensitivity to the $θ_{23}$ octant is found when including the constraint. Posterior distributions of derived quantities can be easily constructed from MCMC results. This work presents the first study of DUNE's sensitivity to the Jarlskog invariant, $J$, a quantity that provides a parametrisation-independent measure of charge-parity violation in the leptonic sector.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Digital State Capacity
Authors:
Patrick Healy,
Simon D. Angus,
Paul Raschky,
Klaus Ackermann,
Nathan Lane,
Weijia Li,
Cynthia Huang
Abstract:
Digital State Capacity is the ability of governments to deploy ICT infrastructure and information systems to implement policy. This paper introduces a new measure of government ICT capacity based on an observable stock of deployable public-sector network infrastructure: public IPv4 address space held by government organisations. These address holdings are key inputs into digital administration bec…
▽ More
Digital State Capacity is the ability of governments to deploy ICT infrastructure and information systems to implement policy. This paper introduces a new measure of government ICT capacity based on an observable stock of deployable public-sector network infrastructure: public IPv4 address space held by government organisations. These address holdings are key inputs into digital administration because they support internet-facing systems, networked information exchange, and coordination across agencies and functions. The core panel covers approximately 150,000 country-entity records classified as government across more than 150 countries from 2019 to 2024 and can be disaggregated by administrative level and government function. In the 2019 to 2024 Admin-1 panel, government IP holdings are observed in 1,681 subnational regions across all years. We validate the measure at the crosscountry and subnational levels and apply it to government tasks related to corruption control and vaccination rollout. In illustrative country-year analysis, higher Digital State Capacity is associated with higher-quality governance and publicservice outcomes in the expected directions, including lower measured corruption and higher vaccination coverage. These associations are descriptive; they demonstrate the empirical relevance of the measure and are not causal estimates.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
Authors:
James Elcock,
William F. Shen,
Xinchi Qiu,
Nicholas D. Lane
Abstract:
Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised f…
▽ More
Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
Operation and performance of ProtoDUNE Dual Phase liquid argon time projection chamber
Authors:
DUNE Collaboration,
S. Abbaslu,
F. Abd Alrahman,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
K. Adhikari,
C. Adriano,
K. Agudelo-Jaramillo,
F. Akbar,
F. Alemanno,
N. S. Alex,
L. Aliaga Soplin,
A. Alqaisi,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
R. Amarinei
, et al. (1341 additional authors not shown)
Abstract:
ProtoDUNE-DP was the largest ever built Liquid Argon Time Projection Chamber (LArTPC) operating in Dual-Phase (DP) mode, with a liquid target and charge read-out placed in the gas. It had an active volume of $6\times6\times6$\,m$^3$ corresponding to an active mass of 300\,t (total LAr mass of 720\,t), constructed at the CERN Neutrino Platform and took data from 2019 to 2020 with cosmic muons. In P…
▽ More
ProtoDUNE-DP was the largest ever built Liquid Argon Time Projection Chamber (LArTPC) operating in Dual-Phase (DP) mode, with a liquid target and charge read-out placed in the gas. It had an active volume of $6\times6\times6$\,m$^3$ corresponding to an active mass of 300\,t (total LAr mass of 720\,t), constructed at the CERN Neutrino Platform and took data from 2019 to 2020 with cosmic muons. In ProtoDUNE-DP the electric drift field is oriented in the vertical direction, causing the electrons to drift vertically towards the anode at the top. The ionization charge is then extracted into the gaseous argon above the liquid surface, amplified by Townsend avalanches, and collected by the charge readout planes. The detector experienced significant technical problems affecting the long-term operation of the Charge Readout Planes, formed by the Large Electron Multipliers, but other critical segments demonstrated required performance including the delivery of -300 kV to the TPC cathode, verification of replaceable charge read-out electronics, and operation of the photon detection system. ProtoDUNE-DP experience resulted in improved designs of the Vertical Drift LArTPC.
△ Less
Submitted 21 July, 2026; v1 submitted 17 July, 2026;
originally announced July 2026.
-
The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
Authors:
Alex Iacob,
Andrej Jovanović,
William F. Shen,
Daniel Burkhardt,
Meghdad Kurmanji,
Nurbek Tastan,
Lorenzo Sani,
Niccolò Alberto Elia Venanzi,
Ambroise Odonnat,
Zeyu Cao,
Bill Marino,
Xinchi Qiu,
Nicholas D. Lane
Abstract:
Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmark, or labeled dataset that remains valid as the agent improves. This ignores a central feature of evolution: species adapt as their environments change with them…
▽ More
Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmark, or labeled dataset that remains valid as the agent improves. This ignores a central feature of evolution: species adapt as their environments change with them. We aim to bring the same principle to recursive self-improvement, making evaluation part of the improvement loop and opening search to evolving evaluators, adversarial objectives, and dynamic utilities that may surpass static benchmarks. We introduce the Red Queen Godel Machine (RQGM), an evolutionary framework for recursive self-improvement under non-stationary utilities. The RQGM makes this possible through controlled utility evolution: search is organized into epochs with a fixed within-epoch evaluation criterion, while the utility can be updated at epoch boundaries, so self-improvement guarantees hold per epoch as the objective evolves across them. We begin by showing that even on verifiable coding tasks, the RQGM improves test pass rate over the prior SOTA by adding a complementary agent-as-a-judge code-review signal. This signal is cheaper and the RQGM uses 1.35x-1.72x fewer tokens. We then turn to scientific paper writing and reviewing, and Olympiad-level proof writing and grading, where the RQGM improves performance over prior self-improving agents: co-evolved writers reach 1.78x-1.86x higher acceptance rates under a diverse agent-as-a-judge panel, while co-evolved graders reach 9% higher ground-truth accuracy. In paper reviewing, the strongest baseline reviewer over-accepts AI-generated papers at up to 1.91x the human rate. The RQGM corrects this by introducing an adversarial objective that discovers reviewers equally stringent on AI and human work.
△ Less
Submitted 29 June, 2026; v1 submitted 24 June, 2026;
originally announced June 2026.
-
Production and installation of wavelength-shifting reflective light enhancers for the Short-Baseline Near Detector
Authors:
R. Acciarri,
L. Aliaga-Soplin,
R. Alvarez-Garrote,
D. Andrade Aldana,
C. Andreopoulos,
A. Antonakis,
S. Balasubramanian,
A. Barnard,
V. Basque,
J. Bateman,
M. C. Bazetto,
A. Beever,
E. Belchior,
M. Betancourt,
A. Bhat,
M. Bishai,
A. Blake,
B. Bogart,
D. Brailsford,
A. Brandt,
S. Brickner,
M. B. Brunetti,
L. Camilleri,
D. Caratelli,
D. Carber
, et al. (172 additional authors not shown)
Abstract:
We report on the design, production, and installation of a wavelength-shifting reflective system on the cathode of the Short-Baseline Near Detector (SBND), a liquid argon time projection chamber located along the Fermilab Booster Neutrino Beam. To increase and homogenize scintillation-light collection, 64 double-sided plates were fabricated from FR4, laminated with specular reflector film and coat…
▽ More
We report on the design, production, and installation of a wavelength-shifting reflective system on the cathode of the Short-Baseline Near Detector (SBND), a liquid argon time projection chamber located along the Fermilab Booster Neutrino Beam. To increase and homogenize scintillation-light collection, 64 double-sided plates were fabricated from FR4, laminated with specular reflector film and coated with 300 $μ$g/cm$^2$ of tetraphenyl butadiene (TPB) wavelength shifter using controlled physical vapor deposition. The coating uniformity was validated through dedicated measurements of deposited mass and profilometry studies. Because exposure to ambient blue/UV light could degrade the TPB, protective filtering and controlled storage conditions were implemented during handling and installation. The coated plates were assembled between conductive meshes for high-voltage compatibility and installed in situ during detector integration. This system constitutes the largest TPB-coated area deployed in a neutrino detector. It operates in conjunction with SBND's photon detection system, which consists of photomultiplier tubes and X-ARAPUCAs. Early light-collection measurements show high uniformity and light response across the detector, supporting improved triggering, calorimetry, and position reconstruction in SBND.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
Probing Nuclear Effects with Transverse Kinematic Imbalance in Muon-neutrino Induced Charged-Current $π^0$ Production on Argon with the MicroBooNE Detector
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
V. Bhelande,
M. Bhattacharya,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (170 additional authors not shown)
Abstract:
Neutrino-nucleus cross-section measurements are needed to improve interaction modeling and to enable precision neutrino oscillation measurements in upcoming experiments such as the Deep Underground Neutrino Experiment (DUNE), Hyper-Kamiokande, and the Short-Baseline Neutrino program. Baryon-resonance neutrino interactions constitute a dominant contribution near the peak of the DUNE neutrino energy…
▽ More
Neutrino-nucleus cross-section measurements are needed to improve interaction modeling and to enable precision neutrino oscillation measurements in upcoming experiments such as the Deep Underground Neutrino Experiment (DUNE), Hyper-Kamiokande, and the Short-Baseline Neutrino program. Baryon-resonance neutrino interactions constitute a dominant contribution near the peak of the DUNE neutrino energy spectrum. We present the first measurement of muon neutrino charged-current resonance-like interactions on argon using transverse kinematic imbalance variables with the MicroBooNE detector. These observables are highly sensitive to the modeling of final-state interactions. This measurement probes kinematic imbalances using the reconstructed momenta of the muon, leading proton, and neutral pion. A comprehensive characterization of the $π^0$-proton final state is presented; however, none of the models considered are able to simultaneously reproduce all measured observables.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs
Authors:
Lorenzo Sani,
Zeyu Cao,
Meghdad Kurmanji,
Alex Iacob,
Andrej Jovanovic,
Yan Gao,
Wanru Zhao,
Nicholas D. Lane
Abstract:
Pre-training Large Language Models (LLMs) typically demands large-scale infrastructure with tightly coupled hardware accelerators. Mixture-of-Experts (MoEs) architectures partially decouple model capacity from per-token compute. This efficiency alone does not make MoE training feasible over ordinary Internet links or loosely connected commodity hardware since active expert routing still assumes hi…
▽ More
Pre-training Large Language Models (LLMs) typically demands large-scale infrastructure with tightly coupled hardware accelerators. Mixture-of-Experts (MoEs) architectures partially decouple model capacity from per-token compute. This efficiency alone does not make MoE training feasible over ordinary Internet links or loosely connected commodity hardware since active expert routing still assumes high-speed datacenter fabrics. Low-communication methods such as DiLoCo and Photon reduce synchronization frequency across distributed sites, mitigating bandwidth constraints, yet still require full model replicas at every site. This creates a mismatch: modern MoEs have sparse data paths, but their distributed training infrastructure remains communication-dense and memory-inefficient, limiting attempts to pool geographically distributed compute. In this work, we introduce FoMoE, a system that breaks the full-replica paradigm by partitioning expert layers across workers and skipping non-resident experts during local training. We demonstrate that FoMoE: (I) reduces communication costs by up to 1.42x over efficient baselines and 45.44x over Distributed Data Parallelism (DDP) via partial expert replication in controlled regimes; (II) achieves empirical throughput speedups of up to 1.4x through the skip-token mechanism; and (III) shows stable routing in the trained regimes and projects the communication/memory benefits to 100B-scale configurations through system modeling.
△ Less
Submitted 20 June, 2026; v1 submitted 17 June, 2026;
originally announced June 2026.
-
PreLort: Prefix-Nested LoRA for Federated Fine-Tuning under Rank Heterogeneity
Authors:
Muhammad Waseem,
Nurbek Tastan,
Andrej Jovanovic,
Nicholas D. Lane,
Nils Lukas,
Karthik Nandakumar,
Samuel Horvath
Abstract:
Federated fine-tuning of large language models using parameter-efficient methods such as LoRA enables privacy-preserving adaptation of foundation models. Heterogeneous hardware resources introduce challenges, as clients with different adapter ranks cannot be directly aggregated. While existing methods enable aggregation under heterogeneous ranks, they fail to control how information is distributed…
▽ More
Federated fine-tuning of large language models using parameter-efficient methods such as LoRA enables privacy-preserving adaptation of foundation models. Heterogeneous hardware resources introduce challenges, as clients with different adapter ranks cannot be directly aggregated. While existing methods enable aggregation under heterogeneous ranks, they fail to control how information is distributed across rank dimensions, leading to suboptimal use of shared low-rank representations. Instead, we propose PreLort: a nested low-rank formulation for federated LoRA that organizes adapter dimensions into a prefix hierarchy. Our approach ensures that lower-rank dimensions encode task-relevant information, while higher-rank dimensions capture additional capacity. Building on this, we introduce (i) a segment-wise aggregation rule that averages only over clients contributing to each rank segment, avoiding dilution from zero-padded lower-rank clients, and (ii) a prefix-nested training strategy that optimizes each adapter under multiple rank truncations, encouraging useful signal to concentrate in low-rank prefix dimensions. Together, these components encourage a consistent low-rank prefix capturing the most task-relevant information, while higher-rank dimensions learn additional capacity. This allows low-rank clients to benefit from richer information contributed by higher-rank clients, as prefix dimensions are consistently learned and aggregated. Experiments demonstrate that our method consistently outperforms prior heterogeneous federated LoRA methods in accuracy and ROUGE-L, while achieving lower or comparable perplexity across multiple base models.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
First Measurement of Sub-GeV $ν_μ$ Charged-Current Coherent Pion Production on Argon in MicroBooNE
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (167 additional authors not shown)
Abstract:
We report a measurement of the charged-current coherent pion production cross section on argon using the MicroBooNE liquid argon time projection chamber exposed to the Booster Neutrino Beam at Fermilab. The measurement uses the MicroBooNE data set corresponding to $1.26 \times 10^{21}$ protons on target with a mean neutrino energy of $0.8$~GeV. The flux-averaged cross section is measured to be…
▽ More
We report a measurement of the charged-current coherent pion production cross section on argon using the MicroBooNE liquid argon time projection chamber exposed to the Booster Neutrino Beam at Fermilab. The measurement uses the MicroBooNE data set corresponding to $1.26 \times 10^{21}$ protons on target with a mean neutrino energy of $0.8$~GeV. The flux-averaged cross section is measured to be $(9.1 \pm 1.2_{\text{stat}} \pm 1.2_\text{syst}) \times 10^{-40}\,\text{cm}^2/\text{Ar}$. This result represents the first measurement of charged-current coherent pion production on argon at sub-GeV neutrino energies. Due to its clean two-body kinematics, where the neutrino interacts coherently with the entire nucleus producing a forward muon and pion with no nuclear breakup, this process provides a useful tool for constraining neutrino flux uncertainties in current and future oscillation experiments such as DUNE.
△ Less
Submitted 11 June, 2026;
originally announced June 2026.
-
DXA-Derived Skeletal Phenotypes and Hip Fracture Risk: A Backdoor-Adjusted Causal Analysis
Authors:
Zixin Shi,
Chen Zhao,
Meiling Zhou,
Kevin A. Maupin,
Joyce H. Keyak,
Nancy E. Lane,
Kuan-Jui Su,
Hui Shen,
Hong-Wen Deng,
Kui Zhang,
Weihua Zhou
Abstract:
Purpose: To compare dual-energy X-ray absorptiometry (DXA)-derived hip skeletal phenotypes in relation to hip fracture risk using prespecified confounder adjustment and to assess whether phenotypes ranked by their backdoor-adjusted average treatment effects (ATEs) improve risk stratification. Methods: We analyzed 21,098 UK Biobank participants with linked health records, hip DXA-derived skeletal m…
▽ More
Purpose: To compare dual-energy X-ray absorptiometry (DXA)-derived hip skeletal phenotypes in relation to hip fracture risk using prespecified confounder adjustment and to assess whether phenotypes ranked by their backdoor-adjusted average treatment effects (ATEs) improve risk stratification. Methods: We analyzed 21,098 UK Biobank participants with linked health records, hip DXA-derived skeletal measures, and prespecified covariates. Sixteen phenotypes spanning bone mineral content (BMC), bone mineral density (BMD), and T-score across hip-related regions were evaluated. Confounder selection was guided by a prespecified directed acyclic graph (DAG). Backdoor-adjusted ATEs were estimated on the absolute risk-difference scale per standard deviation (SD) increase. Effect heterogeneity was evaluated for total femur BMD, and downstream prediction was assessed using clinical variables combined with phenotypes ranked by ATE magnitude. Results: Among 21,098 participants, 115 had hip fractures. All 16 phenotypes showed negative backdoor-adjusted ATEs per SD increase. The largest ATEs were observed for total femur BMC and total femur BMD, each with a risk difference of -0.0047, corresponding to approximately 4.7 fewer hip fractures per 1,000 participants per SD higher phenotype value. Conditional effects of total femur BMD were stronger among older participants and those with lower BMI. In prediction, clinical variables plus the top 11 ATE-ranked phenotypes achieved higher AUC than FRAX with femoral neck BMD (0.842 vs. 0.709), with higher sensitivity (0.748 vs. 0.443) and similar specificity (0.793 vs. 0.777). Conclusion: DXA-derived hip skeletal phenotypes differed in their backdoor-adjusted ATEs. Phenotype-level causal evaluation may help identify informative DXA measures for risk stratification.
△ Less
Submitted 29 May, 2026;
originally announced June 2026.
-
Characterizing the energy resolution of the MicroBooNE LArTPC at the MeV scale using monoenergetic features of $^{208}$Tl decays
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (167 additional authors not shown)
Abstract:
A detailed understanding of the capabilities and fidelity of low-energy reconstruction is crucial for taking advantage of MeV-scale neutrino physics opportunities in liquid argon time projection chambers (LArTPCs). This study presents a measurement of the resolution of reconstructed energy in the MicroBooNE LArTPC at $\approx 1.5$ MeV. The characterization is performed using monoenergetic signals…
▽ More
A detailed understanding of the capabilities and fidelity of low-energy reconstruction is crucial for taking advantage of MeV-scale neutrino physics opportunities in liquid argon time projection chambers (LArTPCs). This study presents a measurement of the resolution of reconstructed energy in the MicroBooNE LArTPC at $\approx 1.5$ MeV. The characterization is performed using monoenergetic signals generated by $2.614$ MeV $γ$-rays from $^{208}$Tl decays undergoing pair production in the detector. The resolution is found to be ($7.52 \pm 0.78 \text{(stat)} \pm 0.92 \text{(syst)}$)%. This value is consistent with the MicroBooNE simulation prediction of ($9.70 \pm 0.65 \text{(stat)}$)% at the $1.6 σ$ level. This study represents the first ever measurement of LArTPC energy resolution at the MeV scale and provides a pathway for monoenergetic energy calibrations in future experiments using LArTPC detectors.
△ Less
Submitted 17 August, 2026; v1 submitted 28 May, 2026;
originally announced May 2026.
-
Beyond Scaling: Agents Are Heading to the Edge
Authors:
Chunlin Tian,
Dongqi Cai,
Wanru Zhao,
Nicholas D. Lane
Abstract:
The bottleneck of useful agentic intelligence has shifted from compressing world knowledge into a single model to executing a coordinated system. This position paper argues that personal-agent architecture must move to the edge because the core properties of agentic intelligence tasks, particularly their structural coupling with high-fidelity local context and the need for zero-latency execution l…
▽ More
The bottleneck of useful agentic intelligence has shifted from compressing world knowledge into a single model to executing a coordinated system. This position paper argues that personal-agent architecture must move to the edge because the core properties of agentic intelligence tasks, particularly their structural coupling with high-fidelity local context and the need for zero-latency execution loops, do not sit well with cloud-centric designs. We develop this claim through three structural shifts. First, the Prefrontal Turn: the main marginal lever of capability has moved from pre-training scale to framework-level executive control. Such control must remain physically close to the environment of action if the agent is to preserve cognitive alignment. Second, the Data-Geography Paradox, the ``dark matter'' of agentic data (local file hierarchies, real-time sensor streams, and transient OS states) degrades, disappears, or loses meaning once prepared for cloud transmission, thereby cutting the agent off from ground-truth context. Third, the interaction-alignment loop, the only economically and ecologically sustainable source of agentic refinement data is the high-fidelity implicit preference signal produced through real-time local interaction. Third, the interaction-alignment loop, the only economically and ecologically sustainable source of agentic refinement data is the high-fidelity implicit preference signal produced through real-time local interaction. We conclude with falsifiable predictions for the next deployment cycle of personal agents.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
Authors:
Nurbek Tastan,
Alex Iacob,
Lorenzo Sani,
Meghdad Kurmanji,
Nicholas D. Lane,
Samuel Horvath,
Karthik Nandakumar
Abstract:
Multi-agent systems can solve complex tasks through collaboration between multiple Large Language Model agents. Existing collaboration frameworks typically operate in either a parallel or a sequential mode. In the parallel mode, agents respond independently to queries followed by aggregation of responses. In contrast, sequential systems allow agents to communicate via a directed topology and refin…
▽ More
Multi-agent systems can solve complex tasks through collaboration between multiple Large Language Model agents. Existing collaboration frameworks typically operate in either a parallel or a sequential mode. In the parallel mode, agents respond independently to queries followed by aggregation of responses. In contrast, sequential systems allow agents to communicate via a directed topology and refine one another step by step. However, both modes are inadequate for achieving the desired objectives of minimizing communication and latency while simultaneously maximizing the accuracy of the final response. In this work, we introduce a hybrid paradigm called Nexa, a trainable response-conditioned policy that bridges the gap between the two modes. Nexa begins with a parallel execution stage, embeds the resulting responses into a shared semantic space, and then predicts a sparse directed acyclic communication graph. If the graph is empty, the system remains purely parallel; if it is non-empty, the system performs one sequential message propagation. The policy is a lightweight transformer model, and the method avoids the need for external LLM judges or reward models, as well as hand-crafted test-time topology search. We formalize this hybrid execution problem, show that the resulting graph is acyclic by construction, and that the framework strictly subsumes pure parallel execution, and present a training procedure based on policy-gradient optimization. Results demonstrate that the response-conditioned policy learned by Nexa under one setting can be reused when the number of agents, the task, or the underlying agent changes, thus emphasizing the generalizability of the learned communication policy.
△ Less
Submitted 14 May, 2026;
originally announced May 2026.
-
EdgeFlowerTune: Evaluating Federated LLM Fine-Tuning Under Realistic Edge System Constraints
Authors:
Jiaxiang Geng,
Yiyi Lu,
Lunyu Zhao,
Yan Gao,
Nicholas D. Lane,
Bing Luo
Abstract:
Federated fine-tuning offers a promising paradigm for adapting large language models (LLMs) on edge devices by leveraging the rich, diverse, and continuously generated data from smartphones and IoT devices without compromising user data privacy. Such edge-side adaptation can improve model personalization, robustness, and responsiveness to local contexts. However, the practical feasibility of feder…
▽ More
Federated fine-tuning offers a promising paradigm for adapting large language models (LLMs) on edge devices by leveraging the rich, diverse, and continuously generated data from smartphones and IoT devices without compromising user data privacy. Such edge-side adaptation can improve model personalization, robustness, and responsiveness to local contexts. However, the practical feasibility of federated LLM fine-tuning on real edge devices remains unclear, as most existing work focuses on cross-silo or simulation-based settings, overlooking the resource and runtime constraints that determine whether a method is deployable on real edge systems. We present EdgeFlowerTune, a deployment-oriented benchmark for federated LLM fine-tuning under realistic edge-system constraints. EdgeFlowerTune jointly evaluates model quality and system costs, including communication, wall-clock latency, memory usage, energy consumption, and robustness to dynamic edge conditions. To compare methods in terms of effectiveness, efficiency, and robustness, EdgeFlowerTune introduces three complementary protocols: Quality-under-Budget, Cost-to-Target, and Robustness. We instantiate EdgeFlowerTune as a real-device platform built on Flower and MobileFineTuner, spanning commercial Android smartphones and NVIDIA edge development boards. Our benchmark results show that accuracy-only evaluation can lead to misleading conclusions: methods with similar final quality may differ substantially in deployability once realistic system constraints are considered. EdgeFlowerTune provides a reproducible benchmark for system-aware evaluation of federated LLM fine-tuning at the edge.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.
-
Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods
Authors:
Wanru Zhao,
Yihong Chen,
Yuzhi Tang,
Wentao Ma,
Shengchao Hu,
Shell Xu Hu,
Alex Iacob,
Abhinav Mehrotra,
Nicholas D. Lane
Abstract:
Data curation is a critical yet under-explored area in large language model (LLM) training. Existing methods, such as data selection and mixing, operate in an offline paradigm, detaching themselves from training. This separation introduces engineering overhead and makes the curation brittle: the entire pipeline must be re-run under model/task shifts. Moreover, offline methods alter data size throu…
▽ More
Data curation is a critical yet under-explored area in large language model (LLM) training. Existing methods, such as data selection and mixing, operate in an offline paradigm, detaching themselves from training. This separation introduces engineering overhead and makes the curation brittle: the entire pipeline must be re-run under model/task shifts. Moreover, offline methods alter data size through hard filtering or resampling, often sacrificing data diversity and harming generalization. We propose to rethink data curation as an online reweighting problem, where sample importance is dynamically adjusted during training via loss weighting rather than static pre-processing. Specifically, we introduce ADAPT (Adaptive Data reweighting for Pretraining and FineTuning), a dynamic online framework that reweights training samples with adaptive per-sample learning rates guided by similarity-based quality signals, without changing the number of training samples. Unlike offline methods that enforce a static data distribution, ADAPT acts as an implicit curriculum learner, progressively shifting focus from coarse-grained patterns to fine-grained semantic distinctions as the model evolves. Experiments on both instruction tuning and large-scale pretraining show that ADAPT consistently outperforms offline selection/mixing and prior online methods, achieving stronger cross-benchmark generalization under equal FLOPs.
△ Less
Submitted 19 April, 2026;
originally announced May 2026.
-
Improved muon energy estimation using a detailed model of multiple Coulomb scattering in the MicroBooNE LArTPC
Authors:
MicroBooNE Collaboration,
P. Abratenko,
D. Andrade Aldana,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (167 additional authors not shown)
Abstract:
We present an improved technique for estimating a muon's energy by measuring the deflections along its path inside the MicroBooNE detector from multiple Coulomb scattering (MCS). This approach implements several innovations that better capture detector non-idealizations compared to previous MCS-based muon energy estimators. As a result, it achieves improved resolution, reduced bias, and better dat…
▽ More
We present an improved technique for estimating a muon's energy by measuring the deflections along its path inside the MicroBooNE detector from multiple Coulomb scattering (MCS). This approach implements several innovations that better capture detector non-idealizations compared to previous MCS-based muon energy estimators. As a result, it achieves improved resolution, reduced bias, and better data-model agreement. Using model simulation, for fully contained events the estimated bias is within 1% and the estimated resolution varies from 4.3% to 10% as muon energy increases from 0.1 GeV to 2 GeV. For events with particles exiting the detector volume, at least a meter of reconstructed muon track, and a muon energy below 2 GeV, the estimated bias is less than 2% and the estimated resolution varies from 7% to 17% over muon energy. These demonstrate significant improvements over the performance of previous work using an MCS-based energy estimator at MicroBooNE, which achieves twice as large a resolution as well as a bias of 20% over the same energy region. Data-model goodness-of-fit studies are used to validate the estimator's performance on data, showing good agreement within model uncertainties.
△ Less
Submitted 14 July, 2026; v1 submitted 4 May, 2026;
originally announced May 2026.
-
Charge readout electronics for the DUNE horizontal drift far detector: design and performance in ProtoDUNE-HD
Authors:
DUNE Collaboration,
S. Abbaslu,
F. Abd Alrahman,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
K. Adhikari,
C. Adriano,
K. Agudelo-Jaramillo,
F. Akbar,
F. Alemanno,
N. S. Alex,
L. Aliaga Soplin,
A. Alqaisi,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
R. Amarinei
, et al. (1346 additional authors not shown)
Abstract:
DUNE (Deep Underground Neutrino Experiment) is a long-baseline neutrino oscillation experiment currently under construction, whose far detectors will be the largest liquid argon time projection chambers ever built. This detector design calls for custom-built cryogenic front-end electronics to meet its performance requirements. This paper describes the charge readout electronics that will be used i…
▽ More
DUNE (Deep Underground Neutrino Experiment) is a long-baseline neutrino oscillation experiment currently under construction, whose far detectors will be the largest liquid argon time projection chambers ever built. This detector design calls for custom-built cryogenic front-end electronics to meet its performance requirements. This paper describes the charge readout electronics that will be used in the DUNE horizontal drift (HD) far detector and presents performance results using data from the ProtoDUNE-HD detector, a 770 ton liquid argon time projection chamber operated at the CERN Neutrino Platform in 2024 that served as the final prototype of the DUNE HD design.
△ Less
Submitted 12 August, 2026; v1 submitted 26 April, 2026;
originally announced April 2026.
-
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs
Authors:
Haoran Wu,
Zeyu Cao,
Yao Lai,
Binglei Lou,
Jiayi Nie,
Can Xiao,
Timi Adeniran,
Przemyslaw Forys,
Kauser Johar,
Catriona Wright,
Junyi Liu,
Kai Shi,
Nicholas D. Lane,
Rika Antonova,
Jianyi Cheng,
Timothy Jones,
Aaron Zhao,
Robert Mullins
Abstract:
Emerging agentic LLM workloads are driving rapidly growing demand on both memory capacity and bandwidth, with different phases of inference (e.g., prefill and decode) imposing distinct requirements. Industry is responding by composing heterogeneous accelerators into single interconnected systems, as exemplified by NVIDIA's Vera Rubin platform, where each device brings its own memory architecture.…
▽ More
Emerging agentic LLM workloads are driving rapidly growing demand on both memory capacity and bandwidth, with different phases of inference (e.g., prefill and decode) imposing distinct requirements. Industry is responding by composing heterogeneous accelerators into single interconnected systems, as exemplified by NVIDIA's Vera Rubin platform, where each device brings its own memory architecture.
This heterogeneity is further compounded by a widening landscape of available memory technologies: high-density on-chip SRAM, HBM, LPDDR, GDDR, and emerging options such as high-bandwidth flash (HBF), each offering different capacity, bandwidth, and power trade-offs.
Identifying the right memory architecture for next-generation inference accelerators requires navigating a vast and rapidly evolving design space, in which the interplay between workload characteristics, NPU design dimensions, and memory system design remains largely underexplored.
To address this challenge, we present MemExplorer, a new memory system synthesizer for heterogeneous NPU systems. MemExplorer provides a unified abstraction for modeling diverse memory technologies across different hierarchy levels (e.g., on-chip and off-chip) and automatically determines an efficient heterogeneous memory system together with NPU design choices (e.g., matrix engine size) to balance throughput and power between prefilling and decoding devices in a multi-device NPU system.
Experimental results show that, under the same power budget for agentic workloads, MemExplorer achieves up to 2.3x higher energy efficiency than the baseline NPU and 3.23x higher than H100 in the prefill-only setting. Under equivalent performance targets in the decode setting, it further delivers up to 1.93x and 2.72x higher power efficiency over the baseline NPU and H100, respectively.
△ Less
Submitted 17 April, 2026;
originally announced April 2026.
-
Task-Centric Personalized Federated Fine-Tuning of Language Models
Authors:
Gabriel U. Talasso,
Meghdad Kurmanji,
Allan M. de Souza,
Nicholas D. Lane,
Leandro A. Villas
Abstract:
Federated Learning (FL) has emerged as a promising technique for training language models on distributed and private datasets of diverse tasks. However, aggregating models trained on heterogeneous tasks often degrades the overall performance of individual clients. To address this issue, Personalized FL (pFL) aims to create models tailored for each client's data distribution. Although these approac…
▽ More
Federated Learning (FL) has emerged as a promising technique for training language models on distributed and private datasets of diverse tasks. However, aggregating models trained on heterogeneous tasks often degrades the overall performance of individual clients. To address this issue, Personalized FL (pFL) aims to create models tailored for each client's data distribution. Although these approaches improve local performance, they usually lack robustness in two aspects: (i) generalization: when clients must make predictions on unseen tasks, or face changes in their data distributions, and (ii) intra-client tasks interference: when a single client's data contains multiple distributions that may interfere with each other during local training. To tackle these two challenges, we propose FedRouter, a clustering-based pFL that builds specialized models for each task rather than for each client. FedRouter uses adapters to personalize models by employing two clustering mechanisms to associate adapters with specific tasks. A local clustering that associate adapters with task data samples and a global one that associates similar adapters from different clients to construct task-centric personalized models. Additionally, we propose an evaluation router mechanism that routes test samples to the best adapter based on the created clusters. Experiments comparing our method with existing approaches across a multitask dataset, FedRouter demonstrate strong resilience in these challenging scenarios performing up to 6.1% relatively better under tasks interference and up to 136% relative improvement under generalization evaluation.
△ Less
Submitted 5 April, 2026; v1 submitted 30 March, 2026;
originally announced April 2026.
-
Supercharging Federated Intelligence Retrieval
Authors:
Dimitris Stripelis,
Patrick Foley,
Mohammad Naseri,
William Lindskog-Münzing,
Chong Shen Ng,
Daniel Janes Beutel,
Nicholas D. Lane
Abstract:
RAG typically assumes centralized access to documents, which breaks down when knowledge is distributed across private data silos. We propose a secure Federated RAG system built using Flower that performs local silo retrieval, while server-side aggregation and text generation run inside an attested, confidential compute environment, enabling confidential remote LLM inference even in the presence of…
▽ More
RAG typically assumes centralized access to documents, which breaks down when knowledge is distributed across private data silos. We propose a secure Federated RAG system built using Flower that performs local silo retrieval, while server-side aggregation and text generation run inside an attested, confidential compute environment, enabling confidential remote LLM inference even in the presence of honest-but-curious or compromised servers. We also propose a cascading inference approach that incorporates a non-confidential third-party model (e.g., Amazon Nova) as auxiliary context without weakening confidentiality.
△ Less
Submitted 26 March, 2026;
originally announced March 2026.
-
Scintillation light calibrations, systematic uncertainties, and triggering efficiency in the MicroBooNE detector
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
L. Arellano,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton
, et al. (169 additional authors not shown)
Abstract:
Scintillation light, produced alongside ionisation charge from particle interactions, plays a critical role in liquid argon time projection chamber (LArTPC) detectors. A detailed understanding of its production and detection mechanisms is essential for robust calibration, systematic uncertainty evaluation, and physics analysis. This article describes the MicroBooNE light simulation, light-based tr…
▽ More
Scintillation light, produced alongside ionisation charge from particle interactions, plays a critical role in liquid argon time projection chamber (LArTPC) detectors. A detailed understanding of its production and detection mechanisms is essential for robust calibration, systematic uncertainty evaluation, and physics analysis. This article describes the MicroBooNE light simulation, light-based triggering schemes, photomultiplier tube gain calibration, light response stability, and light-based systematic uncertainties over the course of five years of data collection. In addition, we present a measurement of scintillation light triggering efficiency, focusing on the lowest-light regime relevant to rare-event searches and low-energy neutrino interactions. Finally, we discuss two notable observations in MicroBooNE's data, both reported here for the first time: an approximately 50% decline in MicroBooNE's light yield over time, concentrated in the first two years of running; and a higher than expected O(200 kHz) rate of single photoelectron noise. The results presented provide an important benchmark of long-term light detection performance in LArTPC neutrino detectors.
△ Less
Submitted 24 March, 2026;
originally announced March 2026.
-
Measurements of the electron neutrino-argon differential cross section without pions in the final state in MicroBooNE
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
L. Arellano,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton
, et al. (168 additional authors not shown)
Abstract:
We present a new measurement of the electron neutrino charged current cross section on argon without pions in the final state. This measurement uses the full MicroBooNE booster neutrino beam dataset of $1.3\times 10^{21}$ protons on target collected at Fermi National Accelerator Laboratory. Events are considered both with and without protons above the kinetic energy visibility threshold. Different…
▽ More
We present a new measurement of the electron neutrino charged current cross section on argon without pions in the final state. This measurement uses the full MicroBooNE booster neutrino beam dataset of $1.3\times 10^{21}$ protons on target collected at Fermi National Accelerator Laboratory. Events are considered both with and without protons above the kinetic energy visibility threshold. Differential cross sections are extracted in proton and electron kinematics, including energy and angle relative to the neutrino beam direction. The relationship between the hadronic and leptonic systems is explored through the angle between the proton and electron directions. The resulting cross sections are compared to a variety of available generator predictions using different models of neutrino interactions. We find good agreement with most models in lepton kinematics and some discrepancies in the hadronic system modeling, particularly in proton angle.
△ Less
Submitted 14 August, 2026; v1 submitted 13 March, 2026;
originally announced March 2026.
-
Building Privacy-and-Security-Focused Federated Learning Infrastructure for Global Multi-Centre Healthcare Research
Authors:
Fan Zhang,
Daniel Kreuter,
Javier Fernandez-Marques,
BloodCounts Consortium,
Gregory Verghese,
Bernard Butler,
Nicholas Lane,
Suthesh Sivapalaratnam,
Joseph Taylor,
Norbert C. J. de Wit,
Nicholas S. Gleadall,
Carola-Bibiane Schönlieb,
Michael Roberts
Abstract:
Collaborative healthcare research across multiple institutions increasingly requires diverse clinical datasets, but cross-border data sharing is strictly constrained by privacy regulations. Federated learning (FL) enables model training while keeping data local; however, many existing frameworks remain proof-of-concept and do not adequately address governance risks such as unauthorised participati…
▽ More
Collaborative healthcare research across multiple institutions increasingly requires diverse clinical datasets, but cross-border data sharing is strictly constrained by privacy regulations. Federated learning (FL) enables model training while keeping data local; however, many existing frameworks remain proof-of-concept and do not adequately address governance risks such as unauthorised participation, misuse, and lack of accountability. In particular, enforceable mechanisms for authentication, authorisation, and accounting (AAA) are often missing, limiting real-world clinical deployment. This paper presents FLA$^3$ (Federated Learning with Authentication, Authorisation, and Accounting), a governance-aware federated learning platform that operationalises regulatory obligations through runtime policy enforcement. FLA$^3$ integrates eXtensible Access Control Markup Language (XACML) compliant attribute-based access control (ABAC), cryptographic accounting, and study-scoped federation directly into the federated learning orchestration layer to enforce institutional sovereignty and protocol adherence. We evaluate FLA$^3$ through two complementary studies. First, we demonstrate operational feasibility by deploying the platform infrastructure across five BloodCounts! Consortium institutions in four countries: United Kingdom, Netherlands, India, and The Gambia. Second, we assess clinical utility using simulated federation of full blood count (FBC) data from 54,446 samples from 35,315 subjects across 25 centres in the INTERVAL study. Results show that FLA$^3$ achieves predictive performance comparable to centralised training while strictly enforcing governance constraints. These results show that enforceable governance can function as a first-class privacy-preserving control, improving trustworthiness for scalable artificial intelligence (AI) in cross-jurisdictional healthcare deployments.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
Improving Generalizability of Hip Fracture Risk Prediction via Domain Adaptation Across Multiple Cohorts
Authors:
Shuo Sun,
Meiling Zhou,
Chen Zhao,
Joyce H. Keyak,
Nancy E. Lane,
Jeffrey D. Deng,
Kuan-Jui Su,
Hui Shen,
Hong-Wen Deng,
Kui Zhang,
Weihua Zhou
Abstract:
Clinical risk prediction models often fail to be generalized across cohorts because underlying data distributions differ by clinical site, region, demographics, and measurement protocols. This limitation is particularly pronounced in hip fracture risk prediction, where the performance of models trained on one cohort (the source cohort) can degrade substantially when deployed in other cohorts (targ…
▽ More
Clinical risk prediction models often fail to be generalized across cohorts because underlying data distributions differ by clinical site, region, demographics, and measurement protocols. This limitation is particularly pronounced in hip fracture risk prediction, where the performance of models trained on one cohort (the source cohort) can degrade substantially when deployed in other cohorts (target cohorts). We used a shared set of clinical and DXA-derived features across three large cohorts - the Study of Osteoporotic Fractures (SOF), the Osteoporotic Fractures in Men Study (MrOS), and the UK Biobank (UKB), to systematically evaluate the performance of three domain adaptation methods - Maximum Mean Discrepancy (MMD), Correlation Alignment (CORAL), and Domain - Adversarial Neural Networks (DANN) and their combinations. For a source cohort with males only and a source cohort with females only, domain-adaptation methods consistently showed improved performance than the no-adaptation baseline (source-only training), and the use of combinations of multiple domain adaptation methods delivered the largest and most stable gains. The method that combines MMD, CORAL, and DANN achieved the highest discrimination with the area under curve (AUC) of 0.88 for a source cohort with males only and 0.95 for a source cohort with females only), demonstrating that integrating multiple domain adaptation methods could produce feature representations that are less sensitive to dataset differences. Unlike existing methods that rely heavily on supervised tuning or assume known outcomes of samples in target cohorts, our outcome-free approaches enable the model selection under realistic deployment conditions and improve generalization of models in hip fracture risk prediction.
△ Less
Submitted 19 February, 2026;
originally announced February 2026.
-
Floe: Federated Specialization for Real-Time LLM-SLM Inference
Authors:
Chunlin Tian,
Kahou Tam,
Yebo Wu,
Shuaihang Zhong,
Li Li,
Nicholas D. Lane,
Chengzhong Xu
Abstract:
Deploying large language models (LLMs) in real-time systems remains challenging due to their substantial computational demands and privacy concerns. We propose Floe, a hybrid federated learning framework designed for latency-sensitive, resource-constrained environments. Floe combines a cloud-based black-box LLM with lightweight small language models (SLMs) on edge devices to enable low-latency, pr…
▽ More
Deploying large language models (LLMs) in real-time systems remains challenging due to their substantial computational demands and privacy concerns. We propose Floe, a hybrid federated learning framework designed for latency-sensitive, resource-constrained environments. Floe combines a cloud-based black-box LLM with lightweight small language models (SLMs) on edge devices to enable low-latency, privacy-preserving inference. Personal data and fine-tuning remain on-device, while the cloud LLM contributes general knowledge without exposing proprietary weights. A heterogeneity-aware LoRA adaptation strategy enables efficient edge deployment across diverse hardware, and a logit-level fusion mechanism enables real-time coordination between edge and cloud models. Extensive experiments demonstrate that Floe enhances user privacy and personalization. Moreover, it significantly improves model performance and reduces inference latency on edge devices under real-time constraints compared with baseline approaches.
△ Less
Submitted 15 February, 2026;
originally announced February 2026.
-
Demonstration and performance of an online data selection algorithm for liquid argon time projection chambers using MicroBooNE
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
L. Arellano,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton
, et al. (169 additional authors not shown)
Abstract:
The MicroBooNE detector is a liquid argon time projection chamber (LArTPC) that produces three-dimensional images of particle interactions using ionization charge collected by anode wire plane arrays and scintillation light collected by a light detection system. In addition to testing long-standing experimental neutrino anomalies and performing measurements of neutrino interactions with argon nucl…
▽ More
The MicroBooNE detector is a liquid argon time projection chamber (LArTPC) that produces three-dimensional images of particle interactions using ionization charge collected by anode wire plane arrays and scintillation light collected by a light detection system. In addition to testing long-standing experimental neutrino anomalies and performing measurements of neutrino interactions with argon nuclei using the Fermilab Booster Neutrino Beam, MicroBooNE aims to develop methodologies for rare beyond the Standard Model and off-beam physics searches. Looking ahead to the upcoming Deep Underground Neutrino Experiment (DUNE), with MicroBooNE serving as a valuable testbed, achieving high sensitivity and livetime for off-beam physics while satisfying data processing and storage constraints will require data-driven, intelligent, and online or real-time data selection techniques. These techniques are essential for reducing data rates and preserving rare signals with high accuracy. In this paper, we describe a fast data selection algorithm suitable for online execution to identify electrons from stopping cosmic ray muons in the MicroBooNE detector utilizing ionization charge information, and present its performance. This represents the first demonstration of online data selection in a LArTPC using real data and charge information exclusively and provides an important proof-of-principle for applying such techniques to other LArTPC experiments such as the Short-Baseline Near Detector and DUNE.
△ Less
Submitted 10 July, 2026; v1 submitted 11 February, 2026;
originally announced February 2026.
-
$f$-FUM: Federated Unlearning via min--max and $f$-divergence
Authors:
Radmehr Karimian,
Amirhossein Bagheri,
Meghdad Kurmanji,
Nicholas D. Lane,
Gholamali Aminian
Abstract:
Federated Learning (FL) has emerged as a powerful paradigm for collaborative machine learning across decentralized data sources, preserving privacy by keeping data local. However, increasing legal and ethical demands, such as the "right to be forgotten", and the need to mitigate data poisoning attacks have underscored the urgent necessity for principled data unlearning in FL. Unlike centralized se…
▽ More
Federated Learning (FL) has emerged as a powerful paradigm for collaborative machine learning across decentralized data sources, preserving privacy by keeping data local. However, increasing legal and ethical demands, such as the "right to be forgotten", and the need to mitigate data poisoning attacks have underscored the urgent necessity for principled data unlearning in FL. Unlike centralized settings, the distributed nature of FL complicates the removal of individual data contributions. In this paper, we propose a novel federated unlearning framework formulated as a min-max optimization problem, where the objective is to maximize an $f$-divergence between the model trained with all data and the model retrained without specific data points, while minimizing the degradation on retained data. Our framework could act like a plugin and be added to almost any federated setup, unlike SOTA methods like (\cite{10269017} which requires model degradation in server, or \cite{khalil2025notfederatedunlearningweight} which requires to involve model architecture and model weights). This formulation allows for efficient approximation of data removal effects in a federated setting. We provide empirical evaluations to show that our method achieves significant speedups over naive retraining, with minimal impact on utility.
△ Less
Submitted 5 February, 2026;
originally announced February 2026.
-
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
Authors:
Andrej Jovanović,
Alex Iacob,
Mher Safaryan,
Ionut-Vlad Modoranu,
Lorenzo Sani,
William F. Shen,
Xinchi Qiu,
Dan Alistarh,
Nicholas D. Lane
Abstract:
Distributed training of foundation models via $\texttt{DDP}$ is limited by interconnect bandwidth. While infrequent communication strategies reduce synchronization frequency, they remain bottlenecked by the memory and communication requirements of optimizer states. Low-rank optimizers can alleviate these constraints; however, in the local-update regime, workers lack access to the full-batch gradie…
▽ More
Distributed training of foundation models via $\texttt{DDP}$ is limited by interconnect bandwidth. While infrequent communication strategies reduce synchronization frequency, they remain bottlenecked by the memory and communication requirements of optimizer states. Low-rank optimizers can alleviate these constraints; however, in the local-update regime, workers lack access to the full-batch gradients required to compute low-rank projections, which degrades performance. We propose $\texttt{LoRDO}$, a principled framework unifying low-rank optimization with infrequent synchronization. We first demonstrate that, while global projections based on pseudo-gradients are theoretically superior, they permanently restrict the optimization trajectory to a low-rank subspace. To restore subspace exploration, we introduce a full-rank quasi-hyperbolic update. $\texttt{LoRDO}$ achieves near-parity with low-rank $\texttt{DDP}$ in language modeling and downstream tasks at model scales of $125$M--$720$M, while reducing communication by $\approx 10 \times$. Finally, we show that $\texttt{LoRDO}$ improves performance even more in very low-memory settings with small rank/batch size.
△ Less
Submitted 17 June, 2026; v1 submitted 4 February, 2026;
originally announced February 2026.
-
Reconstruction of atmospheric neutrinos in DUNE's horizontal-drift far-detector module
Authors:
DUNE Collaboration,
S. Abbaslu,
F. Abd Alrahman,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
C. Adriano,
F. Akbar,
F. Alemanno,
N. S. Alex,
K. Allison,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
P. Amedo,
J. Anderson,
D. A. Andrade,
C. Andreopoulos
, et al. (1325 additional authors not shown)
Abstract:
This paper reports on the capabilities in reconstructing and identifying atmospheric neutrino interactions in one of the Deep Underground Neutrino Experiment's (DUNE) far detector modules, a liquid argon time projection chamber (LArTPC) with horizontal drift (FD-HD) of ionization electrons. The reconstruction is based upon the workflow developed for DUNE's long-baseline oscillation analysis, with…
▽ More
This paper reports on the capabilities in reconstructing and identifying atmospheric neutrino interactions in one of the Deep Underground Neutrino Experiment's (DUNE) far detector modules, a liquid argon time projection chamber (LArTPC) with horizontal drift (FD-HD) of ionization electrons. The reconstruction is based upon the workflow developed for DUNE's long-baseline oscillation analysis, with some necessary machine-learning models' retraining and the addition of features relevant only to atmospheric neutrinos such as the neutrino direction reconstruction. Where relevant, the impact of the detection of the charged particles of the hadronic system is emphasized, and comparisons are carried out between the case when lepton-only information is considered in the reconstruction (as is the case for many neutrino oscillation experiments), versus when all particles identified in the LArTPC were included. Three neutrino direction reconstruction methods have been developed and studied for the atmospheric analyses: using lepton-only information, using all reconstructed particles, and using only correlations from reconstructed hits. The results indicate that incorporating more than just lepton information significantly improves the resolution of both neutrino direction and energy reconstruction. The angle reconstruction algorithms developed in this work result in no strong dependence on particle direction for reconstruction efficiencies or neutrino flavor identification. This comprehensive review of the reconstruction of atmospheric neutrinos in DUNE's FD-HD LArTPC is the first step towards developing a first neutrino oscillation sensitivity analysis, which will ready DUNE for its first measurements.
△ Less
Submitted 9 January, 2026;
originally announced January 2026.
-
Computational Compliance for AI Regulation: Blueprint for a New Research Domain
Authors:
Bill Marino,
Nicholas D. Lane
Abstract:
The era of AI regulation (AIR) is upon us. But AI systems, we argue, will not be able to comply with these regulations at the necessary speed and scale by continuing to rely on traditional, analogue methods of compliance. Instead, we posit that compliance with these regulations will only realistically be achieved computationally: that is, with algorithms that run across the life cycle of an AI sys…
▽ More
The era of AI regulation (AIR) is upon us. But AI systems, we argue, will not be able to comply with these regulations at the necessary speed and scale by continuing to rely on traditional, analogue methods of compliance. Instead, we posit that compliance with these regulations will only realistically be achieved computationally: that is, with algorithms that run across the life cycle of an AI system, automatically steering it toward AIR compliance in the face of dynamic conditions. Yet despite their (we would argue) inevitability, the research community has yet to specify exactly how these algorithms for computational AIR compliance should behave - or how we should benchmark their performance. To fill these gaps, we specify a set of design goals for such algorithms. In addition, we specify a benchmark dataset that can be used to quantitatively measure whether individual algorithms satisfy these design goals. By delivering this blueprint, we hope to give shape to an important but uncrystallized new domain of research - and, in doing so, incite necessary investment in it.
△ Less
Submitted 7 January, 2026;
originally announced January 2026.
-
Cosmic Ray Measurements Using Charge and Light Readout in a Pixelated Liquid Argon Time Projection Chamber
Authors:
SoLAr Collaboration,
N. Anfimov,
A. Branca,
J. Bürgi,
L. Calivers,
P. Carniti,
E. Calvo,
E. Cristaldo,
C. Cuesta,
F. Declich,
R. Diurba,
P. Dunne,
D. A. Dwyer,
J. Evans,
A. C. Ezeribe,
A. Gauch,
I. Gil-Botella,
C. Gotti,
S. Greenberg,
D. Guffanti,
A. Karcher,
J. Kunzmann,
N. Lane,
S. Manthey Corchado,
N. McConkey
, et al. (18 additional authors not shown)
Abstract:
Liquid argon time projection chambers have emerged as a competitive technology for detecting solar neutrinos. The SoLAr collaboration was formed to explore argon detectors with pixelated light and charge readout, aiming for high detection efficiency and improved energy resolution. Building on the success of an initial prototype, we present results obtained with a second SoLAr prototype (V2), a…
▽ More
Liquid argon time projection chambers have emerged as a competitive technology for detecting solar neutrinos. The SoLAr collaboration was formed to explore argon detectors with pixelated light and charge readout, aiming for high detection efficiency and improved energy resolution. Building on the success of an initial prototype, we present results obtained with a second SoLAr prototype (V2), a $30 \times 30 \times 30$ cm$^{3}$ time projection chamber operated in a cryostat containing several hundred kilograms of liquid argon. We report measurements of cosmic-ray muons using both tracking and calorimetry from light and charge sensors, and we highlight the improved performance achieved through combined charge and light reconstruction. These results demonstrate the promise of dual-readout detectors and motivate future prototyping efforts toward kiloton-scale facilities.
△ Less
Submitted 11 December, 2025;
originally announced December 2025.
-
Search for Light Sterile Neutrinos With Two Neutrino Beams at MicroBooNE
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
L. Arellano,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti,
L. Camilleri,
D. Caratelli
, et al. (154 additional authors not shown)
Abstract:
The existence of three distinct neutrino flavours, $ν_{e}$, $ν_μ$, and $ν_τ$, is a central tenet of the Standard Model of particle physics. Quantum-mechanical interference can allow a neutrino of one initial flavour to be detected some time later as a different flavour, a process called neutrino oscillation. Several anomalous observations inconsistent with this three-flavour picture have motivated…
▽ More
The existence of three distinct neutrino flavours, $ν_{e}$, $ν_μ$, and $ν_τ$, is a central tenet of the Standard Model of particle physics. Quantum-mechanical interference can allow a neutrino of one initial flavour to be detected some time later as a different flavour, a process called neutrino oscillation. Several anomalous observations inconsistent with this three-flavour picture have motivated the hypothesis that an additional neutrino state exists which does not interact directly with matter, termed a "sterile" neutrino, $ν_s$. This includes anomalous observations from the LSND and MiniBooNE experiments, consistent with $ν_μ\rightarrowν_{e}$ transitions at a distance inconsistent with the three-neutrino picture. Here, we use data obtained from the MicroBooNE liquid-argon time projection chamber in two accelerator neutrino beams to exclude the single light sterile neutrino interpretation of the LSND and MiniBooNE anomalies at the 95\% confidence level (CL). Additionally, we rule out a significant portion of the parameter space that could explain the gallium anomaly. This is the first measurement to use two accelerator neutrino beams to break a degeneracy between $ν_{e}$ appearance and disappearance that would otherwise weaken the sensitivity to the sterile neutrino hypothesis. We find no evidence for either $ν_μ\rightarrowν_{e}$ flavour transitions or $ν_{e}$ disappearance that would indicate non-standard flavour oscillations. Our results show that previous anomalous observations consistent with $ν_μ\rightarrowν_{e}$ transitions cannot be explained by introducing a single sterile neutrino state.
△ Less
Submitted 7 December, 2025;
originally announced December 2025.
-
FLUX: Efficient Descriptor-Driven Clustered Federated Learning under Arbitrary Distribution Shifts
Authors:
Dario Fenoglio,
Mohan Li,
Pietro Barbiero,
Nicholas D. Lane,
Marc Langheinrich,
Martin Gjoreski
Abstract:
Federated Learning (FL) enables collaborative model training across multiple clients while preserving data privacy. Traditional FL methods often use a global model to fit all clients, assuming that clients' data are independent and identically distributed (IID). However, when this assumption does not hold, the global model accuracy may drop significantly, limiting FL applicability in real-world sc…
▽ More
Federated Learning (FL) enables collaborative model training across multiple clients while preserving data privacy. Traditional FL methods often use a global model to fit all clients, assuming that clients' data are independent and identically distributed (IID). However, when this assumption does not hold, the global model accuracy may drop significantly, limiting FL applicability in real-world scenarios. To address this gap, we propose FLUX, a novel clustering-based FL (CFL) framework that addresses the four most common types of distribution shifts during both training and test time. To this end, FLUX leverages privacy-preserving client-side descriptor extraction and unsupervised clustering to ensure robust performance and scalability across varying levels and types of distribution shifts. Unlike existing CFL methods addressing non-IID client distribution shifts, FLUX i) does not require any prior knowledge of the types of distribution shifts or the number of client clusters, and ii) supports test-time adaptation, enabling unseen and unlabeled clients to benefit from the most suitable cluster-specific models. Extensive experiments across four standard benchmarks, two real-world datasets and ten state-of-the-art baselines show that FLUX improves performance and stability under diverse distribution shifts, achieving an average accuracy gain of up to 23 percentage points over the best-performing baselines, while maintaining computational and communication overhead comparable to FedAvg.
△ Less
Submitted 27 November, 2025;
originally announced November 2025.
-
Measurements of differential charged-current cross sections on argon for electron neutrinos with final-state protons in MicroBooNE
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
L. Arellano,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (156 additional authors not shown)
Abstract:
This work presents single-differential electron-neutrino charged-current cross sections on argon measured using the MicroBooNE detector at the Fermi National Accelerator Laboratory. The analysis uses data recorded when the Neutrinos at the Main Injector beam was operating in both neutrino and antineutrino modes, with exposures of $2 \times 10^{20}$ and $5 \times 10^{20}$ protons on target, respect…
▽ More
This work presents single-differential electron-neutrino charged-current cross sections on argon measured using the MicroBooNE detector at the Fermi National Accelerator Laboratory. The analysis uses data recorded when the Neutrinos at the Main Injector beam was operating in both neutrino and antineutrino modes, with exposures of $2 \times 10^{20}$ and $5 \times 10^{20}$ protons on target, respectively. A selection algorithm targeting electron-neutrino charged-current interactions with at least one proton, one electron, and no pions in the final topology is used to measure differential cross sections as a function of outgoing electron energy, total visible energy, and opening angle between the electron and the most energetic proton. The interaction rate as a function of proton multiplicity is also reported. The total cross section is measured as [4.1 $\pm$ 0.3 (stat.) $\pm$ 1.1 (syst.)]$ $$\times 10^{-39} \mathrm{cm}^{2}/ \mathrm{nucleon}$. The unfolded cross-section measurements are compared to predictions from neutrino event generators commonly employed in the field. Good agreement is seen across all variables within uncertainties.
△ Less
Submitted 5 March, 2026; v1 submitted 21 November, 2025;
originally announced November 2025.
-
Bringing Federated Learning to Space
Authors:
Grace Kim,
Filip Svoboda,
Nicholas Lane
Abstract:
As Low Earth Orbit (LEO) satellite constellations rapidly expand to hundreds and thousands of spacecraft, the need for distributed on-board machine learning becomes critical to address downlink bandwidth limitations. Federated learning (FL) offers a promising framework to conduct collaborative model training across satellite networks. Realizing its benefits in space naturally requires addressing s…
▽ More
As Low Earth Orbit (LEO) satellite constellations rapidly expand to hundreds and thousands of spacecraft, the need for distributed on-board machine learning becomes critical to address downlink bandwidth limitations. Federated learning (FL) offers a promising framework to conduct collaborative model training across satellite networks. Realizing its benefits in space naturally requires addressing space-specific constraints, from intermittent connectivity to dynamics imposed by orbital motion. This work presents the first systematic feasibility analysis of adapting off-the-shelf FL algorithms for satellite constellation deployment. We introduce a comprehensive "space-ification" framework that adapts terrestrial algorithms (FedAvg, FedProx, FedBuff) to operate under orbital constraints, producing an orbital-ready suite of FL algorithms. We then evaluate these space-ified methods through extensive parameter sweeps across 768 constellation configurations that vary cluster sizes (1-10), satellites per cluster (1-10), and ground station networks (1-13). Our analysis demonstrates that space-adapted FL algorithms efficiently scale to constellations of up to 100 satellites, achieving performance close to the centralized ideal. Multi-month training cycles can be reduced to days, corresponding to a 9x speedup through orbital scheduling and local coordination within satellite clusters. These results provide actionable insights for future mission designers, enabling distributed on-board learning for more autonomous, resilient, and data-driven satellite operations.
△ Less
Submitted 18 November, 2025;
originally announced November 2025.
-
Measurement of Exclusive $π^+$--argon Interactions Using ProtoDUNE-SP
Authors:
DUNE Collaboration,
S. Abbaslu,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
C. Adriano,
F. Akbar,
F. Alemanno,
N. S. Alex,
K. Allison,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
P. Amedo,
J. Anderson,
D. A. Andrade,
C. Andreopoulos,
M. Andreotti
, et al. (1304 additional authors not shown)
Abstract:
We present the measurement of $π^{+}$--argon inelastic cross sections using the ProtoDUNE Single-Phase liquid argon time projection chamber in the incident $π^+$ kinetic energy range of 500 -- 800 MeV in multiple exclusive channels (absorption, charge exchange, and the remaining inelastic interactions). The results of this analysis are important inputs to simulations of liquid argon neutrino exper…
▽ More
We present the measurement of $π^{+}$--argon inelastic cross sections using the ProtoDUNE Single-Phase liquid argon time projection chamber in the incident $π^+$ kinetic energy range of 500 -- 800 MeV in multiple exclusive channels (absorption, charge exchange, and the remaining inelastic interactions). The results of this analysis are important inputs to simulations of liquid argon neutrino experiments such as the Deep Underground Neutrino Experiment and the Short Baseline Neutrino program at Fermi National Accelerator Laboratory. They will be employed to improve the modeling of final state interactions within neutrino event generators used by these experiments, as well as the modeling of $π^{+}$--argon secondary interactions within the liquid argon. This is the first measurement of $π^+$--argon absorption at this kinetic energy range as well as the first ever measurement of $π^{+}$--argon charge exchange.
△ Less
Submitted 17 November, 2025;
originally announced November 2025.
-
First Measurement of $π^+$-Ar and $p$-Ar Total Inelastic Cross Sections in the Sub-GeV Energy Regime with ProtoDUNE-SP Data
Authors:
DUNE Collaboration,
S. Abbaslu,
F. Abd Alrahman,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
C. Adriano,
F. Akbar,
F. Alemanno,
N. S. Alex,
L. Aliaga Soplin,
K. Allison,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
P. Amedo,
J. Anderson,
D. A. Andrade
, et al. (1328 additional authors not shown)
Abstract:
The ProtoDUNE-SP detector, a kiloton-scale prototype for the Deep Underground Neutrino Experiment (DUNE) far detector, is the largest liquid argon time projection chamber built to date. Operated at CERN from 2018 to 2020, it collected both cosmic-ray data and a beam consisting of positively-charged particles with discrete momentum settings across a range of 0.3 GeV/$c$ to 7 GeV/$c$. In this letter…
▽ More
The ProtoDUNE-SP detector, a kiloton-scale prototype for the Deep Underground Neutrino Experiment (DUNE) far detector, is the largest liquid argon time projection chamber built to date. Operated at CERN from 2018 to 2020, it collected both cosmic-ray data and a beam consisting of positively-charged particles with discrete momentum settings across a range of 0.3 GeV/$c$ to 7 GeV/$c$. In this letter, we report the total inelastic cross section measurements for $π^+$--Ar and $p$--Ar interactions using selected $π^+$ and proton samples from the 1 GeV/$c$ beam data, spanning kinetic energies of 500--900~MeV and below 450~MeV, respectively. These energy ranges are directly relevant to hadrons produced in DUNE. The measured cross sections are consistent with predictions and provide a dataset that was previously unavailable for argon targets. These measurements are essential for constraining neutrino-argon interaction models and achieving the precision physics goals of the upcoming DUNE experiment.
△ Less
Submitted 26 May, 2026; v1 submitted 14 November, 2025;
originally announced November 2025.
-
An Advanced Two-Stage Model with High Sensitivity and Generalizability for Prediction of Hip Fracture Risk Using Multiple Datasets
Authors:
Shuo Sun,
Meiling Zhou,
Chen Zhao,
Joyce H. Keyak,
Nancy E. Lane,
Jeffrey D. Deng,
Kuan-Jui Su,
Hui Shen,
Hong-Wen Deng,
Kui Zhang,
Weihua Zhou
Abstract:
Hip fractures are a major cause of disability, mortality, and healthcare burden in older adults, underscoring the need for early risk assessment. However, commonly used tools such as the DXA T-score and FRAX often lack sensitivity and miss individuals at high risk, particularly those without prior fractures or with osteopenia. To address this limitation, we propose a sequential two-stage model tha…
▽ More
Hip fractures are a major cause of disability, mortality, and healthcare burden in older adults, underscoring the need for early risk assessment. However, commonly used tools such as the DXA T-score and FRAX often lack sensitivity and miss individuals at high risk, particularly those without prior fractures or with osteopenia. To address this limitation, we propose a sequential two-stage model that integrates clinical and imaging information to improve prediction accuracy. Using data from the Osteoporotic Fractures in Men Study (MrOS), the Study of Osteoporotic Fractures (SOF), and the UK Biobank, Stage 1 (Screening) employs clinical, demographic, and functional variables to estimate baseline risk, while Stage 2 (Imaging) incorporates DXA-derived features for refinement. The model was rigorously validated through internal and external testing, showing consistent performance and adaptability across cohorts. Compared to T-score and FRAX, the two-stage framework achieved higher sensitivity and reduced missed cases, offering a cost-effective and personalized approach for early hip fracture risk assessment.
Keywords: Hip Fracture, Two-Stage Model, Risk Prediction, Sensitivity, DXA, FRAX
△ Less
Submitted 16 October, 2025;
originally announced October 2025.
-
Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
Authors:
Sarina Xi,
Orelia Pi,
Miaomiao Zhang,
Becca Xiong,
Jacqueline Ng Lane,
Nihar B. Shah
Abstract:
There is growing interest in applying artificial intelligence (AI) to automate and support complex decision-making tasks. However, it remains unclear how algorithms compare to human judgment in contexts requiring semantic understanding and domain expertise. We examine this in the context of the judge assignment problem, matching submissions to suitably qualified judges. Specifically, we tackled th…
▽ More
There is growing interest in applying artificial intelligence (AI) to automate and support complex decision-making tasks. However, it remains unclear how algorithms compare to human judgment in contexts requiring semantic understanding and domain expertise. We examine this in the context of the judge assignment problem, matching submissions to suitably qualified judges. Specifically, we tackled this problem at the Harvard President's Innovation Challenge, the university's premier venture competition awarding over \$500,000 to student and alumni startups. This represents a real-world environment where high-quality judge assignment is essential. We developed an AI-based judge-assignment algorithm, Hybrid Lexical-Semantic Similarity Ensemble (HLSE), and deployed it at the competition. We then evaluated its performance against human expert assignments using blinded match-quality scores from judges on $309$ judge-venture pairs. Using a Mann-Whitney U statistic based test, we found no statistically significant difference in assignment quality between the two approaches ($AUC=0.48, p=0.40$); on average, algorithmic matches are rated $3.90$ and manual matches $3.94$ on a 5-point scale, where 5 indicates an excellent match. Furthermore, manual assignments that previously required a full week could be automated in several hours by the algorithm during deployment. These results demonstrate that HLSE achieves human-expert-level matching quality while offering greater scalability and efficiency, underscoring the potential of AI-driven solutions to support and enhance human decision-making for judge assignment in high-stakes settings.
△ Less
Submitted 14 October, 2025;
originally announced October 2025.
-
Identification of low-energy kaons in the ProtoDUNE-SP detector
Authors:
DUNE Collaboration,
S. Abbaslu,
F. Abd Alrahman,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
C. Adriano,
F. Akbar,
F. Alemanno,
N. S. Alex,
K. Allison,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
P. Amedo,
J. Anderson,
D. A. Andrade,
C. Andreopoulos
, et al. (1325 additional authors not shown)
Abstract:
The Deep Underground Neutrino Experiment (DUNE) is a next-generation neutrino experiment with a rich physics program that includes searches for the hypothetical phenomenon of proton decay. Utilizing liquid-argon time-projection chamber technology, DUNE is expected to achieve world-leading sensitivity in the proton decay channels that involve charged kaons in their final states. The first DUNE demo…
▽ More
The Deep Underground Neutrino Experiment (DUNE) is a next-generation neutrino experiment with a rich physics program that includes searches for the hypothetical phenomenon of proton decay. Utilizing liquid-argon time-projection chamber technology, DUNE is expected to achieve world-leading sensitivity in the proton decay channels that involve charged kaons in their final states. The first DUNE demonstrator, ProtoDUNE Single-Phase, was a 0.77 kt detector that operated from 2018 to 2020 at the CERN Neutrino Platform, exposed to a mixed hadron and electron test-beam with momenta ranging from 0.3 to 7 GeV/c. We present a selection of low-energy kaons among the secondary particles produced in hadronic reactions, using data from the 6 and 7 GeV/c beam runs. The selection efficiency is 1\% and the sample purity 92\%. The initial energies of the selected kaon candidates encompass the expected energy range of kaons originating from proton decay events in DUNE (below $\sim$200 MeV). In addition, we demonstrate the capability of this detector technology to discriminate between kaons and other particles such as protons and muons, and provide a comprehensive description of their energy loss in liquid argon, which shows good agreement with the simulation. These results pave the way for future proton decay searches at DUNE.
△ Less
Submitted 9 October, 2025;
originally announced October 2025.
-
MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
Authors:
Alex Iacob,
Andrej Jovanovic,
Mher Safaryan,
Meghdad Kurmanji,
Lorenzo Sani,
Samuel Horváth,
William F. Shen,
Xinchi Qiu,
Nicholas D. Lane
Abstract:
Training large models with distributed data parallelism (DDP) requires frequent communication of gradients across workers, which can saturate bandwidth. Infrequent communication strategies (e.g., Local SGD) reduce this overhead but, when applied to adaptive optimizers, often suffer a performance gap relative to fully synchronous DDP. We trace this gap to a time-scale mismatch: the optimizer's fast…
▽ More
Training large models with distributed data parallelism (DDP) requires frequent communication of gradients across workers, which can saturate bandwidth. Infrequent communication strategies (e.g., Local SGD) reduce this overhead but, when applied to adaptive optimizers, often suffer a performance gap relative to fully synchronous DDP. We trace this gap to a time-scale mismatch: the optimizer's fast-moving momentum, tuned for frequent updates, decays too quickly to smooth gradients over long intervals, leading to noise-dominated optimization. To address this, we propose MT-DAO, a family of optimizers that employs multiple slow- and fast-moving first momenta or the gradient to track update dynamics across different time scales, for which we provide the first convergence guarantees. Empirically, for language-model pre-training, this eliminates the performance gap with DDP, outperforming infrequent-communication baselines in perplexity and reducing iso-token wall-clock time by 6-27% on Ethernet interconnects. At the 720M scale, MT-DAO reaches a target perplexity in 24% fewer steps and 35% less time than the single-momentum DDP baseline. MT-DAO enables effective cross-datacenter training and training over wide geographic areas.
△ Less
Submitted 6 October, 2025;
originally announced October 2025.
-
AIReg-Bench: Benchmarking Language Models That Assess AI Regulation Compliance
Authors:
Bill Marino,
Rosco Hunter,
Christoph Schnabl,
Zubair Jamali,
Marinos Emmanouil Kalpakos,
Mudra Kashyap,
Isaiah Hinton,
Alexa Hanson,
Maahum Nazir,
Felix Steffek,
Hongkai Wen,
Nicholas D. Lane
Abstract:
As governments move to regulate AI, there is growing interest in using Large Language Models (LLMs) to assess whether or not an AI system complies with a given AI Regulation (AIR). However, there is presently no way to benchmark the performance of LLMs at this task. To fill this void, we introduce AIReg-Bench: the first open benchmark dataset designed to test how well LLMs can assess compliance wi…
▽ More
As governments move to regulate AI, there is growing interest in using Large Language Models (LLMs) to assess whether or not an AI system complies with a given AI Regulation (AIR). However, there is presently no way to benchmark the performance of LLMs at this task. To fill this void, we introduce AIReg-Bench: the first open benchmark dataset designed to test how well LLMs can assess compliance with the EU AI Act (AIA). We created this dataset through a two-step process: (1) by prompting an LLM with carefully structured instructions, we generated 120 technical documentation excerpts (samples), each depicting a fictional, albeit plausible, AI system -- of the kind an AI provider might produce to demonstrate their compliance with AIR; (2) legal experts then reviewed and annotated each sample to indicate whether, and in what way, the AI system described therein violates specific Articles of the AIA. The resulting dataset, together with our evaluation of whether frontier LLMs can reproduce the experts' compliance labels, provides a starting point to understand the opportunities and limitations of LLM-based AIR compliance assessment tools and establishes a benchmark against which subsequent LLMs can be compared. The dataset and evaluation code are available at https://github.com/camlsys/aireg-bench.
△ Less
Submitted 6 February, 2026; v1 submitted 1 October, 2025;
originally announced October 2025.
-
Towards mono-energetic virtual $ν$ beam cross-section measurements: A feasibility study of $ν$-Ar interaction analysis with DUNE-PRISM
Authors:
DUNE Collaboration,
S. Abbaslu,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
C. Adriano,
F. Akbar,
F. Alemanno,
N. S. Alex,
K. Allison,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
P. Amedo,
J. Anderson,
D. A. Andrade,
C. Andreopoulos,
M. Andreotti
, et al. (1302 additional authors not shown)
Abstract:
Neutrino-nucleus cross-section measurements are critical for future neutrino oscillation analyses. However, our models to describe them require further refinement, and a deeper understanding of the underlying physics is essential for future neutrino oscillation experiments to realize their ambitious physics goals. Current neutrino cross-section measurements provide clear deficiencies in neutrino i…
▽ More
Neutrino-nucleus cross-section measurements are critical for future neutrino oscillation analyses. However, our models to describe them require further refinement, and a deeper understanding of the underlying physics is essential for future neutrino oscillation experiments to realize their ambitious physics goals. Current neutrino cross-section measurements provide clear deficiencies in neutrino interaction modeling, but almost all are reported averaged over broad neutrino fluxes, rendering their interpretation challenging. Using the DUNE-PRISM concept (Deep Underground Neutrino Experiment Precision Reaction Independent Spectrum Measurement) -- a movable near detector that samples multiple off-axis positions -- neutrino interaction measurements can be used to construct narrow virtual fluxes (less than 100 MeV wide). These fluxes can be used to extract charged-current neutrino-nucleus cross sections as functions of outgoing lepton kinematics within specific neutrino energy ranges. Based on a dedicated simulation with realistic event statistics and flux-related systematic uncertainties, but assuming an almost-perfect detector, we run a feasibility study demonstrating how DUNE-PRISM data can be used to measure muon neutrino charged-current integrated and differential cross sections over narrow fluxes. We find that this approach enables a model independent reconstruction of powerful observables, including energy transfer, typically accessible only in electron scattering measurements, but that large exposures may be required for differential cross-section measurements with few-\% statistical uncertainties.
△ Less
Submitted 9 September, 2025;
originally announced September 2025.
-
Operation of a Modular 3D-Pixelated Liquid Argon Time-Projection Chamber in a Neutrino Beam
Authors:
DUNE Collaboration,
S. Abbaslu,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
C. Adriano,
F. Akbar,
F. Alemanno,
N. S. Alex,
K. Allison,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
P. Amedo,
J. Anderson,
D. A. Andrade,
C. Andreopoulos,
M. Andreotti
, et al. (1299 additional authors not shown)
Abstract:
The 2x2 Demonstrator, a prototype for the Deep Underground Neutrino Experiment (DUNE) liquid argon (LAr) Near Detector, was exposed to the Neutrinos from the Main Injector (NuMI) neutrino beam at Fermi National Accelerator Laboratory (Fermilab). This detector prototypes a new modular design for a liquid argon time-projection chamber (LArTPC), comprised of a two-by-two array of four modules, each f…
▽ More
The 2x2 Demonstrator, a prototype for the Deep Underground Neutrino Experiment (DUNE) liquid argon (LAr) Near Detector, was exposed to the Neutrinos from the Main Injector (NuMI) neutrino beam at Fermi National Accelerator Laboratory (Fermilab). This detector prototypes a new modular design for a liquid argon time-projection chamber (LArTPC), comprised of a two-by-two array of four modules, each further segmented into two optically-isolated LArTPCs. The 2x2 Demonstrator features a number of pioneering technologies, including a low-profile resistive field shell to establish drift fields, native 3D ionization pixelated imaging, and a high-coverage dielectric light readout system. The 2.4 tonne active mass detector is flanked upstream and downstream by supplemental solid-scintillator tracking planes, repurposed from the MINERvA experiment, which track ionizing particles exiting the argon volume. The antineutrino beam data collected by the detector over a 4.5 day period in 2024 include over 30,000 neutrino interactions in the LAr active volume-the first neutrino interactions reported by a DUNE detector prototype. During its physics-quality run, the 2x2 Demonstrator operated at a nominal drift field of 500 V/cm and maintained good LAr purity, with a stable electron lifetime of approximately 1.25 ms. This paper describes the detector and supporting systems, summarizes the installation and commissioning, and presents the initial validation of collected NuMI beam and off-beam self-triggers. In addition, it highlights observed interactions in the detector volume, including candidate muon anti-neutrino events.
△ Less
Submitted 17 June, 2026; v1 submitted 6 September, 2025;
originally announced September 2025.
-
Measurement of single charged pion production in charged-current $ν_μ$-Ar interactions with the MicroBooNE detector
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
L. Arellano,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (156 additional authors not shown)
Abstract:
We present flux-averaged charged-current $ν_μ$ cross-section measurements on argon for final states containing exactly one $π^\pm$ and no other hadrons except nucleons. The analysis uses data from the MicroBooNE experiment in the Booster Neutrino Beam, corresponding to $1.11 \times 10^{21}$ protons on target. Total and single-differential cross-section measurements are provided within a phase spac…
▽ More
We present flux-averaged charged-current $ν_μ$ cross-section measurements on argon for final states containing exactly one $π^\pm$ and no other hadrons except nucleons. The analysis uses data from the MicroBooNE experiment in the Booster Neutrino Beam, corresponding to $1.11 \times 10^{21}$ protons on target. Total and single-differential cross-section measurements are provided within a phase space restricted to muon momenta above 150 MeV, pion momenta above 100 MeV, and muon-pion opening angles smaller than 2.65 rad. Differential cross sections are reported with respect to the scattering angles of the muon and pion relative to the beam direction, their momenta, and their combined opening angle. The differential cross section with respect to muon momentum is based on a subset of selected events with the muon track fully contained in the detector, whereas the cross section with respect to pion momentum is based on a subset of selected events rich in pions that have not hadronically scattered on the argon before coming to rest. The latter has not been measured on argon before. The total cross section is measured as $(3.75~\pm~0.07~\textrm{(stat.)}~\pm~0.80~\textrm{(syst.)}) \times 10^{-38} \, \text{cm}^2/\text{Ar}$ at a mean energy of approximately 0.8 GeV. Comparisons of the measured cross sections with predictions from multiple neutrino-nucleus interaction generators show good overall agreement, except at very forward muon angles.
△ Less
Submitted 11 February, 2026; v1 submitted 3 September, 2025;
originally announced September 2025.
-
Sampling Off-Axis Neutrino Fluxes with the Short-Baseline Near Detector
Authors:
P. Abratenko,
R. Acciarri,
C. Adams,
L. Aliaga-Soplin,
O. Alterkait,
R. Alvarez-Garrote,
D. Andrade Aldana,
C. Andreopoulos,
A. Antonakis,
L. Arellano,
J. Asaadi,
S. Balasubramanian,
A. Barnard,
V. Basque,
J. Bateman,
A. Beever,
E. Belchior,
M. Betancourt,
A. Bhat,
M. Bishai,
A. Blake,
B. Bogart,
D. Brailsford,
A. Brandt,
S. Brickner
, et al. (177 additional authors not shown)
Abstract:
The Short-Baseline Near Detector (SBND), the near detector in the Short-Baseline Neutrino Program at Fermi National Accelerator Laboratory, is located just 110 m from the Booster Neutrino Beam target. Thanks to this close proximity, relative to its 4 m $\times$ 4 m front face, neutrinos enter SBND over a range of angles from $0^{\circ}$ to approximately $1.6^{\circ}$, enabling the detector to samp…
▽ More
The Short-Baseline Near Detector (SBND), the near detector in the Short-Baseline Neutrino Program at Fermi National Accelerator Laboratory, is located just 110 m from the Booster Neutrino Beam target. Thanks to this close proximity, relative to its 4 m $\times$ 4 m front face, neutrinos enter SBND over a range of angles from $0^{\circ}$ to approximately $1.6^{\circ}$, enabling the detector to sample variations in the neutrino flux as a function of angle-a technique known as PRISM, referred to here as SBND-PRISM. In this paper, we show how muon- and electron-neutrino fluxes vary as a function of the neutrino beam axis angle and how this can be exploited to expand the physics potential of SBND. We make use of a model that predicts an angle-dependent electron-neutrino excess signal to illustrate this effect, such as $ν_μ\to ν_e$ oscillations. We present how SBND-PRISM provides a method to add robustness against uncertainties in cross-section modeling and, more generally, uncertainties that do not depend on the spatial position of neutrino interaction inside the detector. The fluxes, along with their associated covariance matrices, are made publicly available with this publication.
△ Less
Submitted 20 April, 2026; v1 submitted 27 August, 2025;
originally announced August 2025.
-
Opportunities and challenges to study solar neutrinos with a Q-Pix pixel readout
Authors:
M. Á. García-Peris,
G. Ruiz,
S. Kubota,
A. Navrer-Agasson,
G. V. Stenico,
E. Gramellini,
R. Guenette,
J. Asaadi,
J. B. R. Battat,
V. A. Chirayath,
E. Church,
Z. Djurcic,
A. C. Ezeribe,
J. N. Gainer,
G. Gansle,
K. Keefe,
N. Lane,
C. Mauger,
Y. Mei,
F. M. Newcomer,
D. R. Nygren,
M. Rooks,
P. Sau,
O. Seidel,
S. Söldner-Rembold
, et al. (2 additional authors not shown)
Abstract:
The study of solar neutrinos presents significant opportunities in astrophysics, nuclear physics, and particle physics. However, the low-energy nature of these neutrinos introduces considerable challenges to isolate them from background events, requiring detectors with low-energy threshold, high spatial and energy resolutions, and low data rate. We present the study of solar neutrinos with a kilot…
▽ More
The study of solar neutrinos presents significant opportunities in astrophysics, nuclear physics, and particle physics. However, the low-energy nature of these neutrinos introduces considerable challenges to isolate them from background events, requiring detectors with low-energy threshold, high spatial and energy resolutions, and low data rate. We present the study of solar neutrinos with a kiloton-scale liquid argon detector located underground, instrumented with a pixel readout using the Q-Pix technology. We explore the potential of using volume fiducialization, directional topological information, light signal coincidence and pulse-shape discrimination to enhance solar neutrino sensitivity. We find that discriminating neutrino signals below 5 MeV is very difficult. However, we show that these methods are useful for the detection of solar neutrinos when external backgrounds are sufficiently understood and when the detector is built using low-background techniques. When building a workable background model for this study, we identify γ background from the cavern walls and from capture of α particles in radon decay chains as both critical to solar neutrino sensitivity and significantly underconstrained by existing measurements. Finally, we highlight that the main advantage of the use of Q-Pix for solar neutrino studies lies in its ability to enable the continuous readout of all low-energy events with minimal data rates and manageable storage for further offline analyses.
△ Less
Submitted 21 July, 2025;
originally announced July 2025.
-
Spatial and Temporal Evaluations of the Liquid Argon Purity in ProtoDUNE-SP
Authors:
DUNE Collaboration,
S. Abbaslu,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
C. Adriano,
F. Akbar,
F. Alemanno,
N. S. Alex,
K. Allison,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
P. Amedo,
J. Anderson,
D. A. Andrade,
C. Andreopoulos,
M. Andreotti
, et al. (1301 additional authors not shown)
Abstract:
Liquid argon time projection chambers (LArTPCs) rely on highly pure argon to ensure that ionization electrons produced by charged particles reach readout arrays. ProtoDUNE Single-Phase (ProtoDUNE-SP) was an approximately 700-ton liquid argon detector intended to prototype the Deep Underground Neutrino Experiment (DUNE) Far Detector Horizontal Drift module. It contains two drift volumes bisected by…
▽ More
Liquid argon time projection chambers (LArTPCs) rely on highly pure argon to ensure that ionization electrons produced by charged particles reach readout arrays. ProtoDUNE Single-Phase (ProtoDUNE-SP) was an approximately 700-ton liquid argon detector intended to prototype the Deep Underground Neutrino Experiment (DUNE) Far Detector Horizontal Drift module. It contains two drift volumes bisected by the cathode plane assembly, which is biased to create an almost uniform electric field in both volumes. The DUNE Far Detector modules must have robust cryogenic systems capable of filtering argon and supplying the TPC with clean liquid. This paper will explore comparisons of the argon purity measured by the purity monitors with those measured using muons in the TPC from October 2018 to November 2018. A new method is introduced to measure the liquid argon purity in the TPC using muons crossing both drift volumes of ProtoDUNE-SP. For extended periods on the timescale of weeks, the drift electron lifetime was measured to be above 30 ms using both systems. A particular focus will be placed on the measured purity of argon as a function of position in the detector.
△ Less
Submitted 27 August, 2025; v1 submitted 11 July, 2025;
originally announced July 2025.
-
AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
Authors:
Preslav Aleksandrov,
Meghdad Kurmanji,
Fernando Garcia Redondo,
David O'Shea,
William Shen,
Alex Iacob,
Lorenzo Sani,
Xinchi Qiu,
Nicola Cancedda,
Nicholas D. Lane
Abstract:
We introduce the Autoregressive Block-Based Iterative Encoder (AbbIE), a novel recursive generalization of the encoder-only Transformer architecture, which achieves better perplexity than a standard Transformer and allows for the dynamic scaling of compute resources at test time. This simple, recursive approach is a complement to scaling large language model (LLM) performance through parameter and…
▽ More
We introduce the Autoregressive Block-Based Iterative Encoder (AbbIE), a novel recursive generalization of the encoder-only Transformer architecture, which achieves better perplexity than a standard Transformer and allows for the dynamic scaling of compute resources at test time. This simple, recursive approach is a complement to scaling large language model (LLM) performance through parameter and token counts. AbbIE performs its iterations in latent space, but unlike latent reasoning models, does not require a specialized dataset or training protocol. We show that AbbIE upward generalizes (ability to generalize to arbitrary iteration lengths) at test time by only using 2 iterations during train time, far outperforming alternative iterative methods. AbbIE's ability to scale its computational expenditure based on the complexity of the task gives it an up to \textbf{12\%} improvement in zero-shot in-context learning tasks versus other iterative and standard methods and up to 5\% improvement in language perplexity. The results from this study open a new avenue to Transformer performance scaling. We perform all of our evaluations on model sizes up to 350M parameters.
△ Less
Submitted 7 August, 2025; v1 submitted 11 July, 2025;
originally announced July 2025.