-
LoRIS: LoRaWAN-based IoT Platform for Sustainability Monitoring in Hotels
Authors:
Yash Pandey,
Angus Gray,
Reza Serati,
Oscar Zhu,
Emil Juvan,
Anna Zinn,
Danyelle Greene,
Qingqing Chen,
Sarah MacInnes,
Siamak Layeghy,
Sara Dolnicar,
Marius Portmann
Abstract:
The hospitality sector is a major source of global greenhouse gas emissions, water stress, and waste generation, yet sustainability reporting in hotels remains constrained by coarse, manually collected operational data. We present LoRIS (LoRaWAN-based IoT platform for sustainability monitoring in hotels), a LoRaWAN-based sensing system that delivers high-resolution measurements of resource consump…
▽ More
The hospitality sector is a major source of global greenhouse gas emissions, water stress, and waste generation, yet sustainability reporting in hotels remains constrained by coarse, manually collected operational data. We present LoRIS (LoRaWAN-based IoT platform for sustainability monitoring in hotels), a LoRaWAN-based sensing system that delivers high-resolution measurements of resource consumption, environmental conditions, and guest behaviour across geographically distributed hotel properties. The architecture follows the canonical LoRaWAN reference model and is built for the operational realities of hospitality deployments: restrictive hotel IT policies, guest privacy expectations, rapid and reversible installation, and multi-year battery operation. Privacy-by-design guides modality selection and deployment zoning, and end-to-end encryption protects data from sensor to dashboard. This system has been running since February 2022 and currently spans 850 sensors of 19 types across 21 sites in Australia and Slovenia, covering both the AU915 and EU868 regulatory regions. The platform has generated over 202 million sensor records and ingests approximately 245,000 uplink messages per day on managed serverless infrastructure. Our system has been successfully used for seven field studies spanning food waste, energy consumption, and water consumption, including controlled intervention experiments that measure environmental outcomes and guest satisfaction in parallel. This system shows that LoRaWAN sensing can be deployed at scale in operational hotels without compromising guest experience or privacy.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
The cycle C9 does not admit uniform mixing
Authors:
Alison Gray,
Pransu Patel,
Isaiah Young
Abstract:
We study continuous-time quantum walks (CTQWs) on cycles. In particular, we prove that the cycle $C_9$ does not admit uniform mixing at any time via algebraic geometry and Gröbner basis techniques to rule out all cyclic 9-roots.
We study continuous-time quantum walks (CTQWs) on cycles. In particular, we prove that the cycle $C_9$ does not admit uniform mixing at any time via algebraic geometry and Gröbner basis techniques to rule out all cyclic 9-roots.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Element-Specific Visualization of Layer-Parity and Twist-Dependent Magnetism in CrSBr
Authors:
Aalok Tiwari,
Shubhada Patil,
Ravi Kumar Bandapelli,
Alevtina Smekhova,
Abhishek Kumar,
Wenhao Liu,
I-Hsuan Kao,
Zhenhong Cui,
Raghvendra Posti,
Brandon Tran,
Zixin Zhai,
Priti Yadav,
Sandy Adhitia Ekahana,
Alexander X. Gray,
Bing Lv,
Vivekanand Shukla,
Florian Kronast,
Simranjeet Singh,
Jyoti Katoch
Abstract:
Van der Waals (vdW) based antiferromagnets (AFMs) are an ideal platform for probing and understanding thickness- and twist-angle-dependent emergent spin phenomena. However, element-specific nanoscale characterization of the spin structure in atomically thin vdW-based AFMs systems and layer-parity effects remain elusive, making them crucial for both fundamental insight into low-dimensional magnetis…
▽ More
Van der Waals (vdW) based antiferromagnets (AFMs) are an ideal platform for probing and understanding thickness- and twist-angle-dependent emergent spin phenomena. However, element-specific nanoscale characterization of the spin structure in atomically thin vdW-based AFMs systems and layer-parity effects remain elusive, making them crucial for both fundamental insight into low-dimensional magnetism and the rational design of spintronic devices based on these materials. Here, we utilize X-ray magnetic circular and linear dichroisms paired with photoemission electron microscopy to resolve the magnetic order in atomically thin CrSBr. Our comprehensive measurements reveal CrSBr magnetic structure at the nanoscale and its dependence on the layer number, surface encapsulation, temperature, and applied field. Moreover, in the orthogonally twisted bilayer configuration, obtained by twisting two CrSBr ferromagnetic monolayers by 90$^\circ$, the magnetic easy axis fundamentally differs from the individual monolayers, unlocking a new pathway for moiré magnetism.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Imaging of van der Waals Materials via Standing-Wave Photoemission Microscopy: Depth-Resolved Electronic Structure of WS2
Authors:
J. R. Paudel,
R. Muzzio,
M. E. Matzelle,
S. Sheikh,
A. Fadul,
A. Tiwari,
F. Salmassi,
E. Gullikson,
K. M. McCreary,
B. T. Jonker,
J. Nieminen,
C. M. Schneider,
F. Kronast,
A. Bansil,
J. Katoch,
A. X. Gray
Abstract:
Two-dimensional van der Waals materials promise electronic, optoelectronic, and quantum technologies, yet depth-resolved characterization remains challenging. Here, we demonstrate standing-wave photoemission electron microscopy (SW-PEEM) for Angstrom-scale spectromicroscopy of monolayer WS2 on a W/C multilayer substrate. Tuning the X-ray standing wave through the monolayer yields a chemical depth…
▽ More
Two-dimensional van der Waals materials promise electronic, optoelectronic, and quantum technologies, yet depth-resolved characterization remains challenging. Here, we demonstrate standing-wave photoemission electron microscopy (SW-PEEM) for Angstrom-scale spectromicroscopy of monolayer WS2 on a W/C multilayer substrate. Tuning the X-ray standing wave through the monolayer yields a chemical depth profile and valence-band modulation with enhanced sensitivity to the top and bottom sulfur layers. X-ray optical modeling determines the structure and field distribution. The measurements reveal an ~0.2 eV shift in sulfur-derived valence-band spectral weight between measurements with enhanced sensitivity to the top and bottom sulfur layers. This shift is unlikely to arise from strong direct substrate hybridization and is instead consistent with sulfur-related surface species, as supported by calculations using a representative elemental-sulfur model. These results establish SW-PEEM as a non-destructive depth-resolved probe, highlighting its potential to probe interfacial coupling, chemical reconstruction, and emergent states in van der Waals and moiré systems.
△ Less
Submitted 25 July, 2026;
originally announced July 2026.
-
Recover, Discover, Plan: Learning Skills and Concepts from Robot Failures
Authors:
Bowen Li,
Mayank Mishra,
Y. Isabel Liu,
Stone Tao,
Nishanth Kumar,
Alexander G. Gray,
Ruwan Wickramarachchi,
Jonathan Francis,
Sebastian Scherer,
Tom Silver
Abstract:
Intelligent robots should not only recover from failures, but also acquire the abstract knowledge needed to avoid them in the future. While reinforcement learning (RL) can learn reactive recovery behaviors, training a separate policy for every distinct failure mode is highly inefficient. We introduce Recovery-Driven Synthesis of Relational Concepts (ReSYNC), the first approach that progressively d…
▽ More
Intelligent robots should not only recover from failures, but also acquire the abstract knowledge needed to avoid them in the future. While reinforcement learning (RL) can learn reactive recovery behaviors, training a separate policy for every distinct failure mode is highly inefficient. We introduce Recovery-Driven Synthesis of Relational Concepts (ReSYNC), the first approach that progressively discovers and refines state abstractions (relational predicates) from failure-recovery experience to support abstract planning. Unlike purely reactive methods, ReSYNC jointly learns skills and concepts through an incremental dual-learning process. In the skill-learning phase, the robot uses RL to learn to recover from failures seen in training tasks. In the concept-learning phase, the robot discovers new relational predicates and refines its abstract planning model to explain and generalize the learned recovery behaviors. This interaction enables ReSYNC to convert local recoveries seen during training into global failure avoidance at test time. Across four simulated domains, we show that ReSYNC's ability to continually expand and refine its abstraction library allows it to solve long-horizon, previously unseen problems, outperforming strong baselines by over 50%. Additionally, we demonstrate sim-to-real transfer of ReSYNC, where it performs real-world non-prehensile manipulation skills and generalizes to unseen scenarios through abstract planning. Overall, ReSYNC represents a significant step toward robots that autonomously acquire abstractions for scalable, failure-aware planning in the physical world.
△ Less
Submitted 16 June, 2026;
originally announced June 2026.
-
Dimensionality-Driven Electronic and Orbital Transitions Mediating Interfacial Magnetism in LaNiO3/CaMnO3 Observed In Situ
Authors:
B-A. Courchene,
A. Hampel,
S. Beck,
J. R. Paudel,
J. D. Grassi,
L. A. Lapinski,
A. M. Derrico,
M. Terilli,
M. Kareev,
C. Klewe,
A. Gloskovskii,
C. Schlueter,
S. K. Chaluvadi,
F. Mazzola,
I. Vobornik,
P. Orgiani,
J. Chakhalian,
A. J. Millis,
A. X. Gray
Abstract:
Emergent magnetic states at oxide interfaces arise from the interplay of charge transfer, orbital reconstruction, and dimensional confinement, offering a route to engineered correlated-electron behavior in nanoscale spintronic materials. Here, we combine in situ synthesis, polarization-dependent angle-resolved photoelectron spectroscopy, X-ray magnetic circular dichroism, and first-principles elec…
▽ More
Emergent magnetic states at oxide interfaces arise from the interplay of charge transfer, orbital reconstruction, and dimensional confinement, offering a route to engineered correlated-electron behavior in nanoscale spintronic materials. Here, we combine in situ synthesis, polarization-dependent angle-resolved photoelectron spectroscopy, X-ray magnetic circular dichroism, and first-principles electronic-structure calculations to investigate LaNiO3/CaMnO3 superlattices. We show that reducing the LaNiO3 thickness drives a metal-insulator transition accompanied by loss of electronic coherence and an orbital-polarization crossover in the ultrathin limit. These changes weaken charge transfer across the interface and suppress the interfacial Mn magnetic moment in CaMnO3, revealing that the emergent ferromagnetic state is directly governed by electronic confinement in LaNiO3. The insulating state and orbital reconstruction are reproduced by density functional theory combined with dynamical mean-field theory. Together, these results establish a direct and tunable coupling among electronic, orbital, and magnetic degrees of freedom in oxide heterostructures.
△ Less
Submitted 30 April, 2026;
originally announced April 2026.
-
Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation
Authors:
Andy Gray
Abstract:
A central question in the debate over large language models is whether transformers can learn rules they have never seen, or whether they can only interpolate: predict new cases from their similarity to training examples. We test this in a controlled setting where interpolation provably fails, so success can only come from computation beyond interpolation. We train small transformers to predict th…
▽ More
A central question in the debate over large language models is whether transformers can learn rules they have never seen, or whether they can only interpolate: predict new cases from their similarity to training examples. We test this in a controlled setting where interpolation provably fails, so success can only come from computation beyond interpolation. We train small transformers to predict the rollout of a cellular automaton whose update rule is pure XOR, and remove one entry of the rule's truth table from all direct supervision. The missing entry's output is never shown to the model; its only trace is indirect, as wrong values corrupt visible predictions at later timesteps. Because XOR parity flips whenever one input bit is changed, every one-bit neighbour of the missing entry carries the opposite label, and we prove that similarity-based predictors, including nearest-neighbour, kernel, and Gaussian-process methods, are forced to the wrong answer. A two-layer transformer can nevertheless recover the missing entry, and circuit extraction confirms it computes XOR exactly. Ablations show the recovery depends on gradient signal propagating through multi-step prediction, and a second, structurally unrelated benchmark on symbolic operator chains exhibits the same capacity under ordinary autoregressive training. Together with a constructive proof that a standard transformer block can implement exact local Boolean rules, these results provide an existence proof that transformers can learn rule structure not directly observed in training and express it explicitly. This rules out the strongest architectural form of the interpolation-only account, the claim that transformers cannot in principle discover and communicate unseen rules, while leaving open when such behaviour arises in large-scale language training.
△ Less
Submitted 29 July, 2026; v1 submitted 17 March, 2026;
originally announced March 2026.
-
Learning Physical Operators using Neural Operators
Authors:
Vignesh Gopakumar,
Ander Gray,
Dan Giles,
Lorenzo Zanisi,
Matt J. Kusner,
Timo Betcke,
Stanislas Pamela,
Marc Peter Deisenroth
Abstract:
Neural operators have emerged as promising surrogate models for solving partial differential equations (PDEs), but struggle to generalise beyond training distributions and are often constrained to a fixed temporal discretisation. This work introduces a physics-informed training framework that addresses these limitations by decomposing PDEs using operator splitting methods, training separate neural…
▽ More
Neural operators have emerged as promising surrogate models for solving partial differential equations (PDEs), but struggle to generalise beyond training distributions and are often constrained to a fixed temporal discretisation. This work introduces a physics-informed training framework that addresses these limitations by decomposing PDEs using operator splitting methods, training separate neural operators to learn individual non-linear physical operators while approximating linear operators with fixed finite-difference convolutions. This modular mixture-of-experts architecture enables generalisation to novel physical regimes by explicitly encoding the underlying operator structure. We formulate the modelling task as a neural ordinary differential equation (ODE) where these learned operators constitute the right-hand side, enabling continuous-in-time predictions through standard ODE solvers and implicitly enforcing PDE constraints. Demonstrated on incompressible and compressible Navier--Stokes equations, our approach achieves better convergence and superior performance when generalising to unseen physics. The method remains parameter-efficient, enabling temporal extrapolation beyond training horizons, and provides interpretable components whose behaviour can be verified against known physics.
△ Less
Submitted 2 April, 2026; v1 submitted 26 February, 2026;
originally announced February 2026.
-
Deep Search for Joint Sources of Gravitational Waves and High-Energy Neutrinos with IceCube During the Third Observing Run of LIGO and Virgo
Authors:
The IceCube Collaboration,
R. Abbasi,
M. Ackermann,
J. Adams,
S. K. Agarwalla,
J. A. Aguilar,
M. Ahlers,
J. M. Alameddine,
S. Ali,
N. M. Amin,
K. Andeen,
C. Argüelles,
Y. Ashida,
S. Athanasiadou,
S. N. Axani,
R. Babu,
X. Bai,
J. Baines-Holmes,
A. Balagopal V.,
S. W. Barwick,
S. Bash,
V. Basu,
R. Bay,
J. J. Beatty,
J. Becker Tjus
, et al. (2193 additional authors not shown)
Abstract:
The discovery of joint sources of high-energy neutrinos and gravitational waves has been a primary target for the LIGO, Virgo, KAGRA, and IceCube observatories. The joint detection of high-energy neutrinos and gravitational waves would provide insight into cosmic processes, from the dynamics of compact object mergers and stellar collapses to the mechanisms driving relativistic outflows. The joint…
▽ More
The discovery of joint sources of high-energy neutrinos and gravitational waves has been a primary target for the LIGO, Virgo, KAGRA, and IceCube observatories. The joint detection of high-energy neutrinos and gravitational waves would provide insight into cosmic processes, from the dynamics of compact object mergers and stellar collapses to the mechanisms driving relativistic outflows. The joint detection of multiple cosmic messengers can also elevate the significance of the common observation even when some or all of the constituent messengers are sub-threshold, i.e. not significant enough to declare their detection individually. Using data from the LIGO, Virgo, and IceCube observatories, including sub-threshold events, we searched for common sources of gravitational waves and high-energy neutrinos during the third observing run of Advanced LIGO and Advanced Virgo detectors. Our search did not identify significant joint sources. We derive constraints on the rate densities of joint sources. Our results constrain the isotropic neutrino emission from gravitational-wave sources for very high values of the total energy emitted in neutrinos (> $10^{52} - 10^{54}$ erg).
△ Less
Submitted 28 January, 2026; v1 submitted 12 January, 2026;
originally announced January 2026.
-
Unifying Deep Predicate Invention with Pre-trained Foundation Models
Authors:
Qianwei Wang,
Bowen Li,
Zhanpeng Luo,
Yifan Xu,
Alexander Gray,
Tom Silver,
Sebastian Scherer,
Katia Sycara,
Yaqi Xie
Abstract:
Long-horizon robotic tasks are hard due to continuous state-action spaces and sparse feedback. Symbolic world models help by decomposing tasks into discrete predicates that capture object properties and relations. Existing methods learn predicates either top-down, by prompting foundation models without data grounding, or bottom-up, from demonstrations without high-level priors. We introduce UniPre…
▽ More
Long-horizon robotic tasks are hard due to continuous state-action spaces and sparse feedback. Symbolic world models help by decomposing tasks into discrete predicates that capture object properties and relations. Existing methods learn predicates either top-down, by prompting foundation models without data grounding, or bottom-up, from demonstrations without high-level priors. We introduce UniPred, a bilevel learning framework that unifies both. UniPred uses large language models (LLMs) to propose predicate effect distributions that supervise neural predicate learning from low-level data, while learned feedback iteratively refines the LLM hypotheses. Leveraging strong visual foundation model features, UniPred learns robust predicate classifiers in cluttered scenes. We further propose a predicate evaluation method that supports symbolic models beyond STRIPS assumptions. Across five simulated and one real-robot domains, UniPred achieves 2-4 times higher success rates than top-down methods and 3-4 times faster learning than bottom-up approaches, advancing scalable and flexible symbolic world modeling for robotics.
△ Less
Submitted 19 December, 2025;
originally announced December 2025.
-
Estimating the prevalence of LLM-assisted text in scholarly writing
Authors:
Andrew Gray
Abstract:
The use of large language models (LLMs) in scholarly publications has grown dramatically since the launch of ChatGPT in late 2022. This usage is often undisclosed, and it can be challenging for readers and reviewers to identify human written but LLM-revised or translated text, or predominantly LLM-generated text. Given the known quality and reliability issues connected with LLM-generated text, the…
▽ More
The use of large language models (LLMs) in scholarly publications has grown dramatically since the launch of ChatGPT in late 2022. This usage is often undisclosed, and it can be challenging for readers and reviewers to identify human written but LLM-revised or translated text, or predominantly LLM-generated text. Given the known quality and reliability issues connected with LLM-generated text, their potential growth poses an increasing problem for research integrity, and for public trust in research.
This study presents a simple and easily reproducible methodology to show the growth in the full text of published papers, across the full range of research, as indexed in the Dimensions database. It uses this to demonstrate that LLM tools are likely to have been involved in the production of more than 10% of all published papers in 2024, based on disproportionate use of specific indicative words, and draws together earlier studies to confirm that this is a plausible overall estimate.
It then discusses the implications of this for the integrity of scholarly publishing, highlighting evidence that use of LLMs for text generation is still being concealed or downplayed by authors, and presents an argument that more comprehensive disclosure requirements are urgently required to address this.
△ Less
Submitted 1 December, 2025;
originally announced December 2025.
-
Evolution of electronic and magnetic properties in Mn- and Co-alloyed ferromagnetic kagome metal Fe3Sn2
Authors:
Prajwal M. Laxmeesha,
Rajesh Dutta,
Rajeev Kumar Rai,
Sharup Sheikh,
Michael F. DiScala,
Uditha M. Jayathilake,
Alexander Velič,
Tarush Tandon,
Tessa D. Tucker,
Christoph Klewe,
Haile Ambaye,
Timothy Charlton,
Tien-Lin Lee,
Eric A. Stach,
Kemp W. Plumb,
Alexander X. Gray,
Steven J. May
Abstract:
Kagome metals are an intriguing class of quantum materials as the presence of both flat bands and Dirac points provides access to functional properties present in strongly correlated and topological materials. To fully harness these electronic features, the ability to tune the Fermi level relative to the band positions is needed. Here we explore the structural, electronic and magnetic impacts of s…
▽ More
Kagome metals are an intriguing class of quantum materials as the presence of both flat bands and Dirac points provides access to functional properties present in strongly correlated and topological materials. To fully harness these electronic features, the ability to tune the Fermi level relative to the band positions is needed. Here we explore the structural, electronic and magnetic impacts of substitutional alloying within ferromagnetic kagome metal Fe3Sn2 in thin films grown by molecular beam epitaxy. Transition metals Mn and Co are chosen as substitutes for Fe to reduce or increase the d-band electron count, thereby moving the Fermi level accordingly. We find that Co is not incorporated into the Fe3Sn2 structure but instead results in a two-phase Fe-Co and (Fe,Co)Sn composite. In contrast, Fe3-xMnxSn2 films are realized with x up to 1.0, retaining crystalline quality comparable to the parent phase. The incorporation of Mn repositions the flat bands relative to the Fermi level in a manner consistent with hole-doping, as revealed by hard x-ray photoemission and density functional theory. The Fe3-xMnxSn2 films retain room temperature ferromagnetism, with x-ray magnetic circular dichroism measurements confirming that the Fe and Mn moments are ferromagnetically aligned. The ability to hole-dope this magnetic kagome metal provides a platform for tuning properties such as anomalous Hall and Nernst responses.
△ Less
Submitted 3 February, 2026; v1 submitted 28 October, 2025;
originally announced October 2025.
-
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering
Authors:
Mary Llewellyn,
Isobel Thornton,
James Bishop,
Annie Gray
Abstract:
LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (i) a sufficient number of evaluations are available for classical inference, and (ii) test prompts are independent. We propose a corrective Bayesian hierarchical model with embedding-space clustering that provides robust performance metrics in limited-data s…
▽ More
LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (i) a sufficient number of evaluations are available for classical inference, and (ii) test prompts are independent. We propose a corrective Bayesian hierarchical model with embedding-space clustering that provides robust performance metrics in limited-data settings while correcting for prompt dependence. We apply the approach to adversarial robustness benchmarks, showing consistent recovery of clustering structure, resulting in more reliable performance metrics, with 4-73% improvements to mean absolute errors and 40-450 unit improvements to expected log posterior densities.
△ Less
Submitted 4 June, 2026; v1 submitted 7 October, 2025;
originally announced October 2025.
-
Redesigning GROMACS Halo Exchange: Improving Strong Scaling with GPU-initiated NVSHMEM
Authors:
Mahesh Doijade,
Andrey Alekseenko,
Ania Brown,
Alan Gray,
Szilárd Páll
Abstract:
Improving time-to-solution in molecular dynamics simulations often requires strong scaling due to fixed-sized problems. GROMACS is highly latency-sensitive, with peak iteration rates in the sub-millisecond, making scalability on heterogeneous supercomputers challenging. MPI's CPU-centric nature introduces additional latencies on GPU-resident applications' critical path, hindering GPU utilization a…
▽ More
Improving time-to-solution in molecular dynamics simulations often requires strong scaling due to fixed-sized problems. GROMACS is highly latency-sensitive, with peak iteration rates in the sub-millisecond, making scalability on heterogeneous supercomputers challenging. MPI's CPU-centric nature introduces additional latencies on GPU-resident applications' critical path, hindering GPU utilization and scalability. To address these limitations, we present an NVSHMEM-based GPU kernel-initiated redesign of the GROMACS domain decomposition halo-exchange algorithm. Highly tuned GPU kernels fuse data packing and communication, leveraging hardware latency-hiding for fine-grained overlap. We employ kernel fusion across overlapped data forwarding communication phases and utilize the asynchronous copy engine over NVLink to optimize latency and bandwidth. Our GPU-resident formulation greatly increases communication-computation overlap, improving GROMACS strong scaling performance across NVLink by up to 1.5x (intra-node) and 2x (multi-node), and up to 1.3x multi-node over NVLink+InfiniBand. This demonstrates the profound benefits of GPU-initiated communication for strong-scaling a broad range of latency-sensitive applications.
△ Less
Submitted 25 September, 2025;
originally announced September 2025.
-
Multilinear and Linear Programs for Partially Identifiable Queries in Quasi-Markovian Structural Causal Models
Authors:
João P. Arroyo,
João G. Rodrigues,
Daniel Lawand,
Denis D. Mauá,
Junkyu Lee,
Radu Marinescu,
Alex Gray,
Eduardo R. Laurentino,
Fabio G. Cozman
Abstract:
We investigate partially identifiable queries in a class of causal models. We focus on acyclic Structural Causal Models that are quasi-Markovian (that is, each endogenous variable is connected with at most one exogenous confounder). We look into scenarios where endogenous variables are observed (and a distribution over them is known), while exogenous variables are not fully specified. This leads t…
▽ More
We investigate partially identifiable queries in a class of causal models. We focus on acyclic Structural Causal Models that are quasi-Markovian (that is, each endogenous variable is connected with at most one exogenous confounder). We look into scenarios where endogenous variables are observed (and a distribution over them is known), while exogenous variables are not fully specified. This leads to a representation that is in essence a Bayesian network where the distribution of root variables is not uniquely determined. In such circumstances, it may not be possible to precisely compute a probability value of interest. We thus study the computation of tight probability bounds, a problem that has been solved by multilinear programming in general, and by linear programming when a single confounded component is intervened upon. We present a new algorithm to simplify the construction of such programs by exploiting input probabilities over endogenous variables. For scenarios with a single intervention, we apply column generation to compute a probability bound through a sequence of auxiliary linear integer programs, thus showing that a representation with polynomial cardinality for exogenous variables is possible. Experiments show column generation techniques to be superior to existing methods.
△ Less
Submitted 2 September, 2025;
originally announced September 2025.
-
GWTC-4.0: Population Properties of Merging Compact Binaries
Authors:
The LIGO Scientific Collaboration,
the Virgo Collaboration,
the KAGRA Collaboration,
A. G. Abac,
I. Abouelfettouh,
F. Acernese,
K. Ackley,
C. Adamcewicz,
S. Adhicary,
D. Adhikari,
N. Adhikari,
R. X. Adhikari,
V. K. Adkins,
S. Afroz,
D. Agarwal,
M. Agathos,
M. Aghaei Abchouyeh,
O. D. Aguiar,
S. Ahmadzadeh,
L. Aiello,
A. Ain,
P. Ajith,
T. Akutsu,
S. Albanesi,
R. A. Alfaidi
, et al. (1783 additional authors not shown)
Abstract:
We detail the population properties of merging compact objects using 158 mergers from the cumulative Gravitational-Wave Transient Catalog 4.0, which includes three types of binary mergers: binary neutron star, neutron star--black hole binary, and binary black hole mergers. We resolve multiple over- and under-densities in the black hole mass distribution: features persist at primary masses of…
▽ More
We detail the population properties of merging compact objects using 158 mergers from the cumulative Gravitational-Wave Transient Catalog 4.0, which includes three types of binary mergers: binary neutron star, neutron star--black hole binary, and binary black hole mergers. We resolve multiple over- and under-densities in the black hole mass distribution: features persist at primary masses of $10\,M_\odot$ and $35\,M_\odot$ with a possible third feature at $\sim 20\,M_\odot$. These are departures from an otherwise power-law-like continuum that steepens above $35\,M_\odot$. Binary black holes with primary masses near $10\,M_\odot$ are more likely to have less massive secondaries, with a mass ratio distribution peaking at $q = 0.74^{+0.13}_{-0.13}$, potentially a signature of stable mass transfer during binary evolution. Black hole spins are inferred to be non-extremal, with 90\% of black holes having $χ< 0.57$, and preferentially aligned with binary orbits, implying many merging binaries form in isolation. However, we find a significant fraction, 0.24-0.42, of binaries have negative effective inspiral spins, suggesting many could be formed dynamically in gas-free environments. We find evidence for correlation between effective inspiral spin and mass ratio, though it is unclear if this is driven by variation in the mode of the distribution or the width. (Abridged)
△ Less
Submitted 17 September, 2025; v1 submitted 25 August, 2025;
originally announced August 2025.
-
GWTC-4.0: Methods for Identifying and Characterizing Gravitational-wave Transients
Authors:
The LIGO Scientific Collaboration,
the Virgo Collaboration,
the KAGRA Collaboration,
A. G. Abac,
I. Abouelfettouh,
F. Acernese,
K. Ackley,
S. Adhicary,
D. Adhikari,
N. Adhikari,
R. X. Adhikari,
V. K. Adkins,
S. Afroz,
D. Agarwal,
M. Agathos,
M. Aghaei Abchouyeh,
O. D. Aguiar,
S. Ahmadzadeh,
L. Aiello,
A. Ain,
P. Ajith,
S. Akcay,
T. Akutsu,
S. Albanesi,
R. A. Alfaidi
, et al. (1787 additional authors not shown)
Abstract:
The Gravitational-Wave Transient Catalog (GWTC) is a collection of candidate gravitational-wave transient signals identified and characterized by the LIGO-Virgo-KAGRA Collaboration. Producing the contents of the GWTC from detector data requires complex analysis methods. These comprise techniques to model the signal; identify the transients in the data; evaluate the quality of the data and mitigate…
▽ More
The Gravitational-Wave Transient Catalog (GWTC) is a collection of candidate gravitational-wave transient signals identified and characterized by the LIGO-Virgo-KAGRA Collaboration. Producing the contents of the GWTC from detector data requires complex analysis methods. These comprise techniques to model the signal; identify the transients in the data; evaluate the quality of the data and mitigate possible instrumental issues; infer the parameters of each transient; compare the data with the waveform models for compact binary coalescences; and handle the large amount of results associated with all these different analyses. In this paper, we describe the methods employed to produce the catalog's fourth release, GWTC-4.0, focusing on the analysis of the first part of the fourth observing run of Advanced LIGO, Advanced Virgo and KAGRA.
△ Less
Submitted 29 June, 2026; v1 submitted 25 August, 2025;
originally announced August 2025.
-
GWTC-4.0: An Introduction to Version 4.0 of the Gravitational-Wave Transient Catalog
Authors:
The LIGO Scientific Collaboration,
the Virgo Collaboration,
the KAGRA Collaboration,
A. G. Abac,
I. Abouelfettouh,
F. Acernese,
K. Ackley,
S. Adhicary,
D. Adhikari,
N. Adhikari,
R. X. Adhikari,
V. K. Adkins,
S. Afroz,
D. Agarwal,
M. Agathos,
M. Aghaei Abchouyeh,
O. D. Aguiar,
S. Ahmadzadeh,
L. Aiello,
A. Ain,
P. Ajith,
S. Akcay,
T. Akutsu,
S. Albanesi,
R. A. Alfaidi
, et al. (1786 additional authors not shown)
Abstract:
The Gravitational-Wave Transient Catalog (GWTC) is a collection of short-duration (transient) gravitational wave signals identified by the LIGO-Virgo-KAGRA Collaboration in gravitational-wave data produced by the eponymous detectors. The catalog provides information about the identified candidates, such as the arrival time and amplitude of the signal and properties of the signal's source as inferr…
▽ More
The Gravitational-Wave Transient Catalog (GWTC) is a collection of short-duration (transient) gravitational wave signals identified by the LIGO-Virgo-KAGRA Collaboration in gravitational-wave data produced by the eponymous detectors. The catalog provides information about the identified candidates, such as the arrival time and amplitude of the signal and properties of the signal's source as inferred from the observational data. GWTC is the data release of this dataset and version 4.0 extends the catalog to include observations made during the first part of the fourth LIGO-Virgo-KAGRA observing run up until 2024 January 31. This paper marks an introduction to a collection of articles related to this version of the catalog, GWTC-4.0. The collection of articles accompanying the catalog provides documentation of the methods used to analyze the data, summaries of the catalog of events, observational measurements drawn from the population, and detailed discussions of selected candidates
△ Less
Submitted 26 June, 2026; v1 submitted 25 August, 2025;
originally announced August 2025.
-
Evaluating an Immersive Analytics Application at an Enterprise Business Intelligence Customer Conference
Authors:
Matthew Brehmer,
Ginger Gloystein,
Bailiang Zhou,
Abby Gray,
Sruthi Pillai,
Ben Medina,
Vidya Setlur
Abstract:
We reflect on an evaluation of an immersive analytics application (Tableau for visionOS) conducted at a large enterprise business intelligence (BI) conference. Conducting a study in such a context offered an opportunistic setting to gather diverse feedback. However, this setting also highlighted the challenge of evaluating usability while also assessing potential utility, as feedback straddled bet…
▽ More
We reflect on an evaluation of an immersive analytics application (Tableau for visionOS) conducted at a large enterprise business intelligence (BI) conference. Conducting a study in such a context offered an opportunistic setting to gather diverse feedback. However, this setting also highlighted the challenge of evaluating usability while also assessing potential utility, as feedback straddled between the novelty of the experience and the practicality of the application in participants' analytical workflows. This formative evaluation with 22 participants allowed us to gather insights with respect to the usability of Tableau for visionOS, along with broader perspectives on the potential for head-mounted displays (HMDs) to promote new ways to engage with BI data. Our experience suggests a need for new evaluation considerations that integrate qualitative and quantitative measures and account for unique interaction patterns with 3D representations and interfaces accessible via an HMD. Overall, we contribute an enterprise perspective on evaluation methodologies for immersive analytics.
△ Less
Submitted 20 August, 2025;
originally announced August 2025.
-
Categorical-algebraic aspects of Heyting semilattices
Authors:
Xabier García-Martínez,
James R. A. Gray,
Michael A. Hoefnagel,
Tim Van der Linden,
Corentin Vienne
Abstract:
This article gives an overview of some key categorical-algebraic properties of the variety of Heyting semilattices, with the aim of correcting a misconception in the literature. We confirm that the category of Heyting semilattices is not algebraically coherent, even though it satisfies a strong version of the so-called Smith is Huq condition (on the equivalence of two types of commutators).
We a…
▽ More
This article gives an overview of some key categorical-algebraic properties of the variety of Heyting semilattices, with the aim of correcting a misconception in the literature. We confirm that the category of Heyting semilattices is not algebraically coherent, even though it satisfies a strong version of the so-called Smith is Huq condition (on the equivalence of two types of commutators).
We also prove that Higgins commutators of normal subobjects are normal, as a consequence of the fact that Heyting semilattices form an arithmetical category. We provide an elementary characterisation of when a pair of subobjects commutes, and use this in the construction of two counterexamples.
We further show that centralisers exist, centralisers of normal monomorphisms are normal monomorphisms, and normal monomorphisms are closed under composition. We study the latter condition in detail. On the other hand, we show that the category of Heyting semilattices does not satisfy normality of unions. Hence, it is not action accessible and so it does not admit all normalisers. In particular, this means that the known implication between action accessibility and the condition requiring the existence of centralisers of normal monomorphisms which are themselves normal, is strict.
△ Less
Submitted 15 August, 2025;
originally announced August 2025.
-
Sloan Digital Sky Survey-V: Pioneering Panoptic Spectroscopy
Authors:
Juna A. Kollmeier,
Hans-Walter Rix,
Conny Aerts,
James Aird,
Pablo Vera Alfaro,
Andrés Almeida,
Scott F. Anderson,
Óscar Jiménez Arranz,
Stefan M. Arseneau,
Roberto Assef,
Shir Aviram,
Catarina Aydar,
Carles Badenes,
Avrajit Bandyopadhyay,
Kat Barger,
Robert H. Barkhouser,
Franz E. Bauer,
Chad Bender,
Felipe Besser,
Binod Bhattarai,
Pavaman Bilgi,
Jonathan Bird,
Dmitry Bizyaev,
Guillermo A. Blanc,
Michael R. Blanton
, et al. (195 additional authors not shown)
Abstract:
The Sloan Digital Sky Survey-V (SDSS-V) is pioneering panoptic spectroscopy: it is the first all-sky, multi-epoch, optical-to-infrared spectroscopic survey. SDSS-V is mapping the sky with multi-object spectroscopy (MOS) at telescopes in both hemispheres (the 2.5-m Sloan Foundation Telescope at Apache Point Observatory and the 100-inch du Pont Telescope at Las Campanas Observatory), where 500 zonal…
▽ More
The Sloan Digital Sky Survey-V (SDSS-V) is pioneering panoptic spectroscopy: it is the first all-sky, multi-epoch, optical-to-infrared spectroscopic survey. SDSS-V is mapping the sky with multi-object spectroscopy (MOS) at telescopes in both hemispheres (the 2.5-m Sloan Foundation Telescope at Apache Point Observatory and the 100-inch du Pont Telescope at Las Campanas Observatory), where 500 zonal robotic fiber positioners feed light from a wide-field focal plane to an optical (R$\sim 2000$, 500 fibers) and a near-infrared (R$\sim 22,000$, 300 fibers) spectrograph. In addition to these MOS capabilities, the survey is pioneering ultra wide-field ($\sim$ 4000~deg$^2$) integral field spectroscopy enabled by a new dedicated facility (LVM-I) at Las Campanas Observatory, where an integral field spectrograph (IFS) with 1801 lenslet-coupled fibers arranged in a 0.5 degree diameter hexagon feeds multiple R$\sim$4000 optical spectrographs that cover 3600-9800 angstroms. SDSS-V's hardware and multi-year survey strategy are designed to decode the chemo-dynamical history of the Milky Way Galaxy and tackle fundamental open issues in stellar physics in its Milky Way Mapper program, trace the growth physics of supermassive black holes in its Black Hole Mapper program, and understand the self-regulation mechanisms and the chemical enrichment of galactic ecosystems at the energy-injection scale in its Local Volume Mapper program. The survey is well-timed to multiply the scientific output from major all-sky space missions. The SDSS-V MOS programs began robotic operations in 2021; IFS observations began in 2023 with the completion of the LVM-I facility. SDSS-V builds upon decades of heritage of SDSS's pioneering advances in data analysis, collaboration spirit, infrastructure, and product deliverables in astronomy.
△ Less
Submitted 9 July, 2025;
originally announced July 2025.
-
Nice exact categories are coexact
Authors:
James Richard Andrew Gray
Abstract:
Several important types of categories have been shown to be both exact and coexact (in the sense of Barr). The first type consists of abelian categories, which due to their self-dual definition, can be seen to be both exact and coexact by Tierney's characterization of them as additive exact categories. The next type consists of elementary toposes which are well-known to be exact, but have also bee…
▽ More
Several important types of categories have been shown to be both exact and coexact (in the sense of Barr). The first type consists of abelian categories, which due to their self-dual definition, can be seen to be both exact and coexact by Tierney's characterization of them as additive exact categories. The next type consists of elementary toposes which are well-known to be exact, but have also been shown to be coexact and coprotomodular by Bourn. In this paper we study a condition weaker than extensivity and equivalent to additivity for pointed categories. We show that for a finitely cocomplete category this condition together with exactness implies coexactness and coprotomodularity. As a special case we obtain that a finitely cocomplete pretopos is coexact.
△ Less
Submitted 27 March, 2026; v1 submitted 29 June, 2025;
originally announced June 2025.
-
Transformers Learn Faster with Semantic Focus
Authors:
Parikshit Ram,
Kenneth L. Clarkson,
Tim Klinger,
Shashanka Ubaru,
Alexander G. Gray
Abstract:
Various forms of sparse attention have been explored to mitigate the quadratic computational and memory cost of the attention mechanism in transformers. We study sparse transformers not through a lens of efficiency but rather in terms of learnability and generalization. Empirically studying a range of attention mechanisms, we find that input-dependent sparse attention models appear to converge fas…
▽ More
Various forms of sparse attention have been explored to mitigate the quadratic computational and memory cost of the attention mechanism in transformers. We study sparse transformers not through a lens of efficiency but rather in terms of learnability and generalization. Empirically studying a range of attention mechanisms, we find that input-dependent sparse attention models appear to converge faster and generalize better than standard attention models, while input-agnostic sparse attention models show no such benefits -- a phenomenon that is robust across architectural and optimization hyperparameter choices. This can be interpreted as demonstrating that concentrating a model's "semantic focus" with respect to the tokens currently being considered (in the form of input-dependent sparse attention) accelerates learning. We develop a theoretical characterization of the conditions that explain this behavior. We establish a connection between the stability of the standard softmax and the loss function's Lipschitz properties, then show how sparsity affects the stability of the softmax and the subsequent convergence and generalization guarantees resulting from the attention mechanism. This allows us to theoretically establish that input-agnostic sparse attention does not provide any benefits. We also characterize conditions when semantic focus (input-dependent sparse attention) can provide improved guarantees, and we validate that these conditions are in fact met in our empirical evaluations.
△ Less
Submitted 18 June, 2025; v1 submitted 16 June, 2025;
originally announced June 2025.
-
Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents
Authors:
Christian Schroeder de Witt,
Klaudia Krawiecka,
Igor Krawczuk,
Ben Hagag,
William L. Anderson,
Peter Belcak,
Ben Bucknall,
Xiaohong Cai,
Ayush Chopra,
Doron Cohen,
Ron F. Del Rosario,
Andis Draguns,
Annie Gray,
Keren Katz,
Vasilios Mavroudis,
Jaron Mink,
Sumeet Ramesh Motwani,
Jonathan Petit,
Leif-Sebastian Rembeck,
Chandler Smith,
John Sotiropoulos,
Steven Young,
Sarah Scheffler,
Mary Llewellyn
Abstract:
AI agents are beginning to interact with each other directly and across internet platforms and physical environments, creating security challenges beyond traditional cybersecurity and AI safety frameworks. Free-form protocols are essential for AI's task generalization but enable new threats like secret collusion and coordinated swarm attacks. Network effects can rapidly spread privacy breaches, di…
▽ More
AI agents are beginning to interact with each other directly and across internet platforms and physical environments, creating security challenges beyond traditional cybersecurity and AI safety frameworks. Free-form protocols are essential for AI's task generalization but enable new threats like secret collusion and coordinated swarm attacks. Network effects can rapidly spread privacy breaches, disinformation, jailbreaks, and data poisoning, while multi-agent dispersion and stealth optimization help adversaries evade oversight - creating novel persistent threats at a systemic level. Despite their critical importance, these security challenges remain understudied, with research fragmented across disparate fields including AI security, multi-agent learning, complex systems, cybersecurity, game theory, distributed systems, and technical AI governance. We introduce multi-agent security, a new field dedicated to securing networks of AI agents against threats that emerge or amplify through their interactions - whether direct or indirect via shared environments - with each other, humans, and institutions, and characterise fundamental security-utility and security-security trade-offs across both distributed and decentralised settings. Our preliminary work (1) taxonomizes the threat landscape arising from interacting AI agents, (2) offers applications to multi-agent security for work across diffuse subfields, and (3) proposes a unified research agenda addressing open challenges in designing secure agent systems and interaction environments. By identifying these gaps, we aim to guide research in this critical area to unlock the socioeconomic potential of large-scale agent deployment, foster public trust, and mitigate national security risks in critical infrastructure and defense contexts.
△ Less
Submitted 29 April, 2026; v1 submitted 4 May, 2025;
originally announced May 2025.
-
Rendering Transparency to Ranking in Educational Assessment via Bayesian Comparative Judgement
Authors:
Andy Gray,
Alma Rahat,
Stephen Lindsay,
Jen Pearson,
Tom Crick
Abstract:
Ensuring transparency in educational assessment is increasingly critical, particularly post-pandemic, as demand grows for fairer and more reliable evaluation methods. Comparative Judgement (CJ) offers a promising alternative to traditional assessments, yet concerns remain about its perceived opacity. This paper examines how Bayesian Comparative Judgement (BCJ) enhances transparency by integrating…
▽ More
Ensuring transparency in educational assessment is increasingly critical, particularly post-pandemic, as demand grows for fairer and more reliable evaluation methods. Comparative Judgement (CJ) offers a promising alternative to traditional assessments, yet concerns remain about its perceived opacity. This paper examines how Bayesian Comparative Judgement (BCJ) enhances transparency by integrating prior information into the judgement process, providing a structured, data-driven approach that improves interpretability and accountability.
BCJ assigns probabilities to judgement outcomes, offering quantifiable measures of uncertainty and deeper insights into decision confidence. By systematically tracking how prior data and successive judgements inform final rankings, BCJ clarifies the assessment process and helps identify assessor disagreements. Multi-criteria BCJ extends this by evaluating multiple learning outcomes (LOs) independently, preserving the richness of CJ while producing transparent, granular rankings aligned with specific assessment goals. It also enables a holistic ranking derived from individual LOs, ensuring comprehensive evaluations without compromising detailed feedback.
Using a real higher education dataset with professional markers in the UK, we demonstrate BCJ's quantitative rigour and ability to clarify ranking rationales. Through qualitative analysis and discussions with experienced CJ practitioners, we explore its effectiveness in contexts where transparency is crucial, such as high-stakes national assessments. We highlight the benefits and limitations of BCJ, offering insights into its real-world application across various educational settings.
△ Less
Submitted 17 March, 2025;
originally announced March 2025.
-
Bayesian Active Learning for Multi-Criteria Comparative Judgement in Educational Assessment
Authors:
Andy Gray,
Alma Rahat,
Tom Crick,
Stephen Lindsay
Abstract:
Comparative Judgement (CJ) provides an alternative assessment approach by evaluating work holistically rather than breaking it into discrete criteria. This method leverages human ability to make nuanced comparisons, yielding more reliable and valid assessments. CJ aligns with real-world evaluations, where overall quality emerges from the interplay of various elements. However, rubrics remain widel…
▽ More
Comparative Judgement (CJ) provides an alternative assessment approach by evaluating work holistically rather than breaking it into discrete criteria. This method leverages human ability to make nuanced comparisons, yielding more reliable and valid assessments. CJ aligns with real-world evaluations, where overall quality emerges from the interplay of various elements. However, rubrics remain widely used in education, offering structured criteria for grading and detailed feedback. This creates a gap between CJ's holistic ranking and the need for criterion-based performance breakdowns.
This paper addresses this gap using a Bayesian approach. We build on Bayesian CJ (BCJ) by Gray et al., which directly models preferences instead of using likelihoods over total scores, allowing for expected ranks with uncertainty estimation. Their entropy-based active learning method selects the most informative pairwise comparisons for assessors. We extend BCJ to handle multiple independent learning outcome (LO) components, defined by a rubric, enabling both holistic and component-wise predictive rankings with uncertainty estimates. Additionally, we propose a method to aggregate entropies and identify the most informative comparison for assessors. Experiments on synthetic and real data demonstrate our method's effectiveness. Finally, we address a key limitation of BCJ, which is the inability to quantify assessor agreement. We show how to derive agreement levels, enhancing transparency in assessment.
△ Less
Submitted 3 September, 2025; v1 submitted 1 March, 2025;
originally announced March 2025.
-
Bilevel Learning for Bilevel Planning
Authors:
Bowen Li,
Tom Silver,
Sebastian Scherer,
Alexander Gray
Abstract:
A robot that learns from demonstrations should not just imitate what it sees -- it should understand the high-level concepts that are being demonstrated and generalize them to new tasks. Bilevel planning is a hierarchical model-based approach where predicates (relational state abstractions) can be leveraged to achieve compositional generalization. However, previous bilevel planning approaches depe…
▽ More
A robot that learns from demonstrations should not just imitate what it sees -- it should understand the high-level concepts that are being demonstrated and generalize them to new tasks. Bilevel planning is a hierarchical model-based approach where predicates (relational state abstractions) can be leveraged to achieve compositional generalization. However, previous bilevel planning approaches depend on predicates that are either hand-engineered or restricted to very simple forms, limiting their scalability to sophisticated, high-dimensional state spaces. To address this limitation, we present IVNTR, the first bilevel planning approach capable of learning neural predicates directly from demonstrations. Our key innovation is a neuro-symbolic bilevel learning framework that mirrors the structure of bilevel planning. In IVNTR, symbolic learning of the predicate "effects" and neural learning of the predicate "functions" alternate, with each providing guidance for the other. We evaluate IVNTR in six diverse robot planning domains, demonstrating its effectiveness in abstracting various continuous and high-dimensional states. While most existing approaches struggle to generalize (with <35% success rate), our IVNTR achieves an average of 77% success rate on unseen tasks. Additionally, we showcase IVNTR on a mobile manipulator, where it learns to perform real-world mobile manipulation tasks and generalizes to unseen test scenarios that feature new objects, new states, and longer task horizons. Our findings underscore the promise of learning and planning with abstractions as a path towards high-level generalization.
△ Less
Submitted 11 May, 2025; v1 submitted 12 February, 2025;
originally announced February 2025.
-
Calibrated Physics-Informed Uncertainty Quantification
Authors:
Vignesh Gopakumar,
Ander Gray,
Lorenzo Zanisi,
Timothy Nunn,
Daniel Giles,
Matt J. Kusner,
Stanislas Pamela,
Marc Peter Deisenroth
Abstract:
Simulating complex physical systems is crucial for understanding and predicting phenomena across diverse fields, such as fluid dynamics and heat transfer, as well as plasma physics and structural mechanics. Traditional approaches rely on solving partial differential equations (PDEs) using numerical methods, which are computationally expensive and often prohibitively slow for real-time applications…
▽ More
Simulating complex physical systems is crucial for understanding and predicting phenomena across diverse fields, such as fluid dynamics and heat transfer, as well as plasma physics and structural mechanics. Traditional approaches rely on solving partial differential equations (PDEs) using numerical methods, which are computationally expensive and often prohibitively slow for real-time applications or large-scale simulations. Neural PDEs have emerged as efficient alternatives to these costly numerical solvers, offering significant computational speed-ups. However, their lack of robust uncertainty quantification (UQ) limits deployment in critical applications. We introduce a model-agnostic, physics-informed conformal prediction (CP) framework that provides guaranteed uncertainty estimates without requiring labelled data. By utilising a physics-based approach, we can quantify and calibrate the model's inconsistencies with the physics rather than the uncertainty arising from the data. Our approach utilises convolutional layers as finite-difference stencils and leverages physics residual errors as nonconformity scores, enabling data-free UQ with marginal and joint coverage guarantees across prediction domains for a range of complex PDEs. We further validate the efficacy of our method on neural PDE models for plasma modelling and shot design in fusion reactors.
△ Less
Submitted 10 June, 2025; v1 submitted 6 February, 2025;
originally announced February 2025.
-
Guaranteed prediction sets for functional surrogate models
Authors:
Ander Gray,
Vignesh Gopakumar,
Sylvain Rousseau,
Sébastien Destercke
Abstract:
We propose a method for obtaining statistically guaranteed prediction sets for functional machine learning methods: surrogate models which map between function spaces, motivated by the need to build reliable PDE emulators. The method constructs nested prediction sets on a low-dimensional representation (an SVD) of the surrogate model's error, and then maps these sets to the prediction space using…
▽ More
We propose a method for obtaining statistically guaranteed prediction sets for functional machine learning methods: surrogate models which map between function spaces, motivated by the need to build reliable PDE emulators. The method constructs nested prediction sets on a low-dimensional representation (an SVD) of the surrogate model's error, and then maps these sets to the prediction space using set-propagation techniques. This results in prediction sets for functional surrogate models with conformal prediction coverage guarantees. We use zonotopes as basis of the set construction, which allow an exact linear propagation and are closed under Cartesian products, making them well-suited to this high-dimensional problem. The method is model agnostic and can thus be applied to complex Sci-ML models, including Neural Operators, but also in simpler settings. We also introduce a technique to capture the truncation error of the SVD, preserving the guarantees of the method.
△ Less
Submitted 19 June, 2025; v1 submitted 30 January, 2025;
originally announced January 2025.
-
Few-shot Policy (de)composition in Conversational Question Answering
Authors:
Kyle Erwin,
Guy Axelrod,
Maria Chang,
Achille Fokoue,
Maxwell Crouse,
Soham Dan,
Tian Gao,
Rosario Uceda-Sosa,
Ndivhuwo Makondo,
Naweed Khan,
Alexander Gray
Abstract:
The task of policy compliance detection (PCD) is to determine if a scenario is in compliance with respect to a set of written policies. In a conversational setting, the results of PCD can indicate if clarifying questions must be asked to determine compliance status. Existing approaches usually claim to have reasoning capabilities that are latent or require a large amount of annotated data. In this…
▽ More
The task of policy compliance detection (PCD) is to determine if a scenario is in compliance with respect to a set of written policies. In a conversational setting, the results of PCD can indicate if clarifying questions must be asked to determine compliance status. Existing approaches usually claim to have reasoning capabilities that are latent or require a large amount of annotated data. In this work, we propose logical decomposition for policy compliance (LDPC): a neuro-symbolic framework to detect policy compliance using large language models (LLMs) in a few-shot setting. By selecting only a few exemplars alongside recently developed prompting techniques, we demonstrate that our approach soundly reasons about policy compliance conversations by extracting sub-questions to be answered, assigning truth values from contextual information, and explicitly producing a set of logic statements from the given policies. The formulation of explicit logic graphs can in turn help answer PCDrelated questions with increased transparency and explainability. We apply this approach to the popular PCD and conversational machine reading benchmark, ShARC, and show competitive performance with no task-specific finetuning. We also leverage the inherently interpretable architecture of LDPC to understand where errors occur, revealing ambiguities in the ShARC dataset and highlighting the challenges involved with reasoning for conversational question answering.
△ Less
Submitted 20 January, 2025;
originally announced January 2025.
-
Search for continuous gravitational waves from known pulsars in the first part of the fourth LIGO-Virgo-KAGRA observing run
Authors:
The LIGO Scientific Collaboration,
the Virgo Collaboration,
the KAGRA Collaboration,
A. G. Abac,
R. Abbott,
I. Abouelfettouh,
F. Acernese,
K. Ackley,
S. Adhicary,
N. Adhikari,
R. X. Adhikari,
V. K. Adkins,
D. Agarwal,
M. Agathos,
M. Aghaei Abchouyeh,
O. D. Aguiar,
I. Aguilar,
L. Aiello,
A. Ain,
P. Ajith,
T. Akutsu,
S. Albanesi,
R. A. Alfaidi,
A. Al-Jodah,
C. Alléné
, et al. (1794 additional authors not shown)
Abstract:
Continuous gravitational waves (CWs) emission from neutron stars carries information about their internal structure and equation of state, and it can provide tests of General Relativity. We present a search for CWs from a set of 45 known pulsars in the first part of the fourth LIGO--Virgo--KAGRA observing run, known as O4a. We conducted a targeted search for each pulsar using three independent ana…
▽ More
Continuous gravitational waves (CWs) emission from neutron stars carries information about their internal structure and equation of state, and it can provide tests of General Relativity. We present a search for CWs from a set of 45 known pulsars in the first part of the fourth LIGO--Virgo--KAGRA observing run, known as O4a. We conducted a targeted search for each pulsar using three independent analysis methods considering the single-harmonic and the dual-harmonic emission models. We find no evidence of a CW signal in O4a data for both models and set upper limits on the signal amplitude and on the ellipticity, which quantifies the asymmetry in the neutron star mass distribution. For the single-harmonic emission model, 29 targets have the upper limit on the amplitude below the theoretical spin-down limit. The lowest upper limit on the amplitude is $6.4\!\times\!10^{-27}$ for the young energetic pulsar J0537-6910, while the lowest constraint on the ellipticity is $8.8\!\times\!10^{-9}$ for the bright nearby millisecond pulsar J0437-4715. Additionally, for a subset of 16 targets we performed a narrowband search that is more robust regarding the emission model, with no evidence of a signal. We also found no evidence of non-standard polarizations as predicted by the Brans-Dicke theory.
△ Less
Submitted 26 September, 2025; v1 submitted 2 January, 2025;
originally announced January 2025.
-
Breaking through the classical Shannon entropy limit: A new frontier through logical semantics
Authors:
Luis A. Lastras,
Barry M. Trager,
Jonathan Lenchner,
Wojciech Szpankowski,
Chai Wah Wu,
Mark S. Squillante,
Alexander Gray
Abstract:
Information theory has provided foundations for the theories of several application areas critical for modern society, including communications, computer storage, and AI. A key aspect of Shannon's 1948 theory is a sharp lower bound on the number of bits needed to encode and communicate a string of symbols. When he introduced the theory, Shannon famously excluded any notion of semantics behind the…
▽ More
Information theory has provided foundations for the theories of several application areas critical for modern society, including communications, computer storage, and AI. A key aspect of Shannon's 1948 theory is a sharp lower bound on the number of bits needed to encode and communicate a string of symbols. When he introduced the theory, Shannon famously excluded any notion of semantics behind the symbols being communicated. This semantics-free notion went on to have massive impact on communication and computing technologies, even as multiple proposals for reintroducing semantics in a theory of information were being made, notably one where Carnap and Bar-Hillel used logic and reasoning to capture semantics. In this paper we present, for the first time, a Shannon-style analysis of a communication system equipped with a deductive reasoning capability, implemented using logical inference. We use some of the most important techniques developed in information theory to demonstrate significant and sometimes surprising gains in communication efficiency availed to us through such capability, demonstrated also through practical codes. We thus argue that proposals for a semantic information theory should include the power of deductive reasoning to magnify the value of transmitted bits as we strive to fully unlock the inherent potential of semantics.
△ Less
Submitted 31 December, 2024;
originally announced January 2025.
-
New Radio Observations of the Supernova Remnant CTA 1
Authors:
Tam Do,
Roland Kothes,
Alex S. Hill,
Andrew Gray,
Patricia Reich,
Wolfgang Reich
Abstract:
We present new radio images of the supernova remnant (SNR) CTA 1 at 1420 and 408 MHz, and in the 21 cm line of H I observed with the Dominion Radio Astrophysical Observatory Synthesis Telescope and at 1420 MHz observed with the Effelsberg 100 m telescope. We confirm previously described continuum features and elaborate further on filamentary features identified using the high-resolution (1') maps…
▽ More
We present new radio images of the supernova remnant (SNR) CTA 1 at 1420 and 408 MHz, and in the 21 cm line of H I observed with the Dominion Radio Astrophysical Observatory Synthesis Telescope and at 1420 MHz observed with the Effelsberg 100 m telescope. We confirm previously described continuum features and elaborate further on filamentary features identified using the high-resolution (1') maps from these new observations. We investigate the abrupt change in sign of rotation measure (RM) across the SNR, using the linear polarization observations in the four bands around 1420 MHz. Following X. H. Sun et al.'s (2011) investigation, we both confirm that the distribution of signs of the RMs for extragalactic sources in the area appears to match that of the shell, as well as combine the data from the four bands to estimate the relative depolarization and the intrinsic rotation measure of the SNR. We do not conclusively reject X. H. Sun et al.'s (2011) claim of a Faraday screen in the foreground causing the distribution of RMs that we observe; however, we do suggest an alternative explanation of a swept-up stellar wind from the progenitor star with a toroidal magnetic field. Finally, we expand on the analysis of the H I observations by applying the Rolling Hough Transform to isolate filamentary structure and better identify H I emission with the SNR. Further constraining the H I velocity channels associated with CTA 1, we use more recent Galactic rotation curves to calculate an updated kinematic distance of 1.09 +/- 0.2 kpc.
△ Less
Submitted 19 December, 2024;
originally announced December 2024.
-
LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation
Authors:
Bowen Li,
Zhaoyu Li,
Qiwei Du,
Jinqi Luo,
Wenshan Wang,
Yaqi Xie,
Simon Stepputtis,
Chen Wang,
Katia P. Sycara,
Pradeep Kumar Ravikumar,
Alexander G. Gray,
Xujie Si,
Sebastian Scherer
Abstract:
Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing benchmarks for NeSy AI fail to provide long-horizon reasoning tasks with complex multi-agent interactions. Furthermore, they are usually constrained by fixed and simplistic logical rules over limited entities, making them…
▽ More
Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing benchmarks for NeSy AI fail to provide long-horizon reasoning tasks with complex multi-agent interactions. Furthermore, they are usually constrained by fixed and simplistic logical rules over limited entities, making them far from real-world complexities. To address these crucial gaps, we introduce LogiCity, the first simulator based on customizable first-order logic (FOL) for an urban-like environment with multiple dynamic agents. LogiCity models diverse urban elements using semantic and spatial concepts, such as IsAmbulance(X) and IsClose(X, Y). These concepts are used to define FOL rules that govern the behavior of various agents. Since the concepts and rules are abstractions, they can be universally applied to cities with any agent compositions, facilitating the instantiation of diverse scenarios. Besides, a key feature of LogiCity is its support for user-configurable abstractions, enabling customizable simulation complexities for logical reasoning. To explore various aspects of NeSy AI, LogiCity introduces two tasks, one features long-horizon sequential decision-making, and the other focuses on one-step visual reasoning, varying in difficulty and agent behaviors. Our extensive evaluation reveals the advantage of NeSy frameworks in abstract reasoning. Moreover, we highlight the significant challenges of handling more complex abstractions in long-horizon multi-agent scenarios or under high-dimensional, imbalanced data. With its flexible design, various features, and newly raised challenges, we believe LogiCity represents a pivotal step forward in advancing the next generation of NeSy AI. All the code and data are open-sourced at our website: https://jaraxxus-me.github.io/LogiCity/
△ Less
Submitted 3 April, 2025; v1 submitted 1 November, 2024;
originally announced November 2024.
-
Search for gravitational waves emitted from SN 2023ixf
Authors:
The LIGO Scientific Collaboration,
the Virgo Collaboration,
the KAGRA Collaboration,
A. G. Abac,
R. Abbott,
I. Abouelfettouh,
F. Acernese,
K. Ackley,
S. Adhicary,
N. Adhikari,
R. X. Adhikari,
V. K. Adkins,
D. Agarwal,
M. Agathos,
M. Aghaei Abchouyeh,
O. D. Aguiar,
I. Aguilar,
L. Aiello,
A. Ain,
T. Akutsu,
S. Albanesi,
R. A. Alfaidi,
A. Al-Jodah,
C. Alléné,
A. Allocca
, et al. (1758 additional authors not shown)
Abstract:
We present the results of a search for gravitational-wave transients associated with core-collapse supernova SN 2023ixf, which was observed in the galaxy Messier 101 via optical emission on 2023 May 19th, during the LIGO-Virgo-KAGRA 15th Engineering Run. We define a five-day on-source window during which an accompanying gravitational-wave signal may have occurred. No gravitational waves have been…
▽ More
We present the results of a search for gravitational-wave transients associated with core-collapse supernova SN 2023ixf, which was observed in the galaxy Messier 101 via optical emission on 2023 May 19th, during the LIGO-Virgo-KAGRA 15th Engineering Run. We define a five-day on-source window during which an accompanying gravitational-wave signal may have occurred. No gravitational waves have been identified in data when at least two gravitational-wave observatories were operating, which covered $\sim 14\%$ of this five-day window. We report the search detection efficiency for various possible gravitational-wave emission models. Considering the distance to M101 (6.7 Mpc), we derive constraints on the gravitational-wave emission mechanism of core-collapse supernovae across a broad frequency spectrum, ranging from 50 Hz to 2 kHz where we assume the gravitational-wave emission occurred when coincident data are available in the on-source window. Considering an ellipsoid model for a rotating proto-neutron star, our search is sensitive to gravitational-wave energy $1 \times 10^{-4} M_{\odot} c^2$ and luminosity $2.6 \times 10^{-4} M_{\odot} c^2/s$ for a source emitting at 82 Hz. These constraints are around an order of magnitude more stringent than those obtained so far with gravitational-wave data. The constraint on the ellipticity of the proto-neutron star that is formed is as low as 1.08, at frequencies above 1200 Hz, surpassing past results.
△ Less
Submitted 11 March, 2025; v1 submitted 21 October, 2024;
originally announced October 2024.
-
Activity Report on the Eighth African School of Fundamental Physics and Applications (ASP2024)
Authors:
Kétévi A. Assamagan,
Mounia Laassiri,
Bobby Acharya,
Christine Darve,
Fernando Ferroni,
Mohamed Chabab,
Farida Fassi,
Kenneth Cecire,
Julia Ann Gray
Abstract:
The African School of Fundamental Physics and Applications, also known as the African School of Physics (ASP), was initiated in 2010, as a three-week biennial event, to offer additional training in fundamental and applied physics to African students with a minimum of three-year university education. Since its inception, ASP has grown to be much more than a school. ASP has become a series of activi…
▽ More
The African School of Fundamental Physics and Applications, also known as the African School of Physics (ASP), was initiated in 2010, as a three-week biennial event, to offer additional training in fundamental and applied physics to African students with a minimum of three-year university education. Since its inception, ASP has grown to be much more than a school. ASP has become a series of activities and events to support academic development of African students, teachers and faculties. We report on the eighth African School of Physics, ASP2024, organized in Morocco, on April 15--19 and July 7--21, 2024. ASP2024 included programs for university students, high school teachers and high school pupils.
△ Less
Submitted 30 September, 2024;
originally announced October 2024.
-
A search using GEO600 for gravitational waves coincident with fast radio bursts from SGR 1935+2154
Authors:
The LIGO Scientific Collaboration,
the Virgo Collaboration,
the KAGRA Collaboration,
A. G. Abac,
R. Abbott,
I. Abouelfettouh,
F. Acernese,
K. Ackley,
S. Adhicary,
N. Adhikari,
R. X. Adhikari,
V. K. Adkins,
D. Agarwal,
M. Agathos,
M. Aghaei Abchouyeh,
O. D. Aguiar,
I. Aguilar,
L. Aiello,
A. Ain,
P. Ajith,
T. Akutsu,
S. Albanesi,
R. A. Alfaidi,
A. Al-Jodah,
C. Alléné
, et al. (1758 additional authors not shown)
Abstract:
The magnetar SGR 1935+2154 is the only known Galactic source of fast radio bursts (FRBs). FRBs from SGR 1935+2154 were first detected by CHIME/FRB and STARE2 in 2020 April, after the conclusion of the LIGO, Virgo, and KAGRA Collaborations' O3 observing run. Here we analyze four periods of gravitational wave (GW) data from the GEO600 detector coincident with four periods of FRB activity detected by…
▽ More
The magnetar SGR 1935+2154 is the only known Galactic source of fast radio bursts (FRBs). FRBs from SGR 1935+2154 were first detected by CHIME/FRB and STARE2 in 2020 April, after the conclusion of the LIGO, Virgo, and KAGRA Collaborations' O3 observing run. Here we analyze four periods of gravitational wave (GW) data from the GEO600 detector coincident with four periods of FRB activity detected by CHIME/FRB, as well as X-ray glitches and X-ray bursts detected by NICER and NuSTAR close to the time of one of the FRBs. We do not detect any significant GW emission from any of the events. Instead, using a short-duration GW search (for bursts $\leq$ 1 s) we derive 50\% (90\%) upper limits of $10^{48}$ ($10^{49}$) erg for GWs at 300 Hz and $10^{49}$ ($10^{50}$) erg at 2 kHz, and constrain the GW-to-radio energy ratio to $\leq 10^{14} - 10^{16}$. We also derive upper limits from a long-duration search for bursts with durations between 1 and 10 s. These represent the strictest upper limits on concurrent GW emission from FRBs.
△ Less
Submitted 21 May, 2025; v1 submitted 11 October, 2024;
originally announced October 2024.
-
Neural Reasoning Networks: Efficient Interpretable Neural Networks With Automatic Textual Explanations
Authors:
Stephen Carrow,
Kyle Harper Erwin,
Olga Vilenskaia,
Parikshit Ram,
Tim Klinger,
Naweed Aghmad Khan,
Ndivhuwo Makondo,
Alexander Gray
Abstract:
Recent advances in machine learning have led to a surge in adoption of neural networks for various tasks, but lack of interpretability remains an issue for many others in which an understanding of the features influencing the prediction is necessary to ensure fairness, safety, and legal compliance. In this paper we consider one class of such tasks, tabular dataset classification, and propose a nov…
▽ More
Recent advances in machine learning have led to a surge in adoption of neural networks for various tasks, but lack of interpretability remains an issue for many others in which an understanding of the features influencing the prediction is necessary to ensure fairness, safety, and legal compliance. In this paper we consider one class of such tasks, tabular dataset classification, and propose a novel neuro-symbolic architecture, Neural Reasoning Networks (NRN), that is scalable and generates logically sound textual explanations for its predictions. NRNs are connected layers of logical neurons which implement a form of real valued logic. A training algorithm (R-NRN) learns the weights of the network as usual using gradient descent optimization with backprop, but also learns the network structure itself using a bandit-based optimization. Both are implemented in an extension to PyTorch (https://github.com/IBM/torchlogic) that takes full advantage of GPU scaling and batched training. Evaluation on a diverse set of 22 open-source datasets for tabular classification demonstrates performance (measured by ROC AUC) which improves over multi-layer perceptron (MLP) and is statistically similar to other state-of-the-art approaches such as Random Forest, XGBoost and Gradient Boosted Trees, while offering 43% faster training and a more than 2 orders of magnitude reduction in the number of parameters required, on average. Furthermore, R-NRN explanations are shorter than the compared approaches while producing more accurate feature importance scores.
△ Less
Submitted 10 October, 2024;
originally announced October 2024.
-
Terahertz-driven parametric excitation of Raman-active phonons in LaAlO$_{3}$
Authors:
M. Basini,
V. Unikandanunni,
F. Gabriele,
M. Cross,
A. M. Derrico,
A. X. Gray,
M. C. Hoffmann,
F. Forte,
M. Cuoco,
S. Bonetti
Abstract:
Achieving parametric excitation in an oscillating physical system involves periodically adjusting one of its parameters to modulate the oscillator's natural frequency. This phenomenon has been observed in numerous systems within physics and engineering, profoundly transforming modern science and technology. Despite rapid progress, the parametric control of collective excitations, such as phonons,…
▽ More
Achieving parametric excitation in an oscillating physical system involves periodically adjusting one of its parameters to modulate the oscillator's natural frequency. This phenomenon has been observed in numerous systems within physics and engineering, profoundly transforming modern science and technology. Despite rapid progress, the parametric control of collective excitations, such as phonons, remains a challenge while promising to generate novel and intriguing effects in a largely unexplored field. Here, we investigate the terahertz (THz) field-induced dynamics of Raman-active phonons in the perovskite structure of LaAlO$_3$ (LAO). Utilizing intense THz pulses, we demonstrate a novel mechanism of parametric phonon excitation marked by substantial subharmonic components. Theoretical analysis can successfully capture the hallmarks of the observed phenomena in a physical scenario with the THz field inducing a parametric coupling between the Raman mode and pairs of acoustic phonon excitations.
△ Less
Submitted 26 January, 2026; v1 submitted 9 October, 2024;
originally announced October 2024.
-
Heterogeneous computing in a strongly-connected CPU-GPU environment: fast multiple time-evolution equation-based modeling accelerated using data-driven approach
Authors:
Tsuyoshi Ichimura,
Kohei Fujita,
Muneo Hori,
Lalith Maddegedara,
Jack Wells,
Alan Gray,
Ian Karlin,
John Linford
Abstract:
We propose a CPU-GPU heterogeneous computing method for solving time-evolution partial differential equation problems many times with guaranteed accuracy, in short time-to-solution and low energy-to-solution. On a single-GH200 node, the proposed method improved the computation speed by 86.4 and 8.67 times compared to the conventional method run only on CPU and only on GPU, respectively. Furthermor…
▽ More
We propose a CPU-GPU heterogeneous computing method for solving time-evolution partial differential equation problems many times with guaranteed accuracy, in short time-to-solution and low energy-to-solution. On a single-GH200 node, the proposed method improved the computation speed by 86.4 and 8.67 times compared to the conventional method run only on CPU and only on GPU, respectively. Furthermore, the energy-to-solution was reduced by 32.2-fold (from 9944 J to 309 J) and 7.01-fold (from 2163 J to 309 J) when compared to using only the CPU and GPU, respectively. Using the proposed method on the Alps supercomputer, a 51.6-fold and 6.98-fold speedup was attained when compared to using only the CPU and GPU, respectively, and a high weak scaling efficiency of 94.3% was obtained up to 1,920 compute nodes. These implementations were realized using directive-based parallel programming models while enabling portability, indicating that directives are highly effective in analyses in heterogeneous computing environments.
△ Less
Submitted 30 September, 2024;
originally announced September 2024.
-
Uncertainty Quantification of Surrogate Models using Conformal Prediction
Authors:
Vignesh Gopakumar,
Ander Gray,
Joel Oskarsson,
Lorenzo Zanisi,
Daniel Giles,
Matt J. Kusner,
Stanislas Pamela,
Marc Peter Deisenroth
Abstract:
Data-driven surrogate models offer quick approximations to complex numerical and experimental systems but typically lack uncertainty quantification, limiting their reliability in safety-critical applications. While Bayesian methods provide uncertainty estimates, they offer no statistical guarantees and struggle with high-dimensional spatio-temporal problems due to computational costs. We present a…
▽ More
Data-driven surrogate models offer quick approximations to complex numerical and experimental systems but typically lack uncertainty quantification, limiting their reliability in safety-critical applications. While Bayesian methods provide uncertainty estimates, they offer no statistical guarantees and struggle with high-dimensional spatio-temporal problems due to computational costs. We present a conformal prediction (CP) framework that provides statistically guaranteed marginal coverage for surrogate models in a model-agnostic manner with near-zero computational cost. Our approach handles high-dimensional spatio-temporal outputs by performing cell-wise calibration while preserving the tensorial structure of predictions. Through extensive empirical evaluation across diverse applications including fluid dynamics, magnetohydrodynamics, weather forecasting, and fusion diagnostics, we demonstrate that CP achieves empirical coverage with valid error bars regardless of model architecture, training regime, or output dimensionality. We evaluate three nonconformity scores (conformalised quantile regression, absolute error residual, and standard deviation) for both deterministic and probabilistic models, showing that guaranteed coverage holds even for out-of-distribution predictions where models are deployed on physics regimes different from training data. Calibration requires only seconds to minutes on standard hardware. The framework enables rigorous validation of pre-trained surrogate models for downstream applications without retraining. While CP provides marginal rather than conditional coverage and assumes exchangeability between calibration and test data, our method circumvents the curse of dimensionality inherent in traditional uncertainty quantification approaches, offering a practical tool for trustworthy deployment of machine learning in physical sciences.
△ Less
Submitted 5 January, 2026; v1 submitted 19 August, 2024;
originally announced August 2024.
-
A Precision Cryogenic Positioning Stage for Detector Dithering and Flexure Compensation
Authors:
Stephen A. Smee,
Stephen C. Hope,
Randolph P. Hammond,
Leon Aslan,
Robert H. Barkhouser,
Katherine G. Smee,
Andrea Bianco,
Christoph Birk,
Maren Cosens,
Aidan C. Gray,
Michele Frangiamore,
Albert C. Harding,
Tyson Hare,
Daniel D. Kelson,
Gerrad Killion,
Nicholas P. Konidaris II,
Alicia Lanz,
Jacob McCloskey,
Andrew B. Newman,
Solange Ramirez,
Gwen C. Rudie,
Andrea Vanella,
Jason E. Williams
Abstract:
This paper presents the design and technical progress of a precision X-Y stage for detector dithering and flexure compensation. The stage is being developed for use in the Magellan InfraRed Multi-Object Spectrograph, MIRMOS. MIRMOS is a very large Nasmyth mounted spectrograph containing a combination of refractive, reflective and diffractive optics mounted on a long cryogenic optical bench. The in…
▽ More
This paper presents the design and technical progress of a precision X-Y stage for detector dithering and flexure compensation. The stage is being developed for use in the Magellan InfraRed Multi-Object Spectrograph, MIRMOS. MIRMOS is a very large Nasmyth mounted spectrograph containing a combination of refractive, reflective and diffractive optics mounted on a long cryogenic optical bench. The instrument utilizes five science cameras, each having a custom x-y stage to control the in-plane detector position within each camera, providing both dithering capability for improved sampling, and flexure compensation to correct for image motion that results from the gravity variant operation of the instrument. Designed to operate at 120~K, the stage will accurately control detector position in two orthogonal degrees of freedom, and have manual fine adjustment features to set detector tip, tilt and piston. The piezo-driven flexure stage provides high-resolution backlash-free motion of the detector and is very compact along the optical path, keeping camera length to a minimum. A magnetoresistive bridge provides position feedback in each degree of freedom, greatly reducing hysteresis, which is common in piezoelectric actuators. The system is designed to operate in open loop using a lookup table keyed to the Nasmyth rotator angle for flexure control. Here, the optomechanical design of the stage, electrical control system, and current performance results from early prototype efforts are presented and discussed.
△ Less
Submitted 19 July, 2024;
originally announced July 2024.
-
Evaluating Ensemble Methods for News Recommender Systems
Authors:
Alexander Gray,
Noorhan Abbas
Abstract:
News recommendation is crucial for facilitating individuals' access to articles, particularly amid the increasingly digital landscape of news consumption. Consequently, extensive research is dedicated to News Recommender Systems (NRS) with increasingly sophisticated algorithms. Despite this sustained scholarly inquiry, there exists a notable research gap regarding the potential synergy achievable…
▽ More
News recommendation is crucial for facilitating individuals' access to articles, particularly amid the increasingly digital landscape of news consumption. Consequently, extensive research is dedicated to News Recommender Systems (NRS) with increasingly sophisticated algorithms. Despite this sustained scholarly inquiry, there exists a notable research gap regarding the potential synergy achievable by amalgamating these algorithms to yield superior outcomes. This paper endeavours to address this gap by demonstrating how ensemble methods can be used to combine many diverse state-of-the-art algorithms to achieve superior results on the Microsoft News dataset (MIND). Additionally, we identify scenarios where ensemble methods fail to improve results and offer explanations for this occurrence. Our findings demonstrate that a combination of NRS algorithms can outperform individual algorithms, provided that the base learners are sufficiently diverse, with improvements of up to 5\% observed for an ensemble consisting of a content-based BERT approach and the collaborative filtering LSTUR algorithm. Additionally, our results demonstrate the absence of any improvement when combining insufficiently distinct methods. These findings provide insight into successful approaches of ensemble methods in NRS and advocates for the development of better systems through appropriate ensemble solutions.
△ Less
Submitted 23 June, 2024;
originally announced June 2024.
-
Valid Error Bars for Neural Weather Models using Conformal Prediction
Authors:
Vignesh Gopakumar,
Joel Oskarrson,
Ander Gray,
Lorenzo Zanisi,
Stanislas Pamela,
Daniel Giles,
Matt Kusner,
Marc Deisenroth
Abstract:
Neural weather models have shown immense potential as inexpensive and accurate alternatives to physics-based models. However, most models trained to perform weather forecasting do not quantify the uncertainty associated with their forecasts. This limits the trust in the model and the usefulness of the forecasts. In this work we construct and formalise a conformal prediction framework as a post-pro…
▽ More
Neural weather models have shown immense potential as inexpensive and accurate alternatives to physics-based models. However, most models trained to perform weather forecasting do not quantify the uncertainty associated with their forecasts. This limits the trust in the model and the usefulness of the forecasts. In this work we construct and formalise a conformal prediction framework as a post-processing method for estimating this uncertainty. The method is model-agnostic and gives calibrated error bounds for all variables, lead times and spatial locations. No modifications are required to the model and the computational cost is negligible compared to model training. We demonstrate the usefulness of the conformal prediction framework on a limited area neural weather model for the Nordic region. We further explore the advantages of the framework for deterministic and probabilistic models.
△ Less
Submitted 20 June, 2024;
originally announced June 2024.
-
Using graph neural networks to reconstruct charged pion showers in the CMS High Granularity Calorimeter
Authors:
M. Aamir,
G. Adamov,
T. Adams,
C. Adloff,
S. Afanasiev,
C. Agrawal,
C. Agrawal,
A. Ahmad,
H. A. Ahmed,
S. Akbar,
N. Akchurin,
B. Akgul,
B. Akgun,
R. O. Akpinar,
E. Aktas,
A. Al Kadhim,
V. Alexakhin,
J. Alimena,
J. Alison,
A. Alpana,
W. Alshehri,
P. Alvarez Dominguez,
M. Alyari,
C. Amendola,
R. B. Amir
, et al. (550 additional authors not shown)
Abstract:
A novel method to reconstruct the energy of hadronic showers in the CMS High Granularity Calorimeter (HGCAL) is presented. The HGCAL is a sampling calorimeter with very fine transverse and longitudinal granularity. The active media are silicon sensors and scintillator tiles readout by SiPMs and the absorbers are a combination of lead and Cu/CuW in the electromagnetic section, and steel in the hadr…
▽ More
A novel method to reconstruct the energy of hadronic showers in the CMS High Granularity Calorimeter (HGCAL) is presented. The HGCAL is a sampling calorimeter with very fine transverse and longitudinal granularity. The active media are silicon sensors and scintillator tiles readout by SiPMs and the absorbers are a combination of lead and Cu/CuW in the electromagnetic section, and steel in the hadronic section. The shower reconstruction method is based on graph neural networks and it makes use of a dynamic reduction network architecture. It is shown that the algorithm is able to capture and mitigate the main effects that normally hinder the reconstruction of hadronic showers using classical reconstruction methods, by compensating for fluctuations in the multiplicity, energy, and spatial distributions of the shower's constituents. The performance of the algorithm is evaluated using test beam data collected in 2018 prototype of the CMS HGCAL accompanied by a section of the CALICE AHCAL prototype. The capability of the method to mitigate the impact of energy leakage from the calorimeter is also demonstrated.
△ Less
Submitted 18 December, 2024; v1 submitted 17 June, 2024;
originally announced June 2024.
-
What makes Models Compositional? A Theoretical View: With Supplement
Authors:
Parikshit Ram,
Tim Klinger,
Alexander G. Gray
Abstract:
Compositionality is thought to be a key component of language, and various compositional benchmarks have been developed to empirically probe the compositional generalization of existing sequence processing models. These benchmarks often highlight failures of existing models, but it is not clear why these models fail in this way. In this paper, we seek to theoretically understand the role the compo…
▽ More
Compositionality is thought to be a key component of language, and various compositional benchmarks have been developed to empirically probe the compositional generalization of existing sequence processing models. These benchmarks often highlight failures of existing models, but it is not clear why these models fail in this way. In this paper, we seek to theoretically understand the role the compositional structure of the models plays in these failures and how this structure relates to their expressivity and sample complexity. We propose a general neuro-symbolic definition of compositional functions and their compositional complexity. We then show how various existing general and special purpose sequence processing models (such as recurrent, convolution and attention-based ones) fit this definition and use it to analyze their compositional complexity. Finally, we provide theoretical guarantees for the expressivity and systematic generalization of compositional models that explicitly depend on our proposed definition and highlighting factors which drive poor empirical performance.
△ Less
Submitted 2 May, 2024;
originally announced May 2024.
-
Depth-resolved profile of the interfacial ferromagnetism in $CaMnO_{3}/CaRuO_{3}$ superlattices
Authors:
J. R. Paudel,
A. Mansouri Tehrani,
M. Terilli,
M. Kareev,
J. Grassi,
R. K. Sah,
L. Wu,
V. N. Strocov,
C. Klewe,
P. Shafer,
J. Chakhalian,
N. A. Spaldin,
A. X. Gray
Abstract:
Emergent magnetic phenomena at interfaces represent a frontier in materials science, pivotal for advancing technologies in spintronics and magnetic storage. In this letter, we utilize a suite of advanced X-ray spectroscopic and scattering techniques to investigate emergent interfacial ferromagnetism in oxide superlattices comprised of antiferromagnetic CaMnO3 and paramagnetic CaRuO3. Our findings…
▽ More
Emergent magnetic phenomena at interfaces represent a frontier in materials science, pivotal for advancing technologies in spintronics and magnetic storage. In this letter, we utilize a suite of advanced X-ray spectroscopic and scattering techniques to investigate emergent interfacial ferromagnetism in oxide superlattices comprised of antiferromagnetic CaMnO3 and paramagnetic CaRuO3. Our findings challenge prior theoretical models by demonstrating that the ferromagnetism extends beyond the interfacial layer into multiple unit cells of CaMnO3 and exhibits an asymmetric profile. Complementary density functional calculations reveal that the interfacial ferromagnetism is driven by the double exchange mechanism, facilitated by charge transfer from Ru to Mn ions. Additionally, defect chemistry, particularly the presence of oxygen vacancies, likely plays a crucial role in modifying the magnetic moments at the interface, leading to the observed asymmetry between the top and bottom CaMnO3 interfacial magnetic layers. Our findings underscore the potential of manipulating interfacial ferromagnetism through point defect engineering.
△ Less
Submitted 2 May, 2024;
originally announced May 2024.
-
Observation of Gravitational Waves from the Coalescence of a $2.5\text{-}4.5~M_\odot$ Compact Object and a Neutron Star
Authors:
The LIGO Scientific Collaboration,
the Virgo Collaboration,
the KAGRA Collaboration,
A. G. Abac,
R. Abbott,
I. Abouelfettouh,
F. Acernese,
K. Ackley,
S. Adhicary,
N. Adhikari,
R. X. Adhikari,
V. K. Adkins,
D. Agarwal,
M. Agathos,
M. Aghaei Abchouyeh,
O. D. Aguiar,
I. Aguilar,
L. Aiello,
A. Ain,
P. Ajith,
S. Akçay,
T. Akutsu,
S. Albanesi,
R. A. Alfaidi,
A. Al-Jodah
, et al. (1771 additional authors not shown)
Abstract:
We report the observation of a coalescing compact binary with component masses $2.5\text{-}4.5~M_\odot$ and $1.2\text{-}2.0~M_\odot$ (all measurements quoted at the 90% credible level). The gravitational-wave signal GW230529_181500 was observed during the fourth observing run of the LIGO-Virgo-KAGRA detector network on 2023 May 29 by the LIGO Livingston Observatory. The primary component of the so…
▽ More
We report the observation of a coalescing compact binary with component masses $2.5\text{-}4.5~M_\odot$ and $1.2\text{-}2.0~M_\odot$ (all measurements quoted at the 90% credible level). The gravitational-wave signal GW230529_181500 was observed during the fourth observing run of the LIGO-Virgo-KAGRA detector network on 2023 May 29 by the LIGO Livingston Observatory. The primary component of the source has a mass less than $5~M_\odot$ at 99% credibility. We cannot definitively determine from gravitational-wave data alone whether either component of the source is a neutron star or a black hole. However, given existing estimates of the maximum neutron star mass, we find the most probable interpretation of the source to be the coalescence of a neutron star with a black hole that has a mass between the most massive neutron stars and the least massive black holes observed in the Galaxy. We provisionally estimate a merger rate density of $55^{+127}_{-47}~\text{Gpc}^{-3}\,\text{yr}^{-1}$ for compact binary coalescences with properties similar to the source of GW230529_181500; assuming that the source is a neutron star-black hole merger, GW230529_181500-like sources constitute about 60% of the total merger rate inferred for neutron star-black hole coalescences. The discovery of this system implies an increase in the expected rate of neutron star-black hole mergers with electromagnetic counterparts and provides further evidence for compact objects existing within the purported lower mass gap.
△ Less
Submitted 26 July, 2024; v1 submitted 5 April, 2024;
originally announced April 2024.
-
IQMDose3D: a software tool for reconstructing the dose in patient using patient planning CT images and the signals measured by IQM detector
Authors:
Aitang Xing,
Gary Goozee,
Alison Gray,
Vaughan Moutrie,
Sankar Arumugam,
Shrikant Deshpande,
Anthony Espinoza,
Vasilis Kondilis,
Marjorie McDonald,
Philip Vial
Abstract:
The integral quality monitor (IQM) system compares the signal measured with a large volume chamber mounted to the linear accelerator's head to the signal calculated using the patient DICOM RT plan for patient-specific quality assurance (PSQA). A method was developed to reconstruct the dose in patients using the signal measured by IQM chamber and patient planning CT images. A software tool named IQ…
▽ More
The integral quality monitor (IQM) system compares the signal measured with a large volume chamber mounted to the linear accelerator's head to the signal calculated using the patient DICOM RT plan for patient-specific quality assurance (PSQA). A method was developed to reconstruct the dose in patients using the signal measured by IQM chamber and patient planning CT images. A software tool named IQMDose3D was implemented to automate this procedure and integrated into the IQM-based PSQA workflow. IQMDose3D enables the physicists to evaluate PSQA by focusing on the clinical perspective by comparing the delivered plan to the approved clinical plan in terms of the clinical goals, dose-volume histogram (DVH) in addition to the three-dimensional (3D) gamma map and gamma pass rate.
△ Less
Submitted 29 March, 2024; v1 submitted 26 March, 2024;
originally announced March 2024.
-
AutoMRISimQA: an automated system for daily quality control of a 3T MRI simulator
Authors:
Aitang Xing,
Gary Goozee,
Gary Liney,
Sankar Arumugam,
Shrikant Deshpande,
Anthony Espinoza,
Alison Gray,
Vasilis Kondilis,
Doaa Elwadia,
Robba Rai,
Lois Holloway
Abstract:
A software system named AutoMRISimQA was developed to monitor the daily performance of a wide-bore 3T scanner(MRI) which was designed and dedicated to radiotherapy simulation. The system can monitor the performance of the MRI simulator not only by using image quality indices such as signal-to-noise ratio (SNR), uniformity, ghosting and contrast but also performing a quick check of geometry accurac…
▽ More
A software system named AutoMRISimQA was developed to monitor the daily performance of a wide-bore 3T scanner(MRI) which was designed and dedicated to radiotherapy simulation. The system can monitor the performance of the MRI simulator not only by using image quality indices such as signal-to-noise ratio (SNR), uniformity, ghosting and contrast but also performing a quick check of geometry accuracy as well as the external lasers quantitatively. It was implemented into the daily clinically workflow in 2013 and has been used for more than 10 years. It was also seamlessly integrated with QAtrack, allowing continuous monitoring of the consistency of the MRI simulator's performance.
△ Less
Submitted 29 March, 2024; v1 submitted 26 March, 2024;
originally announced March 2024.