-
Bootstrap validity in Bayesian semi-parametric models
Authors:
Magid Sabbagh,
David A. Stephens
Abstract:
We discuss Bayesian inference on a low-dimensional targeted parameter in the presence of possibly highly complex nuisance components within the semi-parametric inference framework using an estimating function approach. We obtain a posterior distribution using non-parametric Bayesian methods through the Dirichlet process and the Bayesian bootstrap. We relax the commonly deployed notion of stochasti…
▽ More
We discuss Bayesian inference on a low-dimensional targeted parameter in the presence of possibly highly complex nuisance components within the semi-parametric inference framework using an estimating function approach. We obtain a posterior distribution using non-parametric Bayesian methods through the Dirichlet process and the Bayesian bootstrap. We relax the commonly deployed notion of stochastic equicontinuity and develop a framework leading to posterior inference with good frequentist properties, specifically we demonstrate that the posterior distribution is asymptotically Normal and concentrates at the true value of the parameter. We emphasize the specific assumptions that are required to obtain these results, and how relaxing any of them alters the conclusions. We verify the analytical results in simulation.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
GOATS: The next generation software infrastructure for time-domain astronomy at Gemini/NOIRLab. Application to alerts from Vera C. Rubin Observatory's Legacy Survey of Space and Time
Authors:
Monika Soraisam,
Louis Avner,
Miguel Gómez,
William Vacca,
Bryan Miller,
Andrew Stephens,
Arturo Núñez,
Andrew Adamson,
César Briceño,
Hernán Chacana,
Guillermo Damke,
Nicolás Esquivel,
Paul Hirst,
Kathleen Labrie,
Thomas Matheson,
Chadd Myers,
Robert Nikutta,
Abhijit Saha,
Chris Simpson,
Olesja Smirnova,
D. J. Teal,
Sergio Troncoso,
James Turner,
Sebastián Vicencio,
Hubert Condoretti
, et al. (16 additional authors not shown)
Abstract:
Time-domain and multimessenger astronomy (MMA/TDA) targets demand rapid-response follow-up observations. In many cases, it is the only way to make discoveries and advance our understanding of the astrophysical phenomena, for example, kilonovae accompanying gravitational waves from compact object mergers, shock breakout in supernovae, prompt emission from GRBs, etc. Presently the MMA/TDA follow-up…
▽ More
Time-domain and multimessenger astronomy (MMA/TDA) targets demand rapid-response follow-up observations. In many cases, it is the only way to make discoveries and advance our understanding of the astrophysical phenomena, for example, kilonovae accompanying gravitational waves from compact object mergers, shock breakout in supernovae, prompt emission from GRBs, etc. Presently the MMA/TDA follow-up workflow requires wrangling disparate software packages and user interfaces. We present an end-to-end software tool for the community, the Gemini Observation and Analysis of Targets System (GOATS), which unifies and simplifies the workflow, particularly for Gemini follow-up observations. GOATS achieves this by integrating services from Gemini Observatory and its parent organization, NSF NOIRLab. From a single platform, GOATS enables enhanced target selection via NOIRLab's ANTARES alert broker, triggering of Gemini (and other facilities within the Astronomical Event Observatory Network), automated data retrieval from the Gemini Observatory Archive, and interactive data reduction and analysis through Gemini's DRAGONS software and NOIRLab's Astro Data Lab science platform. GOATS was successfully deployed in an end-to-end demonstration of real-time follow-up of Rubin/LSST alerts with NOIRLab facilities. As part of this demonstration, we selected targets from the Rubin alert stream and triggered follow-up observations within minutes of the Rubin detections. We obtained spectra for several targets and classified them as supernova of various types (Ia, IIP, Ib/c) with redshifts ranging from 0.05 to 0.35. By eliminating the need to manually connect tools and automating repetitive tasks, GOATS lowers the entry barrier and allows users to focus on the scientific interpretation of the observation results.
△ Less
Submitted 26 June, 2026;
originally announced June 2026.
-
Semi- and non-parametric approaches to individualized treatment regimes in the presence of causal mediation
Authors:
Misha Dolmatov,
Erica E. M. Moodie,
David A. Stephens,
Dipankar Bandyopadhyay
Abstract:
Individualized treatment rules (ITRs) map an individual patient's characteristics to their recommended treatment value. Typically, the optimal ITR is defined as the rule which maximizes a mean counterfactual outcome; the resulting ITR maximizes the effect of treatment along all causal pathways to the outcome, including indirect pathways through mediating variables. Although maximizing the total ef…
▽ More
Individualized treatment rules (ITRs) map an individual patient's characteristics to their recommended treatment value. Typically, the optimal ITR is defined as the rule which maximizes a mean counterfactual outcome; the resulting ITR maximizes the effect of treatment along all causal pathways to the outcome, including indirect pathways through mediating variables. Although maximizing the total effect is often sufficient, explicitly incorporating causal mediation in an ITR analysis has several potential benefits such as enhanced interpretability, and additional flexibility in targeting specific causal pathways. For this purpose, we introduce novel Bayesian semiparametric and nonparametric estimators for conditional mediation effects in the presence of multiple mediators and show how they can be used to estimate optimal ITRs. We demonstrate the proposed methodology via an application to optimal kidney allocation with hepatitis C positive donors.
△ Less
Submitted 19 June, 2026;
originally announced June 2026.
-
The Roasting Marshmallows Program with IGRINS on Gemini South V: Atmosphere of MASCARA-1b is Enriched in Refractory Elements
Authors:
Krishna Kanumalla,
Michael R. Line,
Martina Chiarella,
Matteo Brogi,
Peter C. B. Smith,
Jorge A. Sanchez,
Yayaati Chachan,
Joshua Lothringer,
Joost P. Wardenier,
Hayley Beltz,
Carlos Saffe,
Emily K. Deibert,
Megan Weiner Mansfield,
Stefan Pelletier,
Vivien Parmentier,
Yeon-ho Choi,
Swaetha Ramkumar,
Arjun B. Savel,
Luis Welbanks,
Jacob L. Bean,
Vatsal Panwar,
Tomás Azevedo Silva,
Lorenzo Pino,
Yuya Hayashi,
Dongwook Lim
, et al. (53 additional authors not shown)
Abstract:
Ultra-hot Jupiters (UHJs; $T_{\rm eq} \gtrsim 2000$ K) enable simultaneous detection of volatile (ice-forming) and refractory (rock-forming) species in planetary atmospheres, providing a powerful diagnostic of planet formation and atmospheric processing. We present a comprehensive high-resolution cross-correlation spectroscopy (HRCCS) analysis of the UHJ MASCARA-1b ($T_{\rm eq} \approx 2600$ K) us…
▽ More
Ultra-hot Jupiters (UHJs; $T_{\rm eq} \gtrsim 2000$ K) enable simultaneous detection of volatile (ice-forming) and refractory (rock-forming) species in planetary atmospheres, providing a powerful diagnostic of planet formation and atmospheric processing. We present a comprehensive high-resolution cross-correlation spectroscopy (HRCCS) analysis of the UHJ MASCARA-1b ($T_{\rm eq} \approx 2600$ K) using the IGRINS and IGRINS-2 spectrographs. We detect robust (SNR$>$4) signals from H$_2$O, CO, OH, Fe I, Mg I, Ca I, and Ti I, marking the most complete atmospheric inventory of MASCARA-1b to date. Using a chemically consistent atmospheric inference framework, we constrain elemental abundances to a typical precision of $\approx$0.2 dex, retrieving a solar atmospheric metallicity ([M/H]$_\odot$ $= 0.07^{+0.17}_{-0.13}$ $\approx 1.2\times$ solar), a C/O ratio (C/O $= 0.65^{+0.08}_{-0.08}$) consistent with solar value (C/O $=$ 0.59), an enhanced refractory abundance ([R/H]$_\odot$ $= 0.40^{+0.23}_{-0.17} \approx 2.5\times$ solar; $\approx 3.8\times$ stellar), and a moderately super-solar refractory-to-volatile ratio ([R/V]$_\odot$ $= 0.36^{+0.11}_{-0.09}$ $\approx 2.3\times$ solar). Comparison with formation models suggests that MASCARA-1b most likely accreted material between the soot-H$_2$O or H$_2$O-CO snowlines (at 68$\%$ confidence). We additionally find stellar values for atmospheric Ti/Mg and Ca/Mg ratios (at 68$\%$ confidence). The Mg/Fe is also found to be consistent with stellar value at 95$\%$ confidence. Therefore, we do not find strong indication of nightside cold trapping in MASCARA-1b. As homogeneous refractory-to-volatile measurements expand across the UHJ population, particularly with upcoming Extremely Large Telescopes, these diagnostics will enable statistically robust tests of emerging trends in giant planet formation and atmospheric evolution.
△ Less
Submitted 5 June, 2026;
originally announced June 2026.
-
A Panchromatic JWST Spectrum of a Giant Starspot on the Fully Convective M-dwarf TOI-3884
Authors:
C. A. Murray,
L. Garcia,
B. V. Rackham,
Z. Berta-Thompson,
A. D. Feinstein,
S. J. Mercier,
B. Charnay,
L. Hebb,
J. E. Libby-Roberts,
Y. Rotman,
A. Stephens,
M. Timmermans,
L. Welbanks,
K. Barkaoui,
Caleb I. Canas,
M. Delamer,
E. Ducrot,
S. Kanodia,
S. Mahadevan,
J. P. Ninan,
J. de Wit
Abstract:
TOI-3884 b is a rare super-Neptune transiting a fully convective M dwarf that hosts a persistent giant polar spot. Because the planet occults this active region during every transit, the system offers a unique laboratory to directly probe the stellar surface and spot properties. We present seven James Webb Space Telescope (JWST) transits of TOI-3884 b observed with NIRISS and NIRSpec (spanning 0.6…
▽ More
TOI-3884 b is a rare super-Neptune transiting a fully convective M dwarf that hosts a persistent giant polar spot. Because the planet occults this active region during every transit, the system offers a unique laboratory to directly probe the stellar surface and spot properties. We present seven James Webb Space Telescope (JWST) transits of TOI-3884 b observed with NIRISS and NIRSpec (spanning 0.6--5.3$μ$m). While all visits show a recurring spot-crossing signature, each transit exhibits a distinct spot-crossing morphology, enabling us to infer a stellar rotation period of $P$=11.102$\pm$0.003d and tightly constrain the pole-on stellar orientation ($i_{*}$=139.2$\pm$0.3$^{\circ}$, $λ_{*}$=31.1$\pm$0.4$^{\circ}$) and spot properties ($R_{\rm{spot}}=0.576^{+0.006}_{-0.005}$R$_{*}$, $φ_{\rm{spot}}$=-84.69$\pm$0.12$^{\circ}$).We leverage this orbital configuration to measure the first empirical panchromatic spectrum of an M dwarf starspot with JWST, establishing a direct observational benchmark for stellar atmosphere models in the fully convective regime. Comparison with 1D NewEra and SPHINX atmosphere models indicates that the spot is 183$\pm$1K cooler than the photosphere, consistent with previous ground-based measurements and expectations for mid M dwarf spot contrasts. While the models reproduce the observed contrasts at wavelengths longer than 1$μ$m, they significantly underpredict the contrasts at shorter wavelengths. These results demonstrate that M dwarf stellar atmosphere models may not fully capture the wavelength dependence of stellar contamination in transmission spectra and highlight the importance of empirical spot spectra for robust interpretation of planetary atmospheres, particularly in the optical.
△ Less
Submitted 1 June, 2026; v1 submitted 16 March, 2026;
originally announced March 2026.
-
Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions
Authors:
Mame Diarra Toure,
David A. Stephens
Abstract:
In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class. We decompose MI into a per-class vector $C_k(x)=σ_k^{2}/(2μ_k)$, with $μ_k{=}\mathbb{E}[p_k]$ and…
▽ More
In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class. We decompose MI into a per-class vector $C_k(x)=σ_k^{2}/(2μ_k)$, with $μ_k{=}\mathbb{E}[p_k]$ and $σ_k^2{=}\mathrm{Var}[p_k]$ across posterior samples. The decomposition follows from a second-order Taylor expansion of the entropy; the $1/μ_k$ weighting corrects boundary suppression and makes $C_k$ comparable across rare and common classes. By construction $\sum_k C_k \approx \mathrm{MI}$, and a companion skewness diagnostic flags inputs where the approximation degrades. After characterising the axiomatic properties of $C_k$, we validate it on three tasks: (i) selective prediction for diabetic retinopathy, where critical-class $C_k$ reduces selective risk by 34.7\% over MI and 56.2\% over variance baselines; (ii) out-of-distribution detection on clinical and image benchmarks, where $\sum_k C_k$ achieves the highest AUROC and the per-class view exposes asymmetric shifts invisible to MI; and (iii) a controlled label-noise study in which $\sum_k C_k$ shows less sensitivity to injected aleatoric noise than MI under end-to-end Bayesian training, while both metrics degrade under transfer learning. Across all tasks, the quality of the posterior approximation shapes uncertainty at least as strongly as the choice of metric, suggesting that how uncertainty is propagated through the network matters as much as how it is measured.
△ Less
Submitted 28 June, 2026; v1 submitted 24 February, 2026;
originally announced February 2026.
-
Semi-parametric Bayesian inference under Neyman orthogonality
Authors:
Magid Sabbagh,
David A. Stephens
Abstract:
The validity of two-step or plug-in inference methods is questioned in the Bayesian framework. We study semi-parametric models where the plug-in of a non-parametrically modelled nuisance component is used. We show that when the nuisance and targeted parameters satisfy a Neyman orthogonal score property, the approach of cutting feedback through a two-step procedure is a valid way of conducting Baye…
▽ More
The validity of two-step or plug-in inference methods is questioned in the Bayesian framework. We study semi-parametric models where the plug-in of a non-parametrically modelled nuisance component is used. We show that when the nuisance and targeted parameters satisfy a Neyman orthogonal score property, the approach of cutting feedback through a two-step procedure is a valid way of conducting Bayesian inference. Our method relies on a non-parametric Bayesian formulation based on the Dirichlet process and the Bayesian bootstrap. We show that the marginal posterior of the targeted parameter exhibits good frequentist properties despite not accounting for the inferential uncertainty of the nuisance parameter. We adopt this approach in Bayesian causal inference problems where the nuisance propensity score model is estimated to obtain marginal inference for the treatment effect parameter, and demonstrate that a plug-in of the propensity score has a negligible effect on marginal posterior inference for the causal contrast. We investigate the absence of Neyman orthogonality and exploit our findings to show that in conventional two-step procedures, the posterior distribution converges under weaker restrictions than those needed in the frequentist sequel. For a simple family of useful scores, we demonstrate that even in the absence of Neyman orthogonality, the posterior distribution is asymptotically unchanged by the estimation of the nuisance parameter, merely provided the latter estimator is consistent.
△ Less
Submitted 23 February, 2026;
originally announced February 2026.
-
Posterior Uncertainty for Targeted Parameters in Bayesian Bootstrap Procedures
Authors:
Magid Sabbagh,
David A. Stephens
Abstract:
We propose a general method to carry out a valid Bayesian analysis of a finite-dimensional `targeted' parameter in the presence of a finite-dimensional nuisance parameter. We apply our methods to causal inference based on estimating equations. While much of the literature in Bayesian causal inference has relied on the conventional 'likelihood times prior' framework, a recently proposed method, the…
▽ More
We propose a general method to carry out a valid Bayesian analysis of a finite-dimensional `targeted' parameter in the presence of a finite-dimensional nuisance parameter. We apply our methods to causal inference based on estimating equations. While much of the literature in Bayesian causal inference has relied on the conventional 'likelihood times prior' framework, a recently proposed method, the 'Linked Bayesian Bootstrap', deviated from this classical setting to obtain valid Bayesian inference using the Dirichlet process and the Bayesian bootstrap. These methods rely on an adjustment based on the propensity score and explain how to handle the uncertainty concerning it when studying the posterior distribution of a treatment effect. We examine theoretically the asymptotic properties of the posterior distribution obtained and show that our proposed method, a generalized version of the 'Linked Bayesian Bootstrap', enjoys desirable frequentist properties. In addition, we show that the credible intervals have asymptotically the correct coverage properties. We discuss the applications of our method to mis-specified and singly-robust models in causal inference.
△ Less
Submitted 2 February, 2026;
originally announced February 2026.
-
Singular Bayesian Neural Networks
Authors:
Mame Diarra Toure,
David A. Stephens
Abstract:
Bayesian neural networks promise calibrated uncertainty but require $O(mn)$ parameters for standard mean-field Gaussian posteriors. We argue this cost is often unnecessary, particularly when weight matrices exhibit fast singular value decay. By parameterizing weights as $W = AB^{\top}$ with $A \in \mathbb{R}^{m \times r}$, $B \in \mathbb{R}^{n \times r}$, we induce a posterior that is \emph{singul…
▽ More
Bayesian neural networks promise calibrated uncertainty but require $O(mn)$ parameters for standard mean-field Gaussian posteriors. We argue this cost is often unnecessary, particularly when weight matrices exhibit fast singular value decay. By parameterizing weights as $W = AB^{\top}$ with $A \in \mathbb{R}^{m \times r}$, $B \in \mathbb{R}^{n \times r}$, we induce a posterior that is \emph{singular} with respect to the Lebesgue measure, concentrating on the rank-$r$ manifold. This singularity captures structured weight correlations through shared latent factors, geometrically distinct from mean-field's independence assumption. We derive PAC-Bayes generalization bounds whose complexity term scales as $\sqrt{r(m+n)}$ instead of $\sqrt{m n}$, and prove loss bounds that decompose the error into optimization and rank-induced bias using the Eckart-Young-Mirsky theorem. We further adapt recent Gaussian complexity bounds for low-rank deterministic networks to Bayesian predictive means. Empirically, across MLPs, LSTMs, and Transformers on standard benchmarks, our method achieves competitive predictive performance while using up to $33\times$ fewer parameters than 5-member Deep Ensembles. It substantially improves OOD detection and often improves calibration relative to mean-field and perturbation baselines, while Deep Ensembles can still be stronger on in-distribution likelihood-based metrics.
△ Less
Submitted 19 July, 2026; v1 submitted 30 January, 2026;
originally announced February 2026.
-
Paving the Road to the Habitable Worlds Observatory with High-Resolution Imaging I: New and Archival Speckle Observations of Potential HWO Target Stars
Authors:
Zachary D. Hartman,
Catherine A. Clark,
Michael B. Lund,
Kathryn V. Lester,
José A. Caballero,
Steve B. Howell,
David Ciardi,
Sarah Deveny,
Mark E. Everett,
Elise Furlan,
Venu Kalari,
Colin Littlefield,
Andrew W. Stephens,
Jennifer A. Burt,
Guillaume Huber,
Rachel Matson,
Eric E. Mamajek,
Noah Tuchow
Abstract:
One of the key goals of the Habitable Worlds Observatory (HWO) is to directly image about 25 potentially habitable exoplanets and determine their properties. This challenge will require a large survey of nearby, bright stars -- ~100 according to the Astro2020 Decadel Survey. To ensure the success of the mission and to help guide design decisions, the stellar multiplicity of the target stars must b…
▽ More
One of the key goals of the Habitable Worlds Observatory (HWO) is to directly image about 25 potentially habitable exoplanets and determine their properties. This challenge will require a large survey of nearby, bright stars -- ~100 according to the Astro2020 Decadel Survey. To ensure the success of the mission and to help guide design decisions, the stellar multiplicity of the target stars must be well-understood. To this end, we present optical speckle imaging of stars in the NASA Exoplanet Exploration Program (ExEP) provisional HWO star list, which is currently the Tier 1 target list for the HWO Target Stars and Systems Sub-Working Group. We obtained new observations using `Alopeke and Zorro at Gemini Observatory and queried the Exoplanet Follow-up Observing Program Archive for archival observations, resulting in speckle imaging data for 80 of the 164 stars. We confirmed one candidate companion detected previously by Gaia (HD 90089) and obtained an ambiguous detection of a known companion (HD 212330). To examine our sensitivity to companions, we simulated stellar companions down to ~0.1 $M_{\odot}$ for each target and found that 75%-85% would be detected in our speckle images; the remaining simulated companions are either too faint or too close-in, and will require follow-up using other methods such as long-term spectroscopic measurements and space-based techniques. This work represents a first step towards surveying potential HWO targets for close-in stellar companions and helping to inform the target selection process for the HWO direct-imaging survey, bringing us closer towards the discovery of potential habitable worlds.
△ Less
Submitted 8 January, 2026;
originally announced January 2026.
-
Information Borrowing from Partially Compatible Trajectories for Estimation of Dynamic Treatment Regimes
Authors:
Chloe Si,
David A. Stephens,
Erica E. M. Moodie
Abstract:
Dynamic Treatment Regimes (DTRs) provide a systematic framework for optimizing sequential decision-making in chronic disease management, where therapies must adapt to patients' evolving clinical profiles. Inverse probability weighting (IPW) is a cornerstone methodology for estimating regime values from observational data due to its intuitive formulation and established theoretical properties, yet…
▽ More
Dynamic Treatment Regimes (DTRs) provide a systematic framework for optimizing sequential decision-making in chronic disease management, where therapies must adapt to patients' evolving clinical profiles. Inverse probability weighting (IPW) is a cornerstone methodology for estimating regime values from observational data due to its intuitive formulation and established theoretical properties, yet standard IPW estimators face significant limitations, including variance instability and data inefficiency. A fundamental but underexplored source of inefficiency lies in the strict alignment requirement between observed and target treatment trajectories, which fails to account for partial compatibility and discards substantial information from individuals with only minimal deviations from the regime. We propose two novel methodologies that relax the strict inclusion rule through flexible compatibility mechanisms. Both methods provide computationally tractable alternatives that can be easily integrated into existing IPW workflows, offering more efficient approaches to DTR estimation. Theoretical analysis demonstrates that both estimators preserve consistency while achieving superior finite-sample efficiency compared to standard IPW, and comprehensive simulation studies confirm improved stability. We illustrate the practical utility of our methods through an application to HIV treatment data from the AIDS Clinical Trials Group Study 175 (ACTG175).
△ Less
Submitted 25 March, 2026; v1 submitted 10 December, 2025;
originally announced December 2025.
-
Individualized treatment regimens under correlated data with multiple outcomes
Authors:
Misha Dolmatov,
Erica E. M. Moodie,
David A. Stephens,
Dipankar Bandyopadhyay
Abstract:
Precision medicine involves developing individualized treatment regimes (ITRs) which allow for treatment decisions to be tailored to patient characteristics. Naturally, the identification of the optimal regime, that is, the rule which maximizes patient outcomes, is of interest. Several procedures for estimating optimal ITRs from observational data have been proposed; however, relatively few method…
▽ More
Precision medicine involves developing individualized treatment regimes (ITRs) which allow for treatment decisions to be tailored to patient characteristics. Naturally, the identification of the optimal regime, that is, the rule which maximizes patient outcomes, is of interest. Several procedures for estimating optimal ITRs from observational data have been proposed; however, relatively few methods exist for estimating optimal ITRs in the presence of competing risks. Previous approaches either target one particular cause of failure, or rely on singly-robust estimators. We propose a novel doubly-robust regression-based method for estimating optimal ITRs which accounts for the uncertainty related to the unobserved cause of failure by averaging over all possible causes, or targeting the most likely cause. Our approach is straightforward to implement, and we demonstrate an extension to incorporate clustering, motivated by the question of for whom kidney transplantation with hepatitis C virus (HCV)-positive donors is safe, using data from the Organ Procurement and Transplantation Network. Our analysis suggests that a large portion of HCV-negative kidney recipients would see their overall survival unchanged if they were instead provided a kidney from an HCV-positive donor. The estimated treatment rules could be used to provide more efficient allocation of HCV-positive kidneys, increasing the donor pool.
△ Less
Submitted 26 September, 2025;
originally announced September 2025.
-
Multivariate regression with missing response data for modelling regional DNA methylation QTLs
Authors:
Shomoita Alam,
Yixiao Zeng,
Sasha Bernatsky,
Marie Hudson,
Inés Colmegna,
David A. Stephens,
Celia M. T. Greenwood,
Archer Y. Yang
Abstract:
Identifying genetic regulators of DNA methylation (mQTLs) with multivariate models enhances statistical power, but is challenged by missing data from bisulfite sequencing. Standard imputation-based methods can introduce bias, limiting reliable inference. We propose \texttt{missoNet}, a novel convex estimation framework that jointly estimates regression coefficients and the precision matrix from da…
▽ More
Identifying genetic regulators of DNA methylation (mQTLs) with multivariate models enhances statistical power, but is challenged by missing data from bisulfite sequencing. Standard imputation-based methods can introduce bias, limiting reliable inference. We propose \texttt{missoNet}, a novel convex estimation framework that jointly estimates regression coefficients and the precision matrix from data with missing responses. By using unbiased surrogate estimators, our three-stage procedure avoids imputation while simultaneously performing variable selection and learning the conditional dependence structure among responses. We establish theoretical error bounds, and our simulations demonstrate that \texttt{missoNet} consistently outperforms existing methods in both prediction and sparsity recovery. In a real-world mQTL analysis of the CARTaGENE cohort, \texttt{missoNet} achieved superior predictive accuracy and false-discovery control on a held-out validation set, identifying known and credible novel genetic associations. The method offers a robust, efficient, and theoretically grounded tool for genomic analyses, and is available as an R package.
△ Less
Submitted 8 July, 2025;
originally announced July 2025.
-
exoatlas: friendly Python code for exoplanet populations
Authors:
Zach K. Berta-Thompson,
Patcharapol Wachiraphan,
Autumn Stephens,
Mirielle Caradonna,
Catriona Murray,
Valerie Arriero,
Jackson Avery,
Girish M. Duvvuri,
Sebastian Pineda
Abstract:
Planets are complicated. Understanding how they work requires connecting individual objects to the context of broader populations. Exoplanets are easier to picture next to their closest Solar System archetypes, and planets in the Solar System are richer when seen alongside a growing community of known exoplanets in the Milky Way. The `exoatlas` toolkit provides a friendly Python interface for retr…
▽ More
Planets are complicated. Understanding how they work requires connecting individual objects to the context of broader populations. Exoplanets are easier to picture next to their closest Solar System archetypes, and planets in the Solar System are richer when seen alongside a growing community of known exoplanets in the Milky Way. The `exoatlas` toolkit provides a friendly Python interface for retrieving and working with populations of planets, aiming to simplify the process of placing worlds in context.
△ Less
Submitted 2 July, 2025;
originally announced July 2025.
-
Near-Infrared Spectroscopy with IGRINS-2 for Studying Multiple Stellar Populations in Globular Clusters
Authors:
Dongwook Lim,
Young-Wook Lee,
Sol Yun,
Young Sun Lee,
Sang-Hyun Chun,
Heeyoung Oh,
Jae-Joon Lee,
Chan Park,
Sanghyuk Kim,
Ueejeong Jeong,
Hye-In Lee,
Woojin Park,
Youngsam Yu,
Yunjong Kim,
Moo-Young Chun,
Jae Sok Oh,
Sungho Lee,
Jeong-Gyun Jang,
Bi-Ho Jang,
Hyeon Cheol Seong,
Hyun-Jeong Kim,
Cynthia B. Brooks,
Gregory N. Mace,
Hanshin Lee,
John M. Good
, et al. (31 additional authors not shown)
Abstract:
Recent advancements in near-infrared (NIR) spectroscopy have opened new opportunities for studying multiple stellar populations in globular clusters (GCs), particularly for newly discovered clusters in the inner Milky Way. While optical spectroscopy has traditionally played a primary role in detailed chemical abundance studies of GCs, the increasing discovery of GCs in highly reddened environments…
▽ More
Recent advancements in near-infrared (NIR) spectroscopy have opened new opportunities for studying multiple stellar populations in globular clusters (GCs), particularly for newly discovered clusters in the inner Milky Way. While optical spectroscopy has traditionally played a primary role in detailed chemical abundance studies of GCs, the increasing discovery of GCs in highly reddened environments underscores the need for robust NIR spectroscopic methods. To evaluate the utility of high-resolution NIR spectroscopy for studying multiple stellar populations, we observed six stars in M5, a well-studied halo GC, using the recently commissioned IGRINS-2 spectrograph on the Gemini-North telescope. Our chemical abundance measurements in the NIR wavelength range show good agreement with those derived from high-resolution optical spectroscopy, with minor systematic offsets in elements such as Na and Mg. In addition, the measured chemical abundance ratios clearly reproduce the distinctive patterns of multiple stellar populations, including the Na-O anti-correlation. The ability of NIR spectroscopy to measure C, N, and O abundances with high precision further enhances its utility for studying chemical properties of stars and GCs. Our findings demonstrate that IGRINS-2 and similar instruments have significant potential to advance our understanding of GC formation, stellar chemical evolution, and the evolutionary history of the Milky Way.
△ Less
Submitted 3 April, 2025;
originally announced April 2025.
-
An Assessment of the UK Government Clean Energy Strategy for the Year 2030
Authors:
Anthony D. Stephens,
David R. Walwyn
Abstract:
In 2024, the UK Government made two striking announcements on its plans to decarbonise the energy system; it pledged GBP22 billion to establish carbon capture and storage hubs on Teesside and Merseyside and released the Clean Power 2030 Action Plan. This paper questions the validity of both plans, arguing that they do not take adequate account of the consequences of the highly variable nature of w…
▽ More
In 2024, the UK Government made two striking announcements on its plans to decarbonise the energy system; it pledged GBP22 billion to establish carbon capture and storage hubs on Teesside and Merseyside and released the Clean Power 2030 Action Plan. This paper questions the validity of both plans, arguing that they do not take adequate account of the consequences of the highly variable nature of wind and solar generations. Using dynamic models of future UK electricity systems which are designed to take account of these variabilities, it is shown that the Clean Power 2030 Action Plan overestimates the ability of wind and solar generations to decarbonise the electricity system as they increase in size relative to the demand of the electricity system. More importantly, the dynamic models show that most of the achievable decarbonization is the result of increasing wind generation from the current level of around 10 GW to around 20 GW. Increasing wind generation to only 20 GW, rather than to 30 GW as proposed in the Action Plan, should halve the proposed cost, a saving of perhaps GBP 120 billion, with little disbenefit in terms of reduced decarbonization. Furthermore, the dynamic modelling shows that UK gas storage capacity of 7.5 winter days looks hopeless inadequate in comparison with the storage capacities deemed necessary by its continental neighbors. Concern is expressed that a consequence of the Climate Change Act of 2008 requiring the UK to meet arbitrary decarbonization targets is leading government advisors to propose several unproven and therefore highly risky technological solutions.
△ Less
Submitted 18 March, 2025;
originally announced March 2025.
-
An Early Look at the Performance of IGRINS-2 at Gemini-North with Application to the ultrahot Jupiter, WASP-33 b
Authors:
Yeon-Ho Choi,
Ueejeong Jeong,
Jae-Joon Lee,
Hyun-Jeong Kim,
Heeyoung Oh,
Chan Park,
Changwoo Kye,
Luke Finnerty,
Micheal R. Line,
Krishna Kanumalla,
Jorge A. Sanchez,
Peter C. B. Smith,
Sanghyuk Kim,
Hye-In Lee,
Woojin Park,
Youngsam Yu,
Yunjong Kim,
Moo-Young Chun,
Jae Sok Oh,
Sungho Lee,
Jeong-Gyun Jang,
Bi-Ho Jang,
Hyeon Cheol Seong,
Cynthia B. Brooks,
Gregory N. Mace
, et al. (34 additional authors not shown)
Abstract:
Ground-based high-resolution spectroscopy enables precise molecular detections and velocity-resolved atmospheric dynamics, offering a distinct advantage over low-resolution methods for exoplanetary atmospheric studies. IGRINS-2, the successor to IGRINS, features improved throughput and enhanced sensitivity to carbon monoxide by shifting its $\textit{K}$-band coverage by 36 nm to longer wavelengths…
▽ More
Ground-based high-resolution spectroscopy enables precise molecular detections and velocity-resolved atmospheric dynamics, offering a distinct advantage over low-resolution methods for exoplanetary atmospheric studies. IGRINS-2, the successor to IGRINS, features improved throughput and enhanced sensitivity to carbon monoxide by shifting its $\textit{K}$-band coverage by 36 nm to longer wavelengths. IGRINS is a near-infrared high-resolution spectrograph mounted at McDonald, Lowell, and Gemini-South observatories. Our order-drop test shows this added range improves the CO cross-correlation signal-to-noise ratio (SNR) by 2$-$3%, confirming a measurable but modest sensitivity gain. To evaluate its performance, we attempt to investigate the atmospheric characteristics of WASP-33 b. Observations were conducted on 2024 January 7 for a total of 2.43 hours; This includes 1.46 hours in the pre-eclipse phase to capture the planet's thermal emission spectrum. We successfully detect clear cross-correlation signals from molecular species in the dayside atmosphere of WASP-33 b with a combined SNR of 7.4. More specifically, we capture CO, H$_{2}$O, and OH with SNRs of 6.3, 4.7, and 4.2, respectively. These results are consistent with previous studies and demonstrate that IGRINS-2 is well-suited for detailed investigation of exoplanetary atmospheres. We anticipate that future observations with IGRINS-2 will further advance our understanding of exoplanetary atmospheres.
△ Less
Submitted 20 June, 2025; v1 submitted 16 March, 2025;
originally announced March 2025.
-
Keck and Gemini characterization of $Hayabusa2\#$ rendezvous target 1998 KY$_{26}$
Authors:
Bryce T. Bolin,
Christoffer Fremling,
Matthew Belyakov,
Jin Beniyama,
Marco Delbo,
Robert Jedicke,
Ian Wong,
Laura-May Abron,
Keith S. Noll,
Andrew W. Stephens
Abstract:
Near-earth object (NEO) 1998 KY$_{26}$ is a target of the $Hayabusa2\#$ spacecraft, which it will rendezvous with in July 2031. The asteroid is a rapid rotator and has a large out-of-plane nongravitational acceleration. We present deep $g$ and $R$ band imaging obtained with the Keck I/Low Resolution Imaging Spectrometer and visible spectroscopy from Gemini North/Gemini Multi-Object Spectrograph ta…
▽ More
Near-earth object (NEO) 1998 KY$_{26}$ is a target of the $Hayabusa2\#$ spacecraft, which it will rendezvous with in July 2031. The asteroid is a rapid rotator and has a large out-of-plane nongravitational acceleration. We present deep $g$ and $R$ band imaging obtained with the Keck I/Low Resolution Imaging Spectrometer and visible spectroscopy from Gemini North/Gemini Multi-Object Spectrograph taken of 1998 KY$_{26}$ on 2024 June 8-9 when the asteroid was $\sim$0.037 au from the Earth. The asteroid lacks evidence of a dust coma in the deep images and its spectrum most closely resembles Xe-type asteroids, possessing a spectral slope of 6.71$\pm$0.43 $\%$ 100 nm$^{-1}$, and colors $g$-$r$ = 0.63$\pm$0.03, $r$-$i$ = 0.15$\pm$0.03, $i$-$z$ = 0.05$\pm$0.04, and implies a diameter of $\sim$10 m. From our images, we compute a 3$σ$ upper limit on the dust production of 1998 KY$_{26}$ of $<$10$^{-5}$ kg s$^{-1}$, $<$10$^{-2}$ kg s$^{-1}$, and $<$10$^{-1}$ kg s$^{-1}$ assuming $\mathrmμ$m, mm, and cm size dust particles. Additionally, we compare the orbit of 1998 KY$_{26}$ and large nongravitational parameters asteroids to NEO population models and find that the majority, including 1998 KY$_{26}$, likely originated from the inner Main Belt, while the second most numerous group originates from the outer Main Belt, followed by a third group originating from the Jupiter Family Comet population. Given its inner Main Belt origin, its Xe-type spectrum, and rapid rotation, we hypothesize that the nongravitational acceleration of 1998 KY$_{26}$ may be caused by the shedding of large dust grains from its surface due to its rotation rather than H$_2$O vapor outgassing.
△ Less
Submitted 12 April, 2025; v1 submitted 28 January, 2025;
originally announced January 2025.
-
Hydroxyl Lines and Moonlight: a High Spectral Resolution Investigation of NIR skylines from Maunakea to guide NIR spectroscopic surveys
Authors:
Frederick Dauphin,
Andreea Petric,
Étienne Artigau,
Andrew W. Stephens,
Neil James Cook,
Steven Businger,
Nicolas Flagey,
Jennifer Marshall,
Michelle Ntampaka,
Swara Ravindranath,
Laurie Rousseau-Nepton
Abstract:
Subtracting the changing sky contribution from the near-infrared (NIR) spectra of faint astronomical objects is challenging and crucial to a wide range of science cases such as estimating the velocity dispersions of dwarf galaxies, studying the gas dynamics in faint galaxies, measuring accurate redshifts, and any spectroscopic studies of faint targets. Since the sky background varies with time and…
▽ More
Subtracting the changing sky contribution from the near-infrared (NIR) spectra of faint astronomical objects is challenging and crucial to a wide range of science cases such as estimating the velocity dispersions of dwarf galaxies, studying the gas dynamics in faint galaxies, measuring accurate redshifts, and any spectroscopic studies of faint targets. Since the sky background varies with time and location, NIR spectral observations, especially those employing fiber spectrometers and targeting extended sources, require frequent sky-only observations for calibration. However, sky subtraction can be optimized with sufficient a priori knowledge of the sky's variability. In this work, we explore how to optimize sky subtraction by analyzing 1075 high-resolution NIR spectra from the CFHT's SPIRou on Maunakea, and we estimate the variability of 481 hydroxyl (OH) lines. These spectra were collected during two sets of three nights dedicated to obtaining sky observations every five and a half minutes. During the first set, we observed how the Moon affects the NIR, which has not been accurately measured at these wavelengths. We suggest accounting for the Moon contribution at separation distances less than 10 degrees when 1) reconstructing the sky using principal component analysis 2) observing targets at Y JHK mags fainter than ~15 and 3) attempting a sky subtraction better than 1%. We also identified 126 spectral doublets, or OH lines that split into at least two components, at SPIRou's resolution. In addition, we used Lomb-Scargle Periodograms and Gaussian process regression to estimate that most OH lines vary on similar timescales, which provides a valuable input for IR spectroscopic survey strategies. The data and code developed for this study are publicly available.
△ Less
Submitted 6 December, 2024;
originally announced December 2024.
-
Bayesian measurement error modeling of latent time series structure to assess the impact of pollutants on health
Authors:
Yanfei Qu,
David A. Stephens
Abstract:
The association between levels of air pollution and mortality rate is well-established, but quantifying the magnitude of the effect is sometimes complicated by limitations in the data. In this paper, a joint Bayesian hierarchical model is developed to identify significant predictor factors and quantify their impacts on mortalities. To account for potential measurement error in the pollution data,…
▽ More
The association between levels of air pollution and mortality rate is well-established, but quantifying the magnitude of the effect is sometimes complicated by limitations in the data. In this paper, a joint Bayesian hierarchical model is developed to identify significant predictor factors and quantify their impacts on mortalities. To account for potential measurement error in the pollution data, the observed pollutant levels were treated as noisy proxies for the true exposure, therefore requiring a measurement error structure. We illustrate the developed model by performing an analysis of the association between weekly air pollution levels and cardiovascular and respiratory mortality in Los Angeles (LA) County over five years from January 2018 to December 2022. The purpose of the study was to quantify the impact of the main air pollutants (PM2.5, PM10, SO2, NO2, CO, and O3) and of temperature on weekly mortality from four causes: chronic obstructive pulmonary diseases, pneumonia, heart failures, and malignant neoplasms. The results, supported by Bayesian model comparison criteria, indicated that weekly county-level cause-specific mortalities were significantly associated with certain ranges of air pollutants, and that some of these cause-specific mortalities had significant associations with ambient temperature levels. The analysis revealed that several pollutants appeared to be associated with lower mortality; we interpreted this counterintuitive finding from various perspectives.
△ Less
Submitted 10 August, 2026; v1 submitted 1 October, 2024;
originally announced October 2024.
-
Wind lulls and slews; consequences for the stability of future UK electricity systems
Authors:
Anthony D Stephens,
David R Walwyn
Abstract:
As the United Kingdom wind fleet increases in size, wind lulls and slews will increasingly challenge the stability of its electricity system. The paper describes the use of models based on real time records and including solar slews, to investigate the most extreme wind variations likely to be encountered in future, enabling strategies to be devised to mitigate them. Wind lulls are surprisingly fr…
▽ More
As the United Kingdom wind fleet increases in size, wind lulls and slews will increasingly challenge the stability of its electricity system. The paper describes the use of models based on real time records and including solar slews, to investigate the most extreme wind variations likely to be encountered in future, enabling strategies to be devised to mitigate them. Wind lulls are surprisingly frequent, occasionally lasting a week or more, and are always likely to be beyond the capabilities of stored or imported electrical energy to mitigate them. The models indicate that there will be a continuing need for gas powered generation to mitigate wind lulls. Currently, Combined Cycle Gas Turbines (CCGTs) provide most of the dispatchable generation. However, CCGTs are not sufficiently fast acting to cope with the wind and solar slews anticipated in future. The paper suggests that a range of already proven fast-acting sources of dispatchable generation, including Open Cycle Gas Turbines (OCGTs), Internal Combustion Gas-Fired Reciprocating engines (ICGRs) and stored electrical energy systems, should be capable of coping with the largest wind and solar slews likely to be encountered up to the year 2035. Examples are given of the recent introduction of these fast-acting sources of generation which, it is suggested, will progressively replace CCGTs as the wind and solar fleets increase in size. Moreover, we see the pattern of recent investments, summarised in the paper, as a good indication of likely future investments, with OCGT investments mainly serving the 440 kV grid, and ICGRs and stored electrical energy more local networks.
△ Less
Submitted 24 September, 2024;
originally announced September 2024.
-
Planet Hunters NGTS: New Planet Candidates from a Citizen Science Search of the Next Generation Transit Survey Public Data
Authors:
Sean M. O'Brien,
Megan E. Schwamb,
Samuel Gill,
Christopher A. Watson,
Matthew R. Burleigh,
Alicia Kendall,
David R. Anderson,
José I. Vines,
James S. Jenkins,
Douglas R. Alves,
Laura Trouille,
Solène Ulmer-Moll,
Edward M. Bryant,
Ioannis Apergis,
Matthew P. Battley,
Daniel Bayliss,
Nora L. Eisner,
Edward Gillen,
Michael R. Goad,
Maximilian N. Günther,
Beth A. Henderson,
Jeong-Eun Heo,
David G. Jackson,
Chris Lintott,
James McCormac
, et al. (13 additional authors not shown)
Abstract:
We present the results from the first two years of the Planet Hunters NGTS citizen science project, which searches for transiting planet candidates in data from the Next Generation Transit Survey (NGTS) by enlisting the help of members of the general public. Over 8,000 registered volunteers reviewed 138,198 light curves from the NGTS Public Data Releases 1 and 2. We utilize a user weighting scheme…
▽ More
We present the results from the first two years of the Planet Hunters NGTS citizen science project, which searches for transiting planet candidates in data from the Next Generation Transit Survey (NGTS) by enlisting the help of members of the general public. Over 8,000 registered volunteers reviewed 138,198 light curves from the NGTS Public Data Releases 1 and 2. We utilize a user weighting scheme to combine the classifications of multiple users to identify the most promising planet candidates not initially discovered by the NGTS team. We highlight the five most interesting planet candidates detected through this search, which are all candidate short-period giant planets. This includes the TIC-165227846 system that, if confirmed, would be the lowest-mass star to host a close-in giant planet. We assess the detection efficiency of the project by determining the number of confirmed planets from the NASA Exoplanet Archive and TESS Objects of Interest (TOIs) successfully recovered by this search and find that 74% of confirmed planets and 63% of TOIs detected by NGTS are recovered by the Planet Hunters NGTS project. The identification of new planet candidates shows that the citizen science approach can provide a complementary method to the detection of exoplanets with ground-based surveys such as NGTS.
△ Less
Submitted 23 April, 2024;
originally announced April 2024.
-
The Development of Investment Planning Models for the United Kingdoms Wind and Solar Fleets
Authors:
Anthony D Stephens,
David R Walwyn
Abstract:
Previous work has resulted in the development of an energy model able to calculate wind and solar fleet efficiencies. However, for investment planning purposes, it is necessary to calculate from the lowest economically acceptable efficiencies how much wind and solar generation would be economically justified. The paper explains how this objective has been achieved with arrays (investment planning…
▽ More
Previous work has resulted in the development of an energy model able to calculate wind and solar fleet efficiencies. However, for investment planning purposes, it is necessary to calculate from the lowest economically acceptable efficiencies how much wind and solar generation would be economically justified. The paper explains how this objective has been achieved with arrays (investment planning tables) created after carrying out a structured investigation of the behaviour of the electricity system over the whole of its operational range. The tables are then applied to National Grid prediction of the size and composition of the system in the year 2035. A conclusion is reached that wind and solar generation will only be able to supply about 70% of electrical demand, the other 30% being provided by dispatchable sources of generation, which must be sufficiently fast acting to maintain electricity system stability, such as the use of combined cycle gas turbines. This limit on deployment of wind and solar generation restricts their ability to decarbonise the electricity system and is likely to lead in 2035 to a residual of 72 million tonnes per annum of carbon dioxide emissions which wind and solar generations will be unable to address
△ Less
Submitted 14 March, 2024;
originally announced March 2024.
-
Computational Considerations for the Linear Model of Coregionalization
Authors:
Renaud Alie,
David A. Stephens,
Alexandra M. Schmidt
Abstract:
In the last two decades, the linear model of coregionalization (LMC) has been widely used to model multivariate spatial processes. However, it can be a challenging task to conduct likelihood-based inference for such models because of the cubic cost associated with Gaussian likelihood evaluations. Starting from an analogy with matrix normal models, we propose a reformulation of the LMC likelihood t…
▽ More
In the last two decades, the linear model of coregionalization (LMC) has been widely used to model multivariate spatial processes. However, it can be a challenging task to conduct likelihood-based inference for such models because of the cubic cost associated with Gaussian likelihood evaluations. Starting from an analogy with matrix normal models, we propose a reformulation of the LMC likelihood that highlights the linear, rather than cubic, computational complexity as a function of the dimension of the response vector. We describe how those simplifications can be exploited in Gaussian hierarchical models. In addition, we propose a new sparsity-inducing approach to the LMC that introduces structural zeros in the coregionalization matrix in an attempt to reduce the number of parameters in a principled and data-driven way. Our reformulation of the LMC likelihood ensures that our sparse approach comes at virtually no additional cost when included in a Markov chain Monte Carlo (MCMC) algorithm. It is shown, on synthetic data, to significantly improve predictive performance. We also apply our methodology to a dataset comprised of air pollutant measurements from the state of California. We investigate the strength of the correlation among the measurements by providing new insights from our sparse method.
△ Less
Submitted 2 December, 2024; v1 submitted 13 February, 2024;
originally announced February 2024.
-
Population Graph Cross-Network Node Classification for Autism Detection Across Sample Groups
Authors:
Anna Stephens,
Francisco Santos,
Pang-Ning Tan,
Abdol-Hossein Esfahanian
Abstract:
Graph neural networks (GNN) are a powerful tool for combining imaging and non-imaging medical information for node classification tasks. Cross-network node classification extends GNN techniques to account for domain drift, allowing for node classification on an unlabeled target network. In this paper we present OTGCN, a powerful, novel approach to cross-network node classification. This approach l…
▽ More
Graph neural networks (GNN) are a powerful tool for combining imaging and non-imaging medical information for node classification tasks. Cross-network node classification extends GNN techniques to account for domain drift, allowing for node classification on an unlabeled target network. In this paper we present OTGCN, a powerful, novel approach to cross-network node classification. This approach leans on concepts from graph convolutional networks to harness insights from graph data structures while simultaneously applying strategies rooted in optimal transport to correct for the domain drift that can occur between samples from different data collection sites. This blended approach provides a practical solution for scenarios with many distinct forms of data collected across different locations and equipment. We demonstrate the effectiveness of this approach at classifying Autism Spectrum Disorder subjects using a blend of imaging and non-imaging data.
△ Less
Submitted 10 January, 2024;
originally announced January 2024.
-
Symmetry Enforced Fermi Surface Degeneracies Observed in Time-Reversal Symmetry-Breaking Superconductor LaNiGa$_2$
Authors:
Matthew Staab,
Robert Prater,
Sudheer Sreedhar,
Journey Byland,
Eliana Mann,
Davis Zackaria,
Yunshu Shi,
Henry J. Bowman,
Andrew L. Stephens,
Myung-Chul Jung,
Antia S. Botana,
Warren E. Pickett,
Valentin Taufour,
Inna Vishik
Abstract:
LaNiGa$_2$ is superconductor that breaks time-reversal symmetry in the superconducting state without any known nearby magnetism. Recently, single crystals of LaNiGa$_2$ have been synthesized, revealing a nonsymmorphic Cmcm space group. Here, we report measurements of the electronic structure of LaNiGa$_2$ throughout the three-dimensional Brillouin zone (BZ) using angle-resolved photoemission spect…
▽ More
LaNiGa$_2$ is superconductor that breaks time-reversal symmetry in the superconducting state without any known nearby magnetism. Recently, single crystals of LaNiGa$_2$ have been synthesized, revealing a nonsymmorphic Cmcm space group. Here, we report measurements of the electronic structure of LaNiGa$_2$ throughout the three-dimensional Brillouin zone (BZ) using angle-resolved photoemission spectroscopy (ARPES). Our findings show broad consistency with density functional theory (DFT) calculations and provide evidence for degeneracies in the electronic structure that are predicted from the space group. The calculations also predict four Fermi surfaces which cross the purported nodal plane and should therefore form two degenerate pairs. We report evidence for those predicted symmetry enforced degeneracies as well as accidental near degeneracies throughout the BZ. These degeneracies and near-degeneracies may play a role in the pairing mechanism of LaNiGa$_2$. Our results provide insight into the interplay between structure, Fermiology, and superconductivity in unconventional superconductors with nonsymmorphic space group.
△ Less
Submitted 18 December, 2023;
originally announced December 2023.
-
Linking Symptom Inventories using Semantic Textual Similarity
Authors:
Eamonn Kennedy,
Shashank Vadlamani,
Hannah M Lindsey,
Kelly S Peterson,
Kristen Dams OConnor,
Kenton Murray,
Ronak Agarwal,
Houshang H Amiri,
Raeda K Andersen,
Talin Babikian,
David A Baron,
Erin D Bigler,
Karen Caeyenberghs,
Lisa Delano-Wood,
Seth G Disner,
Ekaterina Dobryakova,
Blessen C Eapen,
Rachel M Edelstein,
Carrie Esopenko,
Helen M Genova,
Elbert Geuze,
Naomi J Goodrich-Hunsaker,
Jordan Grafman,
Asta K Haberg,
Cooper B Hodges
, et al. (57 additional authors not shown)
Abstract:
An extensive library of symptom inventories has been developed over time to measure clinical symptoms, but this variety has led to several long standing issues. Most notably, results drawn from different settings and studies are not comparable, which limits reproducibility. Here, we present an artificial intelligence (AI) approach using semantic textual similarity (STS) to link symptoms and scores…
▽ More
An extensive library of symptom inventories has been developed over time to measure clinical symptoms, but this variety has led to several long standing issues. Most notably, results drawn from different settings and studies are not comparable, which limits reproducibility. Here, we present an artificial intelligence (AI) approach using semantic textual similarity (STS) to link symptoms and scores across previously incongruous symptom inventories. We tested the ability of four pre-trained STS models to screen thousands of symptom description pairs for related content - a challenging task typically requiring expert panels. Models were tasked to predict symptom severity across four different inventories for 6,607 participants drawn from 16 international data sources. The STS approach achieved 74.8% accuracy across five tasks, outperforming other models tested. This work suggests that incorporating contextual, semantic information can assist expert decision-making processes, yielding gains for both general and disease-specific clinical assessment.
△ Less
Submitted 8 September, 2023;
originally announced September 2023.
-
The SN 2023ixf Progenitor in M101: II. Properties
Authors:
Schuyler D. Van Dyk,
Sundar Srinivasan,
Jennifer E. Andrews,
Monika Soraisam,
Tamas Szalai,
Steve B. Howell,
Howard Isaacson,
Thomas Matheson,
Erik Petigura,
Peter Scicluna,
Andrew W. Stephens,
Judah Van Zandt,
WeiKang Zheng,
Sang-Hyun Chun,
Alexei V. Filippenko
Abstract:
We follow our first paper with an analysis of the ensemble of the extensive pre-explosion ground- and space-based infrared observations of the red supergiant (RSG) progenitor candidate for the nearby core-collapse supernova SN 2023ixf in Messier 101, together with optical data prior to explosion obtained with the Hubble Space Telescope (HST). We have confirmed the association of the progenitor can…
▽ More
We follow our first paper with an analysis of the ensemble of the extensive pre-explosion ground- and space-based infrared observations of the red supergiant (RSG) progenitor candidate for the nearby core-collapse supernova SN 2023ixf in Messier 101, together with optical data prior to explosion obtained with the Hubble Space Telescope (HST). We have confirmed the association of the progenitor candidate with the SN, as well as constrained the metallicity at the SN site, based on SN observations with instruments at Gemini-North. The internal host extinction to the SN has also been confirmed from a high-resolution Keck spectrum. We fit the observed spectral energy distribution (SED) for the star, accounting for its intrinsic variability, with dust radiative-transfer modeling, which assume a silicate-rich dust shell ahead of the underlying stellar photosphere. The star is heavily dust-obscured, likely the dustiest progenitor candidate yet encountered. We found median estimates of the star's effective temperature and luminosity of 2770 K and 9.0e4 L_Sun, with 68% credible intervals of 2340--3150 K and (7.5--10.9)e4 L_sun. The candidate may have a Galactic RSG analog, IRC -10414, with a strikingly similar SED and luminosity. Via comparison with single-star evolutionary models we have constrained the initial mass of the progenitor candidate from 12 M_sun to as high as 14 M_sun. We have had available to us an extraordinary view of the SN 2023ixf progenitor candidate, which should be further followed up in future years with HST and the James Webb Space Telescope.
△ Less
Submitted 23 April, 2024; v1 submitted 28 August, 2023;
originally announced August 2023.
-
Structured Analysis Reveals Fundamental Mathematical Relationships between Wind and Solar Generations and the United Kingdom Electricity System
Authors:
Anthony D Stephens,
David R Walwyn
Abstract:
The use of wind and solar generation is fundamental to the decarbonisation of the United Kingdom electricity system. However, the optimal level of renewable energy as a proportion of total demand is still being debated. In this paper, several models, whose aims are to predict the efficiency of future system configurations, are explained. The models use historic records from the Gridwatch website f…
▽ More
The use of wind and solar generation is fundamental to the decarbonisation of the United Kingdom electricity system. However, the optimal level of renewable energy as a proportion of total demand is still being debated. In this paper, several models, whose aims are to predict the efficiency of future system configurations, are explained. The models use historic records from the Gridwatch website for the year 2017, which are then scaled accordingly. The model predictions are first demonstrated for the 2035 Scenario as proposed by the National Grid in FES 2022. The analysis reveals that at least one third of the available wind and solar generation will exceed the ability of the electricity system to use it and will have to be shed. By defining an efficiency measure, the Marginal Decarbonisation Efficiency, which quantifies the incremental extent to which wind generation can decarbonise the electricity system, it is shown that the 2035 Scenario will have a low efficiency. Moreover, it will require the use of combined cycle gas turbines, which is at variants with the predictions of the National Grid steady state model. The paper also describes the derivation of a Generic Model, which allows the level of wind energy and dispatchable generation for all system configurations likely to be encountered in future decades, to be calculated without the use of computer models.
△ Less
Submitted 12 July, 2023;
originally announced July 2023.
-
Generalized Random Forests using Fixed-Point Trees
Authors:
David Fleischer,
David A. Stephens,
Archer Y. Yang
Abstract:
We propose a computationally efficient alternative to generalized random forests (GRFs) for estimating heterogeneous effects in large dimensions. While GRFs rely on a gradient-based splitting criterion, which in large dimensions is computationally expensive and unstable, our method introduces a fixed-point approximation that eliminates the need for Jacobian estimation. This gradient-free approach…
▽ More
We propose a computationally efficient alternative to generalized random forests (GRFs) for estimating heterogeneous effects in large dimensions. While GRFs rely on a gradient-based splitting criterion, which in large dimensions is computationally expensive and unstable, our method introduces a fixed-point approximation that eliminates the need for Jacobian estimation. This gradient-free approach preserves GRF's theoretical guarantees of consistency and asymptotic normality while significantly improving computational efficiency. We demonstrate that our method achieves a speedup of multiple times over standard GRFs without compromising statistical accuracy. Experiments on both simulated and real-world data validate our approach. Our findings suggest that the proposed method is a scalable alternative for localized effect estimation in machine learning and causal inference applications
△ Less
Submitted 16 June, 2025; v1 submitted 20 June, 2023;
originally announced June 2023.
-
The impact of directly observed therapy on the efficacy of Tuberculosis treatment: A Bayesian multilevel approach
Authors:
Widemberg S. Nobre,
Alexandra M. Schmidt,
Erica E. M. Moodie,
David A. Stephens
Abstract:
We propose and discuss a Bayesian procedure to estimate the average treatment effect (ATE) for multilevel observations in the presence of confounding. We focus on situations where the confounders may be latent (e.g., spatial latent effects). This work is motivated by an interest in determining the causal impact of directly observed therapy (DOT) on the successful treatment of Tuberculosis (TB); th…
▽ More
We propose and discuss a Bayesian procedure to estimate the average treatment effect (ATE) for multilevel observations in the presence of confounding. We focus on situations where the confounders may be latent (e.g., spatial latent effects). This work is motivated by an interest in determining the causal impact of directly observed therapy (DOT) on the successful treatment of Tuberculosis (TB); the available data correspond to individual-level information observed across different cities in a state in Brazil. We focus on propensity score regression and covariate adjustment to balance the treatment (DOT) allocation. We discuss the need to include latent local-level random effects in the propensity score model to reduce bias in the estimation of the ATE. A simulation study suggests that accounting for the multilevel nature of the data with latent structures in both the outcome and propensity score models has the potential to reduce bias in the estimation of causal effects.
△ Less
Submitted 24 April, 2023;
originally announced April 2023.
-
The two rings of (50000) Quaoar
Authors:
C. L. Pereira,
B. Sicardy,
B. E. Morgado,
F. Braga-Ribas,
E. Fernández-Valenzuela,
D. Souami,
B. J. Holler,
R. C. Boufleur,
G. Margoti,
M. Assafin,
J. L. Ortiz,
P. Santos-Sanz,
B. Epinat,
P. Kervella,
J. Desmars,
R. Vieira-Martins,
Y. Kilic,
A. R. Gomes-Júnior,
J. I. B. Camargo,
M. Emilio,
M. Vara-Lubiano,
M. Kretlow,
L. Albert,
C. Alcock,
J. G. Ball
, et al. (44 additional authors not shown)
Abstract:
Quaoar is a classical Trans-Neptunian Object (TNO) with an area equivalent diameter of 1,100 km and an orbital semi-major axis of 43.3 astronomical units. Based on stellar occultations observed between 2018 and 2021, an inhomogeneous ring (Q1R, Quaoar's first ring) was detected around this body. Aims. A new stellar occultation by Quaoar was observed on August 9th, 2022 aiming to improve Quaoar's s…
▽ More
Quaoar is a classical Trans-Neptunian Object (TNO) with an area equivalent diameter of 1,100 km and an orbital semi-major axis of 43.3 astronomical units. Based on stellar occultations observed between 2018 and 2021, an inhomogeneous ring (Q1R, Quaoar's first ring) was detected around this body. Aims. A new stellar occultation by Quaoar was observed on August 9th, 2022 aiming to improve Quaoar's shape models and the physical parameters of Q1R while searching for additional material around the body. Methods. The occultation provided nine effective chords across Quaoar, pinning down its size, shape, and astrometric position. Large facilities, such as Gemini North and the Canada-France-Hawaii Telescope (CFHT), were used to obtain high acquisition rates and signal-to-noise ratios. The light curves were also used to characterize the Q1R ring (radial profiles and orbital elements). Results. Quaoar's elliptical fit to the occultation chords yields the limb with an apparent semi-major axis of $579.5\pm4.0$ km, apparent oblateness of $0.12\pm0.01$, and area-equivalent radius of $543\pm2$ km. Quaoar's limb orientation is consistent with Q1R and Weywot orbiting in Quaoar's equatorial plane. The orbital radius of Q1R is refined to a value of $4,057\pm6$ km. The radial opacity profile of the more opaque ring profile follows a Lorentzian shape that extends over 60 km, with a full width at half maximum (FWHM) of $\sim5$ km and a peak normal optical depth of 0.4. Besides the secondary events related to the already reported rings, new secondary events detected during the August 2022 occultation in three different data sets are consistent with another ring around Quaoar with a radius of $2,520\pm20$ km, assuming the ring is circular and co-planar with Q1R. This new ring has a typical width of 10 km and a normal optical depth of $\sim$0.004. Like Q1R, it also lies outside Quaoar's classical Roche limit.
△ Less
Submitted 20 April, 2023; v1 submitted 18 April, 2023;
originally announced April 2023.
-
The James Webb Space Telescope Mission
Authors:
Jonathan P. Gardner,
John C. Mather,
Randy Abbott,
James S. Abell,
Mark Abernathy,
Faith E. Abney,
John G. Abraham,
Roberto Abraham,
Yasin M. Abul-Huda,
Scott Acton,
Cynthia K. Adams,
Evan Adams,
David S. Adler,
Maarten Adriaensen,
Jonathan Albert Aguilar,
Mansoor Ahmed,
Nasif S. Ahmed,
Tanjira Ahmed,
Rüdeger Albat,
Loïc Albert,
Stacey Alberts,
David Aldridge,
Mary Marsha Allen,
Shaune S. Allen,
Martin Altenburg
, et al. (983 additional authors not shown)
Abstract:
Twenty-six years ago a small committee report, building on earlier studies, expounded a compelling and poetic vision for the future of astronomy, calling for an infrared-optimized space telescope with an aperture of at least $4m$. With the support of their governments in the US, Europe, and Canada, 20,000 people realized that vision as the $6.5m$ James Webb Space Telescope. A generation of astrono…
▽ More
Twenty-six years ago a small committee report, building on earlier studies, expounded a compelling and poetic vision for the future of astronomy, calling for an infrared-optimized space telescope with an aperture of at least $4m$. With the support of their governments in the US, Europe, and Canada, 20,000 people realized that vision as the $6.5m$ James Webb Space Telescope. A generation of astronomers will celebrate their accomplishments for the life of the mission, potentially as long as 20 years, and beyond. This report and the scientific discoveries that follow are extended thank-you notes to the 20,000 team members. The telescope is working perfectly, with much better image quality than expected. In this and accompanying papers, we give a brief history, describe the observatory, outline its objectives and current observing program, and discuss the inventions and people who made it possible. We cite detailed reports on the design and the measured performance on orbit.
△ Less
Submitted 10 April, 2023;
originally announced April 2023.
-
Bayesian inference for optimal dynamic treatment regimes in practice
Authors:
Daniel Rodriguez Duque,
Erica E. M. Moodie,
David A. Stephens
Abstract:
In this work, we examine recently developed methods for Bayesian inference of optimal dynamic treatment regimes (DTRs). DTRs are a set of treatment decision rules aimed at tailoring patient care to patient-specific characteristics, thereby falling within the realm of precision medicine. In this field, researchers seek to tailor therapy with the intention of improving health outcomes; therefore, th…
▽ More
In this work, we examine recently developed methods for Bayesian inference of optimal dynamic treatment regimes (DTRs). DTRs are a set of treatment decision rules aimed at tailoring patient care to patient-specific characteristics, thereby falling within the realm of precision medicine. In this field, researchers seek to tailor therapy with the intention of improving health outcomes; therefore, they are most interested in identifying optimal DTRs. Recent work has developed Bayesian methods for identifying optimal DTRs in a family indexed by $ψ$ via Bayesian dynamic marginal structural models (MSMs) (Rodriguez Duque et al., 2022a); we review the proposed estimation procedure and illustrate its use via the new BayesDTR R package. Although methods in (Rodriguez Duque et al., 2022a) can estimate optimal DTRs well, they may lead to biased estimators when the model for the expected outcome if everyone in a population were to follow a given treatment strategy, known as a value function, is misspecified or when a grid search for the optimum is employed. We describe recent work that uses a Gaussian process ($GP$) prior on the value function as a means to robustly identify optimal DTRs (Rodriguez Duque et al., 2022b). We demonstrate how a $GP$ approach may be implemented with the BayesDTR package and contrast it with other value-search approaches to identifying optimal DTRs. We use data from an HIV therapeutic trial in order to illustrate a standard analysis with these methods, using both the original observed trial data and an additional simulated component to showcase a longitudinal (two-stage DTR) analysis.
△ Less
Submitted 27 March, 2023;
originally announced March 2023.
-
A Bayesian Non-Stationary Heteroskedastic Time Series Model for Multivariate Critical Care Data
Authors:
Zayd Omar,
David A. Stephens,
Alexandra M. Schmidt,
David L. Buckeridge
Abstract:
We propose a multivariate GARCH model for non-stationary health time series by modifying the variance of the observations of the standard state space model. The proposed model provides an intuitive way of dealing with heteroskedastic data using the conditional nature of state space models. We follow the Bayesian paradigm to perform the inference procedure. In particular, we use Markov chain Monte…
▽ More
We propose a multivariate GARCH model for non-stationary health time series by modifying the variance of the observations of the standard state space model. The proposed model provides an intuitive way of dealing with heteroskedastic data using the conditional nature of state space models. We follow the Bayesian paradigm to perform the inference procedure. In particular, we use Markov chain Monte Carlo methods to obtain samples from the resultant posterior distribution. Due to the natural temporal correlation structure induced on model parameters, we use the forward filtering backward sampling algorithm to efficiently obtain samples from the posterior distribution. The proposed model also handles missing data in a fully Bayesian fashion. We validate our model on synthetic data, and then use it to analyze a data set obtained from an intensive care unit in a Montreal hospital. We further show that our proposed models offer better performance, in terms of WAIC, than standard state space models. The proposed model provides a new way to model multivariate heteroskedastic non-stationary time series data and the simplicity in applying the WAIC allows us to compare competing models.
△ Less
Submitted 15 March, 2023;
originally announced March 2023.
-
Hardness of braided quantum circuit optimization in the surface code
Authors:
Kunihiro Wasa,
Shin Nishio,
Koki Suetsugu,
Michael Hanks,
Ashley Stephens,
Yu Yokoi,
Kae Nemoto
Abstract:
Large-scale quantum information processing requires the use of quantum error correcting codes to mitigate the effects of noise in quantum devices. Topological error-correcting codes, such as surface codes, are promising candidates as they can be implemented using only local interactions in a two-dimensional array of physical qubits. Procedures such as defect braiding and lattice surgery can then b…
▽ More
Large-scale quantum information processing requires the use of quantum error correcting codes to mitigate the effects of noise in quantum devices. Topological error-correcting codes, such as surface codes, are promising candidates as they can be implemented using only local interactions in a two-dimensional array of physical qubits. Procedures such as defect braiding and lattice surgery can then be used to realize a fault-tolerant universal set of gates on the logical space of such topological codes. However, error correction also introduces a significant overhead in computation time, the number of physical qubits, and the number of physical gates. While optimizing fault-tolerant circuits to minimize this overhead is critical, the computational complexity of such optimization problems remains unknown. This ambiguity leaves room for doubt surrounding the most effective methods for compiling fault-tolerant circuits for a large-scale quantum computer. In this paper, we show that the optimization of a special subset of braided quantum circuits is NP-hard by a polynomial-time reduction of the optimization problem into a specific problem called Planar Rectilinear 3SAT.
△ Less
Submitted 1 February, 2023;
originally announced February 2023.
-
A time-dependent Poisson-Gamma model for recruitment forecasting in multicenter studies
Authors:
Armando Turchetta,
Nicolas Savy,
David A. Stephens,
Erica E. M. Moodie,
Marina B. Klein
Abstract:
Forecasting recruitments is a key component of the monitoring phase of multicenter studies. One of the most popular techniques in this field is the Poisson-Gamma recruitment model, a Bayesian technique built on a doubly stochastic Poisson process. This approach is based on the modeling of enrollments as a Poisson process where the recruitment rates are assumed to be constant over time and to follo…
▽ More
Forecasting recruitments is a key component of the monitoring phase of multicenter studies. One of the most popular techniques in this field is the Poisson-Gamma recruitment model, a Bayesian technique built on a doubly stochastic Poisson process. This approach is based on the modeling of enrollments as a Poisson process where the recruitment rates are assumed to be constant over time and to follow a common Gamma prior distribution. However, the constant-rate assumption is a restrictive limitation that is rarely appropriate for applications in real studies. In this paper, we illustrate a flexible generalization of this methodology which allows the enrollment rates to vary over time by modeling them through B-splines. We show the suitability of this approach for a wide range of recruitment behaviors in a simulation study and by estimating the recruitment progression of the Canadian Co-infection Cohort (CCC).
△ Less
Submitted 9 January, 2023;
originally announced January 2023.
-
Detecting correlated errors in twin-field quantum key distribution
Authors:
B. Panchumarthi,
A. Stephens,
M. Beck
Abstract:
We experimentally demonstrate that we can detect correlated errors in a twin-field quantum key distribution (TFQKD) system by using a technique that is related to self-consistent tomography. We implement a TFQKD system based on a fiber-Sagnac loop, in which Alice and Bob encode information in the phase of weak coherent states that propagate in opposite directions around the loop. These states inte…
▽ More
We experimentally demonstrate that we can detect correlated errors in a twin-field quantum key distribution (TFQKD) system by using a technique that is related to self-consistent tomography. We implement a TFQKD system based on a fiber-Sagnac loop, in which Alice and Bob encode information in the phase of weak coherent states that propagate in opposite directions around the loop. These states interfere as they exit the loop and are detected by a third party, Charlie, who reports the results of their measurements to Alice and Bob. We find that it is possible for Alice and Bob to detect correlated state-preparation and measurement errors while trusting only their own individual states, and without trusting Charlie's measurements.
△ Less
Submitted 10 January, 2023; v1 submitted 13 December, 2022;
originally announced December 2022.
-
Totally disconnected semigroup compactifications of topological groups
Authors:
Alexander Stephens,
Ross Stokke
Abstract:
We introduce the notion of an introverted Boolean algebra $\cal B$ of closed-and-open subsets of a topological group $G$, show that the associated Stone space $(ν_{\cal B} G, ν_{\cal B})$ is a totally disconnected semigroup compactification of $G$, and show that every totally disconnected semigroup compactification of $G$ takes this form. We identify and study the universal totally disconnected se…
▽ More
We introduce the notion of an introverted Boolean algebra $\cal B$ of closed-and-open subsets of a topological group $G$, show that the associated Stone space $(ν_{\cal B} G, ν_{\cal B})$ is a totally disconnected semigroup compactification of $G$, and show that every totally disconnected semigroup compactification of $G$ takes this form. We identify and study the universal totally disconnected semigroup compactification, the universal totally disconnected semitopological semigroup compactification and the universal totally disconnected group compactification of $G$. Our main results are obtained independently of Gelfand theory and well-known properties of the (typically non-totally disconnected) universal compactifications $G^{LUC}$, $G^{WAP}$ and $G^{AP}$, though we do employ Gelfand theory to clarify the relationship between these familiar universal compactifications and their totally disconnected counterparts.
△ Less
Submitted 12 July, 2022;
originally announced July 2022.
-
Targeting functional parameters with semiparametric Bayesian inference
Authors:
Vivian Y. Meng,
David A. Stephens
Abstract:
Typical Bayesian inference requires parameter identification via likelihood parameterization, which has invited criticism for being less flexible than the Frequentist framework and subject to misspecification. Though misspecification may be avoided by functional parameter inference under a nonparametric model space, there does not exist a flexible Bayesian semiparametric model that would allow ful…
▽ More
Typical Bayesian inference requires parameter identification via likelihood parameterization, which has invited criticism for being less flexible than the Frequentist framework and subject to misspecification. Though misspecification may be avoided by functional parameter inference under a nonparametric model space, there does not exist a flexible Bayesian semiparametric model that would allow full control over the marginal prior over any general functional parameter. We present the technique of $θ$-augmentation which helps us manipulate nonparametric models into semiparametric ones that directly target any functional parameter. The method allows Bayesian probabilistic statements to be drawn for any estimator that is defined as a functional of the empirical distribution without requiring a likelihood function, thus providing a path to Bayesian analysis in problems like causal inference and censoring where there do not exist well-accepted likelihood functions.
△ Less
Submitted 25 November, 2022; v1 submitted 20 April, 2022;
originally announced April 2022.
-
Causal inference: critical developments, past and future
Authors:
Erica EM Moodie,
David A Stephens
Abstract:
Causality is a subject of philosophical debate and a central scientific issue with a long history. In the statistical domain, the study of cause and effect based on the notion of `fairness' in comparisons dates back several hundred years, and yet statistical concepts and developments that form the area of causal inference are only decades old. In this paper, we review core tenets and methods of ca…
▽ More
Causality is a subject of philosophical debate and a central scientific issue with a long history. In the statistical domain, the study of cause and effect based on the notion of `fairness' in comparisons dates back several hundred years, and yet statistical concepts and developments that form the area of causal inference are only decades old. In this paper, we review core tenets and methods of causal inference and key developments in the history of the field. We highlight connections with traditional `associational' statistical methods, including estimating equations and semiparametric theory, and point to current topics of active research in this crucial area of our field.
△ Less
Submitted 5 April, 2022;
originally announced April 2022.
-
Bayesian Analysis of Sigmoidal Gaussian Cox Processes via Data Augmentation
Authors:
Renaud Alie,
David A. Stephens,
Alexandra M. Schmidt
Abstract:
Many models for point process data are defined through a thinning procedure where locations of a base process (often Poisson) are either kept (observed) or discarded (thinned). In this paper, we go back to the fundamentals of the distribution theory for point processes to establish a link between the base thinning mechanism and the joint density of thinned and observed locations in any of such mod…
▽ More
Many models for point process data are defined through a thinning procedure where locations of a base process (often Poisson) are either kept (observed) or discarded (thinned). In this paper, we go back to the fundamentals of the distribution theory for point processes to establish a link between the base thinning mechanism and the joint density of thinned and observed locations in any of such models. In practice, the marginal model of observed points is often intractable, but thinned locations can be instantiated from their conditional distribution and typical data augmentation schemes can be employed to circumvent this problem. Such approaches have been employed in the recent literature, but some inconsistencies have been introduced across the different publications. We concentrate on an example: the so-called sigmoidal Gaussian Cox process. We apply our approach to resolve contradicting viewpoints in the data augmentation step of the inference procedures therein. We also provide a multitype extension to this process and conduct Bayesian inference on data consisting of positions of two different species of trees in Lansing Woods, Michigan. The emphasis is put on intertype dependence modeling with Bayesian uncertainty quantification.
△ Less
Submitted 10 December, 2024; v1 submitted 13 March, 2022;
originally announced March 2022.
-
End-to-end science operations in the era of extremely large telescopes
Authors:
Olivier R. Hainaut,
Marie Lemoine-Busserolle,
Christophe Dumas,
Robert W. Goodrich,
Bryan W. Miller,
Michael F. Sterzik,
Thomas Bierwirth,
Sidney Wolff,
Andrew W. Stephens,
Gelys Trancho,
Warren Skidmore,
Kim Gillies
Abstract:
Observatory end-to-end science operations is the overall process starting with a scientific question, represented by a proposal requesting observing time, and ending with the analysis of observation data addressing that question, and including all the intermediate steps needed to plan, schedule, obtain, and process these observations. Increasingly complex observing facilities demand a highly effic…
▽ More
Observatory end-to-end science operations is the overall process starting with a scientific question, represented by a proposal requesting observing time, and ending with the analysis of observation data addressing that question, and including all the intermediate steps needed to plan, schedule, obtain, and process these observations. Increasingly complex observing facilities demand a highly efficient science operations approach and at the same time be user friendly to the astronomical user community and enable the highest possible scientific return. Therefore, this process is supported by a collection of tools. In this paper, we describe the overall end-to-end process and its implementation for the three upcoming extremely large telescopes (ELTs), ESO's ELT, the Thirty Meter Telescope (TMT), and the Giant Magellan Telescope (GMT).
△ Less
Submitted 10 March, 2022;
originally announced March 2022.
-
Causal inference under mis-specification: adjustment based on the propensity score
Authors:
David A. Stephens,
Widemberg S. Nobre,
Erica E. M. Moodie,
Alexandra M. Schmidt
Abstract:
We study Bayesian approaches to causal inference via propensity score regression. Much of the Bayesian literature on propensity score methods have relied on approaches that cannot be viewed as fully Bayesian in the context of conventional `likelihood times prior' posterior inference; in addition, most methods rely on parametric and distributional assumptions, and presumed correct specification. We…
▽ More
We study Bayesian approaches to causal inference via propensity score regression. Much of the Bayesian literature on propensity score methods have relied on approaches that cannot be viewed as fully Bayesian in the context of conventional `likelihood times prior' posterior inference; in addition, most methods rely on parametric and distributional assumptions, and presumed correct specification. We emphasize that causal inference is typically carried out in settings of mis-specification, and develop strategies for fully Bayesian inference that reflect this. We focus on methods based on decision-theoretic arguments, and show how inference based on loss-minimization can give valid and fully Bayesian inference. We propose a computational approach to inference based on the Bayesian bootstrap which has good Bayesian and frequentist properties.
△ Less
Submitted 30 January, 2022;
originally announced January 2022.
-
The mass distribution in the Galactic Centre from interferometric astrometry of multiple stellar orbits
Authors:
GRAVITY Collaboration,
R. Abuter,
N. Aimar,
A. Amorim,
J. Ball,
M. Bauböck,
J. P. Berger,
H. Bonnet,
G. Bourdarot,
W. Brandner,
V. Cardoso,
Y. Clénet,
Y. Dallilar,
R. Davies,
P. T. de Zeeuw,
J. Dexter,
A. Drescher,
F. Eisenhauer,
N. M. Förster Schreiber,
A. Foschi,
P. Garcia,
F. Gao,
E. Gendron,
R. Genzel,
S. Gillessen
, et al. (40 additional authors not shown)
Abstract:
The stars orbiting the compact radio source Sgr A* in the Galactic Centre are precision probes of the gravitational field around the closest massive black hole. In addition to adaptive optics assisted astrometry (with NACO / VLT) and spectroscopy (with SINFONI / VLT, NIRC2 / Keck and GNIRS / Gemini) over three decades, since 2016/2017 we have obtained 30-100 mu-as astrometry with the four-telescop…
▽ More
The stars orbiting the compact radio source Sgr A* in the Galactic Centre are precision probes of the gravitational field around the closest massive black hole. In addition to adaptive optics assisted astrometry (with NACO / VLT) and spectroscopy (with SINFONI / VLT, NIRC2 / Keck and GNIRS / Gemini) over three decades, since 2016/2017 we have obtained 30-100 mu-as astrometry with the four-telescope interferometric beam combiner GRAVITY / VLTI reaching a sensitivity of mK = 20 when combining data from one night. We present the simultaneous detection of several stars within the diffraction limit of a single telescope, illustrating the power of interferometry. The new data for the stars S2, S29, S38 and S55 yield significant accelerations between March and July 2021, as these stars pass the pericenters of their orbits between 2018 and 2023. This allows for a high-precision determination of the gravitational potential around Sgr A*. Our data are in excellent agreement with general relativity orbits around a single central point mass, M = 4.30 x 10^6 M_sun with a precision of about +-0.25%. We improve the significance of our detection of the Schwarzschild precession in the S2 orbit to 7 sigma. Assuming plausible density profiles, an extended mass component inside S2's apocentre (= 0.23" or 2.4 x 10^4 R_S) must be 3000 M_sun (1 sigma), or 0.1% of M. Adding the enclosed mass determinations from 13 stars orbiting Sgr A* at larger radii, the innermost radius at which the excess mass beyond Sgr A* tentatively is seen is r = 2.5" >= 10x the apocentre of S2. This is in full harmony with the stellar mass distribution (including stellar-mass black holes) obtained from the spatially resolved luminosity function.
△ Less
Submitted 14 December, 2021;
originally announced December 2021.
-
Gemini/GMOS Transmission Spectroscopy of the Grazing Planet Candidate WD 1856+534 b
Authors:
Siyi Xu,
Hannah Diamond-Lowe,
Ryan J. MacDonald,
Andrew Vanderburg,
Simon Blouin,
P. Dufour,
Peter Gao,
Laura Kreidberg,
S. K. Leggett,
Andrew W. Mann,
Caroline V. Morley,
Andrew W. Stephens,
Christopher E. O'Connor,
Pa Chia Thao,
Nikole K. Lewis
Abstract:
WD 1856+534 b is a Jupiter-sized, cool giant planet candidate transiting the white dwarf WD 1856+534. Here, we report an optical transmission spectrum of WD 1856+534 b obtained from ten transits using the Gemini Multi-Object Spectrograph. This system is challenging to observe due to the faintness of the host star and the short transit duration. Nevertheless, our phase-folded white light curve reac…
▽ More
WD 1856+534 b is a Jupiter-sized, cool giant planet candidate transiting the white dwarf WD 1856+534. Here, we report an optical transmission spectrum of WD 1856+534 b obtained from ten transits using the Gemini Multi-Object Spectrograph. This system is challenging to observe due to the faintness of the host star and the short transit duration. Nevertheless, our phase-folded white light curve reached a precision of 0.12 %. WD 1856+534 b provides a unique transit configuration compared to other known exoplanets: the planet is $8\times$ larger than its star and occults over half of the stellar disc during mid-transit. Consequently, many standard modeling assumptions do not hold. We introduce the concept of a `limb darkening corrected, time-averaged transmission spectrum' and propose that this is more suitable than $(R_{\mathrm{p}, λ} / R_{\mathrm{s}})^2$ for comparisons to atmospheric models for planets with grazing transits. We also present a modified radiative transfer prescription. Though the transmission spectrum shows no prominent absorption features, it is sufficiently precise to constrain the mass of WD 1856+534 b to be > 0.84 M$_\mathrm{J}$ (to $2 \, σ$ confidence), assuming a clear atmosphere and a Jovian composition. High-altitude cloud decks can allow lower masses. WD 1856+534 b could have formed either as a result of common envelope evolution or migration under the Kozai-Lidov mechanism. Further studies of WD 1856+534 b, alongside new dedicated searches for substellar objects around white dwarfs, will shed further light on the mysteries of post-main sequence planetary systems.
△ Less
Submitted 26 October, 2021;
originally announced October 2021.
-
Bayesian Sample Size Calculations for SMART Studies
Authors:
Armando Turchetta,
Erica E. M. Moodie,
David A. Stephens,
Sylvie D. Lambert
Abstract:
In the management of most chronic conditions characterized by the lack of universally effective treatments, adaptive treatment strategies (ATSs) have been growing in popularity as they offer a more individualized approach, and sequential multiple assignment randomized trials (SMARTs) have gained attention as the most suitable clinical trial design to formalize the study of these strategies. While…
▽ More
In the management of most chronic conditions characterized by the lack of universally effective treatments, adaptive treatment strategies (ATSs) have been growing in popularity as they offer a more individualized approach, and sequential multiple assignment randomized trials (SMARTs) have gained attention as the most suitable clinical trial design to formalize the study of these strategies. While the number of SMARTs has increased in recent years, their design has remained limited to the frequentist setting, which may not fully or appropriately account for uncertainty in design parameters and hence not yield appropriate sample size recommendations. Specifically, standard frequentist formulae rely on several assumptions that can be easily misspecified. The Bayesian framework offers a straightforward path to alleviate some of these concerns. In this paper, we provide calculations in a Bayesian setting to allow more realistic and robust estimates that account for uncertainty in inputs through the `two priors' approach. Additionally, compared to the standard formulae, this methodology allows us to rely on fewer assumptions, integrate pre-trial knowledge, and switch the focus from the standardized effect size to the minimal detectable difference. The proposed methodology is evaluated in a thorough simulation study and is implemented to estimate the sample size for a full-scale SMART of an Internet-Based Adaptive Stress Management intervention based on a pilot SMART conducted on cardiovascular disease patients from two Canadian provinces.
△ Less
Submitted 2 August, 2021;
originally announced August 2021.
-
Norm-multiplicative homomorphisms of Beurling algebras
Authors:
Matthew E. Kroeker,
Alexander Stephens,
Ross Stokke,
Randy Yee
Abstract:
We introduce and study "norm-multiplicative" homomorphisms $\varphi: {\cal L}^1(F) \rightarrow {\cal M}_r(G)$ between group and measure algebras, and $\varphi: {\cal L}^1(ω_F) \rightarrow {\cal M}(ω_G)$ between Beurling group and measure algebras, where $F$ and $G$ are locally compact groups with continuous weights $ω_F$ and $ω_G$. Through a unified approach we recover, and sometimes strengthen, m…
▽ More
We introduce and study "norm-multiplicative" homomorphisms $\varphi: {\cal L}^1(F) \rightarrow {\cal M}_r(G)$ between group and measure algebras, and $\varphi: {\cal L}^1(ω_F) \rightarrow {\cal M}(ω_G)$ between Beurling group and measure algebras, where $F$ and $G$ are locally compact groups with continuous weights $ω_F$ and $ω_G$. Through a unified approach we recover, and sometimes strengthen, many of the main known results concerning homomorphisms and isomorphisms between these (Beurling) group and measure algebras. We provide a first description of all positive homomorphisms $\varphi: {\cal L}^1(F) \rightarrow {\cal M}_r(G)$. We state versions of our results that describe a variety of (possibly unbounded) homomorphisms $\varphi: \mathbb{C} F \rightarrow \mathbb{C} G$ for (discrete) groups $F$ and $G$.
△ Less
Submitted 30 July, 2021;
originally announced July 2021.
-
Self-consistent state and measurement tomography with fewer measurements
Authors:
A. Stephens,
J. M. Cutshall,
T. McPhee,
M. Beck
Abstract:
We describe a technique for self consistently characterizing both the quantum state of a single qubit system, and the positive-operator-valued measure (POVM) that describes measurements on the system. The method works with only ten measurements. We assume that a series of unitary transformations performed on the quantum state are fully known, while making minimal assumptions about both the density…
▽ More
We describe a technique for self consistently characterizing both the quantum state of a single qubit system, and the positive-operator-valued measure (POVM) that describes measurements on the system. The method works with only ten measurements. We assume that a series of unitary transformations performed on the quantum state are fully known, while making minimal assumptions about both the density operator of the state and the POVM. The technique returns maximum-likely estimates of both the density operator and the POVM. To experimentally demonstrate the method, we perform reconstructions of over 300 state-measurement pairs and compare them to their expected density operators and POVMs. We find that 95% of the reconstructed POVMs have fidelities of 0.98 or greater, and 92% of the density operators have fidelities that are 0.98 or greater.
△ Less
Submitted 30 June, 2021;
originally announced July 2021.
-
Bayesian inference for continuous-time hidden Markov models with an unknown number of states
Authors:
Yu Luo,
David A. Stephens
Abstract:
We consider the modeling of data generated by a latent continuous-time Markov jump process with a state space of finite but unknown dimensions. Typically in such models, the number of states has to be pre-specified, and Bayesian inference for a fixed number of states has not been studied until recently. In addition, although approaches to address the problem for discrete-time models have been deve…
▽ More
We consider the modeling of data generated by a latent continuous-time Markov jump process with a state space of finite but unknown dimensions. Typically in such models, the number of states has to be pre-specified, and Bayesian inference for a fixed number of states has not been studied until recently. In addition, although approaches to address the problem for discrete-time models have been developed, no method has been successfully implemented for the continuous-time case. We focus on reversible jump Markov chain Monte Carlo which allows the trans-dimensional move among different numbers of states in order to perform Bayesian inference for the unknown number of states. Specifically, we propose an efficient split-combine move which can facilitate the exploration of the parameter space, and demonstrate that it can be implemented effectively at scale. Subsequently, we extend this algorithm to the context of model-based clustering, allowing numbers of states and clusters both determined during the analysis. The model formulation, inference methodology, and associated algorithm are illustrated by simulation studies. Finally, We apply this method to real data from a Canadian healthcare system in Quebec.
△ Less
Submitted 20 June, 2021;
originally announced June 2021.