Statistics
See recent articles
Showing new listings for Friday, 21 August 2026
- [1] arXiv:2608.19224 [pdf, html, other]
-
Title: Causal Inference under Interference with Learned Exposure MappingsComments: 20 pages, 6 figuresSubjects: Methodology (stat.ME); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG); Applications (stat.AP)
Exposure mappings are often assumed to be known in causal spillover analyses. In environmental settings, however, they are typically induced by transport processes that are not directly observed and must instead be learned from pollution data. We study how uncertainty in learned transport processes propagates into exposure mappings and downstream spillover inference under interference. We compare mechanistic transport models with modern operator-learning approaches, including PDE, PINO, FNO, and GeoPT, using both simulation studies and an empirical analysis of California PM$_{2.5}$ data. In simulations, all four transport models achieved nearly identical pollution prediction accuracy, yet estimated spillover effects ranged from 1.78 to 2.27. Models that more accurately recovered the induced exposure mapping also produced spillover estimates closer to the true effect. Disagreement was modest for regional interventions but substantially larger for localized point-source interventions. The California analysis showed the same pattern: competing transport models produced similar predictions of observed PM${2.5}$ concentrations while implying different spillover effects under hypothetical pollution-control interventions. Our findings suggest that predictive agreement alone is insufficient for reliable causal inference when exposure mappings are learned rather than directly observed.
- [2] arXiv:2608.19231 [pdf, html, other]
-
Title: TorchDCM: A Unified PyTorch-Native Package for Discrete Choice ModelingSubjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
Estimating large and simulation-intensive discrete choice models (DCMs) requires repeated evaluation of utilities, probabilities, derivatives, and simulated likelihoods over many observations, alternatives, and draws. Existing DCM software provides mature econometric workflows, while recent GPU-oriented tools accelerate selected models, leaving a gap between econometric coverage and scalable differentiable computation. We introduce TorchDCM, an open Python package for discrete choice modeling that compiles choice data and model specifications into a unified PyTorch-native likelihood engine for estimation, inference, prediction, and structured reporting on CPU or CUDA devices. The package covers the principal econometric functionality available across Biogeme and Apollo, including multinomial, nested, mixed, ordered, latent-variable, and panel likelihoods. It also supports ragged choice sets, constrained parameters, covariance estimation, willingness-to-pay analysis, elasticities, and extensible likelihood components. We evaluate TorchDCM against seven other estimation packages in aligned synthetic and real-data full-estimation experiments. TorchDCM completes all 45 synthetic cases, runs fastest in every comparable synthetic case, and satisfies the prespecified final-log-likelihood tolerance in every comparison with at least two comparable solutions. More precisely, it reduces median runtime by 89.1%-99.7% relative to Biogeme and Apollo across model-data settings. CUDA provides an additional 12.0-71.0x speedup over single-core TorchDCM. These results establish a scalable and reproducible foundation for econometric estimation and differentiable choice-model development. The open-source package and executed examples are available at this https URL.
- [3] arXiv:2608.19275 [pdf, html, other]
-
Title: A Self-Exciting Model of Eddy Formation at SubmesoscaleSubjects: Methodology (stat.ME); Probability (math.PR)
Ocean eddies are highly dynamic structures marked by frequent splitting and merging events. They exhibit complex spatio-temporal clustering that traditional Poisson models fail to capture. In this paper, a novel spatio-temporal Hawkes process is introduced to model self-exciting eddy fields on the basis of high-frequency ocean flow data. We formulate a triggering kernel that couples the self-excitation intensity with spatial deformation caused by the strain rate magnitude. To characterize the asymptotic behavior of the eddy field, we derive a Volterra integral equation that governs the expected eddy intensity over time and space. We develop an Expectation-Maximization (EM) algorithm that treats the unobserved parent-child relationships as latent branching structures for parameter estimation. We further extend this framework to accommodate a time-varying, non-homogeneous background intensity, modifying both the EM updates and the analytical Volterra solution accordingly. Finally, we propose a two-stage simulation framework utilizing a cluster representation algorithm. The simulated empirical paths of eddy formation are compared together with numerical solution of the Volterra equation for mean rate as validation of the Hawkes model.
- [4] arXiv:2608.19282 [pdf, other]
-
Title: Recovering Nonlinear Functions of Latent Variables: A Plausible-Value Neural Network FrameworkEunjeong Song (1), Sehee Hong (1) ((1) Department of Education, Korea University, Seoul, Republic of Korea)Comments: 51 pages, 4 figures, 24 tables. Data, materials, and code: this https URLSubjects: Methodology (stat.ME); Machine Learning (cs.LG)
When factor scores replace true latent scores in nonlinear prediction, measurement error attenuates the recoverable variance of any $k$th-order component of the regression function by $\rho^k$ -- the $k$th power of the score's coefficient of determination -- for any linear score type. This study derives the bound via Hermite polynomial expansion and proposes PV-ANN -- plausible values (posterior draws preserving latent variance) combined with artificial neural networks (learning functional form without prespecification). The bound governs recovery of the latent-scale function, not prediction of the outcome from observed indicators, for which factor scores are already sufficient; the two metrics are therefore predicted to dissociate. An 18-condition simulation supports both predictions: in the nonlinear low-reliability conditions PV-ANN closes about four fifths of the function-shape recovery gap between a factor-score learner and one given the true latent values, and the margin widens as reliability falls, while predictive accuracy is not improved, as the theory requires. A Big Five application illustrates the intended exploratory workflow and delineates boundary conditions under weak signal and measurement model misspecification.
- [5] arXiv:2608.19377 [pdf, html, other]
-
Title: Heteroscedastic Neural Surrogate Modeling for Robust and Rapid Bayesian Inference in Fusion Plasma DiagnosticsComments: 8 pages, 3 figuresSubjects: Computation (stat.CO); Machine Learning (cs.LG)
Bayesian inference via Markov Chain Monte Carlo (MCMC) provides effective parameter estimation, but its real-time application in complex physical systems is hindered by heavy computational bottlenecks and extreme sensitivity to statistical noise. We address this by proposing a neural-network-based probabilistic surrogate framework for rapid and robust MCMC inference. Using fusion plasma Thomson scattering diagnostics as a challenging, noise-dominated testbed, our approach employs a dual-head architecture to simultaneously estimate the expected physical emission spectrum and the channel-wise intrinsic measurement noise variance. By optimizing a Gaussian Negative Log-Likelihood (GNLL) objective, the learned aleatoric uncertainty dynamically buffers the sampler against pathological shot noise. Evaluations demonstrate that this surrogate framework achieves > 1500x acceleration over exact physical forward models, while simultaneously reducing inference error (RMSE) by >20% compared to standard homoscedastic neural baselines, offering a highly promising paradigm for real-time physical analysis.
- [6] arXiv:2608.19383 [pdf, html, other]
-
Title: Causal Generalization of Continuous Treatment Effects under Covariate ShiftSubjects: Methodology (stat.ME); Machine Learning (stat.ML)
Average dose-response functions are widely used to summarize causal effects of continuous treatments, but most existing methods assume that the observed sample represents the target population. We study a covariate-shift setting in which covariates, treatment, and outcome are observed in a labelled source sample, while only covariates are observed in the target sample. We develop a two-sample local polynomial regression framework based on pseudo-outcomes that use source outcomes to address confounding and target covariates to define the population of interest. We further propose a source-to-target extension of distance covariance optimal weighting (DCOW), designed to remove treatment-covariate dependence in the source sample while aligning the weighted source covariate distribution with the target population. A central theoretical contribution is a weight-level analysis of this optimization-based procedure: we show that the population criterion identifies the oracle source-to-target weights and that approximate empirical minimizers, including exact minimizers as a special case, converge uniformly to these weights under regularity conditions. We also establish consistency and asymptotic normality of the resulting estimator. Simulations show that the proposed method improves target dose-response estimation relative to DCOW, generalized-propensity-score weighting, entropy balancing, and unweighted alternatives. We illustrate the method in a county-level analysis of PM2.5 exposure and subsequent heart-disease mortality using a source-target validation design.
- [7] arXiv:2608.19423 [pdf, html, other]
-
Title: Shape-Preserving Covariate Adjustment via Empirical Likelihood in Randomized ExperimentSubjects: Methodology (stat.ME); Statistics Theory (math.ST)
Covariate adjustment improves estimation efficiency in randomized experiments, but standard calibration and augmentation methods, when applied to distribution or survival functions, do not preserve monotonicity---a fundamental property of the estimand. We propose using empirical likelihood with covariate-balancing constraints to construct a covariate-adjusted empirical measure for each treatment arm. Estimators of a broad class of distributional functionals, including cumulative distribution functions, survival functions, quantiles, and restricted mean survival times, are then derived as plug-in functionals of this measure, automatically inheriting proper shape constraints. We establish asymptotic normality with an explicit, guaranteed efficiency gain over unadjusted estimators. The asymptotic distributions are invariant to the randomization scheme, providing a unified inference procedure under simple randomization and all commonly used covariate-adaptive designs satisfying a mild balancing condition. This unified construction, adjusting the empirical measure once and deriving all estimators from it, offers a principled reconciliation of covariate adjustment with shape preservation. Simulations and an application to the SURPASS-4 trial confirm the theoretical gains.
- [8] arXiv:2608.19454 [pdf, other]
-
Title: Dummy RAPM:Representing Low-Minute Players in Regularized Adjusted Plus-MinusSubjects: Applications (stat.AP)
Regularized Adjusted Plus-Minus (RAPM) uses stint-level lineup indicators to estimate player contributions to scoring margin. When low-minute player columns are removed, their stints remain in the data, but the design matrix no longer represents the complete lineup. Dummy RAPM restores this information using five indicators for the number of excluded players on each lineup side. Across 16 NBA seasons, chronological validation selects a 10-minute-per-appearance threshold and a dummy-to-player penalty ratio of 2.2. On held-out March-April games, Dummy RAPM reduces mean season game-margin RMSE from 12.897 to 12.856 and achieves lower RMSE in 13 of 16 seasons. The average reduction is 0.042 points, or 0.30%. Although the improvement in game-level predictive accuracy is small, it is consistent: RAPM performs better when it records how many excluded players are on each side.
- [9] arXiv:2608.19501 [pdf, html, other]
-
Title: A Causal Inference Approach for Evaluating Diagnostic Tests and AI-Enabled Medical Devices: From Effect Modification to Information-Augmented Decision-MakingSubjects: Methodology (stat.ME); Machine Learning (stat.ML)
Diagnostic medical tests and devices provide useful information for evaluating the potential benefits and risks of therapeutic treatments. However, unlike treatments, their impact on health outcomes is generally indirect because measuring diagnostic information typically does not itself affect patient outcomes, which complicates evaluation of their effectiveness. In this work, we develop a causal inference approach for evaluating diagnostic tests by distinguishing explanatory and pragmatic effectiveness. Explanatory effectiveness evaluates whether a diagnostic test result provides treatment-relevant information beyond baseline covariates by explaining additional treatment-effect heterogeneity, which we characterize using a variance-based treatment-effect variable importance measure. Pragmatic effectiveness evaluates whether incorporating this information into personalized treatment decisions improves expected outcomes, a question central to decision-making for patients, clinicians, and other health care stakeholders. We formalize pragmatic effectiveness as a contrast between expected outcomes under optimal personalized treatment rules defined with and without access to the diagnostic test result. We establish identification of the proposed estimand and provide Targeted Maximum Likelihood Estimation (TMLE) and cross-validated TMLE procedures for nonparametric estimation and inference. Simulation studies and a synthetic colorectal cancer application illustrate the proposed estimands and estimation performance. More broadly, this approach provides a causal inference perspective for evaluating AI-enabled devices by distinguishing their dual roles as information-enrichment and personalized decision-optimization tools, clarifying whether AI improves outcomes by expanding information for downstream decisions, improving the decision rule used to act on that information, or through both pathways.
- [10] arXiv:2608.19585 [pdf, html, other]
-
Title: Estimating Negative Income Distributions via Data Fusion with Vine Copula-based ImputationSubjects: Methodology (stat.ME); Applications (stat.AP)
Imputation-based data fusion combines datasets by imputing missing variables in one dataset using information from the other. The validity of imputation-based data fusion relies on accurately preserving both the imputed distributions and the underlying multivariate dependence structure across data sources. However, traditional imputation-based data fusion methods often struggle to preserve the dependence structure, particularly when marginal distributions differ or when complex, nonlinear multivariate dependencies exist. Inadequate preservation of these dependencies can result in distorted joint distributions and biased inference in the fused data. To address this challenge, this study introduces a novel imputation-based data fusion framework that utilises conditional sampling from C-vine and D-vine copulas as the imputation mechanism. This approach flexibly models pairwise and higher-order dependencies while accommodating heterogeneous marginal distributions, thereby generating imputations that more accurately maintain the distribution of the multivariate data. The proposed method is validated through a simulation study using survey data and a real-data application that estimates the negative income distribution in survey data using information from administrative tax records. Results indicate that the proposed method outperforms existing approaches in preserving both the imputed distributions and the dependence structure across datasets.
- [11] arXiv:2608.19596 [pdf, html, other]
-
Title: Martingale R-learner: Estimating Time-varying Heterogeneous Treatment Effects for Time-to-event OutcomesSubjects: Methodology (stat.ME)
Biological research and clinical evidence suggest that treatment response may vary substantially along characteristics, such as comorbidities, genetic variants, environmental, or socio-economic factors. Future precision medicine requires accurate assessment of heterogeneous treatment effects (HTE) to guide optimal clinical decisions at the individual level. We introduce a functional score framework that extends the traditional estimating equations for survival data to nonparametric HTE and generalize the Neyman orthogonality accordingly, thus filling a methodological as well as theoretical gap. Under the Neyman orthogonal functional score framework, we developed the martingale R-learner based on a decomposition of the conditional martingale residuals into residuals of the risk-set propensity score and the marginal martingale, thereby reducing the impact of estimation bias in HTE from nuisance models including (1) marginal survival, and (2) risk-set propensity scores. This enables leveraging advances in machine learning and incorporates flexible estimators for the nuisance functions and attaining the standard optimal nonparametric estimation rate with the oracle property. Numerical experiments demonstrated empirical performance consistent with the theory. We applied the martingale R-learner to estimate the effect of alcohol on dementia using the Honolulu-Asia Aging Study data.
- [12] arXiv:2608.19599 [pdf, html, other]
-
Title: Efficient Poisson Subsampling for the Partially Linear Additive Cox ModelSubjects: Methodology (stat.ME); Computation (stat.CO)
To address the computational and storage challenges often encountered in large-scale survival data analysis, we propose an efficient Poisson subsampling method for the partially linear additive Cox model. This model provides a flexible yet interpretable framework by incorporating linear covariate effects, additive nonparametric components for nonlinear covariates, and a nonparametric baseline hazard function. The proposed method adopts B-spline basis functions to approximate the nonparametric components and employs the decorrelated score technique to construct a Poisson subsampling-based estimation equation, based on which we establish the asymptotic normality of the resulting estimator and derive the optimal subsampling probabilities according to the L-optimality criterion. Furthermore, we design a two-step adaptive algorithm for practical implementation. The proposed approach enables computationally efficient statistical inference for large-scale survival analysis without processing the full dataset. We validate the performance of the proposed method through extensive simulation studies and a real-world application to a lymphoma cancer dataset, demonstrating its efficiency and accuracy in large-scale settings.
- [13] arXiv:2608.19608 [pdf, html, other]
-
Title: Tau-Rho Equality and Other Dependence Measures of a Subclass of Factorizable CopulasSubjects: Statistics Theory (math.ST)
Kendall's tau and Spearman's rho, two widely used dependence measures in statistics and risk management, are often treated as interchangeable, yet can disagree sharply: Schreyer et al.~(2017) established the exact region of attainable $(\tau,\rho)$ pairs. We study the complementary question of equality, namely, identifying nontrivial families of copulas $C$ satisfying $\tau_C=\rho_C$. We prove that this equality holds for every factorizable copula of the form $C_{e,\alpha}\ast C_{\beta,e}$, where $\alpha$ and $\beta$ are piecewise linear monotonic surjections (PLMS). For these PLMS-generated copulas, we study further dependence measures, including Chatterjee's rank correlation coefficient and tail dependence coefficients, revealing some useful algebraic formulas and unexpected phenomena. In particular, Chatterjee's coefficient can exhibit extreme asymmetry.
- [14] arXiv:2608.19631 [pdf, html, other]
-
Title: Variational Goal-Oriented Optimal Experimental Design for Mixed-Distribution Quantities of Interest: Application to Ship Roll SafetySubjects: Methodology (stat.ME); Applications (stat.AP); Computation (stat.CO)
Goal-oriented optimal experimental design (GO-OED) selects experiments according to the expected information gain (EIG) about a quantity of interest (QoI) rather than the full parameter vector. This work develops a variational GO-OED formulation for mixed discrete-continuous QoI laws arising in probabilistic mechanics when thresholding or event-based transformations map a positive-probability set of uncertain inputs to a common value while other inputs produce continuously varying responses. The motivating application is ship roll safety assessment in random waves, where the QoI is the temporal exceedance probability above a prescribed roll-angle threshold. This quantity is zero when no exceedance occurs and varies continuously over positive values otherwise. A purely continuous variational approximation does not dominate a posterior QoI law containing an atom, yielding an infinite Kullback-Leibler divergence and a trivial Barber-Agakov lower bound of $-\infty$. Scoring atom samples using continuous density values instead changes the objective and does not produce a valid lower-bound estimator. We introduce a mixed variational approximation that models the conditional atom probability and continuous component separately, with a normalizing flow used for the latter. An analytical example recovers the correct EIG landscape, while the ship roll application provides stable EIG lower-bound estimates and identifies informative wave conditions for temporal-exceedance-probability inference.
- [15] arXiv:2608.19634 [pdf, html, other]
-
Title: Curvature-Calibrated Quasi-Bayesian Updating for Moment-Restricted ModelsSubjects: Methodology (stat.ME); Econometrics (econ.EM); Statistics Theory (math.ST)
Moment restrictions provide a flexible basis for quasi-Bayesian inference when a full likelihood is unavailable, but the weighting matrix in a quadratic moment criterion determines both the relative importance of the moments and the information scale of posterior updating. We propose curvature-calibrated quasi-Bayesian updating, which uses the inverse of the covariance (or long-run covariance) of the moment conditions evaluated at a self-consistent quasi-posterior center. The resulting fixed-point procedure alternates between covariance estimation and simulation from a fixed-weight quasi-posterior, thereby avoiding parameter-dependent weighting during each simulation run. Under a Bernstein-von Mises condition for the fixed-weight quasi-posterior at the efficient population weight, we show that the calibration map is locally contractive, that its fixed point is consistent at the standard parametric rate, and that the Gaussian approximation continues to hold under the calibrated data-dependent weight, with covariance given by the inverse Godambe information matrix. Under a uniform fourth-moment condition, the scaled quasi-posterior covariance converges to the same matrix, so quasi-posterior and repeated-sampling uncertainty agree to first order. Simulations show improved covariance calibration and interval coverage after a few updates. An application to longitudinal binary-response data illustrates the method with within-subject dependence and overidentified residual moments.
- [16] arXiv:2608.19664 [pdf, html, other]
-
Title: Multivariate Spatio-Temporal Regression with Penalized Model Selection and an Empirical ApplicationComments: 25 pages, 4 figures, 3 tablesSubjects: Methodology (stat.ME); Applications (stat.AP)
This paper develops the statistical foundations of a multivariate general nesting spatio-temporal (MGNST) regression framework for analyzing spatial, temporal, and cross-equation dependence among responses. Four parameter matrices represent spatial lag dependence, spatial error dependence,temporal autoregression, and contemporaneous error covariance. Their off-diagonal elements allow dependence to propagate within and across responses. Matrix restrictions yield eleven specifications encompassing multivariate spatial autoregressive models, multivariate spatial error models, vector autoregressive models with exogenous variables, and independent spatio-temporal regressions as special cases.
We establish identifiability conditions using instrumental-variable rank conditions and introduce penalized likelihood estimation and information criteria based on effective degrees of freedom. Monte Carlo experiments examine three data-generating models, three sample sizes, and three levels of spatial dependence. Correct-selection rates under penalized AIC generally increase with sample size and the strength of spatial dependence, while parameter recovery improves as the sample size increases.
For socioeconomic data from 198 municipalities in Japan's Kansai region, penalized AIC selects the full MGNST model, whereas penalized BIC selects a response-wise independent spatial error model. Despite selecting models of different complexity, both criteria support temporal persistence and spatial error dependence. The pAIC-selected MGNST model reduces strong spatial autocorrelation in the responses to negligible residual levels, demonstrating its usefulness for identifying and comparing multivariate spatio-temporal dependence structures. - [17] arXiv:2608.19716 [pdf, html, other]
-
Title: A Bayesian Time-Varying SEIARD Model for State-Level COVID-19 Transmission and Mortality in the United StatesSubjects: Applications (stat.AP); Probability (math.PR); Computation (stat.CO); Methodology (stat.ME)
We conduct a retrospective analysis of COVID-19 transmission dynamics across U.S. states using a modified population-based Susceptible-Exposed-Infectious-Asymptomatic-Recovered-Deceased (SEIARD) compartmental model. The proposed framework introduces time-varying transmission, reporting, and mortality rates to capture temporal variations in public behavior and policy interventions during the pandemic. In particular, the transmission rate is modeled as a function of population mobility (derived from Google Mobility Reports), with a residual time-decay term capturing the net effect of unobserved factors such as behavioral adaptation and control measures, while reporting is linked to nationwide testing strategies. We employ a Bayesian approach to integrate multiple data sources and quantify uncertainties in model parameters. The model explicitly distinguishes between symptomatic and asymptomatic infectious individuals and links the latent epidemic states to observable quantities, including reported cases and deaths, through a dynamic reporting function. This retrospective modeling framework provides insights into state-level epidemic trajectories and supports data-driven decision-making for optimal allocation of healthcare resources and evaluation of public health interventions during future pandemics. We further apply a clustering analysis to the posterior parameter estimates to identify groups of U.S. states exhibiting similar epidemiological characteristics, revealing substantial regional heterogeneity in transmission intensity, reproduction dynamics, and mortality burden.
- [18] arXiv:2608.19718 [pdf, html, other]
-
Title: Information-Computation Inversion in Pseudo-Marginal MCMCComments: 79 pages, 6 main figures, 1 main table; supplementary material includedSubjects: Methodology (stat.ME)
Pseudo-marginal MCMC is exact under nonnegative unbiased likelihood estimation, but observation design can change both posterior information and the stochastic law of the likelihood estimator. We study fixed-horizon recovery of a declared posterior functional under this joint change. For a broad fresh-estimator pseudo-marginal class, we derive an inverse-weight acceptance bound that yields a lower bound on functional mean-squared error from high retained-weight mass in discrepant states. A two-region corollary establishes an information-computation inversion: stronger exact posterior separation can coexist with worse finite-horizon recovery when the estimator law changes. The obstruction motivates an exact route-conditioned transition that allocates coupled auxiliary inheritance and independent refreshment according to a scientific pair route. In a controlled transcription experiment, a finer observation strengthens likelihood separation while degrading particle-filter reliability and finite-chain behavior. In a fresh 96-state gene-network comparison, route conditioning reduces measured particle-filter wall time by 42.4% in the frozen execution (95% retained-state bootstrap interval 39.6%-45.2%). Recovery uncertainty spans the predeclared noninferiority margin. A prospective Lotka-Volterra experiment supports portability of the retained-state route-restoration mechanism. The results distinguish statistical information, retained-state functional risk, and auxiliary-computation allocation in pseudo-marginal inference.
- [19] arXiv:2608.19722 [pdf, other]
-
Title: Copula-Based Reconstruction and Clustering of Coccidioides Minimum Inhibitory Concentration ProfilesSubjects: Applications (stat.AP); Computation (stat.CO); Other Statistics (stat.OT)
Coccidioidomycosis is a fungal lung infection endemic to parts of the Pacific Northwest and southwestern United States, Mexico, Central America, and South America. A 2017 study by Thompson et al. reported that many analyzed Coccidioides isolates had elevated minimum inhibitory concentration values for fluconazole but comparatively low minimum inhibitory concentration values for other triazole drugs. We constructed 172 synthetic joint minimum inhibitory concentration profiles from the published drug-specific marginal frequency tables. Values sampled from each marginal distribution were paired across drugs using a Gaussian copula, and potential cross-drug patterns were examined using k-means clustering, hierarchical clustering, and Gaussian mixture models. On the selected reconstruction, a forced four-cluster partition consistently identified an upper caspofungin minimum-inhibitory-concentration tail. Across the evaluated dependence scenarios, the gap statistic favored a single cluster in 96% to 100% of nested reconstructions, providing no evidence that the reconstructed multivariate data supported a broader multicluster structure. A separate simulation study examined recovery of an upper-severity-score category defined from prespecified quantiles of a composite minimum-inhibitory-concentration score. The study included 4,000 replications across 20 prevalence and measurement-noise conditions. The multiclass adjusted Rand index ranged from moderate to high depending on the method and prevalence, whereas binary partition agreement for the upper-severity-score category approached zero at 1% prevalence even though recall remained near 1.0. Clustering conclusions therefore depended on the assumed cross-drug dependence structure and on the realized number of observations in the target category.
- [20] arXiv:2608.19749 [pdf, html, other]
-
Title: Causal Survival Forests with Negative ControlsComments: 41 pages, 18 figures, 18 tablesSubjects: Methodology (stat.ME)
We study heterogeneous treatment-effect (HTE) estimation in observational survival studies commonly associated with both censored outcomes and unmeasured confounding. We integrate causal survival forests (CSF) with negative controls (NC) from proximal causal inference and introduce Negative Control Causal Survival Forests (NC-CSF), a flexible nonparametric HTE learner for survival analysis. Our approach uses a loss that incorporates proxy variables and Neyman orthogonalization to train the random forest, thereby mitigating bias from unobserved confounding and gaining robustness to nuisance estimation. Through extensive simulations spanning varying levels of confounding, proxy relevance, and censoring mechanisms, we demonstrate that NC-CSF substantially reduces bias and estimation error relative to existing baselines. We further demonstrate the practical utility of our method on various clinical datasets, where it confirms several existing findings and also reveals new interpretable patterns of treatment-effect heterogeneity. To facilitate practical use, we provide an end-to-end Python implementation of NC-CSF that carefully handles implementation details such as nuisance estimation and clipping.
- [21] arXiv:2608.19767 [pdf, html, other]
-
Title: skchange: Fast and Flexible Algorithms for Changepoint DetectionComments: 6 pages, 3 figuresSubjects: Computation (stat.CO); Machine Learning (cs.LG)
Skchange is an open-source Python library for detecting structural changes in time series. It implements modern change detection algorithms within a unified and extensible framework. The algorithms are modular and composable, and they include changepoint search methods based on both cost minimisation and statistical tests. Key features include the detection of anomalous segments in addition to changepoints; theoretically well-founded fast and approximate search methods; theoretically well-founded algorithms for high-dimensional data, covering settings where either few or many features change simultaneously; utilities for automatic and data-driven penalty calibration, which balances false alarms against missed detections; and a large collection of built-in costs and statistical tests. The design follows established scikit-learn conventions to streamline both user and contributor experience, and Numba is used extensively to achieve high computational performance. Source code and documentation are available at this https URL.
- [22] arXiv:2608.19771 [pdf, html, other]
-
Title: Testing the Validity of Instrumental Variable Sets in Causal Additive Models with Non-Constant EffectsSubjects: Methodology (stat.ME)
Instrumental variable (IV) methods are powerful for causal effect estimation with unmeasured confounding, but in practice researchers often face a set of candidate IVs whose validity is difficult to determine from observational data. This paper studies the problem of testing the validity of IV sets under Causal Additive Models with Non-Constant Effects (CAM-NCE). To address this problem, we propose a testable condition, termed the Cross Auxiliary-based independence Test (CAT) condition, for assessing IV set validity from observational data. We show that, under the completeness condition, if the CAT condition is violated, the corresponding set cannot be a valid IV set. Furthermore, under a cross distributional non-degeneracy condition, we establish that the CAT condition becomes both necessary and sufficient for characterizing valid IV sets under CAM-NCE. We then extend the CAT condition to settings with covariates and develop a practical finite-sample algorithm for testing the validity of candidate IV sets. Extensive experiments on synthetic data and three real-world datasets demonstrate the effectiveness and practical utility of the proposed method.
- [23] arXiv:2608.19815 [pdf, html, other]
-
Title: Trustworthy Decisions in Reliability Set Estimation under Insufficient Model InformationComments: 30 pages,3 figuresSubjects: Statistics Theory (math.ST); Methodology (stat.ME)
Reliability set estimation identifies input regions where a response probability exceeds a target level, bridging estimation and safety-critical decisions. Practitioners typically start with a working model, an imperfect approximation of the true response surface. Relying on this imperfect model may incur decision risk, potentially certifying unsafe regions as safe. We develop a unified framework that turns such a working model into a trustworthy decision rule. First, a modeling-then-calibration procedure decouples estimation from decision. Since the true set is unobservable, we introduce an asymmetric, observable surrogate loss and use a separate calibration set to select a bias-correcting threshold, reducing decision risk and achieving $O_P(1/n)$ volume convergence. Second, we leverage conformal risk control with the surrogate loss to control false inclusion risk, which is the most safety-critical error, at a pre-specified level regardless of working model quality. Together, these calibration procedures show that a separate calibration set is necessary for risk control. Third, an adaptive design concentrates observations on the reliability set and its boundary, improving model quality where errors most affect decisions while controlling budget elsewhere. Numerical studies show not only more accurate set estimates but also calibrated finite-sample risk control that classical plug-in methods lack.
- [24] arXiv:2608.19849 [pdf, html, other]
-
Title: Distributional Extrapolation for InteractionsSubjects: Methodology (stat.ME); Machine Learning (stat.ML)
Predicting combinatorial effects from limited-range observations is a fundamental challenge in many scientific domains, including drug discovery and hyperparameter optimization. We study combinatorial extrapolation, where training data consists of axis-aligned samples with only one active covariate, while test-time inputs involve multiple simultaneously active covariates. We introduce DExtrI, a method for extrapolating interaction effects beyond the support of the training data. We provide theoretical guarantees characterizing when such extrapolation is possible. Empirical results on synthetic and real-world datasets demonstrate that DExtrI successfully generalizes to unseen combinations of covariates. Our approach enables applications such as predicting previously untested drug combinations and improving the efficiency of hyperparameter optimization.
- [25] arXiv:2608.19853 [pdf, html, other]
-
Title: Partial Identification Learning with Categorical Treatments for Individualized Treatment RulesSubjects: Methodology (stat.ME)
We develop a partial identification learning framework for individualized treatment rules (ITRs) with categorical treatments, outcomes, and instrumental variables. Rather than relying on strong causal assumptions required for point identification, our framework leverages causal bounds to characterize the optimal treatment decision. Existing methods for ITR optimization under partial identification are largely restricted to binary treatment settings and the bounds derived by Balke and Pearl under the canonical instrumental variable design. We extend this framework to accommodate a broader class of causal structures as well as scenarios with categorical treatment, outcome, and instrumental variables. We introduce a generalized minimax loss criterion for treatment selection from among more than two options, which minimizes the maximum possible difference between the chosen and the optimal treatment based on partial identification bounds. To construct the ITR, we use a symmetric embedding strategy that maps discrete treatments to the vertices of a regular simplex, avoiding the geometric inconsistencies of standard one-vs-rest approaches. We derive a differentiable, weighted surrogate risk function and show that optimizing it solves the original problem. Furthermore, we provide finite sample convergence rates via an oracle inequality under general regularity conditions, which we show are satisfied by a kernel based implementation. Numerical experiments demonstrate that the framework yields ITRs significantly closer to the oracle ITR compared to existing alternatives in settings with unmeasured confounding.
- [26] arXiv:2608.19874 [pdf, html, other]
-
Title: Combining Concurrent and Historical Functional Linear RegressionSubjects: Statistics Theory (math.ST); Methodology (stat.ME)
We study a function-on-function linear regression model in which the response at time $t$ depends on both the past trajectory of a predictor and its concurrent value. The model combines an $L^2$-historical effect with a point-evaluation effect, and these two coefficient functions are not automatically identifiable. We characterize the resulting non-identifiability and show that the concurrent and historical effects are separately identifiable whenever the covariance eigenfunctions of the covariance operator of the predictor are not pointwise square-summable. This mild and novel condition prevents the concurrent point evaluation from being represented by an $L^2$-historical effect.
Building on an orthogonalized representation of the predictor process, we propose a smoothing-spline estimator for both coefficient functions and establish consistency rates. The rates reveal an interesting trade-off between path regularity and eigenvalue decay: smoother predictor trajectories lead to faster convergence of the historical-effect estimator, whereas rougher trajectories lead to faster convergence of the concurrent-effect estimator. Simulation studies and two real-data applications demonstrate the practical importance of disentangling concurrent and historical effects. - [27] arXiv:2608.19879 [pdf, other]
-
Title: A Repeated Measurements Approach to $SoH$ Battery Modelling of Cyclic Aged Data in a Laboratory EnvironmentComments: 17 pages, 9 figures, 2 tablesSubjects: Methodology (stat.ME); Machine Learning (cs.LG); Systems and Control (eess.SY)
This document describes the application of a first order linearised nonlinear repeated measurements approach to the analysis of battery cell ageing profiles generated under controlled conditions in a laboratory. The primary advantage of the model is it reflects the obvious structure in the data. Consequently, it is a two-component of variance model: variation within ageing profiles (measurement noise) and variation among ageing profiles (test-to-test or cell-to-cell) variation. Novel regularised iterative generalised least squares parameter identification schemes, with optimal hyper-parameter re-estimation, are used to identify the hierarchical nonlinear model. The training data comprised $SoH$ profiles for 10 cells aged at various constant discharge and charge current cycles at a fixed chamber environmental temperature of 25 [$^\circ$C]. Each cell $SoH$ profile is modelled using a simple power law expression, whereas the variation in ageing parameters is modelled using a single knot cubic B-spline. $SoH$ is accurately predicted to $\pm 0.191\%$ for $SOH \in [0,20]$.
- [28] arXiv:2608.19903 [pdf, html, other]
-
Title: Where Does the Union Bound Go? Best-Arm Identification and Strong FWER ControlSubjects: Methodology (stat.ME); Machine Learning (cs.LG); Machine Learning (stat.ML)
In fixed-confidence best-arm identification, proofs often use a union bound across the competing arms. From a multiple-testing point of view this can look puzzling: if the best arm is unique, only one hypothesis of the form ``arm $i$ is best'' can be true. Why then should there be a Bonferroni-type factor of $K-1$? The answer is that there are two natural ways to orient the hypotheses. In one orientation, best-arm identification is literally a strong familywise-error-rate (FWER) problem with $K-1$ true nulls. In the opposite orientation, exactly one null is true, but a pairwise implementation can falsely reject that one null through any of $K-1$ comparisons. Thus the multiplicity has not disappeared; it just pops up in different places. This note makes the equivalence explicit in the terminology of both communities.
- [29] arXiv:2608.20046 [pdf, html, other]
-
Title: Integrating Temporal Disaggregation and Distributed Lag Nonlinear Models for Bayesian Spatio-Temporal Disease Mapping with High-Resolution Environmental ExposuresAlejandro Rozo Posada, Maxime Fajgenblat, Christel Faes, James Colborn, Emanuele Giorgi, Baltazar Candrinho, Thomas NeyensSubjects: Applications (stat.AP)
Environmental conditions are major drivers of malaria transmission, but epidemiological analyses are often constrained by temporal misalignment between health outcomes reported at coarse time scales and environmental exposures available at finer resolutions. Conventional approaches aggregate environmental data to match health outcomes, potentially obscuring delayed and nonlinear relationships. We propose a Bayesian spatio-temporal framework that addresses this limitation through a latent daily disease process linked to observed monthly malaria counts by temporal disaggregation. The framework integrates distributed lag nonlinear models for climatic effects, spatio-temporal random effects, and intervention covariates within a unified hierarchical model.
The methodology was applied to malaria surveillance data from 161 districts in Mozambique between 2017 and 2024, integrating temperature, precipitation, relative humidity, vegetation, elevation, and malaria interventions. Compared with a conventional monthly model, the proposed framework improved predictive accuracy and uncertainty quantification while exploiting the temporal resolution of environmental data. Estimated relationships showed nonlinear associations between climatic variability and malaria incidence, including an optimal temperature range, increasing risk with positive vegetation anomalies, and nonlinear precipitation effects.
By avoiding temporal aggregation of environmental exposures, the framework provides a flexible approach for investigating delayed environmental effects from routine surveillance data and can be extended to other environmentally sensitive diseases with mismatched temporal resolutions. - [30] arXiv:2608.20080 [pdf, html, other]
-
Title: Causal inference via propensity scores for case-control studiesYan Liu, Anita Koushik, Philippe Boileau, Cong Jiang, Miceline Mésidor, Claudia Waddingham, Denis Talbot, Mireille E. SchnitzerSubjects: Methodology (stat.ME)
Propensity score methods for causal inference are increasingly being used in cohort and experimental designs, but their development and uptake in outcome-dependent sampling schemes, such as case-control studies, remains limited. Case-control studies involve the sampling of individuals with and without an outcome of interest with the goal of estimating the effects of past exposures. When the design is observational, statistical adjustment for confounding bias is necessary. In case-control studies, propensity score models can be fit using control data under the assumption that the controls are representative of the source population with respect to their exposure distribution conditional on covariates ("control exchangeability"). In this paper, we first demonstrate that relative effects, such as causal risk ratios, are estimable under three different case-control design variants using control-fitted propensity scores. We appropriate two existing estimators for these designs: inverse probability of treatment weighting and an efficient and doubly robust estimator. We also introduce a novel two-step propensity score caliper-matching procedure for case-control designs. We introduce novel diagnostic tools to verify two necessary types of overlap. We then contrast our estimators using simulated data and apply them to examine the association between regular aspirin use and ovarian cancer risk.
- [31] arXiv:2608.20123 [pdf, html, other]
-
Title: Discrete Diffusion Inference-Time Control with Nested Sequential Monte CarloSubjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
We study inference-time control for text generation in discrete diffusion language models, where the goal is to steer sampling toward sequence-level rewards without retraining. Prior work in this domain has focused on particle-based methods such as best-of-$n$ sampling and bootstrap sequential Monte Carlo, which may suffer from overoptimism and weight degeneracy, respectively. We address these limitations using \emph{nested} sequential Monte Carlo methods. We formulate nested SMC (NSMC) and fully-adapted nested SMC (FA-NSMC) for Feynman--Kac steering, identifying and correcting errors in prior formulations that lead to biased final estimates. We evaluate these methods on toxicity and fluency steering tasks, showing that NSMC and FA-NSMC consistently outperform best-of-$n$ and bootstrap SMC.
- [32] arXiv:2608.20176 [pdf, html, other]
-
Title: The exact Spearman rho-footrule region via optimal transport with applications to finite rankings, mixability, and Chatterjee's rank correlationComments: 46 pages, 7 figuresSubjects: Statistics Theory (math.ST)
We solve the open problem of determining the maximal value of Spearman's rho when Spearman's footrule is prescribed, thereby completing the exact attainable region of these two quantities. To prove this result, we reformulate the underlying copula optimization problem as an optimal transport problem with a linear moment constraint and construct the unique optimal coupling through a matching feasible dual potential and its contact set. Equivalently, this coupling minimizes the variance of $|U-V|$ among all couplings of $U,V\sim\mathcal{U}(0,1)$ with prescribed mean $\mathbb{E}|U-V|$. Our main result admits several applications: First, in the context of finite rankings, we obtain an improved Cauchy--Schwarz inequality between Spearman's footrule distance and the associated quadratic rank difference. Second, in the framework of generalized mixability, we characterize the attainable constant values of $|U'+V'|$, for $U',V'\sim\mathcal{U}(-1/2,1/2)$, and determine the minimal quadratic deviation for a given mean. Third, we derive explicit bounds relating Chatterjee's rank correlation $\xi(X,Y)$, which can detect complex functional dependence of $Y$ on $X$, to the copula correlation ratio--a rank-based fraction of explained variance--by exploiting their conditional i.i.d. representations in terms of Spearman's footrule and Spearman's rho.
- [33] arXiv:2608.20187 [pdf, html, other]
-
Title: Multi-Method Causal Evidence Synthesis: Ranking Candidate Drivers by Convergent Cross-Method Evidence from Observational DataComments: 36 pages, 4 figures, 17 tables. Reference implementation available from the authors on requestSubjects: Methodology (stat.ME); Artificial Intelligence (cs.AI)
Practitioners inferring causality from observational data usually rely on a single method and treat its output as causal truth. Recent tools select an optimal method for a dataset, and recent ensembles aggregate multiple causal-discovery algorithms into one graph, but little work pools evidence across different mathematical traditions, including non-causal ones. We present Multi-Method Causal Evidence Synthesis (MCES), a framework that ranks which candidate drivers in an observational system are most likely relevant to a set of outcomes, and with what strength of evidence. MCES runs eleven methods across eight mathematical traditions on observational panel data and pools their outputs into a Convergent Evidence Score (CES), a linear opinion pool. CES quantifies convergence of evidence across analytical lenses: the degree to which methods with different assumptions point to the same driver-outcome relationship. It does not claim causal identification in the interventionist sense; it supports hypothesis prioritization, not a transferable probability of causation. MCES first applies Structural-Behavioral Decomposition to remove definitional (algebraic) relationships, then runs all methods, normalizes outputs to [0,1], and pools them. We distinguish MCES from method selection, structural ensembles, prediction ensembles, and literature synthesis. Using synthetic data with embedded ground truth, the Sachs protein-signaling benchmark, six Bayesian-network structure benchmarks, and two further synthetic domains, we show MCES ranks true edges near the top (Precision@5 = 1.0, Precision@10 = 0.96 on the primary scenario), with a low empirical rate of null pairs reaching Moderate-or-higher convergence. Our central point is not that the pool beats every individual method, but that no single method is uniformly best across the evaluated scenarios, so MCES offers a method-agnostic default.
- [34] arXiv:2608.20223 [pdf, html, other]
-
Title: Self-Normalizing Denominators in Rational Causal EstimationSubjects: Statistics Theory (math.ST)
Rational causal estimators in linear structural equation models take the form of one covariance polynomial divided by another, and a small denominator is commonly interpreted as weak identification. We show that, under Gaussian sampling, some denominators cannot enter this regime at first order. Their sampling variation is exactly proportional to their magnitude, so the standardized denominator is constant in every sample. Products of powers of nested covariance minors have this property in every dimension and admit an exact Wishart pivot. The converse is complete in dimension two. In dimension three, one mixed family remains open, while a factor-and-rank criterion classifies all denominators with linear or quadratic determinant-free factors and covers instrumental-variable, front-door and proximal formulas. For linear front-door adjustment, Wald inference remains asymptotically valid even as the mediator residual variance vanishes at an arbitrary rate, provided the treatment--mediator coefficient is nonzero. In simulations, proximal Wald coverage fell as a naive treatment--proxy diagnostic strengthened, while front-door coverage stayed nominal, and right-heart-catheterization data distinguished naive from denominator-relevant diagnostics.
- [35] arXiv:2608.20243 [pdf, html, other]
-
Title: A Bayesian Edge-Space Framework for Whole-Connectome Inference in Multisite Autism NeuroimagingComments: 49 pages, 4 main-text figures, 2 main-text tables; supplementary material includes theoretical results, simulation studies, computational diagnostics, and additional figuresSubjects: Methodology (stat.ME); Computation (stat.CO)
Autism spectrum disorder (ASD) is associated with heterogeneous alterations across distributed brain systems, creating challenges for whole-connectome inference. The difficulty arises not only from the large number of connections, but also from dependence among effects indexed by anatomically and functionally related region pairs. We introduce a Bayesian Edge-Space regression framework that treats each participant's connectome as a network-valued response and models the adjusted ASD effect over unordered brain-region pairs. The main methodological contribution is a positive-semidefinite covariance construction defined directly on connections. Anatomical and diagnosis-blind functional similarities are lifted from regions to edge space through a symmetrized endpoint-matching operation that preserves endpoint identity and is invariant to endpoint ordering. An additive Bayesian hierarchy estimates anatomical, functional, and interaction contributions together with multisite adjustments and connection-specific effects. Theoretical results establish covariance validity and continuous nesting of the structured components. Low-rank kernel representations and an exact sufficient-statistic reduction enable whole-connectome computation without preliminary edgewise estimation. Simulations show improved recovery of the effect surface, particularly under weak signals. In the Autism Brain Imaging Data Exchange, the framework identifies widespread reductions together with localized increases in ASD-associated connectivity. This pattern supports heterogeneous reorganization across distributed neural systems rather than uniform hyper- or hypoconnectivity. Under the fitted parameterization, the functional component has the largest structural scale, indicating organization beyond anatomical proximity alone.
- [36] arXiv:2608.20253 [pdf, html, other]
-
Title: GENIE: Generative Neural Inference for EpidemicsLaura M. Guzmán-Rincón, George R.E. Bradley, Joel Kandiah, Kyriakos Flouris, Pietro Liò, Paul J. Birrell, Alexander E. Zarebski, Daniela De AngelisSubjects: Methodology (stat.ME); Quantitative Methods (q-bio.QM)
The SARS-CoV-2 pandemic highlighted the ongoing risk infectious diseases pose to society and the value of reliable information on the likely future burden. When forecasting an epidemic at fine spatial resolution, traditionally used mechanistic compartmental model struggle to capture highly complex granular transmission dynamics, resulting in inaccurate and overconfident forecasts. However, detailed Agent-Based Models (ABMs), are challenging to calibrate and are too computationally expensive to use in real-time. Amortized simulation-based inference promises to overcome this difficulty by exploiting the power of machine learning (ML) to perform approximate forecasting at near-real-time using arbitrarily complex models of epidemics. In this work we introduce Generative Neural Inference for Epidemics (GENIE), a spatio-temporal ML-based framework for high-resolution forecasting of the burden of respiratory pathogens. GENIE is designed to reflect two key characteristics of outbreaks: (i) shared biological mechanisms across locations and (ii) location-specific characteristics affecting transmission dynamics. This results in the model architecture having two modules: (i) a Local Infection Encoder - which learns to represent disease dynamics shared across all locations and (ii) a Local Profile Encoder - which learns location-specific representations. Using simulations from a high-resolution spatio-temporal ABM, GENIE is trained to generate samples from an approximate posterior predictive distribution of future epidemic trajectories. Benchmarked against established statistical and ML models, GENIE demonstrates superior performance across a range of measures including the timing and magnitude of peak hospitalisations.
- [37] arXiv:2608.20255 [pdf, html, other]
-
Title: Transfer Learning in Nonparametric Regression with Deep ReLU NetworksComments: Accepted at the 43rd International Conference on Machine Learning (ICML 2026)Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
This paper develops a general transfer learning framework for nonparametric regression with data consisting of multiple groups. Under the assumption that groups share a common structure along with group-specific deviations in additive form, the proposed method employs a two-stage offset learning procedure: the first stage pools data from all groups to estimate an overall mean function, and the second stage estimates offsets for each group, yielding final group-level estimators through additive combination. Upper bounds on the $\mathcal L_2$ error are established for the proposed framework, covering a broad class of nonparametric estimators under mild complexity and noise conditions. When instantiated with deep ReLU networks, explicit convergence rates are derived under hierarchical composition models, demonstrating the ability to overcome the curse of dimensionality. Conditions that enable positive transfer with faster rates are considered, including learning with simpler functions and data augmentation through pooling samples across groups. Various simulations and real-data experiments further validate the effectiveness of the proposed method.
- [38] arXiv:2608.20260 [pdf, html, other]
-
Title: From Kriging to Spatial AI: Fifty Years of Spatial Statistics for Complex Dependent DataComments: 68 pages; invited review article on the development of spatial statistics from kriging and classical dependence modeling to modern Spatial AISubjects: Methodology (stat.ME)
Spatial statistics has grown from kriging for spatial prediction into a broad framework for learning from complex dependent data. This article traces that development from random fields and spectral methods to Bayesian hierarchical models and scalable computation. It then connects these foundations to Spatial AI, where graph learning and neural networks are being adapted to spatially dependent data. The article introduces the main ideas behind kriging and nonstationarity and explains how data fusion and uncertainty quantification extend spatial inference to more complex settings. The central contribution is a unified account of how these developments lead naturally to new forms of Spatial AI. Rather than treating spatial statistics and machine learning as separate traditions, we show how both learn from dependence while preserving interpretable structure. We also examine how spatial geometry and physical knowledge can guide flexible representation learning and support scientifically meaningful prediction.
- [39] arXiv:2608.20286 [pdf, html, other]
-
Title: Fast high-dimensional mean testing via logistic regressionSubjects: Methodology (stat.ME); Statistics Theory (math.ST)
We propose computationally efficient tests for equality of mean vectors of two or more high-dimensional populations. Central to our approach is an equivalence between equality of means and a zero population logistic regression parameter. We establish this equivalence for independently distributed observations without imposing common distributional assumptions across populations. Our procedure uses logistic Lasso to screen informative variables and an unpenalized logistic refit for inference in the reduced dimension, yielding asymptotically correct size and consistency. For a specified two-sample Gaussian submodel and sparse discriminative class, the test also attains the minimax separation rate. The framework extends to multiple populations through multi-class logistic regression. Simulations demonstrate accurate size control, strong power, and favorable computational scaling compared with existing tests under unbalanced designs and variance heterogeneity. Applications to gene-expression data with more than twenty-two thousand variables illustrate the practical scalability of the proposed procedures.
- [40] arXiv:2608.20321 [pdf, html, other]
-
Title: Large Sample Properties of Higher Order Markov ModelsSubjects: Statistics Theory (math.ST); Probability (math.PR)
We study large-sample properties of higher-order Markov chains on a finite alphabet $\Sigma$ when the order $m_n$ is allowed to grow with the sequence length $n$. By embedding the process into a first-order chain on $\Sigma^{m_n}$ and exploiting return-time decompositions, we establish a central limit theorem for additive functionals $\sum_{t}\! g_n(Y_t^{(n)})$ under natural ergodicity and sparsity conditions. The normalization involves the stationary return time to a suitably chosen state and accommodates triangular arrays with $m_n\!\to\!\infty$ and $m_n/n\!\to\!0$. We further illustrate the assumptions in a binary variable length Markov chain (VLMC), deriving explicit lower bounds on stationary masses that yield a concrete growth regime (e.g., $m_n\log m_n/n \to 0$) ensuring the CLT. These results provide asymptotic foundations for inference in sparse/partitioned higher-order models; including VLMCs and sparse Markov models (SMMs) where the effective dimensionality grows with the sample size.
New submissions (showing 40 of 40 entries)
- [41] arXiv:2608.19291 (cross-list from physics.soc-ph) [pdf, other]
-
Title: Rural Commute PatternsComments: This a brief submitted as a response to COMMUTING IN AMERICA 2022Subjects: Physics and Society (physics.soc-ph); Computation (stat.CO)
Transportation provides access to employment opportunities and essential services such as healthcare services, while urban areas have various transportation options, the situation differs in rural areas. Rural residents often have longer commute distances, limited access to public transit, and extended waiting times for public transportation if they exist, which can significantly impact their access to vital services and job opportunities. This study used data from the 2017 NHTS survey to examine the commuting patterns in rural areas by utilizing multinomial logistic regression to determine how various factors impact the choice of mode of transport in rural areas. Findings from this study revealed a higher dependency, 92.1%, on using personal vehicles when making trips in rural areas. Multinomial logistic regression results showed that socio-demographics, household, and trip characteristics affect the mode of transport used in a trip. Older adults, females, and individuals with higher education levels than high school graduates are less likely to use public transit when making trips. For household characteristics, the availability of vehicles in a household and households with higher income levels have lower probabilities of making a trip using public transit. Longer trip distances reduce the likelihood of a trip using active commuting modes such as walking and biking. These findings provide insights into understanding the transportation behaviors in rural areas and provide knowledge to be used in the planning and developing of transportation projects to promote equitable and accessible transportation in rural areas.
- [42] arXiv:2608.19306 (cross-list from quant-ph) [pdf, html, other]
-
Title: Quantum Gaussian processes for prediction of channel observationsJonas Jäger, Yaroslav Khmelnitskiy, Paolo Braccia, Artur Miroszewski, Diego García-Martín, M. Cerezo, Piotr CzarnikComments: 14 + 7 pages, 5 + 2 figuresSubjects: Quantum Physics (quant-ph); Machine Learning (cs.LG); Machine Learning (stat.ML)
Given a set of input states, we consider the task of predicting the expectation value of a Pauli observable at the output of an unknown quantum evolution, using only a limited number of measurements. Recently, quantum Gaussian process (QGP) regression was introduced for this task across various classes of unitary evolution. Here, we extend the QGP framework beyond unitary dynamics. In particular, we prove convergence of the channel's outputs to a QGP and derive the associated closed-form kernel under a uniform (Lebesgue measure) prior over quantum channels. The kernel's dimensional factor, however, dictates the required observation precision. While manageable when the channel and observable are restricted to small subsystems, exponential suppression precludes learning when the subsystem grows extensively with the system size. Since the Lebesgue prior is overly broad for many applications, we propose an empirical Bayes heuristic that replaces the dimensional factor with a learnable scale parameter while retaining the kernel's state-overlap correlation structure. In numerical simulations of up to 64 qubits, channel QGP regression with the Lebesgue kernel exhibits a strong inductive bias for local channels, enabling faithful extrapolation. For global 64-qubit channels, the rescaled kernel restores learnability, with predictions improving systematically with the shot budget. Results from a noisy quantum computer further demonstrate the robustness of QGP regression under experimental conditions. Beyond regression, we validate QGPs as Bayesian-optimization surrogates for state preparation under noisy XXZ dynamics.
- [43] arXiv:2608.19323 (cross-list from cs.LG) [pdf, html, other]
-
Title: Improved Confidence Estimates for Black-Box Large Language ModelsSubjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). Existing methods, from verbalized confidence to ones requiring multiple generations, are often zero-shot and produce scores quantifying uncertainty without the need for labelled data. Nonetheless, in practice one must always evaluate their performance on a dataset of interest before deployment. In this work we show that, by leveraging this dataset, we consistently outperform these existing scores. Specifically, we build simple classifiers that predict LLM response correctness by using these scores and the correctness of similar queries as features. Our method produces minimal computational overhead, making it a cheap and straightforward enhancement for UQ in LLMs for real-world applications.
- [44] arXiv:2608.19491 (cross-list from cs.LG) [pdf, html, other]
-
Title: DeltaMomentum: A Key-Value based Anisotropic Momentum Update via Delta RuleSubjects: Machine Learning (cs.LG); Computation and Language (cs.CL); Optimization and Control (math.OC); Machine Learning (stat.ML)
Most modern optimizers form their momentum as an exponential moving average (EMA) of past gradients, forgetting every direction at one fixed rate. However, the inputs a deep network sees during training can be highly anisotropic, with a few directions queried frequently while most are seen rarely. Recent methods address this anisotropy by wrapping extra processing around this buffer, leaving the momentum update itself unchanged. We propose DeltaMomentum, which builds direction-awareness into the momentum update rule. The main observation is that the gradient of a linear layer splits into an input that acts as a key and an output-side error that acts as a value. Exploiting the key-value structure, DeltaMomentum updates the momentum buffer by the canonical delta rule, so each direction is forgotten at a rate set by how often it appears. We prove that it is a valid momentum, that it applies the input-side curvature correction without matrix inversion, and that it clears stale directions faster than EMA under both a fixed and a drifting optimum. It is a drop-in replacement for the momentum buffer of any optimizer, its coefficient transfers across widths under $\mu$P, and its extra compute stays between $22.2\%$ and $25.0\%$ of a gated-MLP block's linear cost with no persistent memory. In FineWeb-Edu pretraining, AdamW with DeltaMomentum (DeltaAdamW) reaches AdamW's validation loss in up to $46.39 \pm 4.32\%$ fewer steps at 67M and $22.12 \pm 0.80\%$ at 370M over three seeds, and the gain persists at 1B on a Chinchilla-optimal budget. A Muon baseline tuned under the same protocol sits above DeltaAdamW at both language-model scales, and the gain holds for SGD, ResNet-18, and ViT-Tiny on CIFAR-10. Training-time diagnostics confirm the predicted mechanism, better gradient tracking and healthier input directions.
- [45] arXiv:2608.19584 (cross-list from cs.LG) [pdf, html, other]
-
Title: Kähler landscapes for complex neural network descents and guarantees including a search and destroy of the Calabi-Yau manifoldComments: First versionSubjects: Machine Learning (cs.LG); Differential Geometry (math.DG); Machine Learning (stat.ML)
We study landscapes for complex-parameterized networks. Our approach is motivated with an information-theoretic manifold perspective of the parameter and via classical optimization guarantees although of complex geometric variety such as through Dolbeault asymptotics. The descent path admits a Kähler information metric under a cross-entropy via the Wirtinger Hessian on the log-likelihood potential. We restrict attention to a descent update rule with natural gradient descent via a differentiated loss scaled by the inverse metric, so the descent path remains in the holomorphic tangent bundle. We emphasize Calabi-Yau information manifolds which profane theoretical guarantees via an ill-curvature-conditioned landscape. Under a Calabi-Yau metric, specifically in a non-compact setting with a global potential so defined geometrically rather than invoking the topological requirements of the Calabi conjecture, a wedged nowhere-vanishing holomorphic form is the top exterior product of the Kähler form up to constants, yielding a constant determinant condition. Under a fixed determinant, a metric almost low rank up to an eigenvalue tolerance implies a blow-up effect. Moreover, it has been discovered that negative curvature subverts the loss landscape, specifically sectional curvature, so we expand on this and draw interconnections to negative-definite Ricci curvature. Our arguments primarily exist in a geometric analytic modality, although we establish roots in deep learning theory such as through asymptotics at initialization and connections through failure modes of neural network guarantees under vanishing and negative Ricci curvature.
- [46] arXiv:2608.19643 (cross-list from cs.LG) [pdf, html, other]
-
Title: Time-Uniform Self-Normalized Concentration for Discounted Least Squares: Limits and CorrectionsSubjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Self-normalized concentration inequalities are standard tools in bandit and reinforcement-learning analyses. A widely used weighted extension claims an analogous time-uniform guarantee for discounted least-squares estimators in non-stationary problems. A simple scalar Gaussian counterexample with a fixed parameter shows that the claimed bounded radius is crossed with probability one. For fixed discount and regularization parameters, we further show that, when $\delta\leq1/2$ and $T/\delta$ is sufficiently large, any deterministic anytime boundary valid uniformly over the stated conditionally sub-Gaussian model class must be at least of order $R\sqrt{\log(T/\delta)}$ at some time by horizon $T$; for nondecreasing boundaries, this order is required at time $T$. We identify the proof error: different terminal times use different Gaussian mixing distributions, so the fixed-time mixtures do not form one supermartingale, and the stopping-time argument does not repair this failure. Finally, we show that the weighted inequality remains valid at each fixed deterministic time, give valid finite- and infinite-horizon corrections, and discuss consequences for downstream analyses.
- [47] arXiv:2608.19762 (cross-list from cs.LG) [pdf, html, other]
-
Title: Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamWSubjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC); Machine Learning (stat.ML)
A minibatch can influence training beyond the update at which it is observed because AdamW stores past gradient information in its optimizer states. We study this delayed effect through paired trajectories that differ only in one gradient update and share the same subsequent training sequence. We formulate AdamW as a finite-horizon input--state--output (ISO) system whose state contains the model parameters and first- and second-moment estimates. Linearizing the joint dynamics yields a signed response operator that maps a localized gradient perturbation to its future loss effects, revealing how optimizer memory shapes their magnitude, timing, and sign. We further derive an exact multistep error decomposition and establish first-order finite-horizon accuracy under local smoothness and controlled activation switching. Experiments validate the response mechanism and optimizer-state effects, while repeated-future analyses reveal substantial prospective structure in delayed influence that can be partially recovered from ISO approximations. Code is available at this https URL.
- [48] arXiv:2608.19908 (cross-list from cs.IT) [pdf, html, other]
-
Title: A Layered Simplex Architecture for Large AlphabetsComments: 30 pages, 7 figuresSubjects: Information Theory (cs.IT); Machine Learning (cs.LG); Machine Learning (stat.ML)
Probability estimation over large alphabets under log loss is a well-studied problem, with celebrated methods such as the Good-Turing estimator. We introduce and study a new Bayesian estimator with four notable properties. First, its construction is exceptionally simple: multiply independent uniform draws from the probability simplex coordinate-wise and renormalize. Depth is the only structural parameter, and averaging over depths eliminates the need to tune it. Second, the regret of the resulting mixture, the excess code length it pays relative to a code that knows the source, admits an explicit and efficiently computable expression. Third, despite its simplicity and lack of tuned constants, the estimator is competitive across a diverse set of synthetic and real-text benchmarks with substantially more specialized methods, including Good-Turing. Fourth, the tractability of its regret allows us to identify scaling laws in data, alphabet size, and depth. For Zipf targets with exponent above one, the regret has a simple reading as long as the sample reveals only a small fraction of the alphabet. It closely matches the description length of the set of discovered symbols, at one bit of code per bit of description, plus a further cost per symbol. The data exponent is therefore the rate at which new symbols are discovered.
- [49] arXiv:2608.19930 (cross-list from physics.bio-ph) [pdf, html, other]
-
Title: SoilWaterNow: Soil water nowcasting for mapping plant available water (PAW) across paddocks for improved on-farm decision-makingJournal-ref: GRDC Grains Research Update 2026, Goondiwindi, Australia, 3-4 March, 149-157Subjects: Biological Physics (physics.bio-ph); Geophysics (physics.geo-ph); Applications (stat.AP)
Timely, paddock-scale estimates of plant-available water (PAW) can support crop management in water-limited grain systems. We present the Sydney Soil Water-Energy Balance (SWEB) model, a scalable, physically consistent modelling framework that integrates satellite and climate data with soil properties to estimate crop evapotranspiration and root-zone soil moisture (RZSM) at daily, 30 m resolution. SWEB was validated nationally against different SM monitoring networks across Australia, which demonstrated robust performance across diverse grain-growing environments (correlation coefficients generally ranged from 0.70 to 0.85 for most regions). In future applications, mid-season PAW nowcasts could be combined with water-use-efficiency approaches to estimate water-limited yield potential and inform responsive management decisions, such as nitrogen fertiliser top-up recommendations.
- [50] arXiv:2608.20016 (cross-list from physics.soc-ph) [pdf, html, other]
-
Title: Emergence of cooperation: A reputation-modulated reinforcement learningSubjects: Physics and Society (physics.soc-ph); Disordered Systems and Neural Networks (cond-mat.dis-nn); Adaptation and Self-Organizing Systems (nlin.AO); Populations and Evolution (q-bio.PE); Machine Learning (stat.ML)
Reputation is widely recognized as a key mechanism for sustaining cooperation. However, most existing game-theoretic models treat reputation primarily as an external factor that modulates payoffs, interaction structures, or strategy update rules. In many social contexts, though, reputation operates primarily as information -- it shapes how individuals interpret their own experiences and assess the behavior of others. To bridge this gap, we propose a spatial prisoner's dilemma game grounded in the reinforcement learning paradigm, in which agents equipped with Q-learning integrate both individual and social information via a locally defined reputation metric to guide their decisions. Our results reveal that reputation-modulated learning significantly promotes the emergence of cooperative behavior, and we observe a discontinuous phase transition from full cooperation to full defection as the temptation increases. Cooperation spreads through the nucleation of cooperative clusters, whereas the disintegration of these clusters drives the system into an absorbing state of complete defection. Overall, this study demonstrates that reputation facilitates cooperation not only by providing direct incentives but also by reshaping the social information landscape that agents rely on for learning and adaptation.
- [51] arXiv:2608.20183 (cross-list from cs.LG) [pdf, html, other]
-
Title: Exact Algebraic Computation of Learning Coefficients for Two-Dimensional Singular ModelsGrégoire Sergeant-Perthuis (1), Elias Tsigaridas (2), Jules Tsukahara (2) ((1) CQSB, Sorbonne Université, (2) Ouragan Team, INRIA)Subjects: Machine Learning (cs.LG); Symbolic Computation (cs.SC); Algebraic Geometry (math.AG); Machine Learning (stat.ML)
Classical information criteria such as the Bayesian Information Criterion (BIC) rely on regularity assumptions that break down for singular models, leading to incorrect model selection in settings such as deep learning. The Widely Applicable Bayesian Information Criterion (WBIC) relies on local learning coefficients $\lambda$, which in the analytic case coincides with local Real Log Canonical Thresholds (RLCT) of the Kullback-Leibler divergence of the model, to capture correct marginal likelihood asymptotics. Exact computation of the learning coefficients has been limited to special cases, and only sampling-based estimation methods are generally applicable. We present the first deterministic algorithm that computes local RLCTs exactly for any two-dimensional model whose Kullback-Leibler distance is contact equivalent to a polynomial, derive a bound on its complexity, and demonstrate its effectiveness for a broad class of models, with applications including polynomial neural networks. Beyond providing ground truth to calibrate sampling-based estimators, exact computation reveals algebraic structure in learning coefficients that sampling cannot and out-speeds it in the shallow regime.
- [52] arXiv:2608.20258 (cross-list from cs.LG) [pdf, html, other]
-
Title: DICS: Data-Informed Centroid Splitting for Decision Tree ClassifiersSubjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Decision tree-based models are widely used in machine learning due to their interpretability and strong empirical performance. However, training decision trees can be computationally expensive, particularly for large and high-dimensional datasets, largely due to the exhaustive search over candidate splits at each node. To improve computational efficiency, we propose Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors. By incorporating class-aware structure, DICS significantly reduces the split search space for classification tasks while preserving predictive performance. We further provide theoretical analysis showing that under the stated assumptions, DICS does not degrade the performance of classification trees compared to exhaustive split search. DICS can be incorporated into classification trees, random forests, and gradient-boosting models. Extensive experiments demonstrate that DICS achieves comparable accuracy while substantially reducing training time across synthetic and benchmark datasets, highlighting the benefit of integrating data-informed priors into split selection for scalable classification tree learning.
- [53] arXiv:2608.20279 (cross-list from math.PR) [pdf, html, other]
-
Title: Robustness of random-walk Metropolis for steep potentialsComments: 11 pages, no figuresSubjects: Probability (math.PR); Statistics Theory (math.ST); Computation (stat.CO)
In Markov chain Monte Carlo sampling, light-tailed target distributions present something of a poisoned chalice: their light tails offer good confinement, and tend to imply good mixing properties for natural continuous-time dynamics, but the steepness of their tail decay means that they often fall outside of the scope of modern quantitative convergence theory. For usual gradient-based samplers, this reflects a genuine instability issue, whereby Metropolis acceptance rates can degrade badly. In this work, we study the gradient-free random-walk Metropolis sampler, and show that for a wide range of light-tailed targets, the acceptance probability remains stable for reasonable choices of proposal variance, from which effective and favourable mixing time estimates can be deduced. The analysis relies on a simple relationship between the first and second derivatives of the log-density of the target distribution.
- [54] arXiv:2608.20295 (cross-list from cs.LG) [pdf, html, other]
-
Title: Physical-Support Confidence Sets for Highly Coherent DictionariesSubjects: Machine Learning (cs.LG); Signal Processing (eess.SP); Statistics Theory (math.ST)
Sparse pursuit after dictionary learning can yield a precise atom support even when its physical interpretation is not justified by the calibration data, especially for highly coherent dictionaries where alternative calibration-compatible dictionaries may assign different physical meanings to the same selected support. We develop resolution-aware physical-support inference that jointly accounts for uncertainty in the learned dictionary and in the representation of a deployment signal. Our cross-dictionary confidence correspondence retains calibration-compatible dictionaries and deployment-compatible sparse representations, then projects the surviving explanations onto physical-support space. For local coherent-atom classes with separation scale s, once the deployment data resolve the coherent-block explanation and its atom support, the minimax physical resolution from N calibration signals satisfies $\delta_{\mathrm{opt}}(N,s)\asymp\min\{s,\frac{1}{\sqrt{N}s^2}\}$, with relative resolution governed by the orientation-information scale $Ns^6$. Deployment replication improves physical localization only when orientation changes cannot be absorbed by adjusting the active coefficients. For computation, we introduce active endpoint bracketing (AEB), an adaptive finite-bank procedure that evaluates only candidates that can still affect the physical report and otherwise safely coarsens or abstains. Finite-bank experiments, including a four-region synthetic application, show that a point-valued plug-in selector can be physically overprecise, whereas AEB avoids unsupported refinement with fewer candidate evaluations.
- [55] arXiv:2608.20337 (cross-list from math.PR) [pdf, html, other]
-
Title: Information on trajectories: martingales and random timesSubjects: Probability (math.PR); Information Theory (cs.IT); Machine Learning (cs.LG); Statistics Theory (math.ST)
Accounting for information flow on the path space of trajectories of a nonnegative martingale yields exact variational identities for it, even at arbitrary random times. This recovers the widely used classical concentration inequalities, from Ville to PAC-Bayes, and measures what each one discards. The tail a bound controls is itself a relative entropy, resolved by the chain rule into per-step conditional divergences. The discarded slack has an exact form in each of three geometries: a Gibbs tilt for the Azuma-Hoeffding and PAC-Bayes bounds, the crossing itself for Ville's and for pooled tests, and a dominating certificate for the $L^p$ maximal bound. That certificate's optional-stopping deficit resolves per step into Bregman divergences of the running maximum. On a path-time space, the same identity gains one factor that prices anticipation: an arbitrary random time carries an e-process ``peeking penalty.'' The partition function can be read as a coalescent--a prefix-sharing probability of independent copies--and geometric mixtures of test martingales gain a pooling benefit for multi-model safe testing.
Cross submissions (showing 15 of 15 entries)
- [56] arXiv:2103.04021 (replaced) [pdf, html, other]
-
Title: Asymptotic Theory for IV-Based Reinforcement Learning with Potential EndogeneitySubjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Econometrics (econ.EM); Optimization and Control (math.OC)
In the standard data analysis framework, data is collected (once and for all), and then data analysis is carried out. However, with the advancement of digital technology, decision-makers constantly analyze past data and generate new data through their decisions. We model this as a Markov decision process and show that the dynamic interaction between data generation and data analysis leads to a new type of bias -- reinforcement bias -- that exacerbates the endogeneity problem in standard data analysis. We propose a class of instrument variable (IV)-based reinforcement learning (RL) algorithms to correct for the bias and establish their theoretical properties by incorporating them into a stochastic approximation (SA) framework. Our analysis accommodates iterate-dependent Markovian structures and, therefore, can be used to study RL algorithms with policy improvement. We also provide formulas for inference on optimal policies of the IV-RL algorithms. These formulas highlight how intertemporal dependency of the Markovian environment affects the inference.
- [57] arXiv:2107.07575 (replaced) [pdf, html, other]
-
Title: Optimal tests of the composite null hypothesis arising in mediation analysisComments: 73 pages, 12 figuresSubjects: Statistics Theory (math.ST); Methodology (stat.ME)
The indirect effect of an exposure on an outcome through an intermediate variable can be identified by a product of two regression coefficients under certain causal and regression modeling assumptions. In this context, the null hypothesis of no indirect effect is a composite null hypothesis, as the null holds if either regression coefficient is zero. A consequence is that traditional hypothesis tests are severely underpowered near the origin (i.e., when both coefficients are small with respect to standard errors). We propose hypothesis tests that (i) preserve level alpha type 1 error, (ii) meaningfully improve power when both true underlying effects are small relative to sample size, and (iii) preserve power when at least one is not. One approach gives a closed-form test that is minimax optimal with respect to local power over the alternative parameter space. Another uses sparse linear programming to produce an approximately optimal test for a Bayes risk criterion. We discuss adaptations for performing large-scale hypothesis testing as well as modifications that yield improved interpretability. We provide an R package that implements our proposed methodology.
- [58] arXiv:2407.14778 (replaced) [pdf, html, other]
-
Title: Minimax estimation of functionals in sparse vector model with correlated observationsJournal-ref: Electronic Journal of Statistics 20.2 (2026): 3770-3806Subjects: Statistics Theory (math.ST)
We consider the observations of an unknown $s$-sparse vector ${\boldsymbol \theta}$ corrupted by Gaussian noise with zero mean and unknown covariance matrix ${\boldsymbol \Sigma}$. We propose minimax optimal methods of estimating the $\ell_2$ norm of ${\boldsymbol \theta}$ and testing the hypothesis $H_0: {\boldsymbol \theta}=0$ against sparse alternatives when only partial information about ${\boldsymbol \Sigma}$ is available, such as an upper bound on its Frobenius norm and the values of its diagonal entries to within an unknown scaling factor. We show that the minimax rates of the estimation and testing are leveraged not by the dimension of the problem but by the value of the Frobenius norm of ${\boldsymbol \Sigma}$.
- [59] arXiv:2409.07948 (replaced) [pdf, html, other]
-
Title: Quickest Change Detection Using Mismatched CUSUMComments: Extended version of extended abstract for the Allerton Conference on Communication, Control, and Computing, September 2024Subjects: Statistics Theory (math.ST); Information Theory (cs.IT)
Quickest change detection concerns estimation of an unknown change time \(\tau_a\) from a sequence of partial observations \(\{Y_k:k\ge 0\}\). We consider stopping rules of CUSUM form, \[
\mathcal{X}_{n+1}
=
\max\{0,\mathcal{X}_n+F(Y_{n+1})\},
\quad
\tau_s=\min\{n\ge 0:\mathcal{X}_n\ge \textrm{H}\}, \] where the function \(F\) and threshold \(\textrm{H}\) are design parameters.
The observations and change time are modeled jointly through a hidden Markov model, and \( F\) is selected from a prescribed function class \(\mathcal{G}\) to minimize the weighted criterion \[
\textsf{E}\bigl[
(\tau_s-\tau_a)_+
+
\kappa(\tau_s-\tau_a)_-
\bigr]. \]
When \(\mathcal{G}\) is a linear function class, the optimizer
\(F^*\) is characterized by a convex program, whose dual yields extensions of classical likelihood-ratio constructions. This conclusion is based on analysis that is asymptotic in the regime \(\kappa\to\infty\). We show that the hidden Markov model admits an asymptotically equivalent conditionally independent approximation of the type commonly used in the quickest change detection literature. We then develop the design and asymptotic theory for a substantially broader class of conditionally independent models, so that the resulting conclusions are not tied to the particular POMDP reduction.
Combining renewal theory and large deviations for reflected random walks, we obtain for each $F\in\mathcal{G}$ asymptotically accurate approximations of the optimal threshold and average cost, with error vanishing as \(\kappa\to\infty\). It is found in numerical experiments that the resulting approximations are accurate for moderate values of \(\kappa\). - [60] arXiv:2410.15955 (replaced) [pdf, html, other]
-
Title: The mutual arrangement of Wright-Fisher diffusion path measures and its impact on parameter estimationSubjects: Statistics Theory (math.ST); Probability (math.PR); Populations and Evolution (q-bio.PE)
The Wright-Fisher diffusion is a fundamentally important model of evolution encompassing genetic drift, mutation, and natural selection. Suppose you want to infer the parameters associated with these processes from an observed sample path. Then to write down the likelihood one first needs to know the mutual arrangement of two path measures under different parametrizations; that is, whether they are absolutely continuous, equivalent, singular, and so on. In this paper we give a complete answer to this question by finding the separating times for the diffusion - the stopping time before which one measure is absolutely continuous with respect to the other and after which the pair is mutually singular. In one dimension this extends a classical result of Dawson on the local equivalence between neutral and non-neutral Wright-Fisher diffusion measures. Along the way we also develop new zero-one type laws for the diffusion on its approach to, and emergence from, the boundary. As an application we derive an explicit expression for the joint maximum likelihood estimator of the mutation and selection parameters and show that its convergence properties are closely related to the separating time.
- [61] arXiv:2505.10510 (replaced) [pdf, html, other]
-
Title: Efficient Uncertainty Propagation in Bayesian Two-Step ProceduresSubjects: Methodology (stat.ME)
Bayesian inference provides a principled framework for probabilistic reasoning. If inference is performed in two steps, uncertainty propagation plays a crucial role in accounting for all sources of uncertainty and variability. This becomes particularly important when both aleatoric uncertainty, caused by data variability, and epistemic uncertainty, arising from incomplete knowledge or missing data, are present. Examples include surrogate models and missing data problems. In surrogate modeling, the surrogate is used as a simplified approximation of a resource-heavy and costly simulation. The uncertainty from the surrogate-fitting process can be propagated using a two-step procedure. For modeling with missing data, methods like Multivariate Imputation by Chained Equations (MICE) generate multiple datasets to account for imputation uncertainty. These approaches, however, are computationally expensive, as multiple models must be fitted separately to surrogate parameters respectively imputed datasets. To address these challenges, we propose an efficient two-step approach that reduces computational overhead while maintaining accuracy. By selecting a representative subset of draws or imputations, we construct a mixture distribution to approximate the desired posteriors using Pareto smoothed importance sampling. For more complex scenarios, this is further refined with importance weighted moment matching and an iterative procedure that broadens the mixture distribution to better capture diverse posterior distributions.
- [62] arXiv:2505.20005 (replaced) [pdf, html, other]
-
Title: Existence of penalised likelihood estimates and posterior propriety of separable prior distributions for Gaussian precision matricesComments: 31 pagesSubjects: Statistics Theory (math.ST); Methodology (stat.ME)
Penalised likelihoods are often used for sparse estimation of a Gaussian precision matrix. In high dimensional settings where the matrix dimension is larger than the sample size, the sample covariance matrix $S$ is not of full rank and the maximum likelihood estimate of the precision matrix does not exist. An additional advantage of some penalised likelihood estimates, for example the graphical lasso, is that it can exist even in such high dimensional settings. This paper gives a thorough analysis of the existence of penalised likelihood estimates for positive semidefinite $S$. Specific tail conditions are provided on the diagonal and off-diagonal penalty functions that ensure existence of the estimate. This is also extended to the Bayesian setting where conditions on separable prior distributions are provided that ensure the resulting posterior distribution is proper.
- [63] arXiv:2509.01604 (replaced) [pdf, html, other]
-
Title: A Time-Series Model for Areal Data Using Area-Specific Gaussian Processes with Spatially Correlated HyperparametersAlejandro Rozo Posada, Oswaldo Gressani, Christel Faes, James Colborn, Baltazar Candrinho, Emanuele Giorgi, Thomas NeyensComments: Supplementary material includedSubjects: Methodology (stat.ME)
In many applied settings, areal data are observed repeatedly over long time periods, as commonly occurs in infectious disease surveillance and environmental or demographic monitoring. Accurate characterization of local temporal dynamics and uncertainty is important for monitoring disease trends, identifying local changes, and supporting public health decision-making. Traditional spatio-temporal models generally represent spatial, temporal, and space-time interaction components through structured random effects acting on the latent outcome process. We propose a Bayesian spatio-temporal hierarchical framework in which temporal dynamics are modeled using area-specific Gaussian processes, while spatial dependence is introduced through spatially correlated Gaussian-process covariance hyperparameters. This allows neighboring regions to share information about the characteristics of their temporal dependence while retaining area-specific temporal trajectories, providing an alternative representation of spatio-temporal dependence. Inference is performed using Markov chain Monte Carlo methods. The approach is illustrated using monthly malaria incidence data from three Mozambican provinces and evaluated using the Root Mean Squared Error, Continuous Ranked Probability Score, empirical coverage probability, and credible interval width. Compared with established spatio-temporal models, the proposed framework achieves competitive predictive accuracy and better-calibrated predictive uncertainty. Spatially structuring the temporal hyperparameters also improves predictive interval calibration relative to an equivalent model with independent temporal processes, demonstrating the value of borrowing spatial information on temporal dependence and supporting this formulation as a competitive alternative to conventional outcome-level spatial smoothing for spatio-temporal areal data in disease surveillance.
- [64] arXiv:2509.12587 (replaced) [pdf, html, other]
-
Title: Inverse regression for causal inference with multiple outcomesComments: 80 pages, 5 figuresSubjects: Methodology (stat.ME)
With multiple outcomes in empirical research, a common strategy is to define a composite outcome as a weighted average of the original outcomes. However, the choices of weights are often subjective and can be controversial. We propose an inverse regression strategy for causal inference with multiple outcomes. The key idea is to regress the treatment on the outcomes, which is the inverse of the standard regression of the outcomes on the treatment. Although this strategy is simple and even counterintuitive, it has several advantages. First, testing for zero coefficients of the outcomes is equivalent to testing for zero treatment effects, even though the inverse regression is deemed misspecified. Second, the coefficients of the outcomes provide data-driven weights for defining a composite outcome. Interestingly, these weights maximize the standardized effect size, a measure commonly used for univariate outcomes in empirical studies. We also discuss the associated inference issues. Third, this strategy is applicable to general study designs. We illustrate the theory in both randomized experiments and observational studies. We implement the proposed method in our R package invreg, available at this https URL.
- [65] arXiv:2510.06524 (replaced) [pdf, html, other]
-
Title: Stable Central Limit Theorems for Discrete-Time Lag Martingale Difference Arrays: Applications to Dynamic Causal InferenceSubjects: Statistics Theory (math.ST); Probability (math.PR)
Recent work in dynamic causal inference introduced a class of discrete-time stochastic processes that generalize martingale difference sequences and arrays as follows: the random variates in each sequence have expectation zero given certain lagged filtrations but not given the natural filtration. We formalize this class of stochastic processes and prove stable central limit theorems (CLTs) via martingale-coboundary decomposition, leveraging the classical martingale CLT. We develop a variety of sufficient conditions, including conditions under which the limiting variance has a simple form that depends on variances and covariances of neighboring variates. We demonstrate the application of these results to inference for time-averaged treatment effects in switchback designs and present a simulation study supporting their validity. The CLTs enable various extensions to existing methodology for design-based approaches to dynamic causal inference, including time-lagged effects, random limiting variances, cross-unit dependence, and vector-valued estimands.
- [66] arXiv:2512.03322 (replaced) [pdf, html, other]
-
Title: SSLfmm: An R Package for Semi-Supervised Learning with Mixed MissingnessComments: 20 pages, 2 figureSubjects: Computation (stat.CO); Methodology (stat.ME); Machine Learning (stat.ML)
Partially labelled samples arise when features are observed for all data, but class labels are available for only a subset. In such settings, the mechanism governing label availability may itself contain information relevant to classification, yet it is typically left unmodelled in standard semi-supervised learning procedures. The SSLfmm package implements likelihood-based Gaussian finite-mixture classification in which the label-missingness process is modelled jointly with the class distribution. It supports complete-case, missing completely at random (MCAR), entropy-based missing at random (MAR), and mixed MCAR/MAR analyses. For the mixed mechanism, the source of a missing label may be observed or latent, allowing the same modelling framework to accommodate different forms of information about label availability. A common R interface is provided for model fitting, prediction, performance assessment, simulation, and entropy-based diagnostics. We describe the statistical formulation and software implementation, position SSLfmm relative to existing finite-mixture and semi-supervised learning software, and demonstrate its use through a reproducible simulation comparing observed- and latent-source analyses. A semi-synthetic application to the Blood Transfusion data further illustrates how alternative assumptions about label missingness can be fitted, compared, and diagnosed in practice.
- [67] arXiv:2601.02610 (replaced) [pdf, html, other]
-
Title: Reliable conformal novelty detection at the decision boundaryComments: 48 pages, 17 figures, 1 tableSubjects: Methodology (stat.ME); Machine Learning (stat.ML)
Novelty detection via conformal $p$-values and BH procedure provides distribution-free global false discovery rate (FDR) control. We present here fundamental limits of this approach by showing that it does not produce reliable detection at the decision boundary. We study boundary false discovery rate (bFDR), the probability that the least extreme reported novelty is in fact a null observation. We first show that the support line (SL) procedure, controlling the bFDR in the continuous independent framework, fails to control the bFDR in the conformal case. We then present several modifications of the SL procedure that restore reliability at the decision boundary, by controlling the bFDR, with specific improvements in situations where many novelties are expected (adaptive procedures) and when the calibration sample is too small with respect to the test sample (subsampled procedures). Numerical experiments with both synthetic and real data support our findings and show the relevance of the new proposed approach.
- [68] arXiv:2601.11422 (replaced) [pdf, html, other]
-
Title: Stein's method for the matrix normal distributionComments: 25 pages, 0 figuresSubjects: Statistics Theory (math.ST); Probability (math.PR)
This work presents the first systematic development of Stein's method for matrix distributions. We establish the basic essential ingredients of Stein's method for matrix normal approximation: we derive an extended-generator-based Stein identity from a matrix Ornstein-Uhlenbeck diffusion with two-sided scales, provide an explicit semigroup representation for the solution of the Stein equation, and obtain regularity estimates for the solution. The new methodology is demonstrated in three examples: (i) smooth Wasserstein distance bounds to quantify the matrix central limit theorem (a didactic example), (ii) a Wasserstein distance bound for the matrix normal approximation of the centered matrix $T$ distribution, and (iii) a Stein's method-of-moments approach to estimating the row and column covariance factors of the matrix normal, yielding a flexible class of weighted flip-flop Stein estimators that generalize Dutilleul's classical flip-flop algorithm and naturally accommodate row/column importance weights, systematic missingness, and projection onto structured covariance families. The latter two examples are intrinsically matrix-valued and cannot be treated using naive vectorization.
- [69] arXiv:2605.10069 (replaced) [pdf, html, other]
-
Title: Estimating Consensus Epidemic Trajectories via a Constrained Power Fréchet Mean with Functional RegistrationSubjects: Applications (stat.AP)
In infectious disease modeling during the early phase of a pandemic, SEIR-type compartmental models are standard tools, and they require epidemiological parameters as inputs. Because these parameters are subject to uncertainty, different research groups often report different epidemic curves, and obtaining a representative curve that captures the characteristics of these trajectories is important for decision-making. However, simple pointwise summaries can attenuate epidemic peaks under temporal misalignment and generally do not preserve the dynamical structure of the underlying compartmental models. To address these limitations, we propose a method for summarizing multiple solutions to SEIR-type compartmental models on a functional space by computing a constrained power Fréchet mean with temporal shift registration. In our method, we regard the pairs of exposed and infectious compartments as objects in a Hilbert space, and the consensus curve is defined as the solution to a constrained optimization problem. Differential equation constraints and population constraints are incorporated in the optimization to preserve a partially mechanistic interpretation regarding the infectious compartment. We develop an implementable block-optimization algorithm based on basis function expansion. In simulation studies based on early COVID-19 parameter estimates, the proposed method produced consensus curves with a single epidemic peak, whereas pointwise summaries exhibited attenuated or multiple peaks. We further applied the method to six literature-derived parameter sets from early COVID-19 studies and obtained a representative trajectory with interpretable epidemiological parameters. The proposed approach provides a generalized trajectory-summarization method that includes mean- and median-type estimators and preserves mechanistic interpretability.
- [70] arXiv:2605.22950 (replaced) [pdf, other]
-
Title: Diffusion-based Denoising Beats Vanilla Score Matching in Parameter Estimation: A Theoretical ExplanationSubjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST); Methodology (stat.ME)
Score matching is an alternative to maximum likelihood estimation when the normalizing constant is unknown or too costly to evaluate. However, vanilla score matching has shown to be inefficient relative to maximum likelihood estimation for multimodal distributions with well-separated modes, which are commonly encountered in practical applications. We compare a novel diffusion-based denoising score matching estimator (DDSME) to the vanilla score matching estimator (SME) in this scenario. In particular, we prove statistical guarantees for both estimators, showing that the error bound for the vanilla SME worsens when the separation between the modes increases, which can be avoided in case of the DDSME with suitable hyperparameter tuning. This provides a novel theoretical explanation for the superior behavior of diffusion-based score matching over the vanilla version. We support our theoretical findings by numerical experiments.
- [71] arXiv:2606.01011 (replaced) [pdf, html, other]
-
Title: Semiparametric Efficiency of Residual Correlation Testing under Gaussian Additive Noise ModelsSubjects: Statistics Theory (math.ST); Methodology (stat.ME)
This paper studies conditional independence testing under the Gaussian additive noise model (GANM), where two variables are modeled as nonlinear functions of covariates with independent bivariate Gaussian regression errors. Under this framework, conditional independence can be characterized by the correlation coefficient of the regression errors, which motivates a test based on the Pearson correlation coefficient computed from the fitted residuals. Despite its simple form, the asymptotic behavior and statistical efficiency of the resulting test have not been well understood. In this paper, we develop the semiparametric efficiency theory under GANM and show, surprisingly, that the efficient estimator coincides exactly with the ordinary residual Pearson correlation estimator. We further establish the asymptotic properties of the proposed test and develop the corresponding inference procedure. Simulation studies demonstrate that the proposed method achieves near-oracle efficiency and competitive empirical power while maintaining valid Type I error control. We further apply the proposed test to conditional dependence analysis of U.S. stock returns.
- [72] arXiv:2606.01554 (replaced) [pdf, html, other]
-
Title: Fast Near-Optimal Estimation over Symmetric Norm BallsComments: fixed a bunch of typosSubjects: Statistics Theory (math.ST)
This short note proposes a polynomial-time algorithm for near-optimal Euclidean estimation of a signal constrained to lie in the unit ball of a symmetric norm, where the symmetry is with respect to a known basis and the norm is accessible through an evaluation oracle. We further extend the method to a random-design, moderate-dimensional linear regression setting, where the regression parameter is likewise assumed to belong to a constraint set defined by a symmetric norm.
- [73] arXiv:2606.26804 (replaced) [pdf, html, other]
-
Title: Structured Secant Methods to Select Smoothing Parameters for General Smooth ModelsComments: v2: references updated, clarifications on applicability of Eq. 11Subjects: Methodology (stat.ME); Computation (stat.CO)
General smooth models replace parameters of a regular likelihood with additive models. The models can include parametric terms, Gaussian random effects, and smooth functions of covariates. The latter are parameterized via a reduced-rank spline basis and regularized via weighted quadratic penalties placed on the basis coefficients. Estimates for these weights (i.e., smoothing parameters) can be obtained by optimizing the Laplace-approximate Bayesian marginal likelihood. Existing (second-order) methods require the Hessian of the log-likelihood to solve this optimization problem approximately - exact optimization requires up to fourth order derivatives - which can be difficult to derive and expensive to evaluate. To address these problems, we present a quasi-Newton variant of the second-order Extended Fellner-Schall (EFS) optimization method. Our qEFS method relies on structured limited-memory secant approximations to the Hessian of the log-likelihood and is principally first-order. However, the approximation can also be accumulated for a sub-block of the Hessian, with the remaining columns being constrained to match those of the actual Hessian. The exact columns then provide additional structure for the sub-block approximation, which becomes more accurate as a result. We show that the qEFS method converges to the EFS method under certain conditions and continues to provide good estimates beyond these circumstances, which we illustrate in simulation studies. Secondary tasks involving the Hessian (confidence interval coverage & model selection) require partial approximations to achieve close to nominal performance. We provide Hidden Markov and Tweedie model examples, for which the qEFS method is substantially easier to implement than alternative methods.
- [74] arXiv:2606.27662 (replaced) [pdf, html, other]
-
Title: Design-Aware Variance Reduction for Switchback Experiments: A Comparative StudySubjects: Methodology (stat.ME)
Switchback experiments and other clustered randomized designs are widely used on online platforms, but the clustered, time-dependent nature of these designs can make standard variance reduction methods behave differently than in standard A/B tests. We evaluate design-aware variance reduction methods for switchbacks -- CUPED, CUPAC (ML-based covariate adjustment), and doubly robust (DR) estimators -- relative to a baseline switchback analysis with cluster-robust standard errors. Through a hierarchical simulation framework that varies key regime parameters -- number of clusters, cluster-size imbalance, within-cluster autocorrelation, carryover, and predictive signal strength -- we evaluate validity (false positive rate and confidence interval coverage) and efficiency (standard error reduction, power, and minimum detectable effect as a function of run length). We also include a sensitivity analysis for cross-cluster spillovers to quantify bias and inference degradation under mild interference. The primary outcome is a practitioner-oriented regime map: when CUPED, CUPAC, or DR are most beneficial, and when time and cluster dependence and finite-cluster effects limit improvements.
- [75] arXiv:2607.18652 (replaced) [pdf, html, other]
-
Title: The Price of Hidden Curvature: Improved Lower Bounds for Bandit Convex OptimizationComments: 44 pages, 2 figuresSubjects: Machine Learning (stat.ML); Information Theory (cs.IT); Machine Learning (cs.LG)
We establish improved lower bounds on the minimax expected regret of stochastic bandit convex optimization for $1$-Lipschitz functions on the $d$-dimensional Euclidean ball. For time horizons $n\ge d^{10/3}$, we prove a lower bound of $\Omega(d^{4/3}\sqrt{n})$, the first nontrivial bound that exceeds the $d\sqrt{n}$ dependence of linear bandits, showing that stochastic bandit convex optimization is fundamentally harder than linear bandits. For $d^2\le n\le d^{10/3}$, we obtain a lower bound of $\Omega(\sqrt{d}n^{3/4})$, matching the regret of the algorithm of Flaxman et al. (2005), establishing its optimality in this regime.
The hard class of convex functions we construct takes the following form in dimension $2d$: for an action $a=(a^1,a^2)\in \mathbb{B}^{2d}$, each function is the scaled soft maximum of a "tube", $r^{-1}\|W^\star a^1-\frac{r}{8\varepsilon}a^2 \|$ (hyperparameterized by $\varepsilon,r$), and a squared distance function, $\frac12\|a^1-u^\star\|^2-\frac12\|u^\star\|^2$. Here $u^\star\in\mathbb{R}^d$ is the unknown target determining the minimizer, while $W^\star\in\mathbb{R}^{d\times d}$ hides the region in which the quadratic curvature is observable. Indeed, observations reveal substantial information about $u^\star$ only when the learner acts near the hidden tube $a^2\approx \frac{8\varepsilon}{r}W^\star a^1$; away from it, the tube branch masks the quadratic branch. Thus the learner must pay to uncover the geometry encoded by $W^\star$ before it can effectively exploit the curvature that identifies $u^\star$. Formalizing this tradeoff yields a sample complexity lower bound of $\Omega(\frac{d^{5/2}}{\varepsilon^2}\wedge\frac{d^2}{\varepsilon^4})$ for finding an $\varepsilon$-optimal action, and ultimately the $\Omega(d^{4/3}\sqrt{n}\wedge\sqrt{d}n^{3/4})$ regret lower bound.
The proof was developed by GPT-5.5 Pro and GPT-5.6 Sol Pro under the authors' guidance. - [76] arXiv:2607.21847 (replaced) [pdf, html, other]
-
Title: Distributional Determinantal Point Process for Repulsive Clustering of DistributionsComments: 61 pages, 15 figuresSubjects: Methodology (stat.ME); Machine Learning (cs.LG); Applications (stat.AP); Computation (stat.CO); Machine Learning (stat.ML)
We introduce the distributional determinantal point process (dDPP) as a novel repulsive point process whose atoms are probability distributions rather than points in a real space. The dDPP is constructed via an L-ensemble with a sliced Wasserstein (SW) kernel between distributions. We show its validity as a well-defined point process. In the discrete setting, we derive concentration results for plug-in estimators of the L-ensemble, the correlation kernel, and their determinants given i.i.d. samples from the distributional atoms. Leveraging this framework, we propose a distribution-valued random partition model by way of a repulsive generalized Bayesian mixture model. The model places a dDPP prior over the atoms of the mixing measure and defines a generalized likelihood based on SW distance. To summarize posterior inference, we develop a decision-theoretic approach to report a point estimate of the mixing measure as a Bayes rule under a hierarchical optimal transport utility function. The latter is a natural choice given that the mixing measure is itself a distribution over distributions. We use the proposed framework for inference with single-cell gene expression data and human epilepsy data, producing interpretable and well-separated clusters that reflect meaningful structure in the data.
- [77] arXiv:2607.23287 (replaced) [pdf, html, other]
-
Title: Learning Asymptotics with Convergence-Rate Guarantees using Linear Least SquaresComments: 62 pages, 7 tables, 5 figures; additional referencesSubjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Combinatorics (math.CO); Numerical Analysis (math.NA)
We introduce a new research area that is called Asymptotics Learning Theory (ALT) and combines optimization with asymptotic analysis. In particular, ALT provides a unified approach for computing unknown constants/parameters in proven asymptotic expansions using optimization theory. In this paper, we focus on a general asymptotic form which includes a broad class of asymptotics. Furthermore, we study two powerful numerical methods, namely, sliding Linear Least Squares (sLLSQ) and sliding Tikhonov Linear Least Squares (sT-LLSQ). For these techniques we rigorously prove asymptotic estimates that lead to sufficient conditions for convergence (to the correct values of unknown parameters) and convergence-rate guarantees. Despite their strengths, both methods have also limitations, e.g., slow convergence---or even, counterintuitively, divergence---in some cases. Moreover, we present fundamental applications in analytic combinatorics, a beautiful field of mathematics that deals with asymptotic enumeration of discrete structures using complex analysis. The proposed techniques complement existing approaches, such as the ratio method and its variants. Numerical examples also verify the theoretical results. Finally, we discuss interesting research directions in ALT.
- [78] arXiv:2608.11180 (replaced) [pdf, html, other]
-
Title: Bayesian inference on beta diversity via feature allocation models with imperfect detectionSubjects: Methodology (stat.ME)
Beta diversity quantifies variation in species composition across ecological communities and is fundamental for understanding biodiversity patterns across space and environmental gradients. Statistical inference on beta diversity is challenging: species occurrence data are high dimensional, many species remain unobserved despite extensive sampling, and surveys are subject to imperfect detection. Existing approaches are typically based on empirical dissimilarity indices with limited uncertainty quantification or on models that rely on unrealistic exchangeability and perfect detection assumptions. We introduce a new class of Bayesian feature allocation models for partially exchangeable species occurrence data with imperfect detection. The framework combines latent feature allocation models with occupancy-based detection mechanisms, allowing heterogeneous species compositions across sites, explicitly accounting for false negatives, and accommodating the discovery of previously unobserved species. We develop coherent probabilistic inference for species sharing and beta diversity, deriving explicit posterior and predictive distributions for between-community heterogeneity, including the number of shared species across sites and the number expected under future sampling. These analytical results yield interpretable posterior summaries of compositional heterogeneity and facilitate scalable inference in high-dimensional biodiversity studies. Simulation studies and an application to global fungal biodiversity data demonstrate improved inference on species sharing and between-community diversity.
- [79] arXiv:1808.08763 (replaced) [pdf, html, other]
-
Title: On the convergence of optimistic policy iteration for stochastic shortest path problemComments: 15 pagesSubjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
In this paper, we prove some convergence results of a special case of optimistic policy iteration algorithm for stochastic shortest path problem. We consider both Monte Carlo and $TD(\lambda)$ methods for the policy evaluation step under the condition that the termination state will eventually be reached almost surely.
- [80] arXiv:1905.11395 (replaced) [pdf, html, other]
-
Title: Multi-Modal Graph Interaction for Multi-Graph Convolution Network in Urban Spatiotemporal ForecastingLingyu Zhang, Xu Geng, Zhiwei Qin, Hongjun Wang, Xiao Wang, Ying Zhang, Jian Liang, Guobin Wu, Xuan Song, Yunhai WangSubjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Graph convolution network based approaches have been recently used to model region-wise relationships in region-level prediction problems in urban computing. Each relationship represents a kind of spatial dependency, like region-wise distance or functional similarity. To incorporate multiple relationships into spatial feature extraction, we define the problem as a multi-modal machine learning problem on multi-graph convolution networks. Leveraging the advantage of multi-modal machine learning, we propose to develop modality interaction mechanisms for this problem, in order to reduce generalization error by reinforcing the learning of multimodal coordinated representations. In this work, we propose two interaction techniques for handling features in lower layers and higher layers respectively. In lower layers, we propose grouped GCN to combine the graph connectivity from different modalities for more complete spatial feature extraction. In higher layers, we adapt multi-linear relationship networks to GCN by exploring the dimension transformation and freezing part of the covariance structure. The adapted approach, called multi-linear relationship GCN, learns more generalized features to overcome the train-test divergence induced by time shifting. We evaluated our model on ridehailing demand forecasting problem using two real-world datasets. The proposed technique outperforms state-of-the art baselines in terms of prediction accuracy, training efficiency, interpretability and model robustness.
- [81] arXiv:2212.09391 (replaced) [pdf, html, other]
-
Title: Bounds for matrices of inclusion probabilities in rejective samplingSubjects: Probability (math.PR); Statistics Theory (math.ST)
Motivated by a desire to model n-sparse signals more realistically, we study matrices of inclusion probabilities for conditional Poisson or rejective sampling in the non-asymptotic regime. In particular, we provide bounds in the semi-definite ordering for matrices that collect first and second order inclusion probabilities as their diagonal and off-diagonal entries respectively, and operator norm bounds of their Hadamard products with other matrices. Finally, we demonstrate the usefulness on these bounds, which only involve first order inclusion probabilities, in a toy application from dictionary learning.
- [82] arXiv:2306.11154 (replaced) [pdf, html, other]
-
Title: An Isotonic Mechanism for Overlapping OwnershipSubjects: Computer Science and Game Theory (cs.GT); Theoretical Economics (econ.TH); Applications (stat.AP)
Motivated by the problem of improving peer review at large scientific conferences, this paper studies how to elicit self-evaluations to improve review scores in a natural many-to-many owner-item (e.g., author-paper) situation with overlapping ownership. We design a simple, efficient and truthful mechanism to elicit self-evaluations from item owners that can be used to calibrate their noisy review scores in the existing evaluation process (e.g., papers' review scores from peers).
Our approach starts by partitioning the owner-item relation structure into disjoint blocks, each sharing a common set of co-owners. We then elicit the ranking of items from each owner and employ isotonic regression to produce adjusted item scores, aligning with both the reported rankings and raw item review scores. We prove that truth-telling by all owners is a payoff dominant Nash equilibrium for any valid partition of the overlapping ownership sets under natural conditions. Moreover, the truthfulness depends on eliciting rankings independently within each block, making block partition optimization crucial for improving statistical efficiency. Despite being computationally intractable in general, we develop a nearly linear-time greedy algorithm that provably finds a performant block partition with appealing robust approximation guarantees. Extensive experiments on both synthetic data and real-world conference review data demonstrate the effectiveness of our mechanism in a pressing real-world problem. - [83] arXiv:2511.03618 (replaced) [pdf, html, other]
-
Title: Towards Formalizing Reinforcement Learning Theory: A Robbins-Siegmund ApproachComments: Reinforcement Learning JournalSubjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
In this paper, we formalize the almost sure convergence of $Q$-learning and linear temporal difference (TD) learning with Markovian samples using the Lean 4 theorem prover based on the Mathlib library. $Q$-learning and linear TD are among the earliest and most influential reinforcement learning (RL) algorithms. The investigation of their convergence properties is not only a major research topic during the early development of the RL field but also receives significant attention nowadays. This paper formally verifies their almost sure convergence in a unified framework based on the Robbins-Siegmund theorem. The framework developed in this work can potentially be extended to convergence rates and other modes of convergence. This work thus makes an important step towards fully formalizing convergent RL results. The code is available at this https URL.
- [84] arXiv:2512.11785 (replaced) [pdf, html, other]
-
Title: Universal entrywise eigenvector fluctuations in delocalized spiked matrix models and asymptotics of rounded spectral algorithmsComments: 64 pages, 7 figures. v2: Results expanded to include analysis of multi-spectral and multi-frequency algorithms, as well as other small correctionsSubjects: Probability (math.PR); Data Structures and Algorithms (cs.DS); Statistics Theory (math.ST)
We consider the distribution of the top eigenvector $\widehat{v}$ of a spiked matrix model of the form $H = \theta vv^* + W$, in the supercritical regime where $H$ has an outlier eigenvalue of comparable magnitude to $\|W\|$. We show that, if $v$ is sufficiently delocalized, then the distribution of the individual entries of the projector $\widehat{v}\widehat{v}^*$ (not, we emphasize, merely the inner product $|\langle \widehat{v}, v\rangle|^2$) is universal over a large class of generalized Wigner matrices $W$ having independent entries, depending only on the first two moments of the distributions of the entries of $W$. This complements the observation of Capitaine and Donati-Martin (2021) that these distributions are not universal when $v$ is instead sufficiently localized. Further, for $W$ having entrywise variances close to constant and thus resembling a Wigner matrix, we show by comparing to $W$ drawn from the Gaussian orthogonal or unitary ensembles that averages of entrywise functions of $\widehat{v}\widehat{v}^*$ behave as they would if $\widehat{v}$ had Gaussian fluctuations around a suitable multiple of $v$. We also establish such results for several possibly dependent spiked matrices, showing that, if such matrices are entrywise uncorrelated, then their leading eigenvectors behave as they would with independent Gaussian fluctuations. We apply these results to spectral algorithms with rounding procedures for synchronization problems over the cyclic and circle groups, obtaining the first precise asymptotic error rates for such algorithms. Using our analysis of multiple spiked matrices, we also show that multi-frequency spectral algorithms using estimates from several matrices often have asymptotic error rate superior to that of naive spectral algorithms using just one matrix.
- [85] arXiv:2605.29411 (replaced) [pdf, html, other]
-
Title: The Good, the Bad, and the Ugly of Markov Boundary for Tabular PredictionComments: 12 pages, 9 figures, 2 tables. Accepted at CIKM 2026 (35th ACM International Conference on Information and Knowledge Management)Journal-ref: Proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), November 7-11, 2026, Rome, ItalySubjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Methodology (stat.ME); Machine Learning (stat.ML)
Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redundant. Once the boundary is observed, the target is conditionally independent of the rest of the table. This is a tempting object for tabular prediction, since it names exactly the columns a model should need. Yet modern regressors are still trained on the full feature set. We ask whether the Markov boundary is genuinely useful for prediction on SCM3K, a 3,450-task synthetic SCM benchmark with feature counts from 40 to 1000 and six SCM families, evaluated with six regressors. The answer is more nuanced than the theory suggests. Restricting a regressor to the oracle boundary often improves prediction substantially, and the improvement grows as the feature space becomes larger and sparser. But the natural pipeline of recovering the boundary with causal discovery and training on the recovered mask does not deliver. Existing estimators exhaust the compute budget before reaching the regime where the boundary helps most, and even where they run they rarely beat the full feature set. We trace this to three causes. Discovery optimizes structural recovery rather than prediction. False negatives and false positives carry sharply asymmetric predictive cost. The exact boundary is only one of many feature sets that beat all features. We then develop what these facts imply for prediction-aligned feature selection and for tabular models that learn to use causal structure.
- [86] arXiv:2606.01034 (replaced) [pdf, html, other]
-
Title: A Finite-Calibration Regime Map for LLM Judge PanelsComments: 33 pages, 11 figures, 40 tables. Accepted at WISE 2026Subjects: Computation and Language (cs.CL); Methodology (stat.ME)
Deploying an LLM judge panel spends human labels on fitting a calibrator, constructing candidate judge paths, and validating which candidate to deploy. We study when finite labels should support a low-dimensional stacker or reliability model, and when an unrestricted joint output table is worth its cell-count and unseen-pattern cost. We cast this as a finite-calibration regime map and instantiate it as Finite-Calibration Panel Selection (FCPS), a validation selector over judge path, deployed panel size, and aggregator family with support diagnostics. Across RewardBench, LLMBar, SummEval, and Arena100K with a seven-judge pool, scalar/reliability aggregation has lower MSE than unrestricted joint-table calibration in 16 of 20 real dataset--budget cells by point estimate, while paired 95% intervals exclude zero in 11 cells; richer backoff/shrinkage tables narrow some gaps while preserving the finite-support bottleneck. Controlled calibration-growth data show the opposite regime: when labels contain a six-way interaction, the selected table grows to the interaction-bearing prefix and its MSE falls from 0.224 to 0.061 once unseen mass vanishes. The practical deployment question is whether the next judge's information is estimable under the available human labels.