-
FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation
Authors:
Junjie Luo,
Xuzhe Zhi,
Rui Han,
Abhimanyu Kumbara,
Anand K. Iyer,
Mansur E. Shomali,
Ritu Agarwal,
Guodong Gordon Gao
Abstract:
As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medicati…
▽ More
As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows that subgroup performance gaps align with the proportion of clinically hard cases, and that input-length sensitivity varies across demographics, motivating personalized configurations. Frontier LLMs underperform specialized neural models by 1-6 mg/dL; behavioral events contribute negligibly (approximately 0.1 mg/dL) even under oracle event access. These findings establish that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Improving lepton flavour universality tests with $K_L$ decays
Authors:
G. D'Ambrosio,
A. M. Iyer,
F. Mahmoudi,
S. Neshatpour
Abstract:
Rare kaon decays provide sensitive probes of the flavour structure of the Standard Model and of possible new physics. We perform a global analysis incorporating recent experimental results and updated Standard Model predictions, including the latest measurement of $K^+ \to π^+ ν\barν$ and lepton flavour universality observables in $K^+ \to π^+ \ell^+\ell^-$. The fit favours a best-fit point close…
▽ More
Rare kaon decays provide sensitive probes of the flavour structure of the Standard Model and of possible new physics. We perform a global analysis incorporating recent experimental results and updated Standard Model predictions, including the latest measurement of $K^+ \to π^+ ν\barν$ and lepton flavour universality observables in $K^+ \to π^+ \ell^+\ell^-$. The fit favours a best-fit point close to the Standard Model, while a second local minimum remains phenomenologically relevant. We define benchmark scenarios associated with these two regions and investigate the prospective sensitivity of NA62 and KOTO-II to the new physics parameter space. We consider projected measurements of $K^+ \to π^+ ν\barν$ and lepton flavour universality observables in $K^+ \to π^+ \ell^+\ell^-$ at NA62, and of $K_L \to π^0 ν\barν$, $K_L \to π^0 e^+e^-$, and $K_L \to π^0 μ^+μ^-$ at KOTO-II. We find that KOTO-II has significant potential to probe and discriminate between the viable new physics scenarios, with NA62 providing complementary sensitivity.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Quantum-classical crossover in fault-tolerant quantum dynamics simulation
Authors:
Jinzhao Sun,
Bozhen Zhou,
Jue Xu,
Yuan Yao,
Zhenyu Du,
Zixu Zhang,
Yuntian Gu,
Junxiang Huang,
Shuo Zhou,
Ziruo Wang,
Alexander Yosifov,
Wenzheng Dong,
Yiming Huang,
Daniel Serrano,
Xinzhao Wang,
Tianfeng Feng,
Shreyas Sadugol,
Wenjun Yu,
Zhou You,
Dayue Qin,
Xiao-Ming Zhang,
Yantao Wu,
Aditya Iyer,
You Zhou,
Tongyang Li
, et al. (6 additional authors not shown)
Abstract:
While quantum computers promise to solve classically intractable problems, identifying the point at which fault-tolerant quantum computation outperforms the best classical algorithms for practical applications remains an outstanding challenge. Here we establish a concrete quantum-classical crossover for quantum many-body dynamics under realistic hardware conditions. We introduce a scalable fault-t…
▽ More
While quantum computers promise to solve classically intractable problems, identifying the point at which fault-tolerant quantum computation outperforms the best classical algorithms for practical applications remains an outstanding challenge. Here we establish a concrete quantum-classical crossover for quantum many-body dynamics under realistic hardware conditions. We introduce a scalable fault-tolerant framework that combines coherent observable estimation with a space-time-efficient implementation of non-Clifford rotations, suppressing the residual logical errors that limit existing partially fault-tolerant approaches. A benchmark against state-of-the-art tensor-network and variational Monte Carlo algorithms reveals a concrete crossover for mixed-field Ising dynamics at modest system sizes. For a physical error rate of $p=10^{-3}$, fault-tolerant simulation requires approximately 2 hours and $3.7 \times 10^5$ physical qubits for a 100-site 1D system, whereas tensor network approaches would require about 100 years. For 2D models, where rapid entanglement growth limits the classical evolution time, we project quantum runtimes within minutes. A physical error rate of $p=10^{-4}$ leads to at least an order of magnitude reduction in qubit count ($3.1 \times 10^4$ physical qubits) and runtime (minutes for 1D and seconds for 2D). The reduction in quantum runtime arises from our improved rotation-state injection and co-design of quantum error correction and observable-estimation protocols, which jointly suppress logical-error accumulation and reduce sampling overhead. Our results establish a scalable route towards practical quantum advantage and identify quantitative engineering targets for future fault-tolerant architectures.
△ Less
Submitted 17 July, 2026;
originally announced July 2026.
-
Matter-Field Exchange Generates Entanglement, Not Classical Gravity
Authors:
Nicetu Tibau Vidal,
Aditya Varna Iyer
Abstract:
Aziz and Howl have argued that a hybrid theory with quantum matter and a classical gravitational field can generate entanglement between two massive systems. In their construction, the branch-dependent contribution appears at fourth order through propagators of the quantum matter field in a fixed classical gravitational potential. We analyse this mechanism within the same perturbative QFT framewor…
▽ More
Aziz and Howl have argued that a hybrid theory with quantum matter and a classical gravitational field can generate entanglement between two massive systems. In their construction, the branch-dependent contribution appears at fourth order through propagators of the quantum matter field in a fixed classical gravitational potential. We analyse this mechanism within the same perturbative QFT framework and show that the effect should not be interpreted as classical gravity mediating entanglement in the sense relevant to BMV-type witnesses. The non-separable term relies on a quantum-matter exchange channel between the two interferometers and is present only when the two systems are modelled as excitations of the same matter field. If distinct, non-interconverting matter fields describe the systems, the corresponding cross-propagator is absent and the Aziz-Howl entangling diagram vanishes. The effect relies on coherent propagation amplitudes of the quantum matter field between the two interferometers, together with postselection onto the original localised branch subspace. Moreover, the inference of entanglement is made after projecting the full QFT evolution onto a restricted final subspace containing the original localised $N$-particle branch states. This projection removes precisely the sectors that would record matter-field contamination, mode deformation, or exchange between the two interferometers. We therefore argue that the Aziz-Howl mechanism is a matter-sector cross-talk effect in a classical background, not entanglement mediated by classical gravitational degrees of freedom. Having identified the channel responsible for the Aziz-Howl contribution, we predict that it can be eliminated by using distinct, non-interconverting matter species in the two interferometers, or by inserting a barrier that suppresses matter-field propagation between them.
△ Less
Submitted 3 July, 2026;
originally announced July 2026.
-
KLAS: Using Similarity to Stitch Neural Networks for Improved Accuracy-Efficiency Tradeoffs
Authors:
Debopam Sanyal,
Anantharaman Iyer,
Alind Khare,
Trisha Jain,
Akshay Jajoo,
Myungjin Lee,
Clayton Kerce,
Alexey Tumanov
Abstract:
Given the wide range of deployment targets, flexible model selection is essential for optimizing performance within a given compute budget. Recent work demonstrates that stitching pretrained models within a model family enables cost-effective interpolation of the accuracy-efficiency tradeoff space. Stitching transforms intermediate activations from one pretrained model into another, producing a ne…
▽ More
Given the wide range of deployment targets, flexible model selection is essential for optimizing performance within a given compute budget. Recent work demonstrates that stitching pretrained models within a model family enables cost-effective interpolation of the accuracy-efficiency tradeoff space. Stitching transforms intermediate activations from one pretrained model into another, producing a new interpolated stitched network. Such networks provide a pool of deployment options along the accuracy-efficiency spectrum. However, existing stitching approaches often yield suboptimal tradeoffs and lack generalizability, as they primarily rely on heuristics to select stitch configurations. We argue that constructing improved accuracy-efficiency tradeoffs requires explicitly capturing and leveraging the similarity between pretrained models being stitched. To this end, we introduce KLAS, a novel stitch selection framework that automates and generalizes stitch selection across model families by leveraging KL divergence between intermediate representations. KLAS identifies the most promising binary stitches from the $O(k^2n^2)$ possibilities for $k$ pretrained models of depth $n$. Through comprehensive experiments, we demonstrate that KLAS improves the accuracy-efficiency curve of stitched models at the same finetuning cost as baselines. KLAS achieves up to $1.21\%$ higher ImageNet-1K top-1 accuracy at the same computational cost, or maintains accuracy with a $1.33\times$ reduction in FLOPs.
△ Less
Submitted 27 May, 2026;
originally announced May 2026.
-
Entanglement-facilitated macroscopic cluster formation in quantum many-body dynamics
Authors:
Xiao Wang,
Alexander Yosifov,
Aditya Iyer,
Jinzhao Sun
Abstract:
Metastable quantum many-body dynamics could facilitate the organisation of microscopic degrees of freedom into macroscopic structures. However, the conditions under which this occurs are not well understood. Here we study false-vacuum decay in a 2D quantum Ising model and show that the initial correlation structure can qualitatively change this behaviour. Compared with product-state initialisation…
▽ More
Metastable quantum many-body dynamics could facilitate the organisation of microscopic degrees of freedom into macroscopic structures. However, the conditions under which this occurs are not well understood. Here we study false-vacuum decay in a 2D quantum Ising model and show that the initial correlation structure can qualitatively change this behaviour. Compared with product-state initialisations, correlated false-vacuum states suppress the proliferation of small true-vacuum domains and favour the formation of macroscopic connected clusters. Tree tensor network simulations of lattices up to $25 \times 25$ further reveal that nucleation proceeds predominantly from the boundary. Finite-size scaling demonstrates the dominant connected cluster remains an extensive fraction of the system even as its size increases. By suppressing this edge-assisted nucleation pathway through boundary pinning, we generate large magnetisation fluctuations consistent with macroscopic superposition. This mechanism relies on the 2D nucleation barrier and is absent in 1D systems or product-state quenches. Our results identify correlated state preparation and boundary engineering as complementary techniques for controlling metastable quantum dynamics.
△ Less
Submitted 14 July, 2026; v1 submitted 21 May, 2026;
originally announced May 2026.
-
Validating Coronal Magnetic Field Models Using Gaussian Separation
Authors:
Abhinav G. Iyer,
Michael S. Wheatland,
Brian T. Welsch,
Yang Liu,
S. A. Gilchrist
Abstract:
Nonlinear Force-free Field (NLFFF) models are widely used to investigate coronal magnetic field structure in solar active regions, but methods to validate them remain limited. Here, we use Gaussian separation, recently applied to solar vector magnetogram data, to assess the accuracy of NLFFF models constructed with two methods: optimization and the current-field iteration (CFIT) implementation of…
▽ More
Nonlinear Force-free Field (NLFFF) models are widely used to investigate coronal magnetic field structure in solar active regions, but methods to validate them remain limited. Here, we use Gaussian separation, recently applied to solar vector magnetogram data, to assess the accuracy of NLFFF models constructed with two methods: optimization and the current-field iteration (CFIT) implementation of the Grad-Rubin method. Gaussian separation partitions the photospheric vector magnetic field into three components associated with currents flowing below, above, and passing through the photosphere, respectively. Comparing the photospheric field components due to coronal currents in an NLFFF model with those in the original vector magnetogram data provides a check on the accuracy of the model's coronal currents. We consider NLFFF models constructed for the active region AR 11429. The photospheric signatures of coronal currents in both the models and the vector magnetogram data indicate currents flowing above and parallel to central, sheared polarity inversion lines (PILs), consistent with other recent studies. We find that while both models reproduce the coronal current signatures along the upper section of the main PIL, the CFIT model significantly alters the signature of a flux rope along the lower section of the PIL, including shifting its positive-polarity footpoint. These differences arise from modifications to the vector magnetogram boundary data when solving the NLFFF equations, and from the assumptions underlying the models. We propose Gaussian separation as a useful tool to validate coronal magnetic field models, in addition to existing methods.
△ Less
Submitted 10 May, 2026;
originally announced May 2026.
-
High-Fidelity Surface Splatting-Based 3D Reconstruction from Multi-View Images
Authors:
Nandhana Sunil,
Abhirami R Iyer,
Avirup Mandal
Abstract:
Multi-view mesh reconstruction remains a core challenge in computer graphics and vision, especially for recovering high-frequency geometry from sparse observations. Recent methods such as 3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF) rely on post-processing for mesh extraction, thereby limiting joint optimization of geometry and appearance. Implicit Moving Least Squares (IMLS) ins…
▽ More
Multi-view mesh reconstruction remains a core challenge in computer graphics and vision, especially for recovering high-frequency geometry from sparse observations. Recent methods such as 3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF) rely on post-processing for mesh extraction, thereby limiting joint optimization of geometry and appearance. Implicit Moving Least Squares (IMLS) instead enables direct conversion of point clouds into signed distance and texture fields, supporting end-to-end reconstruction and rendering. However, existing IMLS formulations use exponential kernels that struggle with high-frequency detail. We introduce a compact polynomial kernel with local support and greater flexibility, allowing better control over frequency content and improved geometric fidelity. To further enhance fine details, we incorporate stochastic regularization with Laplacian filtering. Together, these improve the preservation of high-frequency structure while maintaining stable optimization. Experiments show state-of-the-art performance in both surface reconstruction and rendering, yielding more accurate geometry and sharper visuals from multi-view data.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.
-
Tumor-aware augmentation with task-guided attention analysis improves rectal cancer segmentation from magnetic resonance images
Authors:
Aneesh Rangnekar,
Joao Miranda,
Natally Horvat,
Stephanie Chahwan,
Samir Alrayess,
Aditya Apte,
Aditi Iyer,
Eve LoCastro,
Revathi Ravella,
Marc J Gollub,
Iva Petkovska,
Jesse Joshua Smith,
Paul Romesser,
Julio Garcia-Aguilar,
Harini Veeraraghavan,
Joseph O. Deasy
Abstract:
Although self-supervised pretraining is expected to learn broadly transferable representations, its effectiveness across imaging modalities substantially different from the pretraining domain, and on complex tumor-segmentation tasks, remains understudied. Evaluating CT-pretrained transformers on MRI rectal cancer segmentation, we identified two interacting failure modes in CT-to-MRI transfer: (a)…
▽ More
Although self-supervised pretraining is expected to learn broadly transferable representations, its effectiveness across imaging modalities substantially different from the pretraining domain, and on complex tumor-segmentation tasks, remains understudied. Evaluating CT-pretrained transformers on MRI rectal cancer segmentation, we identified two interacting failure modes in CT-to-MRI transfer: (a) inefficient token usage caused by zero-padding to match pretrained input dimensions, and (b) ineffective feature adaptation. We investigated these vulnerabilities using two primary CT-pretrained hierarchical shifted-window transformer backbones, SMIT and Swin UNETR, together with VoCo as a large-scale-pretrained supporting benchmark; these models differ in pretraining objectives and datasets. Mechanistic analysis leveraged an attention dilution index (ADI), an entropy-based metric quantifying attention diverted toward uninformative padding tokens, and centered kernel alignment (CKA) to measure feature reuse during MRI adaptation. ADI increased with zero-padding, while high feature reuse did not necessarily translate to improved downstream accuracy. To mitigate these issues, we introduced two interventions: a tumor-aware augmentation strategy to expand tumor appearance heterogeneity coverage, and an anisotropic cropping strategy to restore token efficiency. Fine-tuning with these strategies on identical rectal MRI datasets yielded detection rates of 91.1% (225/247) and 88.7% (219/247) for the primary SMIT and Swin UNETR backbones, with the supporting VoCo benchmark reaching 90.3% (223/247), demonstrating significantly improved robustness under CT-to-MRI transfer. This study is among the first to examine when pretrained transformers fail to transfer across imaging modalities and demonstrates how targeted mitigation strategies can systematically overcome cross-modality transfer limitations.
△ Less
Submitted 28 June, 2026; v1 submitted 6 May, 2026;
originally announced May 2026.
-
The dark and featureless surface of rocky exoplanet LHS 3844 b from JWST mid-infrared spectroscopy
Authors:
Sebastian Zieba,
Laura Kreidberg,
Brandon P. Coy,
Aaron Bello-Arufe,
Kimberly Paragas,
Xintong Lyu,
Renyu Hu,
Aishwarya Iyer,
Edwin S. Kite,
Daniel D. B. Koll,
Kay Wohlfarth,
Emerson Whittaker,
Heather Knutson,
Robin Wordsworth,
Caroline Morley,
Laura Schaefer
Abstract:
JWST has opened a new era in the study of rocky exoplanets, enabling direct characterization of their surfaces with mid-infrared spectroscopy. Different types of rock have distinct spectral features that are diagnostic of the chemical composition and other physical properties like surface texture. Measurements of these features can provide valuable clues about a planet's geologic history and inter…
▽ More
JWST has opened a new era in the study of rocky exoplanets, enabling direct characterization of their surfaces with mid-infrared spectroscopy. Different types of rock have distinct spectral features that are diagnostic of the chemical composition and other physical properties like surface texture. Measurements of these features can provide valuable clues about a planet's geologic history and interior processes. Here we report a JWST 5-12 micron thermal emission spectrum for the rocky exoplanet LHS 3844 b. It is best matched by a dark, low-silica surface, such as basalt or other olivine-rich materials. The spectrum rules out fresh powder surfaces; however, space weathering can darken the powders and make them more consistent with the data. The data also disfavor trace concentrations of CO$_2$ or SO$_2$ gas (with 5-sigma and 3-sigma upper limits of 100 mbar and 10 microbar, respectively). Taken together, these results are well fit by an old, space-weathered surface with no evidence of accumulated volcanic gases.
△ Less
Submitted 30 April, 2026;
originally announced May 2026.
-
Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting
Authors:
Avinash Paliwal,
Adithya Iyer,
Shivin Yadav,
Muhammad Ali Afridi,
Midhun Harikumar
Abstract:
Precise camera control for reshooting dynamic videos is bottlenecked by the severe scarcity of paired multi-view data for non-rigid scenes. We overcome this limitation with a highly scalable self-supervised framework capable of leveraging internet-scale monocular videos. Our core contribution is the generation of pseudo multi-view training triplets, consisting of a source video, a geometric anchor…
▽ More
Precise camera control for reshooting dynamic videos is bottlenecked by the severe scarcity of paired multi-view data for non-rigid scenes. We overcome this limitation with a highly scalable self-supervised framework capable of leveraging internet-scale monocular videos. Our core contribution is the generation of pseudo multi-view training triplets, consisting of a source video, a geometric anchor, and a target video. We achieve this by extracting distinct smooth random-walk crop trajectories from a single input video to serve as the source and target views. The anchor is synthetically generated by forward-warping the first frame of the source with a dense tracking field, which effectively simulates the distorted point-cloud inputs expected at inference. Because our independent cropping strategy introduces spatial misalignment and artificial occlusions, the model cannot simply copy information from the current source frame. Instead, it is forced to implicitly learn 4D spatiotemporal structures by actively routing and re-projecting missing high-fidelity textures across distinct times and viewpoints from the source video to reconstruct the target. At inference, our minimally adapted diffusion transformer utilizes a 4D point-cloud derived anchor to achieve state-of-the-art temporal consistency, robust camera control, and high-fidelity novel view synthesis on complex dynamic scenes.
△ Less
Submitted 24 April, 2026; v1 submitted 23 April, 2026;
originally announced April 2026.
-
Using spatiotemporal Born rule for testing macroscopic realism: some applications to the pseudo-density matrices and nonclassical temporal correlations
Authors:
Naim Elias Comar,
Lucas C. Céleri,
Mia Stamatova,
Vlatko Vedral,
Aditya Varna Iyer,
Rafael Chaves
Abstract:
We show that, given an evolving quantum system and the quasiprobability distribution generated by the spatiotemporal generalization of the Born rule in pseudo density-matrices (PDMs), this distribution deviates from the sequential measurements probability distribution, given by the Lüders von-Neumann distribution, if and only if the non-signaling in time (NSIT) is violated; equivalently, if and on…
▽ More
We show that, given an evolving quantum system and the quasiprobability distribution generated by the spatiotemporal generalization of the Born rule in pseudo density-matrices (PDMs), this distribution deviates from the sequential measurements probability distribution, given by the Lüders von-Neumann distribution, if and only if the non-signaling in time (NSIT) is violated; equivalently, if and only if the macroscopic realism (MR) is violated. Furthermore, we propose a definition of temporal entanglement according to the structure of the PDMs that is analogous to the definition of spatial entanglement in density matrices, showing that temporal entanglement is necessary for the violation of temporal Bell inequalities and the violation of MR. We employ our results to study the relationship between the negativity of the PDM, temporal entanglement, violation of temporal Bell inequalities, and MR.
△ Less
Submitted 23 March, 2026;
originally announced March 2026.
-
ModTrack: Sensor-Agnostic Multi-View Tracking via Identity-Informed PHD Filtering with Covariance Propagation
Authors:
Aditya Iyer,
Jack Roberts,
Nora Ayanian
Abstract:
Multi-View Multi-Object Tracking (MV-MOT) aims to localize and maintain consistent identities of objects observed by multiple sensors. This task is challenging, as viewpoint changes and occlusion disrupt identity consistency across views and time. Recent end-to-end approaches address this by jointly learning 2D Bird's Eye View (BEV) representations and identity associations, achieving high trackin…
▽ More
Multi-View Multi-Object Tracking (MV-MOT) aims to localize and maintain consistent identities of objects observed by multiple sensors. This task is challenging, as viewpoint changes and occlusion disrupt identity consistency across views and time. Recent end-to-end approaches address this by jointly learning 2D Bird's Eye View (BEV) representations and identity associations, achieving high tracking accuracy. However, these methods offer no principled uncertainty accounting and remain tightly coupled to their training configuration, limiting generalization across sensor layouts, modalities, or datasets without retraining. We propose ModTrack, a modular MV-MOT system that matches end-to-end performance while providing cross-modal, sensor-agnostic generalization and traceable uncertainty. ModTrack confines learning methods to just the \textit{Detection and Feature Extraction} stage of the MV-MOT pipeline, performing all fusion, association, and tracking with closed-form analytical methods. Our design reduces each sensor's output to calibrated position-covariance pairs $(\mathbf{z}, R)$; cross-view clustering and precision-weighted fusion then yield unified estimates $(\hat{\mathbf{z}}, \hat{R})$ for identity assignment and temporal tracking. A feedback-coupled, identity-informed Gaussian Mixture Probability Hypothesis Density (GM-PHD) filter with HMM motion modes uses these fused estimates to maintain identities under missed detections and heavy occlusion. ModTrack achieves 95.5 IDF1 and 91.4 MOTA on \textit{WildTrack}, surpassing all prior modular methods by over 21 points and rivaling the state-of-the-art end-to-end methods while providing deployment flexibility they cannot. Specifically, the same tracker core transfers unchanged to \textit{MultiviewX} and \textit{RadarScenes}, with only perception-module replacement required to extend to new domains and sensor modalities.
△ Less
Submitted 26 March, 2026; v1 submitted 16 March, 2026;
originally announced March 2026.
-
Resolving Interference (RI): Disentangling Models for Improved Model Merging
Authors:
Pratik Ramesh,
George Stoica,
Arun Iyer,
Leshem Choshen,
Judy Hoffman
Abstract:
Model merging has shown that multitask models can be created by directly combining the parameters of different models that are each specialized on tasks of interest. However, models trained independently on distinct tasks often exhibit interference that degrades the merged model's performance. To solve this problem, we formally define the notion of Cross-Task Interference as the drift in the repre…
▽ More
Model merging has shown that multitask models can be created by directly combining the parameters of different models that are each specialized on tasks of interest. However, models trained independently on distinct tasks often exhibit interference that degrades the merged model's performance. To solve this problem, we formally define the notion of Cross-Task Interference as the drift in the representation of the merged model relative to its constituent models. Reducing cross-task interference is key to improving merging performance. To address this issue, we propose our method, Resolving Interference (RI), a light-weight adaptation framework which disentangles expert models to be functionally orthogonal to the space of other tasks, thereby reducing cross-task interference. RI does this whilst using only unlabeled auxiliary data as input (i.e., no task-data is needed), allowing it to be applied in data-scarce scenarios. RI consistently improves the performance of state-of-the-art merging methods by up to 3.8% and generalization to unseen domains by up to 2.3%. We also find RI to be robust to the source of auxiliary input while being significantly less sensitive to tuning of merging hyperparameters. Our codebase is available at: https://github.com/pramesh39/resolving_interference
△ Less
Submitted 13 March, 2026;
originally announced March 2026.
-
TOI-4616 b: a benchmark Earth-sized planet transiting a nearby M4 dwarf
Authors:
F. Zong Lang,
B. O. Demory,
Y. Gomez Maqueo Chew,
Y. Schmid,
M. Timmermans,
F. J. Pozuelos,
M. Gillon,
Artem Y. Burdanov,
Benjamin V. Rackham,
Didier Queloz,
Keivan G. Stassun,
Khalid Barkaoui,
Amaury Triaud,
Julien de Wit,
S. Zuniga-Fernandez,
A. J. Burgasser,
Elsa Ducrot,
Madison G. Scott,
D. Sebastian,
A. Soubkiou,
M. Lendl,
I. Plauchu-Frayn,
U. Schroffenegger,
Erik Meier V.,
P. Pedersen
, et al. (22 additional authors not shown)
Abstract:
Rocky exoplanets are particularly abundant around M-type stars. Their small radii and low luminosities provide favourable conditions for detecting transiting terrestrial planets and probing their atmospheric properties.
We report the discovery and statistical validation of TOI-4616 b, an Earth-sized planet transiting a nearby mid-M dwarf observed by the Transiting Exoplanet Survey Satellite (TES…
▽ More
Rocky exoplanets are particularly abundant around M-type stars. Their small radii and low luminosities provide favourable conditions for detecting transiting terrestrial planets and probing their atmospheric properties.
We report the discovery and statistical validation of TOI-4616 b, an Earth-sized planet transiting a nearby mid-M dwarf observed by the Transiting Exoplanet Survey Satellite (TESS). We confirm the planetary nature of the signal and determine the system parameters by combining TESS photometry with ground-based multi-band transit observations, high-resolution imaging, and optical and near-infrared spectroscopy.
The host star lies at a distance of 28.10 +(-) 0.07 pc and has a radius of 0.1889 +(-)0.0096 solar radii, a mass of 0.1881 +(-) 0.0094 solar masses, and an effective temperature of 3150 +(-) 75 K. TOI-4616 b has a radius of 1.22 Earth radii and an orbital period of 1.55 days. The planet receives an incident flux of approximately 40 times that of Earth, corresponding to an equilibrium temperature of about 525 K. This places TOI-4616 b in a regime intermediate between Earth-sized planets orbiting early M dwarfs and those around ultra-cool hosts.
Statistical validation with the TRICERATOPS framework, supported by high-resolution imaging and chromatic transit constraints, yields a false-positive probability of 0.0135, below the recommended validation threshold of 0.015, confirming TOI-4616 b as a validated planet. Owing to its proximity to Earth, well-constrained stellar properties, and extensive multi-band follow-up, TOI-4616 b constitutes a valuable benchmark system for comparative studies of terrestrial planets around mid-M dwarfs and for future atmospheric investigations.
△ Less
Submitted 11 March, 2026;
originally announced March 2026.
-
Chow-Liu Ordering for Long-Context Reasoning in Chain-of-Agents
Authors:
Naman Gupta,
Vaibhav Singh,
Arun Iyer,
Kirankumar Shiragur,
Pratham Grover,
Ramakrishna B. Bairi,
Ritabrata Maiti,
Sankarshan Damle,
Shachee Mishra Gupta,
Rishikesh Maurya,
Vageesh D. C
Abstract:
Sequential multi-agent reasoning frameworks such as Chain-of-Agents (CoA) handle long-context queries by decomposing inputs into chunks and processing them sequentially using LLM-based worker agents that read from and update a bounded shared memory. From a probabilistic perspective, CoA aims to approximate the conditional distribution corresponding to a model capable of jointly reasoning over the…
▽ More
Sequential multi-agent reasoning frameworks such as Chain-of-Agents (CoA) handle long-context queries by decomposing inputs into chunks and processing them sequentially using LLM-based worker agents that read from and update a bounded shared memory. From a probabilistic perspective, CoA aims to approximate the conditional distribution corresponding to a model capable of jointly reasoning over the entire long context. CoA achieves this through a latent-state factorization in which only bounded summaries of previously processed evidence are passed between agents. The resulting bounded-memory approximation introduces a lossy information bottleneck, making the final evidence state inherently dependent on the order in which chunks are processed.
In this work, we study the problem of chunk ordering for long-context reasoning. We use the well-known Chow-Liu trees to learn a dependency structure that prioritizes strongly related chunks. Empirically, we show that a breadth-first traversal of the resulting tree yields chunk orderings that reduce information loss across agents and consistently outperform both default document-chunk ordering and semantic score-based ordering in answer relevance and exact-match accuracy across three long-context benchmarks.
△ Less
Submitted 10 March, 2026;
originally announced March 2026.
-
NASA's Pandora SmallSat Mission: Simulating the Impact of Stellar Photospheric Heterogeneity and Its Correction
Authors:
Benjamin V. Rackham,
Aishwarya R. Iyer,
Dániel Apai,
Peter McGill,
Yoav Rotman,
Knicole D. Colón,
Brett M. Morris,
Emily A. Gilbert,
Elisa V. Quintana,
Jessie L. Dotson,
Thomas Barclay,
Pete Supsinskas,
Jordan Karburn,
Christina Hedges,
Jason F. Rowe,
David R. Ciardi,
Jessie L. Christiansen,
Trevor O. Foote,
Thomas P. Greene,
Kelsey Hoffman,
Rae Holcomb,
Aurora Y. Kesseli,
Veselin B. Kostov,
Nikole K. Lewis,
James P. Mason
, et al. (6 additional authors not shown)
Abstract:
Stellar photospheric heterogeneity is a dominant astrophysical systematic impacting exoplanet transmission spectroscopy. NASA's Pandora SmallSat Mission is designed to address this challenge through contemporaneous visible photometry and NIR spectroscopy of exoplanet host stars. Here we present an end-to-end simulation study quantifying Pandora's ability to infer stellar photospheric properties an…
▽ More
Stellar photospheric heterogeneity is a dominant astrophysical systematic impacting exoplanet transmission spectroscopy. NASA's Pandora SmallSat Mission is designed to address this challenge through contemporaneous visible photometry and NIR spectroscopy of exoplanet host stars. Here we present an end-to-end simulation study quantifying Pandora's ability to infer stellar photospheric properties and correct stellar contamination using out-of-transit observations. We construct eight representative stellar activity scenarios and generate 160 simulated Pandora datasets, incorporating time-dependent stellar spectra, instrument response, and noise. Given accurate models, Bayesian retrievals of Pandora spectrophotometry recover photospheric temperatures with typical uncertainties of ${\approx}30$ K, with no significant bias. Models with two spectral components (i.e., quiescent photosphere and spots) are strongly favored in 95% of cases; one-component models are preferred when true spot filling factors fall below a detection threshold of ${\approx}0.3$%. We propagate the true and inferred stellar parameters to compute true, inferred, and residual contamination signals under physically motivated spot geometries. For simple spot distributions, contamination signals of $10^2{-}10^3$ ppm are reduced to ${\lesssim}10$ ppm, well below Pandora's expected transmission spectroscopy precision (30$-$100 ppm). For more complex spot distributions, geometric degeneracies limit deterministic corrections, leaving residual contamination at the $10^3$ ppm level that must be mitigated using additional constraints, such as spot-crossing events and joint stellar-planetary retrievals of transmission spectra. These results define regimes in which stellar contamination can be corrected from stellar observations alone and show how Pandora stellar observations can identify cases where additional information is required.
△ Less
Submitted 29 April, 2026; v1 submitted 4 March, 2026;
originally announced March 2026.
-
NASA's Pandora SmallSat Mission: Simulated Modeling and Retrieval of Near-Infrared Exoplanet Transmission Spectra
Authors:
Yoav Rotman,
Peter McGill,
Luis Welbanks,
Benjamin V. Rackham,
Aishwarya Iyer,
Daniel Apai,
Michael R. Line,
Elisa V. Quintana,
Jessie L. Dotson,
Knicole D. Colon,
Thomas Barclay,
Christina Hedges,
Jason F. Rowe,
Emily A. Gilbert,
Brett M. Morris,
Jessie L. Christiansen,
Trevor O. Foote,
Aylin Garcia Soto,
Thomas P. Greene,
Kelsey Hoffman,
Benjamin J. Hord,
Aurora Y. Kesseli,
Veselin B. Kostov,
Megan Weiner Mansfield,
Lindsey S. Wiser
Abstract:
Pandora is a SmallSat mission dedicated to understanding exoplanets and their host stars by disentangling the impact of stellar heterogeneity on exoplanet transmission spectra. Selected as a NASA Astrophysics Pioneers mission in 2021, Pandora will provide simultaneous long-term visible photometric monitoring (0.4--0.7 $μ$m) and low-resolution near-infrared (NIR) spectroscopy (0.9--1.6 $μ$m) of tra…
▽ More
Pandora is a SmallSat mission dedicated to understanding exoplanets and their host stars by disentangling the impact of stellar heterogeneity on exoplanet transmission spectra. Selected as a NASA Astrophysics Pioneers mission in 2021, Pandora will provide simultaneous long-term visible photometric monitoring (0.4--0.7 $μ$m) and low-resolution near-infrared (NIR) spectroscopy (0.9--1.6 $μ$m) of transiting systems for the purposes of monitoring host star variability and characterizing exoplanetary atmospheres. Pandora's year-long prime mission from 2026 to 2027 coincides with the middle of a decade defined by targeted efforts for atmospheric characterization of exoplanets, offering a key opportunity to leverage this new resource to maximize science with JWST and other observatories. Here we investigate Pandora's anticipated performance for the general exoplanet population accessible to transit spectroscopy, from hot Jupiters to temperate sub-Neptunes. By modeling the atmospheres of five test cases broadly consistent with the bulk properties of HD~209458~b, HD~189733~b, WASP-80~b, HAT-P-18~b, and K2-18~b, we find that Pandora may provide abundance constraints as precise as $\sim$1.0\,dex for main atmospheric absorbers such as H$_2$O and CH$_4$. Then, we explore the synergies between Pandora and JWST. Our results suggest that targets with JWST data in the near-infrared can benefit from the addition of Pandora observations and result in more reliable abundance estimates than with JWST data alone. Moreover, Pandora can serve the community by providing precursory observations of targets of interest for JWST atmospheric characterization. We conclude by outlining strategies for the use of Pandora as a standalone observatory and in synergy with JWST.
△ Less
Submitted 4 March, 2026;
originally announced March 2026.
-
CHAI: CacHe Attention Inference for text2video
Authors:
Joel Mathew Cherian,
Ashutosh Muralidhara Bharadwaj,
Vima Gupta,
Anand Padmanabha Iyer
Abstract:
Text-to-video diffusion models deliver impressive results but remain slow because of the sequential denoising of 3D latents. Existing approaches to speed up inference either require expensive model retraining or use heuristic-based step skipping, which struggles to maintain video quality as the number of denoising steps decreases. Our work, CHAI, aims to use cross-inference caching to reduce laten…
▽ More
Text-to-video diffusion models deliver impressive results but remain slow because of the sequential denoising of 3D latents. Existing approaches to speed up inference either require expensive model retraining or use heuristic-based step skipping, which struggles to maintain video quality as the number of denoising steps decreases. Our work, CHAI, aims to use cross-inference caching to reduce latency while maintaining video quality. We introduce Cache Attention as an effective method for attending to shared objects/scenes across cross-inference latents. This selective attention mechanism enables effective reuse of cached latents across semantically related prompts, yielding high cache hit rates. We show that it is possible to generate high-quality videos using Cache Attention with as few as 8 denoising steps. When integrated into the overall system, CHAI is 1.65x - 3.35x faster than baseline OpenSora 1.2 while maintaining video quality.
△ Less
Submitted 17 February, 2026;
originally announced February 2026.
-
SAFuzz: Semantic-Guided Adaptive Fuzzing for LLM-Generated Code
Authors:
Ziyi Yang,
Kalit Inani,
Keshav Kabra,
Vima Gupta,
Anand Padmanabha Iyer
Abstract:
While AI-coding assistants accelerate software development, current testing frameworks struggle to keep pace with the resulting volume of AI-generated code. Traditional fuzzing techniques often allocate resources uniformly and lack semantic awareness of algorithmic vulnerability patterns, leading to inefficient resource usage and missed vulnerabilities. To address these limitations, we present a h…
▽ More
While AI-coding assistants accelerate software development, current testing frameworks struggle to keep pace with the resulting volume of AI-generated code. Traditional fuzzing techniques often allocate resources uniformly and lack semantic awareness of algorithmic vulnerability patterns, leading to inefficient resource usage and missed vulnerabilities. To address these limitations, we present a hybrid testing framework that leverages LLM-guided adaptive fuzzing to detect algorithmic vulnerabilities efficiently. Our system SAFuzz integrates prompt-based behavioral diversification, harness generation with problem-specific oracles, and an LLM-based predictor to enable adaptive resource allocation and dynamic early stopping. Evaluating SAFuzz on CSES algorithmic problems, we improve vulnerability discrimination precision from 77.9% to 85.7% and achieve a 1.71x reduction in time cost compared to SOTA GreenFuzz while maintaining comparable recall. We further observe that combining our approach with existing unit test generation methods yields complementary gains, increasing the bug detection recall from 67.3% to 79.5%.
△ Less
Submitted 10 February, 2026;
originally announced February 2026.
-
SAGE: Agentic Framework for Interpretable and Clinically Translatable Computational Pathology Biomarker Discovery
Authors:
Sahar Almahfouz Nasser,
Juan Francisco Pesantez Borja,
Jincheng Liu,
Sandeep Manandhar,
Shikhar Shiromani,
Mohammad Tanvir Hasan,
Zenghan Wang,
Suman Ghosh,
Jinchu Li,
Xuejian Xu,
Aniket Ramkrishnan Iyer,
Naoto Tokuyama,
Twisha Shah,
Tilak Pathak,
Soundharya Kumaresan,
Yohei Abe,
Himanshu Maurya,
Anant Madabhushi
Abstract:
Engineered image-based biomarkers offer a clinically interpretable alternative to black-box AI in computational pathology, yet their discovery remains largely intuition-driven, guided by fragmented literature rather than rigorous biological validation. We introduce SAGE (Structured Agentic system for hypothesis Generation and Evaluation), a multi-agent framework that grounds biomarker discovery in…
▽ More
Engineered image-based biomarkers offer a clinically interpretable alternative to black-box AI in computational pathology, yet their discovery remains largely intuition-driven, guided by fragmented literature rather than rigorous biological validation. We introduce SAGE (Structured Agentic system for hypothesis Generation and Evaluation), a multi-agent framework that grounds biomarker discovery in biological evidence through three mechanisms: (i) knowledge-graph-anchored hypothesis generation via multi-path ontological reasoning, (ii) a debate-based multi-agent novelty assessment that stress-tests candidate biomarkers against existing literature, and (iii) an end-to-end automated validation pipeline that translates hypotheses directly into executable analyses on multimodal pathology datasets. Together, these components shift biomarker discovery from an intuition-driven, literature-browsing exercise into a structured, traceable reasoning process that clinicians and researchers can inspect, trust, and build upon.
△ Less
Submitted 10 May, 2026; v1 submitted 31 January, 2026;
originally announced February 2026.
-
Fauna Sprout: A lightweight, approachable, developer-ready humanoid robot
Authors:
Fauna Robotics,
:,
Diego Aldarondo,
Ana Pervan,
Daniel Corbalan,
Dave Petrillo,
Bolun Dai,
Aadhithya Iyer,
Nina Mortensen,
Erik Pearson,
Sridhar Pandian Arunachalam,
Emma Reznick,
David Weis,
Jacob Davison,
Samuel Patterson,
Tess Carella,
Michael Suguitan,
David Ye,
Oswaldo Ferro,
Nilesh Suriyarachchi,
Spencer Ling,
Erik Su,
Daniel Giebisch,
Peter Traver,
Sam Fonseca
, et al. (26 additional authors not shown)
Abstract:
Recent advances in learned control, large-scale simulation, and generative models have accelerated progress toward general-purpose robotic controllers, yet the field still lacks platforms suitable for safe, expressive, long-term deployment in human environments. Most existing humanoids are either closed industrial systems or academic prototypes that are difficult to deploy and operate around peopl…
▽ More
Recent advances in learned control, large-scale simulation, and generative models have accelerated progress toward general-purpose robotic controllers, yet the field still lacks platforms suitable for safe, expressive, long-term deployment in human environments. Most existing humanoids are either closed industrial systems or academic prototypes that are difficult to deploy and operate around people, limiting progress in robotics. We introduce Sprout, a developer platform designed to address these limitations through an emphasis on safety, expressivity, and developer accessibility. Sprout adopts a lightweight form factor with compliant control, limited joint torques, and soft exteriors to support safe operation in shared human spaces. The platform integrates whole-body control, manipulation with integrated grippers, and virtual-reality-based teleoperation within a unified hardware-software stack. An expressive head further enables social interaction -- a domain that remains underexplored on most utilitarian humanoids. By lowering physical and technical barriers to deployment, Sprout expands access to capable humanoid platforms and provides a practical basis for developing embodied intelligence in real human environments.
△ Less
Submitted 26 January, 2026;
originally announced January 2026.
-
Probing scalar and pseudoscalar new physics using rare kaon decays
Authors:
G. D'Ambrosio,
A. M. Iyer,
F. Mahmoudi,
S. Neshatpour
Abstract:
Rare kaon decays provide sensitive tests of new physics. In this work, we focus on scalar and pseudoscalar operators, analysing the $K\to π\ell^+\ell^-$ and $K\to \ell^+\ell^-$ decays. We highlight the complementary role of different modes: $K^+\toπ^+\ell^+\ell^-$, in particular the forward-backward asymmetry in the muon channel as a clean probe of scalar effects, the stringent constraints from…
▽ More
Rare kaon decays provide sensitive tests of new physics. In this work, we focus on scalar and pseudoscalar operators, analysing the $K\to π\ell^+\ell^-$ and $K\to \ell^+\ell^-$ decays. We highlight the complementary role of different modes: $K^+\toπ^+\ell^+\ell^-$, in particular the forward-backward asymmetry in the muon channel as a clean probe of scalar effects, the stringent constraints from $K_L\to μ^+μ^-$, and the discovery potential of future measurements of $K_S\to μ^+μ^-$ and $K_L\to π^0 \ell^+\ell^-$. The interplay between charged and neutral modes underscores the complementarity of NA62, the LHCb upgrade, and KOTO-II.
△ Less
Submitted 18 December, 2025;
originally announced December 2025.
-
Solar-cycle variations in meridional flows and rotational shear within the Sun's near-surface shear layer
Authors:
Anisha Sen,
S. P. Rajaguru,
Abhinav Govindan Iyer,
Ruizhu Chen,
Junwei Zhao,
Shukur Kholikov
Abstract:
Using solar-cycle long helioseismic measurements of meridional and zonal flows in the near-surface shear layer (NSSL) of the Sun, we study their spatio-temporal variations and connections to active regions. We find that near-surface inflows towards active latitudes are part of a local circulation with an outflow away from them at depths around 0.97 R, which is also the location where the deviation…
▽ More
Using solar-cycle long helioseismic measurements of meridional and zonal flows in the near-surface shear layer (NSSL) of the Sun, we study their spatio-temporal variations and connections to active regions. We find that near-surface inflows towards active latitudes are part of a local circulation with an outflow away from them at depths around 0.97 R, which is also the location where the deviations in the radial gradient of rotation change sign. These results, together with opposite-signed changes over latitude and depth in the above quantities observed during the solar minimum period, point to the action of the Coriolis force on large-scale flows as the primary cause of changes in the rotation gradient within the NSSL. We also find that such Coriolis force-mediated changes in near-surface flows towards active latitudes only marginally change the amplitude of zonal flow and hence are not likely to be its driving force. Our measurements typically achieve a high signal-to-noise ratio ($>$5$σ$) for near-surface flows but can drop to 3$σ$ near the base (0.95 R) of the NSSL. Close agreements between the depth profiles of changes in rotation gradient and in meridional flows measured from quite different global and local helioseismic techniques, respectively, show that the results are not dependent on the analysis techniques.
△ Less
Submitted 13 December, 2025;
originally announced December 2025.
-
Formal Specification for Fast ACS: Low-Latency File-Based Ordered Message Delivery at Scale
Authors:
Sushant Kumar Gupta,
Anil Raghunath Iyer,
Chang Yu,
Neel Bagora,
Olivier Pomerleau,
Vivek Kumar,
Prunthaban Kanthakumar
Abstract:
Low-latency message delivery is crucial for real-time systems. Data originating from a producer must be delivered to consumers, potentially distributed in clusters across metropolitan and continental boundaries. With the growing scale of computing, there can be several thousand consumers of the data. Such systems require a robust messaging system capable of transmitting messages containing data ac…
▽ More
Low-latency message delivery is crucial for real-time systems. Data originating from a producer must be delivered to consumers, potentially distributed in clusters across metropolitan and continental boundaries. With the growing scale of computing, there can be several thousand consumers of the data. Such systems require a robust messaging system capable of transmitting messages containing data across clusters and efficiently delivering them to consumers. The system must offer guarantees like ordering and at-least-once delivery while avoiding overload on consumers, allowing them to consume messages at their own pace.
This paper presents the design of Fast ACS (an abbreviation for Ads Copy Service), a file-based ordered message delivery system that leverages a combination of two-sided (inter-cluster) and one-sided (intra-cluster) communication primitives - namely, Remote Procedure Call and Remote Memory Access, respectively - to deliver messages. The system has been successfully deployed to dozens of production clusters and scales to accommodate several thousand consumers within each cluster, which amounts to Tbps-scale intra-cluster consumer traffic at peak. Notably, Fast ACS delivers messages to consumers across the globe within a few seconds or even sub-seconds (p99) based on the message volume and consumer scale, at a low resource cost.
△ Less
Submitted 21 November, 2025;
originally announced December 2025.
-
The SPHINX M dwarf Spectral Grid. II. New Model Atmospheres and Spectra to Derive Fundamental Properties of mid-to-late type M-dwarfs
Authors:
Aishwarya R. Iyer,
Michael R. Line,
Philip S. Muirhead,
Jonathan J. Fortney,
Jacqueline K. Faherty
Abstract:
M-dwarfs are the most dominant stars in the Galaxy. Their interiors and atmospheres exhibit complex processes including dust condensation, convective feedback, and magnetic activity-driven heterogeneity. Standard stellar characterization methods often struggle to capture these coupled effects. Part I of this series introduced SPHINX I, a validated grid of self-consistent radiative-convective model…
▽ More
M-dwarfs are the most dominant stars in the Galaxy. Their interiors and atmospheres exhibit complex processes including dust condensation, convective feedback, and magnetic activity-driven heterogeneity. Standard stellar characterization methods often struggle to capture these coupled effects. Part I of this series introduced SPHINX I, a validated grid of self-consistent radiative-convective model atmospheres and spectra for M-dwarfs with up-to-date molecular opacities suitable for early-to-mid M-dwarfs. Here, we present SPHINX II, which extends the model grid to cover mid-to-late type M-dwarfs, including both gray and physically motivated condensate cloud treatments and shorter convective mixing lengths. We validate SPHINX II using 39 benchmark FGK+M binary systems observed with SpeX IRTF (Mann et al. 2014) and apply it to 32 mid-to-late-type M-dwarfs from the SpeX Prism Library. SPHINX II yields improved fits that are statistically consistent with empirical benchmarks, achieving precisions of 0.078 dex in metallicity and 0.13 dex in C/O. Across the model grid, condensate cloud mass peaks between 2100-2400 K, decreasing sharply toward both cooler and hotter temperatures. We find the onset of the cloud-free regime around 2900 K, and below 2100 K, we see formation of deep/buried clouds. As a case study, we also model Trappist-1 and show that even mass-limited silicate grains subtly modify its emergent spectrum, suppressing near-infrared flux and reddening the mid-infrared slope via shallow cloud formation near 1e-2 bar. In sum, SPHINX II provides an improved framework for constraining the fundamental properties of mid-to-late M-dwarfs.
△ Less
Submitted 3 December, 2025; v1 submitted 1 December, 2025;
originally announced December 2025.
-
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
Authors:
Eun Chang,
Zhuangqun Huang,
Yiwei Liao,
Sagar Ravi Bhavsar,
Amogh Param,
Tammy Stark,
Adel Ahmadyan,
Xiao Yang,
Jiaqi Wang,
Ahsan Abdullah,
Giang Nguyen,
Akil Iyer,
David Hall,
Elissa Li,
Shane Moon,
Nicolas Scheffer,
Kirmani Ahmed,
Babak Damavandi,
Rakesh Wanga,
Anuj Kumar,
Rohit Patel,
Xin Luna Dong
Abstract:
We introduce WearVQA, the first benchmark specifically designed to evaluate the Visual Question Answering (VQA) capabilities of multi-model AI assistant on wearable devices like smart glasses. Unlike prior benchmarks that focus on high-quality, third-person imagery, WearVQA reflects the unique challenges of ego-centric interaction-where visual inputs may be occluded, poorly lit, unzoomed, or blurr…
▽ More
We introduce WearVQA, the first benchmark specifically designed to evaluate the Visual Question Answering (VQA) capabilities of multi-model AI assistant on wearable devices like smart glasses. Unlike prior benchmarks that focus on high-quality, third-person imagery, WearVQA reflects the unique challenges of ego-centric interaction-where visual inputs may be occluded, poorly lit, unzoomed, or blurry, and questions are grounded in realistic wearable use cases. The benchmark comprises 2,520 carefully curated image-question-answer triplets, spanning 7 diverse image domains including both text-centric and general scenes, 10 cognitive task types ranging from basic recognition to various forms of reasoning, and 6 common wearables-specific image quality issues. All questions are designed to be answerable using only the visual input and common senses. WearVQA is paired with a rigorous LLM-as-a-judge evaluation framework with 96% labeling accuracy. Open-source and proprietary multi-model LLMs achieved a QA accuracy as low as 24-52% on WearVQA, with substantial drops on lower-quality images and reasoning-heavy tasks. These observations position WearVQA as a comprehensive and challenging benchmark for guiding technical advancement towards robust, real-world multi-model wearables AI systems.
△ Less
Submitted 2 December, 2025; v1 submitted 27 November, 2025;
originally announced November 2025.
-
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
Authors:
Yinwei Dai,
Zhuofu Chen,
Anand Iyer,
Ravi Netravali
Abstract:
Agentic workflows have emerged as a powerful paradigm for solving complex, multi-stage tasks, but serving them at scale is computationally expensive given the many LLM inferences that each request must pass through. Configuration selection, or the cost-aware assignment of workflow agents to specific LLMs, can reduce these costs, but existing approaches bind configuration decisions before request e…
▽ More
Agentic workflows have emerged as a powerful paradigm for solving complex, multi-stage tasks, but serving them at scale is computationally expensive given the many LLM inferences that each request must pass through. Configuration selection, or the cost-aware assignment of workflow agents to specific LLMs, can reduce these costs, but existing approaches bind configuration decisions before request execution, making them ill-suited for the heterogeneous and lengthy execution of workflows. Specifically, system loads can fluctuate rapidly and substantially during a request's lifetime, causing fixed configurations to quickly become suboptimal. We present Aragog, a system that progressively adapts a request's configuration throughout its execution to match runtime dynamics. To make this practical despite the massive space of workflow configurations, Aragog decouples the problem into two core elements -- a one-time routing step that identifies all accuracy-preserving configurations, and a cheap per-stage scheduler that selects among them using up-to-date system observations -- and introduces novel strategies to accelerate each. Across diverse workflows and model families, Aragog increases maximum serving throughput by 50.0--217.0\% and reduces median latency by 32.5--78.9\% at peak request rates, while maintaining accuracy comparable to the most expensive configurations.
△ Less
Submitted 11 December, 2025; v1 submitted 25 November, 2025;
originally announced November 2025.
-
Rethinking Kernel Program Repair: Benchmarking and Enhancing LLMs with RGym
Authors:
Kareem Shehada,
Yifan Wu,
Wyatt D. Feng,
Adithya Iyer,
Gryphon Kumfert,
Yangruibo Ding,
Zhiyun Qian
Abstract:
Large Language Models (LLMs) have revolutionized automated program repair (APR) but current benchmarks like SWE-Bench predominantly focus on userspace applications and overlook the complexities of kernel-space debugging and repair. The Linux kernel poses unique challenges due to its monolithic structure, concurrency, and low-level hardware interactions. Prior efforts such as KGym and CrashFixer ha…
▽ More
Large Language Models (LLMs) have revolutionized automated program repair (APR) but current benchmarks like SWE-Bench predominantly focus on userspace applications and overlook the complexities of kernel-space debugging and repair. The Linux kernel poses unique challenges due to its monolithic structure, concurrency, and low-level hardware interactions. Prior efforts such as KGym and CrashFixer have highlighted the difficulty of APR in this domain, reporting low success rates or relying on costly and complex pipelines and pricey cloud infrastructure. In this work, we introduce RGym, a lightweight, platform-agnostic APR evaluation framework for the Linux kernel designed to operate on local commodity hardware. Built on RGym, we propose a simple yet effective APR pipeline leveraging specialized localization techniques (e.g., call stacks and blamed commits) to overcome the unrealistic usage of oracles in KGym. We test on a filtered and verified dataset of 143 bugs. Our method achieves up to a 43.36% pass rate with GPT-5 Thinking while maintaining a cost of under $0.20 per bug. We further conduct an ablation study to analyze contributions from our proposed localization strategy, prompt structure, and model choice, and demonstrate that feedback-based retries can significantly enhance success rates.
△ Less
Submitted 19 November, 2025;
originally announced November 2025.
-
Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents
Authors:
Divyanshu Saxena,
Rishikesh Maurya,
Xiaoxuan Ou,
Gagan Somashekar,
Shachee Mishra Gupta,
Arun Iyer,
Yu Kang,
Chetan Bansal,
Aditya Akella,
Saravan Rajmohan
Abstract:
The rapid adoption of AI agents across domains has made systematic evaluation crucial for ensuring their usefulness and successful production deployment. Evaluation of AI agents typically involves using a fixed set of benchmarks and computing multiple evaluation metrics for the agent. While sufficient for simple coding tasks, these benchmarks fall short for enterprise-scale agents, where services…
▽ More
The rapid adoption of AI agents across domains has made systematic evaluation crucial for ensuring their usefulness and successful production deployment. Evaluation of AI agents typically involves using a fixed set of benchmarks and computing multiple evaluation metrics for the agent. While sufficient for simple coding tasks, these benchmarks fall short for enterprise-scale agents, where services and requirements evolve continuously and ground-truth examples are sparse. We propose a process of benchmark generation that helps evolve the benchmarks as the requirements change and perform robust evaluation of evolving AI agents. We instantiate this approach for a case study of service migration from one deployment platform to another at a large public enterprise. Our approach relies on semi-structured documents where developers express the high-level intent, and uses state-of-the-art LLMs to generate benchmarks from just a small number of such documents. Overall, this process results in a maintainable evaluation framework, enabling rapid feedback on agent performance and facilitating targeted improvements.
△ Less
Submitted 13 November, 2025;
originally announced November 2025.
-
Bounding interventional queries from generalized incomplete contingency tables
Authors:
Ivano Lodato,
Aditya V. Iyer,
Isaac Z. To
Abstract:
We introduce a method for evaluating interventional queries and Average Treatment Effects (ATEs) in the presence of generalized incomplete contingency tables (GICTs), contingency tables containing a full row of random (sampling) zeros, rendering some conditional probabilities undefined. Rather than discarding such entries or imputing missing values, we model the unknown probabilities as free param…
▽ More
We introduce a method for evaluating interventional queries and Average Treatment Effects (ATEs) in the presence of generalized incomplete contingency tables (GICTs), contingency tables containing a full row of random (sampling) zeros, rendering some conditional probabilities undefined. Rather than discarding such entries or imputing missing values, we model the unknown probabilities as free parameters and derive symbolic expressions for the queries that incorporate them. By extremizing these expressions over all values consistent with basic probability constraints and the support of all variables, we obtain sharp bounds for the query of interest under weak assumptions of small missing frequencies. These bounds provide a formal quantification of the uncertainty induced by the generalized incompleteness of the contingency table and ensure that the true value of the query will always lie within them. The framework applies independently of the missingness mechanism and offers a conservative yet rigorous approach to causal inference under random data gaps.
△ Less
Submitted 7 November, 2025;
originally announced November 2025.
-
CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
Authors:
Jiaqi Wang,
Xiao Yang,
Kai Sun,
Parth Suresh,
Sanat Sharma,
Adam Czyzewski,
Derek Andersen,
Surya Appini,
Arkav Banerjee,
Sajal Choudhary,
Shervin Ghasemlou,
Ziqiang Guan,
Akil Iyer,
Haidar Khan,
Lingkun Kong,
Roy Luo,
Tiffany Ma,
Zhen Qiao,
David Tran,
Wenfang Xu,
Skyler Yeatman,
Chen Zhou,
Gunveer Gujral,
Yinglong Xia,
Shane Moon
, et al. (16 additional authors not shown)
Abstract:
Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-Modal Retrieval-Augmented Generation (MM-RAG) plays a key role in supporting such questions, yet there is still no comprehensive benchmark for this task, especially regarding wearables scenarios. To fill this gap, we pre…
▽ More
Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-Modal Retrieval-Augmented Generation (MM-RAG) plays a key role in supporting such questions, yet there is still no comprehensive benchmark for this task, especially regarding wearables scenarios. To fill this gap, we present CRAG-MM -- a Comprehensive RAG benchmark for Multi-modal Multi-turn conversations. CRAG-MM contains a diverse set of 6.5K (image, question, answer) triplets and 2K visual-based multi-turn conversations across 13 domains, including 6.2K egocentric images designed to mimic captures from wearable devices. We carefully constructed the questions to reflect real-world scenarios and challenges, including five types of image-quality issues, six question types, varying entity popularity, differing information dynamism, and different conversation turns. We design three tasks: single-source augmentation, multi-source augmentation, and multi-turn conversations -- each paired with an associated retrieval corpus and APIs for both image-KG retrieval and webpage retrieval. Our evaluation shows that straightforward RAG approaches achieve only 32% and 43% truthfulness on CRAG-MM single- and multi-turn QA, respectively, whereas state-of-the-art industry solutions have similar quality (32%/45%), underscoring ample room for improvement. The benchmark has hosted KDD Cup 2025, attracting about 1K participants and 5K submissions, with winning solutions improving baseline performance by 28%, highlighting its early impact on advancing the field.
△ Less
Submitted 30 October, 2025;
originally announced October 2025.
-
SHUFFLESPARSE: Learned Shuffles for Structured Sparse Networks
Authors:
Abhishek Tyagi,
Arjun Iyer,
Liam Young,
William H Renninger,
Christopher Kanan,
Yuhao Zhu
Abstract:
Structured weight sparsity accelerates training and inference on modern GPUs, but it trails unstructured dynamic sparse training (DST) in accuracy especially at extreme sparsity. We pinpoint the reason for this difference in performance to a lack of expressivity: a dense layer can implement any pattern of non-zero weights, whereas structured patterns are restricted to only a small set of weight co…
▽ More
Structured weight sparsity accelerates training and inference on modern GPUs, but it trails unstructured dynamic sparse training (DST) in accuracy especially at extreme sparsity. We pinpoint the reason for this difference in performance to a lack of expressivity: a dense layer can implement any pattern of non-zero weights, whereas structured patterns are restricted to only a small set of weight configurations. We introduce SHUFFLESPARSE, a single permutation primitive that applies uniformly across DST-from-scratch and one-shot pruning, and across N:M, block etc. We close most of this gap by learning a single permutation matrix jointly with the structured weight matrix. When used on three different types of structures (block, N:M, and diagonal), SHUFFLESPARSE is able to reduce the structured-vs-unstructured accuracy gap on ViT-B16 (ImageNet-1K) and GPT-2 (WikiText-103) at 90-95% sparsity, while adding minimal inference overhead (< 8.7% inference overhead) and preserving any training acceleration the host structure provides. The same permutation formulation transfers to one-shot 2:4 pruning of pretrained LLMs, where it improves zero-shot accuracy by 4.6 points on LLaMA-2 7B. Together, these results establish learned permutations as a general tool for recovering unstructured-level accuracy from structured sparsity patterns.
△ Less
Submitted 20 July, 2026; v1 submitted 16 October, 2025;
originally announced October 2025.
-
An insight into the rare $Z\rightarrow b \bar{b}γ$ at the HL-LHC
Authors:
T. Thallapalli,
A. Alpana,
A. M. Iyer,
S. Sharma
Abstract:
Studies at the $Z$-pole have played an important role in developing our understanding of the Standard Model (SM). Continuing the explorations in this regime, we consider the possibility of the production of two $b$-quarks and a photon in proton-proton collisions at the HL-LHC. While such a final state is possible in the SM by means of the process $Z\rightarrow b\bar b$ decay with a radiated photon…
▽ More
Studies at the $Z$-pole have played an important role in developing our understanding of the Standard Model (SM). Continuing the explorations in this regime, we consider the possibility of the production of two $b$-quarks and a photon in proton-proton collisions at the HL-LHC. While such a final state is possible in the SM by means of the process $Z\rightarrow b\bar b$ decay with a radiated photon, the focus is on extracting its possible origins due to beyond Standard Model (BSM) physics. The signal topology can be broadly identified as $Z\rightarrow Φγ\rightarrow b\bar bγ$, where $ Φ$ can either be a spin-0 or spin-2 state with a mass less than that of the $Z$ boson.The analysis is characterised by two relatively low $p_T$ jets that are required to be $b$-tagged jets and an isolated photon. The study provides a quantitative framework for their identification and highlights the potential challenges associated with this final state. A range of machine learning architectures is employed to demonstrate the stability and reliability of the discriminating variables. This study highlights the importance of the low-$p_T$ objects in searches at the HL-LHC, hence the need to pay special attention to their identification and efficiencies.
△ Less
Submitted 16 October, 2025;
originally announced October 2025.
-
The Sonora Substellar Atmosphere Models VI. Red Diamondback: Extending Diamondback with SPHINX for Brown Dwarf Early Evolution
Authors:
C. Evan Davis,
Jonathan J. Fortney,
Aishwarya Iyer,
Sagnick Mukherjee,
Caroline V. Morley,
Mark S. Marley,
Michael Line,
Philip S. Muirhead
Abstract:
We extend the Sonora Diamondback brown dwarf evolution models to higher effective temperatures to treat the evolution of younger, higher mass objects. Due to an upper temperature limit of $T_\mathrm{eff}=$2400 K in the original Sonora Diamondback model grid, high mass objects ($M\geq$ 0.05 $M_\mathrm{\odot}=$ 52.4 $M_\mathrm{J}$) were limited to ages of $\gtrsim$ 100 Myr. To include the early evol…
▽ More
We extend the Sonora Diamondback brown dwarf evolution models to higher effective temperatures to treat the evolution of younger, higher mass objects. Due to an upper temperature limit of $T_\mathrm{eff}=$2400 K in the original Sonora Diamondback model grid, high mass objects ($M\geq$ 0.05 $M_\mathrm{\odot}=$ 52.4 $M_\mathrm{J}$) were limited to ages of $\gtrsim$ 100 Myr. To include the early evolution of brown dwarfs at $T_\mathrm{eff}>$ 2400 K, we use existing and new SPHINX cloud-free model atmosphere calculations of temperature structures of M-type atmospheres. These atmospheres range from $T_\mathrm{eff}$ 2000--4000 K, log($g$) 3.0--5.5, and metallicity [M/H] $-$0.5 to $+$0.5. This combination of Diamondback and SPHINX atmospheres, with a transition across $T_\mathrm{eff}$ 2000--2400 K, allows us to calculate evolution tracks, and infrared photometry and colors, for ages $>$ 1 Myr and masses from above the hydrogen burning minimum mass down to planetary masses. The Hayashi phase of massive brown dwarf evolution (ages $<$ 10--100 Myr) at low surface gravity leads to nearly constant $T_\mathrm{eff}$ values, at effective temperatures much lower than would be obtained from simply extrapolating backwards from evolution tracks at older ages.
△ Less
Submitted 9 October, 2025;
originally announced October 2025.
-
COSMIR: Chain Orchestrated Structured Memory for Iterative Reasoning over Long Context
Authors:
Naman Gupta,
Shreeyash Gowaikar,
Arun Iyer,
Kirankumar Shiragur,
Ramakrishna B Bairi,
Rishikesh Maurya,
Ritabrata Maiti,
Sankarshan Damle,
Shachee Mishra Gupta
Abstract:
Reasoning over very long inputs remains difficult for large language models (LLMs). Common workarounds either shrink the input via retrieval (risking missed evidence), enlarge the context window (straining selectivity), or stage multiple agents to read in pieces. In staged pipelines (e.g., Chain of Agents, CoA), free-form summaries passed between agents can discard crucial details and amplify earl…
▽ More
Reasoning over very long inputs remains difficult for large language models (LLMs). Common workarounds either shrink the input via retrieval (risking missed evidence), enlarge the context window (straining selectivity), or stage multiple agents to read in pieces. In staged pipelines (e.g., Chain of Agents, CoA), free-form summaries passed between agents can discard crucial details and amplify early mistakes. We introduce COSMIR (Chain Orchestrated Structured Memory for Iterative Reasoning), a chain-style framework that replaces ad hoc messages with a structured memory. A Planner agent first turns a user query into concrete, checkable sub-questions. worker agents process chunks via a fixed micro-cycle: Extract, Infer, Refine, writing all updates to the shared memory. A Manager agent then Synthesizes the final answer directly from the memory. This preserves step-wise read-then-reason benefits while changing both the communication medium (structured memory) and the worker procedure (fixed micro-cycle), yielding higher faithfulness, better long-range aggregation, and auditability. On long-context QA from the HELMET suite, COSMIR reduces propagation-stage information loss and improves accuracy over a CoA baseline.
△ Less
Submitted 6 October, 2025;
originally announced October 2025.
-
First JWST thermal phase curves of temperate terrestrial exoplanets reveal no thick atmosphere around TRAPPIST-1 b and c
Authors:
Michaël Gillon,
Elsa Ducrot,
Taylor J. Bell,
Ziyu Huang,
Andrew Lincowski,
Xintong Lyu,
Alice Maurel,
Alexandre Revol,
Eric Agol,
Emeline Bolmont,
Chuanfei Dong,
Thomas J. Fauchez,
Daniel D. B. Koll,
Jérémy Leconte,
Victoria S. Meadows,
Franck Selsis,
Martin Turbet,
Benjamin Charnay,
Laetita Delre,
Brice-Olivier Demory,
Aaron Householder,
Sebastian Zieba,
David Berardo,
Achrène Dyrek,
Billy Edwards
, et al. (8 additional authors not shown)
Abstract:
We report JWST/MIRI 15 $μ$m phase curves of TRAPPIST-1 b and c, revealing thermal emission consistent with their irradiation levels, assuming no efficient heat redistribution. We find that TRAPPIST-1 b shows a high dayside brightness temperature (490 $\pm$ 17 K), no significantly detectable nightside emission ($F_{\rm b, Night, max}$ = $39_{-27}^{+55}$ ppm), and no phase offset -- features consist…
▽ More
We report JWST/MIRI 15 $μ$m phase curves of TRAPPIST-1 b and c, revealing thermal emission consistent with their irradiation levels, assuming no efficient heat redistribution. We find that TRAPPIST-1 b shows a high dayside brightness temperature (490 $\pm$ 17 K), no significantly detectable nightside emission ($F_{\rm b, Night, max}$ = $39_{-27}^{+55}$ ppm), and no phase offset -- features consistent with a low-albedo, airless ultramafic rocky surface. TRAPPIST-1 c exhibits a lower dayside brightness temperature (369 $\pm$ 23 K), and a nightside flux statistically indistinguishable from that of TRAPPIST-1 b ($F_{\rm c, Night, max}$ = $62_{-43}^{+60}$ ppm). Atmosphere models with surface pressures $\geq$1 bar and efficient greenhouse effects are strongly disfavoured for both planets. TRAPPIST-1 b is unlikely to possess any substantial atmosphere, while TRAPPIST-1 c may retain a tenuous, greenhouse-poor O$_2$-dominated atmosphere or be similarly airless with a more reflective surface. These results suggest divergent evolutionary pathways or atmospheric loss processes, despite similar compositions. These measurements tightly constrain atmosphere retention in the inner TRAPPIST-1 system.
△ Less
Submitted 2 September, 2025;
originally announced September 2025.
-
Quantitative Nonlinear Optical Polarimetry with High Spatial Resolution
Authors:
Albert Suceava,
Sankalpa Hazra,
Jadupati Nag,
John Hayden,
Safdar Imam,
Zhiwen Liu,
Abishek Iyer,
Mercouri Kanatzidis,
Susan Trolier-McKinstry,
Jon-Paul Maria,
Venkatraman Gopalan
Abstract:
Nonlinear optical microscopy such as in the optical second-harmonic generation (SHG) modality has become a popular tool today for probing materials in the physical and biological sciences. While imaging and spectroscopy are widely used in the microscopy mode, nonlinear polarimetry, which can shed light on materials' symmetry and microstructure, is relatively underdeveloped. This is partly because…
▽ More
Nonlinear optical microscopy such as in the optical second-harmonic generation (SHG) modality has become a popular tool today for probing materials in the physical and biological sciences. While imaging and spectroscopy are widely used in the microscopy mode, nonlinear polarimetry, which can shed light on materials' symmetry and microstructure, is relatively underdeveloped. This is partly because quantitative analytical modeling of the optical SHG response for anisotropic crystals and films largely assumes low-numerical aperture (NA) focusing of light, where the plane-wave approximation is sufficient. Tight focusing provides unique benefits in revealing out-of-plane polarization responses, which cannot be detected by near-plane-wave illumination at normal incidence. Here, we outline a method for quantitatively analyzing SHG polarimetry measurements obtained under high-NA focusing within a microscope geometry. Experiments and simulations of a variety of standard samples, from single crystals to thin films, are in good agreement, including measured and simulated spatial SHG maps of ferroelectric domains. A solution to the inverse problem is demonstrated, where the spatial distribution of an SHG tensor with unknown tensor coefficient magnitudes is determined by experimentally measured polarimetry. The ability to extract the out-of-plane component of the nonlinear polarization in normal incidence is demonstrated, which can be valuable for high-resolution polarimetry of 2D materials, thin films, heterostructures, and uniaxial crystals with a strong out-of-plane response.
Copyright 2025 Optica Publishing Group. Users may use, reuse, and build upon the article, or use the article for text or data mining, so long as such uses are for non-commercial purposes and appropriate attribution is maintained. All other rights are reserved. https://doi.org/10.1364/OPTICA.559060
△ Less
Submitted 30 July, 2025;
originally announced July 2025.
-
On the emergence of quantum memory in non-Markovian dynamics
Authors:
Alexander Yosifov,
Aditya Iyer,
Vlatko Vedral,
Jinzhao Sun
Abstract:
The emergence of memory is a hallmark feature of non-Markovian dynamics. However, the type of memory -- classical or quantum -- required to realize certain dynamics remains unknown. We study the quantum homogenizer as a minimal model of non-Markovian evolution and identify the physical conditions under which genuinely quantum memory becomes necessary. Using entanglement measures and relying only o…
▽ More
The emergence of memory is a hallmark feature of non-Markovian dynamics. However, the type of memory -- classical or quantum -- required to realize certain dynamics remains unknown. We study the quantum homogenizer as a minimal model of non-Markovian evolution and identify the physical conditions under which genuinely quantum memory becomes necessary. Using entanglement measures and relying only on the local dynamics as a witness, we prove both analytically and numerically the type of memory depends not merely on the dynamics itself, but also on the reservoir's initial entanglement structure, and in particular the propagation of non-classical correlations within it. For different bi- or multi-partite reservoir initializations, we establish a correspondence between interaction strength and entanglement generation. We provide physical criteria and an activation lower bound for the onset of quantum memory. The results may inform us how environmental correlations govern the transition from classical to quantum memory in open quantum systems.
△ Less
Submitted 25 November, 2025; v1 submitted 29 July, 2025;
originally announced July 2025.
-
Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions
Authors:
Meng Chen,
Akhil Iyer,
Amy Pavel
Abstract:
Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily lives. However, these models often produce errors that are difficult to detect without sight, posing safety and social risks in scenarios from medication identification to outfit selection. While BLV MLLM users use creative workarounds such as cross-c…
▽ More
Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily lives. However, these models often produce errors that are difficult to detect without sight, posing safety and social risks in scenarios from medication identification to outfit selection. While BLV MLLM users use creative workarounds such as cross-checking between tools and consulting sighted individuals, these approaches are often time-consuming and impractical. We explore how systematically surfacing variations across multiple MLLM responses can support BLV users to detect unreliable information without visually inspecting the image. We contribute a design space for eliciting and presenting variations in MLLM descriptions, a prototype system implementing three variation presentation styles, and findings from a user study with 15 BLV participants. Our results demonstrate that presenting variations significantly increases users' ability to identify unreliable claims (by 4.9x using our approach compared to single descriptions) and significantly decreases perceived reliability of MLLM responses. 14 of 15 participants preferred seeing variations of MLLM responses over a single description, and all expressed interest in using our system for tasks from understanding a tornado's path to posting an image on social media.
△ Less
Submitted 21 July, 2025;
originally announced July 2025.
-
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
Authors:
Junius Pun,
Xilai Dai,
Grace Zgheib,
Mahesh A. Iyer,
Andrew Boutros,
Vaughn Betz,
Mohamed S. Abdelfattah
Abstract:
Flexibility and customization are key strengths of Field-Programmable Gate Arrays (FPGAs) when compared to other computing devices. For instance, FPGAs can efficiently implement arbitrary-precision arithmetic operations, and can perform aggressive synthesis optimizations to eliminate ineffectual operations. Motivated by sparsity and mixed-precision in deep neural networks (DNNs), we investigate ho…
▽ More
Flexibility and customization are key strengths of Field-Programmable Gate Arrays (FPGAs) when compared to other computing devices. For instance, FPGAs can efficiently implement arbitrary-precision arithmetic operations, and can perform aggressive synthesis optimizations to eliminate ineffectual operations. Motivated by sparsity and mixed-precision in deep neural networks (DNNs), we investigate how to optimize the current logic block architecture to increase its arithmetic density. We find that modern FPGA logic block architectures prevent the independent use of adder chains, and instead only allow adder chain inputs to be fed by look-up table (LUT) outputs. This only allows one of the two primitives -- either adders or LUTs -- to be used independently in one logic element and prevents their concurrent use, hampering area optimizations. In this work, we propose the Double Duty logic block architecture to enable the concurrent use of the adders and LUTs within a logic element. Without adding expensive logic cluster inputs, we use 4 of the existing inputs to bypass the LUTs and connect directly to the adder chain inputs. We accurately model our changes at both the circuit and CAD levels using open-source FPGA development tools. Our experimental evaluation on a Stratix-10-like architecture demonstrates area reductions of 21.6% on adder-intensive circuits from the Kratos benchmarks, and 9.3% and 8.2% on the more general Koios and VTR benchmarks respectively. These area improvements come without an impact to critical path delay, demonstrating that higher density is feasible on modern FPGA architectures by adding more flexibility in how the adder chain is used. Averaged across all circuits from our three evaluated benchmark set, our Double Duty FPGA architecture improves area-delay product by 9.7%.
△ Less
Submitted 15 July, 2025;
originally announced July 2025.
-
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Authors:
Gheorghe Comanici,
Eric Bieber,
Mike Schaekermann,
Ice Pasupat,
Noveen Sachdeva,
Inderjit Dhillon,
Marcel Blistein,
Ori Ram,
Dan Zhang,
Evan Rosen,
Luke Marris,
Sam Petulla,
Colin Gaffney,
Asaf Aharoni,
Nathan Lintz,
Tiago Cardal Pais,
Henrik Jacobsson,
Idan Szpektor,
Nan-Jiang Jiang,
Krishna Haridasan,
Ahmed Omran,
Nikunj Saunshi,
Dara Bahri,
Gaurav Mishra,
Eric Chu
, et al. (3410 additional authors not shown)
Abstract:
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal unde…
▽ More
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal understanding and it is now able to process up to 3 hours of video content. Its unique combination of long context, multimodal and reasoning capabilities can be combined to unlock new agentic workflows. Gemini 2.5 Flash provides excellent reasoning abilities at a fraction of the compute and latency requirements and Gemini 2.0 Flash and Flash-Lite provide high performance at low latency and cost. Taken together, the Gemini 2.X model generation spans the full Pareto frontier of model capability vs cost, allowing users to explore the boundaries of what is possible with complex agentic problem solving.
△ Less
Submitted 19 December, 2025; v1 submitted 7 July, 2025;
originally announced July 2025.
-
Multi-boson splashes at future colliders from electroweak compositeness
Authors:
G. Cacciapaglia,
A. Deandrea,
A. M. Iyer,
S. Kulkarni,
A. K Singh
Abstract:
We propose a new collider signature for the composite origin of the electroweak symmetry breaking of the standard model. The Higgs sector consists of new fundamental fermions (hyper-quarks), which confine at a hadronization scale $Λ_{HC} \sim$ few TeV. At energies above $Λ_{HC}$, the Drell-Yan production of the hyper-quarks leads to the production of a few electroweak bosons, in analogy with hadro…
▽ More
We propose a new collider signature for the composite origin of the electroweak symmetry breaking of the standard model. The Higgs sector consists of new fundamental fermions (hyper-quarks), which confine at a hadronization scale $Λ_{HC} \sim$ few TeV. At energies above $Λ_{HC}$, the Drell-Yan production of the hyper-quarks leads to the production of a few electroweak bosons, in analogy with hadron production in QCD at $e^+e^- \to q\bar{q}$ around a few GeV. We show that this regime can be probed at future colliders, namely the proposed 100 TeV hadron collider (FCC-hh) and a 10 TeV muon collider. Together with the direct discovery of electroweak resonances, the multi electroweak boson signature provides a smoking gun for Higgs compositeness.
△ Less
Submitted 10 January, 2026; v1 submitted 24 June, 2025;
originally announced June 2025.
-
The Amazon Nova Family of Models: Technical Report and Model Card
Authors:
Amazon AGI,
Aaron Langford,
Aayush Shah,
Abhanshu Gupta,
Abhimanyu Bhatter,
Abhinav Goyal,
Abhinav Mathur,
Abhinav Mohanty,
Abhishek Kumar,
Abhishek Sethi,
Abi Komma,
Abner Pena,
Achin Jain,
Adam Kunysz,
Adam Opyrchal,
Adarsh Singh,
Aditya Rawal,
Adok Achar Budihal Prasad,
Adrià de Gispert,
Agnika Kumar,
Aishwarya Aryamane,
Ajay Nair,
Akilan M,
Akshaya Iyengar,
Akshaya Vishnu Kudlu Shanbhogue
, et al. (761 additional authors not shown)
Abstract:
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highly-capable multimodal model with the best combination of accuracy, speed, and cost for a wide range of tasks. Amazon Nova Lite is a low-cost multimodal model that is lightning fast for processing images, video, documents…
▽ More
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highly-capable multimodal model with the best combination of accuracy, speed, and cost for a wide range of tasks. Amazon Nova Lite is a low-cost multimodal model that is lightning fast for processing images, video, documents and text. Amazon Nova Micro is a text-only model that delivers our lowest-latency responses at very low cost. Amazon Nova Canvas is an image generation model that creates professional grade images with rich customization controls. Amazon Nova Reel is a video generation model offering high-quality outputs, customization, and motion control. Our models were built responsibly and with a commitment to customer trust, security, and reliability. We report benchmarking results for core capabilities, agentic performance, long context, functional adaptation, runtime performance, and human evaluation.
△ Less
Submitted 17 March, 2025;
originally announced June 2025.
-
Dynamic Sparse Training of Diagonally Sparse Networks
Authors:
Abhishek Tyagi,
Arjun Iyer,
William H Renninger,
Christopher Kanan,
Yuhao Zhu
Abstract:
Recent advances in Dynamic Sparse Training (DST) have pushed the frontier of sparse neural network training in structured and unstructured contexts, matching dense-model performance while drastically reducing parameter counts to facilitate model scaling. However, unstructured sparsity often fails to translate into practical speedups on modern hardware. To address this shortcoming, we propose DynaD…
▽ More
Recent advances in Dynamic Sparse Training (DST) have pushed the frontier of sparse neural network training in structured and unstructured contexts, matching dense-model performance while drastically reducing parameter counts to facilitate model scaling. However, unstructured sparsity often fails to translate into practical speedups on modern hardware. To address this shortcoming, we propose DynaDiag, a novel structured sparse-to-sparse DST method that performs at par with unstructured sparsity. DynaDiag enforces a diagonal sparsity pattern throughout training and preserves sparse computation in forward and backward passes. We further leverage the diagonal structure to accelerate computation via a custom CUDA kernel, rendering the method hardware-friendly. Empirical evaluations on diverse neural architectures demonstrate that our method maintains accuracy on par with unstructured counterparts while benefiting from tangible computational gains. Notably, with 90% sparse linear layers in ViTs, we observe up to a 3.13x speedup in online inference without sacrificing model performance and a 1.59x speedup in training on a GPU compared to equivalent unstructured layers. Our source code is available at https://github.com/horizon-research/DynaDiag/.
△ Less
Submitted 13 June, 2025;
originally announced June 2025.
-
Scaling Human Activity Recognition: A Comparative Evaluation of Synthetic Data Generation and Augmentation Techniques
Authors:
Zikang Leng,
Archith Iyer,
Thomas Plötz
Abstract:
Human activity recognition (HAR) is often limited by the scarcity of labeled datasets due to the high cost and complexity of real-world data collection. To mitigate this, recent work has explored generating virtual inertial measurement unit (IMU) data via cross-modality transfer. While video-based and language-based pipelines have each shown promise, they differ in assumptions and computational co…
▽ More
Human activity recognition (HAR) is often limited by the scarcity of labeled datasets due to the high cost and complexity of real-world data collection. To mitigate this, recent work has explored generating virtual inertial measurement unit (IMU) data via cross-modality transfer. While video-based and language-based pipelines have each shown promise, they differ in assumptions and computational cost. Moreover, their effectiveness relative to traditional sensor-level data augmentation remains unclear. In this paper, we present a direct comparison between these two virtual IMU generation approaches against classical data augmentation techniques. We construct a large-scale virtual IMU dataset spanning 100 diverse activities from Kinetics-400 and simulate sensor signals at 22 body locations. The three data generation strategies are evaluated on benchmark HAR datasets (UTD-MHAD, PAMAP2, HAD-AW) using four popular models. Results show that virtual IMU data significantly improves performance over real or augmented data alone, particularly under limited-data conditions. We offer practical guidance on choosing data generation strategies and highlight the distinct advantages and disadvantages of each approach.
△ Less
Submitted 13 June, 2025; v1 submitted 9 June, 2025;
originally announced June 2025.
-
Experimental Study of Rare Kaon Decays at J-PARC with KOTO and KOTO II
Authors:
J. K. Ahn,
E. Augustine,
L. Bandiera,
J. Bian,
F. Brizioli,
N. Canale,
G. A. Carini,
V. Chobanova,
G. D'Ambrosio,
J. B. Dainton,
S. De Capua,
P. Fedeli,
A. Gianoli,
A. Glazov,
M. Gonzalez,
E. Goudzovski,
M. Homma,
Y. B. Hsiung,
T. Husek,
A. M. Iyer,
E. J. Kim,
C. Kim,
T. K. Komatsubara,
K. Kotera,
M. Kreps
, et al. (40 additional authors not shown)
Abstract:
The rare kaon decay $K_L\toπ^0ν\barν$ is extremely sensitive to new physics, because the contribution to this decay in the Standard Model (SM) is highly suppressed and known very accurately; the branching ratio is $3\times 10^{-11}$ in the SM with a theoretical uncertainty of just 2%. The measurement of this branching ratio could provide essential new information about the flavor structure of the…
▽ More
The rare kaon decay $K_L\toπ^0ν\barν$ is extremely sensitive to new physics, because the contribution to this decay in the Standard Model (SM) is highly suppressed and known very accurately; the branching ratio is $3\times 10^{-11}$ in the SM with a theoretical uncertainty of just 2%. The measurement of this branching ratio could provide essential new information about the flavor structure of the quark sector from the $s\to d$ transition. The decay is being searched for in the KOTO experiment at J-PARC, which has obtained the current best upper limit on the branching ratio of $2.2\times 10^{-9}$; a sensitivity to branching ratios below $10^{-10}$ is achievable by the end of the decade. A next-generation experiment at J-PARC, KOTO II, was proposed in 2024 with 82 members worldwide, including significant contributions from European members. The goal of KOTO II is to measure the $K_L\toπ^0ν\barν$ branching ratio with sensitivity below $10^{-12}$ in the 2030s. Discovery of the decay with $5σ$ significance is achievable at the SM value of the branching ratio. An indication of new physics with a significance of 90% is possible if the observed branching ratio differs by 40% from the SM value. Another important goal of KOTO II is to measure the branching ratio of the unobserved $K_L\to π^0e^+e^-$ decay, which can give an input to flavor structures of new physics. Other rare $K_L$ decays and hidden-sector particles are also in the scope of the study. After 2026, KOTO will be the only dedicated rare kaon decay experiment in the world, and KOTO II is the only future rare kaon decay project currently proposed. We would like to lead a global initiative for the experimental study of rare kaon decays, with significant contributions and support from the European community.
△ Less
Submitted 5 May, 2025;
originally announced May 2025.
-
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
Authors:
Yusen Zhang,
Wenliang Zheng,
Aashrith Madasu,
Peng Shi,
Ryo Kamoi,
Hao Zhou,
Zhuoyang Zou,
Shu Zhao,
Sarkar Snigdha Sarathi Das,
Vipul Gupta,
Xiaoxin Lu,
Nan Zhang,
Ranran Haoran Zhang,
Avitej Iyer,
Renze Lou,
Wenpeng Yin,
Rui Zhang
Abstract:
High-resolution image (HRI) understanding aims to process images with a large number of pixels, such as pathological images and agricultural aerial images, both of which can exceed 1 million pixels. Vision Large Language Models (VLMs) can allegedly handle HRIs, however, there is a lack of a comprehensive benchmark for VLMs to evaluate HRI understanding. To address this gap, we introduce HRScene, a…
▽ More
High-resolution image (HRI) understanding aims to process images with a large number of pixels, such as pathological images and agricultural aerial images, both of which can exceed 1 million pixels. Vision Large Language Models (VLMs) can allegedly handle HRIs, however, there is a lack of a comprehensive benchmark for VLMs to evaluate HRI understanding. To address this gap, we introduce HRScene, a novel unified benchmark for HRI understanding with rich scenes. HRScene incorporates 25 real-world datasets and 2 synthetic diagnostic datasets with resolutions ranging from 1,024 $\times$ 1,024 to 35,503 $\times$ 26,627. HRScene is collected and re-annotated by 10 graduate-level annotators, covering 25 scenarios, ranging from microscopic to radiology images, street views, long-range pictures, and telescope images. It includes HRIs of real-world objects, scanned documents, and composite multi-image. The two diagnostic evaluation datasets are synthesized by combining the target image with the gold answer and distracting images in different orders, assessing how well models utilize regions in HRI. We conduct extensive experiments involving 28 VLMs, including Gemini 2.0 Flash and GPT-4o. Experiments on HRScene show that current VLMs achieve an average accuracy of around 50% on real-world tasks, revealing significant gaps in HRI understanding. Results on synthetic datasets reveal that VLMs struggle to effectively utilize HRI regions, showing significant Regional Divergence and lost-in-middle, shedding light on future research.
△ Less
Submitted 29 April, 2025; v1 submitted 25 April, 2025;
originally announced April 2025.
-
RUKA: Rethinking the Design of Humanoid Hands with Learning
Authors:
Anya Zorin,
Irmak Guzey,
Billy Yan,
Aadhithya Iyer,
Lisa Kondrich,
Nikhil X. Bhattasali,
Lerrel Pinto
Abstract:
Dexterous manipulation is a fundamental capability for robotic systems, yet progress has been limited by hardware trade-offs between precision, compactness, strength, and affordability. Existing control methods impose compromises on hand designs and applications. However, learning-based approaches present opportunities to rethink these trade-offs, particularly to address challenges with tendon-dri…
▽ More
Dexterous manipulation is a fundamental capability for robotic systems, yet progress has been limited by hardware trade-offs between precision, compactness, strength, and affordability. Existing control methods impose compromises on hand designs and applications. However, learning-based approaches present opportunities to rethink these trade-offs, particularly to address challenges with tendon-driven actuation and low-cost materials. This work presents RUKA, a tendon-driven humanoid hand that is compact, affordable, and capable. Made from 3D-printed parts and off-the-shelf components, RUKA has 5 fingers with 15 underactuated degrees of freedom enabling diverse human-like grasps. Its tendon-driven actuation allows powerful grasping in a compact, human-sized form factor. To address control challenges, we learn joint-to-actuator and fingertip-to-actuator models from motion-capture data collected by the MANUS glove, leveraging the hand's morphological accuracy. Extensive evaluations demonstrate RUKA's superior reachability, durability, and strength compared to other robotic hands. Teleoperation tasks further showcase RUKA's dexterous movements. The open-source design and assembly instructions of RUKA, code, and data are available at https://ruka-hand.github.io/.
△ Less
Submitted 17 April, 2025;
originally announced April 2025.
-
Kaon Physics: A Cornerstone for Future Discoveries
Authors:
Jason Aebischer,
Atakan Tugberk Akmete,
Riccardo Aliberti,
Wolfgang Altmannshofer,
Fabio Ambrosino,
Roberto Ammendola,
Antonella Antonelli,
Giuseppina Anzivino,
Saiyad Ashanujjaman,
Laura Bandiera,
Damir Becirevic,
Véronique Bernard,
Johannes Bernhard,
Cristina Biino,
Johan Bijnens,
Monika Blanke,
Brigitte Bloch-Devaux,
Marzia Bordone,
Peter Boyle,
Alexandru Mario Bragadireanu,
Francesco Brizioli,
Joachim Brod,
Andrzej J. Buras,
Dario Buttazzo,
Nicola Canale
, et al. (131 additional authors not shown)
Abstract:
The kaon physics programme, long heralded as a cutting-edge frontier by the European Strategy for Particle Physics, continues to stand at the intersection of discovery and innovation in high-energy physics (HEP). With its unparalleled capacity to explore new physics at the multi-TeV scale, kaon research is poised to unveil phenomena that could reshape our understanding of the Universe. This docume…
▽ More
The kaon physics programme, long heralded as a cutting-edge frontier by the European Strategy for Particle Physics, continues to stand at the intersection of discovery and innovation in high-energy physics (HEP). With its unparalleled capacity to explore new physics at the multi-TeV scale, kaon research is poised to unveil phenomena that could reshape our understanding of the Universe. This document highlights the compelling physics case, with emphasis on exciting new opportunities for advancing kaon physics not only in Europe but also on a global stage. As an important player in the future of HEP, the kaon programme promises to drive transformative breakthroughs, inviting exploration at the forefront of scientific discovery.
△ Less
Submitted 28 March, 2025;
originally announced March 2025.