-
The Memory Hidden in Response Fluctuations: Trajectory-Level Fluctuation-Response Theory and Inequalities for Non-Markovian Jump Dynamics
Authors:
Jiming Zheng,
Zhiyue Lu
Abstract:
Modern experiments often record nonequilibrium dynamics as sequences of discrete events whose rates depend on the realized past. We develop a fluctuation-response theory for such non-Markovian jump processes directly on the observed event record. Memory can destroy a closed master equation for the state probabilities. Each transition count nevertheless obeys an exact stochastic equation. After the…
▽ More
Modern experiments often record nonequilibrium dynamics as sequences of discrete events whose rates depend on the realized past. We develop a fluctuation-response theory for such non-Markovian jump processes directly on the observed event record. Memory can destroy a closed master equation for the state probabilities. Each transition count nevertheless obeys an exact stochastic equation. After the history-dependent mean event tendency is subtracted, the remaining random increment is a martingale increment--the part of the event that cannot be predicted from the past. Martingale increments associated with different transitions and different times are orthogonal. These increments form a complete orthogonal basis for the random deviation of any observable measured from the record, such as a current, occupation time, or event count. The coefficient of a given increment is the event-consequence kernel. It measures how that event changes the predicted final observable, relative to continuing without the event, for the particular history already realized. Multiplying this kernel by the event intensity gives the history-conditioned response to perturbing the corresponding transition rate. Thus, the intensity-normalized response is exactly the expansion coefficient of that event in the observable fluctuation. This identification yields exact finite-time and finite-frequency fluctuation-response relations. Averaging over histories leaves a nonnegative response-heterogeneity gap. The gap measures how strongly the consequence of the same event varies across histories and tests whether a proposed memory coordinate is response-sufficient. Finite-history versions can be estimated from spontaneous trajectories. Bounding the logarithmic rate sensitivity further gives response-kinetic uncertainty relations controlled by dynamical activity.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Chameleon: Robust Defense Against Tor Website Fingerprinting via Many-to-Many Traffic Morphing
Authors:
Yuwen Cui,
Kai Wei,
Kehan Shen,
Ning Wang,
Zhuo Lu,
Yao Liu,
Guangjing Wang
Abstract:
Website fingerprinting (WF) attacks can infer users' browsing activities from encrypted Tor traffic by exploiting side-channel features. Although many WF defenses have been proposed, we find that most existing defenses create learnable web trace mapping features. We further show that robustness against adversarial training does not necessarily imply robustness against defense-aware autoencoder (DA…
▽ More
Website fingerprinting (WF) attacks can infer users' browsing activities from encrypted Tor traffic by exploiting side-channel features. Although many WF defenses have been proposed, we find that most existing defenses create learnable web trace mapping features. We further show that robustness against adversarial training does not necessarily imply robustness against defense-aware autoencoder (DAAE)-based attacks.
To address these limitations, we present Chameleon, a robust WF defense based on many-to-many randomized traffic morphing. Chameleon selects morphing candidates with high intra-class diversity and low inter-class disparity. Chameleon randomly maps each webpage trace to multiple candidates, and allows different webpages to share morphing targets, thereby increasing adversarial uncertainty. For practical Tor deployment, Chameleon introduces a radix-trie-based synchronization mechanism that enables pluggable transport (PT) endpoints to identify consistent morphing traces using packet-direction prefixes, together with trace mutation and normalized prefix matching to reduce overhead. We evaluate Chameleon against six state-of-the-art defenses and five WF attacks on three public datasets in closed- and open-world settings. Compared with Adaptive Tamaraw, Chameleon reduces adversarial-training-based attack accuracy by up to 36.74% while reducing bandwidth and time overhead by 34.12% and 60.38%, respectively. Under DAAE-based RF attacks on GTT23, Chameleon limits attack performance to 35.19% F1-score while Adaptive Tamaraw only limits it to 88.22% F1-score. In the real-world PT bridge evaluation, Chameleon substantially reduces the effectiveness of strong WF attacks while incurring only 16.25% time overhead.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Lattice-data-driven specific heat and isentropic bulk modulus of SU(3) gluon matter at finite temperature
Authors:
Wei Shen,
Zhen-Yan Lu,
Muhammad Waqas,
Xun Chen,
Zhi-Jun Ma,
Guang-Xiong Peng
Abstract:
We investigate the specific heat and isentropic bulk modulus of finite-temperature pure SU(3) gauge matter within a lattice-data-driven phenomenological framework. The equation of state is formulated in terms of a temperature-dependent effective gluon mass constrained { by lattice QCD pressure data as input, allowing the pressure}, trace anomaly, gluon number density, energy per thermally active g…
▽ More
We investigate the specific heat and isentropic bulk modulus of finite-temperature pure SU(3) gauge matter within a lattice-data-driven phenomenological framework. The equation of state is formulated in terms of a temperature-dependent effective gluon mass constrained { by lattice QCD pressure data as input, allowing the pressure}, trace anomaly, gluon number density, energy per thermally active gluonic mode, and derivative-sensitive response functions to be derived in a thermodynamically consistent manner. The resulting pressure and trace anomaly reproduce the characteristic lattice behavior across the deconfinement region, while the effective gluonic degrees of freedom increase rapidly above $T_c$. The normalized specific heat $C_V/T^3$ develops a pronounced enhancement in the vicinity of $T_c$, reflecting the rapid temperature variation of the energy density across the deconfinement region. The isentropic bulk modulus $K_S/T^4$ also rises sharply across the transition region, indicating a substantial stiffening of the equation of state. At high temperatures, both response functions gradually approach values close to their massless conformal Stefan--Boltzmann reference values, with $\left(C_V/T^3\right)_{\rm SB}=32π^2/15\simeq 21.06$ and $\left(K_S/T^4\right)_{\rm SB}=32π^2/135\simeq 2.34$. These findings indicate that the specific heat and isentropic bulk modulus provide complementary constraints on the temperature evolution of nonconformal dynamics in pure SU(3) gauge matter.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
HealMed: Multilingual Evaluation of Large Language Models in Medicine
Authors:
Yingjian Chen,
Fan Gao,
Sherry T. Tong,
Haoyu Zhang,
Aosong Feng,
Kevin W. Jin,
Xing Wu,
Jinghui Lu,
Abdul Samad,
Akbar Faruqi,
Cesar Caraballo,
Cibele Brandão,
Dhruva,
Gupta,
Eunji Jeon,
Gabriel Madera-Santiago,
Geon Lee,
Hugo Toshio Itikawa,
Insook Cho,
Isabelli Martins,
Isarar Siddique,
Israr Ahmed,
Jihyo Kwak,
Kanyakorn Veerakanjana,
Luis Guilherme Cardoso
, et al. (20 additional authors not shown)
Abstract:
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation w…
▽ More
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Learning Highly Dynamic Skills Transition for Quadruped Jumping Through Constrained Space
Authors:
Zeren Luo,
Jiahui Zhang,
Yimin Han,
Ji Ma,
Minghao Lu,
Ioannis Havoutis,
Peng Lu
Abstract:
Although legged animals are capable of performing explosive motions while traversing confined spaces, replicating this behavior in quadrupedal robots has been a longstanding challenge. Here, we propose a hierarchical reinforcement learning pipeline that empowers the robots to perform aggressive locomotion through constrained obstacles--a narrow gate. The imitation learning technique is used to tra…
▽ More
Although legged animals are capable of performing explosive motions while traversing confined spaces, replicating this behavior in quadrupedal robots has been a longstanding challenge. Here, we propose a hierarchical reinforcement learning pipeline that empowers the robots to perform aggressive locomotion through constrained obstacles--a narrow gate. The imitation learning technique is used to train the low-level policy, which mimics the behaviors of real animals and forms a set of diverse skills. The high-level controller, having an awareness of the capability of low-level skills and acquiring the gate information via vision-based detection, determines the suitable maneuvers with collision-free trajectories to traverse it dynamically. Notably, we also verify that this framework can be extended to other highly dynamic tasks. This is one of the first works that perform autonomous and agile aerial gate traversal tasks on ground-walking robots, extending the lifelike agility of legged robots to match that of their biological counterparts.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
MILD: Tractable Terrain Modeling for Learning Improved Bipedal Locomotion on Deformable Surfaces
Authors:
Zeren Luo,
Jiahui Zhang,
Zhe Xu,
Wanyue Li,
Xinqi Li,
Xuechao Chen,
Zhangguo Yu,
Annan Tang,
Peng Lu
Abstract:
Enabling robots to walk on yielding terrain is vital for applications ranging from disaster response to planetary exploration. While bipedal robots hold immense potential, their locomotion on deformable surfaces remains limited as current simulators fail to capture the spatiotemporal heterogeneity of such yielding substrates. We present MILD, featuring a physics-grounded discrete-element contact s…
▽ More
Enabling robots to walk on yielding terrain is vital for applications ranging from disaster response to planetary exploration. While bipedal robots hold immense potential, their locomotion on deformable surfaces remains limited as current simulators fail to capture the spatiotemporal heterogeneity of such yielding substrates. We present MILD, featuring a physics-grounded discrete-element contact solver that accurately simulates spatially varying foot-terrain interactions. Complementing this model, we train a terrain-aware locomotion controller via deep reinforcement learning with latent modulation and proprioceptive estimation. Quantitative comparisons against state-of-the-art methods show our approach generates more diverse and realistic contact scenarios during training, resulting in controllers that exhibit natural adaptation on real deformable surfaces. Through hardware experiments, we demonstrate the system's capability for online terrain identification and adaptation across a wide range of surface stiffness.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Learning Early-to-Final Solution Consistency for MILP Acceleration
Authors:
Guanlin Li,
Chengrui Gao,
Chenguang Wang,
Haopu Shang,
Zherong Zhang,
Ke Xue,
Jixiang Lu,
Weiyong Yang,
Chao Qian
Abstract:
Mixed-Integer Linear Programming (MILP) is a fundamental problem class in operations research and combinatorial optimization, with broad applications to industrial decision-making. Owing to their NP-hardness, however, modern solvers may struggle to find high-quality solutions for challenging MILP instances within practical time limits. Recent learning-based approaches seek to accelerate MILP solvi…
▽ More
Mixed-Integer Linear Programming (MILP) is a fundamental problem class in operations research and combinatorial optimization, with broad applications to industrial decision-making. Owing to their NP-hardness, however, modern solvers may struggle to find high-quality solutions for challenging MILP instances within practical time limits. Recent learning-based approaches seek to accelerate MILP solving by directly predicting high-quality solutions from static instance-level features, such as variable-constraint bipartite graphs. Yet accurate solution prediction from instance features alone is difficult, and these methods largely overlook the information revealed during the solver's search process. In this paper, we find that solutions produced at the early search stage of MILP solvers, which are computationally cheap to obtain, are often structurally close to the solutions found after full-budget search. Motivated by this observation, we propose a new solver-informed paradigm that shifts the learning target from variable assignment to early-to-final consistency: for each variable, we predict whether its early-stage assignment should persist in full-budget solutions. The predicted consistency naturally guides downstream search, for instance by fixing the assignments deemed consistent. At inference time, we further ensemble consistency predictions across multiple early-stage solutions to improve robustness. Experiments across four MILP benchmarks show our method improves prediction-guided search across diverse downstream pipelines. With Gurobi, our proposed method reduces the primal gap by 56.9% on average and closes it completely on combinatorial auction instances. Besides, we transferred the Gurobi-trained model zero-shot to SCIP without adaptation, achieving a 36.4% average gap reduction across benchmarks.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
SN 2021pfs: A Type Ia Supernova Likely Affected by Progenitor Metallicity, as Revealed by Comparison with Its Twin Counterpart
Authors:
Abdusamatjan Iskandar,
Xiaofeng Wang,
Ali Esamdin,
Wenxiong Li,
Xiangyun Zeng,
Ruifeng Huang,
Guoliang Lu,
D. Andrew Howell,
Curtis McCully,
Samuel Wyatt,
Daichi Hiramatsu,
Estefania Padilla Gonzalez,
Craig Pellegrino,
Megan Newsome,
Lluis Galbany,
David J. Sand,
Nathan Smith,
Huei Sears
Abstract:
We present extensive photometric and spectroscopic observations of the normal type Ia supernovae (SNe Ia) 2021pfs, which occurred in the Seyfert 2 galaxy NGC 5427 at a redshift 0.009. SN 2021pfs reached an absolute \textit{B}-band peak magnitude of $M_{\rm max}(B)=$-19.28 $\pm$ 0.40 mag. The mag and a post-peak decline rate of $Δm_{15}(B)=$1.13 $\pm$ 0.06 mag. The observed properties of this nearb…
▽ More
We present extensive photometric and spectroscopic observations of the normal type Ia supernovae (SNe Ia) 2021pfs, which occurred in the Seyfert 2 galaxy NGC 5427 at a redshift 0.009. SN 2021pfs reached an absolute \textit{B}-band peak magnitude of $M_{\rm max}(B)=$-19.28 $\pm$ 0.40 mag. The mag and a post-peak decline rate of $Δm_{15}(B)=$1.13 $\pm$ 0.06 mag. The observed properties of this nearby SN Ia closely resemble those of SN 2011fe, including the main optical spectroscopic features and photometric evolution. Despite their similar decline rates, SN 2021pfs rose more rapidly in the $U$ band but more slowly in the $r$ and $i$ bands compared to SN 2011fe in very early phases. This photometric difference, particularly at short wavelengths, can introduce a systematic uncertainty of up to $\sim$12% in distance estimates. Analysis of the host galaxy's local and global environment shows an environment consistent with producing a higher-metallicity progenitor for SN 2021pfs than that of SN 2011fe.This higher progenitor metallicity may explain the observed photometric discrepancy and the resulting distance between SN 2021pfs and SN 2011fe, though a larger sample of such "twin" SNe Ia is needed to confirm this trend and assess its impact on cosmological measurements.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis
Authors:
Zijiao Chen,
Nicholas Lu,
Xinhui Li,
Jocelyn A. Ricard,
Ce Ju,
Huan H. Wang,
Christian Kindermann,
Jeanette A. Mumford,
Steven Dillmann,
James Kent,
Alejandro de la Vega,
Sanmi Koyejo,
Vince D. Calhoun,
Joshua W. Buckholtz,
Juan Helen Zhou,
Steffen Bollmann,
Russell A. Poldrack
Abstract:
AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimag…
▽ More
AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimaging researcher's computational environment under rules for admissible analyses, required checks and claim scope. In benchmarks, Brain Researcher increased first-choice tool-selection accuracy across seven models by 70.2 percentage points (23.3% without it versus 93.6% with it) and verifiable grounding from 4.6% to 22.0%. In collaborator-led and self-evolving studies, multiverse analyses exposed analytic-choice sensitivity, and scientific review classified claims as accepted, qualified, revised, blocked, rejected or deferred. By linking decisions to evidence and provenance, Brain Researcher embeds methodological judgment within the workflow, not after it.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Signed list edge coloring in graphs of bounded treewidth
Authors:
Li Zhang,
You Lu,
Zhengke Miao,
Yintao Wang
Abstract:
Vizing conjectured that the list edge chromatic number of any graph with maximum degree $Δ$ is at most $Δ+ 1$. This conjecture has been confirmed for several important classes of graphs, in particular, Lang proved that it holds for all graphs of treewidth $3$. In this paper, we introduce the list edge coloring of signed graphs, a framework that generalizes both classical list edge coloring and the…
▽ More
Vizing conjectured that the list edge chromatic number of any graph with maximum degree $Δ$ is at most $Δ+ 1$. This conjecture has been confirmed for several important classes of graphs, in particular, Lang proved that it holds for all graphs of treewidth $3$. In this paper, we introduce the list edge coloring of signed graphs, a framework that generalizes both classical list edge coloring and the signed edge coloring introduced by Behr. We extend Lang's result by proving the signed analogue of Vizing's conjecture for all signed graphs of treewidth $3$, as well as for signed graphs of treewidth $4$ with maximum degree $Δ\ge 10$.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
Authors:
Zhipeng Xu,
Jiahao Lu,
Yining Zheng,
Yuxin Wang,
Xipeng Qiu
Abstract:
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software. We introduce \…
▽ More
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software. We introduce \textbf{SWE-bench Science}, a repository-level benchmark for scientific software engineering comprising 119 tasks from 98 GitHub repositories across 20 scientific domains. Each task is organized into one of three paradigms: Issue-driven, Expert-exploratory, and Engineering-integration. Even the best-performing agent, \textbf{Claude Code with Opus-5 (max), achieves a pass@1 below 50\%}, highlighting the substantial challenges posed by scientific software engineering. We identify four recurring failure mechanisms: deficits in scientific knowledge or abstraction, misguided exploration or surface-level repair, incomplete repair coverage or system integration, and failures to generalize scientific knowledge beyond observed cases in our analysis. We further conduct a paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context. The results show that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success. Together, SWE-bench Science provides a broad testbed for studying both the capabilities and failure mechanisms of coding agents in scientific software engineering.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
GenMatch: An End-to-End Generative Matching Framework for Micro-View Order-Dispatching in Ride-Hailing
Authors:
Chuang Liu,
Yuxueqing Zhang,
Tengfei Lyu,
Zirui Yuan,
Weiqi Hu,
Yanghan Cheng,
Ming Wang,
Li Ma,
Zihao Lu
Abstract:
Micro-View Order-Dispatching assigns available drivers to passenger orders within each dispatch batch and is critical to the service quality and operational efficiency of ride-hailing platforms. Mainstream industrial solutions follow a multi-stage paradigm of model prediction, value calculation, and dispatch matching. Although dispatch quality is determined by the final batch-level assignment, the…
▽ More
Micro-View Order-Dispatching assigns available drivers to passenger orders within each dispatch batch and is critical to the service quality and operational efficiency of ride-hailing platforms. Mainstream industrial solutions follow a multi-stage paradigm of model prediction, value calculation, and dispatch matching. Although dispatch quality is determined by the final batch-level assignment, these stages optimize different intermediate objectives. This cross-stage objective inconsistency means that improving a single stage does not necessarily improve the overall dispatch result. We therefore formulate Micro-View Order-Dispatching as a generative matching problem and propose GenMatch, an end-to-end Generative Matching framework and the first such framework deployed in a real-world production environment. Applying generative modeling to this problem introduces three challenges. First, each dispatch batch forms a dynamic sparse bipartite graph, requiring efficient structured batch-level encoding. Second, replacing the hand-crafted value function requires learning unified business utility from heterogeneous feedback. Third, directly generating an assignment requires tracking the evolving matching state because each selected order-driver pair changes the remaining feasible candidates. GenMatch addresses these challenges with a Context-Aware Bipartite Encoder, a Business-Aware Utility Learner, and a State-Aware Pointer Decoder. Extensive offline evaluations and online A/B tests in five cities across DiDi's international ride-hailing markets show consistent improvements over competitive baselines, confirming the effectiveness and practicality of GenMatch for industrial order-dispatching.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Loss-Resilient Semantic Communication over Packet-Loss Networks at Extreme-Low Bandwidth
Authors:
Shengshi Yao,
Jincheng Dai,
Sixian Wang,
Guo Lu,
Kai Niu,
Wenjun Xu,
Wenjun Zhang,
Ping Zhang
Abstract:
In extreme-low bandwidth network scenarios, generative semantic codecs have emerged as promising solutions to reduce bandwidth cost for visual communications. However, these learned codecs are usually optimized solely for compression efficiency and thus not robust against transmission errors. Corruptions due to packet-loss among these highly compact generative latent representations often cause mo…
▽ More
In extreme-low bandwidth network scenarios, generative semantic codecs have emerged as promising solutions to reduce bandwidth cost for visual communications. However, these learned codecs are usually optimized solely for compression efficiency and thus not robust against transmission errors. Corruptions due to packet-loss among these highly compact generative latent representations often cause more critical degradation in fidelity and realism, intensified by the severe error propagation across the latent contexts and multi-step decoding process. In this paper, we propose ResiGLC, a novel loss-resilient generative latent coding framework designed for robust semantic communication over extreme-low bandwidth packet-loss networks. Motivated by the inherent goal-consistency between generation and compression, we sufficiently exploit the impressive in-context predictive capabilities of language models. Integrated with the masked learning strategy, our model supports arbitrary context modeling of latent codes, which could mitigate the error propagation and handle unpredictable packet loss patterns. At the receiver, a progressive resilient decoding pipeline is presented, which leverages both the contextual relationship of the latent codes and the multi-modal semantic prior in the generative latent space, separately. By jointly optimizing toward both compression efficiency and packet-loss resilience, our proposed progressive decoding mechanism offers graceful performance when dealing with dynamic packet losses. Through extensive experimental evaluations, we establish that under packet-loss network conditions, ResiGLC can effectively improve the loss-resilience in terms of perceptual fidelity and realism qualities with extreme-low bandwidth cost.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering
Authors:
Xinjin Li,
Yudi Xia,
Xi Zhao,
Yiliu Xu,
Yining Liu,
Cheng Lu,
Yujian Long,
Yu Ma,
Jinghan Cao,
Liang Fan,
Yeyun Xu
Abstract:
Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence. Compared with conventional open-ended VQA, educational examples often include structured assessment metadata, diagrams or image contexts, and semantically close answer options, creating strong opportunities for question-option shortcuts. We…
▽ More
Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence. Compared with conventional open-ended VQA, educational examples often include structured assessment metadata, diagrams or image contexts, and semantically close answer options, creating strong opportunities for question-option shortcuts. We develop and evaluate a parameter-efficient adaptation framework for a frozen multimodal large language model in this setting. We introduce GRACE, Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration, a framework that uses the pedagogical state of each question to specialize lightweight language and vision adaptation. The state combines inference-visible subject, grouped skill, grade, visual-context, question-intent, and option-structure cues. GRACE uses factor-specific prompts and lightweight visual adapters, then applies evidence-aware option calibration to score all candidates under a shared multimodal context. On ScienceQA, GRACE improves a shared-adapter baseline from 90.5 percent to 93.1 percent overall accuracy and from 88.7 percent to 91.2 percent on image-context questions. Removing pedagogical composition, option calibration, or the visual adapter reduces overall accuracy by 1.4, 1.0, and 1.5 points, respectively. These controlled results show that structured educational state is an effective routing signal for parameter-efficient multimodal adaptation.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Detectable subhalo impacts in Milky Way streams
Authors:
Junyang Lu,
Elias Bernreuther,
Tongyan Lin,
Vincent S. H. Lee,
Ana Bonaca,
Ethan O. Nadler
Abstract:
Dark matter subhalos leave gravitational imprints in the stellar streams of the Milky Way. Observing individual strong impacts of subhalos offers a compelling way to constrain and discover potentially dark subhalos down to $10^6 M_\odot$, allowing for new tests of the particle physics properties of dark matter. We develop a pipeline and statistical framework to forecast the expected number of dete…
▽ More
Dark matter subhalos leave gravitational imprints in the stellar streams of the Milky Way. Observing individual strong impacts of subhalos offers a compelling way to constrain and discover potentially dark subhalos down to $10^6 M_\odot$, allowing for new tests of the particle physics properties of dark matter. We develop a pipeline and statistical framework to forecast the expected number of detectable subhalo impacts on stellar streams, based on morphological and kinematic data from surveys such as LSST and Via. Starting from a catalog of confirmed stellar streams, we focus our efforts on 14 promising streams that are relatively well-modeled with a particle spray algorithm. Our criteria for a detectable impact is a deviation at 95% CL from the best-fit polynomial proxy model for the stream, which accounts for stream modeling uncertainties and regulates the effect of distant impacts that are degenerate with these uncertainties. Among the 14 streams studied, we find that 5 streams have an expected number of detectable impacts greater than 0.2. With LSST and Via data, the stream Jet has $5.15^{+1.10}_{-0.95}$ expected detectable impacts, followed by Orphan-Chenab ($1.40^{+0.62}_{-0.47}$), ATLAS-Aliqa Uma ($1.25^{+0.60}_{-0.44}$), GD-1 ($0.55^{+0.43}_{-0.28}$), and Palomar 5 ($0.40^{+0.39}_{-0.23}$), where error bars are the 95% containment on the Poisson mean. These values rely on the assumed subhalo population, which can give a factor of few systematic uncertainty in the predictions. We also consider effects of different particle dark matter models on the number of impacts, finding a suppression by a factor of $\sim 4$ for warm dark matter and fuzzy dark matter models at their current mass bounds and an $O(1)$ enhancement for a toy model of self-interacting dark matter.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Image-Guided Pavement Defect Recognition in GPR Data with novel 3D Deep Learning Architecture
Authors:
Yuandong Pan,
Linjun Lu,
Mudan Wang,
Florian Noichl,
Fan Xue,
Brian Sheil,
Lavindra de Silva,
André Borrmann,
Ioannis Brilakis
Abstract:
Ground Penetrating Radar (GPR) is a widely adopted non-destructive sensing technology for subsurface inspection in civil and transportation engineering. Despite its potential for pavement condition assessment, the large-scale application of GPR in automated inspection has two key challenges: the scarcity of annotated real-world datasets and the lack of deep learning models designed for the unique…
▽ More
Ground Penetrating Radar (GPR) is a widely adopted non-destructive sensing technology for subsurface inspection in civil and transportation engineering. Despite its potential for pavement condition assessment, the large-scale application of GPR in automated inspection has two key challenges: the scarcity of annotated real-world datasets and the lack of deep learning models designed for the unique characteristics of 3-Dimensional (3D) GPR data. This study addresses these limitations by firstly introducing a cost-effective data preparation pipeline that integrates orthomosaic Red Green Blue (RGB) imagery with 3D GPR scans to generate annotated 3D GPR datasets. The proposed method uses the aligned segments of RGB and GPR data, using pavement surface images as a reference to transfer labels of surface-visible defects to corresponding GPR segments, enabling efficient large-scale annotation in a real-world dataset collected on a highway section under operation. In addition to the dataset contribution, we propose a specialised 3D Convolutional Neural Network (CNN) architecture incorporating residual connections, mixed convolutional kernel sizes, and both depthwise and channelwise attention mechanisms to enhance feature representation and defect classification. The model is evaluated on binary classification tasks for detecting patch and crack defects in pavement structures. Experimental results demonstrate that the proposed network outperforms baseline architectures across multiple evaluation metrics. Ablation studies further confirm the effectiveness of the designed architectural components. This work contributes a scalable and practical method for real-world dataset generation, along with a novel deep learning framework.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Pretraining Reusable Inference Across Views with Synthetic Task Priors
Authors:
Jielong Lu,
Zhihao Wu,
Jiajun Yu,
Zhaoliang Chen,
Haishuai Wang
Abstract:
Modern pretrained encoders make representations from heterogeneous views increasingly reusable, but the procedure that determines view utility and combines evidence is still relearned for each downstream task. Consequently, knowledge about view relevance, complementarity, reliability, and missingness is repeatedly discarded rather than transferred across tasks. We therefore reformulate multi-view…
▽ More
Modern pretrained encoders make representations from heterogeneous views increasingly reusable, but the procedure that determines view utility and combines evidence is still relearned for each downstream task. Consequently, knowledge about view relevance, complementarity, reliability, and missingness is repeatedly discarded rather than transferred across tasks. We therefore reformulate multi-view learning as learning a reusable, task-conditioned inference procedure rather than a fixed fusion function. Based on this perspective, we propose SIMPLE, a prior-fitted multi-view in-context learner that predicts query labels by conditioning on a small labeled support set. Since existing real-world datasets cover only a limited range of view configurations and task structures, we construct a controllable synthetic task prior in embedding space. It generates diverse support-query episodes with varying class structures, shared and view-specific factors, representation geometries, cross-view dependencies, reliability levels, missingness patterns, and distribution shifts. A hierarchical inference architecture then performs reasoning within views, across views, and across support and query samples. Experiments on multi-view and multi-omics benchmarks demonstrate that the frozen variant of SIMPLE achieves competitive performance without updating the inference backbone, while lightweight adapter calibration attains leading performance on most evaluated datasets. Together, the results under frozen, one-shot, and missing-view settings support the central hypothesis that multi-view reasoning itself can be pretrained and reused, while lightweight adapter calibration provides task-specific alignment when needed.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
What is Missing from AI Post-Training AI: An Empirical Analysis
Authors:
Joy Jia Yin Lim,
Xin Huang,
Hao Peng,
Yaxi Lu,
Xin Cong,
Zhong Zhang,
Maosong Sun,
Yankai Lin
Abstract:
Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. We argue that this picture conflates two distinct capabilities: execution-level capability, iterating within a selected training strategy; and strategy-level capability, revising the high-level jud…
▽ More
Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. We argue that this picture conflates two distinct capabilities: execution-level capability, iterating within a selected training strategy; and strategy-level capability, revising the high-level judgment as experimental evidence accumulates. Analyzing a large corpus of publicly released post-training trajectories, we find that across different tasks, the agent's training strategy is locked in at the very beginning, and the entire remaining budget is spent on local adjustments within the selected strategy. We then examine three natural explanations--missing experience, missing guidance, and insufficient reasoning--with escalating interventions. Extensive experiments show that (1) an experience-driven scaffold improves execution across the board (+12.6 points on GSM8K and +40.8 on HumanEval) but leaves the strategy static; (2) human guidance effectively redirects the initial strategy, yet the agent falls back into local adjustment loops once training starts; and (3) additional inference compute pays off on easier tasks but yields almost no gain on the hardest one. In conclusion, what agents lack is neither experience, guidance, nor reasoning compute, but a mechanism for spontaneously reevaluating their strategy during execution.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Distributed Online Estimation of Spiked Eigenvalues with Adaptive Weighting under Persistent Aspect Ratio Heterogeneity
Authors:
Lu Yan,
Jiang Hu,
Yonghan Zhang,
Xiaoyue Li
Abstract:
We study online estimation of spiked covariance eigenvalues from observations distributed across $L$ nodes with heterogeneous and persistent effective sample sizes. In the proportional high-dimensional regime, local Rayleigh statistics are deterministically distorted by node-specific aspect ratios $c_{\ell,t}=p/N^{\mathrm{eff}}_{\ell,t}$, and direct aggregation of uncorrected statistics converges…
▽ More
We study online estimation of spiked covariance eigenvalues from observations distributed across $L$ nodes with heterogeneous and persistent effective sample sizes. In the proportional high-dimensional regime, local Rayleigh statistics are deterministically distorted by node-specific aspect ratios $c_{\ell,t}=p/N^{\mathrm{eff}}_{\ell,t}$, and direct aggregation of uncorrected statistics converges to the wrong limit. We propose a correct-then-aggregate framework in which each node removes its deterministic bias via an inverse Rayleigh transfer map, and the server fuses corrected estimates using adaptive soft-max weights based on predictable fluctuation metrics, transmitting only $O(k)$ scalars per active node per round. We establish consistency and asymptotic normality of the global estimator, enabling valid online inference, and derive non-asymptotic bounds quantifying how accuracy improves with the number of nodes and their effective sample sizes. The adaptive weights achieve variance reduction comparable to oracle inverse-variance weighting, confirming the data-driven construction is nearly efficient. Simulation studies validate these properties. An application to cross-venue monitoring of a dominant market factor shows the method tracks systemic risk in real time while substantially reducing communication cost relative to a centralized pooled approach.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
3D trapping of a meta-atom in an intensity minimum
Authors:
Bin Lu,
Adeel Afridi,
Nadine Meyer,
Romain Quidant
Abstract:
High-refractive-index particles have recently attracted a growing interest in optical levitation experiments, offering the ability to further engineer optical forces through electromagnetic Mie resonances. Unlike standard silica particles, which are predominantly trapped in the dipole regime and exhibit trap frequencies mainly determined by material density, resonant meta-atoms formed by high-inde…
▽ More
High-refractive-index particles have recently attracted a growing interest in optical levitation experiments, offering the ability to further engineer optical forces through electromagnetic Mie resonances. Unlike standard silica particles, which are predominantly trapped in the dipole regime and exhibit trap frequencies mainly determined by material density, resonant meta-atoms formed by high-index particles enable qualitatively new trapping behaviors. In this work, we experimentally investigate the trapping of resonant silicon particles in an optical standing wave. A direct comparison of silicon and silica highlights the fundamental differences in their optical force scaling and trapping dynamics. Beyond conventional trapping at intensity-maxima, we demonstrate deterministic and stable three-dimensional trapping of silicon nanoparticles in optical intensity minima, a regime that remains inaccessible for silica particles. Drawing a mesoscopic analogy with blue-detuned atom trapping, our results establish meta-atoms as a versatile approach to further extend the optical manipulation tool box towards accessing novel trapping regimes e.g. in close proximity to a surface.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
The progenitors of $z\gtrsim10$ JWST galaxies in the COLIBRE simulations
Authors:
Evgenii Chaikin,
Andrew Pontzen,
Carlos S. Frenk,
Joop Schaye,
Shengdong Lu,
Robert A. Crain,
Anna Durrant
Abstract:
JWST has revealed a large population of luminous galaxies ($M_{\rm UV}\lesssim -20$) at redshifts $z \gtrsim 10$, widely interpreted as posing a challenge to models of galaxy formation within the $Λ$CDM cosmology. Here, we search for counterparts of the JWST galaxies in the COLIBRE simulations of galaxy formation. Although these simulations have not been tuned to reproduce any $z > 0$ observations…
▽ More
JWST has revealed a large population of luminous galaxies ($M_{\rm UV}\lesssim -20$) at redshifts $z \gtrsim 10$, widely interpreted as posing a challenge to models of galaxy formation within the $Λ$CDM cosmology. Here, we search for counterparts of the JWST galaxies in the COLIBRE simulations of galaxy formation. Although these simulations have not been tuned to reproduce any $z > 0$ observations, we find a population of COLIBRE galaxies with properties similar to those of the JWST galaxies, and trace them to their earliest evolutionary phases, $z\simeq25$, to investigate the onset of galaxy formation. We study the evolution of galaxy stellar masses, sizes, star formation rates, UV magnitudes, metallicities, central black hole masses, and molecular gas and dust content, finding good agreement with observationally inferred properties at $z > 10$, except for UV magnitudes and dust masses, which COLIBRE underpredicts and overpredicts, respectively. Our results indicate that the standard galaxy formation physics and $Λ$CDM cosmology adopted in COLIBRE are sufficient to reproduce a broad range of properties of the most extreme $z > 10$ JWST galaxies - including their compact sizes, stellar masses, gas content, and metallicities. We show that the discrepancies with the UV magnitudes and dust masses can both be attributed to the uncertain rate of grain growth at high redshift, possibly alongside a top-heavy stellar initial mass function. These findings provide strong evidence that the standard cosmological model can naturally explain even the most extreme galaxies in the early Universe.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Exact random covers of metric trees: balanced rounding, duality, and sharp thresholds
Authors:
Qi Wu,
Yong Lu
Abstract:
Norin and Turcotte's asymptotically sharp bound for graph burning [J. Combin. Theory Ser. B 168 (2024), 208--235] led them to an exact random-cover conjecture for finite metric trees. Let $U[0,r]$ be the uniform probability measure on $[0,r]$. They conjectured that every finite metric tree $T$ of length $L\ge2r$ admits a probability measure on $0$-good ball covers whose expected radius measure is…
▽ More
Norin and Turcotte's asymptotically sharp bound for graph burning [J. Combin. Theory Ser. B 168 (2024), 208--235] led them to an exact random-cover conjecture for finite metric trees. Let $U[0,r]$ be the uniform probability measure on $[0,r]$. They conjectured that every finite metric tree $T$ of length $L\ge2r$ admits a probability measure on $0$-good ball covers whose expected radius measure is at most $(L/r)U[0,r]$. We prove the conjecture for every finite metric tree.
We recast the bootstrapping calculation of Norin and Turcotte as a zero-error replacement certificate. The resulting local scale reduction, together with a three-piece decomposition and a macro-recursion, produces a fractional marked-ball cover with the exact radius budget. We then pass from the fractional cover to random finite covers by a compact rounding argument. For metric-tree balls, Tamir's balancedness theorem and standard balanced-matrix ideality provide the finite-dimensional integrality input.
We also prove an arbitrary-budget duality criterion. If $0<R\le L$ and $β$ is a finite positive Borel measure on $[0,R]$, then $β$ dominates the expected radius measure of a random $0$-good cover if and only if $σ(T)\le\int_{[0,R]}\max_{v\in T}σ(B_T(v,s))\,dβ(s)$ for every finite positive Borel measure $σ$ on $T$; it is enough to test finite atomic measures. We use this criterion to extend the uniform range to every $r\le L-\operatorname{diam}(T)/2$, determine the exact range for equal-arm metric stars, and derive deterministic bounds, interval rigidity, and a diameter-defect stability estimate.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Remarkable Enhancement of High Harmonic Generation from Superhard Material under High Pressure
Authors:
Zishao Wang,
Tong Wu,
Ziwen Wang,
Shicheng Liu,
Hui Li,
Kun Zhao,
Jian Sun,
Chao Yu,
Ruifeng Lu
Abstract:
High harmonic generation (HHG) in solids offers a pathway to develop compact extreme ultraviolet (EUV) sources crucial for attosecond science and advanced spectroscopy. Here, we demonstrate theoretically that high pressure dramatically enhances HHG in superhard hexagonal tungsten nitride (h-WN6). Compared with solid-state systems at ambient pressure, the reshaped electronic environment under high…
▽ More
High harmonic generation (HHG) in solids offers a pathway to develop compact extreme ultraviolet (EUV) sources crucial for attosecond science and advanced spectroscopy. Here, we demonstrate theoretically that high pressure dramatically enhances HHG in superhard hexagonal tungsten nitride (h-WN6). Compared with solid-state systems at ambient pressure, the reshaped electronic environment under high pressure leads to a unique band-gap widening in h-WN6, which raises the material's damage threshold, allowing the use of stronger laser fields and enabling access to higher-energy bands. This pressure-induced band-gap widening offers a promising strategy to overcome the cutoff limitation of solid-state EUV light sources.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
Authors:
Qingyao Li,
Wenxiang Jiao,
Shuai Shao,
Kangning Zhang,
Yuan Lu,
Yi Guo,
Weiwen Liu,
Weinan Zhang,
Yong Yu
Abstract:
Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. We show that the default remedy, outcome-rewarded RL over the candidate slate, cannot teach it, for a…
▽ More
Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. We show that the default remedy, outcome-rewarded RL over the candidate slate, cannot teach it, for a structural reason we identify and name selector credit starvation: under a broadcast, sequence-level advantage, the few tokens that name the chosen skill carry a vanishing share of the loss, and the credit they inherit is increasingly wrong-signed as trajectories lengthen. A correct choice is punished whenever the execution after it fails, even though the choice itself is among the most valuable decisions in the trajectory. Auditing a completed run's own training artifacts confirms all three properties, each worsening monotonically with horizon. SkillGate removes the failure by construction: it partitions the token support into two disjoint credit channels, outcome credit reaching only execution tokens, and a separate action-local advantage reaching exactly the skill-naming tokens, positive only when a trajectory's single read is the correct one. On five agentic benchmarks under a 16-candidate slate, SkillGate lifts a 9B policy from 40.8% to 53.2% trial success, well ahead of the identical budget spent on outcome reward alone, while cutting exposure to misleading candidates by two thirds and reading fewer skills.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Study of the intrinsic resolution of LaBr3(Ce,Sr) and NaI(Tl) crystals
Authors:
Peiyi Feng,
Xilei Sun,
Zhenghua An,
Dali Zhang,
Xinqiao Li,
Shaolin Xiong,
Hong Lu
Abstract:
The Gravitational wave burst high-energy Elec?tromagnetic Counterpart All-sky Monitor (GECAM) utilizes a large number of LaBr3 and NaI(Tl) crystals as sensitive materials for its gamma-ray detectors. To address the fitting issues of the energy resolution curves in the ground calibration of the GECAM detectors, this work conducts a comprehensive testing and comparative study of the energy resolutio…
▽ More
The Gravitational wave burst high-energy Elec?tromagnetic Counterpart All-sky Monitor (GECAM) utilizes a large number of LaBr3 and NaI(Tl) crystals as sensitive materials for its gamma-ray detectors. To address the fitting issues of the energy resolution curves in the ground calibration of the GECAM detectors, this work conducts a comprehensive testing and comparative study of the energy resolution of 1-inch LaBr3(Ce,Sr) and NaI(Tl) crystals produced from the same batch. We employed a Hard X-ray Calibration Facility (HXCF), a PMT single-photoelectron calibration system, and Geant4 Monte Carlo simulation tools to quantify seven factors influencing energy resolution. The results indicate that the contributions of various components to energy resolution differ, with pho?toelectron statistical fluctuations and intrinsic resolution being predominant. For 100 keV X-rays, the total energy resolution of the LaBr3(Ce,Sr) crystal is 3.71% \pm 0.03% (expressed as 1-σ), with a contribution from photoelectron statistical fluctuations of 2.69% \pm 0.00% and an intrinsic resolution of 2.45% \pm 0.05%. For 100 keV X-rays, the total energy resolution of the NaI(Tl) crystal is 4.41% \pm 0.14% (expressed as 1-σ), with a contribution from photoelectron statistical fluctuations of 3.20% \pm 0.00% and an intrinsic resolution of 2.90% \pm 0.21%. We discussed the sources of intrinsic resolution, and the results indicate that the intrinsic resolution of the LaBr3(Ce,Sr) crystal primarily arises from luminescence non-proportionality, while that of the NaI(Tl) crystal mainly stems from fluctuations during energy transfer. This study emphasizes precise and specific experimental measurements and comparative research, demonstrating that both factors are important contributors to intrinsic resolution.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization
Authors:
Chenle Chen,
Yangbo Wei,
Chao Yao,
Shaoqiang Lu,
Junhong Qian,
Chen Wu,
Lei He
Abstract:
Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document, accepting each candidate through a validation gate. Existing gates rely on verifiable rewards, confining these methods to tasks with an automatic verifier. Replacing the verifier with an LLM-judge gate would lift that restriction, but whether such a gate carries usable signal is untested. We ask a pr…
▽ More
Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document, accepting each candidate through a validation gate. Existing gates rely on verifiable rewards, confining these methods to tasks with an automatic verifier. Replacing the verifier with an LLM-judge gate would lift that restriction, but whether such a gate carries usable signal is untested. We ask a prior question: can we tell, before placing a judge in the loop, whether its scores separate correct from incorrect answers at all? We formalize a reference-free judge as a latent solver -- its verdict rests on agreement with whatever it would itself conclude, so its capacity to evaluate is bounded by its capacity to solve. The model yields a closed-form bound on discriminability (ROC-AUC) in the judge's competence $c$ and answer-space size $k$, a necessary condition $c > 1/k$, and the result that the marginal AUC is confounded by item difficulty while a within-question estimator is not. A non-intervening probe records judge scores on genuine optimization runs without altering any decision. We find discriminability at chance where competence sits near the floor and usable above it; that a judge's benchmark accuracy overstates the competence that matters; and, in a closed-loop study, that the screen predicts which kind of gating error occurs. The result is a cheap pre-deployment diagnostic for judge gates.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
DocClaw: A Unified Agentic System for Intelligent Document Processing
Authors:
Siqi Xiang,
Zhipeng Xu,
Yufei Liu,
Junhao Ji,
Qing Liu,
Zulong Chen,
Zhibo Yang,
Chunyan Miao,
Shijian Lu
Abstract:
Intelligent document processing (IDP) encompasses a broad range of tasks, including optical character recognition (OCR), document question answering (DocQA), and key information extraction (KIE). Despite their distinct objectives, these tasks share a common need to perceive document content, acquire task-relevant information, and progressively refine intermediate results. However, they are typical…
▽ More
Intelligent document processing (IDP) encompasses a broad range of tasks, including optical character recognition (OCR), document question answering (DocQA), and key information extraction (KIE). Despite their distinct objectives, these tasks share a common need to perceive document content, acquire task-relevant information, and progressively refine intermediate results. However, they are typically formulated as separate prediction problems and addressed by task-specific models or processing pipelines. We introduce DocClaw, a unified agentic system that formulates diverse intelligent document processing tasks as a shared process of interaction between an agent and a document. Given a document and a task-specific query, DocClaw follows an appropriate document skill to iteratively identify the information required, invoke relevant tools, and integrate the resulting observations into the desired output. Throughout this process, a structured document state organizes reusable document knowledge and task-specific interaction context, allowing the agent to accumulate, revisit, and progressively refine information as the interaction proceeds. Under this formulation, task-specific requirements are captured by the agent's interpretation of the query objective and the corresponding document skill, while the underlying interaction loop, tool space, and document state are shared across tasks. Extensive experiments across multiple intelligent document processing benchmarks demonstrate that DocClaw effectively handles diverse tasks within a single agentic framework and achieves competitive performance compared with both general-purpose VLMs and task-specific methods.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
ReX-Shot: Single-Image Rephotography via Geometry- and Camera-Grounded Generation
Authors:
Ruiqi Zhang,
Hao Zhu,
Wenhao Zhang,
Qi Zhang,
Junqi Shi,
Ming Lu,
Xun Cao,
Zhan Ma
Abstract:
Single-image rephotography aims to synthesize new shots of a scene from a single reference image with specified viewpoints, focal lengths, and photographic effects, which are intrinsically coupled in imaging. Existing methods typically treat these factors separately and struggle under joint control: novel-view synthesis may introduce geometric distortions under focal-length changes, while super-re…
▽ More
Single-image rephotography aims to synthesize new shots of a scene from a single reference image with specified viewpoints, focal lengths, and photographic effects, which are intrinsically coupled in imaging. Existing methods typically treat these factors separately and struggle under joint control: novel-view synthesis may introduce geometric distortions under focal-length changes, while super-resolution and instruction-guided editing remain confined to 2D and cannot reliably extend detail restoration or appearance control to novel viewpoints. We attribute these limitations to imperfect single-image 3D reconstruction and the sampling limit of continuous focal-length enlargement. To reduce projection bias from geometric errors, we use implicitly transformed foundation-model features for robust target-view guidance. We further formulate focal-length enlargement as a geometry-guided super-resolution problem and exploit generative detail priors to recover details lost during sparse 3D resampling. Built on this 3D-aware generative backbone, we lift photographic-effect control from 2D filtering to 3D-aware appearance editing, preserving content consistency across viewpoints and focal lengths. These components form ReX-Shot, a geometry- and camera-grounded generative framework for single-image rephotography. To our knowledge, ReX-Shot is the first unified framework to jointly control viewpoint, focal length, and parameterized photographic effects from a single image. Experiments show that ReX-Shot outperforms representative baselines across all three controls while enabling near-real-time interactive rephotography.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
First Law of Black Hole Interior Dynamics
Authors:
Ze-Xuan Xiong,
H. Lu
Abstract:
We obtain a local first law of dynamics within the interior geometries that are governed by cosmological solutions connecting the event horizon and the Kasner-like singularity. Evaluating the Iyer-Wald identity between the horizon and near-singularity geometries, we define a Kasner potential and transported response coefficients. We test the first law by both exact and numerical solutions. Our for…
▽ More
We obtain a local first law of dynamics within the interior geometries that are governed by cosmological solutions connecting the event horizon and the Kasner-like singularity. Evaluating the Iyer-Wald identity between the horizon and near-singularity geometries, we define a Kasner potential and transported response coefficients. We test the first law by both exact and numerical solutions. Our formalism provides a self-consistent approach to study the interior dynamics without having to appeal to the outside geometry and the asymptotic charges. It may provide a new approach to study the phase structures of the interior dynamics.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Online Service with Per-Batch Maximum Delay
Authors:
Tianhang Lu,
Runtian Ren,
Shengcai Liu,
Ke Tang
Abstract:
We study online service with one maximum-waiting-time charge per service batch. Requests arrive at points of a finite metric, and a mobile server pays for its movement and, for each service walk, the maximum waiting time among the requests served by that walk. We distinguish elective service, where an encountered request may be left pending, from automatic service, where every encounter serves it.…
▽ More
We study online service with one maximum-waiting-time charge per service batch. Requests arrive at points of a finite metric, and a mobile server pays for its movement and, for each service walk, the maximum waiting time among the requests served by that walk. We distinguish elective service, where an encountered request may be left pending, from automatic service, where every encounter serves it. Although the two semantics have different optimal schedule structures, we prove that their offline optimal values are equal. On a finite line and on an explicitly represented weighted tree, the common offline value is computable by polynomial-time dynamic programming, whereas exact optimization on arbitrary finite metrics is NP-hard.
For the online problem, we prove a metric-independent group-trajectory certificate lemma that charges spatially separated request groups to two parity classes of time windows. It yields deterministic polynomial-time competitive ratios 10 on a line, 12 on a weighted tree, and 20 on an arbitrary finite metric, under both service semantics. With an exact metric-Steiner-tree oracle, the general-metric ratio improves to 12. The polynomial algorithm uses a half-scaled running maximum of terminal-MST weights; the running maximum is necessary because terminal MST weight is not monotone under new arrivals. A fixed two-point line gives a deterministic visible-service lower bound of 3 for every metric class above. Finally, when request locations are hidden until visited, dyadic exploration is 84-competitive on a known finite line. This phenomenon is line-specific: one hidden request gives deterministic and randomized lower bounds 3 and 2 on a line, while a d-leaf unit star gives lower bounds 2d-1 and d.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
MARCUS: Missing-Aware Region Representation with Contextual Urban Signals for Rent Prediction
Authors:
Chenya Huang,
Bin Liang,
Zhidong Li,
Yuxi Lu,
Kunqi Li,
Justin Wang,
Fang Chen
Abstract:
Multimodal urban data has expanded the applications of urban region representation learning, such as functional zone identification and real estate appraisal, but also introduces challenges caused by data incompleteness. Existing studies usually handle missing data through imputation, treating missingness as noise while ignoring its potential semantic value. To address this issue, we propose MARCU…
▽ More
Multimodal urban data has expanded the applications of urban region representation learning, such as functional zone identification and real estate appraisal, but also introduces challenges caused by data incompleteness. Existing studies usually handle missing data through imputation, treating missingness as noise while ignoring its potential semantic value. To address this issue, we propose MARCUS, a missing-aware region representation model that treats missingness as a contextual urban signal. MARCUS models missingness in three stages: Intra Learning jointly encodes observed features and missing patterns, Inter Learning estimates modality reliability to guide cross-modal interaction, and Fusion uses missing-aware and time-aware gating to generate the final region embedding. We apply MARCUS to rent prediction, a task with long-term trends and seasonal fluctuations, using real-world datasets from Sydney and New York. Experimental results show that MARCUS achieves state-of-the-art performance, reducing MAE by 51.35% on Sydney and 12.62% on New York compared with the best baselines. Additional experiments, including an imputation-based ablation study and randomized additional-missingness analysis, further demonstrate the effectiveness of the proposed method.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Beyond Idealized PAHs: Infrared Signatures of Carbon-Chain Defects from Shock Synthesis
Authors:
Xiaoting Tan,
Zhao Wang,
Zu-Jia Lu,
Houda Haidar,
Dimitra Rigopoulou
Abstract:
Polycyclic aromatic hydrocarbons (PAHs) are widely recognized as carriers of the aromatic infrared bands (AIBs). However, most spectral models rely on idealized structures that fail to capture the energetic environments of interstellar PAH formation. This work investigates the infrared (IR) signatures of PAHs formed under shock conditions and explores whether produced defective structures can expl…
▽ More
Polycyclic aromatic hydrocarbons (PAHs) are widely recognized as carriers of the aromatic infrared bands (AIBs). However, most spectral models rely on idealized structures that fail to capture the energetic environments of interstellar PAH formation. This work investigates the infrared (IR) signatures of PAHs formed under shock conditions and explores whether produced defective structures can explain observational features unpredicted by standard, idealized models. We combine two-stage reactive molecular dynamics simulations of PAH formation via condensation and shock processing with density functional theory spectral calculations, and compare our theoretical results with James Webb Space Telescope (JWST) observations of NGC 7023 and MRK 1066. Shock processing produces PAHs featuring fullerene-like carbon skeletons and linear carbon-chain attachments. These structural defects yield distinct IR signatures, including prominent carbon-chain stretching features at 4.6-5.5 micron that is absent in ideal PAHs, and significantly enhanced out-of-plane skeletal modes in the 14.5-20 micron regime. Our findings attribute the observed 5.2 micron band in NGC 7023 and MRK 1066 to carbon-chain vibrations and the 15-18 micron emission to curved skeletal modes, providing observational support for the prevalence of defective, shock-formed PAHs in the interstellar medium.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
MissDiag: Diagnostic Evaluation of Incomplete-Knowledge Robustness in KGQA and KG-RAG
Authors:
Hang Wang,
Hang Dong,
Lu Liu,
Chuanru Ren
Abstract:
Knowledge graph question answering (KGQA) and knowledge-graph-based retrieval-augmented generation (KG-RAG) aim to ground answers in explicit graph evidence, but real-world knowledge graphs are often sparse, outdated, and incomplete. Existing robustness evaluations usually report aggregate changes in answer quality after evidence is removed or perturbed, which measures sensitivity to incomplete su…
▽ More
Knowledge graph question answering (KGQA) and knowledge-graph-based retrieval-augmented generation (KG-RAG) aim to ground answers in explicit graph evidence, but real-world knowledge graphs are often sparse, outdated, and incomplete. Existing robustness evaluations usually report aggregate changes in answer quality after evidence is removed or perturbed, which measures sensitivity to incomplete support but leaves the source of degradation under-specified: the same score change can conflate the type of missing evidence, the response of the evaluated system, and the sensitivity of the answer-matching protocol. To address this gap, we propose \textbf{MissDiag}, a diagnostic evaluation framework for incomplete-knowledge robustness in KGQA and KG-RAG. MissDiag keeps the question and gold answer fixed while applying structurally typed missingness interventions to benchmark-provided support graphs, enabling paired comparisons that decompose robustness changes by evidence type, system response, and evaluation protocol rather than reducing them to a single aggregate score drop. Experiments across multiple system families show that incomplete-knowledge robustness is better understood as a typed degradation phenomenon than as a uniform property: answer-adjacent evidence loss produces the largest observed degradation, source-context removal is often neutral and can be beneficial, and semantic answer matching changes absolute scores while preserving the main typed degradation patterns. By transforming aggregate robustness measurement into typed diagnostic attribution, MissDiag provides a more interpretable basis for comparing, diagnosing, and stress-testing KGQA and KG-RAG systems under incomplete knowledge.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Verblunsky coefficients, CMV matrices and numerical invariants of homogeneous bidisc submodules
Authors:
Yufeng Lu,
Chao Zu
Abstract:
Let $[p]$ be the principal submodule generated by a polynomial in $H^2(\mathbb D^2)$. For homogeneous $p$, the homogeneous slices of $[p]$ admit a weighted OPUC model in which the two wandering vectors are an orthonormal polynomial and its reversal. We show that the associated Verblunsky coefficients determine the singular values of the wandering-projection product and the restricted cross-commuta…
▽ More
Let $[p]$ be the principal submodule generated by a polynomial in $H^2(\mathbb D^2)$. For homogeneous $p$, the homogeneous slices of $[p]$ admit a weighted OPUC model in which the two wandering vectors are an orthonormal polynomial and its reversal. We show that the associated Verblunsky coefficients determine the singular values of the wandering-projection product and the restricted cross-commutator, as well as the non-zero spectrum of the core operator. Toeplitz-determinant and Mahler-measure identities yield exact Fredholm determinants and Schatten estimates, while $[(z-w)^N]$ rules out a uniform Hilbert--Schmidt bound. The same model gives explicit singular values of $[S_z^*,S_w]$ on the homogeneous quotient $H^2(\mathbb D^2)\ominus[p]$; for $p=(z-w)^N$, its squared Hilbert--Schmidt norm is asymptotic to $N$. For arbitrary polynomial generators, we construct a weighted bivariate model with a doubly Toeplitz, block-banded moment matrix and prove $C_p^2|_{\mathscr E_z}=Γ_p^*Γ_p$, relating the core spectrum to the cross-Gram operator between the two edge spaces. We also discuss cyclic-factor obstructions, represent higher numerical invariants by alternating CMV products, and give a quadratic counterexample to their proposed monotonicity.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Depth Anything V4: Dynamic 4D Scene Reconstruction via Riemannian Flow Matching on 4D Gaussian Splatting
Authors:
Jiaming Fan,
Jian Lu,
Jinling Jia,
Chenbin Zhang
Abstract:
We present Depth Anything V4 (DAV4), a framework for dynamic 4D scene reconstruction from monocular video. Our key contribution is the application of Riemannian Flow Matching (RFM) to 4D Gaussian Splatting parameters, defining probability paths directly on non-Euclidean manifolds (scale, rotation, opacity), ensuring all intermediate states are valid. Through controlled experiments, we isolate RFM'…
▽ More
We present Depth Anything V4 (DAV4), a framework for dynamic 4D scene reconstruction from monocular video. Our key contribution is the application of Riemannian Flow Matching (RFM) to 4D Gaussian Splatting parameters, defining probability paths directly on non-Euclidean manifolds (scale, rotation, opacity), ensuring all intermediate states are valid. Through controlled experiments, we isolate RFM's contribution from test-time optimization (TTO) and pre-training. A deterministic MLP baseline with the same data, architecture, and TTO achieves F-score 0.762; RFM achieves 0.806 - the +0.044 gain is RFM's isolated contribution. We provide corrected computational cost analysis: pre-training is 360 GPU-hours, amortizing for large-scale deployment (over 10,000 scenes). Uncertainty is quantified via Negative Gaussian Log-Likelihood and Expected Calibration Error. DAV4 outperforms prior Depth Anything models and per-scene 4D-GS on dynamic reconstruction and novel-view synthesis, while using no human-annotated depth labels as training losses.
△ Less
Submitted 20 August, 2026; v1 submitted 18 August, 2026;
originally announced August 2026.
-
GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction
Authors:
Ziyang Cheng,
Tianshu Tang,
Jinxin Lan,
Xinze Chen,
Yuhan Gong,
Zhichao Liu,
Changzhong Wu,
Yahao Mao,
Zongyan Deng,
Mingxuan Ma,
Huasen Xi,
Yilong Liu,
Yutong Wu,
Xiaofeng Wang,
Yang Wang,
Yun Ye,
Guan Huang,
Xiaojie Jin,
Zheng Zhu,
Jiwen Lu
Abstract:
Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their…
▽ More
Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging the reference-motion corpus, which stops working once feasible behaviors become environment-dependent. We present GigaBrain-WBC-0.5, the first Behavior World Model (BWM) for humanoid whole-body control. Rather than a purely reactive tracker, we train a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next. An automatic terrain-annotation pipeline recovers full 3D contact geometry from retargeted motion, enabling terrain annotation at the scale of existing motion datasets. The predicted distribution is reused at deployment to detect implausible commands online and retract them onto learned behaviors, so the robot attempts tasks in a "best-effort" manner. The result is a unified policy that takes real-time command, interacts with environment, and stays robust to implausible commands, falls, and disturbances. GigaBrain-WBC-0.5 achieves the highest success rate across all four regimes among three large-scale tracker baselines: 81.3% on terrain interaction (4.3x the strongest baseline), 83.1% under implausible commands, and 99.3% fall recovery (16.8x the strongest baseline). Hardware trials show robust interaction under missing supports and disturbances; the Unitree G1 checkpoint transfers to the Maker L01 robot with simple fine-tuning.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud
Authors:
Dheeraj Mohandas Pai,
Lu Xian
Abstract:
Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. Banking is the clearest case: the same agent that answers a question can also change contact details, reset a PIN, or move money, so ordinary customer service is inseparable from authorization, fraud detection, and policy…
▽ More
Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. Banking is the clearest case: the same agent that answers a question can also change contact details, reset a PIN, or move money, so ordinary customer service is inseparable from authorization, fraud detection, and policy compliance. Existing financial-fraud benchmarks classify static transactions or messages, and general agent-safety benchmarks target prompt injection or generic harmful use; none test whether a policy-grounded banking agent safely acts when a caller manipulates identity, authorization, and trust over a conversation. We introduce FraudBench, an executable benchmark built on the $τ^2$-bench dual-control framework and the $τ$-Knowledge banking environment. Both the agent and the simulated caller act through tools over shared, mutable account state, and the agent may grant the caller access to selected tools; the environment exposes a 698-document internal policy corpus that the agent must retrieve from. FraudBench contains 150 authored adversarial scenarios; a frozen public set of 107 (90 across ten fraud mechanisms plus 17 chained adaptive attacks) is used for all reported runs, with 43 further chained attacks held out. Safety is history-dependent: single-control tasks satisfy every precondition but one, and adaptive attacks make a later, locally valid request unsafe because of an earlier probe, admission, or failed attempt. Each scenario is annotated with observable evidence, prohibited actions, safe dispositions, and intervention points. A preliminary single-trial evaluation of four agents on the 107 graded tasks yields attack-security between 49\% and 65\%, with money-mule and first-party fraud the most common cross-model weaknesses.
△ Less
Submitted 2 August, 2026;
originally announced August 2026.
-
DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models
Authors:
Wenxin Duan,
Hanwei Wang,
Zhongying Peng,
Zhonghua Lu,
Jiayi An,
Fan Song,
Yong Liang
Abstract:
Background: Mechanistic elucidation of traditional Chinese medicine (TCM) compound formulas remains a central challenge in the modernization of TCM. Conventional approaches, including data mining and network pharmacology, are insufficient for achieving deep integration between classical TCM theory and modern scientific research. In addition, direct question-answering using general-purpose artifici…
▽ More
Background: Mechanistic elucidation of traditional Chinese medicine (TCM) compound formulas remains a central challenge in the modernization of TCM. Conventional approaches, including data mining and network pharmacology, are insufficient for achieving deep integration between classical TCM theory and modern scientific research. In addition, direct question-answering using general-purpose artificial intelligence large language models is limited by inadequate adaptation to TCM theoretical frameworks and susceptibility to reasoning hallucinations. Consequently, there is an urgent need to develop intelligent analytical methods aligned with the holistic principles of TCM. Objective: To establish a multi-expert intelligent agent framework integrating classical TCM theory with modern life sciences, thereby enabling systematic and interpretable mechanistic analysis of TCM compound formulas, with Guizhi Decoction serving as a representative validation case. Methods: The DeepTCM1.0 framework was constructed based on the general-purpose large language model DeepSeek V3.2. It adopts a three-tier collaborative architecture and a three-round iterative quality-control workflow, simulating the collaborative analytical process of 11 interdisciplinary intelligent agents. The framework was applied to the mechanistic interpretation of Guizhi Decoction from the dual perspectives of classical traditional Chinese medicine theory and modern scientific research. Framework performance was comprehensively evaluated through double-blind five-dimensional scoring, intraclass correlation coefficient (ICC) reliability testing, Mann-Whitney U tests, and effect size analysis. The evaluation employed four independent large language models as evaluators, each conducting five rounds of repeated scoring on five anonymized reports, resulting in a total of 100 independent scoring assessments.
△ Less
Submitted 9 June, 2026;
originally announced August 2026.
-
LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization
Authors:
Ruizhi Zhang,
Jinwei Chen,
Xiangju Lu,
He Yan,
Mo Yu,
Junmin Zhu,
Wei Zhang
Abstract:
Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi-scale benchmark for hallucination detection in long…
▽ More
Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucinations change as the context grows longer. In this study, we propose LongNovel, a multi-scale long-context bilingual (Chinese and English) novel benchmark for hallucination detection. This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter-level data from the BookSum dataset. We design 8 hallucination types and employ a combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories. Furthermore, we manually revise the content in the test set to guarantee data reliability. Extensive experimental results demonstrate that LongNovel is a challenging benchmark. We release LongNovel for future research. https://github.com/BDML-lab/LongNovel
△ Less
Submitted 4 June, 2026;
originally announced August 2026.
-
Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges
Authors:
Yisong Chen,
Yifan Gao,
Sijing Yu,
Chuqing Zhao,
Yang Lu
Abstract:
We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational agents, therapy support tools, prompt engineering, multimodal learning, and ethical considerations. We integrate findings from interdisciplinary studies utilizing diverse data sources such as social media posts, electronic medical records, and multimodal inputs to…
▽ More
We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational agents, therapy support tools, prompt engineering, multimodal learning, and ethical considerations. We integrate findings from interdisciplinary studies utilizing diverse data sources such as social media posts, electronic medical records, and multimodal inputs to enable early detection of depression, suicide risk assessment, personalized therapy support, and psychoeducational content generation. Our review highlights advancements in LLM models and annotation strategies that enhance interpretability and clinical relevance, while we also emphasize the critical role of prompt engineering for domain adaptation. We also discuss emerging multimodal fusion techniques integrating text, speech, and sensor data for improved mental health diagnosis and monitoring. Finally, we address ongoing ethical, sociotechnical, and regulatory challenges, and advocate frameworks to ensure safe, equitable, and accountable deployment of LLMs in real-world mental health care.
△ Less
Submitted 30 May, 2026;
originally announced August 2026.
-
Unique Ergodicity for the Projective Process of the 2D Navier--Stokes Equation with Nondegenerate Noise
Authors:
Zeng Lian,
Rongchang Liu,
Kening Lu
Abstract:
We prove unique ergodicity of the projective process associated with the two dimensional Navier--Stokes equation in vorticity form, with additive diagonal noise acting on every nonzero real Fourier phase and satisfying two sided power law bounds. Consequently, the exact Furstenberg--Khasminskii formula for the top Lyapunov exponent holds.
The main new ingredient is a compact dense mechanism for…
▽ More
We prove unique ergodicity of the projective process associated with the two dimensional Navier--Stokes equation in vorticity form, with additive diagonal noise acting on every nonzero real Fourier phase and satisfying two sided power law bounds. Consequently, the exact Furstenberg--Khasminskii formula for the top Lyapunov exponent holds.
The main new ingredient is a compact dense mechanism for asymptotic generalized coupling in the absence of a Foiaş--Prodi type high-low mode decomposition for the projective dynamics. Using the dense range of the Malliavin derivative and compactness of the state derivative, we construct finite rank perturbations of the Wiener path that compensate, to first order, for perturbations of the initial condition, leaving a residual whose logarithmic growth has negative stationary mean and hence contracts locally at an exponential rate, with a cost controlled by a triangular scheme of blockwise Ramer transformations.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach
Authors:
Lu Xu,
Xu Li,
Linjiang Zheng,
Fan Li,
Riquan Zhang,
Jiaxing Shang
Abstract:
Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models…
▽ More
Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models (LLMs), which excel at language reasoning, bring a promising solution to this issue. However, applying LLMs in this domain presents key challenges such as modal inconsistency, limited classification ability, scarcity of task-specific data for fine-tuning, and lack of domain knowledge. To overcome these challenges, we propose FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis. Specifically, we first perform feature engineering to address modal inconsistency, combining statistical descriptors with physically meaningful flight indicators. This representation is further processed by a Semantic Discretization module, which converts abstract numerical patterns into qualitative descriptions that are more compatible with language reasoning. In addition, since LLMs are not inherently strong classifiers, CatBoost is incorporated as a statistical expert, and its prediction results are injected into the prompt as prior guidance. A contrastive few-shot learning strategy is further adopted to compensate for limited data. Finally, we design structured prompts to embed aviation-specific knowledge into the inference process. Using hard landing, a representative risk event with complex causal mechanisms, as an anchor point, we evaluate FlightLLM on a dataset of 704 real-world A320 flight samples. Experimental results show that the proposed approach achieves competitive classification performance while generating direct and reasonable explanations for event causes.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering
Authors:
Hsiang-Wei Huang,
Fu-Chen Chen,
Li-Wu Tsao,
Cheng-Han Lee,
Che-Chun Su,
Lu Xia,
Ronghui Peng,
Jenq-Neng Hwang,
Min Sun,
Cheng-Hao Kuo
Abstract:
Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key frames for VLM input. However, visual search methods are inefficient because they require visual searc…
▽ More
Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key frames for VLM input. However, visual search methods are inefficient because they require visual search among thousands of video frames for each individual user query. In this work, we propose a memory tree guided key frame selection paradigm for efficient 3D question answering in embodied scenarios. Our method leverages a compact and reusable 3D scene representation, termed MemTree3D, which supports real-time online construction leveraging camera 6-DoF poses. MemTree3D captures multi-level 3D scene information, enabling a Large Language Model to efficiently query and retrieve question-relevant key frames through our scoring-based frame selection without reprocessing the entire video stream. On OpenEQA, our method improves the LLM-Match of GPT-4o by 17.4%, LLaVA-OneVision-7B by 5.8%, outperforms existing visual search methods. Our code is available at https://github.com/hsiangwei0903/MemTree3D
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Traceable Trust for action-ready artificial intelligence in bioscience
Authors:
Huayu Xin,
Yizhi Cai,
Mukilan Deivarajan Suresh,
Gavin Michael Farrell,
Iwona Gajda,
Charlie Harrison,
Conor Houghton,
Mato Lagator,
Yang Lu,
Virginia Portillo,
Reyer Zwiggelaar,
Sebastian Lobentanzer
Abstract:
Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structures, design proteins, rank variants, annotate images, recommend strains and optimise experimental conditions. We argue that the decision to use an AI output to guide laboratory action is a key juncture for trustworthy research and should follow a defined, review…
▽ More
Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structures, design proteins, rank variants, annotate images, recommend strains and optimise experimental conditions. We argue that the decision to use an AI output to guide laboratory action is a key juncture for trustworthy research and should follow a defined, reviewable process. We propose Traceable Trust as a proportionate assessment-and-design framework for this output-to-action boundary. It asks what evidence supports the output, what capability is being claimed, what agency has been delegated, what threshold authorises action, who can override it and how outcomes inform later decisions. We illustrate the framework through three case studies spanning ecosystem resources, project design and laboratory action. Together, the cases show how trust can be documented where AI outputs begin to shape scientific work.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection
Authors:
Bin Li,
Dongdong Wang,
Siyang Lu
Abstract:
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe clas…
▽ More
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conventional calibration metrics indicate good calibration, creating a critical reliability gap for operational monitoring systems. To address this issue, we propose Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection. LoRD learns prediction-route-specific reliability models from latent representations of correctly classified validation samples and estimates prediction reliability through route-wise reconstruction distances. Based on the estimated reliability, LoRD selectively recalibrates high-risk predictions to suppress overconfident errors while preserving reliable predictions. Extensive experiments on four large-scale log benchmark datasets and multiple language model-based detectors demonstrate that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.
△ Less
Submitted 19 August, 2026; v1 submitted 18 August, 2026;
originally announced August 2026.
-
Search for $B$ meson decays to multimuon final states
Authors:
LHCb collaboration,
R. Aaij,
M. Abdelfatah,
A. S. W. Abdelmotteleb,
C. Abellan Beteta,
F. Abudinén,
T. Ackernley,
A. A. Adefisoye,
B. Adeva,
M. Adinolfi,
P. Adlarson,
C. Agapopoulou,
C. A. Aidala,
S. Akar,
K. Akiba,
P. Albicocco,
J. Albrecht,
R. Aleksiejunas,
F. Alessio,
P. Alvarez Cartelle,
S. Amato,
J. L. Amey,
Y. Amhis,
L. An,
L. Anderlini
, et al. (1109 additional authors not shown)
Abstract:
A search for decays of $B$ mesons to final states with four or six muons using $pp$ collision data recorded by the LHCb experiment corresponding to an integrated luminosity of $5.4~\text{fb}^{-1}$ is presented. The decay modes of interest are $B_{(s)}^0 \rightarrow μ^+μ^-μ^+μ^-$, $B^+ \rightarrow K^+μ^+μ^-μ^+μ^-$, $B_{(s)}^0 \rightarrow μ^+μ^-μ^+μ^-μ^+μ^-$ and…
▽ More
A search for decays of $B$ mesons to final states with four or six muons using $pp$ collision data recorded by the LHCb experiment corresponding to an integrated luminosity of $5.4~\text{fb}^{-1}$ is presented. The decay modes of interest are $B_{(s)}^0 \rightarrow μ^+μ^-μ^+μ^-$, $B^+ \rightarrow K^+μ^+μ^-μ^+μ^-$, $B_{(s)}^0 \rightarrow μ^+μ^-μ^+μ^-μ^+μ^-$ and $B^+ \rightarrow K^+μ^+μ^-μ^+μ^-μ^+μ^-$, proceeding via both prompt and long-lived intermediate particles. No evidence for any of the signal modes is found, and upper limits spanning the range of $0.6\times10^{-9}$ to $5.4\times10^{-7}$ at the $95\%$ confidence level are set on their branching fractions, depending on the intermediate-particle masses and lifetimes. In addition, mass-integrated limits across the intermediate-particle lifetime ranges considered in this analysis are determined.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
Authors:
Maosen Zhang,
Jianshuo Dong,
Boting Lu,
Wenyue Li,
Xiaoping Zhang,
Tianwei Zhang,
Jie Zhang,
Han Qiu
Abstract:
LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts alongside user queries creates an attack surface: adversarial inputs can induce models to disclose them. Prior probing studies suggest that leakage-related signals emerge in hidden states, yet the need to extract these states pose…
▽ More
LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts alongside user queries creates an attack surface: adversarial inputs can induce models to disclose them. Prior probing studies suggest that leakage-related signals emerge in hidden states, yet the need to extract these states poses additional deployment challenges. In this paper, we explore whether this internal signal leaves a more accessible ``tell'' before decoding. We propose LeakGauge, which probes this response by appending a suffix that gauges leakage behavior and mapping its prefill token probabilities to an attack-risk score. While a direct gauge uses the initial tokens of confidential content, we find that a content-agnostic one that verbalizes leakage behavior yields more robust signals. Across 11 LLMs, including GLM-5.2 (753B) and Kimi-K3 (2.8T), LeakGauge reaches an AUROC range of 0.944--0.996 on unseen attacks. The signal remains stable when the content changes language or the attack shifts from verbatim to semantic disclosure. By activation-steering interventions, we further show that the risk score is sensitive to an internal leakage-related direction, relating the observable signal to the model's internal representation. In addition, LeakGauge enables an input detector with fewer than 0.5K extra parameters and added latency of 10.34 ms. Code: \href{https://github.com/yeasen-z/LeakGauge}.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Strict Monotonicity of Numerical Invariants for the Submodules $[(z-w)^k]$ in $H^2(\mathbb D^2)$
Authors:
Yin Liu,
Yufeng Lu,
Chao Zu
Abstract:
For $k\geq1$, let $M_k=[(z-w)^k]\subset H^2(\mathbb D^2)$. We first determine the banded Toeplitz matrices associated with the homogeneous components of $M_k$, together with explicit formulas for their determinants and the relevant algebraic cofactors. These formulas lead to a complete description of the spectrum of the core operator: \[ σ(C_{M_k}) =
\{0,1\} \cup \left\{ \pm\frac{k}{n+k}:n\geq1 \r…
▽ More
For $k\geq1$, let $M_k=[(z-w)^k]\subset H^2(\mathbb D^2)$. We first determine the banded Toeplitz matrices associated with the homogeneous components of $M_k$, together with explicit formulas for their determinants and the relevant algebraic cofactors. These formulas lead to a complete description of the spectrum of the core operator: \[ σ(C_{M_k}) =
\{0,1\} \cup \left\{ \pm\frac{k}{n+k}:n\geq1 \right\}. \] In particular, the spectral data determine the parameter $k$.
The determinant and cofactor formulas further yield a unified finite-sum representation for $α_{n,j}^{(k)} =\langle w^jφ_n,z^jψ_n\rangle$, and hence for Yang's higher numerical invariants. We derive an adjacent relation connecting $α_{n,j}^{(k)}$ and $α_{n,j+1}^{(k)}$ by means of an explicit telescoping certificate, and show that the corresponding finite-section transformations are strict contractions. Combining these finite-dimensional estimates with the asymptotic behavior of $α_{n,j}^{(k)}$, we prove the strict monotonicity \[ Σ_0(M_k)> Σ_1(M_k)> Σ_2(M_k)> \cdots . \] The cases $k\geq3$ constitute the new part of the analysis, while the previously known cases $k=1,2$ are recovered within the same framework. Consequently, Yang's monotonicity conjecture holds in strict form for the entire family $\{[(z-w)^k]:k\geq1\}$.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Offline Multi-Agent Reinforcement Learning with a Physics-Informed World Model for Cooperative Mixed Traffic Control
Authors:
Lu Liu,
Chi Xie,
Xi Xiong
Abstract:
This study investigates cooperative control of connected and automated vehicles (CAVs) at partially observable highway bottlenecks in mixed traffic, aiming to mitigate congestion without relying on complete global traffic states or online trial-and-error. We propose a physics-informed world model-based offline multi-agent reinforcement learning framework that reconstructs a physically interpretabl…
▽ More
This study investigates cooperative control of connected and automated vehicles (CAVs) at partially observable highway bottlenecks in mixed traffic, aiming to mitigate congestion without relying on complete global traffic states or online trial-and-error. We propose a physics-informed world model-based offline multi-agent reinforcement learning framework that reconstructs a physically interpretable global traffic state from local CAV observation-action histories, with coupled macroscopic-microscopic traffic dynamics providing physics-based supervision. A probabilistic ensemble world model learns traffic-state transitions and system rewards, while model disagreement quantifies epistemic uncertainty. Multi-step imagined rollouts with pessimistic rewards and uncertainty-driven truncation are then used for offline policy learning. Experiments in a SUMO-based on-ramp bottleneck using approximately $1\times10^6$ offline transitions show that physics supervision improves state reconstruction and world-model prediction accuracy.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Bright dual-pulse betatron X-ray generation from a laser wakefield accelerator
Authors:
Bo Guo,
Yang Wan,
Shuang Liu,
Xiaonan Ning,
Jianfei Hua,
Wei Lu
Abstract:
Pump-probe experiments using dual ultrashort X-ray pulses provide unique opportunities for resolving non-equilibrium dynamics initiated by intense X-ray excitation. Betatron radiation from laser wakefield accelerators offers femtosecond duration, micrometer-scale source size, and intrinsic synchronization with the driving laser, making it a promising candidate for compact ultrafast X-ray sources.…
▽ More
Pump-probe experiments using dual ultrashort X-ray pulses provide unique opportunities for resolving non-equilibrium dynamics initiated by intense X-ray excitation. Betatron radiation from laser wakefield accelerators offers femtosecond duration, micrometer-scale source size, and intrinsic synchronization with the driving laser, making it a promising candidate for compact ultrafast X-ray sources. Here, we experimentally demonstrate a high-flux, dual-pulse betatron X-ray source based on a density-tailored gas-mixture target. Two electron bunches are generated within a single plasma wakefield through ionization-induced and shock-front-triggered injection, subsequently producing twin X-ray pulses. The measured electron spectra and dual-component X-ray angular profiles, together with particle-in-cell simulations, identify the contributions of the two electron populations to the radiation. The total X-ray photon yield reaches the level of 10^{10} photons per shot with a 40-TW laser system. These results establish a compact, single-stage route toward high-flux dual-pulse betatron sources for laboratory-scale ultrafast X-ray spectroscopy.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.