-
Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery
Authors:
Alizer Wong,
Heng Cui,
Yi Tan,
Xiongchao Zhan,
Liang Lin,
Yuxiang Guo,
Zhaorong Dai,
Zixin Zeng,
Wenyuan Li
Abstract:
We present Eureka, a task-conditioned Meta-Agent architecture that compiles long-horizon tasks into dynamic obligation graphs with explicit acceptance semantics. During execution, Eureka forms Macro-Agents with specialized state, memory, operators, tools, verifiers, and local topology via receding-horizon planning, architecture promotion, and minimal-sufficient compilation. When bottlenecks recur,…
▽ More
We present Eureka, a task-conditioned Meta-Agent architecture that compiles long-horizon tasks into dynamic obligation graphs with explicit acceptance semantics. During execution, Eureka forms Macro-Agents with specialized state, memory, operators, tools, verifiers, and local topology via receding-horizon planning, architecture promotion, and minimal-sufficient compilation. When bottlenecks recur, cost-benefit-gated evolution updates the local architecture under constraints. Theoretically, we establish results on regret, planning invalidation, amortization, subtree interfaces, serializability, and verification. Experimentally, Eureka completes 170/170 recursive tasks and generates 3,948 certificates with no false acceptances. Active context compresses median input from 9,490 to 4,005 tokens; incremental processing avoids 65.38% recomputation across 12,000 tasks; 16,000 concurrent executions serialize consistently. The same Meta-Agent instantiates a Theory-Discovery Agent and a Math/Conjecture Agent. The former yields structural results in quantum-process and spacetime theory. The latter identifies bottlenecks in Riemann Hypothesis research and advances a positivity certificate for Suzuki's localized Weil quadratic form to 0 < a <= 69/200 = 0.345, reaching ~99.55% of (log 2)/2. These results suggest that scientific-agent capability depends not only on the base model but on whether an architecture can be formed to match the task's cognitive structure.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Dream2Reward: Transition-Alignment Reward Models from Positive Demonstrations for Robotic Manipulation
Authors:
Haoyu Zhang,
Zecui Zeng,
Bin Wang,
Lusong Li,
Liang Lin,
Long Cheng
Abstract:
Learning robotic policies requires dense rewards that remain informative when behavior departs from successful demonstrations. Progress-based rewards estimate how far an observation has advanced along a nominal successful trajectory, but may remain high after an incorrect transition. We introduce Dream2Reward, which learns a language-conditioned successful latent transition field from positive dem…
▽ More
Learning robotic policies requires dense rewards that remain informative when behavior departs from successful demonstrations. Progress-based rewards estimate how far an observation has advanced along a nominal successful trajectory, but may remain high after an incorrect transition. We introduce Dream2Reward, which learns a language-conditioned successful latent transition field from positive demonstrations. Given the visual history up to a transition start, the model predicts the latent displacement associated with successful execution and scores the observed displacement through signed directional and symmetric magnitude agreement. This transition-level comparison penalizes wrong-direction, overshooting, and stagnant motion even when the resulting observation appears to show progress. Dream2Reward requires no failure annotations, progress labels, or synthetic negatives, and produces a dense causal reward. Across mechanism diagnostics and shared-trajectory evaluations, it provides stronger success-failure separation and more informative feedback on low-quality behavior than progress-based alternatives. Across online and offline policy learning, the same frozen reward model reduces reward hacking and supports stronger downstream performance, including in real-robot manipulation. These results show that comparing realized motion with predicted successful change provides an effective way to convert positive demonstrations into dense rewards for robot learning.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations
Authors:
Liangtao Lin,
Qingang Zhang,
Zhaomeng Zhu,
Tianwei Zhang,
Yonggang Wen
Abstract:
LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same comprehensive harness to every task, which may not be necessary and cause resource wastes. In this paper, we focus on the ident…
▽ More
LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same comprehensive harness to every task, which may not be necessary and cause resource wastes. In this paper, we focus on the identification of optimal harness configurations, and view it as a resource-matching problem between what each task requires and what the harness provides. To measure this match, we classify MCI tasks based on the mathematical representation of the underlying system and rank harness configurations by the amount and type of information they provide. We then construct task-to-harness mappings from two sources: mining research literature and measuring controlled agent execution. Leveraging the measured mapping, we propose a new harness provisioning algorithm: map-guided escalation. It begins with a task-specific harness and expands to full provision only after a failed self-check. We evaluate our method in two representative MCI tasks: in liquid cooling, it improves the agent accuracy from 0.652 under full provision to 0.715 and achieves accuracy comparable to Reflexion with 48% fewer tokens; In power grids, full provision remains accuracy-optimal, while map-based provisioning offers lower-cost alternatives. These findings show that harness provisioning follows a domain-dependent accuracy-cost Pareto frontier rather than a universal optimum.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Learning latent progression states from spatial heterogeneity in uterine histopathology
Authors:
Qiming He,
Yan Liu,
Shuang Ge,
Fan Yang,
Yuxiang Wang,
Ieng Man Zhang,
Jing Yang,
Zihao Jia,
Ajin Hu,
Yexing Zhang,
Zixiu Song,
Qiang Huang,
Xiaoya Zhao,
Zihan Wang,
Xianjing Zheng,
Yijun Zheng,
Liling Lin,
Shuxing Liu,
Bin Bao,
Yue Xie,
Tian Guan,
Yonghong He,
Congrong Liu
Abstract:
Tumor progression is accompanied by changes in architecture, morphology and microenvironmental organization, yet progression-associated heterogeneity is usually compressed into static diagnostic categories in histopathology. Here we present SpaTIE, a uterus-specific computational pathology framework that learns morphology-aware representations and organizes spatial histopathological heterogeneity…
▽ More
Tumor progression is accompanied by changes in architecture, morphology and microenvironmental organization, yet progression-associated heterogeneity is usually compressed into static diagnostic categories in histopathology. Here we present SpaTIE, a uterus-specific computational pathology framework that learns morphology-aware representations and organizes spatial histopathological heterogeneity into progression-associated tumor states. SpaTIE was developed using 10,426 uterine hematoxylin and eosin whole-slide images and evaluated in TCGA-UCEC and TCGA-UCS cohorts. The learned representations formed morphology manifolds, supported diagnostic, molecular and survival-related prediction tasks, and localized attention to informative tumor regions. Beyond supervised prediction, SpaTIE inferred tumor-state axes from cross-sectional morphology without temporal or molecular supervision. These morphology-derived states were spatially coherent and showed associations with clinicopathological variables and survival outcomes, while not simply recapitulating staging or diagnostic labels. Integrative multi-omics analyses linked the inferred states to DNA methylation, somatic copy-number variation, mutation, RNA-seq and RPPA profiles, highlighting molecular programs related to chromatin regulation, copy-number-associated structural variation, receptor tyrosine kinase signaling, cell adhesion, extracellular-matrix remodeling and metabolic adaptation. Progression-guided virtual perturbation further prioritized molecular features coupled to the morphology-derived state organization. Together, these findings suggest that uterine histopathology contains recoverable progression-associated tumor-state information and establish SpaTIE as a framework for connecting spatial morphology with multi-omics-informed tumor-state discovery.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
A Family of Simultaneously Cospectral Trees for Degree-Distance Matrices
Authors:
Limeng Lin,
Quanyu Tang,
Kehua Wang,
Wei Wang
Abstract:
Spectral characterization of graphs for various graph matrices constitutes a central topic in spectral graph theory. Let $G$ be a graph with adjacency matrix $A(G)$, diagonal degree matrix $\Deg(G)$, distance matrix $D(G)$, and transmission matrix \(\Trs(G)\), respectively. Recently, Alfaro and Zapata (2024) introduced the degree-distance matrices \(\Ddegp(G)=\Deg(G)+D(G)\) and \(\Ddeg(G)=\Deg(G)-…
▽ More
Spectral characterization of graphs for various graph matrices constitutes a central topic in spectral graph theory. Let $G$ be a graph with adjacency matrix $A(G)$, diagonal degree matrix $\Deg(G)$, distance matrix $D(G)$, and transmission matrix \(\Trs(G)\), respectively. Recently, Alfaro and Zapata (2024) introduced the degree-distance matrices \(\Ddegp(G)=\Deg(G)+D(G)\) and \(\Ddeg(G)=\Deg(G)-D(G)\), together with the transmission-adjacency matrices \(\Atrsp(G)=\Trs(G)+A(G)\) and \(\Atrs(G)=\Trs(G)-A(G)\). Based on computational evidence for trees on at most \(20\) vertices, they conjectured that all trees are determined by the spectra of \(\Ddegp\) as well as \(\Ddeg\).
In this paper, we disprove these conjectures by constructing an infinite family of pairs of non-isomorphic trees. More precisely, for each integer \(r\ge 3\), we construct a pair of trees on \(17r-15\) vertices which are simultaneously cospectral with respect to the following six matrices \[
A,\quad L,\quad Q,\quad D,\quad \Ddegp,\quad \Ddeg . \] The construction is based on an \(r\)-regularized leaf extension and an equitable-partition reduction. We also record a simple sign-switching observation for transmission-adjacency matrices: if \(G\) is bipartite, then \(\Atrs(G)\) and \(\Atrsp(G)\) are similar via a diagonal \(\{\pm1\}\)-matrix and have the same Smith normal form. Consequently, for trees, the spectral and Smith normal form problems for \(\Atrs\) and \(\Atrsp\) are equivalent.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Measurement of Branching Fraction and Transition Magnetic Moment of the Hyperon Dalitz Decay $Σ^0 \rightarrow Λe^+e^-$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (683 additional authors not shown)
Abstract:
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be…
▽ More
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be $\mathcal{B}(Σ^0 \rightarrow Λe^+e^-) = (6.34 \pm 0.25_{\rm stat.} \pm 0.23_{\rm syst.}) \times 10^{-3}$. This result shows a $2σ$ discrepancy from the theoretical calculation quoted in the PDG, where the uncertainties are statistical and systematic, respectively. In addition to the branching fraction, the transition magnetic moment $μ$ is determined to be $(1.74 \pm 0.03_{\rm stat.} \pm 0.09_{\rm syst.})\,μ_N$, where $μ_N=e/(2m_p)$ represents the nucleon magnetic moment, providing valuable insight into the intrinsic structure of the $Σ^0$ hyperon.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
Authors:
Xinye Li,
Lingshuai Lin,
Lei Wang,
Liuzhou Zhang,
Jialin Cui,
Qingshan Li,
Guanchu Wang,
Qingbin Liu,
Xi Chen,
Jiang Bian,
Wai Lam
Abstract:
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and…
▽ More
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Constraints on ultralight bosons from merging binary and remnant black holes observed during the second and third parts of the fourth LIGO-Virgo-KAGRA observing run
Authors:
The LIGO Scientific Collaboration,
the Virgo Collaboration,
the KAGRA Collaboration,
A. G. Abac,
A. Abe,
I. Abouelfettouh,
F. Acernese,
K. Ackley,
A. Adam,
S. Adhicary,
D. Adhikari,
R. X. Adhikari,
V. K. Adkins,
S. Afroz,
A. Agapito,
D. Agarwal,
M. Agathos,
N. Aggarwal,
S. Aggarwal,
O. D. Aguiar,
I. -L. Ahrend,
L. Aiello,
A. Ain,
P. Ajith,
T. Akutsu
, et al. (1786 additional authors not shown)
Abstract:
We present constraints on ultralight bosons using binary black hole mergers observed in the second and third parts of the fourth LIGO-Virgo-KAGRA observing run. Directed searches are conducted for long-transient gravitational waves from ultralight vector boson clouds around merger remnants, using a hidden-Markov-model (HMM) tracking scheme. We target the remnant black holes formed in the binary co…
▽ More
We present constraints on ultralight bosons using binary black hole mergers observed in the second and third parts of the fourth LIGO-Virgo-KAGRA observing run. Directed searches are conducted for long-transient gravitational waves from ultralight vector boson clouds around merger remnants, using a hidden-Markov-model (HMM) tracking scheme. We target the remnant black holes formed in the binary coalescences that produced GW250114 and GW250207. We find no evidence for such signals from either target. Estimating our search sensitivity at a threshold corresponding to a 1% false alarm probability, we thus disfavor vector boson masses in the range of $[2.80, 3.95]\times 10^{-13}$ eV with greater than 90% confidence. In addition, we derive constraints on ultralight scalar and vector bosons from the inferred high spins of the constituent black holes in three binaries, using events GW240515, GW241113, and GW241225_08. The excluded mass ranges in this approach depend on the assumed black-hole ages. At $10^5$ years, corresponding to typical dynamically formed binaries, we exclude scalar and vector bosons in the ranges $[1.39, 6.94]\times 10^{-13}$ eV and $[0.32, 14.4]\times 10^{-13}$ eV at 90% confidence, respectively.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation
Authors:
Peterson Co,
Sicheng Hu,
Chunxuan Jiao,
Hongyang Cheng,
Yulin Luo,
Yijie Xu,
Sixiang Chen,
Zhongxia Zhao,
Zihao Wang,
DaFeng Chi,
Peidong Liu,
YuTong Chen,
Henghua Liu,
Zhihao Yuan,
Huizhu Jia,
Yuzheng Zhuang,
Tianle Zhang,
Liang Lin,
Huajie Tan,
Shanghang Zhang
Abstract:
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs. Yet their applicability remains difficult to establish because prevailing evaluations emphasize visual quality, task outcomes, or…
▽ More
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs. Yet their applicability remains difficult to establish because prevailing evaluations emphasize visual quality, task outcomes, or coarse rollout-level responsiveness without directly testing simulator fidelity. To address this gap, we evaluate ACWMs through the observable capabilities expected of physical simulators. Accordingly, we formalize Observable Simulator Contract, a minimal contract that any action-conditioned physical simulator should satisfy: supplied actions must induce corresponding agent motion, and environment responses must be grounded in that realized motion. To operationalize this contract, we introduce WorldSimProbe, comprising five controlled suites spanning local control sensitivity, global trajectory variation, source-diverse actions, interaction grounding, and dynamics. Suite-specific evaluators assess simulator-relative calibration, dense action-to-motion correspondence, false-interaction grounding, and primitive-level dynamics. We evaluate six open-source ACWMs on more than 18,000 instances across RoboTwin, ManiSkill, and LIBERO. World-SimProbe reveals systematic action-realization degradation across control variation, structured failures in interaction grounding and dynamics, and benchmark signals consistent with human judgments and downstream outcomes. Together, this capability-based framework provides a transparent, and standardized paradigm for diagnosing ACWM simulator fidelity beyond coarse, task-directed evaluation.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
Authors:
Yuanhe Zhang,
Weiliu Wang,
Jie Ren,
Liang Lin,
Zhenhong Zhou,
Haoran Gao,
Kun Wang,
Chen Li,
Li Sun,
Sen Su
Abstract:
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Loc…
▽ More
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. To mitigate this risk, we propose Distributional Requery Guard (DRG) to detect low-frequency distribution shifts and conditionally request a second recording for semantic recovery. Across six LALMs and multiple audio understanding tasks, ILL reduces accuracy by up to 67 percentage points while receiving a mean human audibility rating of 1.33, close to 1.17 for clean audio; DRG raises mean attacked accuracy from 28.5\% to 46.1\% after clean reacquisition. These findings identify a previously overlooked safety risk for LALMs and provide a foundation for future research on robust audio understanding.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence
Authors:
Ling Lin,
Yang Bai,
Congcong Zhu,
Jiangming Shi,
Meng Wang,
Yang Long,
Jingrun Chen,
Ling Shao,
Huazhu Fu
Abstract:
Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open-ended reasoning process is prone to decision errors and error accumulation, leading to instability in answer quality. To address this, we propose an advantage-guided gating framework that dynamically intervenes in and corrects deviations d…
▽ More
Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open-ended reasoning process is prone to decision errors and error accumulation, leading to instability in answer quality. To address this, we propose an advantage-guided gating framework that dynamically intervenes in and corrects deviations during the reasoning process. Specifically, we model step-by-step reasoning as a finite-horizon decision process and introduce Monte Carlo value evaluation on the reasoning tree to provide intermediate supervision signals. The framework includes Step-Advantage Gate and Trajectory-Advantage Gate, which dynamically select high-value reasoning steps and high-quality complete reasoning trajectories, respectively. During training, we perform supervised learning for the gates using reasoning trees generated via multi-branch sampling, and combine shared-parameter initialization with task-specific heads to achieve cross-task robustness and diversity. During inference, the model greedily selects high-value prefix reasoning steps while choosing the optimal reasoning head based on the problem type, thereby significantly improving the accuracy of the final answer. Furthermore, we constructed the Reasoning-Tree-160k dataset and performed two-stage learning on it. Extensive experiments demonstrate that this advantage-guided gating framework effectively enhances the performance of benchmark MLLMs in visual-based spatial understanding and reasoning tasks. The code is open to the public for research: https://github.com/LingLin-ll/Advantage-Guided-Gate.
△ Less
Submitted 8 August, 2026;
originally announced August 2026.
-
ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB
Authors:
Yifan Wu,
Yuhan Li,
Zhenhua Wang,
Ke Chen,
Lidan Shou,
Zonghao Chen,
Liang Lin,
Huan Li,
Gang Chen
Abstract:
Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge. Our analysis of production workloads in Alibaba AnalyticDB exposes a costly ``provisioning trap'': the fear of catastrophic resource depletion drives users to bl…
▽ More
Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge. Our analysis of production workloads in Alibaba AnalyticDB exposes a costly ``provisioning trap'': the fear of catastrophic resource depletion drives users to blindly over-provision resources, wasting immense monetary budgets without alleviating non-CPU bottlenecks (e.g., I/O saturation). To break this impasse, we propose ScaleSense, a proactive, query-level resource scaling framework. Specifically, it features a multi-faceted query encoder that jointly models plan topologies and hardware specifications. Crucially, a quantile-based resource predictor estimates multi-dimensional physical footprints, acting as a reliable safety net for optimal resource scaling. An auto-scaling controller then navigates the performance-cost Pareto frontier, dynamically tailoring allocations to specific business priorities without requiring model retraining. Evaluations on over 1.36 million production queries show that ScaleSense achieves state-of-the-art prediction accuracy with good prediction interval coverage. By achieving a 76.7% relative improvement in optimal resource configuration selection over the best baseline, this approach addresses the critical performance-cost trade-off while maintaining low-overhead inference latency, confirming its practical performance in production deployments. Under the performance-optimization policy, ScaleSense satisfies user-defined performance requirements while reducing monetary cost by up to 5.22x.
△ Less
Submitted 8 August, 2026;
originally announced August 2026.
-
GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes
Authors:
Ruifeng Zhai,
Renjie Liu,
Guangrun Wang,
Liang Lin
Abstract:
We study the problem of inserting new furniture into indoor scene images. Under masked single-view 2D image-plane conditioning, however, the physical scale of the inserted furniture relative to the scene cannot be uniquely determined, making physically grounded furniture placement underdetermined from image evidence alone. We therefore reformulate the task as a combination of 3D pose inference and…
▽ More
We study the problem of inserting new furniture into indoor scene images. Under masked single-view 2D image-plane conditioning, however, the physical scale of the inserted furniture relative to the scene cannot be uniquely determined, making physically grounded furniture placement underdetermined from image evidence alone. We therefore reformulate the task as a combination of 3D pose inference and geometry-guided image generation, where estimating a geometrically plausible 3D placement is essential for reliable synthesis.
To this end, we propose a two-stage framework. For 3D placement, we introduce GOPI, a generation-oriented 3D pose inference framework that addresses the underdetermined nature of single-view furniture insertion through data-driven iterative inference, producing geometrically plausible object placements. For image generation, we develop a geometry-guided conditioning strategy that projects the inferred 3D pose into the image plane as a pixel-aligned constraint, enforcing consistency between the synthesized image and the underlying 3D geometry.
Experimental results validate the proposed framework from both 3D pose estimation and image synthesis perspectives. For 3D placement, GOPI produces poses with stronger geometric feasibility and better consistency with reference layouts than direct regression and vanilla baselines. For image synthesis, our method preserves alignment with the projected 3D geometry across different furniture scales, showing stable projection-generation alignment across the tested furniture scales.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
CrossTracer: Cross-Embodiment Navigation via VLA Model Reasoning and Trace Residuals Adapting
Authors:
Yao Wang,
Siyuan Wang,
Zhirui Sun,
Wenzheng Chi,
Liang Lin,
Jiankun Wang,
Wenjun Xu
Abstract:
Vision-language-action (VLA) models provide strong semantic priors for robot navigation, but they often ignore embodiment-specific mobility constraints. A path that is semantically plausible for one robot may be physically infeasible for another. We propose CrossTracer, a hierarchical framework for cross-embodiment navigation through adaptive trace residuals. CrossTracer represents navigation plan…
▽ More
Vision-language-action (VLA) models provide strong semantic priors for robot navigation, but they often ignore embodiment-specific mobility constraints. A path that is semantically plausible for one robot may be physically infeasible for another. We propose CrossTracer, a hierarchical framework for cross-embodiment navigation through adaptive trace residuals. CrossTracer represents navigation plans as normalized image-plane waypoints, forming a unified pixel-space interface between semantic reasoning and physical grounding. First, Vision-Language Trace Proposer (VL-Tracer) adapts a pretrained VLA model to predict an initial navigation trace from egocentric observations and flexible goal specifications. Second, CE-Adapter refines this trace by predicting embodiment-conditioned residual corrections from visual traversability cues, robot identity, and the initial trace. To train the refinement module without costly manual annotation, Cross-Embodiment RRT* (CE-RRT*) converts panoptic segmentation into robot-conditioned traversability cost maps and generates cost-minimizing pixel-space traces. We evaluate CrossTracer on the NaviTrace benchmark, which tests whether a model can generate embodiment-consistent navigation traces from egocentric observations, language instructions, and robot embodiment types. CrossTracer achieves a total score of 45.68, outperforming the strongest evaluated general-purpose baseline, Gemini-2.5-Pro, by 10.01 points, corresponding to a 28.1% relative improvement. Real-world deployment on wheeled and legged robots further shows improved navigation success and execution efficiency.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models
Authors:
Haiping Liu,
Qian Zhao,
Lijing Lin,
Jingyuan Sun,
Hongpeng Zhou
Abstract:
This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific technical variation and limiting representation transferability. To avoid directly rec…
▽ More
This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific technical variation and limiting representation transferability. To avoid directly reconstructing such variation, we shift the prediction target from observed gene measurements to latent cell representations and introduce CellWorld, which predicts the latent representations of masked cells from visible spatial context and a limited partial-expression hint. We pretrain four CellWorld variants, spanning 5.74M to 94.56M trainable parameters, on a corpus of 46 million human cells. Our controlled scaling experiments show that performance improves with model capacity, particularly on spatial tasks, while spatial transfer depends more on sufficient optimization and broad biological source diversity than on cell count alone. Across four held-out datasets, even CellWorld-Small, with 5.74M trainable parameters, outperforms every baseline on all 11 linear-probe benchmarks and all seven fine-tuned spatial benchmarks. Most notably, a frozen CellWorld-Large pretrained on only 5\% of the corpus with broad biological source coverage outperforms every fully fine-tuned baseline across all seven spatial benchmarks. Code is available at https://github.com/UoM-HealthAI/CellWorld.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents
Authors:
Tao Wang,
Qihao Yang,
Rongjiao Liang,
Lianghong Lin,
Haitao Wang,
Xinyu Cao,
Tianyong Hao
Abstract:
Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency.…
▽ More
Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency. Existing benchmarks focus on domain knowledge and question answering, largely overlooking intrinsic quality review for professional documents. Such reviews rely heavily on human experts, making them costly and difficult to scale. To bridge this gap, we introduce GB/T-Bench, the first benchmark for the structured review of national standard documents. Its GB/T Review Taxonomy is a hierarchical schema covering document structure, scope alignment, normative modality, terminology consistency, and normative references, with 25 diagnosable error types. A controllable counterexample generation mechanism combines deterministic rules and constrained LLM rewriting to process 488 documents into 7,306 traceable review error instances for evaluation. We also develop a diagnosis-oriented evaluation protocol requiring exact matches on error location, review dimension, and error type, plus document-level coverage metrics. We further propose GB/T-Reviewer, a multi-agent framework that converts review knowledge into specialized skills and coordinates global inspection, targeted diagnosis, rule scanning, and result verification. Experiments with 14 mainstream LLMs reveal a substantial human-LLM gap: the strongest model achieves only 0.3280 CMCS versus 0.6640 for experts. GB/T-Reviewer raises the best CMCS to 0.5094, showing the value of structured skill coordination for rule-intensive document review. This work paves the way for trustworthy AI in standardization and other high-stakes document domains.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
CCAT: Design and Characterization of the 350 GHz Instrument Module
Authors:
Ben Keller,
Jordan Wheeler,
Cody J. Duell,
Darshan A. Patel,
Jason Austermann,
Baird D. Bankovic,
James Burgoyne,
Scott Chapman,
Steve K. Choi,
Rodrigo Freundt,
Min Gao,
Eliza Gazda,
Anthony I. Huber,
Johannes Hubmayr,
Lawrence T. Lin,
Quintin Meyers,
Paul Malachuk,
Alicia Middleton,
Michael D. Niemack,
Tilak M. Patel,
Anna Vaskuri,
Eve Vavagiakis,
Michael R. Vissers,
Samantha Walker,
Yuhan Wang
, et al. (2 additional authors not shown)
Abstract:
The CCAT Collaboration's Prime-Cam instrument will soon be deployed to the Fred Young Submillimeter Telescope (FYST) in Chile's Atacama Desert. Featuring prominently in Prime-Cam's calibration and early science observations will be the 350 GHz instrument module, a broadband camera that will field more than 10,000 microwave kinetic inductance detectors (KIDs) across three detector arrays. Forecasts…
▽ More
The CCAT Collaboration's Prime-Cam instrument will soon be deployed to the Fred Young Submillimeter Telescope (FYST) in Chile's Atacama Desert. Featuring prominently in Prime-Cam's calibration and early science observations will be the 350 GHz instrument module, a broadband camera that will field more than 10,000 microwave kinetic inductance detectors (KIDs) across three detector arrays. Forecasts show this module will be capable of making the most sensitive to-date measurements of polarized dust emission over a large fraction of the sky at this frequency, enabling new galactic polarization science and improved understanding of cosmological foregrounds. In this work we discuss the design of the 350 GHz instrument module, covering aspects of the optics, readout, and detector arrays. We then report on the results of in-lab testing of the fully-integrated module, achieving stable cryogenic performance with a 100 mK focal plane, high detector yield, and a passband comparable to designed specifications. Upon completion of these tests, this module was shipped to the telescope site in Chile for integration in Prime-Cam.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
ContextWeave: A Real-World Workflow Benchmark
Authors:
Bo Wang,
Yuqian Yao,
Enxi Wang,
Luozhijie Jin,
Yang Liu,
Yiran Suo,
Yuxuan Cai,
Enyu Zhou,
Yufei Gao,
Honglin Guo,
Tianyu Huai,
Li Ji,
Zhikai Lei,
Bufan Li,
Lizhi Lin,
Jinxiu Liu,
Jie Yang,
Jiazheng Zhou,
Maosen Zhou,
Pengfang Qian,
Shichun Liu,
Guanshan Liu,
Hao Zheng,
Yunhao Yu,
Hang Yan
, et al. (3 additional authors not shown)
Abstract:
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-mont…
▽ More
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
Training-Free Hashing-Based Attention via Binary Principal Components
Authors:
Daohai Yu,
Zhanpeng Zeng,
Keyu Chen,
Wenhao Li,
Zhifeng Shen,
Luxi Lin,
Ruizhi Qiao,
Xing Sun,
Rongrong Ji
Abstract:
Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches. Existing sparse attention reduce computation by attending to fewer KV pairs, but often suffer from substantial accuracy degradation,…
▽ More
Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches. Existing sparse attention reduce computation by attending to fewer KV pairs, but often suffer from substantial accuracy degradation, require additional training, or rely on expensive hashing. In this work, we present BinaryPC, a training-free, data-aware hashing-based sparse attention for long-context LLMs. BinaryPC constructs compact binary hash codes and corresponding hash function by computing binary principal components of data. Unlike Locality-Sensitive Hashing (LSH) with data-independent random projections or learned non-linear hashing methods, BinaryPC constructs binary codes that explicitly preserve the structural information of data without requiring gradient-based training. Comprehensive experiments across multiple model families and long-context benchmarks show that BinaryPC preserves accuracy relative to full attention while achieving superior performance among sparse and hashing-based baselines. On modern GPUs, BinaryPC improves end-to-end decoding throughput by 3.56$\times$ over the FlashAttention kernel. Our code is available at https://github.com/yudaohai666/BPC.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Spatiotemporal Graph Transformer for Traffic Intelligence in Edge Computing
Authors:
Laha Ale,
Letian Lin,
Na Cao,
Zheng Ma,
Peng Yu
Abstract:
Accurate traffic forecasting is essential for proactive resource management in edge computing, where service demand evolves dynamically across both space and time. In practical cellular edge systems, traffic exhibits strong spatial correlations among neighboring service regions and long-range temporal dependencies driven by user mobility and application behavior. Existing recurrent forecasting app…
▽ More
Accurate traffic forecasting is essential for proactive resource management in edge computing, where service demand evolves dynamically across both space and time. In practical cellular edge systems, traffic exhibits strong spatial correlations among neighboring service regions and long-range temporal dependencies driven by user mobility and application behavior. Existing recurrent forecasting approaches can capture short-term dynamics but often struggle to model long-horizon traffic evolution under non-stationary conditions. To address this challenge, we propose a spatiotemporal graph Transformer framework that jointly models spatial interactions and temporal dependencies for traffic forecasting in edge computing. The framework employs graph neural networks to capture spatial correlations among service regions and leverages Transformer-based self-attention to learn long-range temporal patterns from historical traffic observations. By decoupling spatial representation learning from temporal reasoning, the proposed approach provides an effective mechanism for large-scale spatiotemporal traffic modeling. Extensive experiments on a real-world cellular network dataset demonstrate that the proposed graph Transformer consistently outperforms recurrent graph-based baselines, including GCN-RNN, GCN-LSTM, and GCN-GRU models, across multiple forecasting horizons. The resulting forecasts enable more effective proactive resource provisioning and reduce overload risk compared with reactive management strategies. These results highlight the potential of graph-enhanced attention mechanisms for building intelligent and adaptive edge computing systems.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
Authors:
Lucy Lin,
Ayush Jain,
Yifan Liu,
Katerina Fragkiadaki
Abstract:
Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligne…
▽ More
Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geometry-aware representations to improve spatial reasoning, they continue to lag behind specialist 3D perception systems on grounding and segmentation tasks. We argue that a key limitation is geometry-aware decoding: existing methods communicate 3D predictions through language tokens, proposal selection, or lightweight grounding queries, creating a bottleneck between language reasoning and dense geometric prediction. Building on these insights, we introduce Qwen-3D, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes. Qwen-3D augments visual tokens with 3D Rotary Positional Embeddings, allowing attention to operate directly in 3D scene space rather than across independent image frames and thereby facilitating scalable cross-view and temporal reasoning. To bridge language and geometry, Qwen-3D incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation, unifying referential grounding, instance segmentation, and visual question answering across both images and videos. Across a diverse set of benchmarks, Qwen-3D surpasses existing 3D LMMs and outperforms several large proprietary 2D models. Notably, Qwen-3D achieves these improvements while maintaining strong performance on standard 2D vision-language benchmarks by jointly training on 2D and 3D data.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling
Authors:
Li Lin,
Wujun Xu,
Weiwei Meng,
Kaiwen Xia,
Kang Hao Cheong,
Shuai Wang
Abstract:
Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on current digital cognition. While some efforts are made to enhance VLAs' reasoning capabilities by capturing temporal information, encoding the long-context history causes a…
▽ More
Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on current digital cognition. While some efforts are made to enhance VLAs' reasoning capabilities by capturing temporal information, encoding the long-context history causes an efficiency-decreasing issue. To reconcile the conflict between capturing temporal information and maintaining inference efficiency in VLAs, this paper introduces FibVLA, an efficient framework featuring temporal perception of long-context history. Specifically, we leverage logarithmic hindsight sampling to both proprioceptive states and visual frames to capture long-term temporal dependencies with minimal redundancy. For the action expert, we introduce the flow matching to produce action distributions, and the Fibonacci recurrent inference strategy to generate long-range planning steps based on real-time closed-loop feedback. Experiments demonstrate that FibVLA significantly improves action smoothness and success rates without retraining large-scale visual encoders. Efficiency analysis demonstrates superior real-time responsiveness compared to video-based baselines in real-world evaluations.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
Hierarchical Clustering of Networks via Hierarchical Distance Matrices
Authors:
Li Chen,
Nathaniel Josephs,
Eric D. Kolaczyk,
Lizhen Lin
Abstract:
Clustering populations of networks while recovering their latent hierarchical organization is a fundamental yet largely unexplored problem in network analysis.
To formalize this, we introduce the Hierarchical Distance Matrix, a specific class of population-level distance matrices that encodes latent hierarchical organization through recursively nested distance separation, accommodating unbalance…
▽ More
Clustering populations of networks while recovering their latent hierarchical organization is a fundamental yet largely unexplored problem in network analysis.
To formalize this, we introduce the Hierarchical Distance Matrix, a specific class of population-level distance matrices that encodes latent hierarchical organization through recursively nested distance separation, accommodating unbalanced tree depths.
Building on this framework, we propose a fully data-driven top-down procedure: network hierarchical clustering based on two-sample testing (NHC-TST). The algorithm recursively splits networks via spectral clustering and uses a graph-based two-sample stopping rule. The procedure adaptively determines the branching structure without requiring prior knowledge of the number of clusters or tree depth.
Theoretically, we establish exact recovery of the population-level hierarchical structure and statistical consistency in the empirical procedure.
Simulation studies demonstrate highly accurate recovery of both cluster memberships and hierarchical relationships across a wide range of settings. Applied to a global migration dataset, NHC-TST uncovers interpretable multi-resolution temporal structures that are not revealed by conventional flat clustering approaches.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans
Authors:
Jingwen Yang,
Senmao Wang,
Luoyao Kang,
Runmeng Cui,
Keying Zhang,
Yunjia Bao,
Haifan Gong,
Lin Lin,
Haiyue Jiang
Abstract:
Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly irregular, and interfaces between cartilage and surrounding soft tissues are often ambiguous. Clinical annotations may also include both composite structures containing cartilage and adjacent skin and their corresponding cartilage-only regions, produ…
▽ More
Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly irregular, and interfaces between cartilage and surrounding soft tissues are often ambiguous. Clinical annotations may also include both composite structures containing cartilage and adjacent skin and their corresponding cartilage-only regions, producing nested and overlapping labels. We propose a world-model-based segmentation framework that enables iterative anatomical reasoning beyond conventional feed-forward prediction. Built on an encoder-decoder architecture, the framework introduces a deterministic recurrent state-space model into the intermediate latent space. Multi-scale encoder features and partially decoded representations are fused to form a structural observation that initializes the latent dynamics. During inference, the model performs a three-step latent rollout without ground-truth guidance. Hierarchical anatomical actions update the recurrent state and progressively refine the latent representation. The resulting latent trajectory is projected back into the decoder and combined with high-resolution features to produce the final segmentation. To learn reliable latent transitions, we introduce a balanced hierarchical action objective that addresses foreground sparsity, missing anatomical groups, and imbalance between add and remove operations. Extensive experiments show that the proposed framework consistently improves segmentation accuracy and reduces HD95 by more than 43% for small, irregular, and overlapping auricular structures in CT. These results demonstrate the effectiveness of latent world-model reasoning for challenging medical image segmentation.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
Authors:
Hongyu Chen,
Liang Lin,
Guangrun Wang
Abstract:
Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. At each turn, the model produces a solution t…
▽ More
Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. At each turn, the model produces a solution together with a discrete correctness verdict and a confidence score; it retains the current answer only when the verdict is Correct and confidence exceeds a threshold, and otherwise continues refinement using its own self-verification. Ground-truth correctness is used only to construct training rewards and is never exposed to the policy through refinement prompts or required at inference. SVR is trained with GRPO on fixed-horizon trajectories using rewards that promote solution correctness, calibration-aware self-verification, and stop-ready correct states; adaptive stopping is activated only at inference. On seven mathematical reasoning benchmarks with Qwen3.5-2B, SVR achieves a macro-average accuracy of 0.563 with only 2.99 inference turns on average. In the evaluated complete-system comparison, it exceeds standard GRPO, strong multi-turn baselines, and a fixed-budget oracle-guided score-feedback reference while requiring substantially fewer turns than fixed ten-turn inference. These results demonstrate that learned self-verification can serve as an effective internal control signal for answer retention and adaptive test-time compute allocation.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Schreier-Coset Graph Rewiring
Authors:
Aryan Mishra,
Randy Martinez,
Lizhen Lin
Abstract:
The information flow in the graph neural networks (GNNs) is fundamentally constrained by over-squashing, where structural bottlenecks impede long range information propagation. Graph-rewiring methods, which modify graph topology, have been extensively used to alleviate this. However, existing approaches often introduce prohibitive structural and computational bottlenecks, fail to preserve the crit…
▽ More
The information flow in the graph neural networks (GNNs) is fundamentally constrained by over-squashing, where structural bottlenecks impede long range information propagation. Graph-rewiring methods, which modify graph topology, have been extensively used to alleviate this. However, existing approaches often introduce prohibitive structural and computational bottlenecks, fail to preserve the critical properties of original graphs, and increase the edge counts massively. We introduce a novel method Schreier-Coset Graph Rewiring , a group-theoretic rewiring method that augments the input graph with a Schreier-Coset graph derived from a special linear group. Our method provides theoretical guarantees, a graph that exhibits spectral gap and a bounded effective resistance, creating a low-resistance bypass for long-range communication. Empirical evaluations demonstrate that SCGR reduces effective resistance by 5-40% across various learning tasks, effectively mitigating connectivity bottlenecks while maintaining competitive accuracy.
△ Less
Submitted 29 July, 2026;
originally announced July 2026.
-
Laser power transmission in space: Plasma-based power cell
Authors:
Li Lin,
Michael Keidar
Abstract:
Laser power beaming offers a route to space energy delivery, but semiconductor laser photovoltaic receivers face thermalization, joule heat, and radiative recombination waste, etc. Here we propose a gas-phase plasma power cell that converts vacuum-ultraviolet photons into electrical output through xenon photoionization and magnetically biased charge separation. Particle-in-cell Monte Carlo simulat…
▽ More
Laser power beaming offers a route to space energy delivery, but semiconductor laser photovoltaic receivers face thermalization, joule heat, and radiative recombination waste, etc. Here we propose a gas-phase plasma power cell that converts vacuum-ultraviolet photons into electrical output through xenon photoionization and magnetically biased charge separation. Particle-in-cell Monte Carlo simulations of a low-pressure xenon chamber driven by a 58.4 nm pulsed laser predict a steady-state laser-to-electrical conversion efficiency of 82.25% at 2000 W/m2 average incident power. Energy accounting closes to 1%, with 7.52% photon escape, 8.95% boundary loss, and 1.29% chamber-stored energy. Parameter scans over bias voltage, magnetic field, and gas pressure identify photon absorption and electron confinement as controlling design factors. These results are a proof-of-concept gas-phase receiver architecture for space laser power beaming.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment
Authors:
Shiyu Teng,
Haichen Yu,
Jiaqing Liu,
Hao Sun,
Yu Song,
Shurong Chai,
Ruibo Hou,
Lanfen Lin,
Yen-Wei Chen
Abstract:
Multimodal behavioral analysis offers a scalable approach to assessing depression, anxiety, and stress, yet generic fusion models often ignore the psychometric structure of questionnaire labels. In DASS-21, risk labels are derived from ordered symptom items through fixed item-to-subscale mappings. We propose \textbf{DynaBridge}, a dynamic summary-guided cross-task multimodal framework for DASS-str…
▽ More
Multimodal behavioral analysis offers a scalable approach to assessing depression, anxiety, and stress, yet generic fusion models often ignore the psychometric structure of questionnaire labels. In DASS-21, risk labels are derived from ordered symptom items through fixed item-to-subscale mappings. We propose \textbf{DynaBridge}, a dynamic summary-guided cross-task multimodal framework for DASS-structured mental health assessment. DynaBridge encodes acoustic, visual, and textual cues across multiple sessions and augments them with frozen-LLM-generated DASS-aware summaries as participant-level semantic evidence. It predicts ordinal item distributions, reconstructs depression, anxiety, and stress risk evidence from item-level soft scores, and fuses this evidence with direct multimodal risk predictions. A confidence-aware refinement strategy further incorporates high-confidence semantic cues conservatively. On the official AdoDAS validation split, DynaBridge outperforms the official baseline and representative multimodal methods, achieving 0.5012 mean F1 for D/A/S risk prediction and 0.3216 mean QWK for DASS-21 item prediction. These results show the value of bridging multimodal cues, semantic summaries, and DASS-21 psychometric structure.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning
Authors:
Yuan Zhang,
Jiang Hu,
Zhijian Lai,
Lin Lin,
Zaiwen Wen
Abstract:
Optimization over the Stiefel manifold plays a significant role in various machine learning tasks. Existing methods either use the retraction operators, requiring costly orthonormalization for large-scale matrices, or employ landing methods that rely on careful step size selection and penalty parameter tuning. To address these challenges, we propose a retraction-free and penalty parameter-free alg…
▽ More
Optimization over the Stiefel manifold plays a significant role in various machine learning tasks. Existing methods either use the retraction operators, requiring costly orthonormalization for large-scale matrices, or employ landing methods that rely on careful step size selection and penalty parameter tuning. To address these challenges, we propose a retraction-free and penalty parameter-free algorithm that directly lands on the manifold. By leveraging the strongly-convex-like property of the quadratic penalty function and the proximal smoothness of the Stiefel manifold, we establish global convergence guarantees with the best-known iteration complexities under both constant and diminishing step sizes. Then, we reformulate the low-rank adaptation (LoRA) fine-tuning problem for large language models as a manifold optimization problem, introducing Manifold-LoRA for geometry-accelerated adaptation. This approach employs the proposed landing technique and a carefully designed step size strategy to accelerate the training process. Numerical experiments on benchmark datasets demonstrate the efficiency and strong downstream performance of the proposed method.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Zhinv: Real-time hub-height wind field reconstruction using only local sparse observations
Authors:
Zongwei Zhang,
Chin Chun Ooi,
Lianlei Lin,
Sheng Gao,
Tiantian He,
Yew Soon Ong,
Junkai Wang,
Hangyi Yu,
Jiaqi Zhang,
Hanqing Zhao,
Yu Zhang
Abstract:
The high proportion of wind power connected to the grid places higher demands on fine-grained knowledge of regional wind fields. Since the wind information directly obtainable in actual operations is mostly sparse, discrete, and irregularly distributed local observations, it is difficult to directly meet the needs of tasks such as wind power regulation, wind resource assessment, and low-altitude e…
▽ More
The high proportion of wind power connected to the grid places higher demands on fine-grained knowledge of regional wind fields. Since the wind information directly obtainable in actual operations is mostly sparse, discrete, and irregularly distributed local observations, it is difficult to directly meet the needs of tasks such as wind power regulation, wind resource assessment, and low-altitude environmental perception of continuous regional wind fields. Therefore, we propose Zhinv, an end-to-end reconstruction framework that directly weaves sparse and irregular observations into a fine-grid wind field at hub-height. Experiments in Northeast China, Europe, and Southeast Asia demonstrate that Zhinv can accurately, robustly, and efficiently reconstruct fine-grid wind fields from sparse observations, reducing the error by about 66% compared with Kriging. With local wind-power observations as input, Zhinv enables wind power centers to bypass NWP and complex assimilation processes, supporting direct and real-time wind resource assessment from locally available data.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Coherence from interference: a solvable model of sub-GeV dark matter-nucleus scattering
Authors:
Lynn Lin,
Tongyan Lin,
Momei Fang
Abstract:
How do dark matter-nucleus interactions transition from the regimes of coherent scattering, where single phonons are produced, to that of individual nuclear recoils? Answering this question relies on understanding multiphonon excitations. Multiphonons are important for interpreting low-threshold direct detection experiments, yet are computationally prohibitive to compute. In this paper, we employ…
▽ More
How do dark matter-nucleus interactions transition from the regimes of coherent scattering, where single phonons are produced, to that of individual nuclear recoils? Answering this question relies on understanding multiphonon excitations. Multiphonons are important for interpreting low-threshold direct detection experiments, yet are computationally prohibitive to compute. In this paper, we employ a 1D $N$-site crystal lattice model where dark matter scattering can be computed exactly. We show that the only difference between coherent and incoherent scattering is that conservation of crystal momentum is enforced in coherent scattering. The momentum conservation constraint becomes less important as more phonons are produced, yielding the transition to incoherent scattering. Using numerical calculations of the 1D structure factor, we also obtain quantitative validation of using an incoherent approximation to compute sub-GeV dark matter scattering in realistic 3D crystals.
△ Less
Submitted 27 July, 2026;
originally announced July 2026.
-
MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents
Authors:
Belal S. Alsinglawi,
Weizheng Wang,
Junyi Wu,
Yi Jiang,
Lianhai Lin,
Merouane Debbah,
Izzat Alsmadi
Abstract:
Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception, navigation, collaboration, and reasoning, but few assess whether physical evidence, protocol constraints, and action r…
▽ More
Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception, navigation, collaboration, and reasoning, but few assess whether physical evidence, protocol constraints, and action risk remain coupled during critical decisions. We introduce MulRobBench, an offline, protocol-conditioned benchmark for Vision-Language-Action (VLA) UAV agents in smart-city environments. MulRobBench integrates real UAV multimodal observations, protocol-level security policies, and action-level cyber-physical safety into a unified evaluation framework. The benchmark contains 3,024 samples spanning 17 task taxonomy nodes and 12 scoring dimensions across four stages: operational context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning. Evaluation combines semantic scoring with structural diagnostics, including policy compliance, format compliance, unsafe actions, parsing failures, and dimension-level validity. Across 17 multimodal models, the best semantic protocol-decision score reaches only 0.5141, while the best strict mean scoring-dimension accuracy is 0.1599. A controlled 20-anchor modality-ablation study changes 4-15 action selections per model, confirming that both visual and textual inputs influence decisions. Analysis identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability. MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
GRAPE: Graduated Routing for Articulated Portrait mesh Estimation
Authors:
Yunfei Liu,
Lijian Lin,
Ye Zhu,
Yu Li
Abstract:
Articulated portrait mesh estimation is fundamental to 3D understanding, avatar generation, and immersive interaction. Existing approaches primarily rely on 3D Morphable Models (3DMMs). However, face-centric models suffer from the "floating head" assumption, conflating head pose with global rotation due to the lack of neck kinematics. Conversely, body-centric models lack high-fidelity facial expre…
▽ More
Articulated portrait mesh estimation is fundamental to 3D understanding, avatar generation, and immersive interaction. Existing approaches primarily rely on 3D Morphable Models (3DMMs). However, face-centric models suffer from the "floating head" assumption, conflating head pose with global rotation due to the lack of neck kinematics. Conversely, body-centric models lack high-fidelity facial expression capabilities. Furthermore, current methods struggle to disentangle jaw articulation from expression blendshapes, often over-relying on expressions for mouth opening. These limitations make monocular portrait recovery difficult across representation, supervision, and anatomical parameter estimation. To address these limitations, we introduce GRAPE(Graduated Routing for Articulated Portrait mesh Estimation). We build a Portrait Parametric Model (PPM) with an explicit torso-to-head kinematic chain and a canonical injection step to merge FLAME and the SMPL-X torso. We propose a Progressive Anatomical Alignment (PAA) network, which is composed of a pretrained portrait encoder, a Graduated-Mask Router, and coarse-to-fine experts that follow the portrait anatomical prior. We then train this network with multi-source supervision that combines sparse anatomical keypoints, feature distillation, foreground mask constraints, and relative geometry constraints. Experiments show that GRAPE improves portrait mesh recovery quality, pose alignment, and jaw--expression disentanglement over prior methods. We also demonstrate that our method can benefit the downstream tasks of audio-driven talking-head generation and 3D portrait generation.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Effects of long-chain branching, short-chain branching, and polydispersity on pressure sensitive rheology of polymer melts
Authors:
Lilian Lin,
Matthew Joe,
Heon E. Park
Abstract:
The rheological behavior of polymer melts under high pressure is a critical factor in many industrial processes like injection molding and extrusion, yet it is often inadequately characterized. At operating pressures that can exceed 100 MPa, viscosity can increase by orders of magnitude, making atmospheric-pressure data insufficient for accurate process simulation. This pressure induced viscosity…
▽ More
The rheological behavior of polymer melts under high pressure is a critical factor in many industrial processes like injection molding and extrusion, yet it is often inadequately characterized. At operating pressures that can exceed 100 MPa, viscosity can increase by orders of magnitude, making atmospheric-pressure data insufficient for accurate process simulation. This pressure induced viscosity increase is highly dependent on molecular architectures of the materials. This study aims to deconstruct the influence of specific structural features such as short-chain branching (SCB), long-chain branching (LCB), and polydispersity on the pressure sensitivity of the viscosity of polyethylene. Utilizing a high-pressure sliding plate rheometer (HPSPR) to ensure accurate measurements under uniform shear and pressure, we characterized four distinct polyethylene melts. All samples, regardless of their structure, exhibited piezorheologically simple behavior, allowing the application of time-pressure superposition over the entire shear rate range. A key finding is that the long-chain branched sample, known from the literature to be thermorheologically complex, was found to be piezorheologically simple. This dichotomy is explained by the different physical mechanisms of temperature and pressure. The pressure sensitivity of the viscosity, quantified by the pressure-viscosity coefficient, was found to be strongly dependent on molecular branching. Both SCB and LCB significantly increase the pressure sensitivity while polydispersity had a negligible effect. These results demonstrate that molecular branches are the dominant structural parameter controlling the rheological response of polyethylene to pressure, providing crucial insights for the development of more accurate predictive models for high-pressure polymer processing.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
AdaptICA: Data-Adaptive Transformation Learning for Independent Component Analysis
Authors:
Lida Jalili,
Jingyu Liu,
Vince D. Calhoun,
Li-Hsiang Lin
Abstract:
Independent component analysis (ICA) is widely used to recover latent structure from signal and imaging data, but standard ICA assumes that the observed measurement scale preserves a linear mixing structure. This assumption may fail for features produced through nonlinear preprocessing, such as band-specific power in motor-imagery EEG. We propose AdaptICA, an adaptive transformation-based framewor…
▽ More
Independent component analysis (ICA) is widely used to recover latent structure from signal and imaging data, but standard ICA assumes that the observed measurement scale preserves a linear mixing structure. This assumption may fail for features produced through nonlinear preprocessing, such as band-specific power in motor-imagery EEG. We propose AdaptICA, an adaptive transformation-based framework that jointly learns grouped componentwise transformations and the demixing structure using a profiled mutual-information criterion. Because the transformation and demixing parameters may compensate for one another, their joint estimation introduces new identifiability and asymptotic challenges. We establish identifiability, consistency, and asymptotic normality of the transformation estimator, together with joint strong consistency of the transformation and demixing estimators. AdaptICA selects the transformation structure data-adaptively and includes the identity transformation as a candidate, thereby reducing to standard ICA when no scale adjustment is needed. Extensive simulations support the theoretical results. Applications demonstrate that AdaptICA can recover more independent and interpretable sources when transformation is beneficial while retaining standard ICA when the original measurement scale is adequate.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
Spectral Gap of the Davies Generator for the Mean-Field Heisenberg Model
Authors:
Joao Basso,
Thiago Bergamaschi,
Lin Lin,
Michael Ragone,
Kevin D. Stubbs
Abstract:
The mean-field Heisenberg ferromagnet is a quantum spin model on the complete graph with isotropic spin-1/2 interactions. This non-commuting Hamiltonian is permutation and $\mathsf{SU}(2)$ invariant, and its Gibbs states undergo an $\mathsf{SU}(2)$ symmetry breaking phase transition at inverse temperature $β=2$. We consider the associated Davies generator, a canonical model of open-system thermali…
▽ More
The mean-field Heisenberg ferromagnet is a quantum spin model on the complete graph with isotropic spin-1/2 interactions. This non-commuting Hamiltonian is permutation and $\mathsf{SU}(2)$ invariant, and its Gibbs states undergo an $\mathsf{SU}(2)$ symmetry breaking phase transition at inverse temperature $β=2$. We consider the associated Davies generator, a canonical model of open-system thermalization, and prove tight asymptotic estimates for its spectral gap at all noncritical temperatures. For fixed $β<2$, the gap as a function of number of qubits $n$ is $Θ(1)$, while for fixed $β>2$ the gap is $Θ(n^{-1})$. The matching upper bound of the spectral gap is witnessed by the total magnetization order parameter, suggesting that the low-temperature ($β>2$) slowdown is associated with broken continuous symmetry. Two key ingredients in our approach are a comparison argument, which introduces auxiliary generators to bound dissipation on nontrivial representations of the symmetry groups $\mathsf{SU}(2)$ and $\mathsf{S}_n$, and a decomposition of the space of observables into spherical tensor operators to reveal a form of monotonicity.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
Why Some Quantum States Cannot Be Recovered
Authors:
Yuan Liu,
Linhan Lin,
Ke-Mi Xu
Abstract:
The recovery of quantum information after subsystem loss is a central challenge in quantum information processing. However, some states remain beyond the reach of any recovery strategies. Here we identify the algebraic origin of irrecoverability, the ghost information---correlations encoded in the global state that leave no trace on any accessible subsystem. We introduce a scalar measure quantifyi…
▽ More
The recovery of quantum information after subsystem loss is a central challenge in quantum information processing. However, some states remain beyond the reach of any recovery strategies. Here we identify the algebraic origin of irrecoverability, the ghost information---correlations encoded in the global state that leave no trace on any accessible subsystem. We introduce a scalar measure quantifying its magnitude and prove a universal error floor below which no virtual recovery map can operate, irrespective of resource investment. We further uncover a spectral phase transition in the sampling cost: bounded when the underlying linear map exhibits a spectrum gap, and divergent with a universal exponent in the gapless regime. Together with the universal error floor, this dichotomy organizes all multipartite quantum states into four classes. Moreover, it is revealed that conditional mutual information---the standard entropic diagnostic---is fundamentally irrelevant to virtual recoverability. As an implication, we show that the error floor imposes a detection threshold for loss-tolerant quantum metrology.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation
Authors:
Zhijing Yang,
Haocheng Lin,
Zhihua Xu,
Haojie Li,
Keze Wang,
Liang Lin,
Tianshui Chen
Abstract:
Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamental challenge in automated spatial design. Existing approaches, primarily based on one-shot generation using diffusion models or Large Language Models (LLMs), lack explicit mechanisms for intermediate geometric constraint verification, often resultin…
▽ More
Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamental challenge in automated spatial design. Existing approaches, primarily based on one-shot generation using diffusion models or Large Language Models (LLMs), lack explicit mechanisms for intermediate geometric constraint verification, often resulting in structural collisions and functionally infeasible arrangements under complex room constraints. To address these challenges, we propose Agentic Designer, a progressive, multi-agent framework that formulates structure-aware interior layout generation as an iterative and constraint-verified decision process. By decomposing layout synthesis into modular stages of proposal, verification, and adjustment, the framework coordinates three specialized agents, a Generator, an Evaluator, and a Refiner, through a Progressive Consensus Mechanism. This mechanism enforces stepwise geometric validation and correction before each placement is committed, thereby preventing error accumulation. To facilitate this structure-aware paradigm and standardize evaluation, we establish InStruct, a comprehensive benchmark that integrates a dataset comprising over 18,000 high-quality, parametrically annotated samples with a novel suite of structure-centric metrics. Extensive quantitative evaluations, qualitative analyses, and user studies show that Agentic Designer significantly outperforms state-of-the-art methods, demonstrating substantial improvements in strict structural adherence and functional design coherence.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Modified Mean Curvature Flow in Fuchsian Manifolds
Authors:
Yuk Shing Lam,
Longzhi Lin
Abstract:
In this paper, we show that the modified mean curvature flow starting from an arbitrary graph in a Fuchsian manifold exists for all time and converges smoothly to an equidistant surface of constant mean curvature as $t\to \infty$. This result generalizes earlier work to the modified mean curvature flow setting and removes the restrictive global gradient bound initially required for the standard me…
▽ More
In this paper, we show that the modified mean curvature flow starting from an arbitrary graph in a Fuchsian manifold exists for all time and converges smoothly to an equidistant surface of constant mean curvature as $t\to \infty$. This result generalizes earlier work to the modified mean curvature flow setting and removes the restrictive global gradient bound initially required for the standard mean curvature flow by Huang, Zhou, and the second author \cite{HLZ2020}.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
GWTC-5.0: Tests of General Relativity
Authors:
The LIGO Scientific Collaboration,
the Virgo Collaboration,
the KAGRA Collaboration,
A. G. Abac,
A. Abe,
I. Abouelfettouh,
F. Acernese,
K. Ackley,
A. Adam,
S. Adhicary,
D. Adhikari,
R. X. Adhikari,
V. K. Adkins,
S. Afroz,
A. Agapito,
D. Agarwal,
M. Agathos,
N. Aggarwal,
S. Aggarwal,
O. D. Aguiar,
I. -L. Ahrend,
L. Aiello,
A. Ain,
P. Ajith,
T. Akutsu
, et al. (1800 additional authors not shown)
Abstract:
The signals from the LIGO-Virgo-KAGRA network of gravitational-wave (GW) detectors allow us to perform sensitive tests of general relativity (GR) in the dynamical and strong-field regime of gravity. We present the results of seven tests of GR using the observed binary signals in the fifth GW Transient Catalog (GWTC-5.0), i.e., up to and including the second part of the fourth observing run (O4b).…
▽ More
The signals from the LIGO-Virgo-KAGRA network of gravitational-wave (GW) detectors allow us to perform sensitive tests of general relativity (GR) in the dynamical and strong-field regime of gravity. We present the results of seven tests of GR using the observed binary signals in the fifth GW Transient Catalog (GWTC-5.0), i.e., up to and including the second part of the fourth observing run (O4b). We restrict our analysis to the confident signals, henceforth called events, observed by at least two detectors that have estimated false alarm rates $\le 10^{-3} \ \rm{yr}^{-1}$. These include 72 events from O4b and five events from the first part of the fourth observing run that are now analyzed due to their increased significance from updated search results, bringing the total number of events for tests of GR in the cumulative GWTC to 168. After subtracting the best-fit waveforms, we find the residuals are consistent with detector noise for all events considered. We also find no strong evidence for additional polarizations beyond those predicted by GR. We perform tests of GW generation, improving the constraints on deviations from the GR post-Newtonian coefficients by factors of 1.2-2.6. Finally, we find overall consistency of the remnants with GR using both time- and frequency-domain methods. For GW240621_195059, postmerger data are consistent with the dominant quadrupolar ($\ell=|m|=2$) mode of a Kerr black hole and its first overtone, with spurious high-frequency content preventing a spectroscopic constraint of GR. In the frequency-domain ringdown analysis, the GR prediction lies in the tails of the combined results, possibly due to the limited catalog size. However, the combined results indicate improved consistency with GR over GWTC-4.0, owing to the contribution of GW250114 with a network matched-filter signal-to-noise ratio of 76.9. Overall, we find no evidence for physics beyond GR.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation
Authors:
Yongsen Zheng,
Ruilin Xu,
Guohua Wang,
Liang Lin,
Kwok-Yan Lam
Abstract:
The Matthew effect is a big challenge in Recommender Systems (RSs), where popular items tend to receive increasing attention, while less popular ones are often overlooked, perpetuating existing disparities. Although many existing methods attempt to mitigate Matthew effect in the static or quasi-static recommendation scenarios, such issue will be more pronounced as users engage with the system over…
▽ More
The Matthew effect is a big challenge in Recommender Systems (RSs), where popular items tend to receive increasing attention, while less popular ones are often overlooked, perpetuating existing disparities. Although many existing methods attempt to mitigate Matthew effect in the static or quasi-static recommendation scenarios, such issue will be more pronounced as users engage with the system over time. To this end, we propose a novel framework, Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation (HiCore), aiming to address Matthew effect in the Conversational Recommender System (CRS) involving the dynamic user-system feedback loop. It devotes to learn multi-level user interests by building a set of hypergraphs (i.e., item-, entity-, word-oriented multiple-channel hypergraphs) to alleviate the Matthew effec. Extensive experiments on four CRS-based datasets showcase that HiCore attains a new state-of-the-art performance, underscoring its superiority in mitigating the Matthew effect effectively. Our code is available at https://github.com/zysensmile/HiCore.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
ChemHyperMag: Physics-informed magnetic hypergraph learning improves molecular ADMET prediction
Authors:
Hexiao Ding,
Hongzhao Chen,
Jing Lan,
Yufeng Jiang,
Zihong Luo,
Zehua Xiong,
Tianlong Ruan,
Yunlin Mao,
Nga Chun Ng,
Gwing Kei Yip,
Gerald W. Y. Cheng,
Kate Inyoung Oh,
Jing Cai,
Liang-Ting Lin,
Jung Sun Yoo
Abstract:
Accurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) is important for drug discovery. Most predictors use undirected molecular graphs and pairwise edges. This choice misses asymmetric interactions, nonreversible dynamics, and motif level effects from functional groups and ring systems. We propose ChemHyperMag for multitask ADMET prediction under missing labe…
▽ More
Accurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) is important for drug discovery. Most predictors use undirected molecular graphs and pairwise edges. This choice misses asymmetric interactions, nonreversible dynamics, and motif level effects from functional groups and ring systems. We propose ChemHyperMag for multitask ADMET prediction under missing labels. ChemHyperMag builds a functional group hypergraph from rings, BRICS fragments, Bemis-Murcko scaffolds, and bonds. It also defines a potential driven nonreversible flow guided by electronegativity and Gasteiger partial charges. The resulting circulation is encoded by a Hermitian magnetic Laplacian and processed with a magnetic Chebyshev encoder. We perturb magnetic phases to form stochastic views and train with an InfoNCE objective. Experiments on multiple ADMET benchmarks show improvements over recent methods with fewer labeled samples and no conformers. ChemHyperMag is scalable and provides interpretable directional signals through its magnetic phases.
△ Less
Submitted 22 July, 2026; v1 submitted 19 July, 2026;
originally announced July 2026.
-
Anomalously high deuterium fractionation in a galactic translucent cloud: a challenge to chemical models
Authors:
Gan Luo,
Zhi-Yu Zhang,
Thomas G. Bisbas,
Di Li,
Serena Viti,
Roberto Neri,
Junzhi Wang,
Siyi Feng,
Ningyu Tang,
Daniel R. Rybarczyk,
Lingrui Lin
Abstract:
Deuterated (D-) species have long been proposed to diagnose the physical conditions and chemical evolution of cold dense molecular clouds. While deuterium fractionation has been extensively measured in dense cores, observations in diffuse and translucent clouds remain rare. We report here the detection of DCN and DNC toward a translucent cloud ($A_{\rm V} =1.2\pm0.2$ mag, $n_{\rm H_2}$ =…
▽ More
Deuterated (D-) species have long been proposed to diagnose the physical conditions and chemical evolution of cold dense molecular clouds. While deuterium fractionation has been extensively measured in dense cores, observations in diffuse and translucent clouds remain rare. We report here the detection of DCN and DNC toward a translucent cloud ($A_{\rm V} =1.2\pm0.2$ mag, $n_{\rm H_2}$ = $3.9\pm0.2\times10^2$ cm$^{-3}$) through sensitive absorption observations with the IRAM NOrthern Extended Millimeter Array (NOEMA). This detection reaches the lowest column-density and volume-density regime in which deuteration has been observed so far. Interestingly, the observed DCN/HCN and DNC/HNC abundance ratios ($3.3\pm0.6\times10^{-3}$ and $3.6\pm1.2\times10^{-3}$, respectively), which are more than two orders of magnitude higher than the element abundance [D]/[H] (1.5$\times$10$^{-5}$), suggest an unexpected enhancement of deuterium fractionation in the translucent cloud. These results represent a significant departure from established chemical models considering deuterium fractionation, which predict negligible formation of D-molecules in such environments. Although it remains unclear how D-molecules built up their abundances in translucent gas, a dispersed dense core scenario could potentially explain the observed high deuterium fraction. This interpretation is consistent with the idea proposed by Price et al. (2003) more than two decades ago: a translucent cloud may be a transient, dynamically evolving structure formed through the dissipation of a dense molecular cloud.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation
Authors:
Yongsen Zheng,
Ruilin Xu,
Ziliang Chen,
Guohua Wang,
Mingjie Qian,
Jinghui Qin,
Liang Lin
Abstract:
The Matthew effect is a notorious issue in Recommender Systems (RSs), \emph{i.e.}, the rich get richer and the poor get poorer, wherein popular items are overexposed while less popular ones are regularly ignored. Most methods examine Matthew effect in static or nearly-static recommendation scenarios. However, the Matthew effect will be increasingly amplified when the user interacts with the system…
▽ More
The Matthew effect is a notorious issue in Recommender Systems (RSs), \emph{i.e.}, the rich get richer and the poor get poorer, wherein popular items are overexposed while less popular ones are regularly ignored. Most methods examine Matthew effect in static or nearly-static recommendation scenarios. However, the Matthew effect will be increasingly amplified when the user interacts with the system over time. To address these issues, we propose a novel paradigm, Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation (HyCoRec), which aims to alleviate the Matthew effect in conversational recommendation. Concretely, HyCoRec devotes to alleviate the Matthew effect by learning multi-aspect preferences, \emph{i.e.}, item-, entity-, word-, review-, and knowledge-aspect preferences, to effectively generate responses in the conversational task and accurately predict items in the recommendation task when the user chats with the system over time. Extensive experiments conducted on two benchmarks validate that HyCoRec achieves new state-of-the-art performance and the superior of alleviating Matthew effect. Our code is available at https://github.com/zysensmile/HyCoRec.
△ Less
Submitted 19 July, 2026;
originally announced July 2026.
-
Cross-Coordinate Correspondence Pruning for Image-to-Point Cloud Registration
Authors:
Xin Liu,
Rong Qin,
Huipeng Lin,
Leizhi Shu,
Jin Wu,
Chi-Man Vong,
Liang Lin,
Jufeng Yang
Abstract:
Recent detection-free approaches have shown significant efficacy in image-to-point cloud (I2P) registration by employing a coarse-to-fine matching pipeline. In the coarse stage, down-sampled image features and voxelized point cloud features are typically fused to establish initial coarse correspondences for subsequent refinement. However, existing methods largely overlook the critical role of poin…
▽ More
Recent detection-free approaches have shown significant efficacy in image-to-point cloud (I2P) registration by employing a coarse-to-fine matching pipeline. In the coarse stage, down-sampled image features and voxelized point cloud features are typically fused to establish initial coarse correspondences for subsequent refinement. However, existing methods largely overlook the critical role of point cloud density, which fundamentally dictates the quality of coarse correspondences and the final registration results. Specifically, excessively sparse point clouds lead to an insufficient number of inliers, while overly dense ones often introduce a high outlier ratio. Consequently, this creates an inherent density trade-off, thereby significantly limiting the registration accuracy of current approaches. For mitigating this trade-off, we propose a novel Cross-Coordinate Correspondences Pruning (CCP) strategy to acquire sufficient inliers while ensuring a low outlier ratio. To minimize interference from inter-modal coordinate discrepancies, we first project cross-coordinate coarse correspondences to the 2D image coordinate system for spatial unification. Subsequently, a lightweight pruning network is responsible for predicting the inlier confidences, which are used to filter coarse outliers, from coordinate geometric and modal feature dimensions. To maximize inlier recall, we further design a Multi-Density Point Ensemble (MDPE) strategy that consolidates and deduplicates pruned coarse correspondences across varying point cloud densities. Our method achieves a significant performance improvement, surpassing existing state-of-the-art methods by at least 8.6% in Registration Recall across various benchmarks.
△ Less
Submitted 19 July, 2026;
originally announced July 2026.
-
PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
Authors:
Yang Liu,
Weixing Chen,
Xinshuai Song,
Tao Pu,
Siwen Mo,
Yongjie Bai,
Zihao Chen,
Qianran Sun,
Liruo Zhong,
Ying Shen,
Liang Lin
Abstract:
Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, semantic verification, and persistent experience across heterogeneous embodiments. We present PhyAgentOS, a runtime foundation delivering scheduling, verification, memory, benchmarking, and safety as system-level services. I…
▽ More
Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, semantic verification, and persistent experience across heterogeneous embodiments. We present PhyAgentOS, a runtime foundation delivering scheduling, verification, memory, benchmarking, and safety as system-level services. Its Session-Centered Runtime treats a session, not an action, as the minimum unit of scheduling, compatibility preflight, supervised execution, evidence collection, and acceptance. To decouple cognition from physical execution, the cognition-physics boundary is a file system: the State-as-a-File protocol materializes cross-layer state as Markdown with YAML, yielding inspectable, versionable records without code dependencies between Agent and Runtime layers. These views form a unified cognitive state space aligning intent, capabilities, environment, execution, and experience. The SessionVerifier distinguishes execution termination from semantic task completion via evidence-grounded verdicts of success, failure, or replan. Verified outcomes are consolidated through epistemic memory into reusable knowledge and corrective lessons, closing a trial-and-error loop without retraining. Benchmarking reuses the deployment session and verification path, so results trace to real execution. Layered safety constrains both policy-driven and agent-driven execution: preflight, action bridges, SafetyGuard, heartbeat monitoring, and target-local constraints. Validation is progressive: games test cognitive planning, simulation adds dynamics and control, real robots add hardware noise, with the cognitive layer held constant. PhyAgentOS is benchmarked on Optimus-67, StarDojo, and DST-Dojo, validated on 19+ simulated and physical embodiments, and gains on LIBERO, Calvin, and RoboCasa365 across multiple VLA models.
△ Less
Submitted 18 July, 2026;
originally announced July 2026.
-
Gravitational Effective Theories with Maximal Supersymmetry and a Peculiar Parity
Authors:
Justin Berman,
Simon Caron-Huot,
Aditi V. Chandra,
Henriette Elvang,
Aidan Herderschee,
Loki L. Lin,
Roger Morales
Abstract:
We study the space of four-dimensional ultraviolet completions for $\mathcal{N}=8$ supergravity that are described at low energies by weakly-coupled effective field theories (EFTs) with maximal supersymmetry and $\mathrm{SU}(4)\times\mathrm{SU}(4)$ R-symmetry. We show that tree-level factorization of the 4-, 5-, and 6-point EFT scattering amplitudes, together with a certain ``peculiar parity'' con…
▽ More
We study the space of four-dimensional ultraviolet completions for $\mathcal{N}=8$ supergravity that are described at low energies by weakly-coupled effective field theories (EFTs) with maximal supersymmetry and $\mathrm{SU}(4)\times\mathrm{SU}(4)$ R-symmetry. We show that tree-level factorization of the 4-, 5-, and 6-point EFT scattering amplitudes, together with a certain ``peculiar parity'' condition, leads to nonlinear constraints on the 4-point Wilson coefficients. This peculiar parity is a property that can only be imposed on a subset of scalar amplitudes. Combining the nonlinear constraints with positivity, we find that the allowed region of 4-point Wilson coefficients is reduced to a non-convex domain with two sharp corners: one being the closed superstring Virasoro--Shapiro amplitude, the other an infinite spin tower amplitude exchanging states of every spin at the same mass. We show both numerically and analytically that requiring a finite number of states near the first mass level leaves only the Virasoro--Shapiro amplitude.
△ Less
Submitted 15 July, 2026;
originally announced July 2026.
-
Joint On-and-Off Policy Learning for Vision-and-Language Navigation
Authors:
Qingrong He,
Lin Zhao,
Kevin Zheng,
Liang Lin
Abstract:
Vision-and-Language Navigation (VLN) necessitates an embodied agent to navigate in the physical world by adhering to natural language instructions. Recent advancements in Vision-Language Models (VLM) have propelled the development of VLM-based VLN methods with two predominant paradigms: (1) imitation learning (IL) on expert demonstrations, followed by the Dataset Aggregation (DAgger) algorithm to…
▽ More
Vision-and-Language Navigation (VLN) necessitates an embodied agent to navigate in the physical world by adhering to natural language instructions. Recent advancements in Vision-Language Models (VLM) have propelled the development of VLM-based VLN methods with two predominant paradigms: (1) imitation learning (IL) on expert demonstrations, followed by the Dataset Aggregation (DAgger) algorithm to bolster error recovery capabilities; (2) reinforcement learning (RL) driven by verifiable rewards to enhance reasoning and exploration. A notable gap is the absence of integration between these two distinct paradigms. This paper introduces JOP-VLN, a novel VLN framework that synergistically combines off-policy imitation learning and on-policy exploration within a three-stage training pipeline. Initially, IL is employed on expert demonstrations to acquire basic navigation skills. Subsequently, the DAgger algorithm is utilized to generate heuristic exploration trajectories, which are then used for imitation learning to improve error recovery capabilities. Finally, a joint on-and-off policy learning framework is implemented, featuring high-entropy trajectory sampling to enhance RL training efficiency and an error-correction-prioritized trajectory sorting strategy for effective error correction. Extensive experiments demonstrate the efficacy of JOP-VLN, achieving success rates of 69.9% and 68.0% on the VLN-CE R2R and RxR benchmarks, respectively, setting a new state-of-the-art on R2R. Project page: https://qingrongh.github.io/JOP-VLN.
△ Less
Submitted 15 July, 2026;
originally announced July 2026.
-
Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents
Authors:
Yaopei Zeng,
Congchao Wang,
JianHang Chen,
Nan Wang,
Yurui Chang,
Lu Lin
Abstract:
LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final failure is observed. Reliable deployment therefore requires \emph{step-level confidence estimation}: a calibrated probability that each proposed action is productive, avai…
▽ More
LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final failure is observed. Reliable deployment therefore requires \emph{step-level confidence estimation}: a calibrated probability that each proposed action is productive, available \emph{before} the action is executed. Existing LLM confidence estimators are designed to score a response from the given prompt, but agent confidence also depends on execution consequences: whether similar actions in similar situations actually advanced the task after the environment responded. We introduce the \method (\methodshort), a self-evolving critic framework in which an LLM critic accumulates evidence from its own past judgments and their observed consequences. After each trajectory, a hindsight LLM that sees the full execution feedback votes on whether each step was productive. The resulting pseudo-labels populate a memory bank from which related productive and unproductive experiences are retrieved into the critic's prompt whenever a similar step recurs. \methodshort requires no training and uses no ground truth step labels. Across three agent benchmarks and three critic backbones, \methodshort attains the best calibration (ECE and Brier) and ranking (AUC) in every dataset--critic combination, reducing ECE by up to $54\%$ relative to the strongest training-free baseline.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
Rethinking Monocular Depth Embedding for Generalized Stereo Matching
Authors:
Libo Lin,
Shuangli Du,
Minghua Zhao,
Zhenzhen You,
Shun Lv,
Yiguang Liu
Abstract:
Generally, monocular methods capture rich contextual priors but lack geometric precision, whereas stereo methods are geometrically accurate yet struggle in textureless and occluded regions. Several approaches attempt to combine their strengths to enhance the generalization of stereo matching (SM) by aligning monocular depth with stereo information. However, establishing a stable and generalizable…
▽ More
Generally, monocular methods capture rich contextual priors but lack geometric precision, whereas stereo methods are geometrically accurate yet struggle in textureless and occluded regions. Several approaches attempt to combine their strengths to enhance the generalization of stereo matching (SM) by aligning monocular depth with stereo information. However, establishing a stable and generalizable alignment is challenging, and unreliable monocular cues can substantially degrade performance. This paper rethinks monocular depth embedding. First, to prevent shortcut learning, we reduce branch coupling instead of expanding network width. Second, we construct soft constraints instead of hard ones from monocular depth to improve tolerance to monocular depth errors. Based on the principles, we integrate monocular information into both feature extraction and GRU iterations. Specifically, the monocular depth map is fused with the RGB image to sharpen depth boundary perception and suppress matching ambiguities. The fused image is then used for feature extraction, allowing the contextual features to encode global geometric information. Furthermore, the monocular depth gradient feature is employed to guide disparity updates, helping to escape local oscillations. Finally, to address the boundary blurring of supervised disparity caused by data augmentation, we propose an edge confidence estimation method and an edge-aware loss function. Our method achieves state-of-the-art (SOTA) performance on multiple standard benchmarks, demonstrating excellent generalization while improving accuracy. The code is available at https://github.com/linliboabc-maker/stereo-matching-digital.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.