-
Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training
Authors:
Shangbo Yuan,
Jie Xu,
Xiaofeng Zhu,
Na Zhao
Abstract:
Recently, open-vocabulary 3D object detection (3D-OVD) has gained increasing attention for its ability to detect unseen objects in 3D scenes. Existing approaches typically adopt a two-stage pipeline that first discovers novel objects using foundation models and then trains a 3D-OVD model based on these discovered objects. Although effective, this pipeline often suffers from inaccurate localization…
▽ More
Recently, open-vocabulary 3D object detection (3D-OVD) has gained increasing attention for its ability to detect unseen objects in 3D scenes. Existing approaches typically adopt a two-stage pipeline that first discovers novel objects using foundation models and then trains a 3D-OVD model based on these discovered objects. Although effective, this pipeline often suffers from inaccurate localization and mismatched classification during the discovery stage, which subsequently limits the performance of the model training stage. To address these limitations, we advocate for improving both the reliability of novel object discovery and the robustness of model training, and propose an innovative framework. Specifically, for reliable discovery, our co-distillation strategy distills high-quality novel objects by applying Hungarian matching over a comprehensive score that incorporates geometric consistency, structural objectness, and semantic certainty. To enhance robust model training, we further propose a dual-guidance learning scheme, incorporating a scene-awareness-guided uncertainty regularization for the regression head and an LLM-guided hierarchical alignment for the classification head, effectively mitigating the negative effects of imprecise 3D bounding boxes and semantic ambiguity. Extensive experiments on SUN RGB-D and ScanNetV2 demonstrate that our method achieves significant performance gains over state-of-the-art approaches. Code is available at https://github.com/shangboyuan/Co-3DGT
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
The Symmetry and Linear Stability of Convex 1+5 Coorbital Central Configurations with Homogeneous Potential
Authors:
Yiyang Deng,
Jiangtao Xu
Abstract:
For the planar Newtonian 1+N-body problem when the N masses tend to zero, the corresponding relative equilibria become coorbital around the dominant mass. In this work, we focus on convex central configurations in the planar 1+N coorbital problem. For the 1+5 coorbital problem with the homogeneous potential, we prove that any convex coorbital central configuration with symmetric masses must have a…
▽ More
For the planar Newtonian 1+N-body problem when the N masses tend to zero, the corresponding relative equilibria become coorbital around the dominant mass. In this work, we focus on convex central configurations in the planar 1+N coorbital problem. For the 1+5 coorbital problem with the homogeneous potential, we prove that any convex coorbital central configuration with symmetric masses must have an axis of symmetry. Furthermore, under explicit restrictions on the angular variables in a homogeneous potential, we prove the linear stability of both convex symmetric 1+5 and convex 1+N coorbital central configurations.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction
Authors:
Kangdi Wang,
Yusheng Dai,
Jin Xu
Abstract:
Continuous-latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high-frequency loss, phase incoherence, and stereo-image collapse. These share a structural root: waveform autoencoders lack an explicit frequency axis, leaving no handle for targeted per-band correction. Among five matched-budget represent…
▽ More
Continuous-latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high-frequency loss, phase incoherence, and stereo-image collapse. These share a structural root: waveform autoencoders lack an explicit frequency axis, leaving no handle for targeted per-band correction. Among five matched-budget representations, the complex STFT achieves the lowest full-band and high-frequency spectral distances, providing direct access to magnitude and phase at every bin. Building on this, we present ear-VAE2, a complex-spectral autoencoder with cross-channel interaction. Spec-SnakeBeta learns a periodic activation per frequency bin with frequency-dependent initialization, outperforming other activation variants while using fewer parameters than the fully independent variant. Duplex-Aware Refiner applies band-specific corrections to magnitude and phase following duplex theory of sound localization. On the 546-track Song Describer Dataset, ear-VAE2 achieves the best point estimates on five of seven reconstruction metrics. The Duplex-Aware Refiner reduces Mel Distance by 19.4% and uses ~45% fewer residual-output dimensions than the Unconstrained Refiner, while also lowering spectral distances, spatial-cue errors, and receiving higher ratings from professional engineers. The downstream generator using ear-VAE2 latents achieves better point estimates on all 12 automatic metrics.Demo page is available at https://eps-acoustic-revolution-lab.github.io/EAR_VAE2/.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale
Authors:
Xiaohan Huang,
Qingqing Long,
Xiaolei Du,
Siyu Pu,
Jiawen Xu,
Haotian Chen,
Chenyang Zhao,
Jinbiao Liu,
Xuezhi Wang,
Hao Wang,
Hengshu Zhu,
Yuanchun Zhou
Abstract:
Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation. This limitation stems from the fragmentation of scientific data across heterogeneous repositories and from dataset representations designed primarily for human use. To address this limitation, we introduce the Scientific Data Ski…
▽ More
Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation. This limitation stems from the fragmentation of scientific data across heterogeneous repositories and from dataset representations designed primarily for human use. To address this limitation, we introduce the Scientific Data Skill (SciDSK), an agent-ready representation that packages dataset-specific knowledge and operational guidance as a reusable agent skill. A SciDSK integrates dataset descriptions, scientific context, file organization, usage procedures, quality checks, and provenance information while retaining the underlying data in its original repository. We define a structured SciDSK specification and develop a systematic construction pipeline that grounds each SciDSK in authoritative dataset records and associated supporting materials. We further establish the Scientific Data Skill Bank, a unified platform that publishes SciDSK resources across six scientific disciplines and supports package access, persistent identification, and traceability to source datasets. We evaluate SciDSK through a retrieval benchmark for dataset discovery and controlled cases for dataset interpretation. The results show that SciDSK improves agent-driven dataset discovery and provides more precise and actionable support for dataset interpretation. These findings support the value of organizing dataset-specific knowledge in an agent-ready representation.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Early Planet Formation in Embedded Disks (eDisk). XXIV: Systematic Investigation of Disk Structures based on Visibility Analysis
Authors:
Mayank Narang,
Jerry Xu,
Leslie W. Looney,
Nagayoshi Ohashi,
Anika Khandavalli,
Patrick Sheehan,
Jonathan P. Williams,
Shigehisa Takakuwa,
Jes K. Jørgensen,
Ilseung Han,
Woojin Kwon,
Zhi-Yun Li,
Nguyen Thi Phuong,
John J. Tobin
Abstract:
The dust continuum emission from young protostellar disks encodes key information about their mass distribution and early evolution, yet uniform high-resolution comparative studies remain limited. We present a systematic uv-plane analysis of parametric intensity models applied to ALMA Band-6 (1.3 mm) observations of 23 disks (19 protostellar systems with 4 being in binary) from the eDisk sample, s…
▽ More
The dust continuum emission from young protostellar disks encodes key information about their mass distribution and early evolution, yet uniform high-resolution comparative studies remain limited. We present a systematic uv-plane analysis of parametric intensity models applied to ALMA Band-6 (1.3 mm) observations of 23 disks (19 protostellar systems with 4 being in binary) from the eDisk sample, spanning Gaussian profiles to power-law cores with exponential tails (PLCT), including asymmetric extensions. Gaussian models generally fail to reproduce the centrally peaked emission and extended outer structure observed in most disks, whereas the PLCT framework provides a significantly improved description of radial brightness profiles. Incorporating azimuthal asymmetries further reduces residuals in 15 of 17 inclined disks, indicating that departures from axisymmetry are common at early stages. Only two disks, L1489 IRS and Oph IRS63, exhibit clear gap and ring substructures, while most appear smooth at the spatial resolution and sensitivity of our observations. These systems are among the most evolved in the sample, and the absence of flat-spectrum sources limits the evolutionary range probed, {suggesting that the detection of prominent gaps and rings is not common} in the earliest phases of disk evolution. Using a uniform definition of disk radius based on the 95\% enclosed flux, we find a positive correlation with stellar mass, $R_{\rm disk} \propto M_{\star}^{1.5 \pm 0.1}$, with disks in binary systems systematically smaller than those around isolated protostars. While the models capture overall morphology and large-scale asymmetries, distinguishing intrinsic structures from radiative transfer effects in optically thick regions remains challenging.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Universal Machine-learning Molecular Dynamics at the Speed of Empirical Potentials
Authors:
Tiancheng Li,
Jianming Xue,
Linfeng Zhang,
Duo Zhang,
Han Wang
Abstract:
No interatomic potential has offered universality across chemistry, near-first-principles accuracy and the speed of empirical potentials at once. Here we introduce DPA4C, an equivariant potential whose architecture and compressed CUDA operators are co-designed under deployment constraints to pursue accuracy and efficiency together. Five variants spanning a 49-fold parameter range form the high-thr…
▽ More
No interatomic potential has offered universality across chemistry, near-first-principles accuracy and the speed of empirical potentials at once. Here we introduce DPA4C, an equivariant potential whose architecture and compressed CUDA operators are co-designed under deployment constraints to pursue accuracy and efficiency together. Five variants spanning a 49-fold parameter range form the high-throughput end of the measured accuracy--throughput frontier. The largest variant approaches the accuracy of the MACE-Omat models at about two orders of magnitude higher measured throughput. The most compact reduces the energy, force and stress errors of the fastest existing universal MLIP by 61.4%, 48.1% and 34.3% at 1.92 times its saturated throughput. All five variants complete multimillion-atom simulations on a single GPU and run molecular dynamics for 2.048 billion atoms on 1,024 16-GB NVIDIA V100 GPUs at 83.3--91.2% weak-scaling efficiency. Compared with the MEAM empirical potential, DPA4C-Nano reaches 1.8 and 2.5 times the saturated throughput in single-GPU scans on the same V100 hardware for diamond carbon and FCC copper, respectively. DPA4C therefore brings quantum-trained universal accuracy into a regime of speed and system size previously associated with empirical potentials.
△ Less
Submitted 19 August, 2026; v1 submitted 19 August, 2026;
originally announced August 2026.
-
Perturbation theory of mesoscale plasmonic waveguide with an analytical treatment of nonclassical electromagnetic boundary condition
Authors:
Jiling Xue,
Haitao Liu
Abstract:
The optical modes of mesoscale plasmonic waveguides (MPWs) are significantly affected by nonclassical quantum effects, which can be comprehensively described by the nonclassical electromagnetic boundary condition (NEBC) formulated with the surface-response Feibelman d-parameters. In this paper, a perturbation theory for the nonclassical waveguide modes (NWMs) supported by MPWs under the NEBC is pr…
▽ More
The optical modes of mesoscale plasmonic waveguides (MPWs) are significantly affected by nonclassical quantum effects, which can be comprehensively described by the nonclassical electromagnetic boundary condition (NEBC) formulated with the surface-response Feibelman d-parameters. In this paper, a perturbation theory for the nonclassical waveguide modes (NWMs) supported by MPWs under the NEBC is proposed. In this theory, by adopting the classical waveguide modes (CWMs) under the classical electromagnetic boundary condition (CEBC) as the basis functions and treating the NEBC as a first-order perturbation, a general expression of the propagation constant of the NWM with an analytical dependence on the NEBC is derived. This theory transparently reveals the underlying general relation between the nonclassical effects and the propagation properties of the NWMs, thereby providing an effective tool for the understanding and design of MPW devices, as well as for the experimental measurement of the d-parameters.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Sharp Sobolev Approximation on General Domains by Linearized Shallow Networks with Analytic Activations
Authors:
Jia Li,
Tong Mao,
Jinchao Xu
Abstract:
We study Sobolev approximation on bounded domains by linearized shallow neural networks whose inner parameters are prescribed independently of the target function. Our main step is a one-dimensional construction for analytic activations. We prove that quasi-Chebyshev parameter sets with univariate resolution $m$ generate fixed feature spaces attaining the sharp $H^r$-to-$H^s$ approximation order…
▽ More
We study Sobolev approximation on bounded domains by linearized shallow neural networks whose inner parameters are prescribed independently of the target function. Our main step is a one-dimensional construction for analytic activations. We prove that quasi-Chebyshev parameter sets with univariate resolution $m$ generate fixed feature spaces attaining the sharp $H^r$-to-$H^s$ approximation order $m^{-(r-s)}$ for a class of analytic activations satisfying a quantitative non-cancellation condition on their Taylor coefficients. Combining this result with the ridge-function lifting theorem in [SIAM J. Math. Anal. 30 (1998), pp. 155-189] and its extension to arbitrary quasi-uniform direction sets established in this work, we construct tensor-product-type parameter sets that attain the sharp rate $$\|f-f_n\|_{L^2(Ω)}\lesssim n^{-\frac rd}\|f\|_{H^r(Ω)},\quad f\in H^r(Ω)$$ for all $r>0$. In contrast to the finite-difference construction in [Neural Comput. 8 (1996), pp. 164-177], whose explicit admissibility condition may require an extremely small parameter scale, the proposed parameter sets remain distributed over fixed intervals and are therefore more amenable to practical computation.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Authors:
Liya Zhu,
Xin Ma,
Tao Liu,
Haodong Wang,
Ge Zhang,
Jingzhe Ding,
Qingshui Gu,
Yongjie Zhong,
Jinxiang Meng,
Yuan Gao,
Yunqiu Zhou,
Hao Zhu,
Jifeng He,
Yongzhi Liao,
Xinyi Zhang,
Chaoxin Li,
Yi Zhu,
Xi Lin,
Duju Zeng,
Xiang Gao,
Wen Zhang,
Yunyang Wang,
Duo Wang,
Huan Zhou,
Zuo Wang
, et al. (13 additional authors not shown)
Abstract:
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-va…
▽ More
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Agent Lightning v1.0: Towards Harnessed Agentic RL
Authors:
Zhiyuan He,
Siwei Zhang,
Zhiwen Zhou,
Yuqing Yang,
Yu Kang,
Yuge Zhang,
Luna K. Qiu,
Tin Yan Tsui,
Jiahang Xu,
Chong Luo
Abstract:
Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through an LLM endpoint proxy, an approach later adopted by frameworks such as verl Uni-Agent, AReaL 2.0, slime, and Polar. We refer to th…
▽ More
Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through an LLM endpoint proxy, an approach later adopted by frameworks such as verl Uni-Agent, AReaL 2.0, slime, and Polar. We refer to this paradigm as harnessed agentic RL, where the deploy-time harness directly participates in model post-training. Harnessed agentic RL differs fundamentally from traditional agentic RL: the harness, rather than the training engine, owns the environment interaction loop, while the trainer observes only sequences of LLM request-response pairs. This introduces challenges in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling, which can substantially affect training stability and effectiveness. We present Agent Lightning v1.0, a lightweight framework for harnessed agentic RL implemented in approximately 3,500 lines of code. It supports arbitrary agent harnesses and serves as a practical testbed for studying these challenges. We evaluate it on instruction-following, search, and coding agents, and provide a complete reproducible pipeline for coding-agent RL. Using only 6K training examples and modest compute, RL improves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%, a 14.6-point absolute gain. We release the complete workflow and training scripts to facilitate reproducible research on harnessed agentic RL.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design
Authors:
Xuefeng Liu,
Mingxuan Cao,
Xiao Luo,
Songhao Jiang,
Tobin Sosnick,
Jinbo Xu,
Louis Maher,
Rick Stevens
Abstract:
Biomolecular design underpins applications from molecular recognition to therapeutics and synthetic biology, yet de novo interaction design remains challenging-especially for DNA/RNA, underexplored non-protein modalities with scarce, heterogeneous complex data and sharper geometric and chemical constraints. We introduce MCTH (Monte Carlo Tree Hallucination), an inference-only framework that casts…
▽ More
Biomolecular design underpins applications from molecular recognition to therapeutics and synthetic biology, yet de novo interaction design remains challenging-especially for DNA/RNA, underexplored non-protein modalities with scarce, heterogeneous complex data and sharper geometric and chemical constraints. We introduce MCTH (Monte Carlo Tree Hallucination), an inference-only framework that casts all-atom sequence-structure co-design as uncertainty-aware planning over hallucinated states from pretrained folding and inverse-folding models, with optional biophysical control within the same decision loop. MCTH treats these models as frozen black-box operators and uses Monte Carlo Tree Search to allocate a fixed inference budget across competing design trajectories, incorporating model confidence and uncertainty, as well as cross-expert consensus/disagreement when multiple predictors are available. Across protein-RNA, protein-DNA, protein-protein, and protein-ligand design, matched-budget experiments show that adaptive search improves over simpler sampling and cycling strategies, while held-out AlphaFold3 and Chai-1 evaluations demonstrate transfer beyond the search-time oracle. MCTH provides a shared planning layer across modalities while allowing task-specific folding, inverse-folding, and biophysical modules, requiring no fine-tuning or backpropagation through component models.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Ribet bimodules and principally polarized superspecial abelian varieties with quaternion action
Authors:
Jiangwei Xue,
Xiangning Yang
Abstract:
In an influential paper [K. Ribet, Bimodules and abelian surfaces, in Algebraic number theory, 359-407, Adv. Stud. Pure Math., Vol. 17, 1989] on the bad reduction of Shimura curves, Ribet studies certain superspecial abelian surfaces over $\overline{\mathbb{F}}_p$ with quaternion multiplication by a maximal order $\mathcal{O}$ in an indefinite quaternion $\mathbb{Q}$-algebra ramified at $p$. In pa…
▽ More
In an influential paper [K. Ribet, Bimodules and abelian surfaces, in Algebraic number theory, 359-407, Adv. Stud. Pure Math., Vol. 17, 1989] on the bad reduction of Shimura curves, Ribet studies certain superspecial abelian surfaces over $\overline{\mathbb{F}}_p$ with quaternion multiplication by a maximal order $\mathcal{O}$ in an indefinite quaternion $\mathbb{Q}$-algebra ramified at $p$. In particular, he classifies the $p$-divisible groups of such $\mathcal{O}$-abelian surfaces by classifying $(\mathcal{O}_p, \mathcal{O}_p)$-bimodules $L_p$ that are free over $\mathbb{Z}_p$ (i.e.bilattices) under an additional admissible assumption. In this paper, we generalize Ribet's result by removing the admissible assumption and producing a complete classification of $(\mathcal{O}_p, \mathcal{O}_p)$-bilattices $L_p$. Equip the right order $\mathcal{O}_p$ with the canonical involution, and suppose additionally that the left order $\mathcal{O}_p$ is equipped with an orthogonal involution $*$. We derive the necessary and sufficient condition for the existence of a perfect quaternion hermitian form $\langle~,~\rangle_p:L_p\times L_p\to \mathcal{O}_p$ on the right $\mathcal{O}_p$-lattice $L_p$ inducing the given involution $*$ on the left order $\mathcal{O}_p$, and give a complete classification of such self-dual quaternion hermitian $(\mathcal{O}_p, *, \mathcal{O}_p)$-bilattices $(L_p, \langle~,~\rangle_p)$. Globally, we apply these classification results to the study of the existence of principal polarizations on superspecial abelian varieties over $\overline{\mathbb{F}}_p$ equipped with $\mathcal{O}$-action.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Temperature-Induced Reorganization of Supported Zn$_3$ Clusters on Cu(111): From Minimum-Energy Structures to Finite-Temperature Ensembles
Authors:
Jiayan Xu,
Zheng Yu,
Abhirup Patra,
Amar Deep Pathak,
Sharan Shetty,
Detlef Hohl,
Roberto Car
Abstract:
Understanding the nature of catalytic active sites under reaction conditions remains a central challenge in heterogeneous catalysis. In industrial copper/zinc oxide/alumina catalysts for methanol synthesis, small Zn-based species at the Cu interface have long been proposed as active-site candidates, yet their atomic-scale structure and stability remain controversial. Computational studies typicall…
▽ More
Understanding the nature of catalytic active sites under reaction conditions remains a central challenge in heterogeneous catalysis. In industrial copper/zinc oxide/alumina catalysts for methanol synthesis, small Zn-based species at the Cu interface have long been proposed as active-site candidates, yet their atomic-scale structure and stability remain controversial. Computational studies typically identify such species from optimized 0 K structures, assuming that minimum-energy configurations remain representative under reaction conditions. Here, we combine machine-learning-interatomic-potential-accelerated global optimization, molecular dynamics, and enhanced-sampling free-energy calculations to investigate supported Zn$_3$(OH)$_3$ and Zn$_3$(OH)$_2$CHOO clusters on Cu(111)-based surfaces from 0 to 450 K. While compact triangular configurations are generally favored among minimum-energy structures at 0 K, finite-temperature free-energy calculations reveal a pronounced shift toward extended linear configurations with increasing temperature. This transition is driven primarily by entropic stabilization and cannot be inferred from potential energies alone. Molecular dynamics further shows substantial cluster mobility on pristine Cu(111), indicating that long-term persistence depends not only on configurational stability but also on surface mobility. Surface Zn alloying strongly suppresses diffusion, thereby stabilizing isolated interfacial Zn species. Together, these results show that thermodynamically relevant structures of supported Zn-based clusters can differ fundamentally from static 0 K predictions because of competing enthalpic and entropic effects. Our findings highlight the limitations of identifying catalytic active sites solely from 0 K structures and underscore the importance of explicit finite-temperature sampling in catalyst modeling.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation
Authors:
Cong Zhao,
Shuai Tian,
Xu Zhang,
Baocheng Ni,
Xinguo Song,
Xueying Sun,
Shu Jiang,
Shouchang Yang,
Bo Tang,
Jin Deng,
Ge Zhu,
YongCheng Wang,
Jin Xu,
Ri Yang
Abstract:
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps acr…
▽ More
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Unlocking Multi-Component Bulk-Materials Molecular Dynamics with a Small-Footprint Machine Learning Interatomic Potential
Authors:
Yucheng Ouyang,
Xin Chen,
Ying Liu,
Lifang Wang,
Xingyu Gao,
Xiawei Du,
Jianierken Habudelihan,
Haifeng Song,
Huimin Cui,
Xiaobing Feng,
Jingling Xue
Abstract:
Bulk materials, as opposed to nanomaterials, require molecular dynamics (MD) simulations on a large spatial scale (~10^9 atoms or more) to adequately capture their atomic-scale physical properties. Previously, the introduction of machine-learning interatomic potentials (MLIPs) has extended MD to this scale, but even single-component bulk systems require tens of thousands of GPUs on high-end superc…
▽ More
Bulk materials, as opposed to nanomaterials, require molecular dynamics (MD) simulations on a large spatial scale (~10^9 atoms or more) to adequately capture their atomic-scale physical properties. Previously, the introduction of machine-learning interatomic potentials (MLIPs) has extended MD to this scale, but even single-component bulk systems require tens of thousands of GPUs on high-end supercomputers. However, multi-component bulk MD simulations remain barely achievable, as the HBM footprint of existing MLIPs - already substantial for single-component systems - grows explosively in multi-component scenarios. This paper proposes an MLIP with a small HBM footprint - less than 3% that of existing MLIPs - unlocking multi-component bulk MD using only hundreds of GPUs. This is achieved by first identifying feature vectors and intermediate tensors as the two primary contributors to HBM footprints in existing MLIPs. To address these two sources, the dimensionality of the feature vectors has been reduced by introducing physical and chemical knowledge, and intermediate tensors have been eliminated by aggressively fusing all kernels into a single mega-kernel. In evaluation, the proposed MLIP has used 144 NVIDIA A100 GPUs to perform MD simulations on a 6-component bulk system with 1.14x10^9 atoms, while previously such MD simulation spatial scale has been restricted to unary systems and typically achieved on high-end supercomputers equipped with tens of thousands of GPUs.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
First measurements of the branching fractions of $J/ψ$ and $ψ(3686) \to Σ^{0} \barΣ^{0}η$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (750 additional authors not shown)
Abstract:
Based on $(10087 \pm 44) \times 10^6$ $J/ψ$ and $(2712 \pm 14) \times 10^6$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, the hadronic decays $J/ψ\to Σ^{0} \barΣ^{0} η$ and $ψ(3686) \to Σ^{0} \barΣ^{0} η$ are observed for the first time. The corresponding branching fractions are measured to be…
▽ More
Based on $(10087 \pm 44) \times 10^6$ $J/ψ$ and $(2712 \pm 14) \times 10^6$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, the hadronic decays $J/ψ\to Σ^{0} \barΣ^{0} η$ and $ψ(3686) \to Σ^{0} \barΣ^{0} η$ are observed for the first time. The corresponding branching fractions are measured to be $\mathcal{B}(J/ψ\to Σ^{0} \barΣ^{0}η)= (7.5 \pm 0.3 \pm 0.8) \times 10^{-5}$ and $\mathcal{B}(ψ(3686) \to Σ^{0} \barΣ^{0}η)= (1.3\pm 0.1 \pm 0.1) \times 10^{-5}$, respectively, where the first uncertainties are statistical, and the second systematic. The ratio $\text{Q} \approx \frac{\mathcal{B}(ψ(3686) \to Σ^{0} \barΣ^{0} η)}{\mathcal{B}(J/ψ\to Σ^{0} \barΣ^{0} η)}$ is determined to be $(17.3 \pm 1.5 \pm 1.7)\%$, which is con sistent with the 12\%-rule within 3.0$σ$.~No significant intermediate states or threshold enhancements are observed in the $Σ^0$($\barΣ^{0}$)$η$ and $Σ^0$$\barΣ^{0}$ invariant mass spectra.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Measurement of Branching Fraction and Transition Magnetic Moment of the Hyperon Dalitz Decay $Σ^0 \rightarrow Λe^+e^-$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (683 additional authors not shown)
Abstract:
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be…
▽ More
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be $\mathcal{B}(Σ^0 \rightarrow Λe^+e^-) = (6.34 \pm 0.25_{\rm stat.} \pm 0.23_{\rm syst.}) \times 10^{-3}$. This result shows a $2σ$ discrepancy from the theoretical calculation quoted in the PDG, where the uncertainties are statistical and systematic, respectively. In addition to the branching fraction, the transition magnetic moment $μ$ is determined to be $(1.74 \pm 0.03_{\rm stat.} \pm 0.09_{\rm syst.})\,μ_N$, where $μ_N=e/(2m_p)$ represents the nucleon magnetic moment, providing valuable insight into the intrinsic structure of the $Σ^0$ hyperon.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Superspecial Points on Shimura Curves
Authors:
Yasuhiro Terakado,
Jiangwei Xue,
Chia-Fu Yu
Abstract:
Let $X$ be the Shimura curve attached to an indefinite quaternion $\mathbb{Q}$-algebra $B$ with a maximal order $O_B$. This paper investigates the reduction $X\otimes \mathbb{F}_p$ of $X$ modulo an arbitrary prime $p$, focusing particularly on its superspecial locus. We give an explicit criterion for the existence of superspecial $\mathbb{F}_q$-rational points on $X$. Furthermore, we compute both…
▽ More
Let $X$ be the Shimura curve attached to an indefinite quaternion $\mathbb{Q}$-algebra $B$ with a maximal order $O_B$. This paper investigates the reduction $X\otimes \mathbb{F}_p$ of $X$ modulo an arbitrary prime $p$, focusing particularly on its superspecial locus. We give an explicit criterion for the existence of superspecial $\mathbb{F}_q$-rational points on $X$. Furthermore, we compute both the number of geometric superspecial points and the number of $\mathbb{F}_p$-rational superspecial points, through the Eichler class number formula and the Selberg trace formula. As a key ingredient, we classify the Dieudonné modules attached to superspecial $O_B$-abelian surfaces, which generalizes Ribet's classification of admissible quaternion bimodules of rank $2$ by dropping the admissible hypothesis. These results generalize Deuring's explicit formula for supersingular elliptic curves over $\mathbb{F}_p$ and give the Shimura-curve analogue of the Ibukiyama-Katsura formulas for principally polarized superspecial abelian surfaces over $\mathbb{F}_p$.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
SEER: Long-Context Reasoning via Selective Visual-Text Compression
Authors:
Jiawei Xu,
Zhilin Zhai,
Jinrui Fang,
Ruohan Xu,
Mingfei Lu,
Yi Zhang,
Guanchu Wang,
Tianlong Chen,
Ying Ding
Abstract:
Long-context reasoning remains computationally expensive for large language models due to the quadratic complexity of attention over text tokens. Visual-text compression offers a promising alternative by rendering text into images and processing them with vision-language models, often reducing token usage. However, existing approaches apply uniform compression regardless of query relevance, potent…
▽ More
Long-context reasoning remains computationally expensive for large language models due to the quadratic complexity of attention over text tokens. Visual-text compression offers a promising alternative by rendering text into images and processing them with vision-language models, often reducing token usage. However, existing approaches apply uniform compression regardless of query relevance, potentially sacrificing precision where detailed extraction is required. We present SEER, a framework that learns to select query-relevant images through visual scanning and retrieve textual content only where needed, combining the efficiency of visual compression with the precision of text-based reasoning. Through supervised fine-tuning on tool-interaction trajectories, SEER learns adaptive tool invocation for selection and retrieval. Experiments on long-context benchmarks show that SEER improves extraction precision through selective text retrieval while retaining average prompt-token savings relative to full-text baselines. On LongBench, SEER achieves 51.11% average accuracy, outperforming the visual-text baseline Glyph-9B by 2.33 points and Qwen3-8B by 3.49 points. Code can be accessed at https://github.com/jiaweixu98/SEER
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation
Authors:
Wenhao Yuan,
Chenchen Lin,
Wentao Hu,
Jian Chen,
Jinfeng Xu,
Shujie Li,
Edith Cheuk Han Ngai
Abstract:
\textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, under client heterogeneity, the conventional static split strategy may be suboptimal because clients can differ in data distributions, adaptation dynamics, and representation learning progress, making a single split point insufficient to accommodate client-speci…
▽ More
\textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, under client heterogeneity, the conventional static split strategy may be suboptimal because clients can differ in data distributions, adaptation dynamics, and representation learning progress, making a single split point insufficient to accommodate client-specific training states. In this paper, we propose \textsc{FedSGA}, a \textbf{S}ufficiency-\textbf{G}uided \textbf{A}daptive split \textbf{Fed}erated learning framework that addresses this question through client-specific shallow sufficiency estimation. First, we introduce a client-specific adaptation channel based on private prompt tokens, which tracks local adaptation dynamics separately from the shared backbone and provides a lightweight signal for detecting whether client adaptation remains active. To further avoid repeated online probing over multiple candidate depths, we design a shallow sufficiency estimator that combines cross-client semantic alignment, temporal interface stability, and prompt-state variation to estimate whether the shallowest split is already sufficient. Finally, we introduce a split-compatible interface harmonization module that projects activations from different split depths into a shared semantic space, improving the comparability of heterogeneous client interfaces before server-side prediction. Extensive experiments on multiple heterogeneous benchmarks demonstrate the effectiveness of \textsc{FedSGA} in improving model performance compared with state-of-the-art methods while reducing unnecessary client-side computation.
△ Less
Submitted 20 August, 2026; v1 submitted 16 August, 2026;
originally announced August 2026.
-
Superalgebras and Algebras with Involution: Classifying Cubic Codimension Sequences
Authors:
Yan-Hong Bao,
Jiang-Nan Xu,
Yuan-Feng Zhang
Abstract:
A $\varphi$-algebra is either a superalgebra or an algebra with involution. In this paper, we study $\operatorname{T}^\varphi$-ideals associated with unital $\varphi$-algebras whose $\varphi$-codimension sequence exhibits cubic polynomial growth. As a consequence, we obtain a complete classification of all $\varphi$-codimension sequences of cubic growth for unital $\varphi$-algebras. Furthermore,…
▽ More
A $\varphi$-algebra is either a superalgebra or an algebra with involution. In this paper, we study $\operatorname{T}^\varphi$-ideals associated with unital $\varphi$-algebras whose $\varphi$-codimension sequence exhibits cubic polynomial growth. As a consequence, we obtain a complete classification of all $\varphi$-codimension sequences of cubic growth for unital $\varphi$-algebras. Furthermore, we explicitly determine a minimal-degree multilinear generator for every $\operatorname{T}^\varphi$-ideal associated with unital $\varphi$-algebras whose $\varphi$-codimension growth is at most quadratic.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
Off-Grid Position Optimization under Mutual Coupling in Fluid Antenna Arrays
Authors:
Jingyuan Xu,
Jian Dang,
Zaichen Zhang
Abstract:
Fluid antenna arrays exploit continuous antenna repositioning within a finite aperture to provide geometry diversity beyond grid-constrained port selection. Every displacement, however, changes both the radiation response and the multiport mutual-impedance network, coupling geometry optimization with the source-voltage constraint. This paper develops an electromagnetic-aware (EM-aware) beamforming…
▽ More
Fluid antenna arrays exploit continuous antenna repositioning within a finite aperture to provide geometry diversity beyond grid-constrained port selection. Every displacement, however, changes both the radiation response and the multiport mutual-impedance network, coupling geometry optimization with the source-voltage constraint. This paper develops an electromagnetic-aware (EM-aware) beamforming framework for planar fluid antenna arrays. Phase retrieval converts an amplitude-only shaped-beam specification into an aperture-compatible complex target, and an EM-aware orthogonal matching pursuit (OMP) method selects grid-constrained initial antenna positions. Continuous refinement then alternates exact voltage-constrained current optimization with movement-constrained projected adaptive moment estimation (Adam) updates of all physical antenna positions. Across independently perturbed symmetric dual-beam targets, the proposed method consistently improves the average mainlobe signal-to-noise ratio (SNR) and reduces the peak sidelobe level (PSLL) over a uniform array and discrete port selection.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
Sparse Port Selection under Mutual Coupling in Fluid Antenna Arrays
Authors:
Jingyuan Xu,
Haoyu Liang,
Zaichen Zhang,
Jian Dang
Abstract:
Fluid antenna systems obtain spatial degrees of freedom by reconfiguring antenna positions within a confined region, a principle that extends to beamforming: shaped beams can be synthesized using far fewer radio-frequency feeds than candidate antenna positions. When the candidates are densely arranged, however, electromagnetic mutual coupling changes the relationship among terminal voltages, induc…
▽ More
Fluid antenna systems obtain spatial degrees of freedom by reconfiguring antenna positions within a confined region, a principle that extends to beamforming: shaped beams can be synthesized using far fewer radio-frequency feeds than candidate antenna positions. When the candidates are densely arranged, however, electromagnetic mutual coupling changes the relationship among terminal voltages, induced currents, and radiated fields, so an uncoupled model no longer describes the hardware and may activate an unsuitable set of ports, distorting the synthesized pattern. This paper develops a mutual-coupling-aware framework that converts the desired beam amplitude into a finite-aperture-compatible complex target and models the complete antenna lattice as a coupled multiport network, selecting the active ports and their source voltages through the coupled voltage-to-field response. Inactive candidate ports remain part of the network and carry induced currents, and every compared design is evaluated through the same electromagnetic model under the same source-voltage budget. Numerical results show that the mutual-coupling-aware design improves both the average mainlobe signal-to-noise ratio (SNR) and the peak sidelobe level (PSLL) over coupling-unaware selection and a fixed array, demonstrating that mutual coupling should be exploited in the design itself rather than compensated only in the final evaluation.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
Hamiltonian paths in the permutation digraphs $P(n,n-2)$
Authors:
Jiaxin Guo,
Ming Duan,
Jie Xue
Abstract:
For $1\leq k<n$, let $P(n,k)$ be the directed overlap graph whose vertices are the $k$-permutations of $[n]$ and whose arcs are the $(k+1)$-permutations. Isaak proved that $P(n,n-2)$ has no directed Hamiltonian cycle for $n\geq4$ and asked whether it nevertheless has a directed Hamiltonian path. We answer this question affirmatively by showing that $P(n,n-2)$ has a Hamiltonian path.
For $1\leq k<n$, let $P(n,k)$ be the directed overlap graph whose vertices are the $k$-permutations of $[n]$ and whose arcs are the $(k+1)$-permutations. Isaak proved that $P(n,n-2)$ has no directed Hamiltonian cycle for $n\geq4$ and asked whether it nevertheless has a directed Hamiltonian path. We answer this question affirmatively by showing that $P(n,n-2)$ has a Hamiltonian path.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
Self-Synchronized Terahertz and X-Ray Free-Electron Lasers from a Single Pre-Bunched Electron Beam
Authors:
Yin Kang,
Kaiqing Zhang,
Zhen Wang,
Cheng Yu,
Zhangfeng Gao,
Wencai Cheng,
Hang Luo,
Yue Wang,
Hanghua Xu,
Xiaoqing Liu,
Jinguo Wang,
Huan Zhao,
Yanyan Zhu,
Yongmei Wen,
Fei Gao,
Yangyang Lei,
Chengcheng Xiao,
Liping Sun,
Yongfang Liu,
Jiaqiang Xu,
Weiyi Yin,
Xingtao Wang,
Taihe Lan,
Zheng Qi,
Tao Liu
, et al. (5 additional authors not shown)
Abstract:
Ultrafast pump-probe spectroscopy combining intense terahertz (THz) and X-ray pulses is a critical tool for investigating complex structural and electronic dynamics in materials. However, current setups combining THz sources and X-ray free-electron lasers (FELs) often suffer from high system complexity, inherent timing jitter, or limited THz pulse properties. Here, we experimentally demonstrate th…
▽ More
Ultrafast pump-probe spectroscopy combining intense terahertz (THz) and X-ray pulses is a critical tool for investigating complex structural and electronic dynamics in materials. However, current setups combining THz sources and X-ray free-electron lasers (FELs) often suffer from high system complexity, inherent timing jitter, or limited THz pulse properties. Here, we experimentally demonstrate the generation of intrinsically synchronized, strong-field, narrow-band THz and X-ray FELs from a single pre-bunched electron beam. Sequentially passing the beam through X-ray and THz amplifiers reveals a highly synergistic process: the initial periodic THz density modulation notably boosts the X-ray FEL pulse energy, while robustly surviving the intense X-ray emission to drive high-power, narrow-band THz radiation. Originating from the same electron bunch, the two pulses inherently maintain a precise, constant time delay. This jitter-free scheme establishes a highly reliable platform tailored for both X-ray-pump/THz-probe and THz-pump/X-ray-probe experiments.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
MEDR: Query-Independent Frame Selection via Multi-Signal Event Modeling and Dynamic Rescoring
Authors:
Xinlei Pu,
Weijie Shi,
Wen Yang,
Yi Cao,
Hao Chen,
Yuanjun Liu,
Wenwei Ding,
Jia Zhu,
Jiajie Xu
Abstract:
Frame selection is a fundamental component of multimodal large language models, enabling long videos to be processed under limited visual-token and computational budgets. Uniform sampling preserves temporal coverage but may miss informative content that appears only briefly. To alleviate this limitation, query-dependent methods can retrieve question-relevant frames. However, because the selected f…
▽ More
Frame selection is a fundamental component of multimodal large language models, enabling long videos to be processed under limited visual-token and computational budgets. Uniform sampling preserves temporal coverage but may miss informative content that appears only briefly. To alleviate this limitation, query-dependent methods can retrieve question-relevant frames. However, because the selected frames depend on the current question, the same visual input cannot be directly shared across different questions, and frame selection must be repeated in multi-turn video dialogue. This motivates us to seek a query-independent frame selection method that preserves the reusability of a fixed visual input while improving the coverage of informative events beyond uniform sampling. We propose Multi-Signal Event Modeling and Dynamic Rescoring (MEDR), a training-free and query-independent frame selection method. Multi-Signal Event Modeling organizes complementary visual, motion, and text signals into signal-specific temporal events. Dynamic Rescoring then iteratively reevaluates each candidate relative to the current selected set, updating its score according to frame-level signal strength, additional event coverage, and temporal proximity. The resulting fixed frame set is constructed without observing the query and can be reused across different questions. On the standard benchmark evaluations, MEDR improves model accuracy by 0.63%-0.89% on Video-MME. On the long-video subset of LongVideoBench, it improves accuracy by up to 1.23% with Qwen3-VL-8B. MEDR further improves overall accuracy by 0.53%, while reusing exactly the same frame set for every question about a video.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
Infrared Spectroscopy of Cyanonaphthalenes under Interstellar Relevant Conditions and Their Potential Connection with Astronomical Aromatic Infrared Bands
Authors:
Jiaqi Xin,
Jianzhi Xu,
Piero Ferrari,
Gao-Lei Hou
Abstract:
Context. Aromatic infrared bands (AIBs) are widely observed in diverse astrophysical environments and are generally attributed to vibrational emission from polycyclic aromatic hydrocarbons (PAHs). The recent interstellar detection of 1-cyanonaphthalene (1-CNN) and 2-cyanonaphthalene (2-CNN) has motivated detailed infrared spectroscopic studies of cyano-substituted PAHs. Aims. We aim to characteriz…
▽ More
Context. Aromatic infrared bands (AIBs) are widely observed in diverse astrophysical environments and are generally attributed to vibrational emission from polycyclic aromatic hydrocarbons (PAHs). The recent interstellar detection of 1-cyanonaphthalene (1-CNN) and 2-cyanonaphthalene (2-CNN) has motivated detailed infrared spectroscopic studies of cyano-substituted PAHs. Aims. We aim to characterize the infrared spectra and vibrational modes of neutral 1-CNN and 2-CNN under cold and gas-phase conditions and to assess their possible spectroscopic relevance to the astronomical AIBs. Methods. The gas-phase infrared spectra of neutral 1-CNN and 2-CNN were measured in a cold molecular beam using ion-dip spectroscopy. The observed bands were assigned with the aid of harmonic and anharmonic calculations at the B3LYP/N07D level. Infrared emission spectra were subsequently simulated from the experimental spectra within a single-photon approximation framework. Results. We report the infrared spectra of neutral 1-CNN and 2-CNN measured under cold and gas-phase conditions relevant to the interstellar medium. Their vibrational features were assigned in detail, including fundamental vibrations as well as overtone and combination bands. The simulated emission spectra exhibit features in several wavelength regions associated with prominent AIBs, including the aromatic CH stretching region near 3.3 micron, the CC stretching region near 6.2 micron, the mixed CH in-plane bending and CC stretching region at 8.6-8.9 microns, and the CH out-of-plane bending region between 10 and 15 microns. Conclusions. The present spectra provide laboratory reference data for small cyano-substituted PAHs and offer useful clues for interpreting selected AIB regions. These results suggest that cyanonaphthalene molecules are promising contributors to the aromatic infrared bands.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Best Reaction Target To Determine Proton Distribution Radii of Atomic Nuclei
Authors:
Jun-Yao Xu,
Bao-Hua Sun,
Isao Tanihata,
Satoru Terashima,
Jian-Wei Zhao,
Ji-Chao Zhang,
Ge Guo,
Shi-Tao Wang,
Lei Shen,
Jun Su,
Xiao-Dong Xu,
Andrej Prochazka,
Guang-Shuai Li,
Xiu-Lin Wei,
Chang-Jian Wang,
Feng Wang,
Meng Wang,
Jing Wang,
Liu-Chun He,
Chuan-Ye Liu,
Wen-Jian Lin,
Wei-Ping Lin,
Zhong Liu,
Pei-Pei Ren,
Yu Zhang
, et al. (7 additional authors not shown)
Abstract:
We found that a heavy target such as Pb is most suitable for determining the proton distribution radii of unstable nuclei through charge-changing cross-section ($σ_\text{cc}$) measurements. As a heavy ion probe, low-$Z$ targets are routinely used to determine nucleon distribution radii of unstable isotopes. This approach has recently been extended to study proton distribution radii from…
▽ More
We found that a heavy target such as Pb is most suitable for determining the proton distribution radii of unstable nuclei through charge-changing cross-section ($σ_\text{cc}$) measurements. As a heavy ion probe, low-$Z$ targets are routinely used to determine nucleon distribution radii of unstable isotopes. This approach has recently been extended to study proton distribution radii from $σ_\text{cc}$ measurements. However, empirical scaling factors have to be introduced to apply the Glauber models. In the present work, we systematically investigated the scaling factor using 39 new $σ_\text{cc}$ data of 18 $p$-shell nuclei on hydrogen, carbon, silver, and lead targets at around 240 MeV/nucleon. Together with the existing data, we reveal a universal dependence of the scaling factor on both the masses of target nuclei and the separation energies of projectile nuclei. The scaling factors decrease with increasing target-nucleus mass and converge to 1 for the highest-$Z$ target, making the scaling unnecessary. We conclude that instead of a low-$Z$ target, employing a heavy target such as Pb in $σ_\text{cc}$ measurements is the best option to determine the proton distribution radii of unstable nuclei.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures
Authors:
Haibo HU,
Lianming Huang,
Qiao Li,
Nan Guan,
Chun Jason Xue
Abstract:
Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure. In a full autonomous driving stack, this problem is even more pronounced: legacy vehicle platforms were provisioned for modular pipelines, yet afte…
▽ More
Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure. In a full autonomous driving stack, this problem is even more pronounced: legacy vehicle platforms were provisioned for modular pipelines, yet after several planning-related functions are absorbed into a unified VLA model, part of the original CPU budget becomes underutilized, while the visual encoder and the main reasoning path still concentrate most computation and memory demand on the GPU. As a result, directly deploying VLA together with the rest of the onboard system can be hard under realistic GPU memory constraints. To address this issue, we present a hybrid CPU--GPU inference framework with flexible resource scheduling for autonomous driving. Our design partitions the VLA backbone at the block-layer granularity, executes the visual encoder and LLM prefix on the GPU, and offloads the LLM suffix to the CPU through a cross-frame asynchronous pipeline, thereby exposing a schedulable boundary for redistributing compute and memory pressure across heterogeneous processors. We evaluate the proposed framework on two representative driving VLA models, Orion and MindDrive. On Bench2Drive, our method reduces average latency from 521ms to 408.0ms for Orion and from 443ms to 306.2ms for MindDrive, corresponding to 21.7% and 30.9% reduction, respectively. For Orion, the estimated peak GPU memory is further reduced from 45GB to 29GB. In real-vehicle deployment under coexistence with Autoware.Universe, native Orion cannot run because the onboard GPU memory budget is insufficient, whereas the hybrid version runs successfully together with the full vehicle stack.
△ Less
Submitted 18 June, 2026;
originally announced August 2026.
-
HW-Router: Hardware-Aware Routing for Scalable Multi-LLM Serving
Authors:
Ahasan Kabir,
Jiaqi Xue,
Mengxin Zheng,
Qian Lou
Abstract:
Modern large language model (LLM) serving platforms deploy multiple models across different GPUs, requiring routers to direct incoming queries to appropriate LLMs. However, existing routing approaches primarily rely on static model attributes such as size or FLOPs to estimate serving costs. This static cost modeling fails to capture the dynamic behavior of real deployments, where the same model ca…
▽ More
Modern large language model (LLM) serving platforms deploy multiple models across different GPUs, requiring routers to direct incoming queries to appropriate LLMs. However, existing routing approaches primarily rely on static model attributes such as size or FLOPs to estimate serving costs. This static cost modeling fails to capture the dynamic behavior of real deployments, where the same model can exhibit vastly different inference latencies depending on hardware type (e.g., H100 vs. V100), current system load (e.g., running and waiting queue lengths), and resource contention (e.g., KV-cache usage and GPU utilization). Such hardware-agnostic routing leads to suboptimal decisions, resulting in SLO violations, queue buildup, and underutilized GPUs. To address these challenges, we present HW-Router, a dynamic routing framework that integrates real-time hardware signals into model selection to enable accurate latency prediction and intelligent, SLO-aware routing decisions. Our approach incorporates model-specific features (architecture, size, input length) alongside hardware metrics including queue lengths, KV-cache utilization, and recent TTFT/TPOT performance, and uses a lightweight latency predictor to estimate per-model-per-GPU serving time. Evaluations across diverse workloads show that HW-Router achieves 3.4-3.9x lower end-to-end latency, 46-48 percentage points higher SLO attainment, 6-8x lower GPU load skew, and a 3.1-3.4x reduction in waiting-queue fraction compared to state-of-the-art router baselines, CARROT and IRT, with only ~200 us of additional routing overhead and no loss in output quality. These results highlight the importance of real-time hardware feedback for scalable, predictable, and well-balanced multi-LLM serving. Code is available at https://github.com/UCF-ML-Research/HW-Router.
△ Less
Submitted 10 June, 2026;
originally announced August 2026.
-
WARA: Toward Automated Wireless Optimization Research with Closed-Loop LLM Agents
Authors:
Yuan Guo,
Yilong Chen,
Chao Hu,
Xianghao Yu,
Liang Hong,
Jie Xu
Abstract:
Large language model (LLM) agents are increasingly capable of tool use, code execution, artifact inspection, and iterative revision, creating new opportunities for automating scientific and engineering research. To the best of our knowledge, this paper presents the first end-to-end autoresearch framework for the wireless domain, with a focus on wireless resource allocation optimization. We propose…
▽ More
Large language model (LLM) agents are increasingly capable of tool use, code execution, artifact inspection, and iterative revision, creating new opportunities for automating scientific and engineering research. To the best of our knowledge, this paper presents the first end-to-end autoresearch framework for the wireless domain, with a focus on wireless resource allocation optimization. We propose the Wireless AutoResearch Agent (WARA), a closed-loop multi-agent system for automated wireless optimization research. Given only an initial topic, WARA decomposes the workflow into three phases: research gap identification and problem proposal, wireless optimization modeling, algorithm design and experimentation, and research deliverable construction. Across these phases, WARA uses artifact-mediated control: upstream artifacts are consumed as inputs, structured outputs are stored for downstream use, and controller-managed gates validate consistency among models, algorithms, experiments, and claims. When validation fails, WARA repairs only the responsible artifact instead of restarting the whole workflow. We present a representative wireless resource allocation case study showing how WARA converts an initial topic into a complete research package with executable evidence and a synthesized technical manuscript. We further design a structured LLM-based ScoringAgent to evaluate manuscript-level research validity and optimization research maturity. Comparative results show that WARA substantially outperforms one-shot LLM generation and approaches the quality profile of recently accepted peer-reviewed technical papers. These results indicate that closed-loop artifact control is a promising path toward end-to-end LLM-assisted wireless optimization research. The source code is available at https://github.com/guoyuan-dotcom/WARA_CUHKSZ.
△ Less
Submitted 6 June, 2026;
originally announced August 2026.
-
Singer-Informed Vocal Source Separation for Multi-Singer Music Mixtures
Authors:
Jocelyn Xu,
Minje Kim
Abstract:
Music source separation systems typically extract a single vocal track and do not distinguish between multiple singers. We study singer-informed vocal source separation for multi-singer mixtures. Our framework introduces a short enrollment recording of a target singer to guide separation through a learned embedding. The singer embedding is incorporated using feature concatenation or feature-wise l…
▽ More
Music source separation systems typically extract a single vocal track and do not distinguish between multiple singers. We study singer-informed vocal source separation for multi-singer mixtures. Our framework introduces a short enrollment recording of a target singer to guide separation through a learned embedding. The singer embedding is incorporated using feature concatenation or feature-wise linear modulation (FiLM), enabling the model to focus on the target singer while suppressing interference. We construct a duet dataset based on DAMP-VSEP with quality filtering and non-overlapping enrollment segments. Experiments on solo and duet settings show that while baseline models perform well for single-singer mixtures, the proposed method improves target-singer extraction in multi-singer cases, increasing target-singer SI-SDR from 0.33 dB to 5.58 dB. Fréchet Audio Distance (FAD) further shows improved perceptual quality and better alignment with target audio distributions. Code and checkpoints are available at https://github.com/jocelynxu01/singer-separation-paper.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
THRIVE: Therapeutic Humanoid Robot In Virtual Environment
Authors:
Jin Xu,
Yu-Ping Chen,
Ayanna Howard
Abstract:
This paper presents THRIVE (Therapeutic Humanoid Robot In Virtual Environment), an at-home rehabilitation platform that integrates a suite of virtual-reality upper-body rehabilitation games, a real-time camera-based motion-tracking system, and a socially interactive robot therapist. The system is designed for therapy and intervention in children with upper-limb motor impairments, which can be impr…
▽ More
This paper presents THRIVE (Therapeutic Humanoid Robot In Virtual Environment), an at-home rehabilitation platform that integrates a suite of virtual-reality upper-body rehabilitation games, a real-time camera-based motion-tracking system, and a socially interactive robot therapist. The system is designed for therapy and intervention in children with upper-limb motor impairments, which can be improved through consistent, task-specific practice. THRIVE features a set of newly designed, engaging games that target functional reaching, grasping, and object-manipulation movements through customizable popping, hitting, catching, and grabbing tasks, while the camera-based tracking system captures the child's kinematic performance during play. A robot therapist - deployable either as a physical robotic coach or as a remote-presence virtual agent - delivers adaptive, dynamic feedback to motivate the child and guide their movements toward therapeutic goals. THRIVE decouples the therapeutic games from the robot embodiment, extending the platform to support various embodiments and different robots within one modular system. This robot-agnostic design makes THRIVE affordable, scalable, and readily adaptable for sustained use in the home, offering a practical pathway to more consistent and engaging upper-limb therapy for children with motor function impairments.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Authors:
Kai Chen,
Jifeng Ding,
Ning Ding,
Jiaye Ge,
Lixin Gu,
Yicheng Gu,
Qipeng Guo,
Ermo Hua,
Haian Huang,
Haozheng Hou,
Jie Hou,
Xiangyu Hong,
Che Jiang,
Minxi Jin,
Cheng Liang,
Dahua Lin,
Dawei Liu,
Kuikun Liu,
Chengqi Lv,
Haijun Lv,
Han Lv,
Ningsheng Ma,
Biqing Qi,
Jianmin Qian,
Shiya Su
, et al. (22 additional authors not shown)
Abstract:
We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning. Using hidden states as cache and carrier, reasoners repeatedly query memory for required knowledge-vectors, while the knowledge is transmitted back to reasoning operators. Through this knowledge-reas…
▽ More
We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning. Using hidden states as cache and carrier, reasoners repeatedly query memory for required knowledge-vectors, while the knowledge is transmitted back to reasoning operators. Through this knowledge-reasoning-separation architecture, Mobius achieves better knowledge compression and reasoning efficiency. Built upon Mobius-v0 architecture: 1) Our 7B model trained-from-scratch achieves similar downstream score as a 7B Transformer baseline with 62.6% of baseline's training data. 2) Our Intern-S2-Mobius, continually-pretrained from Qwen3.5-35B, achieves similar downstream score while delivering nearly 4x end-to-end inference speedup.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Equivalent norms and $\varphi$-transform of matrix-weighted anisotropic Besov-type and Triebel-Lizorkin-type spaces
Authors:
Tengfei Bai,
Pengfei Guo,
Jingshi Xu
Abstract:
We introduce matrix-weighted anisotropic Besov-type and Triebel-Lizorkin-type spaces associated with an expansive matrix $A$. Inspired by the $\A_p$-dimensions of matrix weight of Bu et al. (2025), we study the properties of matrix weight associated with $A$. Using the nice properties of matrix weight, we obtain that these spaces are equivalent with their corresponding averaging spaces and establi…
▽ More
We introduce matrix-weighted anisotropic Besov-type and Triebel-Lizorkin-type spaces associated with an expansive matrix $A$. Inspired by the $\A_p$-dimensions of matrix weight of Bu et al. (2025), we study the properties of matrix weight associated with $A$. Using the nice properties of matrix weight, we obtain that these spaces are equivalent with their corresponding averaging spaces and establish the discrete $\varphi$-transform of these spaces. Finally, we introduce matrix-weighted anisotropic Triebel-Lizorkin spaces for the limiting case $p=\infty$ and obtain their $\varphi$-transform. The relation between matrix-weighted anisotropic Triebel-Lizorkin spaces and matrix-weighted anisotropic Besov-type and Triebel-Lizorkin-type spaces is also studied.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Intern-S2-Preview: Scientific Agentic Foundation Model
Authors:
Lei Bai,
Jiaqi Cao,
Chiyu Chen,
Guanzhou Chen,
Kai Chen,
Guangran Cheng,
Erfei Cui,
Xuanlang Dai,
Shengyuan Ding,
Shangheng Du,
Yanhui Duan,
Yue Fan,
Youqing Fang,
Quan Gan,
Yuanyuan Gao,
Jiaye Ge,
Lixin Gu,
Yuzhe Gu,
Qipeng Guo,
Junjun He,
Xin Hong,
Ming Hu,
Zhouqi Hua,
Haian Huang,
Junhao Huang
, et al. (100 additional authors not shown)
Abstract:
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tas…
▽ More
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Bricks that every removable edge is solitary
Authors:
Jinxin Xue,
Jun Ge,
Fuliang Lu,
Yaxian Zhang
Abstract:
A brick is a 3-connected graph $G$ such that $G-u-v$ has a perfect matching for any two distinct vertices $u,v\in V(G)$. An edge $e$ in a matching covered graph $G$ is removable if $G-e$ is matching covered. We say that a removable edge $e$ in a brick $G$ is $b$-invariant if $b(G-e)=b(G)=1$, where $b(H)$ denotes the number of bricks in the tight cut decomposition of a matching covered graph $H$. A…
▽ More
A brick is a 3-connected graph $G$ such that $G-u-v$ has a perfect matching for any two distinct vertices $u,v\in V(G)$. An edge $e$ in a matching covered graph $G$ is removable if $G-e$ is matching covered. We say that a removable edge $e$ in a brick $G$ is $b$-invariant if $b(G-e)=b(G)=1$, where $b(H)$ denotes the number of bricks in the tight cut decomposition of a matching covered graph $H$. An edge of a graph is solitary if it lies in precisely one perfect matching.
Lucchesi and Murty proposed the problem of characterizing bricks, distinct from $K_4$, $\overline{C_6}$ and the Petersen graph, in which every $b$-invariant edge is solitary. Note that every $b$-invariant edge is removable. In this paper, we strengthen the condition by requiring that every removable edge is solitary. We show that every nonsolid brick satisfying this strengthened condition can be obtained by repeatedly splicing odd wheels (up to multiple edges). Moreover, properties of such bricks imply that "repeatedly splicing odd wheels" cannot be replaced by "repeatedly splicing copies of $K_4$".
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
High-precision measurement of the space-like $η^\prime$ transition form factor
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (758 additional authors not shown)
Abstract:
Using a data sample corresponding to an integrated luminosity of $20.3\ \text{fb}^{-1}$, collected with the BESIII detector at a center-of-mass energy of $3.773\ \text{GeV}$ at the BEPCII collider, we report a precision measurement of the product $Q^2|F(Q^2)|$, where $F(Q^2)$ is the single-virtual space-like transition form factor of the $η'$ meson and $Q^2$ is the squared momentum transfer of the…
▽ More
Using a data sample corresponding to an integrated luminosity of $20.3\ \text{fb}^{-1}$, collected with the BESIII detector at a center-of-mass energy of $3.773\ \text{GeV}$ at the BEPCII collider, we report a precision measurement of the product $Q^2|F(Q^2)|$, where $F(Q^2)$ is the single-virtual space-like transition form factor of the $η'$ meson and $Q^2$ is the squared momentum transfer of the tagged virtual photon. The transition form factor is extracted from the differential Born cross section of the two-photon fusion processes $e^+e^- \to e^+e^-γγ^* \to e^+e^-η^\prime$ using a single-tag technique, where only one scattered lepton is detected. The measurement covers $Q^2 \in [0.1, 6.0]$ GeV$^2$, achieving unprecedented precision, better than $3.0\%$ for $Q^2 < 1.5$ GeV$^2$, and providing the first direct determination at $Q^2 < 0.3$ GeV$^2$.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization
Authors:
Yuanyu Li,
Jintao Xu,
Zijiang Liu,
Yongzhi Qi,
Ningxuan Kang,
Jianshen Zhang,
Wei Qi,
Chen Xie,
Zuo-Jun Max Shen
Abstract:
Neural combinatorial optimization (NCO) relies on parallel solution sampling for training, yet existing methods fail to fully exploit the rich information latent in a co-sampled solution group. Preference-optimization methods anchor on the single best solution and discard fine-grained quality and structural signal from all other peers-a failure we term gradient signal polarization. Mean-based base…
▽ More
Neural combinatorial optimization (NCO) relies on parallel solution sampling for training, yet existing methods fail to fully exploit the rich information latent in a co-sampled solution group. Preference-optimization methods anchor on the single best solution and discard fine-grained quality and structural signal from all other peers-a failure we term gradient signal polarization. Mean-based baselines instead weight peers uniformly, so structurally near-identical peers flood the baseline with redundant information and keep gradient variance high-a failure we term baseline redundancy. We propose SSPO (Structure-Aware Similarity-Weighted Preference Optimization), which scores all $B$ sampled solutions jointly through a dissimilarity-weighted leave-one-out baseline: structurally distinct peers receive higher weight, resolving both failures in a single mechanism. The baseline uses zero-parameter, problem-adaptive solution embeddings built from the encoder's existing node representations. Experiments on TSP, EFL, and JSP benchmarks show consistent gains over prior best-anchor and uniform-weight baselines. A direct comparison against uniform RLOO on TSP and EFL confirms that structure-aware weighting is the primary driver of improvement. The SSPO-trained EFL policy has been deployed in a production facility-location system at JD$\mathord{.}$com, confirming practical viability at scale.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Sharp Phase Transition for Ellipsoid Fitting
Authors:
Sofia de la Cerda,
Aaron Potechin,
Madhur Tulsiani,
Jeff Xu
Abstract:
We resolve the ellipsoid fitting conjecture of Saunderson, Chandrasekaran, Parrilo, and Willsky up to a vanishing factor. Concretely, for $m$ independent Gaussian points in dimension $d$, we show that with high probability,
for $m \leq (1-o_d(1)) \cdot d^2/4$, there exists a centered ellipsoid passing through all $m$ points;
for $m\geq (1+o_d(1) )\cdot d^2/4$, no such ellipsoid exists.
This…
▽ More
We resolve the ellipsoid fitting conjecture of Saunderson, Chandrasekaran, Parrilo, and Willsky up to a vanishing factor. Concretely, for $m$ independent Gaussian points in dimension $d$, we show that with high probability,
for $m \leq (1-o_d(1)) \cdot d^2/4$, there exists a centered ellipsoid passing through all $m$ points;
for $m\geq (1+o_d(1) )\cdot d^2/4$, no such ellipsoid exists.
This confirms that the ellipsoid fitting problem has a sharp phase transition at $d^2/4$.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents
Authors:
Junliang Liu,
Ruoyu Li,
Wenxin Tang,
Jingyu Xiao,
Zhenyu Liu,
Jingheng Xu,
Laizhong Cui
Abstract:
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publishers: a static skill may steer an otherwise correct task onto an unnecessarily costly trajectory. Prior work studies selection manipulation, malicious skill instructions…
▽ More
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publishers: a static skill may steer an otherwise correct task onto an unnecessarily costly trajectory. Prior work studies selection manipulation, malicious skill instructions, and tool-chain resource amplification largely separately, leaving their end-to-end composition unclear. We introduce Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack that couples these stages. Under shared semantic cover, a description establishes relevance during selection, while an aligned body reuses that rationale to fabricate plausible dependencies during planning. CDH attracts an attacker-controlled coordinator alongside legitimate skills, recruits unnecessary benign skills into a bounded detour, and then re-enters the original route to preserve task completion. We evaluate it across multiple LLM backends and 491 held-out tasks under single-task and multi-turn conditions. On DeepSeek-V4-Pro, the matched coordinator is selected in 80.02% of tasks; among coordinator-hit runs that complete tasks, token consumption and end-to-end execution time increase by 66.91% and 92.45%, respectively, while aggregate task completion remains comparable. Thus, correct outcomes do not guarantee trajectory integrity or cost safety.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Reconfiguring Geovisualization in the Age of Generative AI: Insights from Domain Experts
Authors:
Mengyi Wei,
Chenyu Zuo,
Jiaying Xue,
Nianhua Liu,
Dongsheng Chen,
Shengkai Wang,
Yu Feng,
Liqiu Meng
Abstract:
GenAI is increasingly integrated into geovisualization, yet its broader implications for professional practice are insufficiently understood. To examine these implications, we conducted semi-structured interviews with 20 geovisualization experts. The interviews were structured around four broad analytical domains: Data, Ideation, Prototyping, and Iteration, while also encouraging participants to r…
▽ More
GenAI is increasingly integrated into geovisualization, yet its broader implications for professional practice are insufficiently understood. To examine these implications, we conducted semi-structured interviews with 20 geovisualization experts. The interviews were structured around four broad analytical domains: Data, Ideation, Prototyping, and Iteration, while also encouraging participants to reflect on issues that extend beyond these activities. Our findings show that GenAI expands the capabilities of geovisualization, particularly in terms of data handling, creative exploration, and rapid prototyping, but does not simply remove existing constraints. Instead, key bottlenecks are shifting from production to judgment and verification. As routine technical tasks become more automated, professional value increasingly depends on spatial reasoning, contextual interpretation, aesthetic and ethical judgment, and the ability to assess whether AI-generated outputs are appropriate for use. At the same time, GenAI introduces new challenges regarding provenance, interpretability, and accountability, raising questions about how responsibility should be distributed across models, developers, practitioners, institutions, and users. These shifts are particularly significant in geovisualization because spatial representations are constrained by geographic reality and must balance scientific validity, visual expression, and technical implementation. We therefore argue that responsible GenAI in geovisualization requires domain-specific approaches to spatial validation, provenance, uncertainty communication, human oversight, and accountable use. This study provides an expert-grounded perspective on how GenAI is reconfiguring geovisualization as a practice of spatial knowledge production. It also identifies implications for future professional practice, education, system design, and governance.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Proof-Valid Caching under Premise Erasures: Local Structural Limits and Shared-Workload Gains
Authors:
Jianfeng Xu
Abstract:
We study reliable query recovery under independent premise erasures in semantically transparent caching systems, where every cached object must be a logical consequence of the premise base. Recovery succeeds only when the query remains derivable from surviving premises and the cache. Under a deterministic canonical-witness regime, we prove a query-local projection theorem and an exact residual-lea…
▽ More
We study reliable query recovery under independent premise erasures in semantically transparent caching systems, where every cached object must be a logical consequence of the premise base. Recovery succeeds only when the query remains derivable from surviving premises and the cache. Under a deterministic canonical-witness regime, we prove a query-local projection theorem and an exact residual-leaf law: recovery fails exactly when an erased base leaf retains a cache-free path to the query. Single-query design becomes weighted partial path interception. For shared workloads, we introduce semantic modules and derive exact reliability laws under joint and maximal-error criteria. The shared-module cache is exactly optimal under exact module routing and homogeneous costs, whereas optimal selection in general derivation DAGs is NP-complete at depth two. Against a coded benchmark recovering workload-relevant leaf payloads, MDS parity caching is optimal up to one packet. Leaf-only transparency incurs a first-order overhead inversely proportional to the erasure rate; shared modules multiply that inverse-erasure-rate scaling by the module-to-leaf cost ratio divided by the number of protected leaves. A Datalog instance and Monte Carlo checks illustrate the theory. For derivation-structured content, the results provide exact stochastic-erasure counterparts of function-correcting storage and an exact distributional quantification of maximal recoverability.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
Authors:
Kai Yang,
Jingwei Xu,
Wanyu Wang,
Kai-Yuan Guo,
Zhenbo Yu,
Yi Wang,
Yu Qiao
Abstract:
On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation. Although prior work has characterized the subspace geometry of aggregate updates, the stepwise variation of this geometry and its relationship to model performance remain unclear. We i…
▽ More
On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation. Although prior work has characterized the subspace geometry of aggregate updates, the stepwise variation of this geometry and its relationship to model performance remain unclear. We introduce Principal-Subspace Overlap, a dimension-corrected measure of individual rollout updates relative to the dominant singular subspaces of pretrained weights. Despite low average overlap, transient spikes often precede performance degradation. To address this, we propose GCPO (Geometrically Constrained Policy Optimization), which applies hard bilateral orthogonal projections to constrain updates to the complementary subspaces, preventing such excursions by construction. Across mathematical reasoning, code generation, and tool-use tasks on Qwen3-8B and GLM4-9B, GCPO consistently outperforms GRPO and recent variants, including DAPO and GSPO, improving over the base models and the strongest baseline by up to 27.69 and 2.37 points, respectively. Furthermore, GCPO preserves general capabilities, eliminates response-length inflation, and stabilizes policy entropy. Our findings provide a new diagnostic lens and a principled design perspective for stable reinforcement learning post-training.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Spin lifetime anisotropy in graphene induced by the SiO2 interface
Authors:
Aron W. Cummings,
Chunhao Guo,
Andrew Grieder,
Shihao Tu,
Mayank Gupta,
Junqing Xu,
Juan Marmolejo-Tejada,
Yuan Ping
Abstract:
Understanding how common dielectric substrates influence the spin transport properties of graphene is essential for advancing graphene-based spintronic technologies. Here we use a comprehensive set of numerical simulations to reveal how a SiO$_2$ substrate modifies the spin texture and governs spin relaxation in graphene. Using first-principles density matrix dynamics simulations, as well as tight…
▽ More
Understanding how common dielectric substrates influence the spin transport properties of graphene is essential for advancing graphene-based spintronic technologies. Here we use a comprehensive set of numerical simulations to reveal how a SiO$_2$ substrate modifies the spin texture and governs spin relaxation in graphene. Using first-principles density matrix dynamics simulations, as well as tight-binding (TB) transport simulations, we quantify the effects of electron-phonon scattering, impurity scattering, and electrostatic disorder on the spin relaxation process. We find that a 2D SiO$_2$ substrate induces a predominantly Rashba-type helical spin texture in graphene, leading to a spin lifetime anisotropy of 1/2. Meanwhile, bulk SiO$_2$ breaks in-plane symmetry in graphene, leading to anisotropic in-plane and out-of-plane components in the spin texture, which we capture with a newly-developed TB model of graphene. Transport simulations under realistic disorder conditions reveal a spin lifetime anisotropy between 0.5 and 1, similar to what is seen in measurements of graphene spin valves on a SiO$_2$ substrate. Our results reveal a more complex picture of spin relaxation at the ubiquitous graphene/SiO$_2$ interface, beyond the standard Rashba model, providing critical insight for interpreting experiments and guiding substrate engineering for graphene spintronics.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
$|ΔI|=3/2$ non-leptonic hyperon decays in covariant baryon chiral perturbation theory
Authors:
Jie Xu,
Jun-Xu Lu,
Rui-Xiang Shi,
Li-Sheng Geng
Abstract:
Inspired by the recent BESIII measurements of non-leptonic hyperon decays, we reexamine their $|ΔI|=3/2$ amplitudes in covariant baryon chiral perturbation theory with the extended-on-mass-shell renormalization scheme. Using the same restricted set of diagrammatic topologies as in the early analyses in heavy baryon chiral perturbation theory, we assess the effects of relativistic corrections, expl…
▽ More
Inspired by the recent BESIII measurements of non-leptonic hyperon decays, we reexamine their $|ΔI|=3/2$ amplitudes in covariant baryon chiral perturbation theory with the extended-on-mass-shell renormalization scheme. Using the same restricted set of diagrammatic topologies as in the early analyses in heavy baryon chiral perturbation theory, we assess the effects of relativistic corrections, explicit decuplet baryons, different spin-$3/2$ coupling schemes, and pion-loop contributions. Our results show that relativistic effects alone lead to only mild changes, while decuplet contributions, especially in the consistent-coupling scheme, significantly improve the fit quality. Pion-loop contributions further reduce the $χ^2$ values. We highlight the importance of the consistent coupling scheme in the decuplet sector for describing the selected $|ΔI|=3/2$ amplitudes.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
Neural Tree Collaborative Filtering: Rethinking Graph Collaborative Filtering as Tree Collaborative Filtering with Curvature-Aware Propagation Depth
Authors:
Jinfeng Xu,
Zheyu Chen,
Ziyue Peng,
Shuo Yang,
Jinze Li,
Wenhao Yuan,
Jian Chen,
Edith C. H. Ngai
Abstract:
Graph Collaborative Filtering (GCF) has become the dominant paradigm in modern recommender systems by modeling user-item interactions as a bipartite graph and propagating embeddings through a fixed number of message-passing layers. However, applying a uniform propagation depth to every node ignores a fundamental property of real interaction graphs: nodes differ substantially in their local connect…
▽ More
Graph Collaborative Filtering (GCF) has become the dominant paradigm in modern recommender systems by modeling user-item interactions as a bipartite graph and propagating embeddings through a fixed number of message-passing layers. However, applying a uniform propagation depth to every node ignores a fundamental property of real interaction graphs: nodes differ substantially in their local connectivity, so peripheral nodes quickly suffer from over-smoothing while hub-like nodes remain under-explored beyond their immediate neighborhood. In this paper, we revisit GCF from a tree-structured perspective and propose Neural Tree Collaborative Filtering (NTCF), a framework that re-interprets each node's local neighborhood as a rooted tree and assigns a node-specific propagation depth based on a closed-form local-degree-imbalance score that serves as a discrete Ricci-curvature proxy. We provide a theoretical analysis showing that (i) NTCF strictly generalizes NGCF, degenerating to NGCF when all curvature-induced depth adjustments vanish (a lower bound on its representation power), and (ii) the curvature-aware schedule retains strictly more discriminative information at deep layers on positively-curved (peripheral) nodes than uniform-depth propagation. NTCF can achieve higher performance than most widely used GCF backbone models and can be integrated into existing advanced self-supervised models as a backbone, replacing their original backbone to achieve enhanced performance. Extensive experiments on three public datasets demonstrate the superiority of NTCF.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
4D-WAM: 4D Consistent World Modeling for Autonomous Driving
Authors:
Jiacheng Fu,
Yibo Yuan,
Meng Tian,
Yue Li,
Jiangtong Zhu,
Jianhua Han,
Yueyi Zhang,
Jianwu Fang,
Jianru Xue,
Hang Xu,
Zhiwei Xiong
Abstract:
Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. However, existing WAMs are typically trained with video data, which is only 2D projections of the underlying 4D driving scene. Consequently, WAMs fail to understand and capture the structure of 4D scenes and thus generate visu…
▽ More
Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. However, existing WAMs are typically trained with video data, which is only 2D projections of the underlying 4D driving scene. Consequently, WAMs fail to understand and capture the structure of 4D scenes and thus generate visually plausible yet 4D inconsistent future predictions that mislead downstream planning. To alleviate this issue, we present 4D-WAM, a model that leverages geometric foundation models for training-time supervision to enable 4D consistent world modeling. Specifically, we feed WAM-predicted future frames into a geometric foundation model, and use 4D-aware responses to define a 4D consistency loss. This loss encourages the model to understand, represent, and predict physically consistent 4D scenes during training, without additional inference cost. Moreover, we identify an early-decision phenomenon in WAMs and propose a decision-oriented timestep sampling strategy that emphasizes supervision at early, high-noise stages, where driving decisions are primarily formed. By propagating 4D supervision to this critical decision-formation phase, the proposed strategy further improves trajectory planning. Extensive experiments demonstrate that 4D-WAM effectively models 4D consistent scene evolution and achieves state-of-the-art performance on challenging NAVSIM-v1 and NAVSIM-v2 benchmarks.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Towards Expert-level Medical AI for Real-time Video Consultations
Authors:
Mahvish Nagda,
Jihyeon Lee,
Matthew Thompson,
Chunjong Park,
Tim Strother,
Valentin Liévin,
Roma Ruparel,
Akshay Goel,
Teya Bergamaschi,
Suhana Bedi,
Meet Shah,
Pavel Dubov,
Liviu Panait,
Toshiyuki Fukuzawa,
Sam Schmidgall,
Craig Schiff,
Joseph Xu,
Aliya Rysbek,
Yana Lunts,
Jan Freyberg,
Rebecca Hemengway,
Sunny Virmani,
David Racz,
Carey Radebaugh,
Joëlle Barral
, et al. (15 additional authors not shown)
Abstract:
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated fea…
▽ More
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
Authors:
Yuling Shi,
Jinghan Xu,
Kelin Fu,
Wenhao Zeng,
Shilin He,
Lei Zhang,
Yue Liu,
Zelin Zhao,
Terry Yue Zhuo,
Jialun Cao,
Siyu Ye,
Tianyu Liu,
Kai Cai,
Shing-Chi Cheung,
Xiaodong Gu
Abstract:
As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found that nearly 60% of unsolved SWE-bench Verified instances contain flawed tests -- either overly narrow tests that reject correct solutions or overly broad tests that check unstated req…
▽ More
As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found that nearly 60% of unsolved SWE-bench Verified instances contain flawed tests -- either overly narrow tests that reject correct solutions or overly broad tests that check unstated requirements -- and that frontier models can verbatim reproduce gold patches from training data. Code refactoring, which requires coordinated, behavior-preserving changes across many files, offers a substantially harder and more realistic test of agent capability, yet remains underserved by current benchmarks. We introduce SWE-Bench ProMax, an expert-curated, multilingual code refactoring benchmark of 170 instances drawn from real commits across seven programming languages (Python, Java, TypeScript, Go, C, C++, and Rust). Every instance undergoes rigorous, multi-stage curation that directly addresses the quality problems identified in prior benchmarks: issue descriptions are rewritten from scratch to provide precise, unambiguous specifications, and test suites are manually reviewed to remove overly narrow and overly broad tests. Tasks with insufficient complexity or limited cross-file scope are filtered out, yielding a benchmark of challenging, large-scale refactoring tasks that average 11.4 modified files and 261.6 lines of code per instance, substantially exceeding the scale of existing benchmarks. Experiments with frontier models under two agent scaffolds show that the best model achieves only 41.2% resolve rate, confirming that SWE-Bench ProMax presents a meaningful and unsaturated challenge for current AI coding agents. Our benchmark is available at https://huggingface.co/datasets/swe-bench-promax/SWE-Bench-ProMax.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.