-
ALOHA IRDCs Molecular Line Follow-up: I. Gas properties and kinematics
Authors:
Jinjin Xie,
Yaoting Yan,
Zhiyuan Ren,
Jarken Esimbek,
Di Li,
Yan Duan,
Gary A. Fuller,
Nicolas Peretto,
Jingwen Wu,
Wenjin Yang,
Christian Henkel,
Xuepeng Chen,
Qianru He,
Yongxiong Wang,
Keping Qiu,
Ningyu Tang,
Sijia Peng,
Chao-Wei Tsai,
Pham Ngoc Diep,
Hauyu Baobab Liu,
Busaba Kramer,
Kee-Tae Kim,
Ken'ichi Tatematsu,
Mark G. Rawlings,
Maria Jesus Jimenez Donaire
, et al. (87 additional authors not shown)
Abstract:
Infrared Dark Clouds are ideal sites for investigating the initial conditions of massive star and cluster formation. The A Lei Of the Habitat and Assembly of Infrared Dark Clouds (ALOHA IRDCs), a James Clerk Maxwell Telescope (JCMT) Large Program, has mapped nearby IRDCs with SCUBA-2. Complementary molecular line observations are needed to characterise the physical, kinematic, and chemical propert…
▽ More
Infrared Dark Clouds are ideal sites for investigating the initial conditions of massive star and cluster formation. The A Lei Of the Habitat and Assembly of Infrared Dark Clouds (ALOHA IRDCs), a James Clerk Maxwell Telescope (JCMT) Large Program, has mapped nearby IRDCs with SCUBA-2. Complementary molecular line observations are needed to characterise the physical, kinematic, and chemical properties of the dense gas. We aim to determine the thermal, kinematic, and chemical properties of clumps identified in the ALOHA IRDCs, and to assess their evolutionary status and level of star-forming activity. We performed single-pointing K-band and W-band observations towards 56 ALOHA IRDCs clumps using the Effelsberg 100-m and Yebes 40-m telescopes, respectively. We derived NH3 kinetic temperatures using the hyperfine group ratio (HFGR) method and identified infall and shock signatures from HCO+, H13CO+, SiO, and HNCO profiles. Water masers and NH2D emission were used as complementary tracers of chemical evolution and star formation. The clumps exhibit kinetic temperatures of 15-29 K. We detect NH2D emission towards 18 sources, with NH2D centroid velocities consistent with NH3, indicating both species trace the same dense gas component. More than half of the clumps display blue-asymmetric HCO+ profiles, identifying them as infall candidates. Water masers are detected in 22 sources, with prominent velocity ranges and variability. Broad SiO emission (>~20 km/s) indicates strong shocks, while narrower extents (<~6km/s) likely trace large-scale interactions or low-velocity shocks. The widespread infall signatures, shock tracers, masers, and NH2D emission suggest that relatively quiescent, chemically young material can coexist with dynamically active gas affected by early protostellar feedback, providing insight into the coupled physical and chemical evolution of massive IRDC clumps.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Learning Early-to-Final Solution Consistency for MILP Acceleration
Authors:
Guanlin Li,
Chengrui Gao,
Chenguang Wang,
Haopu Shang,
Zherong Zhang,
Ke Xue,
Jixiang Lu,
Weiyong Yang,
Chao Qian
Abstract:
Mixed-Integer Linear Programming (MILP) is a fundamental problem class in operations research and combinatorial optimization, with broad applications to industrial decision-making. Owing to their NP-hardness, however, modern solvers may struggle to find high-quality solutions for challenging MILP instances within practical time limits. Recent learning-based approaches seek to accelerate MILP solvi…
▽ More
Mixed-Integer Linear Programming (MILP) is a fundamental problem class in operations research and combinatorial optimization, with broad applications to industrial decision-making. Owing to their NP-hardness, however, modern solvers may struggle to find high-quality solutions for challenging MILP instances within practical time limits. Recent learning-based approaches seek to accelerate MILP solving by directly predicting high-quality solutions from static instance-level features, such as variable-constraint bipartite graphs. Yet accurate solution prediction from instance features alone is difficult, and these methods largely overlook the information revealed during the solver's search process. In this paper, we find that solutions produced at the early search stage of MILP solvers, which are computationally cheap to obtain, are often structurally close to the solutions found after full-budget search. Motivated by this observation, we propose a new solver-informed paradigm that shifts the learning target from variable assignment to early-to-final consistency: for each variable, we predict whether its early-stage assignment should persist in full-budget solutions. The predicted consistency naturally guides downstream search, for instance by fixing the assignments deemed consistent. At inference time, we further ensemble consistency predictions across multiple early-stage solutions to improve robustness. Experiments across four MILP benchmarks show our method improves prediction-guided search across diverse downstream pipelines. With Gurobi, our proposed method reduces the primal gap by 56.9% on average and closes it completely on combinatorial auction instances. Besides, we transferred the Gurobi-trained model zero-shot to SCIP without adaptation, achieving a 36.4% average gap reduction across benchmarks.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
On a classical zero-sum invariant
Authors:
Alfred Geroldinger,
Wenkai Yang
Abstract:
Let $G$ be a nontrivial, finite abelian group. Then $ν(G)$ is the smallest integer $\ell$ such that every zero-sum free sequence $T$ over $G$ of length at least $\ell$ has the following property: all nonzero elements of $G$ that do not occur as a subsequence sum of $T$ lie in a proper coset of some subgroup of $G$. We study the invariant $ν(G)$, which was introduced in Zero-Sum Theory in the 1960s…
▽ More
Let $G$ be a nontrivial, finite abelian group. Then $ν(G)$ is the smallest integer $\ell$ such that every zero-sum free sequence $T$ over $G$ of length at least $\ell$ has the following property: all nonzero elements of $G$ that do not occur as a subsequence sum of $T$ lie in a proper coset of some subgroup of $G$. We study the invariant $ν(G)$, which was introduced in Zero-Sum Theory in the 1960s.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
A characterization of tight ($ k, 0 $)-stable graphs
Authors:
Yuqi Xu,
Weihua Yang,
Xiaxia Guan
Abstract:
Let k and l be two non-negative integers with k > l. A graph G is (k,l)-stable if alpha(G - S) >= alpha(G) - l for every subset S of V(G) with |S| = k, where alpha(G) denotes the independence number of G. Dong and Wu established that alpha(G) <= floor((n - k + 1)/2) + l for a (k, l)-stable graph G, where n is the order of G. A (k, l)-stable graph G is tight if alpha(G) = floor((n - k + 1)/2) + l.…
▽ More
Let k and l be two non-negative integers with k > l. A graph G is (k,l)-stable if alpha(G - S) >= alpha(G) - l for every subset S of V(G) with |S| = k, where alpha(G) denotes the independence number of G. Dong and Wu established that alpha(G) <= floor((n - k + 1)/2) + l for a (k, l)-stable graph G, where n is the order of G. A (k, l)-stable graph G is tight if alpha(G) = floor((n - k + 1)/2) + l. In this paper, we provide a complete characterization of tight (k, 0)-stable graphs for k >= 4. In particular, we prove that tight (k, 0)-stable graphs are K_{k+1} and K_{k+2} for k >= 5, which not only extends the result of Liu, Song and Wang [J. Graph Theory 110(2) (2025), 193-199] from k >= 24 to k >= 5, but also proves the conjecture of Dong and Luo [Electron. J. Comb. 32(4) (2025), 4-45] once more.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Sublime Transfer Printing of Three-Dimensional Nanostructure Ensembles
Authors:
Lei Chen,
Hao Wang,
Wang Zhang,
Fu Fan,
Peng Liu,
Xiaoxue Bi,
John You En Chan,
Cheng-Feng Pan,
Bochang Wu,
Zhengchao Liu,
Rou Yun Teo,
Hongtao Wang,
Huigao Duan,
Joel K. W. Yang
Abstract:
High-resolution three-dimensional (3D) nanostructures for visible-light photon manipulation provide unique and bespoke capabilities in optics and photonics. However subwavelength nanofabrication and reliable ensemble manipulation of the 3D prints onto arbitrary substrates remain challenging. Here, we introduce sublime transfer strategy tailored for transfer printing ensembles of delicate 3D printe…
▽ More
High-resolution three-dimensional (3D) nanostructures for visible-light photon manipulation provide unique and bespoke capabilities in optics and photonics. However subwavelength nanofabrication and reliable ensemble manipulation of the 3D prints onto arbitrary substrates remain challenging. Here, we introduce sublime transfer strategy tailored for transfer printing ensembles of delicate 3D printed nanostructures. This strategy enables conformal, damage-free integration of arrays of 3D structures on diverse substrates. Naphthalene acts as a transient stamp to encapsulate the structures during transfer and placement. We rely on the low sublimation temperature of naphthalene to release the structures reliably with nearly zero stress, preventing mechanical damage and positional misalignment. This approach is broadly applicable to integrate diverse nanostructures and photonic devices onto various substrates, and enabling inorganic architectures through ensemble uniform post-processing, including 2.5D photonic crystals on flexible PDMS, diffractive optical elements on curved lenses, spiral phase plates on CMOS chips, multilayer achromatic metalens on optical fiber facet, as well as 3D glass photonic crystals and optical topological resonators on anti-stiction quartz.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
KC-BFPRL: Knowledge-Guided Multi-UAV Collaboration for Grassland Restoration via Bilevel Formerpointer-Based Reinforcement Learning
Authors:
Dongbin Jiao,
Xianyi Wang,
Yuchen Yuan,
Weibo Yang,
Peng Yang,
Peng Zhao,
Zhanhuan Shang,
Shi Yan
Abstract:
Multi-unmanned aerial vehicle (UAV) systems provide scalable service platforms for large-scale environmental tasks, such as grassland ecosystem restoration. However, coordinating fleet operations requires solving the restoration area maximization problem (RAMP). This non-linear combinatorial optimization challenge is complicated by payload-dependent energy dynamics and heterogeneous ecological deg…
▽ More
Multi-unmanned aerial vehicle (UAV) systems provide scalable service platforms for large-scale environmental tasks, such as grassland ecosystem restoration. However, coordinating fleet operations requires solving the restoration area maximization problem (RAMP). This non-linear combinatorial optimization challenge is complicated by payload-dependent energy dynamics and heterogeneous ecological degradation. We propose a novel knowledge-guided collaborative bilevel formerpointer reinforcement learning framework (KC-BFPRL) to address this complexity. Using a hierarchical paradigm, KC-BFPRL decomposes RAMP into global task allocation and local restoration planning, with the latter further divided into upper-level trajectory planning and lower-level restoration area allocation. Our specialized architecture pairs featuring a Transformer-based encoder that fuses static environmental features with dynamic UAV states, and a Pointer Network decoder trained via a robust actor-critic framework. By embedding ecological priority rules and heuristic logic, KC-BFPRL achieves a structured warm-start, solving the RL cold-start problem while ensuring strict constraint satisfaction. Extensive experiments demonstrate that KC-BFPRL consistently outperforms state-of-the-art baselines, achieving superior objective values and efficiency. It maintains a $0.00\%$ optimality gap in the most complex scenarios U8-R160 and operates nearly three times faster than MAPDP, validating its robustness, scalability, and real-time applicability for large-scale automated ecological restoration.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
On Brezis' open problem 2.2
Authors:
Hong-Ge Chen,
Yong Liu,
Juncheng Wei,
Wen Yang
Abstract:
We prove that the global minimizer of the Ginzburg-Landau energy in the disk of radius $R$ with boundary value $ u(x)=\frac{x}{|x|}$ is the degree-one radial solution of the planar Ginzburg--Landau equation. This gives an affirmative answer to Open Problem~2.2 in Brezis' open-problem list. This is achieved by comparing the radial solution $f$ in the disk with the degree-one radial solution $F$ in…
▽ More
We prove that the global minimizer of the Ginzburg-Landau energy in the disk of radius $R$ with boundary value $ u(x)=\frac{x}{|x|}$ is the degree-one radial solution of the planar Ginzburg--Landau equation. This gives an affirmative answer to Open Problem~2.2 in Brezis' open-problem list. This is achieved by comparing the radial solution $f$ in the disk with the degree-one radial solution $F$ in the whole plane. Multiplying a disk competitor by $F/f$ enables us to use the known minimality of the whole-plane vortex without changing the boundary trace. The difference of the two energies can be decomposed into Fourier modes. Every nonzero mode is nonnegative, and the zero mode is then handled by a Picone type identity.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion
Authors:
Yu He,
Weikai Yang
Abstract:
Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable procedural knowledge. However, existing methods often consolidate experience based on semantic similarity or LLM judgments, which may merge superficially related but behaviorally incompatible strategies and thereby degrade performance. To address the issue, we propo…
▽ More
Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable procedural knowledge. However, existing methods often consolidate experience based on semantic similarity or LLM judgments, which may merge superficially related but behaviorally incompatible strategies and thereby degrade performance. To address the issue, we propose SkillCommit, an online skill evolution framework that continuously transforms experience into a hierarchical library of reusable skills. Each new experience is initially preserved as an instance-specific patch, retaining the behavior validated in its local context. As related skills accumulate, SkillCommit abstracts those sharing a common behavioral mechanism into higher-level skills. Specifically, for each incoming skill, embedding-based retrieval first identifies candidate related skills. Cross-instance replay and an LLM-based mechanism check determine whether these skills transfer across cases and share a common underlying mechanism. Candidates that pass both checks are abstracted into a higher-level skill and committed only if it preserves the validated behavior of all constituent skills. Experiments on RuleArena, OpenExempt and KOR-Bench demonstrate that SkillCommit consistently improves agent performance across diverse domains. Moreover, the learned skills transfer across model scales and families, enabling cross-model experience transfer.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
Stochastic Liouville-transport theory of light-atom interaction noise in thermal atomic vapors
Authors:
Shaoxin Yuan,
Bin Wu,
Mingyong Jing,
Chaoyang Hu,
Yan Peng,
Tingting Li,
Xingya Li,
Wenguang Yang,
Junyao Xie,
Zongkai Liu,
Hao Zhang,
Linjie Zhang,
Liantuan Xiao,
Suotang Jia
Abstract:
Atom-light interaction noise can limit thermal-vapor sensing. Existing theories often treat internal-state dynamics, finite-mode atomic motion, and stochastic renewal separately, obscuring their coupled contributions to measured noise. We develop a general stochastic Liouville-transport theory, tested against polarization-resolved resonant Cs D$_2$ spectra. Joint experiment-theory analysis identif…
▽ More
Atom-light interaction noise can limit thermal-vapor sensing. Existing theories often treat internal-state dynamics, finite-mode atomic motion, and stochastic renewal separately, obscuring their coupled contributions to measured noise. We develop a general stochastic Liouville-transport theory, tested against polarization-resolved resonant Cs D$_2$ spectra. Joint experiment-theory analysis identifies atom-light noise below approximately 100 kHz as transit-dominated. Ballistic motion through the finite Gaussian mode modulates both the coupling-weighted effective atom number and trajectory-dependent Rabi coupling, producing predominantly common-mode noise. Boundary renewal introduces atoms with independently sampled ground-state sublevels, generating differential population fluctuations with opposite effects on the circular channels. Under an applied longitudinal magnetic field, experiment and theory show the same qualitative nonmonotonic change in common-mode suppression, supporting Zeeman redistribution of the channel responses. The framework can analyze noise in other thermal-atom sensors, including Rydberg-atom electric-field measurements.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
MEDR: Query-Independent Frame Selection via Multi-Signal Event Modeling and Dynamic Rescoring
Authors:
Xinlei Pu,
Weijie Shi,
Wen Yang,
Yi Cao,
Hao Chen,
Yuanjun Liu,
Wenwei Ding,
Jia Zhu,
Jiajie Xu
Abstract:
Frame selection is a fundamental component of multimodal large language models, enabling long videos to be processed under limited visual-token and computational budgets. Uniform sampling preserves temporal coverage but may miss informative content that appears only briefly. To alleviate this limitation, query-dependent methods can retrieve question-relevant frames. However, because the selected f…
▽ More
Frame selection is a fundamental component of multimodal large language models, enabling long videos to be processed under limited visual-token and computational budgets. Uniform sampling preserves temporal coverage but may miss informative content that appears only briefly. To alleviate this limitation, query-dependent methods can retrieve question-relevant frames. However, because the selected frames depend on the current question, the same visual input cannot be directly shared across different questions, and frame selection must be repeated in multi-turn video dialogue. This motivates us to seek a query-independent frame selection method that preserves the reusability of a fixed visual input while improving the coverage of informative events beyond uniform sampling. We propose Multi-Signal Event Modeling and Dynamic Rescoring (MEDR), a training-free and query-independent frame selection method. Multi-Signal Event Modeling organizes complementary visual, motion, and text signals into signal-specific temporal events. Dynamic Rescoring then iteratively reevaluates each candidate relative to the current selected set, updating its score according to frame-level signal strength, additional event coverage, and temporal proximity. The resulting fixed frame set is constructed without observing the query and can be reused across different questions. On the standard benchmark evaluations, MEDR improves model accuracy by 0.63%-0.89% on Video-MME. On the long-video subset of LongVideoBench, it improves accuracy by up to 1.23% with Qwen3-VL-8B. MEDR further improves overall accuracy by 0.53%, while reusing exactly the same frame set for every question about a video.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents
Authors:
Wei-Chieh Huang,
Weizhi Zhang,
Yuchen Wu,
Yankai Chen,
Eric Hanchen Jiang,
Wooseong Yang,
Yiwei Yang,
Henry Peng Zou,
Hanrong Zhang,
Ying Nian Wu,
Haolun Wu,
Kai-Wei Chang,
Philip S. Yu,
Xue Liu,
Aylin Caliskan
Abstract:
Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which memory is represented and stored, should be used under different operating regimes. We present a controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text re…
▽ More
Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which memory is represented and stored, should be used under different operating regimes. We present a controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms. Across three backbone models and four benchmark suites spanning user-centric question answering and agent-centric decision-making, we instrument 26 performance and efficiency metrics under a unified harness. Our results show that no single substrate consistently dominates: broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context. Scalability introduces a further routing axis, as substrates that perform well at moderate history lengths can become costly or brittle at longer horizons. These findings motivate substrate routing as a necessary component of adaptive agent memory systems and provide empirical guidance for designing efficient, reliable, and regime-aware long-term memory for LLM agents. Code will be made available upon acceptance.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Authors:
Yiwei Li,
Wanli Yang,
Hexiang Tan,
Xiangzhou Huang,
Zhengyu Chen,
Ziran Li,
Borun Chen,
Shanglin Lei,
Huaisheng Zhu,
Hao Tian,
Fei Sun,
Xunliang Cai,
Jingang Wang
Abstract:
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic…
▽ More
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
Authors:
Peng Ling,
Yingda Yin,
Lingting Zhu,
Weikai Chen,
Shengju Qian,
Zeyu Hu,
Xin Wang,
Wenming Yang
Abstract:
While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational bottlenecks during inference. Existing token pruning methods primarily rely on diversity-based selection, discarding similar tokens to maximize dispersion. However, in 3D environments, this approach frequently drops rep…
▽ More
While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational bottlenecks during inference. Existing token pruning methods primarily rely on diversity-based selection, discarding similar tokens to maximize dispersion. However, in 3D environments, this approach frequently drops representative prototype tokens in favor of outliers, breaking the multi-view consistencies and geometric structures essential for spatial reasoning. In this paper, we propose a paradigm shift for 3D VLM token pruning: from maximizing diversity to preserving visual evidence coverage. We introduce CoverPrune, a training-free framework that formulates inference-time token pruning as an Optimal Transport (OT) problem. To overcome the intractable combinatorial subset selection inherent in this formulation, we design the Feature-Spatial-Temporal (FST) transport cost and target capacity, along with an efficient Spatial-Guided Greedy Selection (SGS) algorithm to approximate the OT objective. Furthermore, we propose CoverPrune-Lite, an accelerated variant utilizing spatially structured local matching for minimal overhead. Extensive experiments across multiple 3D visual-spatial reasoning benchmarks demonstrate that our methods achieve state-of-the-art token efficiency, maintaining robust reasoning performance even under highly aggressive pruning budgets. Visit our project website at https://github.com/Brucess/CoverPrune.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Sharp Lower Bounds on the Haraux Function Beyond Reflexivity
Authors:
Weifeng Yang
Abstract:
We prove that the sharp $\frac{1}{2}$ lower bound for the Haraux function holds for every maximally monotone operator of type~(NI) on an arbitrary real Banach space. This resolves the nonreflexive extension raised by the recent reflexive result. We establish an exact decomposition at each graph point, where the local contribution to the Haraux function and a nonnegative residual together equal…
▽ More
We prove that the sharp $\frac{1}{2}$ lower bound for the Haraux function holds for every maximally monotone operator of type~(NI) on an arbitrary real Banach space. This resolves the nonreflexive extension raised by the recent reflexive result. We establish an exact decomposition at each graph point, where the local contribution to the Haraux function and a nonnegative residual together equal $\frac{1}{2}$ times the weighted squared displacement. Since the equivalence between type~(NI) and quasidensity provides graph points whose residuals tend to zero, this decomposition also yields the sharp bound without requiring a graph point at which the residual vanishes. Moreover, for every operator with a nonempty graph, this decomposition yields a lower bound involving the residual infimum. For maximally monotone operators, this decomposition also yields a new characterization of type~(NI) in terms of the Haraux function. Finally, on $c_0$, we give a maximally monotone operator of type~(NI) for which the residual infimum is zero at some target but is not attained. This shows that the existence of a graph point at which the residual vanishes is strictly stronger than the vanishing of the residual infimum required in our proof.
△ Less
Submitted 17 August, 2026; v1 submitted 13 August, 2026;
originally announced August 2026.
-
Dual-Manifold Geometry Guided Representation Learning: Adaptive Coupling between Kernel and Data Spaces
Authors:
Wencong Zhang,
Yue Zhang,
Meiyan Huang,
Wei Yang,
Qianjin Feng
Abstract:
Deep representation learning has primarily focused on how features evolve across network layers, while largely overlooking the structured geometry embedded in network parameters. We introduce a dual-manifold perspective in which each convolutional layer contains two coupled geometric spaces: a Kernel Manifold induced by convolutional filters and a Data Manifold characterized by intermediate featur…
▽ More
Deep representation learning has primarily focused on how features evolve across network layers, while largely overlooking the structured geometry embedded in network parameters. We introduce a dual-manifold perspective in which each convolutional layer contains two coupled geometric spaces: a Kernel Manifold induced by convolutional filters and a Data Manifold characterized by intermediate feature representations. Because these manifolds share the same channel space, parameter geometry can provide complementary structural information to guide feature evolution. Based on this insight, we propose Kernel-Guided Feature Transform (KGFT), a lightweight module that derives a geometric guidance matrix from the kernel Gram matrix and uses it to transform the covariance structure of feature representations. Unlike conventional attention mechanisms that reweight feature responses, KGFT explicitly reshapes feature relationships by transferring geometric information from the kernel manifold to the data manifold. To accommodate network hierarchy, we further introduce Exploit and Explore modes with a depth-aware scheduling strategy and a learnable guidance strength that adaptively controls the contribution of geometric transformation. This design promotes geometric alignment in shallow layers while encouraging feature diversity in deeper layers, without imposing excessive constraints on representation learning. Theoretical analysis establishes the validity of the proposed transformation and characterizes its effect on feature covariance. Extensive experiments across CNN- and Transformer-based architectures, including ResNet, ViT, and LLaMA-7B, demonstrate consistent improvements on image classification and arithmetic reasoning tasks, validating the generality and effectiveness of kernel-guided dual-manifold representation learning. Code will be publicly available.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
Authors:
Wooseong Yang,
Wei-Chieh Huang,
Weizhi Zhang,
Yu Wang,
Philip S. Yu,
Junhyun Lee
Abstract:
Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's internal representations, or require architectural homogeneity for latent-level transfer. We identify the entity grounding problem in…
▽ More
Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's internal representations, or require architectural homogeneity for latent-level transfer. We identify the entity grounding problem in cross-architecture communication: cross-attention bridges that transfer continuous representations across different LLM families suffer from rare-token compression collapse, where entity identity is lost in the continuous bottleneck (bridge-only F1 ~30%). We propose XBRIDGE, a decode-free communication protocol that addresses this through two mechanisms. Lexical Anchor Mapping (LAM) maps the sender's original context tokens to the receiver's vocabulary, providing discrete entity anchors. A Latent Enrichment Bridge (LEB) lets the receiver query the sender's hidden states for contextual enrichment. The entity anchors ground the bridge's contextual signals to specific entities through the receiver's own self-attention. Across three model families (Llama, Qwen, and Mistral), seven benchmarks, and both communication directions, XBRIDGE outperforms text-based communication on all seven tasks for each model pair while achieving 11x lower latency, and in a same-architecture setting it also exceeds a KV-sharing baseline on six of seven tasks. LEB requires only 264M trainable parameters (3.8% of the receiver), is trained on a small balanced sample set, and adds negligible inference overhead.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Nonradial stable solutions near the Joseph--Lundgren threshold
Authors:
Shibing Chen,
Yong Liu,
Juncheng Wei,
Wen Yang
Abstract:
We study positive stable solutions of the supercritical Lane--Emden equation in the first Joseph--Lundgren interval. For a family of dimensions, we construct nonradial stable entire solutions with exponent close to the upper endpoint of this interval. This disproves a radiality conjecture of Chan and Wei. The construction begins with a smooth positive nonconstant solution on the sphere, which is o…
▽ More
We study positive stable solutions of the supercritical Lane--Emden equation in the first Joseph--Lundgren interval. For a family of dimensions, we construct nonradial stable entire solutions with exponent close to the upper endpoint of this interval. This disproves a radiality conjecture of Chan and Wei. The construction begins with a smooth positive nonconstant solution on the sphere, which is obtained by matching a polar cap to an inner neck. A sharp expansion of the lowest shifted eigenvalue proves that the resulting singular cone is strictly stable. Finally, a minimal-solution and rescaling argument replaces the cone by a smooth stable entire solution while preserving its sphere variation. To the best of our knowledge, this is the first nontrivial example of nonradial stable solutions for the Emden-Fowler equation.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
DREAM Technical Report
Authors:
Bin Zhang,
Bowen Zheng,
Chao Yi,
Chengyu Lai,
Dian Chen,
Dimin Wang,
Gaoyang Guo,
Jialin Zhu,
Jian Wu,
Jing Yu,
Jiuning Lin,
Lingqing Zhang,
Lingyun Zheng,
Mao Zhang,
Mingming Pan,
Ruiquan Lan,
Shuai Zhong,
Wen Chen,
Wendong Zhang,
Xiaodong Zhu,
Xuan Chen,
Xunke Xi,
Yifan Lu,
Yiheng Wang,
Yue Zeng
, et al. (52 additional authors not shown)
Abstract:
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine…
▽ More
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.
△ Less
Submitted 13 August, 2026; v1 submitted 10 August, 2026;
originally announced August 2026.
-
One-Time Training for All Grains: Open-Set Grain Recognition and Quantitative Analysis
Authors:
Qihe Su,
Mengyu Sun,
Yuxi Ke,
Zhuoyan Jiang,
Wanneng Yang,
Chenglong Huang,
Ziyuan Yang
Abstract:
Advances in crop breeding have introduced an increasing number of grain varieties, creating a growing demand for efficient variety recognition and quantitative analysis. However, existing methods are typically trained on a fixed variety set, and incorporating newly introduced varieties requires additional data collection and model retraining. To address this limitation, we propose GROW, a framewor…
▽ More
Advances in crop breeding have introduced an increasing number of grain varieties, creating a growing demand for efficient variety recognition and quantitative analysis. However, existing methods are typically trained on a fixed variety set, and incorporating newly introduced varieties requires additional data collection and model retraining. To address this limitation, we propose GROW, a framework for Grain Recognition and quantitative analysis in Open sets Without retraining. GROW first performs class-agnostic grain localization, converting mixed-grain images into individual instances for variety-wise counting and phenotypic measurement. It then combines visual embeddings and morphological descriptors into fused grain descriptors stored in an extensible GrainBank. Query grains are recognized through rank-similarity weighted top-k retrieval, and newly introduced varieties are incorporated by appending their descriptors without updating the deployed models. Extensive experiments under progressive variety expansion, varying grain densities, and background domain shifts demonstrate the scalability, robustness, and adaptability of GROW. Compared with joint retraining, GROW reduced the average category-registration time from 4153 s to only 39 s while maintaining competitive recognition performance. These results demonstrate that GROW provides an efficient and maintainable solution for extensible grain recognition, counting, and phenotypic analysis without repeated model retraining.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation
Authors:
Yilei Hua,
Beibei Jing,
Ce Zheng,
Hanyu Zhou,
Yawei Luo,
Wei Yang
Abstract:
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in…
▽ More
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity. To overcome this bottleneck, we ground motion editing directly within text-to-motion generation across data, architecture, and inference. At the data level, we develop a closed-loop synthesis-and-verification pipeline that produces Omni-MoEdit, a large-scale dataset spanning body-part, amplitude, temporal, action, and style edits. At the architectural level, we introduce UniMoFlow, a unified latent flow-matching model that shares broad semantic and kinematic knowledge between generation and editing. At the inference level, SAFE (Source-Anchored Flow Editing) complements UniMoFlow with controllable, source-anchored refinement. Furthermore, we augment standard evaluations with semantics-aware metrics to account for valid edits that inherently deviate from a single ground-truth reference. Extensive experiments demonstrate improved target-text alignment, edit effectiveness, and cycle consistency, while maintaining competitive source fidelity and text-to-motion generation quality.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Wafer-scale monolithic integration of Ce:YIG films and magneto-optical isolators on silicon
Authors:
Tianchi Zhang,
Yucong Yang,
Weihao Yang,
JieJun Su,
Tianyi Ma,
Xuan Zhao,
Junxian Wang,
Di Wu,
Zhenyuan Ren,
Yi Shuai,
Zixuan Wei,
Lei Bi
Abstract:
Silicon integrated cerium doped yttrium iron garnet (Ce:YIG) thin films are promising candidates for integrated nonreciprocal photonic devices, cryogenic photonic modulators and optical computing applications. However, previously reported Ce:YIG thin film on silicon is limited to milimeter sizes. Wafer-scale integration and non-destructive characterization of high quality Ce:YIG thin films on sili…
▽ More
Silicon integrated cerium doped yttrium iron garnet (Ce:YIG) thin films are promising candidates for integrated nonreciprocal photonic devices, cryogenic photonic modulators and optical computing applications. However, previously reported Ce:YIG thin film on silicon is limited to milimeter sizes. Wafer-scale integration and non-destructive characterization of high quality Ce:YIG thin films on silicon has been elusive. Here, we report growth of 4-inch wafer-scale Ce:YIG thin films on silicon substrates by radio-frequency magnetron sputtering. Strong Faraday effect of 2318 deg/cm, low propagation loss of 80 dB/cm and excellent thickness uniformity of 3.5% is demonstrated across the 4-inch silicon wafer. Furthermore, a custom designed wafer-scale, non-destructive magneto-ellipsometry was established to characterize the film thickness, optical constants and magneto-optical constants across the wafer. Wafer-scale integration of ring resonator type magneto-optical isolators are also demonstrated. Our work demonstrates a step forward toward wafer-scale heterogeneous integration and characterization of magneto-optical thin films on silicon, providing material candidates for non-reciprocal photonic device arrays, magneto-optical in-memory computing networks and integrated magneto-optic magnetometers.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding
Authors:
Donghui Feng,
Fengxi Zhang,
Changsheng Gao,
Wenhan Yang,
Qi Wang,
Qunshan Gu,
Hongwei Hu,
Zhengxue Cheng,
Li Song
Abstract:
Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature compression critical under bandwidth and computation constraints. Existing ViT feature codecs typically flatten heterogeneous global and patch tokens into an L x C pseudo image, causing entropy models to mainly capture…
▽ More
Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature compression critical under bandwidth and computation constraints. Existing ViT feature codecs typically flatten heterogeneous global and patch tokens into an L x C pseudo image, causing entropy models to mainly capture sequence-axis dependencies while overlooking the native two-dimensional patch-grid structure. In this paper, we show that ViT patch tokens retain strong local spatial correlations on the original grid. To exploit this structural prior, we propose the Visual Token Codec (VTC), a dual-path learned codec that separates global and patch tokens into dedicated coding paths. Global tokens are compressed with a lightweight factorized prior, whereas patch tokens are encoded on the patch-token grid using a spatial-channel context entropy model. To support intermediate-layer compression and practical rate adaptation, VTC further incorporates feature-matching supervision after subsequent ViT blocks and variable-rate modules within a single codec. Experiments on DINOv2 and SAM3 show that VTC consistently outperforms representative ViT feature coding baselines on classification, segmentation, and detection tasks. At 90% of uncompressed-feature performance, VTC reduces bitrate by 15.7x-37.4x across these tasks. We further provide intermediate-layer rate-utility analyses for practical transmission- and storage-oriented deployment scenarios.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
Sharp $L^2$ Estimates for $(2+1)$-dimensional oscillatory integral operators with homogeneous binomial phases
Authors:
Chu-hee Cho,
Jin Bong Lee,
Chan Woo Yang
Abstract:
We study oscillatory integral operators in $(2+1)$-dimensions with a homogeneous binomial phase \[
Φ(x,y,t)=x^{k-k_P}t^{k_P}+y^{k-k_Q}t^{k_Q}, \qquad 1\le k_P<k_Q<k. \] For compactly supported smooth amplitudes, we establish sharp \(L^2(\R)\to L^2(\R^2)\) estimates with logarithmic losses occurring only in certain critical cases. The proof is based on scale-dependent Phong--Stein estimates.
We study oscillatory integral operators in $(2+1)$-dimensions with a homogeneous binomial phase \[
Φ(x,y,t)=x^{k-k_P}t^{k_P}+y^{k-k_Q}t^{k_Q}, \qquad 1\le k_P<k_Q<k. \] For compactly supported smooth amplitudes, we establish sharp \(L^2(\R)\to L^2(\R^2)\) estimates with logarithmic losses occurring only in certain critical cases. The proof is based on scale-dependent Phong--Stein estimates.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
Authors:
Yuqi Zhang,
Cheng Chen,
Yuyu Guo,
Wenjie Yang,
Lingchen Meng,
Peng Di,
Hang Yu,
Zuxuan Wu,
Yu-Gang Jiang
Abstract:
Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions either resort to aggressive token pruning, risking irreversible information loss, or adopt efficient but less precise architectures, while largely ignoring the equally vital textual component. We introduce VLZip, a framewor…
▽ More
Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions either resort to aggressive token pruning, risking irreversible information loss, or adopt efficient but less precise architectures, while largely ignoring the equally vital textual component. We introduce VLZip, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer. At its core, VLZip hierarchically distills visual and textual segments into compact, layer-specific "soft prefixes" and injects them into each decoder layer's hidden states, drastically shortening the attention sequence while preserving fine-grained global context. To address deficient evaluations in the field, we also introduce LongVLBench, a new benchmark derived from video narratives that demands holistic, narrative-level reasoning. Extensive experiments show VLZip achieves leading performance on long-context multimodal reasoning, enabling training up to 120K tokens, a 6x increase over the baseline, and inference beyond 280K tokens with significantly reduced memory, while demonstrating the memory scalability to handle up to 2M tokens. By excelling at extreme context lengths where existing methods collapse, VLZip establishes an efficient and powerful new standard for long-context multimodal AI. Code is available at https://github.com/ShareLab-SII/VLZip.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
Authors:
Yi Shu,
Tianyu Peng,
Yingzhuo Deng,
Wen Yang,
Jun Lin,
Changming Xie,
Xinyu Yu,
Jiajun Zhang
Abstract:
Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic representation space of speech dialogue models continuously evolves, while conventional speech supervision remains unchanged, leading to semantic inconsistency be…
▽ More
Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic representation space of speech dialogue models continuously evolves, while conventional speech supervision remains unchanged, leading to semantic inconsistency between hidden representations and speech targets and degrading speech stability and naturalness. To address these issues, we propose DialectS2S, an end-to-end speech dialogue model for Chinese dialects. We first develop a scalable dialect speech dialogue synthesis pipeline for efficient data construction. We further introduce a two-stage post-training strategy with self-aligned speech supervision, which aligns the semantic content of speech supervision with the evolved semantic representations of the model to improve dialect speech generation quality. Experimental results show that DialectS2S consistently outperforms existing baselines across multiple Chinese dialects in speech dialogue, achieving substantial improvements in dialect consistency, response quality, and speech intelligibility. Our work provides an efficient and scalable solution for end-to-end speech dialogue modeling in low-resource dialect scenarios. To facilitate future research and practical applications, we fully open-source the DialectS2S framework, including model checkpoints, training datasets, and fine-tuning code.
△ Less
Submitted 14 August, 2026; v1 submitted 8 August, 2026;
originally announced August 2026.
-
Anisotropic Particle Transport from a Pulsar Wind Nebula Revealed by Einstein Probe and LHAASO
Authors:
Zhen Cao,
F. Aharonian,
Y. X. Bai,
Y. W. Bao,
D. Bastieri,
X. J. Bi,
Y. J. Bi,
W. Bian,
J. Blunier,
A. V. Bukevich,
C. M. Cai,
W. Y. Cao,
Zhe Cao,
J. Chang,
J. F. Chang,
E. S. Chen,
G. H. Chen,
H. K. Chen,
L. F. Chen,
Liang Chen,
Long Chen,
M. J. Chen,
M. L. Chen,
Q. H. Chen,
S. Chen
, et al. (320 additional authors not shown)
Abstract:
Pulsar wind nebulae (PWNe) are major cosmic ray accelerators, yet the mechanisms transporting high-energy particles into the interstellar medium remain elusive. Building on the LHAASO discovery of an ultra-high-energy (UHE) $γ$-ray source near the bow-shock PWN powered by the pulsar PSR J1740+1000, we present a joint Einstein Probe (EP) and LHAASO study of this system. EP observations reveal an ex…
▽ More
Pulsar wind nebulae (PWNe) are major cosmic ray accelerators, yet the mechanisms transporting high-energy particles into the interstellar medium remain elusive. Building on the LHAASO discovery of an ultra-high-energy (UHE) $γ$-ray source near the bow-shock PWN powered by the pulsar PSR J1740+1000, we present a joint Einstein Probe (EP) and LHAASO study of this system. EP observations reveal an extended X-ray tail far exceeding the structure previously seen by XMM-Newton. Updated LHAASO observations show that the $γ$-ray emission is elongated, with its major axis aligned with the extended X-ray tail revealed by EP. This is the first detection of an X-ray pulsar tail associated with a spatially coincident extended UHE $γ$-ray emission. The X-ray and $γ$-ray spectrum can be well explained with a single population of relativistic electrons via synchrotron and inverse Compton radiation, respectively, removing the need for particle re-acceleration during propagation. The results unambiguously show that electrons/positrons above 100 TeV are escaping from the PWN. Instead of the immediate, isotropic diffusion into ambient interstellar medium that is typically assumed, these particles are transported anisotropically over at least $\sim$10 pc, either guided by the background magnetic field or carried by an advective outflow.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Designer Codes from GALA: Compact, Self-Dual, and Rate-1/2 QEC on Reconfigurable Atom Arrays
Authors:
Willers Yang,
Casey Duckering,
Arpit Dua
Abstract:
High rate quantum low-density parity-check codes on reconfigurable neutral-atom arrays can reduce the overhead of quantum error correction, but near-term devices support only hundreds of qubits with limited reconfigurability from a few crossed acousto-optic deflectors (AOD). A practical code must be compact in addition to low-overhead, with checks and logical gates mapping onto hardware-compatible…
▽ More
High rate quantum low-density parity-check codes on reconfigurable neutral-atom arrays can reduce the overhead of quantum error correction, but near-term devices support only hundreds of qubits with limited reconfigurability from a few crossed acousto-optic deflectors (AOD). A practical code must be compact in addition to low-overhead, with checks and logical gates mapping onto hardware-compatible physical instructions. We introduce the GALA codes, or Group-Action Lifts with Active orthogonality, that lifts over a product group $G = H_k \times C_m$ (or $H_k \ltimes C_m^k$). The small non-abelian factor $H_k$ supplies active orthogonality, reaching $1/2$ rate with above-weight distance, while the large abelian factor $C_m$ supplies symmetries that give code automorphisms and explicit, simple AOD move schedules. Hardware compatibility and logical capability thereby become customizable inputs to a code search rather than properties verified post-hoc, making GALA designer codes by construction. The GALA family contains several previously discovered rate-1/2 Kasai codes of Ref. [arXiv:2601.08824, arXiv:2604.16209] while exposing simpler parameter bounds, logical operations, and ZX-dual variants with AOD-compatible fold-transversal Clifford gates. Our search yields a compact self-dual $[[132, 30, 12]]$ with a small number of $4$-cycles (almost girth-6) and below $10^{-8}$ logical error rate (LER) for memory at $10^{-3}$ physical error rate, 3.1ms syndrome-extraction cycle and transversal Clifford gates; a girth-6, rate-$1/2$ $[[672, 336, 12]]$ with a 6.76ms cycle with below $10^{-10}$ LER (extrapolated) and rate-$1/2$ barrier-breaking $[[1752, 880, 14]]$ and $[[2232, 1120, 16]]$ with exactly certified distances greater than check weights and all smaller than previously known hardware compatible rate-1/2 codes.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Universal Concept Disruption for SAM3 Image Segmentation
Authors:
Hao Wang,
Yuxuan Zhang,
Wei Yang
Abstract:
SAM3 extends promptable segmentation from geometry-driven mask prediction to open-vocabulary concept segmentation, where a text-conditioned grounding model decides whether a concept is present and segments all matching instances. While this presence-gated design improves concept-level prediction, its adversarial robustness remains unexplored. In this paper, we introduce Universal Concept Disruptio…
▽ More
SAM3 extends promptable segmentation from geometry-driven mask prediction to open-vocabulary concept segmentation, where a text-conditioned grounding model decides whether a concept is present and segments all matching instances. While this presence-gated design improves concept-level prediction, its adversarial robustness remains unexplored. In this paper, we introduce Universal Concept Disruption (UCD), the first universal cross-concept adversarial attack tailored to SAM3 image segmentation. UCD learns a single bounded image perturbation from (image, noun-phrase) pairs and attacks SAM3 as an integrated concept-grounding system. It jointly disrupts the text-conditioned input path, maximizes divergence in prompt-shared visual features, suppresses the final presence-gated concept scores, and corrupts the spatial validity of retained masks through area collapse and clean-mask Dice disruption. Across SACo-Gold, LVIS, RefCOCO, PhraseCut, and OpenImages datasets, UCD consistently outperforms all baselines under a matched evaluation protocol, reducing average mask AP from 59.43 to 18.73 and average cgF1 from 50.32 to 20.49. The learned perturbation also transfers to SAM3.1 and to SAM3 video inference without re-optimization, while prompt ensembling, lightweight head fine-tuning, and temporal filtering provide limited recovery.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Photometric Distances for Metal-poor Giants and a Search for Hypervelocity Stars with LAMOST DR13 and Gaia DR3
Authors:
Shuai Xu,
Haibo Yuan,
Bowen Huang,
Xiao Kai,
Wuming Yang
Abstract:
Hypervelocity stars (HVSs) are stars with velocities high enough to escape the Milky Way, but their identification depends sensitively on distance estimates, particularly for distant giants. In this work, we search for metal-poor HVS candidates by combining LAMOST DR13 spectroscopy with Gaia DR3 astrometry. We calibrate a metallicity-dependent color--absolute-magnitude relation for normal metal-po…
▽ More
Hypervelocity stars (HVSs) are stars with velocities high enough to escape the Milky Way, but their identification depends sensitively on distance estimates, particularly for distant giants. In this work, we search for metal-poor HVS candidates by combining LAMOST DR13 spectroscopy with Gaia DR3 astrometry. We calibrate a metallicity-dependent color--absolute-magnitude relation for normal metal-poor giants using a high-quality reference sample, and apply it to derive photometric distances for 41{,}331 stars. The relation reproduces the reference absolute magnitudes with a scatter of 0.24\,mag, corresponding to an intrinsic distance uncertainty of $\sim$9.7\%. Combining these distances with Gaia proper motions and LAMOST radial velocities, we identify 13 initially unbound candidates under the Galactic potential from \citet{McMillan2017}. Spectral inspection indicates that several are chromospherically active binaries or other non-standard systems for which a giant-star calibration is unreliable; removing these contaminants leaves nine metal-poor HVS candidates. Backward orbit integrations suggest that one candidate is most consistent with a disk origin, while three have trajectories suggestive of an association with the Sagittarius stream.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Dual-polarization control of broadband nonreciprocal thermal radiation by combining local and nonlocal metasurfaces
Authors:
Shuang Xia,
Mengqi Liu,
Wenjian Wan,
Jialong Wang,
Weihao Yang,
Chaoran Wang,
Huiqin Ma,
Jun Qin,
Hua Li,
Yuan Wang,
Lei Bi,
Chengwei Qiu,
Xiaobo Yin
Abstract:
Nonreciprocal thermal radiation offers a route to decouple spectral directional absorptivity and emissivity, thereby enabling new paradigms in thermal-photonic systems. However, in magneto-optical platforms, the intrinsic gyroelectric response generally confines observable nonreciprocity to transverse-magnetic (TM) polarization, while the transverse-electric (TE) response is absent. In this work,…
▽ More
Nonreciprocal thermal radiation offers a route to decouple spectral directional absorptivity and emissivity, thereby enabling new paradigms in thermal-photonic systems. However, in magneto-optical platforms, the intrinsic gyroelectric response generally confines observable nonreciprocity to transverse-magnetic (TM) polarization, while the transverse-electric (TE) response is absent. In this work, we experimentally demonstrate, for the first time, a local thermal metasurface strategy to activate TE-polarized nonreciprocity by creating artificial gyromagnetic response in a gyroelectric semiconductor platform. We further extend this mechanism to broadband dual-polarization operation employing a nonlocal thermal metasurface, which combines a resonator supercell with gradient-doped epsilon-near-zero magneto-optical multilayers. Pronounced absorptivity contrast is maintained over 22-27 μm for TE polarization and 19-27 μm for TM polarization. This platform provides a mechanism-based route to achieve broadband and dual-polarization nonreciprocal thermal absorption, opening new opportunities for advancing radiative energy-conversion devices.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
Authors:
Feier Wu,
Wanke Xia,
Xu He,
Zilang Zhou,
Si Chen,
Dongxia Liu,
Liyang Chen,
Qimeng Wu,
Zhengbo Zhang,
Wenming Yang,
Zhiyong Wu
Abstract:
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatia…
▽ More
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
An Inertial Block Proximal Linearized Method with Adaptive Momentum for Nonconvex and Nonsmooth Optimization
Authors:
Weifeng Yang
Abstract:
In this paper, we consider a class of multiblock nonconvex nonsmooth optimization problems, which covers many applications such as the analysis of pre-earthquake anomalies and machine learning. To solve this class of problems, we propose the inertial block proximal linearized method with two-phase adaptive momentum (IBPL$^+$-TP). Compared to the current methods, our method possesses three main adv…
▽ More
In this paper, we consider a class of multiblock nonconvex nonsmooth optimization problems, which covers many applications such as the analysis of pre-earthquake anomalies and machine learning. To solve this class of problems, we propose the inertial block proximal linearized method with two-phase adaptive momentum (IBPL$^+$-TP). Compared to the current methods, our method possesses three main advantages: (1) it introduces a two-phase adaptive momentum strategy to effectively update the extrapolation parameters, (2) it allows using two different extrapolation points to accelerate the convergence, (3) it allows the extrapolation parameters of these two extrapolation points to be independent of and unconstrained by all other parameters. While maintaining the above advantages, we prove that our method ensures the monotonic convergence of the objective function of this class of problems, and we also prove that the sequence generated by our method globally converges to a critical point, as well as establish the convergence rate of our method. To demonstrate the effectiveness of our method, we apply it to solve two nonconvex and nonsmooth machine learning problems, namely sparse nonnegative matrix factorization with $\ell_0$-constraints and sparse nonnegative CP decomposition with $\ell_0$-constraints. The numerical experimental results on solving these problems show that our method outperforms several state-of-the-art methods.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
Nondegeneracy and Morse Index of Ginzburg--Landau Vortices
Authors:
Manuel del Pino,
Yong Liu,
Monica Musso,
Juncheng Wei,
Wen Yang
Abstract:
We prove that the standard degree-two and degree-three vortex solutions of the Ginzburg-Landau equation are nondegenerate. Their Morse indices are also computed. The proof relies on new explicit upper and lower bounds of the modulus of these solutions and a comparison argument. It is expected that our method can be generalized to study higher degree solutions.
We prove that the standard degree-two and degree-three vortex solutions of the Ginzburg-Landau equation are nondegenerate. Their Morse indices are also computed. The proof relies on new explicit upper and lower bounds of the modulus of these solutions and a comparison argument. It is expected that our method can be generalized to study higher degree solutions.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
CommBench: Can LLMs Write Correct and Efficient GPU Communication Code?
Authors:
Shuang Ma,
Yuyi Li,
Yihan Zhang,
Hezhi Xie,
Danyang Chen,
Shuyang Ji,
Ziming Mao,
Cheng Ji,
Ansha Prashanth,
Wenting Yang,
Yiran Wang,
Chihan Cui,
Pei Yu Lin,
Ion Stoica,
Yang Zhou
Abstract:
Training and serving large language models (LLMs) rely heavily on high-performance GPU communication, yet implementing efficient GPU communication primitives requires deep expertise in GPU architectures, networking hardware, and distributed communication patterns, making them particularly challenging for code generation models. We present CommBench, a comprehensive benchmark for GPU communication…
▽ More
Training and serving large language models (LLMs) rely heavily on high-performance GPU communication, yet implementing efficient GPU communication primitives requires deep expertise in GPU architectures, networking hardware, and distributed communication patterns, making them particularly challenging for code generation models. We present CommBench, a comprehensive benchmark for GPU communication programming, consisting of over 100 expert-curated tasks spanning point-to-point communication, collective operations, expert-parallel communication, compute--communication fusion, and communication utility functions, with reference implementations either written by GPU communication experts or distilled from production codebases. We further introduce a cheat-resistant evaluation framework that automatically compiles, executes, and validates generated code on multi-GPU systems, and a unified metric that jointly measures functional correctness and communication performance. Evaluating leading frontier and open-source code generation models on both intra-node NVLink and inter-node RDMA platforms reveals that even the strongest model, GPT-5.5, correctly implements and achieves competitive performance on only 30.7\% of the benchmark tasks. Our results expose a substantial gap between current LLMs and expert-written GPU communication code, establishing CommBench as a challenging benchmark for advancing AI-assisted systems programming.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
Pressure induced magnetic-field-free superconducting diode effect in NbSe2 flake
Authors:
Shihao Zhu,
Tian Le,
Cuiying Pei,
Changhua Li,
Yi Liao,
Yi Zhao,
Lingxiao Zhao,
Qi Wang,
Juefei Wu,
Qilian Zhang,
Yueshen Wu,
Tonghuan Fu,
Xujie Lü,
Wenge Yang,
Jie Shen,
Jun Li,
Yulin Chen,
Xiao Lin,
Wen-Yu He,
Yanpeng Qi
Abstract:
The superconducting diode effect (SDE) is a fascinating nonreciprocal phenomenon where the critical current is different for opposite current directions. It is widely believed that realizing SDE requires breaking both inversion symmetry (IS) and time-reversal symmetry (TRS), which are usually achieved via heterostructure engineering and applying external magnetic fields. Here, we report a pressure…
▽ More
The superconducting diode effect (SDE) is a fascinating nonreciprocal phenomenon where the critical current is different for opposite current directions. It is widely believed that realizing SDE requires breaking both inversion symmetry (IS) and time-reversal symmetry (TRS), which are usually achieved via heterostructure engineering and applying external magnetic fields. Here, we report a pressure-induced magnetic-field-free SDE in NbSe2 flakes without any heterostructures. We show that pressure alone breaks the IS, as confirmed by the second harmonic generation. Crucially, upon applying an out-of-plane magnetic field (B), the SDE exhibits even-in-B behavior, implying the absence of explicit TRS breaking. This finding challenges the prevailing theoretical paradigm and demonstrates that a magnetic-field-free SDE can emerge without explicitly breaking TRS. Thereby, our work establishes pressure engineering as a powerful tool for inducing nonreciprocal superconductivity and designing versatile, magnetic-field-free superconducting devices.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse
Authors:
Shuxiao Xie,
Shuyang Xie,
Yuan Cao,
Dezhi Ran,
Wei Yang,
Tao Xie
Abstract:
A bfloat16 transformer can train normally for many steps and then collapse abruptly. Distinct low-precision errors can trigger the same failure, leaving unclear whether each source needs its own repair or one shared route can be blocked. We isolate a reproduced GPT-2-class collapse to the streaming-softmax accumulator, where fp32 accumulation repairs it, and use the fault as an assay for moving co…
▽ More
A bfloat16 transformer can train normally for many steps and then collapse abruptly. Distinct low-precision errors can trigger the same failure, leaving unclear whether each source needs its own repair or one shared route can be blocked. We isolate a reproduced GPT-2-class collapse to the streaming-softmax accumulator, where fp32 accumulation repairs it, and use the fault as an assay for moving controlled errors across sources. Errors placed outside attention still drive the same query-key (QK) spectral runaway, while correcting only QK keeps training stable with the source fault active. This source-channel dissociation shows that fault source is not failure channel. It holds across the tested architectures and scales and reproduces on a second GPU architecture. A causal probe projects each update off the current QK weights' leading three singular directions: the query projection's largest singular value stays at 11.1, whereas removing equal energy elsewhere leaves it at 237. The QK channel therefore drives the early runaway rather than merely tracking it. Entry depends on temporal sign-coherence across steps, not aggregate deviation. QK-Guard closes the channel with a dormant controller that switches on parameter-free QK normalization when attention-logit saturation begins. It contains every tested runaway and matches always-on QK normalization over 60k steps, while non-QK actions at the same trigger fail. The results support intervention at the shared QK locus rather than separate repair at each fault source.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning
Authors:
Jingqi Tian,
Haoji Zhang,
Lin Chen,
Hongbo Jin,
Haonan Xu,
Tianrui Zhu,
Xingming Shui,
Shilin Ma,
Wenjing Yang,
Yansong Tang
Abstract:
Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large language model can adapt its reasoning effort to each question. We propose AdaThinkV, an adaptive framework for video reasoning that learns whether to reason explicitly without offline difficulty labels, manually tuned conf…
▽ More
Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large language model can adapt its reasoning effort to each question. We propose AdaThinkV, an adaptive framework for video reasoning that learns whether to reason explicitly without offline difficulty labels, manually tuned confidence thresholds, or an external router. During reinforcement learning, AdaThinkV samples matched rollouts in explicit reasoning and direct answering modes for each prompt. ThinkGain estimates the prompt-level utility of explicit reasoning by balancing its accuracy gain against additional response length, providing supervision for both conditional response generation and autonomous mode selection. For difficult prompts, limited rollout exploration can yield groups in which every response is unsuccessful and accuracy rewards show little variation, providing insufficient signal for learning. We therefore introduce Variance Recovery Policy Optimization (VRPO), which retains and progressively expands these groups to recover informative signals from prompts that are difficult yet solvable. At inference, AdaThinkV selects a response mode and generates the response in a single autoregressive sequence. Across a unified suite of video reasoning evaluations, AdaThinkV achieves a mean accuracy of 40.79 with an average of 257.20 output tokens, outperforming the strongest evaluated adaptive baseline by 2.98 points while using 22.7% fewer tokens. Project page: https://trilarflagz.github.io/AdaThinkV/
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
Symmetry-Projected Weakly Compatible Multiparameter Quantum Sensing
Authors:
G. R. Jin,
Z. Y. Zhou,
W. Yang
Abstract:
Achieving joint quantum-enhanced precision in multiparameter sensing requires both high sensitivity and measurement compatibility. These two aspects are characterized by the quantum Fisher information matrix (QFIM) and the Uhlmann curvature matrix (UCM), respectively, with weak compatibility corresponding to the vanishing of the relevant UCM elements. Here, we develop a symmetry-projection framewo…
▽ More
Achieving joint quantum-enhanced precision in multiparameter sensing requires both high sensitivity and measurement compatibility. These two aspects are characterized by the quantum Fisher information matrix (QFIM) and the Uhlmann curvature matrix (UCM), respectively, with weak compatibility corresponding to the vanishing of the relevant UCM elements. Here, we develop a symmetry-projection framework that classifies phase generators into subspace-preserving and subspace-changing sectors. For probe states confined to a symmetry subspace, symmetry projection imposes a common block-diagonal structure on the QFIM and UCM, rendering cross-sector parameters simultaneously free from information cross-talk and measurement incompatibility. When the subspace-changing generators act as scalars within the occupied subspace, the corresponding QFIM block reduces to four times the symmetrized covariance matrix, even for mixed probe states. For parity-protected collective $\mathrm{SU}(2)$ systems, this structure singles out the transverse anti-squeezed quadrature and the longitudinal mean-spin direction as natural optimal sensing axes. Applied to a dissipative one-axis twisting model, the dynamically generated probe state exhibits identically vanishing UCM elements for transverse--longitudinal parameter pairs, while maintaining nearly balanced, Heisenberg-scaled QFIM components over a broad transient window. Our work opens a route to symmetry-protected, weakly compatible multiparameter sensing in interacting quantum many-body systems.
△ Less
Submitted 7 August, 2026; v1 submitted 3 August, 2026;
originally announced August 2026.
-
Direct Detection of Light Self-Interacting Dark Matter via Electronic Collective Excitations
Authors:
Wen-Na Yang,
Wu-Long Xu,
Wenyu Wang,
Ning Liu
Abstract:
Models of light dark matter often invoke a light mediator to facilitate interactions with the Standard Model. If sufficiently light, this mediator can induce long-range self-interactions among dark matter particles, offering a compelling resolution to small-scale structure anomalies. However, direct detection of light self-interacting dark matter (SIDM) remains challenging for conventional detecto…
▽ More
Models of light dark matter often invoke a light mediator to facilitate interactions with the Standard Model. If sufficiently light, this mediator can induce long-range self-interactions among dark matter particles, offering a compelling resolution to small-scale structure anomalies. However, direct detection of light self-interacting dark matter (SIDM) remains challenging for conventional detectors. In this work, we investigate the sensitivity of searches for light SIDM accelerated by high-energy cosmic rays in silicon detectors. Leveraging the electronic collective excitations, we derive 90\% C.L. exclusion limits using public SENSEI and DAMIC-M ionization data. Our constraints can cover a portion of the light SIDM parameter space favored by galactic small-scale anomalies.
△ Less
Submitted 10 August, 2026; v1 submitted 2 August, 2026;
originally announced August 2026.
-
Radio-Gamma-Ray Properties and High-Energy Implications for Fermi Blazars
Authors:
Xu-Hong Ye,
Wen-Xin Yang,
Guo-Hai Chen,
Zhi-Yuan Pei,
Yong-Yun Chen,
Yi Liu,
Denis Bastieri,
Jun-Hui Fan
Abstract:
Radio and $γ$-ray emissions in blazars, a subclass of active galactic nuclei (AGNs), provide important insight into their high-energy radiation processes. We studied the relation between radio and $γ$-ray emissions using a large sample of 1687 \textit{Fermi} blazars, based on the Radio Fundamental Catalogue and the latest Third Data Release of the Fourth \textit{Fermi} AGN Catalogue. A clear corre…
▽ More
Radio and $γ$-ray emissions in blazars, a subclass of active galactic nuclei (AGNs), provide important insight into their high-energy radiation processes. We studied the relation between radio and $γ$-ray emissions using a large sample of 1687 \textit{Fermi} blazars, based on the Radio Fundamental Catalogue and the latest Third Data Release of the Fourth \textit{Fermi} AGN Catalogue. A clear correlation between radio and $γ$-ray fluxes for both BL Lacertae objects (BL Lacs) and flat-spectrum radio quasars (FSRQs) suggests a synchrotron self-Compton (SSC) contribution to both subclasses. The ratio of $γ$-ray and radio emissions, $γ$-ray loudness ($G_{\rm r}$), is further examined with the $γ$-ray photon index ($Γ_γ$) and the synchrotron peak frequency ($ν_{\rm{peak}}$). An anti-correlation between $G_{\rm r}$ and $Γ_γ$ is explained by the shift of the spectral energy distribution rather than the Compton cooling effect. We found that $G_{\rm r}$ shows a positive dependence on $ν_{\rm{peak}}$ for low-synchrotron-peaked BL Lacs (LBLs) and FSRQs, in line with the SSC-contributed scenario, although additional external Compton contributions may account for the substantial scatter observed in LBLs and FSRQs. In contrast, high-synchrotron-peaked BL Lacs (HBLs) reach the plateau of $G_{\rm r}$ between $\log (ν_{\rm peak}/{\rm Hz}) \simeq15.5-16$, possibly indicating the transition from the Thomson to the Klein--Nishina (KN) regime. Interpreting this feature within a one-zone SSC framework could constrain the magnetic field strength of $-4.14 < \log (B/{\rm G}) < -1.69$ for those HBLs affected by the KN suppression.
△ Less
Submitted 2 August, 2026;
originally announced August 2026.
-
DGA$_2$D: Directed Graph-Guided Automated Algorithm Design with Large Language Models
Authors:
Jiale Zhao,
Zimu Chen,
Sirui Mao,
Wentao Yang,
Yuxiang Bai,
Liyuanjun Lai
Abstract:
The rapid development of Large Language Models (LLMs) has opened new avenues for Automated Heuristic Design (AHD) for solving NP-hard combinatorial optimization problems (COPs). However, existing LLM-driven AHD methods are largely confined to rigid solver templates, relegating the search process to isolated module tuning. Transitioning to fully autonomous, system-level algorithm design is essentia…
▽ More
The rapid development of Large Language Models (LLMs) has opened new avenues for Automated Heuristic Design (AHD) for solving NP-hard combinatorial optimization problems (COPs). However, existing LLM-driven AHD methods are largely confined to rigid solver templates, relegating the search process to isolated module tuning. Transitioning to fully autonomous, system-level algorithm design is essential but fraught with low reliability of generated operators, extremely large search spaces, and ineffective credit assignment. To overcome these drawbacks, this paper proposes a Directed Graph-Guided Automated Algorithm Design framework, termed DGA$_2$D. It structures the open-ended program space as a directed graph, where each node represents a functional operator that can be instantiated using one of multiple candidate code implementations, while directed walks constitute complete algorithmic pipelines. A first-order path-dependent credit assignment mechanism is introduced to evaluate code variations strictly based on their topological context. Extensive experiments across 12 distinct COPs, ranging from complex scheduling to routing, demonstrate the consistent empirical advantages of DGA$_2$D. It reduces the average normalized gap by up to 10.96 percentage points compared to state-of-the-art LLM baselines.
△ Less
Submitted 1 August, 2026;
originally announced August 2026.
-
Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams
Authors:
Fengxiang Wang,
Qiuyang Yu,
Yueying Li,
Mingshuo Chen,
Chengchi Fei,
Kaiyi Xu,
Lixin Gu,
Wangxu Wei,
Junchao Gong,
Lipeng Ma,
Jiong Wang,
Fenghua Ling,
Wenlong Zhang,
Xue Yang,
Wenjing Yang,
Ben Fei,
Long Lan
Abstract:
Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. Existing remote sensing benchmarks largely rely on static, post-hoc, and expert-processed products, such as gridded reanalysis data, which are difficult to align with operational disaster scenari…
▽ More
Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. Existing remote sensing benchmarks largely rely on static, post-hoc, and expert-processed products, such as gridded reanalysis data, which are difficult to align with operational disaster scenarios where hazards evolve rapidly and decisions must be made under strict time constraints. To bridge this gap, we introduce Obshazard-bench, a real-time, observation-driven benchmark for evaluating disaster intelligence in MLLMs. Unlike image-centric or post-event benchmarks, Obshazard-bench directly integrates raw, high-frequency satellite sounding streams from diverse satellite sensors with concurrent ground-station observations, historical disaster records, and socio-economic indicators, bypassing delayed expert-processing and physical-inversion pipelines. The benchmark covers 8 major disaster categories and 28 sub-categories across more than 60 countries, incorporating over 120 historically documented extreme-event cases and thousands of lifecycle-oriented VQA samples. Moreover, Obshazard-bench further defines a three-stage evaluation taxonomy aligned with the operational disaster workflow: Predictive Crisis Anticipation for pre-disaster risk detection and early forecasting, Active Evolution Reasoning for in-situ disaster tracking and termination prediction, and Multi-faceted Impact Quantification for post-disaster magnitude deduction, humanitarian burden estimation, and socio-economic impact assessment. Experiments on representative general-purpose and Earth-focused foundation models reveal substantial limitations in transforming raw multi-channel physical observations into temporally grounded and decision-relevant disaster reasoning.
△ Less
Submitted 24 June, 2026;
originally announced August 2026.
-
First-principles carrier mobility and optical absorption of strained ZnO with self-consistent Hubbard interactions
Authors:
Hong-Guk Min,
Wooil Yang,
Sabyasachi Tiwari,
Feliciano Giustino,
Young-Woo Son
Abstract:
Carrier mobility and optical absorption are key performance parameters of oxide semiconductors in transparent and flexible displays. We use a newly developed density-functional perturbation theory with a self-consistent Hubbard correction (DFPT+U) to study phonon-limited electron transport and phonon-assisted optical absorption in strained zinc oxide (ZnO). This parameter-free approach accounts fo…
▽ More
Carrier mobility and optical absorption are key performance parameters of oxide semiconductors in transparent and flexible displays. We use a newly developed density-functional perturbation theory with a self-consistent Hubbard correction (DFPT+U) to study phonon-limited electron transport and phonon-assisted optical absorption in strained zinc oxide (ZnO). This parameter-free approach accounts for electron-phonon interactions and on-site correlation effects simultaneously. Electronic structures and phonon dispersions are computed under three distinct uniaxial strain directions. Uniaxial tensile strain up to 4.8% along [\bar110] is found to increase the room-temperature electron mobility by 19% while leaving visible-range optical absorption essentially unchanged. These results demonstrate that moderate strain can selectively enhance carrier transport without degrading optical transparency, and establish DFPT+U as an effective framework for predicting strain-dependent transport and optical properties in wide-band-gap oxides with implications for strain-engineered display and optoelectronic applications.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
On Sirakov's equal-frequency uniqueness conjecture
Authors:
Hong-Ge Chen,
Yong Liu,
Juncheng Wei,
Wen Yang
Abstract:
Let $N\in\{2,3\}$, $0<μ_1\leqμ_2$, and $0<β<μ_1$. We prove that the equal-frequency two-component cubic Schrödinger system \[
-Δu+u=μ_1u^3+βuv^2,
\qquad
-Δv+v=μ_2v^3+βu^2v
\quad\text{in }\mathbb{R}^N \] has exactly one positive solution in $H^1(\mathbb{R}^N)\times H^1(\mathbb{R}^N)$ modulo simultaneous translations. More precisely, every positive solution is a simultaneous translate of the…
▽ More
Let $N\in\{2,3\}$, $0<μ_1\leqμ_2$, and $0<β<μ_1$. We prove that the equal-frequency two-component cubic Schrödinger system \[
-Δu+u=μ_1u^3+βuv^2,
\qquad
-Δv+v=μ_2v^3+βu^2v
\quad\text{in }\mathbb{R}^N \] has exactly one positive solution in $H^1(\mathbb{R}^N)\times H^1(\mathbb{R}^N)$ modulo simultaneous translations. More precisely, every positive solution is a simultaneous translate of the synchronized state constructed from the unique positive radial solution of $-Δw+w=w^3$ in $\mathbb{R}^N$. This settles Sirakov's equal-frequency uniqueness conjecture throughout the weak-coupling range.
The main difficulty in the proof is to exclude radial solutions for which the ratio of the normalized components is nonconstant. After normalization, the two components satisfy scalar equations with a common potential. We construct a weighted Pohozaev functional for the system together with a correction term and prove that both the corrected functional and the associated weighted functional are strictly positive. Combining these sign properties with a radial flux identity and an auxiliary quotient associated with the component ratio forces synchronization.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems
Authors:
Yu Cui,
Wuli Yang,
Yirui Shi,
Junhao Xia,
Hui Jiang,
Lei Gao,
Chenfu Bao
Abstract:
Autonomous multi-agent systems (AMAS) built on large language models (LLMs), such as Hermes, increasingly rely on inference-time harnesses to coordinate reasoning and action. Constructing these harnesses requires substantial engineering effort and computational resources, as they are iteratively optimized over a combinatorial search space while co-evolving with the underlying LLM. Inference-time h…
▽ More
Autonomous multi-agent systems (AMAS) built on large language models (LLMs), such as Hermes, increasingly rely on inference-time harnesses to coordinate reasoning and action. Constructing these harnesses requires substantial engineering effort and computational resources, as they are iteratively optimized over a combinatorial search space while co-evolving with the underlying LLM. Inference-time harnesses therefore constitute valuable intellectual property (IP). Although prior work has investigated IP leakage in static multi-agent systems with pre-configured architectures, it remains unclear whether similar risks arise in AMAS, where harness behavior emerges dynamically during inference. To address this gap, we introduce Agent Harness Distillation (AHD), a framework for studying the security risks arising from inference-time harness extraction in AMAS. We formalize harness extraction as a new security problem and develop an evaluation framework for quantifying such risks. AHD extracts inference-time harness capabilities from a target agent through black-box interactions and consists of two stages. In the pre-distillation stage, AHD infers inference-time harness behaviors from the responses of the target agent and constructs an initial harness. In the post-distillation stage, AHD iteratively refines the initial harness to align with the behavioral patterns of the target agent. Experiments on real-world AMAS across multiple backbone LLMs demonstrate the effectiveness of AHD and reveal substantial IP leakage risks. We further propose a deception-based defense that reduces harness extraction effectiveness while preserving the utility of the protected agent. Our findings uncover a previously underexplored security threat to AMAS.
△ Less
Submitted 19 August, 2026; v1 submitted 30 July, 2026;
originally announced July 2026.
-
One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting
Authors:
Rui Tang,
Wentao Yang,
Peirong Zhang,
Yongxin Shi,
Shun Zhang,
Huiguo He,
Lianwen Jin
Abstract:
Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for Multimodal Large Language Models (MLLMs), the previous multi-patch paradigm often introduces redundant noise and localization ambiguity, particularly for dense or small text instances. To address this, we propose Single-…
▽ More
Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for Multimodal Large Language Models (MLLMs), the previous multi-patch paradigm often introduces redundant noise and localization ambiguity, particularly for dense or small text instances. To address this, we propose Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement. To accurately identify this anchor without oracle labels, we introduce Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discrete visual-token selection using patch-level rewards. To further improve representation robustness and localization precision, we introduce Directional Embedding Alignment (DEA) to suppress unstable norm bias by decoupling feature magnitude and direction, and Patch-Enhanced Decoding (PED) to fuse the routed anchor with language semantics and cross-attend over the full-image feature map for geometry-aware boundary regression beyond coordinate-space surrogates. Extensive experiments demonstrate that SPaTS consistently and significantly outperforms both frontier closed-source MLLMs and OCR MLLMs. Code is available at https://github.com/eeNickTang/SPaTS.
△ Less
Submitted 3 August, 2026; v1 submitted 30 July, 2026;
originally announced July 2026.
-
Unconventional and Fragile Magnetic Exciton in a van der Waals Quantum Magnet
Authors:
Kai-Xuan Zhang,
Min Zhang,
Minjae Kim,
Yong-Hyun Kim,
Junghyun Kim,
Heejun Yang,
Pyeongjae Park,
Chaebin Kim,
Mangesh Diware,
Junik Hwang,
Youjin Lee,
Byeong-Gwan Cho,
Hyeong-Do Kim,
Tae-Yeong Koo,
Chunhua Chen,
Mingtao Li,
Xujie Lü,
Wenge Yang,
Kee-Hoon Kim,
Seung-Ho Baek,
Hyeonsik Cheong,
Sung-Keun Lee,
Beom Hyun Kim,
Christopher Lane,
Jian-Xin Zhu
, et al. (3 additional authors not shown)
Abstract:
The recently discovered magnetic exciton in the van der Waals (vdW) antiferromagnet NiPS3 exemplifies these phenomena, exhibiting several distinctive characteristics. Despite extensive investigation, much of its physics remains unresolved, with key questions about why the NiPS3 magnetic exciton is so sharp and optically bright despite the nominally spin-forbidden transition, posing significant cha…
▽ More
The recently discovered magnetic exciton in the van der Waals (vdW) antiferromagnet NiPS3 exemplifies these phenomena, exhibiting several distinctive characteristics. Despite extensive investigation, much of its physics remains unresolved, with key questions about why the NiPS3 magnetic exciton is so sharp and optically bright despite the nominally spin-forbidden transition, posing significant challenges to a proper understanding and practical manipulation of the exciton. An urgent question is to what extent it is due to chemical disorder, magnetic weakening, lattice modification, or intrinsic instability of the bright exciton itself: answers to which will put stringent constraints on possible theoretical models. Here we address these questions using hydrostatic pressure as a clean, continuous, reversible, and in-situ tuning parameter. We find that the sharp photoluminescence peak is drastically suppressed by as little as 0.4 GPa and completely quenched by 1.5 GPa, with demonstrating its reversibility. Crucially, this bright-to-dark conversion occurs without magnetic, crystallographic, or electronic reconstruction despite an increase in the Neel temperature, as established by Raman, X-ray absorption, nuclear magnetic resonance spectroscopy, and first-principles many-body calculations. Our results demonstrate that the optical brightness of the magnetic exciton is independent of chemical disorder, lattice expansion, and weakening of magnetic order, indicating that a higher-order correlated mechanism governs the bright exciton. We further propose experimentally constrained microscopic scenarios involving exciton pairing, crystal-field-controlled spin-orbit mixing, and symmetry breaking, providing a framework for future tests of entangled magnetic exciton in correlated quantum magnets.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Authors:
Shawn Li,
Wei Yang,
Jike Zhong,
Jiate Li,
Jiawei Yang,
You Qin,
Ryan Rossi,
Franck Dernoncourt,
Roger Zimmermann,
Yue Wang,
Zhengzhong Tu,
Vicente Ordonez,
Mohit Bansal,
Yue Zhao
Abstract:
Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content,…
▽ More
Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4$\times$4 to 16$\times$16), we find that \textbf{zero-shot VLMs largely lack geometric reasoning}: only one of five frontier models (GPT-5.5) exceeds random baseline on 4$\times$4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves $>$97\% on 4$\times$4, \textbf{all models collapse on larger grids}: GPT-5.5 drops from 70\% to near-random on 8$\times$8, and even fine-tuned models fall below 5\% on 12$\times$12. This ``scaling cliff'' suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. \ours{} establishes scalable geometric reasoning as an open challenge for vision-language models.
△ Less
Submitted 3 August, 2026; v1 submitted 30 July, 2026;
originally announced July 2026.
-
A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
Authors:
Wenhao Yang,
Runzhi He,
Minghui Zhou
Abstract:
Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate coding agents' behavior, spanning from a total ban, mandatory disclosure, to verification gates and human sign-offs. Yet, whether coding agents read and follow those rules, and behave in open source repositories, remains unknown. To estimate real-world rule compli…
▽ More
Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate coding agents' behavior, spanning from a total ban, mandatory disclosure, to verification gates and human sign-offs. Yet, whether coding agents read and follow those rules, and behave in open source repositories, remains unknown. To estimate real-world rule compliance of coding agents, we curate 106 issues from 49 repositories containing AI contribution rules into RepoComplianceBench. We judge the trajectory of each run against the repository's rules, measuring whether the agent refuses to contribute, discloses its assistance truthfully, clears the required verification gates, or escalates critical steps to a human. We also test if extra prompts, rule disclosure, or feedback from the compliance verifier help with the situation. Our experiments on four frontier models show that today's agents almost never proactively retrieve the contribution rules. Agents pick up disclosure and verification with reminder prompts, rule quotes, and verifier feedback; however, they never refuse to contribute in AI-banned repositories under any condition we tested. The status reveals that verification and disclosure issues are solvable with existing mechanisms, yet enforcing bans and human escalations remains an open problem.
△ Less
Submitted 29 July, 2026;
originally announced July 2026.
-
Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing
Authors:
Fengxiang Wang,
Jiangnan Huang,
Mingshuo Chen,
Yueying Li,
Yang Shi,
Junwei Luo,
Haoyu Wang,
Yansheng Li,
Jing Zhang,
Haiyan Zhao,
Wenjing Yang
Abstract:
Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-relevant evidence is often sparse, local, and spatially dispersed across extremely large visual contexts. A natural solution is to equip MLLMs with zoom-in tools for active local inspect…
▽ More
Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-relevant evidence is often sparse, local, and spatially dispersed across extremely large visual contexts. A natural solution is to equip MLLMs with zoom-in tools for active local inspection. However, through a pilot study on XLRS-Bench, we find that zoom-in is only partially effective: it resolves easy and medium-level tasks with locally recoverable evidence, but saturates on hard cases requiring global search, multi-region comparison, path planning, or dispersed-evidence reasoning. Motivated by this finding, we move beyond single-tool zoom-in and introduce GeoMTVR, a large-scale Geospatial Multi-Tool Visual Reasoning dataset built from wide-area satellite imagery. GeoMTVR contains 13K UHR VQA samples with interleaved reasoning trajectories, diverse visual tool calls, and returned visual observations, enabling models to learn question decomposition, tool selection, regional inspection, object-level grounding, auxiliary visual reasoning, and cross-tool evidence integration. Beyond supervised fine-tuning, we propose a tool-attention-focused reinforcement learning algorithm that concentrates optimization on critical tool-use decisions, including when to invoke tools, which tool to select, where to apply it, and how to interpret tool outputs. By combining SFT on GeoMTVR with our RL algorithm, we develop GeoLens, a multi-tool visual reasoning MLLM for UHR RS. Experiments show that GeoLens consistently outperforms direct reasoning and single-tool zoom-in baselines, achieving stronger accuracy, better evidence grounding, and more efficient tool-use trajectories.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.