-
Scalar dark matter in space-based gravitational-wave detectors: center-of-mass motion, size breathing, and TDI projection
Authors:
Rui-Yang Hu,
Yuan-Zhi Li,
An-Qi Wang,
Zong-Ru Zou,
Fa-Peng Huang,
Cheng-Gang Qin
Abstract:
Ultralight scalar dark matter can make space-based gravitational-wave detectors respond through both the scalar charge of freely falling test masses and scalar-induced changes of local solid length scales. Existing space-detector forecasts usually model the former as a center-of-mass force, while ground-based interferometer studies show that scalar fields can also act through material and optical-…
▽ More
Ultralight scalar dark matter can make space-based gravitational-wave detectors respond through both the scalar charge of freely falling test masses and scalar-induced changes of local solid length scales. Existing space-detector forecasts usually model the former as a center-of-mass force, while ground-based interferometer studies show that scalar fields can also act through material and optical-path transduction. We ask which part of a local material response survives after one-way Doppler measurements are assembled into delayed time-delay-interferometry observables. To this end, we formulate center-of-mass motion and endpoint-size breathing in a common link-response notation for LISA-, Taiji-, and TianQin-like detectors. The main result is a projection rule: in the equal-arm, identical-endpoint, common-field limit, endpoint breathing enters Michelson-$X$ as a common-mode link perturbation and is removed from the retained channel. Its leading leakage is controlled by finite scalar wave vector, unequal or time-dependent arms, nonidentical endpoint response, or auxiliary readouts, and carries extra geometric and delay suppressions beyond the local size response. We then give reproducible noise, sensitivity, and network-combination formulas, and quote multi-mission improvements only under explicit independent-stream and scalar-coherence assumptions. The result provides a controlled baseline for deciding when test-mass breathing can be neglected and when instrument-specific material response must be modeled.
△ Less
Submitted 21 August, 2026;
originally announced August 2026.
-
Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs
Authors:
Yifan Lu,
Adinath Dukre,
Abhijit Das,
Ziyun Zou,
Haolin Yang,
Yutong Xie,
Imran Razzak
Abstract:
Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insufficiently grounded in image evidence. Mitigation methods applied during decoding offer a practical solution, but they typically lack anatomical awareness or rely heavily on ground trut…
▽ More
Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insufficiently grounded in image evidence. Mitigation methods applied during decoding offer a practical solution, but they typically lack anatomical awareness or rely heavily on ground truth annotations, which limits their applicability. We propose Counterfactual Anatomy-guided Spatial-Temporal decoding (CAST), a framework that operates entirely during inference and requires no manual annotations for anatomically grounded hallucination mitigation. CAST automatically discovers anatomical regions relevant to the given query through broad medical segmentation. It then selects a compact, causally informative area using counterfactual intervention based on the drop in answer likelihood under occlusion. Guided by this chosen region, CAST performs a unified contrastive decoding process, combining classifier-free guidance to correct spatial attention with stepwise temporal contrast to regulate generation dynamics. Experiments on the SLAKE and MIMIC-CXR datasets across three Med-VLMs demonstrate that CAST consistently outperforms strong baselines and surpasses decoding strategies reliant on ground truth. Our results indicate that compact, automatically selected regions provide highly effective contrastive guidance without expert annotations, offering a practical and generalizable solution for improving spatial grounding and reducing hallucinations. Code is available at https://github.com/csyifan/CAST.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Nonstandard Solution for Anomaly Cancellation as Seesaw Neutrino Origin in the SM
Authors:
Zi-Yue Zou,
Chia-Wei Liu,
Zhong-Lv Huang,
Xiao-Gang He
Abstract:
For fixed Standard Model (SM) non-Abelian representations of 15 chiral fermions with arbitrary hypercharges, anomaly cancellation admits the usual assignment and a distinct nonstandard solution. In the latter, the colored exotic quark and exotic lepton weak doublets and exotic lepton singlet have zero hypercharge, whereas the two colored exotic quark singlets carry opposite hypercharges $-q$ and…
▽ More
For fixed Standard Model (SM) non-Abelian representations of 15 chiral fermions with arbitrary hypercharges, anomaly cancellation admits the usual assignment and a distinct nonstandard solution. In the latter, the colored exotic quark and exotic lepton weak doublets and exotic lepton singlet have zero hypercharge, whereas the two colored exotic quark singlets carry opposite hypercharges $-q$ and $q$. The exotic neutral lepton singlets naturally play the role of the heavy neutrinos. The minimal model with two exotic lepton copies gives a rank-two seesaw, the minimal seesaw model, with one massless active neutrino and predicts $m_{ββ} = 1.2 \text{-} 4.1 \, \mathrm{meV}$ for normal ordering or $15.9\text{-}48.9 \,\mathrm{meV}$ for inverted ordering. In a direct SM realization, generating exotic quark masses through the SM Higgs mechanism forces the exotic quarks to carry electric charges $\pm 1/2$. A separate $\mathrm{SU}(2)_{L'}$ realization of the nonstandard solution can allow exotic quarks from several TeV to $10\,\mathrm{TeV}$ with order-one Yukawa couplings. In this model, the charged exotic quarks and leptons carry electric charges $\pm q$. In both cases, the lightest exotic quark and lepton are stable, but suitable choices of their charges and masses can satisfy experimental constraints.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation
Authors:
Chenrun Wang,
Mingxuan Zhu,
Tiancheng Huang,
Wenjie Li,
Yujie Zhang,
Zichen Zhu,
Zhiying Zou,
Kai Yu,
Lu Chen
Abstract:
With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability t…
▽ More
With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability to provide unified and reliable assessments across a coherent distribution of generated ideas. To address this challenge, we propose LigBench, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions. In addition, we introduce PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation. Extensive experiments demonstrate that LigBench achieves stable and interpretable evaluations, significantly improving alignment with expert judgments. Furthermore, models trained on PAIR-IQ exhibit enhanced ranking accuracy and robustness, establishing a principled standard for scalable and objective research idea assessment.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
InterSAGE: The Secure and Verifiable Interoperability Protocol for An Internet of Agents
Authors:
Zhenhua Zou,
Sheng Guo,
Qiuyang Zhan,
Lepeng Zhao,
Shuo Li,
Zhuotao Liu
Abstract:
The emerging Internet of Agents enables LLM-powered agents to discover peers, invoke tools, and delegate tasks across organizational boundaries. Existing protocols increasingly define how agents exchange messages, but not how an agent proves its identity, authorization, advertised capabilities, or accountability after delegation. We present InterSAGE, a trust-native protocol suite that supplies th…
▽ More
The emerging Internet of Agents enables LLM-powered agents to discover peers, invoke tools, and delegate tasks across organizational boundaries. Existing protocols increasingly define how agents exchange messages, but not how an agent proves its identity, authorization, advertised capabilities, or accountability after delegation. We present InterSAGE, a trust-native protocol suite that supplies this missing security substrate alongside, rather than in place of, communication protocols. InterSAGE comprises four layers: Persistent Identity, Discovery, Trust Negotiation, and Accountability. Its four core primitives are: (1) Agent Identity Cards that bind developer, code package, operator, and deployment context; (2) capability-aware discovery using DID-bound Verifiable Credential manifests; (3) trust negotiation combining monotonic capability attenuation with two-tier access control; and (4) kernel-mediated cryptographic audit trails that bind usage, delegation, and execution traces to agent identity without a consensus ledger. InterSAGE is designed to complement MCP, A2A, ANP, and AG-UI, allowing communication protocols to evolve independently while keeping trust semantics explicit, portable, and verifiable. We compare InterSAGE with more than 50 efforts spanning agent protocols, decentralized identity, OAuth/OIDC extensions, zero-trust governance, delegation, and audit architectures. We show that no prior architecture jointly enforces persistent identity, capability-aware discovery, trust negotiation, and accountability as a unified four-layer trust substrate for secure agent interoperability.
△ Less
Submitted 13 August, 2026; v1 submitted 13 August, 2026;
originally announced August 2026.
-
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
Authors:
Peng Cai,
Zhaofan Zou,
Shifa Liu,
Yikun Wang,
Jiawei Tang,
Kaicheng Yang,
Meng Tong,
MingKun Jiang,
Zhongjiang He,
Hao Sun
Abstract:
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documen…
▽ More
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.
△ Less
Submitted 18 August, 2026; v1 submitted 13 August, 2026;
originally announced August 2026.
-
Projection measurement of the comb basis through free-electron-photon interactions
Authors:
Zihang Zou,
Feng-Xiao Sun,
Yunquan Liu,
Qiongyi He
Abstract:
Free electrons, driven by rapid advances in photon-induced near-field electron microscopy, have emerged as a promising platform for quantum information processing, including quantum computing and quantum sensing. However, conventional measurements that rely on the electron energy loss spectrum (EELS) are inherently destructive to electron qubits, thereby constraining their applicability. In this L…
▽ More
Free electrons, driven by rapid advances in photon-induced near-field electron microscopy, have emerged as a promising platform for quantum information processing, including quantum computing and quantum sensing. However, conventional measurements that rely on the electron energy loss spectrum (EELS) are inherently destructive to electron qubits, thereby constraining their applicability. In this Letter, we propose a scheme that performs projection measurement on the electron comb basis, where high measurement precision can be achieved with bright squeezed vacuum states and strong PINEM couplings. Notably, this approach is not only nondestructive to electron qubits but also maximally incompatible with energy measurements, enabling alternative quantum information applications, such as quantum error mitigation and Einstein-Podolsky-Rosen steering detection. Our findings open an avenue towards a systematic understanding of quantum free electrons and towards the development of nondestructive free electron quantum information tasks.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
DA-NBV: A Direction-Aware Next-Best-View Planner for Efficient 3D Reconstruction of Ships at Sea
Authors:
Jiaming Chen,
Juntao Yang,
Zhentao Zou,
Qi Ming,
Yi Yu,
Zhihang Zhong,
Xue Yang,
Xue Jiang,
Yue Zhou
Abstract:
Accurate 3D reconstruction of ships at sea is important for maritime supervision, damage assessment, and autonomous maritime operations. Although 3D reconstruction has advanced considerably, high-quality data acquisition still largely relies on manually designed trajectories or skilled operators, resulting in high costs and limited scalability. Next-best-view (NBV) planning automates this process…
▽ More
Accurate 3D reconstruction of ships at sea is important for maritime supervision, damage assessment, and autonomous maritime operations. Although 3D reconstruction has advanced considerably, high-quality data acquisition still largely relies on manually designed trajectories or skilled operators, resulting in high costs and limited scalability. Next-best-view (NBV) planning automates this process by selecting subsequent viewpoints based on the current state. However, existing NBV policies mainly model spatial occupancy while overlooking directional observation history. This limitation is particularly problematic for ships: their complex superstructures and severe self-occlusions require observations from multiple viewpoints, and insufficient directional coverage often yields incomplete reconstructions. These challenges are further amplified at sea, where wave-induced heave, roll, and pitch continuously alter the ship's pose and surface visibility. Meanwhile, wind disturbances and limited onboard power impose stricter requirements on scanning efficiency. To address these challenges, we propose DA-NBV, a direction-aware NBV policy that augments the conventional occupancy state with directional observation statistics. We introduce a learnable Position Advantage Field (PAF) that uses directional information to guide viewpoint selection. The policy further adopts a locally constrained action space and a nonlinear coverage-shaping reward to improve scanning efficiency. We also develop the ship-oriented SeaShip-3D dataset and a configurable sea-state simulation environment. Experiments under varying heave, roll, and pitch conditions show that DA-NBV improves reconstruction completeness by approximately 3 percentage points and reduces Chamfer distance by 43% while achieving higher path efficiency.
△ Less
Submitted 8 August, 2026;
originally announced August 2026.
-
Depth-Guided Video Object Counting in Crowded Scenes
Authors:
Yuanjing Xu,
Xinyan Liu,
Weidong Chen,
Zixuan Zou,
Linhao Zhang,
Zhuangzhe Meng,
Antoni B. Chan,
Weigang Zhang
Abstract:
Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, limiting their discriminative ability in crowded and occluded conditions. To address this, we propose a Depth-Guided Detector (DG-Det) along with a general post-processing pipeline.…
▽ More
Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, limiting their discriminative ability in crowded and occluded conditions. To address this, we propose a Depth-Guided Detector (DG-Det) along with a general post-processing pipeline. By integrating depth cues with multi-scale RGB-D cross-attention and explicit occlusion prediction, our method enhances spatial understanding and achieves robust detection in crowded and occluded scenes. Furthermore, we introduce a unified de-duplication framework to eliminate cross-frame redundant counting. To facilitate future research, we also release a new RGB-D Video Object Counting dataset featuring depth information and multiple object categories persequence. Extensive experiments demonstrate that our method achieves a 62.01\% reduction in MAE compared to existing baselines, and also produces consistent improvements in RMSE. We provide the source code at https://github.com/streamer-AP/DG-Net and the dataset at https://huggingface.co/datasets/aerospace123/RGBD-VideoCount.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models
Authors:
Yuanruyi,
Yue Cao,
Haojia Gao,
Guanqiu Guo,
Ziyuezhang,
Shangqin,
Junbo Tan,
Bokui Chen,
Zhuo Zou,
Xueqian Wang
Abstract:
Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challenged by physical events under occlusion, where later predictions may depend on object evidence that is no longer available in the current view. Addressing this challenge requires historical evidence not only to be preserved but also to remain acces…
▽ More
Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challenged by physical events under occlusion, where later predictions may depend on object evidence that is no longer available in the current view. Addressing this challenge requires historical evidence not only to be preserved but also to remain accessible when it becomes relevant to a subsequent prediction. Existing approaches mainly enlarge the temporal context, cache generic video features, or impose explicit object-centric states, thereby improving the capacity or structure of retained history. However, they do not directly address how relevant historical evidence can be selectively retrieved and integrated into a pretrained predictor without interfering with its native latent workspace. Accordingly, we introduce HERA (Historical Evidence Routing Adapter), a framework for routing retained historical evidence into a frozen latent predictor, and instantiate it with Register-Routed Patch Memory (RRPM), a lightweight adapter comprising a Structured Memory Bank, Memory Registers, and Workspace Registers. On the IntPhys2 Main split, HERA with RRPM improves the pairwise AvgSurprise accuracy of V-JEPA 2-G from 52.57% to 54.35%. Subgroup analysis shows particularly strong improvements on fixed-camera continuity, from 46.15% to 57.69%, and fixed-camera immutability, from 46.15% to 63.46%. These results support historical evidence routing as a practical adaptation strategy for physical prediction in latent world models.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
Retrieve in Time, Correct in Frequency
Authors:
Yuze Fan,
Yue Cao,
Pengjie Gao,
Haojia Gao,
Guangqiu Guo,
Ziyue Zhang,
Junbo Tan,
Bokui Chen,
Zhuo Zou,
Xueqian Wang
Abstract:
Frozen vision-language-action (VLA) policies generate temporally extended action chunks, but long-horizon manipulation remains vulnerable to accumulated execution error and visual aliasing across task stages. Successful rollouts provide useful corrective evidence, yet current frame retrieval can return progress-misaligned actions,while direct replay or time-domain fusion can overwrite the reactive…
▽ More
Frozen vision-language-action (VLA) policies generate temporally extended action chunks, but long-horizon manipulation remains vulnerable to accumulated execution error and visual aliasing across task stages. Successful rollouts provide useful corrective evidence, yet current frame retrieval can return progress-misaligned actions,while direct replay or time-domain fusion can overwrite the reactive structure of the policy proposal. We introduce Retrieve in Time, Correct in Frequency (RTCF), a training-free test-time correction framework that improves frozen VLA performance with low model-side overhead.RTCF separates which experience to retrieve from which part of its action to transfer. Progressive Memory Alignment (PMA) causally aligns the growing visual execution history with complete successful trajectories through incrementally updated monotonic frontiers, jointly identifying a relevant memory and the current aligned memory position without stage labels. From the aligned action chunk,RTCF transfers a coefficient-wise-clipped low-frequency residual on motion channels. Higher-frequency components and gripper decisions remain inherited from the frozen policy. Across four LIBERO suites and 2,000 episodes per condition, RTCF raises aggregate success from 86.4% to 88.4% and improves LIBERO-Long from 61.6% to 68.6%.These gains require no parameter updates, repeated VLA inference, or additional GPU resources: correction can be performed on the client CPU after a single policy invocation, and the median latencies sum to only 10.99 ms per action chunk
△ Less
Submitted 6 August, 2026; v1 submitted 5 August, 2026;
originally announced August 2026.
-
LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching
Authors:
Siwei Yu,
Han Guo,
Zhenwei Shi,
Zhengxia Zou
Abstract:
Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending dense matching to global-scale remote sensing remains challenging because image pairs may differ in acquisition time, season, viewpoint, spatial resolution, and land-cover state. The resulting large geometric offsets, partial overlap, and intrinsica…
▽ More
Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending dense matching to global-scale remote sensing remains challenging because image pairs may differ in acquisition time, season, viewpoint, spatial resolution, and land-cover state. The resulting large geometric offsets, partial overlap, and intrinsically unmatchable regions make direct dense correspondence prediction unreliable and inefficient. We thus reformulate dense matching as localization-and-registration: first localizing the matchable overlap and affine geometry, then refining dense residuals within the aligned frame. Based on this formulation, we propose LoRetta, a foundation model coupling matchability-aware affine localization with guided dense registration. We also introduce LEVIR-GM, a global-scale multi-temporal optical matching benchmark with dataset-native matchability labels (103K aligned, 827K augmented pairs, six continents, five years, 0.5-1024 m resolution). We further establish a unified evaluation protocol for sparse, semi-dense, and dense matchers. On LEVIR-GM, LoRetta achieves an area under the curve (AUC) of 83.3%, outperforming the strongest baseline RoMa v2 by 1.6 points, with larger percentage of correct keypoints (PCK) gains of 6.5 and 8.2 points at 1 and 2 pixels, while reducing inference latency by 47.8%. Astronaut-to-satellite and unmanned aerial vehicle (UAV)-to-satellite geolocalization experiments further demonstrate its transferability as a reusable geometric aligner.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis
Authors:
Zhishan Zou
Abstract:
Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid implausible attribute combinations, preserve an everyday photographic character, and expose quality-control decisions at scale. We present Poplar, a reproducible Spe…
▽ More
Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid implausible attribute combinations, preserve an everyday photographic character, and expose quality-control decisions at scale. We present Poplar, a reproducible Specify--Render--Inspect pipeline for human-centric image dataset synthesis. Specify samples structured attributes under commonsense constraints and verbalizes them as photography-oriented prompts. Render uses a realism-adapted image generator across composition-aware aspect ratios and retries obvious technical failures. Inspect applies a single structured vision--language review to each candidate, preserving the original prompt while rejecting intrinsic image defects or material prompt mismatches. Using Poplar, we construct Poplar-9K: 9,401 curated human-centric image--text pairs retained from 11,765 reviewed candidates (79.9\% acceptance). We release the dataset together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human-centric collections.
△ Less
Submitted 1 August, 2026;
originally announced August 2026.
-
FootprintNet: State-Transition-Guided Dynamic Footprint Learning for Multi-temporal Remote Sensing Change Detection
Authors:
Haotian Zhang,
Hao Chen,
Han Guo,
Zhengxia Zou,
Zhenwei Shi
Abstract:
Despite substantial progress in remote sensing multi-temporal change detection (MTCD), most existing MTCD methods still represent the dynamic process at each spatial location over the entire observation period using a single change category associated with the final observation. This implicit single-change assumption limits their ability to characterize regions of recurrent change closely related…
▽ More
Despite substantial progress in remote sensing multi-temporal change detection (MTCD), most existing MTCD methods still represent the dynamic process at each spatial location over the entire observation period using a single change category associated with the final observation. This implicit single-change assumption limits their ability to characterize regions of recurrent change closely related to human activities. To address this limitation, we introduce Urban Building Dynamics Detection (UBDD), which identifies building-change dynamic footprints, i.e., the temporal intervals in which changes occur, from multi-temporal imagery and produces pixel-wise classification masks. For regions undergoing two or more changes, UBDD introduces an independent multi-change class for unified representation, thereby enabling unified modeling of single- and multi-change processes. Furthermore, we propose FootprintNet, which abstracts building-change processes as interactions between latent states and actions, and imposes state-action transition constraints to guide the learning of causally coherent change trajectories. It further exploits temporal change-boundary cues to enhance feature contrast across boundary sides, thereby improving the discrimination among different dynamic footprints and enabling accurate detection of dynamic footprints. Moreover, we introduce the Building Change Dynamics Score (BCDS) to address the inability of conventional metrics to reflect the temporal proximity between predicted footprints and labels. It evaluates predictions according to their preservation of change semantics and temporal offsets from the corresponding labels. Extensive experiments on TSCD, MUDS, and WUSU demonstrate that FootprintNet outperforms current state-of-the-art methods. The code is available at https://github.com/zmoka-zht/FootprintNet.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Diverse Morphologies of GRB X-Ray Plateaus within a Common Magnetar Framework
Authors:
Xiao-Fei Dong,
Yong-Feng Huang,
Nurimangul Nurmamat,
Chen Deng,
Ze-Cheng Zou,
Fan Xu,
Abdusattar Kurban,
Chen Du,
Chen-Ran Hu,
Jin-Jun Geng
Abstract:
The origin of the X-ray plateau phase in gamma-ray bursts (GRBs) remains an open problem. In particular, it is unclear whether GRBs with different temporal morphologies (i.e., with a rising, flat, or decaying plateau) arise from a common underlying mechanism. Although magnetar energy injection is a leading explanation, previous studies have primarily inferred magnetar properties on a burst-by-burs…
▽ More
The origin of the X-ray plateau phase in gamma-ray bursts (GRBs) remains an open problem. In particular, it is unclear whether GRBs with different temporal morphologies (i.e., with a rising, flat, or decaying plateau) arise from a common underlying mechanism. Although magnetar energy injection is a leading explanation, previous studies have primarily inferred magnetar properties on a burst-by-burst basis and have not tested the model at the population level. Here we perform the first hierarchical population inference of magnetar parameters for a uniform sample of 185 long GRBs with X-ray plateaus within a conditional Poisson point-process framework. It is found that the observed plateau population is well reproduced by physically plausible magnetar populations. The inferred parameter distributions show no strong statistical separation among subclasses with different plateau morphologies. Nevertheless, all subclasses show a substantial intrinsic luminosity scatter, $σ_{L,\rm int}\sim0.5$--1.0 dex, whereas the intrinsic duration scatter remains considerably smaller. The results provide a population-level test of the magnetar interpretation of GRB X-ray plateaus, showing that the observed diversity of plateau morphologies does not require distinct magnetar populations.
△ Less
Submitted 25 July, 2026;
originally announced July 2026.
-
CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data
Authors:
Peiguang Li,
Yongwei Zhou,
Juncheng Diao,
Yuchun Fan,
Jian Yang,
Jianxiao Yang,
Zhongda Su,
Shuguang Jiao,
Xiao Wei,
Zhiye Zou,
Gan Dong,
Zhizhao Zeng,
Rongxiang Weng,
Jingang Wang,
Xunliang Cai
Abstract:
Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance. However, these pipelines typically rely on singular optimization objectives, which inevitably narrows distributional diversity and marginalizes long-tail knowledge, thereby restricting data coverage and underutilizing the…
▽ More
Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance. However, these pipelines typically rely on singular optimization objectives, which inevitably narrows distributional diversity and marginalizes long-tail knowledge, thereby restricting data coverage and underutilizing the vast potential of the open web. To address this limitation, we propose a novel curation paradigm that shifts from linear pruning to the joint optimization of quality, redundancy, and diversity. This framework synergizes dual-track cleaning (rule-based and model-driven) with hybrid deduplication (n-gram and semantic), while employing a multi-objective sampler to balance informational quality with distributional breadth. Applying this framework to Common Crawl, we construct CuraWeb, a 2T-token English corpus. Unlike existing resources, CuraWeb establishes an industrial-grade standard for data curation by recovering a more holistic data distribution with enhanced diversity and minimal redundancy, achieving broader coverage of long-tail knowledge across diverse domains. Experimental evaluations at the 3B scale demonstrate that CuraWeb significantly outperforms state-of-the-art baselines, yielding an average performance gain of 1.8\% across a wide range of benchmarks, particularly in knowledge-intensive and reasoning tasks.
△ Less
Submitted 28 June, 2026;
originally announced July 2026.
-
Learning Spatiotemporal Decision Priors for Efficient Path Planning under Partial Observability
Authors:
Yi Liu,
Hongda Zhang,
Leyao Zou,
Chunlei Meng,
Ziqing Zhou,
Yuning Chen,
Zhuo Zou,
Lida Xu,
Zhongxue Gan,
Chun Ouyang
Abstract:
Path planning under partial observability remains challenging because an agent must make long-horizon navigation decisions from only locally bounded observations. Nevertheless, historical trajectories contain reusable experience-guided directional preferences. Classical planners, however, typically solve each instance from scratch and lack an explicit mechanism to exploit such transferable decisio…
▽ More
Path planning under partial observability remains challenging because an agent must make long-horizon navigation decisions from only locally bounded observations. Nevertheless, historical trajectories contain reusable experience-guided directional preferences. Classical planners, however, typically solve each instance from scratch and lack an explicit mechanism to exploit such transferable decision knowledge, often leading to redundant node expansions and locally myopic search behaviors. Motivated by this limitation, this paper proposes ImiPath, a prior-guided learning framework that distills reusable spatiotemporal decision priors from demonstration trajectories and uses them as experience-informed directional guidance to bias planners toward reliable and promising search directions under partial observability. Specifically, ImiPath first constructs a local spatiotemporal observation representation, which encodes the spatial information of the local environment and the temporal information of historical trajectories. The SpatioTemporal-Attention Policy Network (STAPNet) then transforms this representation into dicision priors. These priors are further incorporated into heterogeneous planners as directional guidance, biasing the search toward locally promising regions. Extensive experiments demonstrate that ImiPath achieves competitive path quality and improves search efficiency by reducing redundant node expansions under local observability. Additional physical experiments on a magnetic microrobot platform further validate the adaptability and practical deployment potential of the proposed framework.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
Authors:
Zhengyu Zou,
Hao Li,
Kuixuan Jiao,
Liu Liu,
Tingyang Xiao,
Xiaolin Zhou,
Fangzhou Hong,
Zhizhong Su,
Dingwen Zhang,
Ziwei Liu
Abstract:
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic rec…
▽ More
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic scenes. In this paper, we propose IGGT4D, a streaming instance-grounded geometry Transformer for online 4D scene understanding. IGGT4D processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. This enables long-sequence feed-forward reconstruction with geometry-instance consistency in dynamic environments. To address the lack of high-quality 4D supervision, we further construct InsScene4D-147K, a large-scale dataset spanning real/synthetic and static/dynamic scenes, with RGB images, depth, poses, and temporally consistent instance masks generated by an automated geometry-guided annotation pipeline. Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Quantum-Enhanced Multi-Objective Optimization
Authors:
Maolin Luo,
Jiapei Zhuang,
Zuoheng Zou,
Man-Hong Yung
Abstract:
Multi-objective combinatorial optimization requires identifying Pareto-optimal trade-off solutions among conflicting objectives, often making it more demanding than its single-objective counterpart. Although quantum multi-objective optimization methods have begun to emerge, most existing quantum optimization workflows are still built around single-objective or fixed-scalarization settings. Buildin…
▽ More
Multi-objective combinatorial optimization requires identifying Pareto-optimal trade-off solutions among conflicting objectives, often making it more demanding than its single-objective counterpart. Although quantum multi-objective optimization methods have begun to emerge, most existing quantum optimization workflows are still built around single-objective or fixed-scalarization settings. Building on existing weighted-sum QAOA approaches to quantum multi-objective optimization, we propose QEMOO, a quantum-enhanced multi-objective optimization framework that combines Pareto-based selection and warm-started QAOA sampling in a multi-round protocol under the same total shot budget. We further introduce a PBI-inspired adaptive direction-update scheme to improve coverage in strongly conflicting benchmark regimes. Across three benchmark stages, QEMOO improves Pareto-front hypervolume over the single-pass weighted-sum QAOA baseline under matched shot budgets, suggesting a practical route toward shot-efficient quantum-assisted multi-objective optimization and its future applications.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Private Approximation of Graph Spectra and Cuts via Spectral Amplifiers
Authors:
Chenglin Fan,
Jingcheng Liu,
Pan Peng,
Hangyu Xu,
Zongrui Zou
Abstract:
We study the problem of releasing a synthetic graph that approximates the sizes of all cuts of an input graph under edge-level differential privacy. If one insists on purely additive error, the optimal worst-case error is $\widetildeΘ(n^{3/2})$. If one allows a small multiplicative slack, an information-theoretic exponential-time mechanism achieves nearly linear additive error, but the best known…
▽ More
We study the problem of releasing a synthetic graph that approximates the sizes of all cuts of an input graph under edge-level differential privacy. If one insists on purely additive error, the optimal worst-case error is $\widetildeΘ(n^{3/2})$. If one allows a small multiplicative slack, an information-theoretic exponential-time mechanism achieves nearly linear additive error, but the best known polynomial-time algorithms have substantially larger error. We give a polynomial-time $(\varepsilon,δ)$-differentially private algorithm which, for every $n$-vertex unweighted graph $G$, outputs a non-negative weighted synthetic graph $\widetilde G$ such that, with high probability, every cut $S\subseteq V(G)$ satisfies \[
|w_G(S)-w_{\widetilde G}(S)|
\le
γw_G(S)+\widetilde O_{\varepsilon,δ,γ}(n^{13/12+o(1)}). \] This improves the previous polynomial-time worst-case bound $\widetilde O(n^{5/4+o(1)})$ of Aamand et al. (ICML 2025) for mixed multiplicative/additive private cut approximation.
The main technical ingredient is a new set of private spectral primitives for bounded-degree graphs, one of them gives spectral error $\widetilde O_δ((nd)^{1/4}/\sqrt\varepsilon)$ in estimating the graph Laplacian for graphs of maximum degree $d$, being the first to beat the standard $\min\{2d,\widetilde O_δ(\sqrt{n}/\varepsilon)\}$ baseline in the high-degree regime. We further develop a primitive with a sharper error dependence on $n$ and $d$ for the downstream cut approximation. Combined with a new edge-sensitive terminal cut oracle with additive error $\widetilde O(n+(n^2M)^{1/3})$ on graphs with $M$ edges, this yields the final worst-case $\widetilde O(n^{13/12+o(1)})$ private cut-release error.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
RCLUPPr: a new randomized CholeskyQR with LU preconditioning
Authors:
Haoran Guan,
Zhenyu Zou,
Yufeng Wei,
Yipei Chen,
Peiting You,
Yuwei Fan
Abstract:
In this work, we present the comprehensive rounding error analysis of RCLUPPr proposed in \cite{RCLUPP}, which is a novel randomized CholeskyQR-type algorithm performing LU decomposition with partial pivoting (LUPP decomposition) directly on the tall-skinny $X\in\mathbb{R}^{m\times n}$ with $m \ge n$ and $\mbox{rank}(X)=n$. In contrast to the existing RCLUPP in \cite{RCLUPP}, which applies matrix…
▽ More
In this work, we present the comprehensive rounding error analysis of RCLUPPr proposed in \cite{RCLUPP}, which is a novel randomized CholeskyQR-type algorithm performing LU decomposition with partial pivoting (LUPP decomposition) directly on the tall-skinny $X\in\mathbb{R}^{m\times n}$ with $m \ge n$ and $\mbox{rank}(X)=n$. In contrast to the existing RCLUPP in \cite{RCLUPP}, which applies matrix sketching before LUPP decomposition, RCLUPPr places LUPP decomposition as a preconditioning step first, significantly reducing error propagation. Our analysis rigorously proves that RCLUPPr enjoys markedly better applicability to the ill-conditioned matrices than the existing CholeskyQR-type algorithms and remains stable and accurate in the mixed-precision arithmetic. We further propose practical acceleration strategies in the real implementations of RCLUPPr. Extensive numerical experiments on the real-world problems confirm the theoretical results in this work, demonstrating the robustness and practicality of RCLUPPr in the single, double, and the mixed-precision architecture.
△ Less
Submitted 16 July, 2026;
originally announced July 2026.
-
Nexus: Native Mesh Generation with Diffusion
Authors:
Hanxiao Wang,
Ying-Tian Liu,
Yuan-Chen Guo,
Qi-Yuan Feng,
Zi-Xin Zou,
Ding Liang,
Biao Zhang,
Yan-Pei Cao
Abstract:
Generating high-quality triangle meshes is essential for film, gaming, and interactive 3D applications. Mainstream methods rely on mesh serialization and autoregressive processes, which stuggles in effective inference and is sensitive to error accumulation. In this paper, we present Nexus, a diffusion method that achieves holistic mesh generation via decoupled vertex and topology generation. First…
▽ More
Generating high-quality triangle meshes is essential for film, gaming, and interactive 3D applications. Mainstream methods rely on mesh serialization and autoregressive processes, which stuggles in effective inference and is sensitive to error accumulation. In this paper, we present Nexus, a diffusion method that achieves holistic mesh generation via decoupled vertex and topology generation. First, we view mesh vertices as sparse voxels organized as an octree and adopt a diffusion model to generate the vertices in a coarse-to-fine manner. Second, for topology modeling, we propose Spacetime Interval, as an extension of Spacetime Distance to encode arbitrary edge and face topology into continuous per-vertex embeddings. It allows for a global and efficient recovery of complex topology. We then employ a diffusion model to generate the continuous embeddings on the generated vertices. Extensive experiments on the Objaverse and Toys4K datasets and in-the-wild images demonstrate that our method outperforms state-of-the-art autoregressive and two-stage baselines, effectively circumventing the inherent limitations of sequential mesh modeling. A blind user study from 3D practitioners confirms strong perceptual preference for our results.
△ Less
Submitted 15 July, 2026;
originally announced July 2026.
-
Precision quantum simulation of magnon spectra and interactions
Authors:
Trond I. Andersen,
Nikita Astrakhantsev,
Jeronimo Martinez,
Will Morong,
Johannes Motruk,
Dario Rossi,
Brayden Ware,
Bryce Kobrin,
Weijie Wu,
Elizabeth Bennewitz,
Manuel Rudolph,
Tom Westerhout,
Amira Abbas,
Rajeev Acharya,
Laleh Aghababaie Beni,
Ross Alcaraz,
Sayra Alcaraz,
Markus Ansmann,
Frank Arute,
Kunal Arya,
Walt Askew,
Juan Atalaya,
Christopher Ayala,
Ryan Babbush,
Brian Ballard
, et al. (307 additional authors not shown)
Abstract:
Quantum simulation promises to advance materials discovery by accurately simulating complex states of matter, their microscopic excitations, and macroscopic response functions. The central challenge in resolving the underlying interacting dynamics is to combine high-fidelity evolution with the sophisticated control necessary to manipulate individual quasi-particles in quantum many-body states. Her…
▽ More
Quantum simulation promises to advance materials discovery by accurately simulating complex states of matter, their microscopic excitations, and macroscopic response functions. The central challenge in resolving the underlying interacting dynamics is to combine high-fidelity evolution with the sophisticated control necessary to manipulate individual quasi-particles in quantum many-body states. Here, we report on high-precision simulation of both linear and non-linear response functions in a 2D XY spin-1/2 magnet using an analog-digital superconducting processor of up to 97 qubits. By interleaving digital gates with analog evolution precisely characterized via Hamiltonian learning, we selectively excite magnons at tunable energy densities. Measuring first the linear magnon response -- a central probe in neutron-scattering experiments -- we extract temperature-dependent spectra and lifetimes. Our results reveal stark variations in magnon decay rates across the Brillouin zone, with enhancement near van Hove singularities and suppression for edge-localized modes. Next, we perform a suite of nonlinear measurements, including the study of self-scattering mechanisms, as well as pump-probe spectroscopy to directly characterize the magnon interactions. While matrix-product state simulations capture the dynamics well in either small systems or at low temperatures, their predictions become inaccurate away from these limits. This work demonstrates precise simulation of the interacting dynamics in quantum magnets, and provides key insights into quasi-particles and their microscopic scattering mechanisms.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
TSSM: Triaxial State Space Model for Global Station Weather Forecasting with Temporal-Variable-Historical Modeling
Authors:
Songru Yang,
Zili Liu,
Tao Han,
Ben Fei,
Fenghua Ling,
Lei Bai,
Chang Liu,
Xiangyang Ji,
Zhenwei Shi,
Zhengxia Zou
Abstract:
Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions. Despite efforts to exploit look-back windows, existing methods show limited accuracy gains and struggle with extreme events and error accumulation. These limitations stem from overreliance on short-term patterns, which are insufficient to capture chaotic weather dynamics, especially…
▽ More
Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions. Despite efforts to exploit look-back windows, existing methods show limited accuracy gains and struggle with extreme events and error accumulation. These limitations stem from overreliance on short-term patterns, which are insufficient to capture chaotic weather dynamics, especially under partial observations. To address this problem, we propose a novel Triaxial State Space Model (TSSM) with a history-enhanced Temporal-VariableHistorical paradigm, which incorporates period-aligned historical weather data to compensate for long-term, large-scale periodic, and full-window weather patterns beyond the temporal lookback window. Specifically, TSSM stacks historical samples into period-aligned batches, where forecasting is causally supported by historical and current observations. Temporal, variable, and historical scanning are designed to capture axial temporal dependencies, variable correlations, and historical evolution. This structure is hierarchically shared to model seasonal to extreme events while alleviating misalignment across historical patterns. TSSM achieves SOTA performance on Weather-5K, the largest station weather dataset to date, with 10% and 61% gains in accuracy and extreme event metrics, and obtains 95% best or second-best results on human-involved datasets. Its advantages are more pronounced in long-horizon and iterative forecasting, reaching a 37.5% gain at 240h and up to 103.5% under a 48h times 5 iterative setting. Moreover, TSSM retains > 90% performance under up to 80% missing observations, compared with < 43% for baselines, demonstrating robustness and practical potential for reliable GSWF in global in-situ observation networks.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Authors:
Zhishan Zou,
Guoyan Sun,
Zhiwei Wei,
Jiancheng Pan,
Yujie Li,
Mugen Peng,
Wenjia Xu
Abstract:
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial…
▽ More
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we introduce SIS-Bench, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation. SIS-Bench organizes evaluation along two complementary dimensions, space and self, and a three-level hierarchy of perception, memory, and reasoning. It contains 4,856 question--answer pairs across 13 tasks derived from 1,646 real-world UAV videos through a task-conditioned construction pipeline with expert verification. Extensive evaluations reveal that current MLLMs exhibit fundamental limitations in modeling dynamic and agent-centered processes. In particular, we observe a clear imbalance between spatial cognition and self-awareness, as well as a progressive performance degradation across cognitive levels. Motivated by these findings, we further explore a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion. Experimental results show that modeling agent motion consistently improves perception and memory performance, not only in spatial cognition but also in self-awareness, and generalizes to downstream UAV decision-making tasks. Our results highlight the importance of self-awareness for advancing embodied spatial intelligence, and provide both a new benchmark and empirical evidence for motion-aware self-in-space modeling.
△ Less
Submitted 14 July, 2026; v1 submitted 14 July, 2026;
originally announced July 2026.
-
SLIDER: Sparse History-Guided Aerial Robot Target Search using Sliding Local Maps
Authors:
Xiaolei Hou,
Zheng Pan,
Hua Lan,
Zhenghao Zou,
Yinhong Chen,
Chenxi Zhu,
Yang Lyu,
Jinwen Hu,
Chunhui Zhao
Abstract:
Efficient exploration and target search in large-scale unknown environments remain challenging for aerial robots due to the demands of broad spatial coverage, fine-grained perception, and real-time decision-making. This paper presents SLIDER, a lightweight and memory-efficient framework that avoids reliance on globally dense maps by combining a local sliding map with sparse global history informat…
▽ More
Efficient exploration and target search in large-scale unknown environments remain challenging for aerial robots due to the demands of broad spatial coverage, fine-grained perception, and real-time decision-making. This paper presents SLIDER, a lightweight and memory-efficient framework that avoids reliance on globally dense maps by combining a local sliding map with sparse global history information. A novel observation quality evaluation method is proposed, leveraging historical poses and sensor models to assess point cloud data in real-time, enabling efficient frontier detection. To support scalable and responsive planning, an incremental viewpoint clustering strategy dynamically adapts to local updates, significantly reducing the number of candidate targets and decreasing computational load. A sparse global topological map is incrementally maintained to assist global planning and cost evaluation. Extensive simulations and real-world experiments demonstrate that the proposed system outperforms state-of-the-art methods in memory usage, decision latency, and search efficiency.
△ Less
Submitted 11 July, 2026;
originally announced July 2026.
-
EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation
Authors:
Dexiang Hong,
Yijie Guo,
Weidong Chen,
Xinyan Liu,
Zixuan Zou,
Zhendong Mao,
Yongdong Zhang
Abstract:
Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion. In this challenge, the main difficulty is that the visual and affective attributes available in the training data are not explicitly provided at test time. Without these attributes, the generator has to decide not only what to depict, but also how…
▽ More
Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion. In this challenge, the main difficulty is that the visual and affective attributes available in the training data are not explicitly provided at test time. Without these attributes, the generator has to decide not only what to depict, but also how the target emotion should be expressed through color, lighting, brushwork, composition, line, and layout. This creates a control gap between the available test prompt and the fine-grained conditions needed for emotion-aware artistic generation. To bridge this gap, we propose EmoStyle, a Z-Image-based framework that converts the input prompt into a structured generation state. An LLM reasoner first predicts affective cues (valence-arousal, dominant emotion, and therapeutic-effect labels) and an aspect-ratio decision. Instead of using these predictions only as additional prompt text, we encode the affective fields into an affective condition vector and inject it into the denoising blocks through AdaLN-style modulation. This allows the inferred control variables to directly guide the generation of intermediate features. Since emotional expression is also style-dependent, we further train a dedicated LoRA adapter for each artistic style bucket and select the corresponding expert during inference, enabling the same affective cues to be rendered with bucket-specific priors for color, texture, brushwork, and composition. Finally, a lightweight VLM-guided candidate selection step ranks the generated images based on prompt alignment, style consistency, emotional expression, and visual quality. In Track 1 of the AffectiveArt Challenge 2026, our USTC\_PI\_LAB\_TEAM submission achieved first place.
△ Less
Submitted 11 July, 2026;
originally announced July 2026.
-
HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM Inference
Authors:
Hui Dong,
Yanzhao Li,
Jie Gao,
Chunlu Li,
Zhiyuan Zhang,
Yupeng Sun,
Zhenyuan Chen,
Zhiqiang Zou
Abstract:
We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM inference on Ascend NPUs, while maintaining the online softmax state in FP16. To our knowledge, HiFA4 is the first Ascend-HIF4-targeted design of this kind evaluated on standard NLP benchmarks.
HiFA4 combines two mechanisms. Smooth-QK applies a calibration-sta…
▽ More
We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM inference on Ascend NPUs, while maintaining the online softmax state in FP16. To our knowledge, HiFA4 is the first Ascend-HIF4-targeted design of this kind evaluated on standard NLP benchmarks.
HiFA4 combines two mechanisms. Smooth-QK applies a calibration-static per-channel equivalent rescaling to Q and K after RoPE, transferring quantization difficulty from K to Q without per-tile online reduction at inference. P-Reordering accumulates the softmax normalizer from the same quantized attention weights P_hat used in the PV GEMM, rather than from a higher-precision reconstruction. We show that this inconsistent formulation introduces a coherent output-scaling error, and validate the effect on a Qwen3-8B Layer-0 MMLU trace, where all 3.6M measured attention tiles exhibit net probability-mass loss with median epsilon_bar = -0.064. P-Reordering also allows the normalizer to be fused into the PV Cube GEMM.
Across five LLMs, HiFA4 reduces quantization-induced decision drift. On Qwen3-8B, it recovers 37.5% of the accuracy gap introduced by direct HIF4 quantization, narrows the sample-weighted accuracy loss from 1.12 pp to 0.70 pp, reduces BF16-inconsistent MMLU predictions from 16.3% to 8.2%, and cuts MMLU accuracy regressions by 57% (1071 to 465). On Gemma2-9B, mild smoothing keeps HiFA4 within 0.7 pp of BF16 while reducing MMLU regressions by 27%. On LLaMA3.1-8B, Mistral-7B, and Phi-4B, where Smooth-QK is disabled, P-Reordering with the adopted Q-Mean auxiliary still reduces full-set MMLU regressions by 41-52%. A preliminary instruction-scheduling analysis projects a 35.4% critical-path latency reduction relative to BF16 by fusing the softmax normalizer into the PV Cube GEMM; on-hardware validation is left to future work.
△ Less
Submitted 5 July, 2026;
originally announced July 2026.
-
Revisiting $\bar B^0 \rightarrow Λ_c^+ \bar p$ decay with higher twist corrections
Authors:
Zhou Rui,
Zhi-Tian Zou,
Ying Li
Abstract:
We investigate the single-charmed baryonic decays $\bar B^0 \to Λ_c^+ \bar p$ and $\bar B^0 \to \barΛ_c^- p$, which receive contributions from both $W$-emission and $W$-exchange topologies, within the framework of perturbative QCD (PQCD). Higher-power corrections associated with the hadronic light-cone distribution amplitudes (LCDAs) of both the initial- and final-state hadrons are systematically…
▽ More
We investigate the single-charmed baryonic decays $\bar B^0 \to Λ_c^+ \bar p$ and $\bar B^0 \to \barΛ_c^- p$, which receive contributions from both $W$-emission and $W$-exchange topologies, within the framework of perturbative QCD (PQCD). Higher-power corrections associated with the hadronic light-cone distribution amplitudes (LCDAs) of both the initial- and final-state hadrons are systematically taken into account. We find that these higher-twist contributions play an important role in baryonic $B$ decays and cannot be neglected. A sizable destructive interference between the $W$-emission and $W$-exchange amplitudes is observed, which significantly reduces the predicted branching fraction of $\bar B^0 \to Λ_c^+ \bar p$ and leads to improved agreement with experimental measurements. The doubly Cabibbo-suppressed decay $\bar B^0 \to \barΛ_c^- p$ is studied for the first time. Its branching fraction is predicted to be of order $10^{-8}$, placing it within the reach of future high-luminosity experiments. We further present the first theoretical predictions for the decay asymmetry parameters of both channels, which provide additional observables for testing the underlying decay dynamics and can be confronted with future experimental data.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving
Authors:
Chen Yang,
Yuhao Wei,
Ze Xu,
Ziheng Zou,
Shuang Liang,
Delin Ouyang,
Lingfeng Qi,
Jie Li,
Guofa Li
Abstract:
Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving intentions and remain insufficient for geometrically accurate, future-aware, and multi-view-grounded planning. To address these limitations, we develop the Layer-Wise Wo…
▽ More
Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving intentions and remain insufficient for geometrically accurate, future-aware, and multi-view-grounded planning. To address these limitations, we develop the Layer-Wise World-Model-Guided Driving framework (LWDrive). LWDrive is a VLM planning framework that refines coarse trajectories through layer-wise world-model guidance. Instead of treating the VLM output as the final trajectory, LWDrive uses it as an intent-aware coarse plan, expands a diverse candidate space around it, and progressively refines the candidates through a Foresight Cascade Planner (FCP). Specifically, we introduce future-frame generation supervision to encourage the VLM to learn forward-looking scene representations, thereby injecting planning-relevant predictive dynamics into its internal hidden states. Built upon these world-model-supervised representations, FCP exploits VLM features across multiple layers and integrates historical temporal states, Action-Query representations, and current-frame multi-view Bird's-Eye-View (BEV) features to refine candidate trajectories in a coarse-to-fine manner. This design enables progressive correction of spatial positions and motion trends while grounding trajectory refinement with multi-view scene cues and preserving the high-level driving intention produced by the large model. Finally, a score head evaluates the refined candidates and selects the best trajectory as the final planning output. Experiments show that LWDrive achieves a score of 92.0 on the NAVSIM benchmark and 89.6 on NAVSIM-v2. Code and models will be made publicly available.
△ Less
Submitted 30 June, 2026; v1 submitted 29 June, 2026;
originally announced June 2026.
-
State Space Models Meet Remote Sensing: A Survey
Authors:
Qinzhe Yang,
Chenyang Liu,
Jia Xu,
Zhenwei Shi,
Zhengxia Zou
Abstract:
State Space Models (SSMs), designed for long-range modeling, offer linear computational complexity and strong capabilities in capturing long-range dependencies. In the field of remote sensing, SSMs have gained popularity due to their effectiveness in addressing unique challenges such as dense visual predictions, multi-modal remote sensing data, and temporal remote sensing data, which have also yie…
▽ More
State Space Models (SSMs), designed for long-range modeling, offer linear computational complexity and strong capabilities in capturing long-range dependencies. In the field of remote sensing, SSMs have gained popularity due to their effectiveness in addressing unique challenges such as dense visual predictions, multi-modal remote sensing data, and temporal remote sensing data, which have also yielded significant advancements in customized architectures. This paper presents a comprehensive review of SSM-based approaches in remote sensing, covering most of the relevant studies since SSMs were first introduced to the field. We offer a multi-dimensional analysis examining SSM applications in remote sensing tasks and discussing advancements in architecture design. This paper not only synthesizes the rapid progress in SSM-based research but also identifies key challenges and future opportunities. By providing a detailed perspective, this paper aims to serve as a foundational resource for remote sensing researchers, offering actionable insights to foster further advancements in this evolving domain. We will keep tracing related works at https://github.com/QinzheYang/Awesome-RS-State-Space-Model.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models
Authors:
Qinzhe Yang,
Keyan Chen,
Jia Xu,
Zhenwei Shi,
Zhengxia Zou
Abstract:
The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision models, particularly recent ViT-based foundation models in dense prediction tasks. Instance segmentation, a typical dense visual prediction task in the remote sensing field, faces similar challenges. In this paper, inspired by the recent advances of k…
▽ More
The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision models, particularly recent ViT-based foundation models in dense prediction tasks. Instance segmentation, a typical dense visual prediction task in the remote sensing field, faces similar challenges. In this paper, inspired by the recent advances of knowledge distillation in large language models, we introduce RS4D - a new remote sensing instance segmentation method with linear computational complexity, which addresses the inefficiency of long sequence modeling through distilled state space modeling (SSM). We propose an adaptive noise and masking knowledge distillation training method for pre-training lightweight SSM backbones, which effectively compresses knowledge from the vast self-attention space into a compact, dense linear state space. We also design a remote sensing image instance segmentation architecture based on this lightweight visual encoder, where we explore variants of three different backbones and two segmentation heads. Extensive experiments are conducted on multiple benchmark datasets, including SSDD, WHU, and NWPU. Compared to ViT-based approaches, our proposed SSM backbone achieves an 8x reduction in parameters and a 9x reduction in FLOPs while maintaining comparable or superior accuracy to both ViT- and CNN-based instance segmentation methods. The implementation codes have been publicly available at https://github.com/QinzheYang/RS4D.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection
Authors:
Qinzhe Yang,
Dongyu Wang,
Haohan Niu,
Jia Xu,
Zhenwei Shi,
Zhengxia Zou
Abstract:
Remote sensing object detection has advanced rapidly with the development of large-scale benchmarks and modern detection architectures. However, existing datasets and detectors remain fragmented. Most benchmarks focus on limited categories, fixed spatial resolutions, or a single sensor, while detectors still struggle to work across different sensors and categorical systems. In this paper, we intro…
▽ More
Remote sensing object detection has advanced rapidly with the development of large-scale benchmarks and modern detection architectures. However, existing datasets and detectors remain fragmented. Most benchmarks focus on limited categories, fixed spatial resolutions, or a single sensor, while detectors still struggle to work across different sensors and categorical systems. In this paper, we introduce LEVIRDet-159, the largest and most comprehensive remote sensing object detection dataset to date, with 159 categories, 2.56 million bounding boxes, and 700k fine-grained annotations under a multi-level taxonomy. In each key scale dimension, LEVIRDet-159 exceeds the corresponding largest existing remote sensing object detection dataset, containing approximately (7x) more images, (6x) more object instances, and (4x) more categories. Based on this dataset, we design LEVIRDetNet, a scale-hierarchy-aware detection foundation model for universal remote sensing object detection. LEVIRDetNet couples online visual Ground Sampling Distance (GSD) prediction, GSD-conditioned query modulation and allocation, and a hierarchy-aware detection head for mixed-granularity remote sensing supervision. Under stringent evaluation settings, LEVIRDetNet demonstrates strong cross-domain generalization. Even without target-domain training or fine-tuning, it achieves state-of-the-art detection performance on 9 external benchmarks, improving the strongest fully supervised competing methods by 5.02 mAP on average under each benchmark's primary metric. We hope this study will facilitate the development of strongly generalizable remote sensing object detection across diverse category systems, spatial resolutions, and sensor platforms. The dataset and trained models will be released at https://qinzheyang.github.io/LEVIRDet/, accompanying the final paper.
△ Less
Submitted 4 July, 2026; v1 submitted 23 June, 2026;
originally announced June 2026.
-
Isometric free finite group actions on non-positively curved 3-manifolds
Authors:
Zhengyu Zou
Abstract:
Let $M$ be a closed orientable $3$-manifold admitting a metric of nonpositive sectional curvature (an NPC metric), and let $G$ be a finite group acting freely on $M$ by orientation-preserving diffeomorphisms. Previous results showed that $M$ admits a $G$-invariant NPC metric except possibly when $M$ is a graph manifold. In this paper, we resolve the remaining case by proving that $M$ also admits a…
▽ More
Let $M$ be a closed orientable $3$-manifold admitting a metric of nonpositive sectional curvature (an NPC metric), and let $G$ be a finite group acting freely on $M$ by orientation-preserving diffeomorphisms. Previous results showed that $M$ admits a $G$-invariant NPC metric except possibly when $M$ is a graph manifold. In this paper, we resolve the remaining case by proving that $M$ also admits a $G$-invariant NPC metric when $M$ is a graph manifold. This result advances our understanding in dimension $3$ of the question posed by Schoen-Yau about Nielsen realization for NPC $3$-manifolds.
△ Less
Submitted 14 July, 2026; v1 submitted 20 June, 2026;
originally announced June 2026.
-
Training-free sparse attention based on cumulative energy filtering
Authors:
Chunlu Li,
Yixuan Pan,
Bai Du,
Zhenyuan Chen,
Yanzhao Li,
Hui Dong,
Hui Wang,
Zhiqiang Zou
Abstract:
Sparse attention accelerates Diffusion Transformers (DiTs) for video generation by computing only the important tokens while skipping the rest. The token selection strategy is key to balancing sparsity and accuracy. We formulate the token filtering process as a dual-goal optimization problem: maximizing sparsity and minimizing accuracy degradation. Existing algorithms cannot fulfill both objective…
▽ More
Sparse attention accelerates Diffusion Transformers (DiTs) for video generation by computing only the important tokens while skipping the rest. The token selection strategy is key to balancing sparsity and accuracy. We formulate the token filtering process as a dual-goal optimization problem: maximizing sparsity and minimizing accuracy degradation. Existing algorithms cannot fulfill both objectives simultaneously. For example, Top-p only considers the accuracy constraint, while Top-k maintains a fixed computational budget but loosens the accuracy constraint.
This paper demonstrates that maintaining a fixed recall rate is sufficient for ensuring accuracy, whereas a fixed threshold is suboptimal for reducing computational cost. Therefore, we propose a dynamic thresholding scheme to improve sparsity while maintaining the same level of accuracy. Furthermore, our algorithm is deeply integrated with Flash Attention (FA), eliminating the need for any additional masking computation overhead. Experimental results on Wan 2.2 validate that, compared to the BLASST algorithm which is also integrated with FA, our dynamic thresholding strategy enhances sparsity from 61.42\% to 82\% with a VBench metric drop of less than 5\%. This results in an approximate 15\% in attention computation and a $1.61\times$ increase in computational efficiency, which is 1.18x higher than that of BLASST.
△ Less
Submitted 15 June, 2026;
originally announced June 2026.
-
Graphical conditional generative modeling for digital twin modeling
Authors:
Zongren Zou,
Théo Bourdais,
Ricardo Baptista,
Houman Owhadi
Abstract:
Digital twin modeling, including control and data assimilation under model uncertainty, often faces an open-ended fidelity problem: adding variables, data streams, and time scales can indefinitely increase model complexity, ultimately producing systems that are difficult to maintain, validate, interpret, and use for stress or safety testing. As an alternative, one can seek parsimonious stochastic…
▽ More
Digital twin modeling, including control and data assimilation under model uncertainty, often faces an open-ended fidelity problem: adding variables, data streams, and time scales can indefinitely increase model complexity, ultimately producing systems that are difficult to maintain, validate, interpret, and use for stress or safety testing. As an alternative, one can seek parsimonious stochastic surrogate models built only on the variables needed to describe the relevant quantities of interest. We introduce a framework for discovering such variables from observational data by identifying which candidate inputs influence the full conditional law of a target quantity, rather than only its conditional mean. This distinction is essential in stochastic, coarse-grained, or partially observed systems, where dependencies may appear through changes in variability, tail behavior, multimodality, or uncertainty rather than through deterministic functional relationships. The framework couples conditional generative modeling, which learns the conditional distribution of the target given candidate inputs, with Gaussian-process-based analysis of variance (through kernel mode decomposition), which enables iterative pruning of non-influential inputs and interpretable structure discovery. In control settings, the resulting surrogate can be interpreted as a learned Markov decision process: the method identifies not only a transition model, but also the state, action, and memory variables needed to make the learned dynamics effectively Markovian. Across examples involving stochastic dynamical systems, missing variables, PDE control, reinforcement learning, and economic data, the discovered structures yield interpretable stochastic surrogates whose downstream performance is comparable to models trained on the full variable set.
△ Less
Submitted 15 June, 2026;
originally announced June 2026.
-
TrustedARI: Towards Trust-Native Agentic Routing Infrastructure for Agentic AI
Authors:
Qi Li,
Zhenhua Zou,
Shuo Li,
Mingwei Xu,
Zhuotao Liu
Abstract:
AI agents increasingly access external models, tools, and services through Agentic Routing Infrastructure (ARI) to manage the overhead of heterogeneous interfaces and fragmented subscriptions. Yet, the architecture of ARI introduces fundamental trust risks: it obtains plaintext access to agent queries and service responses, while leaving agents unable to verify that their queries are routed to int…
▽ More
AI agents increasingly access external models, tools, and services through Agentic Routing Infrastructure (ARI) to manage the overhead of heterogeneous interfaces and fragmented subscriptions. Yet, the architecture of ARI introduces fundamental trust risks: it obtains plaintext access to agent queries and service responses, while leaving agents unable to verify that their queries are routed to intended service providers or that requests and responses remain untampered. To address this problem, we present TrustedARI, the first trust-native agentic routing infrastructure for agentic AI. Architecturally, TrustedARI is built upon three core innovations: (i) an ARI-adapted three-party TLS handshake that enables the agent and ARI to jointly authenticate the service provider through role-specific distribution of TLS key materials; (ii) a privacy-preserving query-construction protocol that allows the agent and ARI to collaboratively construct well-formed queries without exposing their respective private inputs; and (iii) a verifiable billing protocol that supports fair usage-based settlement while preserving the integrity and confidentiality of service responses.
We implemented and extensively evaluated a prototype of TrustedARI to validate its performance. Experiments confirm that TrustedARI is highly efficient: our ARI-adapted handshake protocol reduces communication overhead by 39.34% compared to the existing three-party TLS handshake. Furthermore, the privacy-preserving query-construction protocol imposes negligible overhead-averaging 0.19 seconds in computation time and 0.58 MB in communication costs-while the verifiable billing protocol speeds up proof generation by 28.20x. Crucially, TrustedARI is readily deployable without any modification to the service providers.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
Hiding the Trees in the Forest: Building Network Covert Channels with Hash-Based Covert Carrier Filtering
Authors:
Zexiao Zou,
Zhiqiang Wang,
Baoxu Liu,
Yuyang Han,
Yan Zhang
Abstract:
As an effective anti-censorship mechanism, network covert channels can provide data privacy protection and ensure communication security. However, the covertness of existing network covert channels primarily depends on the secrecy of their covert algorithms. With the increasing depth of research in this field, the difficulty of breaking such algorithms has gradually decreased. Once the algorithm i…
▽ More
As an effective anti-censorship mechanism, network covert channels can provide data privacy protection and ensure communication security. However, the covertness of existing network covert channels primarily depends on the secrecy of their covert algorithms. With the increasing depth of research in this field, the difficulty of breaking such algorithms has gradually decreased. Once the algorithm is exposed, the network covert channel can be easily detected by adversaries. To address this issue, this paper proposes a covert carrier filtering strategy based on the hash. In this strategy, a key-dependent filtering rule is introduced during the construction of the network covert channel, enabling the communicating parties to randomly and dynamically filter a sparse subset from the carrier set as the covert carrier set. This strategy not only enhances the randomness of carrier selection but also tightly couples the covertness of the network covert channel with the security of the key. We employ machine learning-based traffic analysis methods to experimentally validate the strategy in two types of network covert channels: network storage and timing covert channels. The experimental results demonstrate that the proposed strategy significantly improves the detection resistance of network covert channels. When the filter key size exceeds six bits, the impact on the detection effect of the classifier becomes quite significant. Furthermore, the processing delay for a single packet is less than 8 $μs$, indicating the feasibility of deploying the proposed strategy in high-speed network environments.
△ Less
Submitted 9 June, 2026;
originally announced June 2026.
-
X-rays breaking out of pre-explosion ejecta mark a supernova's first light
Authors:
Weimin Yuan,
Qiu-Ju Huang,
Jin-Ping Zhu,
Yun-Wei Yu,
Dong Xu,
Chen Zhang,
Zhuo Li,
Yuan Liu,
Tao An,
Giulia Gianfagna,
Weikang Zheng,
Guowang Du,
Xing Liu,
Ji-An Jiang,
Johan P. U. Fynbo,
Alexei S. Pozanenko,
Junjie Jin,
Yi Yang,
Jinsong Deng,
Hui Sun,
Guang-Lei Wu,
Yu-Hao Zhang,
Bao Wang,
Yu Wang,
Xiangyu Wang
, et al. (108 additional authors not shown)
Abstract:
Massive stars die as core-collapse supernovae, whose optical light emerges days after the implosion. Theory predicts that the initial collapse-driven shock, upon breaking through the star and dense circumstellar medium, emits a brief thermal flash of soft X-rays and ultraviolet. Yet these elusive first signals have remained largely undetected, owing to limited wide-field soft X-ray monitoring. Her…
▽ More
Massive stars die as core-collapse supernovae, whose optical light emerges days after the implosion. Theory predicts that the initial collapse-driven shock, upon breaking through the star and dense circumstellar medium, emits a brief thermal flash of soft X-rays and ultraviolet. Yet these elusive first signals have remained largely undetected, owing to limited wide-field soft X-ray monitoring. Here we report the discovery of a soft X-ray flash, EP260321a, followed days later by a broad-lined supernova from an envelope-stripped progenitor. Its X-ray spectrum, best modeled with blackbody, establishes it as the long-sought archetypal shock breakout. The burst's duration and energetics place the breakout at a radius of 300 solar radii, tracing a dense surrounding shell and revealing abrupt mass ejection within the final month before collapse.
△ Less
Submitted 9 July, 2026; v1 submitted 8 June, 2026;
originally announced June 2026.
-
QueryAgent-R1: Bridging Query Generation and Product Retrieval for E-Commerce Query Recommendation
Authors:
Dike Sun,
Zheng Zou,
Jingtong Zang,
Qi Sun,
Huaipeng Zhaoand Tao Luo,
Xiaoyi Zeng
Abstract:
Query recommendation in e-commerce search aims to proactively suggest queries that match users' potential interests. However, existing methods mainly optimize query-level relevance, while neglecting whether the retrieved products align with users' downstream preferences. This mismatch often leads to high query click through rates (CTR) but low product conversion rates (CVR). To bridge this gap, we…
▽ More
Query recommendation in e-commerce search aims to proactively suggest queries that match users' potential interests. However, existing methods mainly optimize query-level relevance, while neglecting whether the retrieved products align with users' downstream preferences. This mismatch often leads to high query click through rates (CTR) but low product conversion rates (CVR). To bridge this gap, we propose QueryAgent-R1, a memory-augmented agentic framework that improves end-to-end alignment via chain-of-retrieval optimization. Our QueryAgent-R1 grounds query generation in real inventory retrieval, allowing the agent to validate and refine queries based on retrieved products. We also design a consistency reward in the agentic reinforcement learning (RL) process to jointly optimize query relevance and downstream engagement. In addition, we construct a memory abstraction module for efficient user profiling. To support offline evaluation, we construct two datasets based on both proprietary industrial data and public datasets, on which QueryAgent-R1 consistently outperforms strong baselines. Moreover, on a large scale production platform, QueryAgent-R1 improves Query CTR by 2.9% and guided CVR by 3.1% in online A/B tests.
△ Less
Submitted 3 June, 2026;
originally announced June 2026.
-
Macroscopic Spin GHZ States with a Levitated Ferromagnet
Authors:
Xueqi Ni,
Zhixing Zou,
Ping Koy Lam,
Tao Wang,
Jiangbin Gong
Abstract:
The generation of macroscopic quantum states can drive both fundamental physics and quantum technologies. This work proposes a top-down approach to the generation of macroscopic spin GHZ states using a levitated ferromagnet, where a strong locking between the collective spin and the lattice rotation enables mechanical control of the collective spin. We quantify the metrological advantage of the re…
▽ More
The generation of macroscopic quantum states can drive both fundamental physics and quantum technologies. This work proposes a top-down approach to the generation of macroscopic spin GHZ states using a levitated ferromagnet, where a strong locking between the collective spin and the lattice rotation enables mechanical control of the collective spin. We quantify the metrological advantage of the resulting macrospin superposition state by showing that Heisenberg scaling of the quantum Fisher information is achievable. Roles of symmetry and geometry are analyzed in terms of decoherence due to gas collisions, identifying accessible conditions for experimental realization. The usefulness of a macrospin superposition state of a levitated cylindrical ferromagnet in testing spin-dependent wavefunction collapse models is also discussed.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
Probing the imaginary parts and their $q^2$ dependences for the tau $g-2$ and EDM
Authors:
Xin-Yu Du,
Xiao-Gang He,
Zhong-Lv Huang,
Chia-Wei Liu,
Zi-Yue Zou
Abstract:
The $τ$ anomalous magnetic dipole moment (MDM) $a_τ= (g-2)_τ/2$ and electric dipole moment (EDM) $d_τ$, are precision probes of electroweak dynamics and possible new physics sources, yet both remain weakly constrained experimentally. Treated as generalized form factors, these quantities exhibit a generic $q^2$ dependence for an off-shell interacting photon. For timelike momentum transfer above the…
▽ More
The $τ$ anomalous magnetic dipole moment (MDM) $a_τ= (g-2)_τ/2$ and electric dipole moment (EDM) $d_τ$, are precision probes of electroweak dynamics and possible new physics sources, yet both remain weakly constrained experimentally. Treated as generalized form factors, these quantities exhibit a generic $q^2$ dependence for an off-shell interacting photon. For timelike momentum transfer above the $τ^+τ^-$ threshold, $q^2 = s > 4m_τ^2$, the form factors can acquire absorptive imaginary parts. We investigate how such a $q^2$ dependence and the associated imaginary parts are generated from two complementary perspectives: the model-independent Standard Model Effective Field Theory (SMEFT) and a UV-complete Two-Higgs-Doublet Model (2HDM). The effective framework reveals the intimate correlation between $a_τ$ and $d_τ$. New CP-violating interactions which generate a non-zero $d_τ$, can also generically have non-zero contributions to $a_τ$, thereby deeply linking their phenomenological studies. Within the 2HDM, we demonstrate that sizable imaginary parts and significant $q^2$ running can be generated at levels accessible by $e^+e^-$ colliders. Motivated by these features, we propose experimental methods to extract both the real and imaginary components of the dipole form factors. Utilizing these techniques, we show that Belle II and the Super Tau-Charm Facility (STCF) can improve current bounds on $a_τ$ by more than one order of magnitude. Finally, we highlight that combining measurements across the distinct center-of-mass energies of Belle II and STCF provides a unique, previously unexplored avenue to explicitly obtain information about the $q^2$ evolution of these dipole form factors.
△ Less
Submitted 10 June, 2026; v1 submitted 31 May, 2026;
originally announced June 2026.
-
FedQHD: Closed-Form Function-Space Federated Reinforcement Learning
Authors:
Yuchen Hou,
Yongshan Chen,
Zhuowen Zou,
Calvin Yeung,
Mohsen Imani,
Tian Lan,
Mahdi Imani
Abstract:
Federated reinforcement learning enables decentralized agents to collaboratively improve policies or value estimates without exchanging raw trajectories. However, FedAvg-style parameter averaging is not function-space consistent: when clients use heterogeneous encoders or even identical nonlinear networks, averaged parameters need not correspond to the weighted average of client value functions in…
▽ More
Federated reinforcement learning enables decentralized agents to collaboratively improve policies or value estimates without exchanging raw trajectories. However, FedAvg-style parameter averaging is not function-space consistent: when clients use heterogeneous encoders or even identical nonlinear networks, averaged parameters need not correspond to the weighted average of client value functions in any common function space. We propose FedQHD, a federated Q-learning method using hyperdimensional (random-feature) state encoders with a linear readout, so that Q-functions are nonlinear in state yet linear in trainable parameters. This linear structure enables closed-form aggregation. With a shared encoder, the function-space consensus update coincides exactly with weighted averaging of local readout matrices. With heterogeneous encoders, the server constructs a global teacher by averaging client Q-values on a shared anchor-state set, and each client compiles this teacher into its local representation via a single ridge projection. We formalize the federation gap -- the error incurred when compiling a federated teacher into a heterogeneous client representation -- relative to a client-specific oracle projection. We show that this gap decomposes into subspace misalignment, anchor-set conditioning, and regularization bias. We further identify the anchor-to-dimension ratio $m \geq D_i$ as the well-conditioned regime in which the gap reduces to a multiple of the encoder heterogeneity floor. On four continuous-state, discrete-action control benchmarks, FedQHD matches or outperforms FedAvg-style baselines and distillation-based alternatives while requiring substantially less computation, and the empirical dependence of the federation gap on encoder dimension matches our theoretical analysis.
△ Less
Submitted 27 May, 2026;
originally announced May 2026.
-
SAFEVPR: Patch-Based Conformal Verification for Safe Cross-Condition Sequence Visual Place Recognition
Authors:
Ha Sier,
Jiaqiang Zhang,
Zhuo Zou,
Xianjia Yu,
Tomi Westerlund
Abstract:
Sequence-based visual place recognition (VPR) for SLAM and robot relocalization must decide whether the retrieved top-1 candidate is safe to accept. Conformal prediction is a natural framework for this accept/reject decision, but its finite-sample guarantees rely on exchangeability between calibration and deployment (test) data, which is violated under cross-condition deployment. We introduce SAFE…
▽ More
Sequence-based visual place recognition (VPR) for SLAM and robot relocalization must decide whether the retrieved top-1 candidate is safe to accept. Conformal prediction is a natural framework for this accept/reject decision, but its finite-sample guarantees rely on exchangeability between calibration and deployment (test) data, which is violated under cross-condition deployment. We introduce SAFEVPR, a non-trainable verification-and-calibration pipeline for safe cross-condition sequence VPR. SAFEVPR replaces the standard backbone cosine similarity with a mutual-nearest-neighbour (MNN) patch-matching score computed from frozen DINOv2 ViT features, and replaces flat Learn-Then-Test calibration with Mondrian conformal LTT, fitting separate Bonferroni-corrected thresholds across score bins. Under exchangeability, these thresholds would provide finite-sample false-discovery-rate (FDR) control; under condition shift, we evaluate empirical validity per deployment. Across 23 cross-condition setups from Oxford RobotCar, NCLT, and St Lucia datasets, using three frozen VPR backbones, SAFEVPR is empirically valid on 23/23 setups at target FDR alpha = 0.10, achieving mean accepted FDR 0.014 and mean true-positive rate (TPR) 0.75. The results show that raw discrimination alone is not sufficient for conformal validity: AnyLoc-VLAD and Super-Point+LightGlue reach comparable area under the receiver operating characteristic curve (AUROC) but fail more setups under the same calibration. On textureless repetitive scenery, SAFEVPR safely abstains rather than accepting unreliable matches. Code is available at https://github.com/Hasar12139/SafeVPR.
△ Less
Submitted 27 May, 2026;
originally announced May 2026.
-
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
Authors:
Calvin Yeung,
Prathyush Poduval,
Ali Zakeri,
Zhuowen Zou,
Mohsen Imani
Abstract:
Text-to-image diffusion models generate images through an iterative denoising process, so internal neural layers produce trajectories of activations rather than single static representations. Sparse autoencoders (SAEs) have recently been used to decompose diffusion activations into interpretable feature directions, but most approaches analyze activations at individual timesteps or condition on tim…
▽ More
Text-to-image diffusion models generate images through an iterative denoising process, so internal neural layers produce trajectories of activations rather than single static representations. Sparse autoencoders (SAEs) have recently been used to decompose diffusion activations into interpretable feature directions, but most approaches analyze activations at individual timesteps or condition on time rather than learning directly from full activation trajectories. In this work, we introduce residualized temporal SAEs for diffusion activation trajectories. We collect activations across denoising time, fit linear predictors between neighboring timesteps, and represent each trajectory using an initial activation together with residual components not explained by these linear dynamics. Training an SAE on this residualized representation encourages sparse latents to capture structure beyond what is linearly predictable. The residualized decoder directions can be mapped back into activation space, allowing each latent to be analyzed as a feature trajectory over denoising time. Through reconstruction and ablation studies, spatiotemporal feature analysis, and qualitative steering experiments on Stable Diffusion~1.5, we show that residualized temporal SAEs provide a useful framework for studying temporally structured diffusion activations.
△ Less
Submitted 26 May, 2026;
originally announced May 2026.
-
Constraining the Supernova Remnant Environment of FRB 190520B with Dispersion Measure and Scattering Timescale
Authors:
Jia-Peng Wei,
Chen Deng,
Gwenael Giacinti,
Ze-Cheng Zou,
Chen-Ran Hu,
Yong-Feng Huang,
Jin-Jun Geng
Abstract:
FRB 190520B is a repeating fast radio burst source whose large dispersion measure (DM) and temporal broadening suggest a dense and evolving local environment. In this work, we test the possibility that FRB 190520B originates from the core-collapse of a massive star so that its central engine is embedded in a supernova remnant (SNR) expanding into a wind environment, whose evolution is described by…
▽ More
FRB 190520B is a repeating fast radio burst source whose large dispersion measure (DM) and temporal broadening suggest a dense and evolving local environment. In this work, we test the possibility that FRB 190520B originates from the core-collapse of a massive star so that its central engine is embedded in a supernova remnant (SNR) expanding into a wind environment, whose evolution is described by the self-similar solution. We use the observed DM and scattering timescale of FRB 190520B to constrain the physical parameters of its surrounding SNR and host-galaxy DM. Twenty typical cases are considered, arising from four ejecta profiles and five scattering prescriptions. It is found that only 6 cases are retained and provide acceptable fits. All retained cases have a shallow ejecta profile and a young source age of $t_0=79.8$--$169.8~{\rm yr}$. The ejecta mass is inferred to be large for all six cases, while the kinetic energy and mass-loss rate span a wide range. The secular DM evolution is reproduced better than the detailed scattering evolution. The up-drift behavior of the scattering residual suggests an additional component or more complicated structures inside the SNR. All retained cases are self-consistent within the adopted scattering theory and the circum-burst medium becomes transparent for GHz bursts before the inferred source ages.
△ Less
Submitted 26 May, 2026;
originally announced May 2026.
-
Global Average Treatment Effects for Individualized Randomization Experiments with Aggregate Data
Authors:
Shuguang Yu,
Ting Li,
Yuchen Lu,
Chengchun Shi,
Fan Zhou,
Zhichao Zou,
Peng Zhen,
Hongtu Zhu
Abstract:
Individualized randomized experiments are central to online platforms for optimizing personalized decisions in complex environments. In two-sided markets, however, standard treatment effect estimation is often invalid due to strong temporal and cross-unit interference, a challenge compounded when only aggregated data are available because of privacy or system constraints. To address these issues,…
▽ More
Individualized randomized experiments are central to online platforms for optimizing personalized decisions in complex environments. In two-sided markets, however, standard treatment effect estimation is often invalid due to strong temporal and cross-unit interference, a challenge compounded when only aggregated data are available because of privacy or system constraints. To address these issues, we identify the Global Average Treatment Effect (GATE) using only group-level data from treatment and control groups. We first establish identification conditions based on aggregated observations, and then propose the Individualized Randomized Experiment Varying Coefficient Decision Process (IRE-VCDP) model, which accounts for interference through supply-demand dynamics. Building on this framework, we develop a complete procedure for estimation and statistical inference of the GATE, along with theoretical guarantees for the proposed test. Extensive simulations and real-world experiments using data from a leading ridesharing platform demonstrate the effectiveness of our approach.
△ Less
Submitted 26 May, 2026;
originally announced May 2026.
-
SFR-Net: Learning Scale-Frustum Representations for Ultra-Wide Area Remote Sensing Image Segmentation
Authors:
Chuyu Zhong,
Keyan Chen,
Qinzhe Yang,
Bowen Chen,
Zhengxia Zou,
Zhenwei Shi
Abstract:
Pixel count and geographical coverage are two key characteristics of remote sensing images. Existing remote sensing image segmentation methods typically focus on images with either a small pixel count or a large pixel count but limited geographical coverage. In this paper, we introduce a novel segmentation task targeting ultra-wide area (UWA) remote sensing images, characterized by both a large pi…
▽ More
Pixel count and geographical coverage are two key characteristics of remote sensing images. Existing remote sensing image segmentation methods typically focus on images with either a small pixel count or a large pixel count but limited geographical coverage. In this paper, we introduce a novel segmentation task targeting ultra-wide area (UWA) remote sensing images, characterized by both a large pixel count and extremely wide geographical coverage. The core challenges of UWA segmentation lie in simultaneously handling ground objects with significantly varying scales and maintaining long-range contextual semantic continuity. To address these challenges, we propose the Scale-Frustum Representation Network (SFR-Net). Inspired by the viewing frustums of remote sensing images captured from different altitudes, we construct scale-frustum representations, enabling unified modeling of ground objects and contextual features at different scales. Furthermore, we design a cascaded cross-scale fusion mechanism to effectively integrate these representations, enhancing local semantic understanding while ensuring long-range contextual continuity. Experimental results on GID and FBPS demonstrate that SFR-Net achieves state-of-the-art performance, improving mIoU by 1.72% and 4.29%, respectively, over the strongest competing methods. In addition, the proposed scale-frustum representations can be integrated into generic segmentation networks to improve both segmentation accuracy and convergence speed. The implementation code will be publicly available at https://github.com/ChuyuZhong/SFR-Net.
△ Less
Submitted 25 May, 2026;
originally announced May 2026.
-
Non-equilibrium pathway to mesoscale ordering in ethanol-water binary liquid
Authors:
Xinyue Jiang,
Yating Shang,
Jianhui Li,
Zhaoyong Zou,
Yanxia Zuo,
Yuqun Xie
Abstract:
Ethanol-water mixtures are a classic example of thermodynamic non-ideality, yet the structural origin of their pronounced anomalies, such as volume contraction and a large negative excess entropy, has remained a long-standing puzzle. Here, we demonstrate these anomalies are not equilibrium properties but calorimetric fingerprint of an arrested phase transition. By imposing periodic thermal oscilla…
▽ More
Ethanol-water mixtures are a classic example of thermodynamic non-ideality, yet the structural origin of their pronounced anomalies, such as volume contraction and a large negative excess entropy, has remained a long-standing puzzle. Here, we demonstrate these anomalies are not equilibrium properties but calorimetric fingerprint of an arrested phase transition. By imposing periodic thermal oscillations, we drive a 50% (v/v) ethanol-water system along a complete hierarchical self-assembly pathway that progressed from ethanol clusters to water-containing droplets, then to acicular flakes, and finally to micron-scale ordered ethanol aggregates. Fluorescence spectroscopy, two-dimensional correlation analysis and nuclear magnetic resonance revealed the underlying non-equilibrium molecular mechanism: a periodic perturbation of the water-dominated hydrogen-bond network initiates a ethanol-water coexistence intermediate, ultimately leading to the stable ordered assembly of an ethanol-rich phase. Our finding demonstrated that periodic physical perturbations capable drive spontaneous ordering across multiple length scales in a simple binary mixture, providing a kinetic perspective on the structural origin of solution non-ideality, and carry general implications for self-assembly strategies in soft matter.
△ Less
Submitted 25 May, 2026;
originally announced May 2026.
-
ReMATF: Recurrent Motion-Adaptive Multi-scale Turbulence Mitigation for Dynamic Scenes
Authors:
Zhiming Liu,
Zhicheng Zou,
Nantheera Anantrasirichai
Abstract:
Atmospheric turbulence severely degrades video quality by introducing distortions such as geometric warping, blur, and temporal flickering, posing significant challenges to both visual clarity and temporal consistency. Current state-of-the-art methods are based on transformer, 3D architectures and require multi-frame input, but their large computational cost and memory usage limit real-time deploy…
▽ More
Atmospheric turbulence severely degrades video quality by introducing distortions such as geometric warping, blur, and temporal flickering, posing significant challenges to both visual clarity and temporal consistency. Current state-of-the-art methods are based on transformer, 3D architectures and require multi-frame input, but their large computational cost and memory usage limit real-time deployment, especially in resource-constrained scenarios. In this work, we propose ReMATF, a lightweight recurrent framework that restores videos using only two frames at a time while preserving spatial detail and temporal stability. ReMATF combines a multi-scale encoder-decoder with temporal warping and a motion-adaptive temporal fusion module that performs per-pixel fusion between the warped previous output and the current prediction to enhance coherence without enlarging the temporal window. This design reduces flicker, sharpens details, and remains efficient. Experiments on synthetic and real turbulence datasets show consistent improvements in PSNR/SSIM and perceptual quality (LPIPS), along with substantially faster inference than multi-frame transformer baselines, making ReMATF suitable turbulence mitigation in resource-constrained scenarios.
△ Less
Submitted 20 May, 2026;
originally announced May 2026.