-
Carbon reductions through optimized solar heat gain glass properties considering future climate and grid emissions: case study of Chicago's residential buildings
Authors:
Yiwei Lyu,
Jialiang Xiang,
Holly Samuelson
Abstract:
Existing resources leave confusion over the benefits of high versus low Solar Heat Gain Coefficient (SHGC) windows for energy performance in residential buildings retrofits in cold climates. Additionally, few studies have considered the impact of expected future climate conditions and time-variable grid emission rates on energy-related metrics. Utilizing the ResStock, residential building stock mo…
▽ More
Existing resources leave confusion over the benefits of high versus low Solar Heat Gain Coefficient (SHGC) windows for energy performance in residential buildings retrofits in cold climates. Additionally, few studies have considered the impact of expected future climate conditions and time-variable grid emission rates on energy-related metrics. Utilizing the ResStock, residential building stock models from the National Renewable Energy Laboratory (NREL), this study investigates retrofits increasing the SHGC of windows in Chicago, a cold US city. The results indicate that increasing window SHGC increases summer cooling needs; however, in most cases, this effect is more than offset by reduced winter heating needs. This balance is particularly beneficial considering the state's expected long-run marginal carbon emission rates. The study also examines the combined effects of high SHGC with improved window insulation values, demonstrating that such strategic window retrofits not only enhance overall building energy performance but also contribute to greater emission reductions. On average, the current Chicago residences (n = 4,826) save 4.6 % on heating and cooling carbon emissions by increasing the SHGC of the windows. If we assume that those homes are upgraded with heat pumps (electrification), a popular retrofit that reduces heating-related carbon emissions in particular, the increased window SHGC saves 2.5 % of long-run marginal carbon emissions. These results provide new insight into the carbon benefits of higher SHGC replacement windows in a cold climate. The benefits are significant, even considering future trends of a warming climate, higher demand grid emissions, and building electrification.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Limiting absorption principle for time-harmonic elastic scattering of plane waves from diffraction gratings
Authors:
Jianli Xiang,
Guanghui Hu
Abstract:
We establish the limiting absorption principle for time-harmonic elastic scattering of plane waves by a periodic rigid diffraction grating. By perturbing the frequency with a small positive imaginary part, we regularize the ill-posed problem at propagative wavenumbers (that is, when uniqueness fails under the classical Rayleigh expansion condition) and characterize the limiting solution via a sing…
▽ More
We establish the limiting absorption principle for time-harmonic elastic scattering of plane waves by a periodic rigid diffraction grating. By perturbing the frequency with a small positive imaginary part, we regularize the ill-posed problem at propagative wavenumbers (that is, when uniqueness fails under the classical Rayleigh expansion condition) and characterize the limiting solution via a singular perturbation result from functional analysis. The limiting solution satisfies the original scattering problem together with an additional constraint that ensures uniqueness. Both incident pressure and shear waves are considered, and the same constraint condition is obtained in both cases. The results provide a rigorous selection mechanism for physically admissible solutions at resonance frequencies. Our framework extends naturally to the Neumann (cavity) boundary condition as well as other transmission conditions, in particular when guided waves exist in periodic structures.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Direct Sum and Direct Product Decompositions of Multivariate Functions
Authors:
Hua-Lin Huang,
Yiming Liu,
Jianhua Xiang,
Yu Ye
Abstract:
This paper addresses the problem of whether or not a vector-valued multivariate functions can be expressed as a sum or a product of vector-valued functions in disjoint sets of variables through a proper invertible linear change of variables. The crux is an invariant algebra, the so-called center, that we introduce for a set of multivariate functions with second order partial derivatives. We thus p…
▽ More
This paper addresses the problem of whether or not a vector-valued multivariate functions can be expressed as a sum or a product of vector-valued functions in disjoint sets of variables through a proper invertible linear change of variables. The crux is an invariant algebra, the so-called center, that we introduce for a set of multivariate functions with second order partial derivatives. We thus provide simple criteria and algorithms for simultaneous additive and multiplicative decompositions of any set of multivariate functions with minor analytic conditions. This is applied to the factorization problem of multivariate homogeneous polynomials, in particular those that are products of linear forms.
△ Less
Submitted 17 July, 2026;
originally announced August 2026.
-
What the Detector Can See: Evaluating CPS Anomaly Detectors Independently of the Decision Rule
Authors:
Peiran Shi,
Jian Xiang,
Xiang Zhang,
Chenglong Fu
Abstract:
Anomaly detectors are often the last line of defense for cyber-physical systems (CPS). But detectors built in very different ways, from deep neural networks to invariant templates, are usually compared using precision, recall, or F1 at a single operating point. These scores mix two separate things: how well the detector represents the physical process, and how well its alarm threshold is set. We t…
▽ More
Anomaly detectors are often the last line of defense for cyber-physical systems (CPS). But detectors built in very different ways, from deep neural networks to invariant templates, are usually compared using precision, recall, or F1 at a single operating point. These scores mix two separate things: how well the detector represents the physical process, and how well its alarm threshold is set. We therefore treat a CPS anomaly detector as a two-stage pipeline: Stage 1 maps observations to residuals, and Stage 2 maps residuals to alarms. Instead of scoring only the final alarms, we evaluate Stage 1 directly using normalized residual energy, which has an exact connection to the Kullback-Leibler divergence from the trained-normal reference distribution. Because it does not depend on a specific alarm rule, it can separately measure attack separation, stability across the train-test gap, and the compactness with which a detector encodes the plant.
Without any per-detector tuning, we apply this evaluation to five detectors -- GDN, FuSAGNet, TranAD, NSIBF, and GeCo -- across three CPS benchmarks: SWaT, WADI, and HAI. Although the detectors have similar ROC-AUC values on SWaT, their performance differs by more than an order of magnitude at a common false-alarm rate. Rankings also change across testbeds: TranAD ranks first on HAI but last on SWaT, while NSIBF ranks first on WADI but last on HAI. On WADI, localized attacks can evade detectors that pool evidence across all channels, helping explain why NSIBF outperforms methods that do well on other benchmarks. These results show that detection failure can come from different sources: a weak representation, poor threshold calibration, or an attack with little physical effect. A decision-rule-free analysis helps separate these causes.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Authors:
Dongfang Li,
Xiaodong Luo,
Ruoyu Sun,
Xuhui Chen,
Linyuan Qiu,
Jian Meng,
Zhengxuan Lu,
Yiting Wang,
Yucheng Xie,
Tao Guo,
Tianxiang Fang,
Jing Li,
Sihang Chen,
Shihao Hong,
Chang Liu,
Weihua Dai,
Zirong Zeng,
Ziwei Zhu,
Zhuohan Wang,
Zhengjun Yue,
Igor Vasilyev,
Min Liu,
Weijian Sun,
Xin Chen,
Yingmeng Gao
, et al. (40 additional authors not shown)
Abstract:
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on…
▽ More
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
△ Less
Submitted 19 August, 2026; v1 submitted 22 July, 2026;
originally announced July 2026.
-
High-Energy Microresonator Soliton Generation
Authors:
Zhenhua Guo,
Sushant Kumar,
Xue Dong,
Yi Zhang,
Jiewei Xiang,
Arunima Nauriyal,
Junchi Zhang,
Elias Veilleux,
Jaime Cardenas,
William H. Renninger
Abstract:
Kerr resonators generate stable frequency combs in a compact platform with applications in coherent communications, sensing, quantum information processing, and astrophysics. Ultrashort pulses can be generated, moreover, with a wavelength and repetition rate flexibility inaccessible by traditional mode-locked lasers, which is desirable for high peak-power applications including in biomedicine and…
▽ More
Kerr resonators generate stable frequency combs in a compact platform with applications in coherent communications, sensing, quantum information processing, and astrophysics. Ultrashort pulses can be generated, moreover, with a wavelength and repetition rate flexibility inaccessible by traditional mode-locked lasers, which is desirable for high peak-power applications including in biomedicine and materials processing. However, for these applications, single-pulse energies will need to be significantly improved beyond the fJ level characteristic of sources today. Here we describe and demonstrate a simple approach for increasing the pulse energy in anomalous dispersion Kerr microresonators. Through scaling laws based on a mean-field cavity model and supported by experimentally-accurate numerical simulations, we show that pulse energy scales strongly with output coupling if supported by sufficient drive power. In strongly over-coupled cavities, femtosecond pulses can be stabilized with pulse energy beyond the pJ level. With this theoretical framework, by pumping a 12 GHz Si3N4 cavity with 30% output coupling with 2.6-ps time-lens generated pulses, we observe stable 74-fs pulses with a record output pulse energy of 6 pJ. High energy Kerr resonators are anticipated to improve performance for current applications and compliment mode locked lasers for high peak power applications.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation
Authors:
Ruicheng Li,
Qixiu Li,
Ruichun Ma,
Yu Deng,
Lin Luo,
Zhiying Du,
Jianfeng Xiang,
Huizhi Liang,
Ruicheng Wang,
Jiaolong Yang,
Baining Guo
Abstract:
Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by conditioning on past images or language summaries. Vision-based memory approaches address this by conditioning on sampled past image frames, but they are computationally expensive and fundamentally limited when temporal event…
▽ More
Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by conditioning on past images or language summaries. Vision-based memory approaches address this by conditioning on sampled past image frames, but they are computationally expensive and fundamentally limited when temporal events are visually ambiguous, e.g., pushing a button multiple times with small movements. We propose FM-VLA, a VLA model with force-based memory, enabling temporal context reasoning for non-Markovian, contact-rich manipulation. We encode force histories into compact force memory tokens with a variational autoencoder (VAE) pretrained with force time series reconstruction. By projecting force latent representations and short state history as additional conditioning tokens to the action expert module, we enable VLAs to leverage accumulated contact event history to guide manipulation. We evaluate FM-VLA on three memory-dependent tasks, including finding a hidden block, pressing a button, and wiping a dish for a specific number of times. Our lightweight force memory achieves over 80% success rate with minimal inference overhead, significantly outperforming baseline approaches. Project page: https://qft-333.github.io/FM-VLA-Page/
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement
Authors:
Lingyu Kong,
Ruicheng Li,
Ruicheng Wang,
Sicheng Xu,
Chengtang Yao,
Jianfeng Xiang,
Jiaolong Yang
Abstract:
Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structure, especially in fine details, like thin structures and small objects. We attribute this limitation to an architectural mismatch: most current models decode 3D geometry within a 2D parameterization, where feature intera…
▽ More
Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structure, especially in fine details, like thin structures and small objects. We attribute this limitation to an architectural mismatch: most current models decode 3D geometry within a 2D parameterization, where feature interactions are governed by image-plane proximity rather than true 3D spatial relationships. This inadvertently mixes features from geometrically distant surfaces, resulting in over-smoothed geometry particularly around thin or elongated structure. In this paper, we propose MoGe-3, a fine-detail monocular geometry estimation model with Self-Guided Sparse 3D Refinement (SSR) that lifts monocular geometry modeling from 2D image space to 3D space for high-fidelity metric-scale point maps. MoGe-3 lifts the coarse point map from a foundation base model onto a sparse voxel shell and refines it via SSR. The SSR employs sparse convolutions that aggregate features based on 3D spatial locality, avoiding feature mixing across depth discontinuities. Extensive experiments on diverse datasets demonstrate that MoGe-3 significantly outperforms existing approaches in recovering fine detailed 3D geometry across both quantitative metrics and qualitative visualizations. Project page: https://qft-333.github.io/moge3page/
△ Less
Submitted 21 July, 2026; v1 submitted 20 July, 2026;
originally announced July 2026.
-
From Classification to Consistent Templates: Multiple Permuted-Label Classifier Encoding for Biometric Template Protection
Authors:
Baogang Song,
Zhongshu Zhao,
Qianrong Zheng,
Jianwen Xiang,
Dongdong Zhao
Abstract:
Biometric template protection (BTP) must secure stored templates while tolerating intra-class variations. Existing methods rely on protected-domain similarity matching, error correction, or predefined-template mappings, potentially retaining exploitable similarity structures, introducing helper-data risks, depending on artificial targets, or coupling protection to specific modalities. Storing only…
▽ More
Biometric template protection (BTP) must secure stored templates while tolerating intra-class variations. Existing methods rely on protected-domain similarity matching, error correction, or predefined-template mappings, potentially retaining exploitable similarity structures, introducing helper-data risks, depending on artificial targets, or coupling protection to specific modalities. Storing only cryptographic hash digests eliminates directly comparable representations and conceals pre-hash templates, but hash-based exact-match verification requires genuine samples to generate identical intermediate templates before hashing. Identity classification is naturally suited to this requirement because it maps variable biometric samples to stable and discriminative identity-level outputs. Based on this insight, we propose Multiple Permuted-Label Classifier Encoding (MPLCE). Through classifier-specific label permutations, MPLCE assigns each identity different labels across multiple classifiers. The predicted labels are encoded and concatenated to form an intermediate template, preventing repeated encodings of a single identity label and enlarging the effective candidate space while preserving classification consistency. The template is randomized with an application-specific XOR string and cryptographically hashed, enabling exact-match verification without error correction codes or biometric-dependent helper data. Using modality-specific classifiers, MPLCE retains the same template generation and protection procedure across modalities. On four face and two iris datasets, MPLCE achieves competitive performance, including a GAR of 98.61\% at a FAR of 5.51\(\times\)10\textsuperscript{-5}\% on YTF and a GAR of 99.10\% at a FAR of 0.00\% on CASIA-Iris-Lamp. Security analyses and attack evaluations support its irreversibility, revocability, and unlinkability under the threat model.
△ Less
Submitted 15 July, 2026;
originally announced July 2026.
-
Giant magnetocaloric effect at low fields in triangular-lattice NdMgAl$_{11}$O$_{19}$
Authors:
Yantao Cao,
He Sun,
Zhendong Fu,
Zhaoming Tian,
Huiqian Luo,
Junsen Xiang,
Peijie Sun,
Jinkui Zhao,
Hanjie Guo
Abstract:
Magnetic refrigeration in the sub-Kelvin regime requires refrigerant materials to retain a large magnetic entropy at low temperatures by suppressing magnetic ordering. Quantum spin liquids (QSLs), which evade long-range magnetic ordering while retaining strong quantum fluctuations to the lowest temperatures, therefore provide a promising platform for realizing high-performance magnetic refrigerant…
▽ More
Magnetic refrigeration in the sub-Kelvin regime requires refrigerant materials to retain a large magnetic entropy at low temperatures by suppressing magnetic ordering. Quantum spin liquids (QSLs), which evade long-range magnetic ordering while retaining strong quantum fluctuations to the lowest temperatures, therefore provide a promising platform for realizing high-performance magnetic refrigerants. Here, we investigate the magnetic ground state and the magnetocaloric effect of the hexaaluminate, NdMgAl$_{11}$O$_{19}$, in which the Nd$^{3+}$ ions form a network of triangular lattices. Magnetic susceptibility and specific heat measurements indicate a magnetically dynamic state down to 50~mK, consistent with a QSL state. Specific heat measurements further reveal substantial magnetic entropy retained below 50~mK. Quasi-adiabatic demagnetization measurements demonstrate a superior cooling performance of NdMgAl$_{11}$O$_{19}$, which can be cooled to 113~mK from 1.9~K by only a small magnetic field change of 2~T. The outstanding refrigeration performance is attributed to the persistent spin fluctuations associated with the QSL-like ground state, together with a large effective \textit{g} factor and the smallness of the exchange interactions along the easy-axis direction. This study demonstrates that frustration, combined with strong spin-orbit coupling and crystal-electric-field effect in the rare earth magnets provides a promising design principle for next-generation cryogenic magnetic refrigerants.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Knowledge-Constrained Shape Optimization with a Mixture-of-Experts Neural Operator for High-Confidence Design
Authors:
Wenhao Fan,
Yuanwei Bin,
Jianghan Gu,
Wenfa Luo,
Jiao Xiang,
Yuntian Chen,
Shiyi Chen
Abstract:
Engineering shape optimization faces challenges in both expert-dependent problem setup and surrogate-model reliability. In practical aerodynamic design, optimization settings such as editable regions, deformation ranges, and design-preservation constraints are typically specified manually by experienced engineers, while surrogate-based optimization may become unreliable for heterogeneous geometry…
▽ More
Engineering shape optimization faces challenges in both expert-dependent problem setup and surrogate-model reliability. In practical aerodynamic design, optimization settings such as editable regions, deformation ranges, and design-preservation constraints are typically specified manually by experienced engineers, while surrogate-based optimization may become unreliable for heterogeneous geometry databases and out-of-distribution designs. To address these challenges, we propose a knowledge-constrained shape-optimization framework that translates knowledge-based constraints and user intent into quantifiable parameters of DFFD-based deformation operators, enabling engineering-aware and controllable constrained optimization. We further develop a Mixture-of-Experts Neural Operator (MoE-NO) to improve drag prediction and trend consistency over heterogeneous aerodynamic datasets. Based on the MoE-NO encoder and Mahalanobis distance, an uncertainty-estimation strategy is introduced to detect out-of-distribution geometries and selectively trigger physics-solver feedback for local sample enrichment. Experiments on in-house MPV, SUV, and Sedan datasets show that MoE-NO achieves a test-set MAPE of $1.16\%$ and a trend-prediction accuracy of $94.34\%$, outperforming the best baseline results of $1.52\%$ and $90.34\%$, respectively. Vehicle shape-optimization experiments further yield CFD-validated drag coefficient reductions of approximately $4\%$ to $10\%$.
△ Less
Submitted 6 July, 2026;
originally announced July 2026.
-
ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation
Authors:
Yunhan Xu,
Qifeng Wu,
Xunjin Li,
Yuanwei Bin,
Qingsong Yao,
Jianghang Gu,
Guan Wang,
Weihao Lv,
Huiyu Yang,
Wenfa Luo,
Jiao Xiang,
Yuntian Chen,
Shiyi Chen
Abstract:
Computer-aided design (CAD) for industrial components requires long-horizon procedural modeling, robust feature dependencies, editable parametric geometry, and production-grade B-Rep execution. Existing text-to-CAD methods have made promising progress in generating CAD programs from natural-language descriptions, but they still struggle when user prompts are ambiguous, underspecified, or only desc…
▽ More
Computer-aided design (CAD) for industrial components requires long-horizon procedural modeling, robust feature dependencies, editable parametric geometry, and production-grade B-Rep execution. Existing text-to-CAD methods have made promising progress in generating CAD programs from natural-language descriptions, but they still struggle when user prompts are ambiguous, underspecified, or only describe high-level design intent. They also rarely exploit expert procedural knowledge naturally available in industrial workflows, such as CATIA operation recordings, macro logs, drawing notes, and engineering descriptions. We present ArtisanCAD, a skill-guided industrial CAD agent with expert-grounded knowledge distillation. The core of ArtisanCAD is CAD intermediate representation (CAD-IR), an executable procedural representation that encodes parameters, ordered operations, MCP tool bindings, dependencies, generated entities, and verification rules. CAD-IR plays two key roles: it first serves as the carrier for distilling expert CAD procedures into reusable parameterized skills; then it provides a procedural scaffold that turns vague or intermediate-level prompts into complete executable CAD operations. ArtisanCAD retrieves expert-derived skills, instantiates and revises CAD-IR, executes the resulting procedure through a dedicated CATIA-MCP backend, and uses multi-view visual feedback for iterative refinement, and finally generates production-ready B-Rep models. On the Text2CAD benchmark, CAD-IR improves generation from intermediate prompts by reducing mean Chamfer Distance from $14.83$ to $9.88$, showing its ability to bridge ambiguous textual intent and executable CAD construction. On four complex automotive components, CAD-IR enables expert CATIA recordings to be distilled into reusable skills, allowing ArtisanCAD to generate editable CATIA-native B-Rep models for new variant requests.
△ Less
Submitted 7 July, 2026; v1 submitted 6 July, 2026;
originally announced July 2026.
-
Physically-guided Image Generation for Multi-Projection Mapping
Authors:
Xingyun Liu,
Yuqi Li,
Jinhui Xiang,
Pinyan Tang,
Chong Wang
Abstract:
Projection Mapping (PM) enables seamless superimposition of digital content onto real-world 3D objects, serving as a fundamental technique for immersive visualization, digital twins, and interactive art. Although text-to-image diffusion models have greatly facilitated customized content creation, directly integrating them into practical PM pipelines remains challenging due to the mismatch between…
▽ More
Projection Mapping (PM) enables seamless superimposition of digital content onto real-world 3D objects, serving as a fundamental technique for immersive visualization, digital twins, and interactive art. Although text-to-image diffusion models have greatly facilitated customized content creation, directly integrating them into practical PM pipelines remains challenging due to the mismatch between idealized 2D generation and physical constraints. To bridge this gap, this paper formalizes two application-level generative paradigms: the cooperative paradigm (harmonizing generated semantics with physical attributes) and the adversarial paradigm (eliminating surface interference via radiometric compensation). Based on this, we propose ConPhyG, a unified controllable physically-guided generative multi-projection mapping framework that enables creators to interactively adjust physical constraints and flexibly switch generative paradigms. In cooperative mode, multi-dimensional physical priors (per-pixel gamut, depth, and edges) are injected into the diffusion process. In adversarial mode, the framework releases the generative potential and applies bounded numerical optimization for multi-projector radiometric compensation. It allows users to dynamically switch constraints to balance artistic freedom with physical feasibility. Furthermore, we extend ConPhyG to 360-degree multi-view consistent PM using a sequential generation strategy. Quantitative and qualitative evaluations on a real-world four-projector setup demonstrate that ConPhyG significantly outperforms state-of-the-art methods in geometric alignment, gamut utilization, and semantic fidelity.
△ Less
Submitted 21 June, 2026;
originally announced June 2026.
-
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Authors:
DeepSeek-AI,
Anyi Xu,
Bangcai Lin,
Bing Xue,
Bingxuan Wang,
Bingzheng Xu,
Bochao Wu,
Bowei Zhang,
Chaofan Lin,
Chen Dong,
Chenchen Ling,
Chengda Lu,
Chenggang Zhao,
Chengqi Deng,
Chengyu Hou,
Chenhao Xu,
Chenze Shao,
Chong Ruan,
Conner Sun,
Damai Dai,
Daya Guo,
Dejian Yang,
Deli Chen,
Donghao Li,
Dongjie Ji
, et al. (294 additional authors not shown)
Abstract:
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention arc…
▽ More
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4.
△ Less
Submitted 26 April, 2026;
originally announced June 2026.
-
LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis
Authors:
Chengfu Liu,
Dongyang Hou,
Junwu Xiang,
Cheng Yang,
Xuezhi Cui,
Zeyuan Wang,
Liangtian Liu,
Zelang Miao
Abstract:
Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and high-level geoscientific semantics, while general-purpose vision-language models (VLMs) suffer from perceptual limitations and domain hallucinations in complex geological scenarios. To address these challenges, we propose an instruction-drive…
▽ More
Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and high-level geoscientific semantics, while general-purpose vision-language models (VLMs) suffer from perceptual limitations and domain hallucinations in complex geological scenarios. To address these challenges, we propose an instruction-driven agentic framework comprising three components. First, LandslideBench, a multimodal fine-grained dataset with seven subtype labels, high-resolution imagery, pixel-level masks, and high-quality textual descriptions, is constructed via multi-VLM cross-validation and interactive annotation. Then, LandslideVLM, a landslide-oriented VLM, is fine-tuned via LoRA on LandslideBench to enhance geological semantic understanding. Finally, LandslideAgent, a domain rule-enhanced agent taking LandslideVLM as its cognitive backbone, employs a dual-rule controller incorporating structured report metadata constraints and cross-validation identification constraints to regulate automated tool invocation. Experiments demonstrate that LandslideBench provides effective baselines across five mainstream models on fine-grained classification and semantic segmentation. LandslideVLM achieves accuracy improvements of 10.96%, 32.87%, and 15.91% on landslide discrimination, fine-grained classification, and semantic description quality, respectively. LandslideAgent further enables autonomous multi-source spatial data inference, realizing full-process intelligence for landslide identification and analysis.
△ Less
Submitted 17 June, 2026;
originally announced June 2026.
-
Continuous Cross-Domain Traffic State Prediction via Memory-Augmented Graph Liquid Time-Constant Networks
Authors:
Jinrong Xiang,
Ming Xu
Abstract:
Traffic state prediction is a fundamental task in intelligent transportation systems. In practical applications, some regions suffer from limited traffic observations due to insufficient sensing infrastructure, making cross-domain knowledge transfer an important solution for data-scarce traffic prediction. However, existing cross-domain traffic prediction methods still face several limitations, in…
▽ More
Traffic state prediction is a fundamental task in intelligent transportation systems. In practical applications, some regions suffer from limited traffic observations due to insufficient sensing infrastructure, making cross-domain knowledge transfer an important solution for data-scarce traffic prediction. However, existing cross-domain traffic prediction methods still face several limitations, including coarse-grained source-target adaptation, limited capability in handling unseen target-domain patterns, and insufficient modeling of continuous traffic dynamics under irregular or heterogeneous temporal conditions. To address these issues, this paper proposes a continuous cross-domain traffic prediction framework, termed Memory-Augmented Graph Liquid Time-Constant Network (MA-GLTC). Specifically, we first construct spatio-temporal units (STUs) to decompose traffic networks into transferable local units, enabling fine-grained knowledge alignment across domains. Then, a graph liquid time-constant network (GLTC) is developed to model graph-coupled traffic evolution in continuous time. Different from generic graph neural ODE-based models, GLTC introduces graph-coupled recurrent conductance into liquid time-constant dynamics, allowing node states to evolve with leakage, adaptive time constants, and neighborhood-aware feedback. Furthermore, a Memory-based Transfer Storage (MTS) mechanism is designed to preserve source-domain knowledge, retrieve matched traffic patterns, and update reliable target-domain patterns when unseen states emerge. Experiments on five public traffic datasets demonstrate that MA-GLTC consistently outperforms representative innerdomain and cross-domain baselines in both short-term and longterm prediction tasks. Compared with the second-best method, MA-GLTC reduces the average prediction errors by 3.02%, 0.33%, 8.92%, 10.09%, and 2.11%, respectively.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
Averaging principles for nonautonomous multiscale McKean-Vlasov stochastic systems
Authors:
Jie Xiang,
Huijie Qiao
Abstract:
This paper investigates a class of nonautonomous multiscale McKean-Vlasov stochastic systems. By leveraging the nonautonomous Poisson equation, we rigorously establish both strong and weak averaging principles, accompanied by explicit convergence rates. Notably, the coefficients of the averaging equations derived in the general case retain dependence on the scaling parameter $\varepsilon$. However…
▽ More
This paper investigates a class of nonautonomous multiscale McKean-Vlasov stochastic systems. By leveraging the nonautonomous Poisson equation, we rigorously establish both strong and weak averaging principles, accompanied by explicit convergence rates. Notably, the coefficients of the averaging equations derived in the general case retain dependence on the scaling parameter $\varepsilon$. However, under the additional assumptions that the fast-scale coefficients are either asymptotically convergent or time-periodic, we demonstrate that the slow component converges, in the strong or weak sense, to averaging equations with coefficients independent of $\varepsilon$.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
One Model, Multiple Goals: Adaptive Multi-Objective Learning for E-commerce Dialogue Systems
Authors:
Mingzhe Li,
Jing Xiang,
Enguo Zhou,
Lang Gao,
Tai Li,
Qishen Zhang,
Xiangliang Zhang,
Xiuying Chen
Abstract:
Dialogue systems in e-commerce scenarios often need to satisfy multiple objectives: accurately reasoning over user profiles (e.g., eligibility, credit limit) to ensure correct decision-making and user state interpretation, while also generating natural and faithful responses. These goals are complementary but not identical. In this work, we propose MORE, an adaptive Multi-Objective REinforcement l…
▽ More
Dialogue systems in e-commerce scenarios often need to satisfy multiple objectives: accurately reasoning over user profiles (e.g., eligibility, credit limit) to ensure correct decision-making and user state interpretation, while also generating natural and faithful responses. These goals are complementary but not identical. In this work, we propose MORE, an adaptive Multi-Objective REinforcement learning framework that jointly optimizes reasoning accuracy and linguistic naturalness. Our preliminary experiments show that directly mixing rewards with diverging optimization dynamics can cause oscillations and unstable learning. Thus, instead of optimizing a single mixed reward, we treat reasoning functions as constraints that guide policy optimization. At inference time, the system directly generates responses without explicit reasoning steps, while still benefiting from reasoning-enhanced scaffold and avoiding additional inference overhead. To better balance linguistic objectives during response generation, we introduce an adaptive multi-reward mechanism that aggregates signals such as fluency and naturalness and dynamically reweighs them via gradient feedback. We evaluate MORE on two real-world dialogue systems at ByteDance and the MultiWOZ 2.2 benchmark, where it consistently outperforms strong baselines. In 14-day online experiments on ByteDance production traffic, MORE improves overall and reached conversion by 16.53% and 30.09%, while increasing user satisfaction and reducing handoff rates. Notably, in a human-machine comparison, MORE recovers about 60% of the incremental conversion lift achieved by human agents.
△ Less
Submitted 8 June, 2026;
originally announced June 2026.
-
SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter
Authors:
Powei Chang,
Jinpeng Zhang,
Chaoqun Sun,
MiniWell Tsao,
Lianrui Li,
Jianxiang Xiang,
Chenyu Wang,
Yukang Gao,
Dongying Kong
Abstract:
Reinforcement learning with verifiable rewards (RLVR) often adopts GRPO-style group-relative updates, sampling multiple rollouts per prompt to construct normalized learning signals. However, merely increasing the number of rollouts does not reliably strengthen learning: under GRPO-style group normalization, per-rollout policy-gradient features can concentrate into a low-rank, signed geometry, caus…
▽ More
Reinforcement learning with verifiable rewards (RLVR) often adopts GRPO-style group-relative updates, sampling multiple rollouts per prompt to construct normalized learning signals. However, merely increasing the number of rollouts does not reliably strengthen learning: under GRPO-style group normalization, per-rollout policy-gradient features can concentrate into a low-rank, signed geometry, causing substantial cancellation during aggregation and weakening the effective update. We address this failure mode with SALT, a Subspace-Adaptive geometry pLug-in componenT that uses sample-wise gradient geometry to reweight the coefficients of group-relative updates. SALT estimates a dominant shared subspace from the mini-batch Gram geometry, decomposes group-relative coefficients into shared and residual channels, and adaptively amplifies the residual channel when signed cancellation is severe. Across diverse reasoning-oriented RLVR benchmarks and model scales, SALT improves effective update geometry and performance without modifying the reward model or the rollout sampling procedure
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
Overview of the ClinicalSkillQA 2026 Shared Task on Continuous Perception and Procedural Reasoning in Clinical Skill Assessment
Authors:
Xiyang Huang,
Renxiong Wei,
Yihuai Xu,
Zhiyuan Chen,
Keying Wu,
Jiayi Xiang,
Buzhou Tang,
Yanqing Ye,
Jinyu Chen,
Cheng Zeng,
Min Peng,
Qianqian Xie,
Sophia Ananiadou
Abstract:
This paper presents an overview of the ClinicalSkillQA 2026 shared task, which was organized with the BioNLP Workshop at ACL 2026. The goal of this shared task is to evaluate continuous perception and procedural reasoning in clinical skill assessment by requiring systems to reconstruct the correct temporal order of shuffled clinical key frames and generate rationales grounded in clinical workflow…
▽ More
This paper presents an overview of the ClinicalSkillQA 2026 shared task, which was organized with the BioNLP Workshop at ACL 2026. The goal of this shared task is to evaluate continuous perception and procedural reasoning in clinical skill assessment by requiring systems to reconstruct the correct temporal order of shuffled clinical key frames and generate rationales grounded in clinical workflow knowledge. The benchmark contains 200 test-only instances sampled from clinical skill videos, covering three emergency-care procedures. Each instance is annotated with the ground-truth temporal order and an expert-verified rationale. A total of seven teams participated in the task, collectively making 90 submissions, with four teams providing system description papers. Systems are evaluated using Task Accuracy, Pairwise Accuracy, and BERTScore, which measure exact sequence reconstruction, local temporal consistency, and rationale quality, respectively. In this paper, we describe the task setup, dataset construction, and evaluation criteria. We further summarize the methodologies adopted by participating teams and present a comprehensive analysis of the submitted systems. The official results suggest that current models still struggle with continuous perception and procedural reasoning, especially when they must integrate visual evidence, temporal structure, and clinical workflow knowledge.
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
EIVE: End-to-End Instance-Specific Visual Explanations for Detection Transformers
Authors:
Jianlin Xiang,
Yanshan Li,
Linhui Dai
Abstract:
Visual explainability for object detection remains challenging due to the multi-instance nature of detection. Existing approaches predominantly adopt post-hoc paradigms, such as gradient-based or perturbation-based explanation methods, to interpret pretrained detectors. However, these methods require additional gradient computation or repeated model inference, resulting in limited efficiency. To a…
▽ More
Visual explainability for object detection remains challenging due to the multi-instance nature of detection. Existing approaches predominantly adopt post-hoc paradigms, such as gradient-based or perturbation-based explanation methods, to interpret pretrained detectors. However, these methods require additional gradient computation or repeated model inference, resulting in limited efficiency. To address this issue, we propose an End-to-end Instance-specific Visual Explanation framework (EIVE) that directly generates instance-level saliency maps following the forward pass of Detection Transformer (DETR)-like models. Specifically, we reformulate the cross-attention mechanism in the decoder as an instance-level feature attribution pathway, so that the cross-attention of each object query corresponds to the visual attribution of its predicted instance. Based on this formulation, we design a cross-layer hybrid consensus fusion (CLHCF) module to aggregate cross-attention signals across decoder layers, producing stable and compact explanations. The explanation process of EIVE requires neither gradient computation nor input perturbation, yielding high computational efficiency, and applies to single- and multi-scale DETR-like object detectors. Finally, we present an attention-aware joint training strategy (AAJTS) as a training-oriented application, which imposes spatial constraints on cross-attention patterns to encourage stable and concentrated attribution representations, thereby improving both interpretability and detection performance. Experiments on MS COCO 2017, ExDark, and Cityscapes demonstrate that EIVE produces high-quality instance-level saliency maps and achieves performance comparable to, or better than, state-of-the-art post-hoc methods across standard metrics, while substantially improving explanation efficiency. Code is available at https://github.com/xjlDestiny/EIVE.git.
△ Less
Submitted 31 May, 2026;
originally announced June 2026.
-
Lightning Plus Polynomial Approximation: Optimal Root-Exponential Convergence for Singular Functions in Corner Domains
Authors:
Shuhuang Xiang,
Jun Xiang,
Shunfeng Yang,
Yuee Zhong
Abstract:
This paper presents a rigorous convergence analysis for the lightning plus polynomial approximation scheme, which employs rational approximations constructed with preassigned tapered, exponentially clustered poles. This pole placement strategy was originally introduced by Trefethen and his collaborators for the resolution of corner singularities. Ample numerical results indicate that this scheme a…
▽ More
This paper presents a rigorous convergence analysis for the lightning plus polynomial approximation scheme, which employs rational approximations constructed with preassigned tapered, exponentially clustered poles. This pole placement strategy was originally introduced by Trefethen and his collaborators for the resolution of corner singularities. Ample numerical results indicate that this scheme achieves root-exponential convergence, and in particular attains the same optimal convergence rate as the best rational approximation to $x^α$ on $[0,1]$ established by Stahl.% which is conjectured in [SIAM J. Numer. Anal., 61:2580-2600, 2023].
In this work, we establish optimal root-exponential convergence for the class of prototype functions of the form $g(z)z^α$ or $g(z)z^α\log z$, where $g$ is analytic on a neighborhood of the sector domain. These results confirm the validity of Conjectures 3.1 and 5.3 stated in [SIAM J. Numer. Anal., 61:2580-2600, 2023], and demonstrate that the choice $σ_{\mathrm{opt}} =\frac{\sqrt{2(2 - β)}π}{\sqrtα}$ achieves the theoretically optimal convergence rate $\mathcal{O}\left(e^{-\sqrt{2(2 - β)Nα}π}\right)$. Notably, for the specific case of $β= 0$, the scheme recovers Stahl's optimal convergence rate for $x^α$. Furthermore, working within the decomposition framework for corner domains proposed by Gopal and Trefethen, this paper provides a rigorous proof of optimal root-exponential convergence for lightning plus polynomial approximation problems on corner domains, and explicitly derives the optimal pole clustering parameter.
△ Less
Submitted 7 July, 2026; v1 submitted 28 May, 2026;
originally announced May 2026.
-
Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation
Authors:
Ning Wu,
Rui Liu,
Xinkun Lin,
Weixing Chen,
Jinxi Xiang,
Tao Wei,
Lina Yao,
Mingjie Li
Abstract:
Distilling demonstration effects into hidden-space interventions offers a lightweight alternative to full finetuning. However, existing multimodal variants are mostly evaluated on short-form tasks, where outputs end after a few tokens. Extending these methods to long-form generation exposes a fundamental yet underexamined limitation: token-level distillation implicitly treats all output tokens as…
▽ More
Distilling demonstration effects into hidden-space interventions offers a lightweight alternative to full finetuning. However, existing multimodal variants are mostly evaluated on short-form tasks, where outputs end after a few tokens. Extending these methods to long-form generation exposes a fundamental yet underexamined limitation: token-level distillation implicitly treats all output tokens as equally informative, but long-form outputs are dominated by high-frequency template and grammatical tokens, while the tokens that actually determine output quality are sparsely distributed. In medical report generation (MRG), two such decisive tokens stand out: pathology-related tokens that determine diagnostic content, and the end-of-sequence (EOS) event that determines termination. Both receive insufficient supervision under uniform cross-entropy, and autoregressive decoding further compounds the problem by drifting away from teacher-forced trajectories. We propose DIVE, a frozen-backbone distillation framework that addresses long-form report generation through two complementary mechanisms matched to these failures. Decisive-token supervision restores supervision balance by upweighting the cross-entropy contribution of pathology-related tokens and the EOS event, ensuring that content fidelity and termination are learned during training rather than imposed at decoding time. State-conditioned dynamic steering replaces fixed open-loop residuals with hidden-state-dependent adapters, allowing the injected signal to adapt as decoding drifts. Experiments on MIMIC-CXR and CheXpert Plus with two medical VLM backbones show that DIVE consistently ranks among the strongest methods across lexical and clinical-proxy metrics. Our method achieves the best BLEU-4, ROUGE-L, and RadGraph F1 in all dataset--backbone settings, while remaining competitive on coarse label-level CheXbert F1.
△ Less
Submitted 26 May, 2026;
originally announced May 2026.
-
Approaching physical limits of latent dimensionality in optical computing
Authors:
Zhenyu Zhao,
Zijun Qiu,
Xuan Hu,
Yao Zhou,
Jinlong Xiang,
Youlve Chen,
Chaojun Xu,
Yuchen Yin,
Tao Lin,
Yikai Su,
Xuhan Guo
Abstract:
The physical implementation of artificial intelligence requires mapping computational processes onto the dynamic physical processes of the underlying computing platform. The photonic processors offer an intrinsically parallel and low energy framework for this mapping, however, a mismatch between the potential computing capability of a bounded optical domain and the human accessible manipulation ra…
▽ More
The physical implementation of artificial intelligence requires mapping computational processes onto the dynamic physical processes of the underlying computing platform. The photonic processors offer an intrinsically parallel and low energy framework for this mapping, however, a mismatch between the potential computing capability of a bounded optical domain and the human accessible manipulation range sets a hard integration density ceiling on existing architectures. Here, we address this challenge by investigating the integration density limits in photonic processors through exploring the fundamental physical limits on the latent dimensionality for maximum expressivity of a bounded optical domain. These physical limits potentially serve as universal metrics for evaluating optical computing capacity. To validate these, we design and realize ultracompact multimode photonic processors approaching these limits: a 2.2 um by 8 um processor achieves 86.7 % accuracy in experiment for iris flower classification, and a 20.6 um by 44.8 um processor reaches 92.9% accuracy in handwritten digit recognition. Finally, we scale this architecture to highly complex tasks by implementing a generative diffusion model for image synthesis. By grounding photonic processor design in the wave physics origin of latent dimensionality, our results supply the missing theoretical reference point for optical computing architecture.
△ Less
Submitted 22 May, 2026;
originally announced May 2026.
-
Distributed Image Compression with Multimodal Side Information at Extremely Low Bitrates
Authors:
Guojun Xu,
Mingyang Zhang,
Jianwen Xiang,
Cheng Tan,
Yanchao Yang,
Junwei Zhou
Abstract:
Distributed Image Compression (DIC) is crucial for multi-view transmission, especially when operating at extremely low bitrates (< 0.1 bpp). Its core challenge is effectively utilizing side information to achieve high-quality reconstruction under strict bitrate budgets. However, existing DIC approaches struggle to exploit global context and object-level details from side information, leading to lo…
▽ More
Distributed Image Compression (DIC) is crucial for multi-view transmission, especially when operating at extremely low bitrates (< 0.1 bpp). Its core challenge is effectively utilizing side information to achieve high-quality reconstruction under strict bitrate budgets. However, existing DIC approaches struggle to exploit global context and object-level details from side information, leading to local blurring and the loss of fine details in the reconstruction. To address these limitations, we propose a Multimodal DIC framework (MDIC), which, for the first time, leverages side information in a multimodal manner into the DIC paradigm, effectively preserving fine-grained local details and enhancing global perceptual quality in reconstructed images. Specifically, we introduce a text-to-image diffusion-based decoder conditioned on textual side information extracted from correlated images to capture shared global semantics. Moreover, we design a feature-mask generator, supervised by a multimodal fine-grained alignment task, to strengthen the exploitation of visual side information. The generated mask serves two purposes: first, it guides the extraction of fine-grained details from losslessly transmitted side information to preserve the semantic consistency of reconstructed details; second, it regulates the extraction of clustered feature representations from the quantized VQ-VAE embeddings, compensating for category information lost under the extreme compression of the primary image. Extensive experiments on the widely used KITTI Stereo and Cityscapes datasets demonstrate that MDIC achieves state-of-the-art perceptual quality at extremely low bitrates.
△ Less
Submitted 21 May, 2026;
originally announced May 2026.
-
Scalable Environments Drive Generalizable Agents
Authors:
Jiayi Zhang,
Fanqi Kong,
Guibin Zhang,
Maojia Song,
Zhaoyang Yu,
Jianhao Ruan,
Jinyu Xiang,
Bang Liu,
Chenglin Wu,
Yuyu Luo
Abstract:
Generalizable agents should adapt to diverse tasks and unseen environments beyond their training distribution. This position paper argues that such generalization requires environment scaling: expanding the distribution of executable rule-sets that agents interact with, rather than only increasing trajectories or tasks within fixed benchmarks. Current scaling practices largely focus on collecting…
▽ More
Generalizable agents should adapt to diverse tasks and unseen environments beyond their training distribution. This position paper argues that such generalization requires environment scaling: expanding the distribution of executable rule-sets that agents interact with, rather than only increasing trajectories or tasks within fixed benchmarks. Current scaling practices largely focus on collecting more experience or broader task sets under fixed interaction rules, leaving agents brittle when underlying interfaces, dynamics, observations, or feedback signals change. The core challenge is therefore a world-level distribution shift: agents need systematic exposure to environments with meaningfully different executable rule-sets. To clarify this challenge, we propose a unified taxonomy that separates trajectory scaling, task scaling, and environment scaling by their primary deliverables and by what changes in the executable rule-set. Building on this taxonomy, we synthesize construction paradigms for scalable environments, contrasting programmatic generators that prioritize controllability and verifiability with generative world models that offer broader coverage and open-endedness. We further outline how environment scaling can be coupled with stateful learning mechanisms, emphasizing learned update rules for cross-environment adaptation. We conclude by discussing alternative perspectives and argue that scalable environments provide the essential substrate for measurable and controllable progress toward robust general agents.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
Ascend-RaBitQ: Heterogeneous NPU-CPU Acceleration of Billion-Scale Similarity Search with 1-bit Quantization
Authors:
Fujun He,
Chuyue Ye,
Huaxiang Cai,
Zetao Lv,
Baolong Cui,
Wenru Yan,
Chao Zhan,
Zigang Zhang,
Hao Yi,
Jie Xiang,
Xiabing Li,
Yuhang Gai,
Ziyang Zhang,
Pengfei Zheng,
Yunfei Du
Abstract:
Vector similarity search is a critical component of modern AI systems, but traditional CPU-based implementations face fundamental scalability bottlenecks for billion-scale corpora due to prohibitive computational overhead and memory bandwidth limitations. While Neural Processing Units (NPUs) offer orders-of-magnitude higher compute density, existing CPU/GPU-optimized 1-bit RaBitQ quantization impl…
▽ More
Vector similarity search is a critical component of modern AI systems, but traditional CPU-based implementations face fundamental scalability bottlenecks for billion-scale corpora due to prohibitive computational overhead and memory bandwidth limitations. While Neural Processing Units (NPUs) offer orders-of-magnitude higher compute density, existing CPU/GPU-optimized 1-bit RaBitQ quantization implementations cannot be directly ported to NPU architectures due to fundamental hardware mismatches, and homogeneous design paradigms struggle to simultaneously balance accuracy, memory footprint, and performance.
This paper presents Ascend-RaBitQ, the first heterogeneous NPU-CPU optimized IVF-RaBitQ system for billion-scale vector search, built on the core insight that decoupling coarse ranking (NPU) from fine ranking (CPU) allows each stage to leverage its optimal hardware, breaking the long-standing accuracy-memory-performance trade-off. We propose a three-stage heterogeneous execution path comprising AI Core-accelerated coarse ranking on 1-bit quantized vectors, on-device AI CPU Top-k processing, and host CPU fine re-ranking on full-precision vectors. We introduce four NPU architecture-native optimizations: fused AIC-AIV operators for parallel distance computation, computation flow restructuring to exploit rotation orthogonality, fine-grained index block-level load balancing that breaks query boundaries, and intra-NPU pipeline parallelism between AI Core and AI CPU to mask Top-k latency. Evaluation on standard datasets shows that Ascend-RaBitQ achieves 3.0X to 62.8X faster index construction than the CPU baseline, up to 11.7X throughput improvement over the fastest CPU IVF-RaBitQ implementation, and over two orders of magnitude over the mathematically equivalent CPU baseline, while demonstrating encouraging scalability on distributed multi-NPU systems.
△ Less
Submitted 14 June, 2026; v1 submitted 15 May, 2026;
originally announced May 2026.
-
Harnessing Agentic Evolution
Authors:
Jiayi Zhang,
Yongfeng Gu,
Jianhao Ruan,
Maojia Song,
Yiran Peng,
Zhiguang Han,
Jinyu Xiang,
Zhitao Wang,
Caiyin Yang,
Yixi Ouyang,
Bang Liu,
Chenglin Wu,
Yuyu Luo
Abstract:
Agentic evolution has emerged as a powerful paradigm for improving programs, workflows, and scientific solutions by iteratively generating candidates, evaluating them, and using feedback to guide future search. However, existing methods are typically instantiated either as fixed hand-designed procedures that are modular but rigid, or as general-purpose agents that flexibly integrate feedback but c…
▽ More
Agentic evolution has emerged as a powerful paradigm for improving programs, workflows, and scientific solutions by iteratively generating candidates, evaluating them, and using feedback to guide future search. However, existing methods are typically instantiated either as fixed hand-designed procedures that are modular but rigid, or as general-purpose agents that flexibly integrate feedback but can drift in long-horizon evolution. Both forms accumulate rich evidence over time, including candidates, feedback, traces, and failures, yet lack a stable interface for organizing this evidence and revising the mechanism that drives future evolution. We address this limitation by formulating agentic evolution as an interactive environment, where the accumulated evolution context serves as a process-level state. We introduce AEvo, a harnessed meta-editing framework in which a meta-agent observes this state and acts not by directly proposing the next candidate, but by editing the procedure or agent context that controls future evolution. This unified interface enables AEvo to steer both procedure-based and agent-based evolution, making accumulated evidence actionable for long-horizon search. Empirical evaluations on agentic and reasoning benchmarks show that AEvo outperforms five evolution baselines, achieving a 26 relative improvement over the strongest baseline. Across three open-ended optimization tasks, AEvo further outperforms four evolution baselines and achieves state-of-the-art performance under the same iteration budget.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Map2World: Segment Map Conditioned Text to 3D World Generation
Authors:
Jaeyoung Chung,
Suyoung Lee,
Jianfeng Xiang,
Jiaolong Yang,
Kyoung Mu Lee
Abstract:
3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D w…
▽ More
3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D world generation conditioned on user-defined segment maps of arbitrary shapes and scales, ensuring global-scale consistency and flexibility across expansive environments. To further enhance the quality, we propose a detail enhancer network that generates fine details of the world. The detail enhancer enables the addition of fine-grained details without compromising overall scene coherence by incorporating global structure information. We design the entire pipeline to leverage strong priors from asset generators, achieving robust generalization across diverse domains, even under limited training data for scene generation. Extensive experiments demonstrate that our method significantly outperforms existing approaches in user-controllability, scale consistency, and content coherence, enabling users to generate 3D worlds under more complex conditions.
△ Less
Submitted 1 May, 2026;
originally announced May 2026.
-
Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis
Authors:
Jiahong Xiang,
Xiaoyang Xu,
Xiaopan Chu,
Hongliang Tian,
Yuqun Zhang
Abstract:
Autonomous agents for automated program repair represent a promising frontier in software engineering, yet their effectiveness is often hindered by reliance on post-mortem, coarse-grained execution feedback. While integrating traditional interactive debuggers seems a natural solution, their low-level, line-by-line interaction paradigm turns out to be cost-inefficient for LLM-based agents, leading…
▽ More
Autonomous agents for automated program repair represent a promising frontier in software engineering, yet their effectiveness is often hindered by reliance on post-mortem, coarse-grained execution feedback. While integrating traditional interactive debuggers seems a natural solution, their low-level, line-by-line interaction paradigm turns out to be cost-inefficient for LLM-based agents, leading to exhausted budgets and unproductive loops. To mitigate this, we introduce Agent-centric Debugging Interface (ADI), a novel agent-centric debugging interface designed for cost-efficient, end-to-end autonomous interaction. Specifically, Agent-centric Debugging Interface realizes a function-level interaction paradigm, powered by our Frame Lifetime Trace, a comprehensive data structure encapsulating a function's stateful execution trace, and a set of high-level navigational commands. Our extensive evaluation on the SWE-bench benchmark demonstrates the effectiveness and efficiency of ADI. By simply equipping a basic agent with ADI, it successfully resolves 63.8\% of the tasks on the SWE-bench Verified set, even slightly outperforming the highly optimized and high-investment Claude-Tools agent, at an average cost of USD 1.28 per task with Claude-Sonnet-3.7. Furthermore, we demonstrate ADI's generality by integrating it as a plug-and-play component into existing SOTA agents, delivering consistent gains ranging from 6.2\% to 18.5\% on the resolved tasks. These results indicate that Agent-centric Debugging Interface can provide a general and efficient enhancement for existing autonomous agents.
△ Less
Submitted 27 April, 2026;
originally announced April 2026.
-
Validating a Deep Learning Algorithm to Identify Patients with Glaucoma using Systemic Electronic Health Records
Authors:
John Xiang,
Rohith Ravindranath,
Sophia Y. Wang
Abstract:
We evaluated whether a glaucoma risk assessment (GRA) model trained on All of Us national data can identify patients at high probability of glaucoma using only systemic electronic health records (EHR) at an independent institution. In this cross-sectional study, 20,636 Stanford patients seen from November 2013 to January 2024 were included (15% with glaucoma). A pretrained GRA model was fine-tuned…
▽ More
We evaluated whether a glaucoma risk assessment (GRA) model trained on All of Us national data can identify patients at high probability of glaucoma using only systemic electronic health records (EHR) at an independent institution. In this cross-sectional study, 20,636 Stanford patients seen from November 2013 to January 2024 were included (15% with glaucoma). A pretrained GRA model was fine-tuned on the Stanford cohort and tested on a held-out set using demographics, systemic diagnoses, medications, laboratory results, and physical examination measurements as inputs. The best model achieved AUROC 0.883 and PPV 0.657. Calibration was consistent with clinical risk: the highest prediction decile showed the greatest glaucoma diagnosis rate (65.7%) and treatment rate (57.0%). Performance improved with more trainable layers up to 15 and with additional data. An EHR-only GRA model may enable scalable and accessible pre-screening without specialized imaging.
△ Less
Submitted 22 April, 2026;
originally announced April 2026.
-
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
Authors:
Chaojie Mao,
Chen-Wei Xie,
Chongyang Zhong,
Haoyou Deng,
Jiaxing Zhao,
Jie Xiao,
Jinbo Xing,
Jingfeng Zhang,
Jingren Zhou,
Jingyi Zhang,
Jun Dan,
Kai Zhu,
Kang Zhao,
Keyu Yan,
Minghui Chen,
Pandeng Li,
Shuangle Chen,
Tong Shen,
Yu Liu,
Yue Jiang,
Yulin Pan,
Yuxiang Tuo,
Zeyinzi Jiang,
Zhen Han,
Ang Wang
, et al. (33 additional authors not shown)
Abstract:
We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivity tools. While contemporary diffusion models excel at aesthetic generation, they frequently encounter critical bottlenecks in rigorous design workflows that demand absolute controllability, complex typography rendering,…
▽ More
We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivity tools. While contemporary diffusion models excel at aesthetic generation, they frequently encounter critical bottlenecks in rigorous design workflows that demand absolute controllability, complex typography rendering, and strict identity preservation. To address these challenges, Wan-Image features a natively unified multi-modal architecture by synergizing the cognitive capabilities of large language models with the high-fidelity pixel synthesis of diffusion transformers, which seamlessly translates highly nuanced user intents into precise visual outputs. It is fundamentally powered by large-scale multi-modal data scaling, a systematic fine-grained annotation engine, and curated reinforcement learning data to surpass basic instruction following and unlock expert-level professional capabilities. These include ultra-long complex text rendering, hyper-diverse portrait generation, palette-guided generation, multi-subject identity preservation, coherent sequential visual generation, precise multi-modal interactive editing, native alpha-channel generation, and high-efficiency 4K synthesis. Across diverse human evaluations, Wan-Image exceeds Seedream 5.0 Lite and GPT Image 1.5 in overall performance, reaching parity with Nano Banana Pro in challenging tasks. Ultimately, Wan-Image revolutionizes visual content creation across e-commerce, entertainment, education, and personal productivity, redefining the boundaries of professional visual synthesis.
△ Less
Submitted 23 April, 2026; v1 submitted 21 April, 2026;
originally announced April 2026.
-
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
Authors:
Xinwei He,
Yansong Zheng,
Qianru Han,
Zhichuan Wang,
Yuxuan Cai,
Yang Zhou,
Jingbo Xia,
Yulong Wang,
Jinhai Xiang,
Xiang Bai
Abstract:
Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent space, previous work typically adapts the CLIP encoder to build view-based 3D descriptors. Despite CLIP's strong generalization ability, its lack of fine-grainedness prompted us to explore the potential of a more recent…
▽ More
Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent space, previous work typically adapts the CLIP encoder to build view-based 3D descriptors. Despite CLIP's strong generalization ability, its lack of fine-grainedness prompted us to explore the potential of a more recent self-supervised encoder-DINO. To address this, we propose DINO Eats CLIP (DEC), a novel framework for dynamic multi-view integration that is regularized by synthesizing data for unseen classes. We first find that simply mean-pooling over view features from a frozen DINO backbone gives decent performance. Yet, further adaptation causes severe overfitting on average view patterns of known classes. To combat it, we then design a module named Chunking and Adapting Module (CAM). It segments multi-view images into chunks and dynamically integrates local view relations, yielding more robust features than the standard pooling strategy. Finally, we propose Virtual Feature Synthesis (VFS) module to mitigate bias towards known categories explicitly. Under the hood, VFS leverages CLIP's broad, pre-aligned vision-language space to synthesize virtual features for unseen classes. By exposing DEC to these virtual features, we greatly enhance its open-set discrimination capacity. Extensive experiments on standard open-set 3DOR benchmarks demonstrate its superior efficacy.
△ Less
Submitted 21 April, 2026;
originally announced April 2026.
-
Proactive Detection of GUI Defects in Multi-Window Scenarios via Multimodal Reasoning
Authors:
Xinyao Zhang,
Rui Wang,
Jinhao Cui,
Haotian Huang,
Wei Xue,
Wenhua Hu,
Jianwen Xiang,
Rui Hao
Abstract:
Multi-window mobile scenarios, such as split-screen and foldable modes, make GUI display defects more likely by forcing applications to adapt to changing window sizes and dynamic layout reflow. Existing detection techniques are limited in two ways: they are largely passive, analyzing screenshots only after problematic states have been reached, and they are mainly designed for conventional full-scr…
▽ More
Multi-window mobile scenarios, such as split-screen and foldable modes, make GUI display defects more likely by forcing applications to adapt to changing window sizes and dynamic layout reflow. Existing detection techniques are limited in two ways: they are largely passive, analyzing screenshots only after problematic states have been reached, and they are mainly designed for conventional full-screen interfaces, making them less effective in multi-window settings.We propose an end-to-end framework for GUI display defect detection in multi-window mobile scenarios. The framework proactively triggers split-screen, foldable, and window-transition states during app exploration, uses Set-of-Mark (SoM) to align screenshots with widget-level interface elements, and leverages multimodal large language models with chain-of-thought prompting to detect, localize, and explain display defects. We also construct a benchmark of GUI display defects using 50 real-world Android applications.Experimental results show that multi-window settings substantially increase the exposure of layout-related defects, with text truncation increasing by 184% compared with conventional full-screen settings. At the application level, our method detects 40 defect-prone apps with a false positive rate of 10.00% and a false negative rate of 11.11%, outperforming OwlEye and YOLO-based baselines. At the fine-grained level, it achieves the best F1 score of 87.2% for widget occlusion detection.
△ Less
Submitted 21 April, 2026;
originally announced April 2026.
-
Dual-stream Spatio-Temporal GCN-Transformer Network for 3D Human Pose Estimation
Authors:
Jiawen Duan,
Jian Xiang,
Zhiqiang Li,
Linlin Xue,
Wan Xiang
Abstract:
3D human pose estimation is a classic and important research direction in the field of computer vision. In recent years, Transformer-based methods have made significant progress in lifting 2D to 3D human pose estimation. However, these methods primarily focus on modeling global temporal and spatial relationships, neglecting local skeletal relationships and the information interaction between diffe…
▽ More
3D human pose estimation is a classic and important research direction in the field of computer vision. In recent years, Transformer-based methods have made significant progress in lifting 2D to 3D human pose estimation. However, these methods primarily focus on modeling global temporal and spatial relationships, neglecting local skeletal relationships and the information interaction between different channels. Therefore, we have proposed a novel method,the Dual-stream Spatio-temporal GCN-Transformer Network (MixTGFormer). This method models the spatial and temporal relationships of human skeletons simultaneously through two parallel channels, achieving effective fusion of global and local features. The core of MixTGFormer is composed of stacked Mixformers. Specifically, the Mixformer includes the Mixformer Block and the Squeeze-and-Excitation Layer ( SE Layer). It first extracts and fuses various information of human skeletons through two parallel Mixformer Blocks with different modes. Then, it further supplements the fused information through the SE Layer. The Mixformer Block integrates Graph Convolutional Networks (GCN) into the Transformer, enhancing both local and global information utilization. Additionally, we further implement its temporal and spatial forms to extract both spatial and temporal relationships. We extensively evaluated our model on two benchmark datasets (Human3.6M and MPI-INF-3DHP). The experimental results showed that, compared to other methods, our MixTGFormer achieved state-of-the-art results, with P1 errors of 37.6mm and 15.7mm on these datasets, respectively.
△ Less
Submitted 19 April, 2026;
originally announced April 2026.
-
PIEDet: Prototype-Driven Intrinsically Explainable Object Detection
Authors:
Jianlin Xiang,
Linhui Dai,
Xue Yang,
Chaolei Yang,
Yanshan Li
Abstract:
Existing object detectors typically make predictions in a black-box manner and struggle to simultaneously provide discriminative evidence for their predictions, which limits their deployment in safety-critical scenarios. To explain model predictions, existing post-hoc explanation methods mostly rely on gradient-based or perturbation-based operators. These methods not only introduce additional memo…
▽ More
Existing object detectors typically make predictions in a black-box manner and struggle to simultaneously provide discriminative evidence for their predictions, which limits their deployment in safety-critical scenarios. To explain model predictions, existing post-hoc explanation methods mostly rely on gradient-based or perturbation-based operators. These methods not only introduce additional memory and computational overhead but also make it difficult to ensure that the generated explanations faithfully reflect the model's internal decision-making process. To address these limitations, we propose PIEDet, a prototype-driven intrinsically explainable object detection framework. PIEDet innovatively embeds class prototypes as explicit discriminative units into the classification branch of a one-stage detector, thereby improving detection performance while providing intrinsic interpretability. First, PIEDet constructs hierarchical class prototypes at different detection levels, enabling the model to learn scale-aware class-semantic representations. Second, we propose a prototype-driven feature learning method consisting of prototype regularization and a region-to-prototype matching loss. The former enhances the inter-class discriminability of the prototypes, while the latter encourages prototype responses to focus on object regions. Finally, we introduce a scale-aligned hierarchical prototype supervision mechanism that assigns scale-matched supervision signals to different detection levels, thereby enhancing the scale specificity of the hierarchical prototypes. On the ExDark, RTTS, and VOC2012-FOG datasets, PIEDet improves mAP@0.5 over the baseline by 4.7%, 1.6%, and 4.8%, respectively, while demonstrating superior computational efficiency. Compared with mainstream post-hoc explanation methods, PIEDet achieves a better balance between explanation quality and explanation cost.
△ Less
Submitted 14 August, 2026; v1 submitted 15 April, 2026;
originally announced April 2026.
-
Beyond Reconstruction: Reconstruction-to-Vector Diffusion for Hyperspectral Anomaly Detection
Authors:
Jijun Xiang,
Tao Wang,
Jiayi Wang,
Pengxiang Wang,
Cheng Chen,
Nian Wang
Abstract:
While Hyperspectral Anomaly Detection (HAD) excels at identifying sparse targets in complex scenes, existing models remain trapped in a scalar "reconstruction-as-endpoint" paradigm. This reliance on ambiguous scalar residuals consistently triggers sub-pixel anomaly vanishing during spatial downsampling, alongside severe confirmation bias when unpurified anomalies corrupt training weights. In this…
▽ More
While Hyperspectral Anomaly Detection (HAD) excels at identifying sparse targets in complex scenes, existing models remain trapped in a scalar "reconstruction-as-endpoint" paradigm. This reliance on ambiguous scalar residuals consistently triggers sub-pixel anomaly vanishing during spatial downsampling, alongside severe confirmation bias when unpurified anomalies corrupt training weights. In this paper, we propose Reconstruction-to-Vector Diffusion (R2VD), which fundamentally redefines reconstruction as a manifold purification origin to establish a novel residual-guided generative dynamics paradigm. Our framework introduces a four-stage pipeline: (1) a Physical Prior Extraction (PPE) stage that mitigates early confirmation bias via dual-stream statistical guidance; (2) a Guided Manifold Purification (GMP) stage utilizing an OmniContext Autoencoder (OCA) to extract purified residual maps while preserving fragile sub-pixel topologies; (3) a Residual Score Modeling (RSM) stage where a Diffusion Transformer (DiT), guarded by a Physical Spectral Firewall (PSF), effectively isolates cross-spectral leakage; and (4) a Vector Dynamics Inference (VDI) stage that robustly decouples targets from backgrounds by evaluating high-dimensional vector interference patterns instead of conventional scalar errors. Comprehensive evaluations on eight datasets confirm that R2VD establishes a new state-of-the-art, delivering exceptional target detectability and background suppression. The code is available at https://github.com/Bondojijun/R2VD.
△ Less
Submitted 14 April, 2026; v1 submitted 13 April, 2026;
originally announced April 2026.
-
VCC-DSA: A Novel Vascular Consistency Constrained DSA Imaging Model for Motion Artifact Suppression
Authors:
Rongjun Ge,
Weilong Mao,
Jian Lu,
Rong Yan,
Yikun Zhang,
Peng Yuan,
Jun Xiang,
Hui Tang,
Guanyu Yang,
Yudong Zhang,
Yang Chen,
Shuo Li
Abstract:
Digital Subtraction Angiography (DSA) is a clinically significant imaging technique for diagnosing cerebrovascular disease, as gold-standard. However, the artifacts caused by motion of high-attenuation tissues such as bones, teeth, and catheters, seriously reduce the visibility of blood vessels. This paper presents a novel Vascular Consistency Constrained DSA Imaging Model (VCC-DSA) for robust mot…
▽ More
Digital Subtraction Angiography (DSA) is a clinically significant imaging technique for diagnosing cerebrovascular disease, as gold-standard. However, the artifacts caused by motion of high-attenuation tissues such as bones, teeth, and catheters, seriously reduce the visibility of blood vessels. This paper presents a novel Vascular Consistency Constrained DSA Imaging Model (VCC-DSA) for robust motion suppression and precise vascular imaging with the following designs: 1) We specially design a Learning-based Subtraction Mapping Paradigm, so that the ill-posed problem of existing learning-based methods can be solved to enhance the stability of the algorithm. 2) Our model effectively develops Residual Dense Blocks and details-shortcut to improve the performance under complex structures, such as moving bones overlapping with blood vessels, and small features, like peripheral vessels. 3) An innovative Vascular Consistency Strategy is proposed to extract intrinsically consistency from the various relative motions in mask-live images, so that spontaneously distils the vascular structure with contrast-agent development and robustly suppress motion artifacts, and also naturally alleviates the high matching requirements of data. 4) We creatively design a Mixup-based Data Self-evolution Strategy for data-intra self-enhancement in training loop, so that the training data gains dynamically optimized to promote model better learning the vascular features, and excluding the irrelevant structures in live/mask image and even the inevitable-artifacts/fake-structure in label. Prospectively, to further evaluate practical value, an actual general anesthesia animal experiment is specially conducted, besides the assessment on human clinical data. Compared with other method, our model improves the PSNR and SSIM by 73.4% and 8.56%, respectively.
△ Less
Submitted 12 April, 2026;
originally announced April 2026.
-
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos
Authors:
Xiyang Huang,
Jiawei Lin,
Keying Wu,
Jiaxin Huang,
Kailai Yang,
Renxiong Wei,
Cheng zeng,
Jiayi Xiang,
Ziyan Kuang,
Min Peng,
Qianqian Xie,
Sophia Ananiadou
Abstract:
Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but overlook a harder capability required for expert procedural judgment: tracking how ongoing interactions update the procedural state and thereby determine the correctness of later actions. We introduce SiMing-Bench, the first benchmark for evaluating this…
▽ More
Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but overlook a harder capability required for expert procedural judgment: tracking how ongoing interactions update the procedural state and thereby determine the correctness of later actions. We introduce SiMing-Bench, the first benchmark for evaluating this capability from full-length clinical skill videos. It targets rubric-grounded process-level judgment of whether interaction-driven state updates preserve procedural correctness across an entire workflow. SiMing-Bench is instantiated with SiMing-Score, a physician-annotated dataset of real clinical skill examination videos spanning cardiopulmonary resuscitation, automated external defibrillator operation, and bag-mask ventilation, each paired with a standardized step-wise rubric and dual-expert labels. Across diverse open- and closed-source MLLMs, we observe consistently weak agreement with physician judgments. Moreover, weak performance on rubric-defined intermediate steps persists even when overall procedure-level correlation appears acceptable, suggesting that coarse global assessment substantially overestimates current models' procedural judgment ability. Additional analyses with binary step judgment and step-aligned clips indicate that the bottleneck is not merely fine-grained scoring or temporal localization, but modeling how continuous interactions update procedural state over time.
△ Less
Submitted 10 April, 2026;
originally announced April 2026.
-
A Generative Foundation Model for Multimodal Histopathology
Authors:
Jinxi Xiang,
Mingjie Li,
Siyu Hou,
Yijiang Chen,
Xiangde Luo,
Yuanfeng Ji,
Xiang Zhou,
Ehsan Adeli,
Akshay Chaudhari,
Curtis P. Langlotz,
Kilian M. Pohl,
Ruijiang Li
Abstract:
Accurate diagnosis and treatment of complex diseases require integrating histological, molecular, and clinical data, yet in practice these modalities are often incomplete owing to tissue scarcity, assay cost, and workflow constraints. Existing computational approaches attempt to impute missing modalities from available data but rely on task-specific models trained on narrow, single source-target p…
▽ More
Accurate diagnosis and treatment of complex diseases require integrating histological, molecular, and clinical data, yet in practice these modalities are often incomplete owing to tissue scarcity, assay cost, and workflow constraints. Existing computational approaches attempt to impute missing modalities from available data but rely on task-specific models trained on narrow, single source-target pairs, limiting their generalizability. Here we introduce MuPD (Multimodal Pathology Diffusion), a generative foundation model that embeds hematoxylin and eosin (H&E)-stained histology, molecular RNA profiles, and clinical text into a shared latent space through a diffusion transformer with decoupled cross-modal attention. Pretrained on 100 million histology image patches, 1.6 million text-histology pairs, and 10.8 million RNA-histology pairs spanning 34 human organs, MuPD supports diverse cross-modal synthesis tasks with minimal or no task-specific fine-tuning. For text-conditioned and image-to-image generation, MuPD synthesizes histologically faithful tissue architectures, reducing Fréchet inception distance (FID) scores by 50% relative to domain-specific models and improving few-shot classification accuracy by up to 47% through synthetic data augmentation. For RNA-conditioned histology generation, MuPD reduces FID by 23% compared with the next-best method while preserving cell-type distributions across five cancer types. As a virtual stainer, MuPD translates H&E images to immunohistochemistry and multiplex immunofluorescence, improving average marker correlation by 37% over existing approaches. These results demonstrate that a single, unified generative model pretrained across heterogeneous pathology modalities can substantially outperform specialized alternatives, providing a scalable computational framework for multimodal histopathology.
△ Less
Submitted 4 April, 2026;
originally announced April 2026.
-
A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction
Authors:
Jinxi Xiang,
Siyu Hou,
Yuchen Li,
Ryan Quinton,
Xiaoming Zhang,
Feyisope Eweje,
Xiangde Luo,
Yijiang Chen,
Zhe Li,
Colin Bergstrom,
Ted Kim,
Sierra Willens,
Francesca Maria Olguin,
Matthew Abikenari,
Andrew Heider,
Sanjeeth Rajaram,
Joel Neal,
Maximilian Diehn,
Xiang Zhou,
Ruijiang Li
Abstract:
Spatial transcriptomics (ST) enables gene expression mapping within anatomical context but remains costly and low-throughput. Hematoxylin and eosin (H\&E) staining offers rich morphology yet lacks molecular resolution. We present \textbf{\ours} (\textbf{S}patial \textbf{T}ranscriptomics and hist\textbf{O}logy \textbf{R}epresentation \textbf{M}odel), a foundation model trained on 1.2 million spatia…
▽ More
Spatial transcriptomics (ST) enables gene expression mapping within anatomical context but remains costly and low-throughput. Hematoxylin and eosin (H\&E) staining offers rich morphology yet lacks molecular resolution. We present \textbf{\ours} (\textbf{S}patial \textbf{T}ranscriptomics and hist\textbf{O}logy \textbf{R}epresentation \textbf{M}odel), a foundation model trained on 1.2 million spatially resolved transcriptomic profiles with matched histology across 18 organs. Using a hierarchical architecture integrating morphological features, gene expression, and spatial context, STORM bridges imaging and omics through robust molecular--morphological representations. STORM enhances spatial domain discovery, producing biologically coherent tissue maps, and outperforms existing methods in predicting spatial gene expression from H\&E images across 11 tumor types. The model is platform-agnostic, performing consistently across Visium, Xenium, Visium HD, and CosMx. Applied to 23 independent cohorts comprising 7,245 patients, STORM significantly improves immunotherapy response prediction and prognostication over established biomarkers, providing a scalable framework for spatially informed discovery and clinical precision medicine.
△ Less
Submitted 4 April, 2026;
originally announced April 2026.
-
World Reasoning Arena
Authors:
PAN Team,
Qiyue Gao,
Kun Zhou,
Jiannan Xiang,
Zihan Liu,
Dequan Yang,
Junrong Chen,
Arif Ahmad,
Cong Zeng,
Ganesh Bannur,
Xinqi Huang,
Zheqi Liu,
Yi Gu,
Yichi Yang,
Guangyi Liu,
Zhiting Hu,
Zhengzhong Liu,
Eric Xing
Abstract:
World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM benchmarks remain narrowly focused on next-state prediction and visual fidelity, overlooking the richer simulation capabilities required for intelligent behavior. To address this gap, we introduce WR-Arena, a comprehensive be…
▽ More
World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM benchmarks remain narrowly focused on next-state prediction and visual fidelity, overlooking the richer simulation capabilities required for intelligent behavior. To address this gap, we introduce WR-Arena, a comprehensive benchmark for evaluating WMs along three fundamental dimensions of next world simulation: (i) Action Simulation Fidelity, the ability to interpret and follow semantically meaningful, multi-step instructions and generate diverse counterfactual rollouts; (ii) Long-horizon Forecast, the ability to sustain accurate, coherent, and physically plausible simulations across extended interactions; and (iii) Simulative Reasoning and Planning, the ability to support goal-directed reasoning by simulating, comparing, and selecting among alternative futures in both structured and open-ended environments. We build a task taxonomy and curate diverse datasets designed to probe these capabilities, moving beyond single-turn and perceptual evaluations. Through extensive experiments with state-of-the-art WMs, our results expose a substantial gap between current models and human-level hypothetical reasoning, and establish WR-Arena as both a diagnostic tool and a guideline for advancing next-generation world models capable of robust understanding, forecasting, and purposeful action. The code is available at https://github.com/MBZUAI-IFM/WR-Arena.
△ Less
Submitted 26 March, 2026;
originally announced March 2026.
-
Quantum Neural Physics: Solving Partial Differential Equations on Quantum Simulators using Quantum Convolutional Neural Networks
Authors:
Jucai Zhai,
Muhammad Abdullah,
Boyang Chen,
Fazal Chaudry,
Paul N. Smith,
Claire E. Heaney,
Yanghua Wang,
Jiansheng Xiang,
Christopher C. Pain
Abstract:
Neural Physics recasts local discretisations of partial differential equations (PDEs) as fixed convolutional operators, providing a physics-preserving alternative to data-driven surrogate modelling in scientific machine learning. However, existing realizations remain largely confined to classical AI hardware and do not directly connect to quantum structured operator design. To bridge this gap, we…
▽ More
Neural Physics recasts local discretisations of partial differential equations (PDEs) as fixed convolutional operators, providing a physics-preserving alternative to data-driven surrogate modelling in scientific machine learning. However, existing realizations remain largely confined to classical AI hardware and do not directly connect to quantum structured operator design. To bridge this gap, we introduce a \emph{Quantum Neural Physics} framework and develop a Hybrid Quantum-Classical CNN Multigrid Solver (HQC-CNNMG). The proposed method maps analytically prescribed stencil operators to local quantum convolutional primitives and embeds them within a classical multilevel W-cycle architecture, combining the operator-centric view of scientific ML with the numerical rigor of multigrid solvers. Using amplitude encoding together with the Linear Combination of Unitaries (LCU) and the Quantum Fourier Transform (QFT), the resulting local quantum operators admit logarithmic-depth implementation, with circuit depth scaling as $\mathcal{O}(\log K)$ for an encoded block of size $K$ under the idealized parallel circuit model considered here. Numerical experiments on Poisson, transient diffusion, convection--diffusion, and incompressible Navier--Stokes problems demonstrate numerical consistency, stable multilevel behaviour, and workflow-level feasibility on noiseless simulators. Comparisons with representative quantum linear solver paradigms further show that the main strength of HQC-CNNMG lies in its balanced trade-off among local circuit depth, numerical robustness, and compatibility with PDE structure, rather than in fully quantum global inversion.
△ Less
Submitted 20 June, 2026; v1 submitted 25 March, 2026;
originally announced March 2026.
-
AirSimAG: A High-Fidelity Simulation Platform for Air-Ground Collaborative Robotics
Authors:
Yangjie Cui,
Xin Dong,
Boyang Gao,
Jinwu Xiang,
Daochun Li,
Zhan Tu
Abstract:
As spatial intelligence continues to evolve, heterogeneous multi-agent systems-particularly the collaboration between Unmanned Aerial Vehicles (UAVs) and Unmanned Ground Vehicles (UGVs), have demonstrated strong potential in complex applications such as search and rescue, urban surveillance, and environmental monitoring. However, existing simulation platforms are primarily designed for single-agen…
▽ More
As spatial intelligence continues to evolve, heterogeneous multi-agent systems-particularly the collaboration between Unmanned Aerial Vehicles (UAVs) and Unmanned Ground Vehicles (UGVs), have demonstrated strong potential in complex applications such as search and rescue, urban surveillance, and environmental monitoring. However, existing simulation platforms are primarily designed for single-agent dynamics and lack dedicated frameworks for interactive air-ground collaborative simulation. In this paper, we present AirsimAG, a high-fidelity air-ground collaborative simulation platform built upon an extensively customized AirSim framework. The platform enables synchronized multi-agent simulation and supports heterogeneous sensing and control interfaces for UAV-UGV systems. To demonstrate its capabilities, we design a set of representative air-ground collaborative tasks, including mapping, planning, tracking, formation, and exploration. We further provide quantitative analyses based on these tasks to illustrate the platform effectiveness in supporting multi-agent coordination and cross-modal data consistency. The AirsimAG simulation platform is publicly available at https://github.com/BIULab-BUAA/AirSimAG.
△ Less
Submitted 24 March, 2026;
originally announced March 2026.
-
LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic Inference
Authors:
Junkun Jiang,
Ho Yin Au,
Jingyu Xiang,
Jie Chen
Abstract:
Human motion is highly expressive and naturally aligned with language, yet prevailing methods relying heavily on joint text-motion embeddings struggle to synthesize temporally accurate, detailed motions and often lack explainability. To address these limitations, we introduce LabanLite, a motion representation developed by adapting and extending the Labanotation system. Unlike black-box text-motio…
▽ More
Human motion is highly expressive and naturally aligned with language, yet prevailing methods relying heavily on joint text-motion embeddings struggle to synthesize temporally accurate, detailed motions and often lack explainability. To address these limitations, we introduce LabanLite, a motion representation developed by adapting and extending the Labanotation system. Unlike black-box text-motion embeddings, LabanLite encodes each atomic body-part action (e.g., a single left-foot step) as a discrete Laban symbol paired with a textual template. This abstraction decomposes complex motions into interpretable symbol sequences and body-part instructions, establishing a symbolic link between high-level language and low-level motion trajectories. Building on LabanLite, we present LaMoGen, a Text-to-LabanLite-to-Motion Generation framework that enables large language models (LLMs) to compose motion sequences through symbolic reasoning. The LLM interprets motion patterns, relates them to textual descriptions, and recombines symbols into executable plans, producing motions that are both interpretable and linguistically grounded. To support rigorous evaluation, we introduce a Labanotation-based benchmark with structured description-motion pairs and three metrics that jointly measure text-motion alignment across symbolic, temporal, and harmony dimensions. Experiments demonstrate that LaMoGen establishes a new baseline for both interpretability and controllability, outperforming prior methods on our benchmark and two public datasets. These results highlight the advantages of symbolic reasoning and agent-based design for language-driven motion synthesis.
△ Less
Submitted 12 March, 2026;
originally announced March 2026.
-
Systematic study of superheavy nuclei within a microscopic collective Hamiltonian: Impact of quantum shape fluctuations
Authors:
X. Q. Yang,
R. Y. Hu,
R. N. Mao,
J. Xiang,
Z. P. Li
Abstract:
The even-even superheavy nuclei with $104 \leqslant Z \leqslant 126$ and $N\leqslant 258$ have been investigated using a microscopic five-dimensional collective Hamiltonian (5DCH) based on constrained triaxial relativistic Hartree-Bogoliubov calculations with the PC-PK1 density functional. The 5DCH approach effectively captures the characteristic of isospin dependence of nuclear binding energies,…
▽ More
The even-even superheavy nuclei with $104 \leqslant Z \leqslant 126$ and $N\leqslant 258$ have been investigated using a microscopic five-dimensional collective Hamiltonian (5DCH) based on constrained triaxial relativistic Hartree-Bogoliubov calculations with the PC-PK1 density functional. The 5DCH approach effectively captures the characteristic of isospin dependence of nuclear binding energies, two-nucleon separation energies, and $α$-decay energies across isotopic chains and demonstrates consistent accuracy as $Z$ increases, underscoring the model's predictive power. The collective potentials, average quadrupole deformations, and characteristic collective observables: $E(2^+_1)$, $R_{42}$, and $B(E2; 2^+_1\to 0^+_1)$ reveal a shape transition from well-prolate deformation around $N=150$ and $N=210$ to medium-deformed $γ$-soft shape around $N=176$ and $N=246$, and finally to a spherical shape near $N=184$ and $N=258$ for the isotopic chains with $104\leqslant Z\leqslant 118$. Oblate deformations are favored for $Z\geqslant 120$ isotopes around $N=178$. Remarkably, for a substantial range of transitional superheavy nuclei with $N\gtrsim184$ and $N\gtrsim240$, no $0^+$ states bounded by the fission saddles are predicted within their very shallow potential wells due to quantum shape fluctuations (QSFs). Additionally, sharp variations predicted for two-neutron separation energies $S_{2n}$ and $α$-decay energies $Q_α$ at $N=184$ and $258$ in mean-field calculations are significantly reduced and shifted to $N=182$ and $256$ in the 5DCH calculations, which is caused by the rapid evolution of the dynamical correlation energies related to QSFs around the nuclear spherical shells.
△ Less
Submitted 10 March, 2026;
originally announced March 2026.
-
Learning Context-Adaptive Motion Priors for Masked Motion Diffusion Models with Efficient Kinematic Attention Aggregation
Authors:
Junkun Jiang,
Jie Chen,
Ho Yin Au,
Jingyu Xiang
Abstract:
Vision-based motion capture solutions often struggle with occlusions, which result in the loss of critical joint information and hinder accurate 3D motion reconstruction. Other wearable alternatives also suffer from noisy or unstable data, often requiring extensive manual cleaning and correction to achieve reliable results. To address these challenges, we introduce the Masked Motion Diffusion Mode…
▽ More
Vision-based motion capture solutions often struggle with occlusions, which result in the loss of critical joint information and hinder accurate 3D motion reconstruction. Other wearable alternatives also suffer from noisy or unstable data, often requiring extensive manual cleaning and correction to achieve reliable results. To address these challenges, we introduce the Masked Motion Diffusion Model (MMDM), a diffusion-based generative reconstruction framework that enhances incomplete or low-confidence motion data using partially available high-quality reconstructions within a Masked Autoencoder architecture. Central to our design is the Kinematic Attention Aggregation (KAA) mechanism, which enables efficient, deep, and iterative encoding of both joint-level and pose-level features, capturing structural and temporal motion patterns essential for task-specific reconstruction. We focus on learning context-adaptive motion priors, specialized structural and temporal features extracted by the same reusable architecture, where each learned prior emphasizes different aspects of motion dynamics and is specifically efficient for its corresponding task. This enables the architecture to adaptively specialize without altering its structure. Such versatility allows MMDM to efficiently learn motion priors tailored to scenarios such as motion refinement, completion, and in-betweening. Extensive evaluations on public benchmarks demonstrate that MMDM achieves strong performance across diverse masking strategies and task settings. The source code is available at https://github.com/jjkislele/MMDM.
△ Less
Submitted 8 March, 2026;
originally announced March 2026.
-
Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents
Authors:
Jiahong Xiang,
Wenxiao He,
Xihua Wang,
Hongliang Tian,
Yuqun Zhang
Abstract:
The Rust programming language presents a steep learning curve and significant coding challenges, making the automation of issue resolution essential for its broader adoption. Recently, LLM-powered code agents have shown remarkable success in resolving complex software engineering tasks, yet their application to Rust has been limited by the absence of a large-scale, repository-level benchmark. To b…
▽ More
The Rust programming language presents a steep learning curve and significant coding challenges, making the automation of issue resolution essential for its broader adoption. Recently, LLM-powered code agents have shown remarkable success in resolving complex software engineering tasks, yet their application to Rust has been limited by the absence of a large-scale, repository-level benchmark. To bridge this gap, we introduce Rust-SWE-bench, a benchmark comprising 500 real-world, repository-level software engineering tasks from 34 diverse and popular Rust repositories. We then perform a comprehensive study on Rust-SWE-bench with four representative agents and four state-of-the-art LLMs to establish a foundational understanding of their capabilities and limitations in the Rust ecosystem. Our extensive study reveals that while ReAct-style agents are promising, i.e., resolving up to 21.2% of issues, they are limited by two primary challenges: comprehending repository-wide code structure and complying with Rust's strict type and trait semantics. We also find that issue reproduction is rather critical for task resolution. Inspired by these findings, we propose RUSTFORGER, a novel agentic approach that integrates an automated test environment setup with a Rust metaprogramming-driven dynamic tracing strategy to facilitate reliable issue reproduction and dynamic analysis. The evaluation shows that RUSTFORGER using Claude-Sonnet-3.7 significantly outperforms all baselines, resolving 28.6% of tasks on Rust-SWE-bench, i.e., a 34.9% improvement over the strongest baseline, and, in aggregate, uniquely solves 46 tasks that no other agent could solve across all adopted advanced LLMs.
△ Less
Submitted 26 February, 2026;
originally announced February 2026.
-
RA-Nav: A Risk-Aware Navigation System Based on Semantic Segmentation for Aerial Robots in Unpredictable Environments
Authors:
Ziyi Zong,
Xin Dong,
Jinwu Xiang,
Daochun Li,
Zhan Tu
Abstract:
Existing aerial robot navigation systems typically plan paths around static and dynamic obstacles, but fail to adapt when a static obstacle suddenly moves. Integrating environmental semantic awareness enables estimation of potential risks posed by suddenly moving obstacles. In this paper, we propose RA- Nav, a risk-aware navigation framework based on semantic segmentation. A lightweight multi-scal…
▽ More
Existing aerial robot navigation systems typically plan paths around static and dynamic obstacles, but fail to adapt when a static obstacle suddenly moves. Integrating environmental semantic awareness enables estimation of potential risks posed by suddenly moving obstacles. In this paper, we propose RA- Nav, a risk-aware navigation framework based on semantic segmentation. A lightweight multi-scale semantic segmentation network identifies obstacle categories in real time. These obstacles are further classified into three types: stationary, temporarily static, and dynamic. For each type, corresponding risk estimation functions are designed to enable real-time risk prediction, based on which a complete local risk map is constructed. Based on this map, the risk-informed path search algorithm is designed to guarantee planning that balances path efficiency and safety. Trajectory optimization is then applied to generate trajectories that are safe, smooth, and dynamically feasible. Comparative simulations demonstrate that RA-Nav achieves higher success rates than baselines in sudden obstacle state transition scenarios. Its effectiveness is further validated in simulations using real- world data.
△ Less
Submitted 19 February, 2026;
originally announced February 2026.
-
World Action Models are Zero-shot Policies
Authors:
Seonghyeon Ye,
Yunhao Ge,
Kaiyuan Zheng,
Shenyuan Gao,
Sihyun Yu,
George Kurian,
Suneel Indupuru,
You Liang Tan,
Chuning Zhu,
Jiannan Xiang,
Ayaan Malik,
Kyungmin Lee,
William Liang,
Nadun Ranawaka,
Jiasheng Gu,
Yinzhen Xu,
Guanzhi Wang,
Fengyuan Hu,
Avnish Narayan,
Johan Bjorck,
Jing Wang,
Gwanghyun Kim,
Dantong Niu,
Ruijie Zheng,
Yuqi Xie
, et al. (11 additional authors not shown)
Abstract:
State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce DreamZero, a World Action Model (WAM) built upon a pretrained video diffusion backbone. Unlike VLAs, WAMs learn physical dynamics by predicting future world states and actions, using video as a dense representation of how th…
▽ More
State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce DreamZero, a World Action Model (WAM) built upon a pretrained video diffusion backbone. Unlike VLAs, WAMs learn physical dynamics by predicting future world states and actions, using video as a dense representation of how the world evolves. By jointly modeling video and action, DreamZero learns diverse skills effectively from heterogeneous robot data without relying on repetitive demonstrations. This results in over 2x improvement in generalization to new tasks and environments compared to state-of-the-art VLAs in real robot experiments. Crucially, through model and system optimizations, we enable a 14B autoregressive video diffusion model to perform real-time closed-loop control at 7Hz. Finally, we demonstrate two forms of cross-embodiment transfer: video-only demonstrations from other robots or humans yield a relative improvement of over 42% on unseen task performance with just 10-20 minutes of data. More surprisingly, DreamZero enables few-shot embodiment adaptation, transferring to a new embodiment with only 30 minutes of play data while retaining zero-shot generalization.
△ Less
Submitted 17 February, 2026;
originally announced February 2026.